Hiro development journal

Requiring attributable regression tests before candidate evaluation

Early baseline-contrast gate implemented, tested, and loaded into the live queue Machine-readable JSON

Executive summary

Hiro's candidate builder previously required a generated targeted test to pass on the candidate but deferred the untouched-baseline counterfactual until Stage 4. This allowed expensive evaluations to begin before discovering that a test already passed on baseline or could not execute there.

Stage 3 now copies the exact candidate-authored tests into a detached checkout of the candidate's immutable base commit. A candidate can be frozen as ready only when those tests collect, execute, and produce a JUnit-recorded assertion failure on baseline while passing on the candidate.

Import errors, collection errors, setup or runtime errors, absent reports, zero collected tests, and skip-only runs are explicitly inconclusive rather than evidence of improvement. Tests that pass on baseline are classified separately as failing to demonstrate a change.

The bounded repair loop now receives the contrast classification, outcome counts, and compact collection and execution output, together with a repair focus tailored to the failure. Hiro was restarted on the new commit, its implementation-time dirty-repository breaker was audit-reset, and the scheduler selected a live candidate under the new code.

After a live packet showed that Qwen twice imported newly invented candidate-only router functions, the builder was strengthened again: it now supplies an exact baseline module-to-public-symbol allowlist and explicitly requires every production import in a targeted test to come from that list.

Work completed

Early counterfactual gate

Completed
  • Extended CandidateBuilder validation after candidate syntax, collection, and targeted tests pass.
  • Creates a uniquely named detached Git worktree at the exact base commit under the external candidate-worktree root.
  • Copies only the exact targeted tests from the candidate workspace into that untouched baseline and validates collection before execution.
  • Runs pytest with a bounded JUnit artifact and removes the detached baseline worktree in a finally block.
  • Candidate readiness now requires both local candidate success and an improvement_demonstrated baseline contrast.

Assertion-aware evidence classification

Completed
  • Added a shared targeted-contrast module used by both Stage 3 construction and Stage 4 evaluation.
  • Improvement is demonstrated only when at least one non-skipped case executed, JUnit recorded one or more failures, JUnit recorded zero errors, and pytest returned the ordinary test-failure code.
  • A clean baseline pass is classified as already_passes_baseline rather than inconclusive.
  • Missing or malformed JUnit evidence, no executable cases, all-skipped cases, test errors, and inconsistent return-code/report combinations are inconclusive.
  • Stage 4 now repeats this stronger JUnit-aware check as defense in depth for frozen packets.

Bounded test repair

Completed
  • Preserved the existing maximum of two repairs after the initial complete candidate attempt.
  • Each failed attempt still resets to the clean base commit; repairs must return a complete revised patch and tests rather than editing a contaminated failed attempt.
  • Repair context now contains the baseline status, reason, outcome counts, and compact stdout and stderr from collection and execution.
  • The local candidate agent receives a specific focus: create an assertion contrast when baseline already passes, repair imports to use baseline-existing entrypoints after collection failure, or remove errors and skips for other inconclusive executions.
  • The prompt now explicitly forbids treating candidate-only imports, setup failures, collection failures, skipped tests, or empty collections as improvement evidence.
  • The full code-owned task contract is now carried from the incident spec into CandidateBuildRequest, the Qwen request, and the frozen packet as the authoritative behavioral oracle.
  • The request also contains baseline_test_imports, a bounded AST-derived map of public functions and classes that exist in the untouched baseline production files; private helpers and test-local symbols are excluded.

Live service transition

Completed
  • Committed the implementation before service activation so new candidates pin a stable base revision.
  • The old scheduler opened its infrastructure breaker after three blocked_dirty_repository events during the edit window, correctly refusing to work against changing source.
  • Stopped only the existing Hiro API process tree, preserved the loaded Qwen model server, and restarted Hiro through the checked-in path-normalizing detached launcher.
  • Reset the resolved dirty-repository breaker through ContinuousQueue.clear_infrastructure_failures, producing an append-only circuit_breaker_reset event.
  • The restarted scheduler selected a preference-comparison incident and began candidate construction on the new code with Qwen connected.
  • That live frozen packet was pinned to 3a86749, preserved the full preference-comparison task contract, and classified attempts 0 and 2 as baseline collection/import failures because their tests imported candidate-only router functions. No invalid test was credited.
  • The live evidence motivated the final 0b2a784 baseline_test_imports constraint. Hiro was restarted again, the interrupted lease was released only after its exact worker process had been stopped, and stale candidate state is rebuilt on the active revision by the queue's existing upgrade recovery.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Candidate builder and evaluator tests passed The final focused suite contains 60 passing builder, evaluator, sandbox, continuous-queue, and agenda tests. Coverage includes real detached baseline assertion contrast, tests that already pass on baseline, repair-context diagnostics, strict return-code/JUnit classification, errors, skipped tests, absent reports, task-contract propagation, and exclusion of private or test-local symbols from the baseline import allowlist.
Sandbox, continuous queue, and agenda integration tests passed The integration-focused subset is included in the 60 focused passes and confirms that stronger candidate packet requirements do not break queue revision handling, sandbox orchestration, or agenda execution.
Full Hiro regression suite passed 638 tests passed in 183.28 seconds after the final baseline-symbol constraint.
Live service restart and scheduler passed Hiro restarted through scripts/start_hiro.py, reported healthy with qwen/qwen3.8-27b connected, reset the resolved breaker to zero failures, and selected a scheduler-owned candidate on the new commit.
Live candidate contrast packet passed The scheduler produced sandbox-20260817021750-570da205-01 on base 3a86749 with the complete task contract frozen. Attempts 0 and 2 were explicitly inconclusive because baseline collection returned 2 for candidate-only router imports; attempt 1 failed local validation. After adding the exact-import constraint, sandbox-20260817022645-55020c0e-01 was pinned to final base 0b2a784 with its writing-assistance contract; it failed earlier at empty-plan, ambiguous-edit, and candidate-local validation gates and therefore never reached baseline contrast. Neither packet advanced.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 140 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps