Hiro development journal

Diagnosing overnight promotion starvation and bounding builder repair

Repairs integrated, qualified, restarted, and exercised; live retries now carry the exact failed patch Machine-readable JSON

Executive summary

The overnight autonomous run produced no promotion. The scheduler remained active, but the supervisor recorded 59 failed construction attempts and generated 11 consecutive versions of the same builder-repair task, allowing one failure lineage to consume the entire promotion window.

The first positive construction result arrived at 06:11 Pacific, after the overnight expectation had already failed. That builder-repair candidate independently entered canary and passed its 0-, 5-, and 15-minute checkpoints, but a direct causal trial proved its patch did not improve construction: the originating weather candidate failed with the exact same harness-owned replay error.

Implemented a bounded repair-lineage policy in an isolated worktree. The same builder failure may receive at most three repair generations; prior terminal evidence is carried into later generations, and exhaustion parks the originating artifact without evaluating its idea merit so the ranked queue can advance.

Replaced the hard-coded candidate-builder compatibility label with a fingerprint of the actual construction code. A real builder change now reactivates eligible blocked artifacts, while unrelated repository commits do not pretend the builder changed.

Moved the continuous-improvement turn onto a worker thread. Synchronous pytest canaries and governor subprocesses can no longer freeze Hiro's FastAPI event loop and benchmark API while they run.

Added a causal builder-repair canary. A repair must now use its frozen candidate revision to reconstruct the originating failed improvement and clear that exact construction boundary before it can finish timed canary or reach the governor.

Prevented builder-version changes from reactivating historical repair artifacts. The first repaired scheduler turn automatically superseded all 11 obsolete repair generations while retaining legitimate originating improvements.

The final focused lifecycle suite passed 70 tests and the full pinned-runtime repository suite passed 731 tests with two pre-existing unknown-marker warnings. Frozen qualification passed five consecutive eight-test cycles on both intermediate and final integrated revisions. No evaluator threshold, timed-canary requirement, full-suite gate, or governor authority was weakened.

Extended unsupervised observation exposed one more upstream defect: supervisor reflection packets reported changed_files as empty and candidate_output as empty even when a failed candidate had made a concrete code edit. Qwen therefore received pytest text but not the code that caused it, and independently recreated the same defective regex.

Revision 5065952 now recovers the immutable failed edit from the frozen candidate packet and carries a bounded patch excerpt plus its verified changed-file manifest into the next autonomous attempt. The complete repository suite passed 734 tests, and a new five-cycle qualification passed all 40 end-to-end checks.

After restart, a real transit-directions candidate failed construction and its independently scheduled retry contained two changed files and 2,870 characters of the actual failed patch. This live receipt proves the new process no longer depends on a human observer to explain what the preceding attempt changed.

The first patch-carrying retry still copied the production edit byte-for-byte. A new builder invariant now rejects an identical prior production patch inside CandidateBuilder and uses the local repair loop instead of consuming another supervisor attempt.

Live exercise then showed Qwen changing only a comment label while preserving identical behavior. The invariant was strengthened again to compare Python semantic tokens while ignoring comments and formatting. Final revision fe46222 passed 736 repository tests and 40/40 frozen qualification checks.

Work completed

Overnight funnel audit

Completed
  • Qwen 3.8 27B, Hiro, and their listener ports remained alive. This ruled out a simple model-process or service-process outage.
  • The durable ledger showed 59 construction failures, one candidate-ready result, 12 completed artifact-blocked episodes, and no candidate implementation during the overnight window.
  • Eleven supervisor_builder_repair_escalated events created repair generations for the same candidate-validation fingerprint. Each terminal repair was followed by a fresh episode that rediscovered the same global watchdog condition and created another repair generation.
  • The main API accepted connections but did not respond while the scheduler synchronously ran 92 canary tests. After that subprocess completed, the API responded normally and the checkpoint receipt was recorded.

Bounded repair lineage and queue fairness

Implemented and tested
  • Added a three-generation limit for one builder-version and failure-fingerprint pair.
  • Each later generation receives bounded outcome evidence from preceding terminal repair generations, including failure stage, type, reason, and disposition, so it is instructed not to repeat those edit shapes.
  • When the generation budget is exhausted, the root is moved to artifact_blocked with idea_merit_evaluated false and retriable_after_builder_change true. The active supervisor episode completes, releasing candidate capacity for the next ranked idea.
  • The queue records a dedicated supervisor_builder_repair_budget_exhausted event and the configured generation limit for dashboard and forensic use.

Real builder compatibility

Implemented and tested
  • The old compatibility value was a static date/version string. A promoted change to candidate_builder.py could therefore leave blocked artifacts labeled with the same version and never retried.
  • Compatibility is now a deterministic SHA-256-derived fingerprint over candidate_builder.py, autonomous_sandbox.py, continuous_engine.py, and supervisor.py.
  • Tests prove unrelated file changes leave the fingerprint stable and a builder-code change alters it. The queue's existing one-at-a-time artifact reactivation policy now receives this real fingerprint.

API responsiveness during improvement work

Implemented and tested
  • Candidate construction and canary/governor tests contain synchronous subprocess work even though the outer scheduler entrypoint is asynchronous.
  • The scheduler now runs an improvement turn on a dedicated worker thread with its own async event loop. The main FastAPI event loop remains available for chat, health, and benchmark requests.
  • A timing regression test injects blocking improvement work and proves the calling event loop remains responsive.

Causal validation for builder repairs

Implemented, tested, and exercised diagnostically
  • The candidate that reached canary added candidate_output metadata to local_candidate_agent, but CandidateBuilder consumes the plan's patches and ignores that field. Generic candidate-builder tests therefore passed without demonstrating improved construction.
  • A direct isolated trial loaded candidate revision 35402751447e44050597038f679368ba1171976e and retried the originating weather improvement. It returned the same candidate_validation_failure from the harness-owned captured boundary replay.
  • Builder-repair canary checkpoint zero now launches a candidate-worktree subprocess that retries the originating queue idea with the frozen builder candidate. The receipt is stored in canary metrics and the expensive causal trial runs only once.
  • A failed causal trial rejects the repair before promotion. Later timed checkpoints reuse the durable receipt rather than rerunning construction.

Obsolete repair cleanup

Implemented and exercised live
  • Artifact reactivation now skips records whose mechanism is builder_repair. Those records are prior implementation attempts, not independent improvement goals.
  • Queued builder-repair records without an owning active platform-repair episode are superseded atomically with a dedicated orphaned_builder_repair_superseded event.
  • After the final restart, the first scheduler turn superseded all eleven historical repair generations and left zero queued repair artifacts while continuing the originating bounded candidate episode.

Causal patch feedback between autonomous attempts

Implemented, fully tested, restarted, and exercised live
  • The extended run showed repeated schedule candidates using the same regex word-boundary shape. The durable reflection packet contained the assertion failure but declared no changed files and no candidate output, so the next clean-baseline attempt could not inspect the preceding edit.
  • Failed-candidate results now load the code-owned frozen candidate packet, verify its candidate identity, recover its changed-file manifest, and extract a bounded excerpt of the final attempted patch.
  • Supervisor reflection retains up to 4,000 characters of this candidate patch. The next Qwen request receives both the exact test failure and the exact implementation shape that produced it.
  • A live post-restart transit candidate generated a retry packet with core/response_envelope.py, its harness-owned autonomous test, and a 2,870-character patch excerpt. No manual queue-row edit or per-candidate prompt was used.

Mechanical rejection of repeated failed patches

Implemented, fully tested, qualified, and exercised live
  • A live retry reproduced its preceding production patch exactly even though the failed patch was present in the prompt. Prompt guidance alone was therefore insufficient to guarantee a new strategy.
  • The candidate request now carries a bounded prior_failed_candidate_output field. CandidateBuilder checks proposed production edits before application and rejects a repeated failed edit into its internal repair loop.
  • The first live version correctly rejected the byte-identical edit, but the next model response changed only a comment label. The final detector tokenizes Python snippets and ignores comments, whitespace, indentation tokens, and newlines while retaining behavioral code tokens.
  • The live ledger records the duplicate rejection in attempt three, followed by a semantically identical comment-only edit that exposed the need for token-level comparison. Final tests cover byte-identical, comment-only, and genuinely different passing implementations.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Overnight SQLite ledger audit failed objective, diagnosis complete The run recorded 59 construction failures and 11 repeated builder-repair escalations before the first candidate-ready result. No overnight candidate implementation was present.
Focused supervisor, engine, scheduler, and builder tests passed 108 focused tests passed in 27.84 seconds. Additional focused runs covered bounded generation exhaustion, supervisor release, compatibility fingerprinting, repair-history propagation, and nonblocking scheduling.
Full repository regression suite passed The final integrated repair passed 731 tests in 423.03 seconds under Hiro's pinned runtime. Two existing generated autonomous tests emitted unknown hiro_contract marker warnings; there were no failures.
Live autonomous canary progression and causal contrast generic checkpoints passed; causal validation failed The live builder-repair candidate passed independently scheduled checkpoints at 0, 5, and 15 minutes. A separate candidate-worktree trial then retried the originating improvement and reproduced the exact same harness-owned failure, proving the candidate was not a functional builder improvement. The source upgrade invalidated and requeued that obsolete canary before its 60-minute governor decision.
Frozen five-cycle qualification passed The final revision 76a2f8c162ba2196dc2393964f881eb2a898385d passed five consecutive cycles of eight end-to-end lifecycle tests. The frozen packet is qualification-20260823T142249Z.json with SHA-256 022faed4dc25084ba8975c0ebee0ee407e4f7c5c86bab2813ece87a80a9eb2f5.
Checked-launcher restart and responsiveness passed The checked start_hiro.cmd launcher restarted Hiro on the final revision in 3.406 seconds. Ports 8000, 8001, and 8765 share the new Hiro process; Qwen 3.8 remains on 8080. Benchmark API requests continued returning HTTP 200 during multi-minute Qwen construction.
Live obsolete-repair cleanup passed The first final-revision scheduler turn emitted eleven orphaned_builder_repair_superseded events and left zero queued builder-repair artifacts. The originating improvement resumed independently.
Failed-patch reflection focused and lifecycle suites passed The direct supervisor/continuous-engine suite passed 62 tests. The wider candidate-builder, continuous-engine, supervisor, active-loop, governor, and scheduler suite passed 120 tests.
Full repository regression after failed-patch reflection repair passed Revision 5065952da53350f2610f19671837151405c9bbca passed 734 tests in 409.54 seconds. The same two existing generated tests emitted unknown hiro_contract marker warnings; there were no failures.
Final frozen five-cycle qualification passed Revision 5065952da53350f2610f19671837151405c9bbca passed five consecutive cycles of eight end-to-end lifecycle checks, including controlled real governor promotion, restart idempotency, and rollback. Packet qualification-20260823T151821Z.json has SHA-256 7e404ace6460c20611e1ccb99079fc3eae04420e9a937b89a7c335604a06cae1.
Live autonomous failed-patch carry-forward passed After a checked-launcher restart, a real transit-directions candidate failed its harness-owned replay. Its queued attempt-two reflection recorded core/response_envelope.py and the autonomous test plus 2,870 characters of the exact attempted implementation. The retry was scheduled independently for 08:31:57 Pacific.
Live duplicate-patch rejection passed and exposed comment-only evasion The transit attempt-three frozen result contains the explicit error that its production patch exactly repeated the prior failed candidate. CandidateBuilder invoked local repair rather than accepting it. The subsequent model plan changed only the rule comment, which motivated semantic-token comparison.
Final semantic-duplicate regression suite passed Revision fe46222538f44f4686361bc4bdda224175202b57 passed 736 tests in 412.48 seconds with the same two existing unknown-marker warnings and no failures.
Final semantic-duplicate five-cycle qualification passed Revision fe46222538f44f4686361bc4bdda224175202b57 passed all 40 end-to-end checks. Packet qualification-20260823T161705Z.json has SHA-256 f0d0f6fd3d9a8f12d8858526dcf50d45ff0f1bc1720753883da4603d51a86ae5.

Current state

Next steps