Executive summary
The overnight autonomous run produced no promotion. The scheduler remained active, but the supervisor recorded 59 failed construction attempts and generated 11 consecutive versions of the same builder-repair task, allowing one failure lineage to consume the entire promotion window.
The first positive construction result arrived at 06:11 Pacific, after the overnight expectation had already failed. That builder-repair candidate independently entered canary and passed its 0-, 5-, and 15-minute checkpoints, but a direct causal trial proved its patch did not improve construction: the originating weather candidate failed with the exact same harness-owned replay error.
Implemented a bounded repair-lineage policy in an isolated worktree. The same builder failure may receive at most three repair generations; prior terminal evidence is carried into later generations, and exhaustion parks the originating artifact without evaluating its idea merit so the ranked queue can advance.
Replaced the hard-coded candidate-builder compatibility label with a fingerprint of the actual construction code. A real builder change now reactivates eligible blocked artifacts, while unrelated repository commits do not pretend the builder changed.
Moved the continuous-improvement turn onto a worker thread. Synchronous pytest canaries and governor subprocesses can no longer freeze Hiro's FastAPI event loop and benchmark API while they run.
Added a causal builder-repair canary. A repair must now use its frozen candidate revision to reconstruct the originating failed improvement and clear that exact construction boundary before it can finish timed canary or reach the governor.
Prevented builder-version changes from reactivating historical repair artifacts. The first repaired scheduler turn automatically superseded all 11 obsolete repair generations while retaining legitimate originating improvements.
The final focused lifecycle suite passed 70 tests and the full pinned-runtime repository suite passed 731 tests with two pre-existing unknown-marker warnings. Frozen qualification passed five consecutive eight-test cycles on both intermediate and final integrated revisions. No evaluator threshold, timed-canary requirement, full-suite gate, or governor authority was weakened.
Extended unsupervised observation exposed one more upstream defect: supervisor reflection packets reported changed_files as empty and candidate_output as empty even when a failed candidate had made a concrete code edit. Qwen therefore received pytest text but not the code that caused it, and independently recreated the same defective regex.
Revision 5065952 now recovers the immutable failed edit from the frozen candidate packet and carries a bounded patch excerpt plus its verified changed-file manifest into the next autonomous attempt. The complete repository suite passed 734 tests, and a new five-cycle qualification passed all 40 end-to-end checks.
After restart, a real transit-directions candidate failed construction and its independently scheduled retry contained two changed files and 2,870 characters of the actual failed patch. This live receipt proves the new process no longer depends on a human observer to explain what the preceding attempt changed.
The first patch-carrying retry still copied the production edit byte-for-byte. A new builder invariant now rejects an identical prior production patch inside CandidateBuilder and uses the local repair loop instead of consuming another supervisor attempt.
Live exercise then showed Qwen changing only a comment label while preserving identical behavior. The invariant was strengthened again to compare Python semantic tokens while ignoring comments and formatting. Final revision fe46222 passed 736 repository tests and 40/40 frozen qualification checks.