Executive summary
Completed and activated a simplified continuous-improvement process centered on reproducible failures, harness-owned tests, isolated candidates, paired evaluation, timed canary, independent promotion, and automatic rollback.
Candidate construction failures no longer become negative verdicts about the underlying idea. Exhausted build attempts move to artifact_blocked with idea_merit_evaluated false and become eligible again only after a relevant builder change.
Valid comparative evaluation failures remain rejected evidence. Security regressions remain terminal, infrastructure failures receive cooldown retries, and canary or governor failures retain their own evidence and revision paths.
The candidate builder now mechanically normalizes harmless generated whitespace and common edit-shape mistakes, uses the exact harness-owned production replay instead of allowing a candidate to author its own approval test, and rejects repeated failed patch semantics.
A durable supervisor watches the improvement factory itself. Recursive model-authored builder-repair tasks were retired; a bounded run of construction failures now stops queue manufacture and exposes a diagnostic incident without consuming idea merit.
Corrected a watchdog defect discovered during implementation: legitimate evaluation rejections no longer count as candidate-builder failures or move the pipeline toward a false global shutdown.
Corrected a Windows linked-worktree defect discovered by release qualification. Qualification pytest roots now anchor at the short common Git checkout instead of inheriting an arbitrarily long release-worktree path.
Preserved all eight previously promoted commits on the live branch. The final release is a fast-forward from the live revision rather than a replacement or history rewrite.
The exact activated revision a11954ef646e7f43404af7909678073af701edae passed 750 repository tests and five consecutive ten-test end-to-end qualification cycles. Hiro restarted healthy with Qwen 3.8 connected and the existing durable queue intact.
A separate read-only whole-code inspection then found that the system is not yet a complete autonomous production loop: candidate execution needs stronger operating-system isolation, the current timed checks are synthetic rather than product-traffic probation, promotion is not transactionally joined to runtime restart verification, and post-promotion rollback monitoring is not wired.