Executive summary
The non-promotion interval was caused by a structural queue defect, not by Qwen model selection or a lack of improvement ideas. Historical interaction-audit records lacked the full synthetic prompt needed to reproduce their failures, yet remained eligible for candidate construction and were repeatedly revived after unrelated repository revisions.
The repair adds an explicit replay-completeness invariant and an append-only migration. It superseded 239 prompt-incomplete historical audit records and left zero such records actionable. Artifact-blocked work is now reconsidered only after a candidate-builder compatibility change, rather than after every Git revision.
Twelve research signals from the user's scheduled ChatGPT research conversation were reduced into six locally defined, testable mechanisms and admitted to the ranked queue: evaluator blind spots, measured-null calibration, strategy reconsideration, recursive iteration, security drift, and algorithmic invention.
The implementation passed 77 focused tests, all 718 repository tests, and five exact-revision pipeline qualification cycles. Hiro restarted on the qualified revision with Qwen 3.8 27B at a 16,384-token context window. Its first post-repair audit candidate used the complete synthetic prompt, produced a real patch, ran the isolated replay, and received a bounded retry for one concrete failed assertion rather than failing from absent evidence.