Executive summary
A full-path audit repaired multiple real defects in candidate construction, evaluation reachability, runner ownership, causal comparison, and fresh-canary handling. The latest Hiro regression suite passed 672 tests in 187.60 seconds.
The target arithmetic proposal reached deterministic replay and corrected public and held-out evaluation, but an independent fresh canary returned a bare result that violated the response contract. The candidate was not promoted.
Subsequent model-generated revisions failed during isolated construction. Repeatedly adding arithmetic-specific instructions to rescue that one incident would overfit the harness and would not demonstrate a general self-improvement capability.
Hiro was therefore paused. The approval architecture is retained as useful infrastructure, while the candidate-construction strategy now requires a deliberate reset around generic primitives, explicit reachability evidence, portfolio proposals, and governed stop rules.