Executive summary
Hiro no longer treats every failed candidate transaction as evidence that an idea was bad. The queue now distinguishes construction failure, evaluation rejection, safety rejection, infrastructure blockage, disproven hypotheses, legacy outcomes, supersession, and implementation.
Candidate repair attempts now restart from a clean isolated baseline, receive the prior failure evidence, and must produce a complete alternative candidate. Python candidates can use AST-bounded top-level symbol replacement instead of fragile text matching.
Continuous candidates now prove targeted improvement by running the candidate-authored test against an untouched baseline: it must execute and fail assertions there, then pass on the candidate. Public and held-out suites enforce global non-regression, invariants, categories, and repeated cross-suite latency evidence.
A shadow replay found eleven retained candidates that satisfy the corrected targeted and global evidence contract. Four were rejected because baseline already passed, and twenty-four remained inconclusive rather than receiving false credit.
A real Qwen3.8 retry of a previously failed memory candidate produced a candidate-ready packet after one clean repair. All 597 repository tests passed and the live circuit breaker remains closed.