Executive summary
Hiro's candidate builder previously required a generated targeted test to pass on the candidate but deferred the untouched-baseline counterfactual until Stage 4. This allowed expensive evaluations to begin before discovering that a test already passed on baseline or could not execute there.
Stage 3 now copies the exact candidate-authored tests into a detached checkout of the candidate's immutable base commit. A candidate can be frozen as ready only when those tests collect, execute, and produce a JUnit-recorded assertion failure on baseline while passing on the candidate.
Import errors, collection errors, setup or runtime errors, absent reports, zero collected tests, and skip-only runs are explicitly inconclusive rather than evidence of improvement. Tests that pass on baseline are classified separately as failing to demonstrate a change.
The bounded repair loop now receives the contrast classification, outcome counts, and compact collection and execution output, together with a repair focus tailored to the failure. Hiro was restarted on the new commit, its implementation-time dirty-repository breaker was audit-reset, and the scheduler selected a live candidate under the new code.
After a live packet showed that Qwen twice imported newly invented candidate-only router functions, the builder was strengthened again: it now supplies an exact baseline module-to-public-symbol allowlist and explicitly requires every production import in a targeted test to come from that list.