Executive summary
Phase 3F-VS tested whether Hiro can establish that a numeric capability gap belongs to real production behavior and that a proposed candidate surface causally controls the measured behavior before reproduction is authorized.
The new binder uses exact repository paths and symbols, static call paths, code-owned runtime receipts, source hashes, and per-path causal classification. Claim IDs and source-paper wording do not select or certify bindings.
The two Phase 3F-VC false positives were confirmed as lab-only measurements of Daylab suite selection. Their numeric results remain valid, but they are now MEASURED_GAP_UNBOUND rather than production deficiencies.
Five other preserved probes use a verified adapter that reaches Hiro's real local-model router. Their proposed agent, memory, and metacognition surfaces were not on the measured execution path, and all five had already returned NO_CURRENT_GAP.
The remaining preserved hypothesis had no executable hidden-state interface and terminated SUBJECT_BINDING_UNCERTAIN with NO_TRANSFER_SURFACE.
Two immutable historical non-meta improvements passed positive-control binding at the exact modified symbols: reflection synthesis and memory-candidate normalization. Both were DIRECT_CONTROL, proving the binder does not merely reject every input.
No preserved fresh hypothesis remains VIABLE_FOR_REPRODUCTION. No discovery, reproduction, candidate construction, production write, promotion, or meta-improvement occurred.
The complete Hiro suite passed 883 tests with one expected skip. Hiro is healthy with the local model connected and its loaded and checkout revisions aligned at 28589bf3b63444e2a79d7706c4b71a8e9dc1ca48.