Executive summary
Phase 3F-VR added and qualified a generic baseline-only probe boundary for measuring Hiro's current behavior before research reproduction is considered.
The qualification reused the same eight preserved Phase 3F-V hypotheses. It performed no discovery, source treatment, reproduction, candidate construction, production mutation, promotion, or meta-improvement.
Seven of eight fresh hypotheses compiled, executed, and passed independent validation. Two established a measured current gap, five established no current gap, and one terminated as PROBE_NOT_EXECUTABLE because Hiro's current runtime does not expose the required hidden-state interface.
Three historical controls also executed. They produced one measured gap and two no-current-gap outcomes, showing that the apparatus measures behavior rather than inferring results from whether code already exists.
Every executed probe froze cases, metric, aggregation, threshold, resource bounds, and interfaces before execution; retained raw observations and resource receipts; and underwent independent metric recalculation from hashed artifacts.
The verified probe results were fed only into the preserved Phase 3F-V viability assessment. The two measured-gap hypotheses became VIABLE_FOR_REPRODUCTION, five became NO_CURRENT_GAP, and the unexecutable hypothesis remained GAP_UNPROVEN.
No reproduction was executed. The viable findings are evidence-backed inputs for a separate future authorization, not candidates or approved improvements.
The complete Hiro repository suite passed 873 tests with one expected skip. Hiro was restarted and is healthy with the local model connected and its loaded and checkout revisions aligned at ed102076d9d9f0b3c4702bdf3196e2eb15edfb64.
Phase 3F-VR is empirically demonstrated: Hiro can execute preregistered baseline-only probes, preserve raw observations, independently calculate metrics, and determine whether proposed current gaps actually exist.