Hiro development journal

Phase 3F-VR proves current-capability measurement

Phase 3F-VR implementation, full Hiro validation, public journal tests, and the production journal build completed successfully Machine-readable JSON

Executive summary

Phase 3F-VR added and qualified a generic baseline-only probe boundary for measuring Hiro's current behavior before research reproduction is considered.

The qualification reused the same eight preserved Phase 3F-V hypotheses. It performed no discovery, source treatment, reproduction, candidate construction, production mutation, promotion, or meta-improvement.

Seven of eight fresh hypotheses compiled, executed, and passed independent validation. Two established a measured current gap, five established no current gap, and one terminated as PROBE_NOT_EXECUTABLE because Hiro's current runtime does not expose the required hidden-state interface.

Three historical controls also executed. They produced one measured gap and two no-current-gap outcomes, showing that the apparatus measures behavior rather than inferring results from whether code already exists.

Every executed probe froze cases, metric, aggregation, threshold, resource bounds, and interfaces before execution; retained raw observations and resource receipts; and underwent independent metric recalculation from hashed artifacts.

The verified probe results were fed only into the preserved Phase 3F-V viability assessment. The two measured-gap hypotheses became VIABLE_FOR_REPRODUCTION, five became NO_CURRENT_GAP, and the unexecutable hypothesis remained GAP_UNPROVEN.

No reproduction was executed. The viable findings are evidence-backed inputs for a separate future authorization, not candidates or approved improvements.

The complete Hiro repository suite passed 873 tests with one expected skip. Hiro was restarted and is healthy with the local model connected and its loaded and checkout revisions aligned at ed102076d9d9f0b3c4702bdf3196e2eb15edfb64.

Phase 3F-VR is empirically demonstrated: Hiro can execute preregistered baseline-only probes, preserve raw observations, independently calculate metrics, and determine whether proposed current gaps actually exist.

Work completed

Generic baseline-probe contract

Completed
  • Added a declarative contract requiring probe and claim identity, immutable provenance, current capability, exact observable, metric, frozen cases and repetitions, aggregation, gap rule, artifacts, resource bounds, interfaces, isolation, determinism classification, seeds, and preregistration identity.
  • The semantic compiler selects reusable measurement primitives from hypothesis and observable fields. It does not branch on claim IDs, source titles, or expected results.
  • Malformed plans terminate explicitly, and a hidden-state hypothesis terminates as PROBE_NOT_EXECUTABLE when the current runtime cannot expose the required observations without changing production functionality.

Isolated execution and independent validation

Completed
  • Current Hiro interfaces execute frozen synthetic cases in a disposable workspace. A restricted, network-denied worker calculates raw rows and a result without production authority.
  • A separate host validator independently reconstructs the declared metric from the frozen specification and subject observations, checks every case and repetition identity, confirms the original threshold, and validates artifact hashes and resource receipts.
  • The implementation reuses only the safe isolation, hashing, immutable-output, and restricted-execution substrate from Phase 3F-R. Baseline probes retain a distinct schema and are never represented as reproduction experiments.

Eight-hypothesis and control qualification

Completed
  • All seven executable fresh probes used six frozen cases and one repetition. Both adaptive-suite relevance hypotheses measured an irrelevant-selection rate of 0.8333333333333334 against a frozen gap threshold of 0.30.
  • Five model-contract probes measured a success rate of 1.0 against frozen gap thresholds of 0.85 and therefore correctly returned NO_CURRENT_GAP.
  • The hidden-state uncertainty hypothesis did not fabricate evidence and terminated with the explicit reason CURRENT_RUNTIME_HIDDEN_STATE_INTERFACE_UNAVAILABLE.
  • Historical full-context and bounded-reasoning-depth controls returned NO_CURRENT_GAP at success rate 1.0. The adaptive-reasoning-cost control returned MEASURED_GAP at routing accuracy 0.6666666666666666 against a frozen 0.90 threshold.

Measured-evidence viability feedback

Completed
  • Added a verified probe-evidence handoff to the existing Phase 3F-V assessment. The handoff verifies the immutable probe report and binds exact outcome, metric, and frozen decision rule without granting reproduction authority.
  • The final downstream funnel contains two VIABLE_FOR_REPRODUCTION, five NO_CURRENT_GAP, and one GAP_UNPROVEN among the eight preserved hypotheses.
  • A new immutable final report links the unchanged probe report to the downstream viability report and records two VIABLE_FOR_REPRODUCTION_AFTER_MEASUREMENT outcomes.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused probe and viability tests passed 15 tests passed. The only warning was the existing inaccessible pytest cache directory.
Fresh probe qualification passed Seven probes preregistered, compiled, executed, and validated; two returned MEASURED_GAP, five returned NO_CURRENT_GAP, one returned PROBE_NOT_EXECUTABLE, and no build, runtime, validation, or resource failure occurred.
Independent artifact verification passed Ten evidence packages were verified across seven fresh probes and three historical controls with zero verification reason codes.
Measured-evidence viability feedback passed The same eight preserved hypotheses terminated as two VIABLE_FOR_REPRODUCTION, five NO_CURRENT_GAP, and one GAP_UNPROVEN. No reproduction authority was exercised.
Complete Hiro repository suite passed 873 tests passed, one expected test was skipped, and seven non-failing warnings were reported in 407.04 seconds.
Live runtime identity passed Hiro reported healthy with the local model connected and checkout and loaded runtime both at ed102076d9d9f0b3c4702bdf3196e2eb15edfb64.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 180 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not modified because dependency maintenance was outside this session.

Current state

Next steps