Hiro development journal

Phase 3F-VS binds measured gaps to real production surfaces

Phase 3F-VS implementation, full Hiro validation, public journal tests, and the production journal build completed successfully Machine-readable JSON

Executive summary

Phase 3F-VS tested whether Hiro can establish that a numeric capability gap belongs to real production behavior and that a proposed candidate surface causally controls the measured behavior before reproduction is authorized.

The new binder uses exact repository paths and symbols, static call paths, code-owned runtime receipts, source hashes, and per-path causal classification. Claim IDs and source-paper wording do not select or certify bindings.

The two Phase 3F-VC false positives were confirmed as lab-only measurements of Daylab suite selection. Their numeric results remain valid, but they are now MEASURED_GAP_UNBOUND rather than production deficiencies.

Five other preserved probes use a verified adapter that reaches Hiro's real local-model router. Their proposed agent, memory, and metacognition surfaces were not on the measured execution path, and all five had already returned NO_CURRENT_GAP.

The remaining preserved hypothesis had no executable hidden-state interface and terminated SUBJECT_BINDING_UNCERTAIN with NO_TRANSFER_SURFACE.

Two immutable historical non-meta improvements passed positive-control binding at the exact modified symbols: reflection synthesis and memory-candidate normalization. Both were DIRECT_CONTROL, proving the binder does not merely reject every input.

No preserved fresh hypothesis remains VIABLE_FOR_REPRODUCTION. No discovery, reproduction, candidate construction, production write, promotion, or meta-improvement occurred.

The complete Hiro suite passed 883 tests with one expected skip. Hiro is healthy with the local model connected and its loaded and checkout revisions aligned at 28589bf3b63444e2a79d7706c4b71a8e9dc1ca48.

Work completed

Production-subject binding contract

Completed
  • Defined PRODUCTION_SUBJECT, PRODUCTION_BOUNDARY_ADAPTER, LAB_ONLY_SUBJECT, and SUBJECT_BINDING_UNCERTAIN as explicit, mutually exclusive subject classes.
  • The contract records capability identity, probe entrypoint, measurement adapter, production callable, normal runtime entrypoint and callers, relevant configuration, probe and production traces, proposed paths, source hashes, and runtime receipt.
  • A numeric MEASURED_GAP is now MEASURED_GAP_UNBOUND unless the subject is production-real, the probe and production paths reach it, and a candidate surface has verified causal control.

Deterministic call and surface verification

Completed
  • Built a repository call graph from Python syntax, resolving module-level and function-local imports without executing source text.
  • Binding verification requires exact paths and symbols to exist, a static probe-to-subject path, a normal-production path where applicable, and a code-owned runtime call receipt.
  • Candidate surfaces are classified DIRECT_CONTROL, BOUNDED_UPSTREAM_CONTROL, BOUNDED_DOWNSTREAM_CONTROL, CONTEXT_ONLY, or NO_TRANSFER_SURFACE. Only the first three can support reproduction viability.

Eight-hypothesis reassessment

Completed
  • Reused the accepted Phase 3F-VR measurements and did not rerun source discovery or either Phase 3F-VC reproduction.
  • Classified two measured gaps as LAB_ONLY_SUBJECT plus CONTEXT_ONLY, five no-gap probes as PRODUCTION_BOUNDARY_ADAPTER plus CONTEXT_ONLY, and one unexecutable probe as SUBJECT_BINDING_UNCERTAIN plus NO_TRANSFER_SURFACE.
  • The resulting funnel contains zero MEASURED_PRODUCTION_GAP and zero VIABLE_FOR_REPRODUCTION findings.

Historical positive controls

Completed
  • Selected passing historical controls structurally from immutable eligible candidate artifacts rather than hard-coded identities.
  • Failure attribution bound the evaluator and normal runtime to synthesize_reflection; its changed file directly implements the measured subject and the preserved score improved from 0.1667 to 0.8333.
  • Candidate selection bound the evaluator and normal runtime to normalize_candidates; its changed file directly implements the measured subject and the preserved score improved from 0.0 to 1.0.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Canonical Phase 3F-VS report passed The final immutable report verified with zero reason codes, eight preserved hypotheses, two passing historical controls, and no downstream authority.
Focused binding and compatibility tests passed 24 tests passed, covering lab-only rejection, real production-boundary adaptation, direct control, mandatory runtime receipts, exact modified-symbol selection, updated probe compilation, viability normalization, and Phase 3F-VC compatibility.
Complete Hiro repository suite passed 883 tests passed, one expected test was skipped, and six non-failing pre-existing marker warnings were reported in 443.15 seconds.
Live runtime identity passed Hiro reported healthy with the local model connected and checkout and loaded runtime both at 28589bf3b63444e2a79d7706c4b71a8e9dc1ca48.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 182 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was outside this session's scope.

Current state

Next steps