Hiro development journal

Phase 3F-V finds the missing pre-reproduction gap probe

Phase 3F-V implementation, full Hiro validation, public journal tests, and the production journal build completed successfully Machine-readable JSON

Executive summary

Phase 3F-V added a non-authoritative viability boundary before research reproduction. It inspects current production capabilities, requires measurable gap evidence, checks for a bounded implementation surface, and returns an explicit pre-reproduction disposition.

The six original Phase 3F findings were used only as a diagnostic corpus. An initial diagnostic exposed a real false-positive defect: generic runtime failures and hypothetical new modules were accepted as evidence of a current bounded opportunity.

The minimum repair required a numeric baseline tied to cited telemetry, existing repository evidence for required input signals, an existing production path, and a production change rather than evaluation-only instrumentation. The preserved six plans then produced zero false viable findings and six GAP_UNPROVEN outcomes.

A fresh bounded campaign considered 72 sources, retained 14 after excluding prior material, validated 32 claims, and produced eight explicit Hiro-transfer feasibility plans. Twenty-four claims terminated NOT_HIRO_RELEVANT.

All eight fresh plans were assessed and independently audited. The model proposed all eight as viable, but deterministic validation rejected all eight as GAP_UNPROVEN because none established the required current numeric baseline. Seven cited no runtime measurement and one cited telemetry without reproducing its count in the baseline observation.

Three independent model audits also incorrectly treated absence of a specific mechanism as proof of an improvement gap. The deterministic evidence contract correctly vetoed those judgments, demonstrating why model agreement alone is insufficient.

The first unresolved boundary is CURRENT_CAPABILITY_PROBE to MEASURED_GAP. Hiro can describe a cheap baseline probe but does not yet have a safe declarative primitive that executes and freezes that baseline before deciding viability.

No reproduction experiment ran, no candidate was constructed, no promotion was requested, and no meta-improvement began. The complete Hiro repository suite passed 864 tests with one expected skip.

Hiro was restarted and is healthy with the local model connected and its checkout and loaded runtime aligned at revision 6f317c9242b78e613249fc42eb046632c064a13a.

Phase 3F-V is not demonstrated. The evidence identifies a gap-assessment interface failure rather than a downstream reproduction, candidate, governor, or promotion failure.

Work completed

Original six-finding diagnostic

Completed
  • Repository inspection established that the context-projection, adaptive-reasoning-cost, and bounded-reasoning-depth mechanisms already existed, but no claim-specific current performance baseline proved a remaining optimization gap.
  • Repository inspection established that structured action-sequence search had no typed, scoreable production action-sequence contract and that graph retrieval had no relation or edge metadata in the current recall interface.
  • The earliest reliable stop for the first three findings was gap assessment; for the remaining three it was transfer-surface assessment. Their historical Phase 3F dispositions were not changed.

Capability, gap, and transfer contract

Completed
  • Added explicit capability states for absent, present-with-gap, present-without-gap, and uncertain capabilities.
  • Added structured current-capability evidence, gap hypothesis, baseline method, source mechanism, expected direction, cheap-probe description, target component, permitted paths, required input evidence, blast radius, holdout design, and resource observables.
  • Repository evidence includes AST-derived function behavior such as called functions, branch count, and await count rather than relying only on symbol names and docstrings.
  • Deterministic validation rejects invented evidence references, evidence-reference strings used as paths, speculative new production modules, evaluation-only changes, unmeasured gap prose, and prohibited improvement-controller scope.
  • The policy grants no reproduction, candidate-construction, promotion, or meta-improvement authority.

Fresh bounded qualification

Completed
  • The normal source collector considered arXiv, GitHub, and Moltbook material. Seventeen sources from the prior corpus were excluded by immutable source identity; no title or claim-specific allowlist was used.
  • Seventy-two sources were considered and fourteen retained, all from arXiv. Thirty-two claims validated; feasibility planning found eight local-ready explicit transfer hypotheses and classified twenty-four as not Hiro-relevant.
  • The final fixed-input viability report assessed all eight relevant plans and returned eight GAP_UNPROVEN outcomes and zero viable findings.
  • The campaign stopped before reproduction exactly as authorized.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused viability and fresh-source tests passed 10 tests passed; the only warning was the existing inaccessible pytest cache directory.
Original six-plan diagnostic after repair passed Six relevant plans terminated GAP_UNPROVEN and zero terminated VIABLE_FOR_REPRODUCTION; the immutable report SHA-256 is ff49c97ae264553ce2b6b9834e5cd95cd461deedd91fbe8a90623b53ef7f616f.
Fresh fixed-input qualification passed All eight relevant fresh plans completed assessment and audit; eight terminated GAP_UNPROVEN and zero became viable. The final immutable report SHA-256 is df1df34eedfa0a1fe768304cde5946ace7f5e39080d7f1c84ae904c7cf76e1ae.
Complete Hiro repository suite passed 864 tests passed, one expected test was skipped, and seven non-failing warnings were reported in 514.95 seconds.
Live runtime identity passed Hiro reported healthy with the local model connected and checkout and loaded runtime both at 6f317c9242b78e613249fc42eb046632c064a13a.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 179 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not modified because dependency maintenance was outside this session.

Current state

Next steps