Hiro development journal

Phase 3F-R demonstrates generic plan-to-reproduction

All six preserved fresh Phase 3F plans compiled, preregistered, executed in isolation, and produced independently verified evidence without fixed corpus or claim-ID assumptions Machine-readable JSON

Executive summary

Phase 3F-R addressed only the first failed Phase 3F boundary: converting fresh valid declarative experiment plans into preregistered, isolated, independently verified reproduction outcomes. No new discovery, candidate construction, promotion, corroboration-policy change, or meta-improvement was authorized.

The six exact Phase 3F feasibility plans and the historical Phase 3A-R report were frozen before implementation. The source feasibility report retained SHA-256 41b3c72119629ab8c339bc04ab77b2517b2078d94ad73aa2d70680336731992f.

The audit separated historical fixture constraints—exactly fourteen plans, fixed claim IDs, a claim-ID protocol registry, and Phase 3A-R-specific naming—from legitimate safety constraints such as preregistration, immutable provenance, code-owned primitives, resource limits, OS isolation, network denial, raw-artifact retention, and independent recalculation.

An additive generic compiler now selects from a safe primitive catalog using declared mechanism semantics rather than claim identity or source title. Unsupported mechanisms terminate explicitly instead of receiving fabricated experiments.

All six preserved plans compiled and preregistered before execution. All six then ran through the existing OS-isolated restricted executor with calculation-network denial and no model bridge, production write, candidate, or promotion authority.

All six result packages passed identity, sample-count, baseline/treatment execution, criterion-integrity, resource, artifact-hash, and independent metric-recalculation checks. All six bounded transfer experiments returned REPRODUCTION_SUPPORTED.

The immutable historical Phase 3A-R report reverified with zero reason codes. All fourteen legacy specifications remained compatible and all thirteen valid historical evidence packages rehashed successfully; no prior evidence was rewritten.

The final generic report verified from disk with zero reason codes. The complete repository suite passed 851 tests with one expected skip and seven warnings.

Hiro was restarted on committed revision 20a51e90ac18105c046eacd75dea68cb449f5437, and checkout and loaded runtime identities match. The original Phase 3F campaign remains stopped.

The final disposition is PHASE 3F-R DEMONSTRATED — GENERIC PLAN-TO-REPRODUCTION WORKS. These synthetic transfer results make the preserved plans eligible for a separately authorized downstream phase; they do not themselves warrant a production candidate.

Work completed

Fixture-assumption audit

Completed
  • Identified the exact-cardinality guard requiring fourteen plans and the following claim-ID registry membership check as the original Phase 3F blockers.
  • Classified exact corpus size, claim IDs, phase names, and claim-selected prewritten implementations as historical qualification fixtures rather than runtime safety requirements.
  • Retained source and claim hashes, frozen hypotheses and criteria, result-free preregistration, code-owned execution, hidden ground truth where applicable, isolation, bounded resources, independent recalculation, and explicit authority denial as mandatory constraints.
  • Preserved the historical runner and verifier for backward compatibility rather than rewriting or invalidating Phase 3A-R evidence.

Generic declarative compiler

Implemented
  • Added a schema validator for claim identity, source provenance, Hiro-local hypothesis, transferable mechanism, observable, falsifier, assets, isolation requirements, resource estimates, and feasibility state.
  • Compilation freezes explicit baseline, treatment, bounded data-only cases, metric, acceptance criterion, falsification criterion, sample strategy, resource bounds, isolation requirements, expected artifacts, compiler identity, semantic selection signals, and source-plan hash.
  • The initial reusable safe primitives cover context projection, planning cost/accuracy, reasoning depth, structured search, and graph retrieval.
  • Primitive selection excludes claim identity and source title. An arbitrary unknown claim ID with the same mechanism selects the same primitive; an unknown mechanism terminates as UNSUPPORTED_EXPERIMENT_TYPE.
  • Input cases cannot carry command, script, executable, or Python-source fields. The isolated worker accepts only the code-owned primitive allowlist and JSON fixtures.

Immutable execution and independent verification

Completed
  • Each plan received an immutable preregistration hash and compiled-spec hash before a result directory was created.
  • The existing restricted executor copied a fixed metric worker into each disposable workspace, denied calculation-network access, enforced time and output boundaries, and retained raw and calculated JSON artifacts.
  • Host-side code independently recomputed accuracy, effect, mean cost, cost reduction, precision gain, low-similarity relational recovery, and result-count ratios from raw rows.
  • A standalone disk verifier rehashed the manifest, every evidence package, raw measurements, calculated result, preregistration identity, isolation receipt, resource receipt, and authority boundary.

Fresh six-plan qualification

Passed
  • Context-asymmetric routing compiled to context projection and produced a paired accuracy effect of 1.0 with a treatment cost ratio of 1.05.
  • Adaptive reasoning cost compiled to planning cost/accuracy and retained perfect labeled accuracy while reducing the operation-cost proxy by 0.8043.
  • Under-reasoning compiled to reasoning-depth comparison and produced a paired completion-accuracy effect of 1.0 under the frozen depth contrast.
  • Structured exploration compiled to matched-budget search and achieved 1.0 versus 0.25 optimal-path success.
  • Both GraphMemix transfer hypotheses compiled to bounded graph retrieval. Each produced a 0.5 precision gain, ten low-similarity relevant recoveries, and unchanged retrieval count.
  • All six outcomes are bounded synthetic Hiro-local transfer evidence. None reproduces a source paper's full external benchmark or proves that a production change is warranted.

Compatibility and first-divergence repairs

Completed
  • The historical report retained SHA-256 423a41156b6f15c1cba465df6bda904bc6c1813bea734eaec3f04ab1c4c2820d, and its fourteen specifications and thirteen valid evidence packages reverified.
  • A focused validator test found that a malformed plan with no resource object was initially mislabeled RESOURCE_POLICY_VIOLATION. The minimum repair made resource classification contingent on a present resource plan, restoring PLAN_SCHEMA_INVALID as the first reason.
  • The first fresh execution preceded explicit resource receipts. It was preserved, resource verification was added without changing plans or results, and a new fully preregistered versioned run became the final evidence.
  • No hypothesis, acceptance criterion, case, or observed result was altered to obtain success.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Fresh generic qualification funnel passed 6 received, 6 compiled, 6 preregistered, 6 executed, 6 valid, and 6 REPRODUCTION_SUPPORTED; zero schema, underspecification, unsupported, build, runtime, invalid, inconclusive, contradictory, or resource failures.
Final evidence verifier passed The final report reverified with zero reason codes, six immutable preregistrations and evidence packages, historical compatibility true, and candidate and promotion authority false. Report SHA-256: 84760ceee954a91d760c47541f390e72a51673941c10978366926c06c76778e2.
Historical compatibility passed The original Phase 3A-R report hash remained unchanged; all 14 legacy specifications were contract-compatible and all 13 valid evidence packages rehashed successfully.
Focused reproduction, isolation, evidence, and feasibility tests passed 22 tests passed. One non-failing warning concerned the inaccessible pytest cache directory.
Complete Hiro repository suite passed 851 tests passed with one expected skip and seven warnings in 452.15 seconds.
Live runtime identity passed After the checked restart, checkout and loaded runtime both matched committed revision 20a51e90ac18105c046eacd75dea68cb449f5437 and Hiro reported healthy with the local model connected.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 177 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session.

Current state

Next steps