Hiro development journal

Phase 3A-R demonstrates trustworthy reproduction execution

Evidence acquisition works; thirteen experiments completed validly and seven produced admissible supporting evidence Machine-readable JSON

Executive summary

Qualified only Hiro's ability to execute the fourteen accepted Phase 3C experiment plans and turn their observations into claim-specific evidence. Candidate construction, production changes, promotion, activation, rollback, corroboration semantics, claim extraction, discovery ranking, and meta-improvement were unchanged.

Froze all fourteen source claims, provenance spans and hashes, Hiro-local hypotheses, baselines, treatments, measurements, acceptance and falsification criteria, asset plans, resource estimates, sample strategies, and complete experiment specifications before observing any result.

Built a declarative reproduction runner that separates experiment code from production candidates. Synthetic fixtures and hidden ground truth are code-owned; local-model subjects receive neither ground truth nor the other experimental condition.

Executed numerical calculation inside disposable WSL2/Bubblewrap workspaces with OS isolation and network denial. A typed loopback bridge supplied synthetic cases to the fixed local model where a plan required model behavior, without exposing source text, credentials, arbitrary code, or production-write authority.

Thirteen plans were experimentally executable and all thirteen completed with valid, independently recalculated measurements. Seven satisfied their preregistered local criteria and six validly did not. Negative results were preserved as successful evidence acquisition rather than experiment failures.

One plan was classified EXPERIMENT_INVALID before build because it attempted to attribute failures away from model size while providing only one model and no independent model-size comparator. The plan and reason remain preserved rather than being weakened.

The first execution attempt for the planner-handoff experiment failed before measurement because a batched model response did not preserve the frozen case identities. The immutable v1 failure remains recorded; a minimal preregistered v2 transport repair requested one response per unchanged case and then completed validly.

The unchanged corroboration truth table produced seven PASS, zero FAIL, and seven DEFER decisions. Every PASS and two valid DEFER results have complete source-to-decision provenance audits.

The final disposition is PHASE 3A-R DEMONSTRATED — EVIDENCE ACQUISITION WORKS. Seven evidence packages are eligible as inputs to a later separately authorized candidate-construction phase, but no candidate was constructed or promoted.

Work completed

Immutable preregistration

Completed
  • The accepted fourteen-plan Phase 3C feasibility report was used as the only qualification corpus; no replacement plan or new discovery result was selected.
  • Each preregistration preserves the exact source claim, supporting source location and hashes, local transfer hypothesis, baseline, treatment, measurement, acceptance criterion, falsifier, required assets, expected resource use, isolation controls, and sample policy.
  • Every preregistration passed the existing evidence-package contract, which rejects missing fields and any result written before execution.
  • The frozen manifest SHA-256 is 944f9942daaadc2715832f6ad7fb6684526c0d8da4e654ba582a09029c81ac5a.

Experiment-only construction

Implemented
  • Added bounded code-owned protocols for planner handoff, failure attribution, resource scheduling, planning complexity, structured artifacts, residual verification, misleading-premise behavior, answer selection, judge rarity, repository-context test generation, and dense-context reliability.
  • Protocols use declarative fixtures and measurements. They cannot write Hiro production code, create a candidate, request promotion, or inherit authority from source material.
  • The experiment worker is copied into a disposable research workspace and consumes only the frozen specification and typed subject observations.
  • Safe assertion experiments accept only a narrow Python assertion grammar and run in the OS sandbox. Other answer-quality experiments compare against hidden deterministic expected values rather than an unrestricted LLM evaluator.

Execution and integrity validation

Completed
  • Thirteen experiment builds were attempted and all thirteen built successfully. All thirteen executable plans completed and passed integrity validation.
  • Validation checked baseline and treatment execution, exact case counts, experiment and attempt identity, preregistration and specification hashes, unchanged acceptance criteria, raw-artifact identity, and deterministic result recalculation.
  • Every numerical result was calculated in the isolated worker and independently recalculated from raw rows by the host validator. All calculations matched.
  • The local model was treated only as a subject. Synthetic prompts excluded hidden answers and the other condition; model identity and telemetry were retained.
  • No experiment downloaded an asset, required a resource reclassification, or touched production.

Observed plan defect

Preserved as invalid
  • One failure-cause plan proposed distinguishing evidence, tool-use, and constraint failures from model-size or reasoning-length effects.
  • The supplied assets contained one local model and no independently varied model-size condition, making the causal comparison unidentifiable.
  • The plan was frozen, classified EXPERIMENT_INVALID with reason PLAN_DEFECT_MODEL_SIZE_COMPARATOR_UNAVAILABLE, and sent to corroboration as unavailable evidence, which correctly deferred it.
  • No weaker hypothesis, substitute metric, or post-result plan was created.

First execution divergence and versioned repair

Repaired
  • Planner-handoff attempt v1 failed before experiment execution because one batched treatment response omitted or duplicated a case identity.
  • The original failure is immutable in the first qualification report, whose SHA-256 is 548c2d51b17220fdc4212240b8b48500c44db20d6f26108e74b840eff6ded5ea.
  • The minimum v2 repair changed only response transport from one batch to one schema-bound response per frozen case. Cases, sample count, prompts, hidden ground truth, model configuration, acceptance criterion, falsifier, and preregistration hash remained unchanged.
  • The repair record was frozen before v2 execution, explicitly denies outcome-based rerunning, and has SHA-256 5181092c9b6b806f551dbe1e5da6d2e58f5b618bf7a9d7eae9784b1ab99dcb6e.
  • Attempt v2 completed validly and met the preregistered criterion.

Reproduction outcomes

Completed
  • Seven experiments supported their local hypotheses: planner/executor handoff, causal graph attribution, resource-aware scheduling, structured artifact handoff, residual-guided verification, misleading-premise answer correction, and combined frequency/quality selection.
  • Six valid experiments did not support their local hypotheses: graph versus flat model attribution, simple versus multi-step planning, misleading-premise routing, judge reliability by candidate rarity, repository-context unit-test generation, and dense-context reliability.
  • Several non-supporting experiments produced 100 percent accuracy in both conditions. The system retained the zero-delta result and deferred corroboration instead of manufacturing an improvement.
  • There were no opposite-direction effects crossing a preregistered threshold, so REPRODUCTION_CONTRADICTED remained zero. There were no inconclusive, build-failed, final runtime-failed, or resource-reclassification outcomes.

Existing corroboration integration

Completed without semantic changes
  • Valid supporting evidence plus the independently established healthy regression context mapped to PASS/CORROBORATED.
  • Valid non-supporting evidence mapped to DEFER/INSUFFICIENT_EVIDENCE rather than being misreported as contradiction.
  • The invalid plan mapped to DEFER/EVIDENCE_UNAVAILABLE. No result mapped to FAIL because no valid opposite-direction evidence or attributable regression occurred.
  • Final counts were seven PASS, zero FAIL, and seven DEFER.
  • Nine complete provenance audits cover every PASS and two valid DEFER results. Each chain records source, claim, hypothesis, preregistration, experiment implementation, raw observations, calculated result, criterion, interpretation, integrity status, and corroboration decision.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused Phase 3A-R and upstream evidence tests passed 21 tests passed. One non-failing warning concerned the inaccessible pytest cache directory.
Pre-execution regression context passed 820 repository tests passed in 421.99 seconds before experiment execution. Six non-failing warnings concerned five existing unregistered marks and the inaccessible pytest cache directory.
Final full Hiro regression suite passed 821 tests passed in 418.60 seconds after the transport repair and evidence verifier were complete, with the same six non-failing warnings.
Frozen evidence-package verification passed The verifier rehashed all thirteen valid evidence packages, raw measurement artifacts, calculated result artifacts, the source manifest, and all seven PASS audits. It returned zero reason codes. The final report SHA-256 is 423a41156b6f15c1cba465df6bda904bc6c1813bea734eaec3f04ab1c4c2820d.
Final reproduction funnel passed 14 plans frozen; 13 builds attempted and successful; 13 executed and valid; 7 supported; 6 not supported; 0 contradicted; 0 inconclusive; 1 invalid plan; 0 final build/runtime failures; 0 resource reclassifications; corroboration 7 PASS, 0 FAIL, 7 DEFER.
Public journal tests and build passed npm run test:hiro passed. npm run build generated and validated 173 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session.

Current state

Next steps