Hiro development journal

Phase 3D demonstrates evidence-to-candidate qualification

Evidence-to-candidate conversion works; two independently supported findings produced promotion-eligible candidates without promotion or activation Machine-readable JSON

Executive summary

Qualified only Hiro's ability to convert the seven accepted Phase 3A-R supporting evidence packages into bounded, non-meta candidate implementations. Discovery, claim extraction, reproduction evidence, corroboration semantics, governor policy, promotion thresholds, activation, and Phase 2 promotion machinery were not changed.

Froze all seven evidence inputs, transfer assessments, four candidate specifications, baseline measurements, development cases, acceptance thresholds, resource limits, prohibited scope, and a separate evaluator-owned holdout vault before autonomous candidate construction.

Correctly concluded that three supported mechanisms did not justify production changes: planner handoff and structured artifacts were already present, while residual-guided verification lacked independently labeled residual authority.

The autonomous builder attempted four low-risk candidates in disposable external worktrees. Three valid implementations were constructed; one concurrency candidate never satisfied its frozen construction contract.

All three valid implementations passed deterministic implementation-fidelity assessment, preregistered target-gain evaluation, and frozen holdout evaluation. One premise-correction implementation then correctly failed its hard 60-line resource limit after changing 261 production lines.

The causal failure-attribution and quality-aware memory-selection candidates each passed hard resource limits, the actual complete repository regression suite, and the existing governor's read-only validation boundary.

Both eligible candidates have complete immutable provenance chains from source evidence through governor decision. Their exact candidate revisions remain isolated and unmerged.

The qualification report verified with zero reason codes. Production HEAD was identical before and after qualification; zero promotion transactions and zero activations occurred.

The final disposition is PHASE 3D DEMONSTRATED — EVIDENCE-TO-CANDIDATE WORKS. This session stopped at PROMOTION_ELIGIBLE and did not authorize or perform promotion.

Work completed

Immutable evidence and transfer assessment

Completed
  • Verified the existing Phase 3A-R report and all seven authorized supporting evidence-package identities before qualification.
  • Produced an explicit transfer assessment for every finding: demonstrated mechanism, Hiro relevance, possible target component, transfer rationale, assumptions, falsifier, smallest implementation, blast radius, and whether production modification was warranted.
  • Preserved three SUPPORTED_BUT_NO_ACTIONABLE_IMPLEMENTATION outcomes instead of forcing every supported research result into code.
  • No new finding was introduced and no negative or deferred Phase 3A-R result was substituted for an authorized input.

Preregistered candidate and holdout protocols

Completed
  • Froze four candidate specifications before implementation: causal failure attribution, bounded trip-search concurrency, grounded premise correction, and quality-aware memory selection.
  • Each specification records its evidence identity, target capability and component, exact intended mechanism, baseline, metric, minimum gain, development cases, holdout design, permitted files, prohibited scope, resource limits, rollback expectation, and low-risk classification.
  • Holdout cases were stored in a separate immutable vault before build. Expected holdout outputs were withheld from the candidate builder.
  • Candidate acceptance thresholds were not lowered and holdout cases were not changed after results were observed.

Isolated autonomous construction

Completed
  • Used Hiro's actual candidate builder and local candidate agent in external disposable Git worktrees with path-scoped authority.
  • Candidate source material was provided as structured validated evidence rather than executable instruction authority.
  • The original candidate builds and target/holdout evaluation ran through the existing WSL2/Bubblewrap restricted executor with OS isolation and network denial.
  • Immutable candidate packets preserve exact file snapshots, per-file hashes, candidate diffs, construction attempts, baseline contrast, scope audit, and targeted construction tests.

Independent candidate evaluation

Completed
  • Verified implementation fidelity structurally and behaviorally before treating measured improvement as meaningful; comments or claimed intent were never accepted as proof.
  • All three valid implementations passed their frozen development and holdout thresholds. The causal-attribution candidate improved from 1/6 to 5/6 development cases and passed 4/5 fresh holdouts. Premise correction and memory selection each achieved their preregistered perfect target and holdout scores.
  • Resource evaluation separated production complexity from candidate test code and measured latency at the operation boundary. Causal attribution changed 20 production lines and memory selection changed 12; both added no model calls, network destinations, or persistent writes.
  • Premise correction stopped at RESOURCE_CONSTRAINT_FAIL because 261 production lines exceeded its fixed maximum of 60. Full regression and governor assessment correctly did not run for it.

Complete regression and governor boundary

Completed
  • Only the two candidates passing fidelity, target, holdout, and resource gates incurred the complete repository regression cost.
  • Causal failure attribution completed with 837 tests passed and six non-failing warnings in 494.44 seconds. Quality-aware memory selection completed with 837 tests passed and six non-failing warnings in 465.11 seconds.
  • The existing continuous governor validated both candidates under unchanged policy and returned PASS for each.
  • The qualification runner invoked only the governor's read-only validation method. It contains no promotion call, created no promotion transaction, performed no activation, and did not restart Hiro.

First-divergence repairs

Completed
  • Preserved every failed run in a versioned immutable directory and repaired only the first observed qualification boundary before replay.
  • Bound candidate identifiers to the existing 64-character contract; separated production/test resource accounting and microbenchmark noise; and supplied edit-shape guidance compatible with the existing one-patch-per-path contract.
  • Made consolidation replay immutable per-file snapshots and include both tracked changes and untracked candidate tests in identity verification.
  • A complete-suite replay exposed missing Windows process-contract values and a missing ignored pinned-runtime junction in fresh worktrees. The existing child-environment normalization and runtime provisioning path were minimally repaired; the exact two launcher tests passed before both complete suites were rerun successfully.
  • No repair changed evidence, holdouts, candidate behavior requirements, acceptance criteria, corroboration semantics, governor policy, or promotion thresholds.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Canonical Phase 3D report verification passed The verifier rehashed the report and both eligible provenance audits, returned valid true with zero reason codes, and confirmed two promotion-eligible candidates, zero promotion transactions, and zero activations. Final report SHA-256: 9e6c4b835a59390c8772945867b66d8fdd3f6e2b7ed4fc23cdebc0fbe7a1cc50.
Final qualification funnel passed 7 supported findings; 7 valid transfers; 3 no-action dispositions; 4 candidate specs and build attempts; 3 successful builds; 3 fidelity, target, and holdout passes; 2 resource passes and 1 resource failure; 2 full-regression passes; 2 governor passes; 2 promotion-eligible candidates.
Focused qualification and upstream regression tests passed 77 tests passed in 31.92 seconds. One non-failing warning concerned the inaccessible pytest cache directory.
Eligible candidate complete repository suites passed Both exact candidate worktrees passed the actual complete suite: 837 tests in 494.44 seconds and 837 tests in 465.11 seconds, each with six non-failing warnings.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 174 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session.

Current state

Next steps