Hiro development journal

Graduating Hiro's Stage 4 proposal-only evaluation cycle

Validated and published Machine-readable JSON

Executive summary

Hiro satisfied the sealed Stage 4 graduation criteria after an autonomous real-Qwen campaign of natural bounded repair candidates, repeated eligible-candidate latency trials, interruption recovery, and independent proposal review.

Eleven natural candidate attempts produced ten eligible candidates that passed public, external held-out, regression, invariant, latency, integrity, and sealed-oracle checks. The one deficient candidate was correctly rejected, preserved, and replaced by a generalized correction.

All eleven proposal-gate decisions matched the independently sealed oracle. The campaign recorded zero false accepts, false rejects, incomplete evaluation runs, integrity failures, model failures, infrastructure failures, merges, promotions, or deployments.

The append-only graduation ledger contains 50 started and 50 completed evaluation runs with 400 observations. Five repeated latency trials remained eligible, with worst public and held-out p95 ratios of 0.9863 and 1.0665 against the sealed 1.10 ceiling.

Stage 4 is graduated. Hiro is ready to enter Stage 5 reviewed integration and monitoring, but no candidate has been approved, merged, promoted, or deployed.

Work completed

Sealed natural-candidate graduation campaign

Passed
  • The graduation thresholds were written and hashed before candidate construction: at least ten eligible natural candidates, perfect gate-decision accuracy, five repeated eligible latency trials, three new interruption boundaries, no integrity or authority failures, and evidence-based bounded proposals.
  • Qwen 3.6 35B-A3B constructed candidates for numeric clamping, inclusive windows, retry backoff, boolean configuration parsing, safe ratios, cache freshness, override precedence, chunk boundaries, ordered tag normalization, and canonical priority mapping.
  • Each candidate began at the same sealed baseline commit in a distinct external worktree and was limited to one implementation path and one targeted-test path.

Behavioral rejection and autonomous correction

Correctly rejected and corrected
  • The initial boolean-parser candidate handled normalized true and false strings but did not generalize to the complete common-token contract exercised by held-out evaluation.
  • The Stage 4 gate rejected it for held-out score, confidence, invariant, and category evidence even though its public and targeted tests passed. Its frozen packet and recommendation remain immutable and are excluded from the ten eligible-candidate count.
  • A new candidate used an explicit generalized hypothesis for true, yes, on, 1 and false, no, off, 0 tokens plus unknown and non-string errors. It passed public, held-out, regression, invariant, integrity, latency, and sealed-oracle checks.

Proposal quality and authority audit

Passed
  • All eleven recommendations cited matching public, held-out, latency, regression, and evaluation identifiers and remained explicitly subject to human approval.
  • All eleven frozen candidates addressed one hypothesis, changed only their two predeclared paths, included their targeted test, remained within the two-repair ceiling, and passed packet and snapshot integrity checks.
  • No candidate was merged, promoted, deployed, committed into the fixture baseline, or granted service or scheduler authority.

Interruption and idempotency coverage

Passed
  • A partially recorded held-out run resumed to the exact unique repetition set without losing or duplicating observations.
  • Already completed public and held-out evaluations were reused instead of rerun.
  • A simulated interruption after proposal recording but before recommendation freeze resumed to one proposal and one frozen recommendation without duplicate proposal records.
  • During campaign continuation, an attempted second process correctly refused an already owned external worktree. The original single owner completed and froze the candidate, which the continuation then reused. This wrapper-level contention did not create an evaluation run or enter the successful-run count.

Latency, thermal, and repository validation

Passed
  • Five repeated evaluations of the eligible clamp candidate all remained eligible. The worst public p95 ratio was 0.9863 and the worst held-out p95 ratio was 1.0665, both below 1.10.
  • Observed GPU temperature ranged from 36 to 52 degrees Celsius; the campaign's 80-degree thermal pause was never reached.
  • The independent repository suite completed with 279 passing tests in 55.27 seconds.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Natural candidate outcomes passed Eleven attempts yielded ten eligible sealed-oracle passes and one correctly rejected sealed-oracle failure; all eleven gate decisions were correct.
Append-only evaluation ledger passed 50 runs started, 50 runs completed, 400 observations recorded, and zero incomplete runs remained.
Repeated eligible latency passed Five of five trials remained eligible; maximum public and held-out p95 ratios were 0.9863 and 1.0665 under the 1.10 limit.
Interruption recovery passed Held-out observation resume, completed-evaluation reuse, and recommendation-freeze resume all completed without duplicated observations or proposals.
Candidate and recommendation audit passed Zero candidate-integrity, recommendation-integrity, scope, targeted-test, repair-bound, proposal-evidence, or proposal-authority failures across eleven attempts.
Failure classification passed Zero model or infrastructure failures were included in the campaign; the one rejected behavioral candidate was preserved and excluded from the eligible count.
Repository-wide regression suite passed 279 tests passed in 55.27 seconds.
Integration authority boundary passed Merge, promotion, and deployment counts remained zero; the fixture source repository stayed clean with its single sealed baseline commit.
Hiro journal generation and frontend build passed Timestamped-entry tests passed, the generator produced and validated 59 journal pages, and the TypeScript and Vite production build completed successfully.

Current state

Next steps