Hiro development journal

Hiro's improvement factory now separates pipeline health from candidate merit

Core process implemented, qualified, and activated; independent inspection found critical completion gaps Machine-readable JSON

Executive summary

Completed and activated a simplified continuous-improvement process centered on reproducible failures, harness-owned tests, isolated candidates, paired evaluation, timed canary, independent promotion, and automatic rollback.

Candidate construction failures no longer become negative verdicts about the underlying idea. Exhausted build attempts move to artifact_blocked with idea_merit_evaluated false and become eligible again only after a relevant builder change.

Valid comparative evaluation failures remain rejected evidence. Security regressions remain terminal, infrastructure failures receive cooldown retries, and canary or governor failures retain their own evidence and revision paths.

The candidate builder now mechanically normalizes harmless generated whitespace and common edit-shape mistakes, uses the exact harness-owned production replay instead of allowing a candidate to author its own approval test, and rejects repeated failed patch semantics.

A durable supervisor watches the improvement factory itself. Recursive model-authored builder-repair tasks were retired; a bounded run of construction failures now stops queue manufacture and exposes a diagnostic incident without consuming idea merit.

Corrected a watchdog defect discovered during implementation: legitimate evaluation rejections no longer count as candidate-builder failures or move the pipeline toward a false global shutdown.

Corrected a Windows linked-worktree defect discovered by release qualification. Qualification pytest roots now anchor at the short common Git checkout instead of inheriting an arbitrarily long release-worktree path.

Preserved all eight previously promoted commits on the live branch. The final release is a fast-forward from the live revision rather than a replacement or history rewrite.

The exact activated revision a11954ef646e7f43404af7909678073af701edae passed 750 repository tests and five consecutive ten-test end-to-end qualification cycles. Hiro restarted healthy with Qwen 3.8 connected and the existing durable queue intact.

A separate read-only whole-code inspection then found that the system is not yet a complete autonomous production loop: candidate execution needs stronger operating-system isolation, the current timed checks are synthetic rather than product-traffic probation, promotion is not transactionally joined to runtime restart verification, and post-promotion rollback monitoring is not wired.

Work completed

Explicit outcome separation

Implemented and active
  • The live queue distinguishes implemented, evaluation-rejected, security-rejected, hypothesis-disproven, infrastructure-blocked, superseded, and artifact-blocked outcomes.
  • Artifact-blocked means patch or test manufacture failed and idea merit was not evaluated. These records are not included in the ordinary rejected count.
  • Evaluation rejection requires a valid candidate and evidence from public and held-out comparisons. Construction text, patch errors, inconclusive baseline contrast, and malformed test evidence cannot masquerade as an idea verdict.
  • Infrastructure faults use a separate retry circuit breaker and do not consume the candidate's bounded engineering-attempt budget.

Reproducible behavior before implementation

Implemented and active
  • The rotating everyday interaction audit exercises thirty-two questions across sixteen assistant categories and retains a deterministic replay fixture and task contract for every admitted failure.
  • Stochastic one-off audit results enter an observing state. A second independent observation is required before the automatic audit admits a patch-authorizing queue item.
  • Legacy single-sample audit records are superseded rather than allowed to consume current candidate capacity.
  • Markdown emphasis is removed for semantic phrase checks so presentation punctuation does not create false improvement incidents.

Harness-owned candidate evidence

Implemented and active
  • For captured interaction failures, the model proposes only production changes. The candidate builder replaces model-authored replay tests with a deterministic test generated from the pre-candidate code-owned fixture.
  • The identical target test must collect on the untouched baseline, reach a marked assertion, fail there by AssertionError, and pass only when the production-reachable boundary changes.
  • Tests that merely classify or log the bad response do not qualify as a response improvement.
  • Repairs start from a clean baseline and receive the exact prior failure plus a bounded excerpt of the failed implementation, while semantic duplicates are mechanically rejected.

Candidate construction reliability

Implemented and active
  • Generated text files are normalized mechanically before validation so trailing whitespace does not consume an autonomous repair attempt.
  • A common multiline insert-before mistake is converted into an exact modify operation while still requiring the old block to resolve in the isolated workspace.
  • Every pytest phase receives an isolated, bounded basetemp outside deep candidate paths, preventing cross-run cache contamination and Windows nested-path failures.
  • Candidate, evaluator, integration, governor, canary, and qualification test roots use purpose-specific short runtime locations.

Independent evaluation and promotion

Implemented and active
  • Candidates run paired public and held-out evaluation, independent regression tests, security comparison, and targeted baseline contrast before Stage 5 integration.
  • Only the exact changed-file manifest verified by frozen Stage 5 evidence reaches the promotion governor; an allowlist is not treated as proof that every permitted file changed.
  • Low-risk and moderate-risk canaries retain their authorized risk-based checkpoints. The stable governor rechecks the complete suite and permits only a fast-forward from the unchanged candidate baseline.
  • Rollback creates an additive Git revert and preserves the failed post-promotion evidence.

Improvement-factory health controller

Implemented and active
  • The durable supervisor records candidate episodes, strategies, failure fingerprints, exact evidence, and candidate-ready receipts in append-only tables.
  • A bounded no-canary watchdog now counts construction failures for one builder compatibility revision. Evaluation rejections are explicitly excluded because they prove the evaluator is working, not that candidate construction is broken.
  • When the construction budget is exhausted, the active artifact is parked without a merit verdict and queue manufacture stops with diagnose_candidate_pipeline as the required action.
  • The Observatory now displays pipeline health, the construction-failure count and threshold, and whether diagnosis is required.

Release reconciliation and activation

Completed
  • The original workspace contained an unrelated uncommitted benchmark experiment. It was preserved and excluded from the process release rather than accidentally committed.
  • The live .hd3 checkout contained eight promoted commits beyond the initial workspace revision. A clean integration branch was based on that live head so none of those improvements was discarded.
  • The final live update was a fast-forward from 4c15149 to a11954e. The durable queue database, local configuration, credentials, model runtime, and historical artifacts remained in place.
  • The exact live checkout received the frozen qualification packet for a11954e and an exact safe-directory entry required by its existing Windows ownership boundary.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused state-machine and release tests passed The combined release passed 79 focused supervisor, queue, Observatory, API, and linked-worktree qualification tests.
Full repository regression suite passed The exact combined revision passed 750 tests in 461.77 seconds. Four generated autonomous tests emitted pre-existing unknown hiro_contract marker warnings; there were no failures.
Frozen five-cycle pipeline qualification passed Revision a11954ef646e7f43404af7909678073af701edae passed five consecutive cycles of ten end-to-end checks, for 50 successful lifecycle checks. The frozen packet is qualification-20260825T154151Z.json with SHA-256 7157a0153cf4f1dbfc59cfa7ae50f4bb8f17aa4429fec7d4b9b9ecfb0fd0eb84.
Windows linked-worktree qualification passed after repair The first detached release qualification exposed Git's administrative-path limit. A new regression test proves qualification now anchors pytest under the primary common checkout, and all five subsequent cycles passed.
Live fast-forward and restart passed The clean live checkout fast-forwarded from 4c15149 to a11954e. The checked hidden launcher restarted Hiro in 3.406 seconds without restarting Qwen.
Live health and qualification passed Ports 8000, 8001, and 8765 are owned by the new Hiro process. Health is OK, Qwen 3.8 27B is connected, the live revision matches the qualified revision, and the pipeline watchdog is not blocked.
Durable queue preservation passed The post-restart ledger retained 487 ideas, 35 actionable records, seven implementations, nine artifact-blocked records, and the existing scheduled infrastructure retry.
Independent whole-code inspection completed with critical findings A separate read-only inspector reviewed the exact activated revision across architecture, security boundaries, candidate construction, evaluation, promotion, runtime lifecycle, concurrency, persistence, observability, legacy code, and tests. It produced a severity-ordered remediation plan with five critical system-boundary findings and eleven high-priority findings.
Public journal timestamped-entry tests passed npm run test:hiro passed in the clean journal checkout.
Public journal production build passed after installing locked dependencies The initial clean-checkout build generated and validated all 160 journal pages but could not find the uninstalled local TypeScript compiler. After npm ci installed the package-lock dependencies, npm run build completed the generator, timestamp and alias validation, TypeScript build, and Vite production bundle successfully.

Current state

Next steps