Hiro development journal

Designing the missing autonomous improvement supervisor

Implementation design complete; no Hiro source change in this session Machine-readable JSON

Executive summary

The promotion drought should be fixed by adding a distinct upstream ImprovementSupervisor while leaving the independent ContinuousGovernor intact. The supervisor owns progress toward a valid candidate; the governor continues to own promotion authority.

The current engine performs bounded retries but does not retain an improvement episode. After a failed candidate it stores a truncated previous reason, schedules a retry, and returns the record to the globally ranked queue. This permits many fresh incidents to displace reflective work and prevents failure evidence from accumulating into a coherent strategy.

The repair needs four capabilities: sticky candidate ownership, a structured attempt ledger, cross-candidate failure clustering with builder-repair escalation, and automatic resume of the originating candidate after the builder repair qualifies.

The change must be proven with end-to-end autonomous qualification. A scripted candidate must fail its first construction attempt, improve from exact causal feedback, pass isolated validation, traverse canary, and reach the real governor without a human make-next action or manual state edit. A second scenario must demonstrate detection and repair of a shared builder defect.

Work completed

Add persistent improvement episodes

Designed
  • Create hiro/improvement/supervisor.py with an ImprovementSupervisor that maintains one active episode independently of ordinary queue priority.
  • Persist episode_id, originating idea, objective, current strategy, attempt number, progress evidence, failure fingerprint, builder version, status, and next allowed action in SQLite and the append-only event ledger.
  • When an episode is active, run its due reflective attempt before selecting a fresh queue record. Release the lock only after candidate_tests_passed, reproduced failure disproven, explicit non-progress, systemic escalation, or a bounded terminal decision.
  • Do not use make-next as the implementation. Human priority remains a separate user control; episode ownership must be an auditable autonomous state transition.

Replace generic retries with causal reflection packets

Designed
  • Record each attempt's exact failing assertion, baseline output, candidate output, changed-file diff summary, validator stage, test command, return code, and bounded stderr/stdout excerpt.
  • Before a revision attempt, create a compact code-owned reflection packet containing what changed, what improved, what remained wrong, and which constraint the next strategy must satisfy.
  • Require successive strategies to declare whether they are correcting an edit, changing the mechanism, narrowing scope, or expanding the permitted implementation surface. Identical retries without new evidence do not count as progress.
  • Qwen 3.8 27B receives the complete replay plus the structured reflection packet within the 16K context budget; raw external research text remains outside candidate authority.

Escalate systemic failures into builder repairs

Designed
  • Fingerprint construction failures by stage and causal class, including malformed edit, missing symbol preservation, incorrect contract invocation, encoding corruption, baseline-already-passes, and production-reachability mismatch.
  • When three independent ideas exhibit the same fingerprint, or twelve serious construction attempts yield no candidate_tests_passed event, pause fresh candidate churn and open one platform-repair episode.
  • Evaluate the builder repair on a frozen diagnostic set drawn from the triggering failures plus held-out construction cases. It must increase construction success without changing evaluator thresholds, governor policy, or expected task answers.
  • Advance the builder compatibility version only after the repair passes focused tests, the full suite, and five clean pipeline qualification cycles. Then resume the original candidate episode automatically.

Expose and enforce funnel health

Designed
  • Add supervisor state, construction-pass rate, paired-evaluation entry rate, canary entry rate, and governor decision rate to the benchmark page.
  • A watchdog opens platform diagnosis after six active hours or twelve serious construction attempts without candidate_tests_passed, whichever occurs first.
  • A candidate episode may use three ordinary reflection rounds and one mechanism revision while measurable progress continues. Repeating the same failure without progress escalates or terminates rather than looping indefinitely.
  • Promotion cadence is reported only from governor-completed implementations; investigation and model-activity counts are never presented as upgrades.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Current queue-selection inspection diagnostic finding Failed candidates return to the globally ranked queued state with a retry time and only a bounded previous-failure field; no persisted episode forces the reflective retry to retain ownership.
Current retry-context inspection diagnostic finding The next candidate receives a previous failure reason and revision counter, but not a normalized attempt ledger containing exact assertion, baseline/candidate contrast, diff summary, progress delta, and failure fingerprint.
Governor separation passed by design The proposed supervisor sits upstream and has no path to ContinuousGovernor.promote; only a canary-complete frozen candidate can enter the existing governor call path.
Repository tests not run This session produced an implementation design and made no Hiro source change.

Current state

Next steps