Hiro development journal

Why two supervised promotions did not become autonomous throughput

Supervision-boundary diagnosis complete; no Hiro source change in this session Machine-readable JSON

Executive summary

The two recent promotions were valid code improvements, but they were not evidence that Hiro had achieved a self-sustaining autonomous improvement loop. Both emerged from long-lived candidates that were repeatedly made-next, rebuilt after platform changes, and carried through invalid harness or governor outcomes during direct supervision.

The arithmetic candidate had been in the queue since August 17. It was explicitly made-next three times, crossed multiple builder and harness revisions, passed candidate tests several times, and was finally promoted on August 22 after a governor full-suite failure was diagnosed as infrastructure-related and retried.

The prompt-injection candidate was also made-next three times. Its successful path required a sequence of platform corrections for editing, missing task context, a double-applied response boundary, production reachability, an invalid regression expectation, and a governor manifest mismatch. Twelve repository commits occurred during the concentrated supervision interval before its final promotion.

After supervision stopped, Hiro reverted to breadth-first queue churn. In the current post-repair window, 28 ideas received one investigation each and only three ideas reached three investigations. None reached candidate_tests_passed. The missing component is therefore an encoded supervisory control loop that holds focus, clusters systemic failures, repairs the builder or harness, and resumes the same candidate with accumulated causal feedback.

Work completed

Reconstruct the arithmetic promotion

Completed
  • Incident d2074a2645707df25a06f74e entered the queue on August 17 and received three explicit make-next events.
  • Before promotion it encountered malformed patch plans, dirty-repository infrastructure blocks, candidate validation failures, an evaluation regression, a failed canary, several builder-change requeues, and a governor full-suite rejection.
  • Candidate tests passed at events 2636, 2940, 3011, and 3057 across successive revisions. The final candidate completed canary checkpoints and was promoted at event 3085 as candidate_implemented_after_infrastructure_retry.
  • This history demonstrates genuine iterative improvement, but the iteration was sustained by external prioritization and repeated platform interventions rather than autonomous task persistence.

Reconstruct the prompt-injection promotion

Completed
  • Incident fe44ffbe49335e3aa0d5f30d originated on August 12 and was made-next three times during the successful recovery period.
  • Two attempts were explicitly voided after editor defects, and another was voided because candidate construction lacked the complete audit prompt and task-specific guidance.
  • Later attempts exposed a double-applied response boundary, evaluator-only reachability, a platform regression test that required defective behavior, and a governor manifest mismatch. Each was diagnosed and corrected before the same improvement thread resumed.
  • The final candidate changed core/response_envelope.py and its autonomous contract test, passed the full suite with 715 tests, and was promoted at event 3285 after the manifest defect was voided.

Compare autonomous behavior after supervision

Completed
  • Since the replay migration began at event 3607, 28 distinct ideas received exactly one investigation and three ideas received three investigations.
  • The three repeatedly attempted ideas exhausted bounded construction and became artifact-blocked; the remaining scheduler capacity moved across new high-priority interaction incidents.
  • No post-repair idea reached candidate_tests_passed, canary, governor, or implementation.
  • The system retained complete replay evidence and healthy Qwen service, so the discontinuity is task persistence and supervisory adaptation rather than model availability or the earlier stale-prompt defect.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Successful-candidate event reconstruction passed Both promotion histories were reconstructed from append-only queue events, including make-next actions, candidate revisions, canary checkpoints, governor outcomes, platform voids, and final implementation events.
Post-supervision attempt distribution diagnostic finding Twenty-eight ideas received one investigation and three ideas received three; none reached candidate_tests_passed.
Supervision-era repository activity diagnostic finding Twelve repository commits occurred during the concentrated August 22 supervision interval, including causal candidate guidance, full audit context, end-to-end contract validation, production reachability, Qwen context, and governor manifest corrections.
Repository tests not run This session performed a read-only historical and live diagnosis and made no Hiro source change.

Current state

Next steps