Hiro development journal

Diagnosing the post-correction promotion bottleneck

Diagnostic session completed; no Hiro source change made Machine-readable JSON

Executive summary

The continuous queue is active, its circuit breaker is closed, and external discovery continues to refresh. The absence of a new promotion is not caused by a stopped scheduler.

The live queue reports 230 ideas, 137 actionable records, 73 retrying records, one active candidate, three implemented outcomes, 42 rejected outcomes, and 17 artifact-blocked outcomes.

The newest post-correction candidate demonstrated useful behavior by returning the correct arithmetic result, 17 × 24 = 408, but construction rejected it because the generated contract required at least 15 characters and the correct concise response contained 13.

That candidate never reached public and held-out evaluation. This is direct evidence of a remaining construction-contract and repair-feedback defect rather than evidence that the proposed improvement lacked merit.

A focused regression run passed 58 tests, showing that the queue controller and the recently corrected attribution machinery are internally consistent while leaving a semantic calibration gap in generated task contracts and candidate repairs.

Work completed

Live queue and scheduler inspection

Completed
  • Queried the live continuous-improvement API and confirmed one active candidate, a closed circuit breaker, ongoing retry activity, and continued external-source packet creation.
  • The active worker is using repository revision 908cb83, the revision containing the corrected continuous approval evidence changes.
  • The most recent scheduler messages continue to report completed ticks and three implemented outcomes, so the displayed lack of new promotions is not a dashboard-only consequence of a dead worker.

Outcome separation

Completed
  • Separated the three implemented records from 31 evaluation rejections, four disproven canaries, one construction failure, six legacy rejections, and 17 currently artifact-blocked records.
  • Two implemented records are user-authorized platform repairs. One record is a prior automatic Stage 5 promotion, so the dashboard's total of three implemented outcomes should not be interpreted as three new autonomous promotions.
  • Recent valid evaluator rejections include held-out epistemics regression and excessive latency. Those are materially different from construction-stage failures and remain legitimate negative evidence.

Post-correction candidate autopsy

Completed
  • Inspected the frozen packet for the newest arithmetic candidate. Its final repair changed the real router entrypoint and produced the correct deterministic result 17 × 24 = 408.
  • The candidate-authored regression test then failed only because len(result) was 13 while the generated task contract specified minimum_characters 15.
  • The candidate exhausted its construction repairs and was marked candidate_failed without receiving a baseline contrast, public evaluation, held-out evaluation, canary, or governor decision.
  • Earlier attempts in the same packet also exposed repair-quality problems: one test invoked the model path and received an HTTP 400, and another called an asynchronous function without awaiting it. The final attempt fixed those issues but did not adapt to the remaining length-only failure.

Recent artifact blockage pattern

Completed
  • Within the latest 100 queue events, five candidates exhausted repairs into artifact-blocked status and 25 more scheduled another candidate revision.
  • Three of those five terminal artifact blocks were stale or ambiguous edit anchors in core/router.py. Another combined an ambiguous anchor with invalid local-agent JSON, and one baseline run raised a non-contract TypeError.
  • The latest arithmetic case did apply its patch successfully, making the overly rigid test contract the immediate blocker rather than an edit-anchor failure.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Live benchmark API passed The continuous-improvement endpoint returned HTTP 200 and a current queue snapshot.
Queue liveness passed The circuit breaker is closed, one candidate is active, 73 records are retrying, and recent events show continuing investigation and construction attempts.
Focused pipeline regression suite passed 58 tests passed in 66.49 seconds across candidate builder, candidate evaluator, targeted contrast provenance, continuous engine, and continuous governor coverage.
Newest candidate evidence failed before evaluation The candidate returned the correct arithmetic value but failed its generated minimum-character assertion. No global evaluator or promotion decision was reached.
Hiro source mutation not performed This session was diagnostic. The Hiro working tree was not modified.

Current state

Next steps