Hiro development journal

A realistic promotion cadence after the queue repair

Live throughput diagnosis complete; no source change in this session Machine-readable JSON

Executive summary

The stale-replay repair restored active, causally complete candidate construction, but it has not yet restored a healthy promotion funnel. In approximately three hours after restart, the queue began 36 investigations, confirmed 34 potentials, scheduled 30 candidate revisions, marked three artifacts blocked, and reproduced two cases as already fixed. No candidate reached paired evaluation or promotion.

Counting distinct ideas in the same interval gives 30 investigated and 28 confirmed. Twenty-seven distinct ideas received at least one construction retry. Candidate validation failure appeared in 28 retry events; three also reported that targeted tests already passed on the untouched baseline.

The honest current promotion expectation is therefore effectively zero until the construction-to-evaluation transition improves. The system is busy and Qwen is healthy, but activity at five-to-six-minute intervals is not equivalent to improvement throughput.

For a functioning version of this architecture, a reasonable operating target is one small, attributable promotion per 12 to 24 hours of active runtime, with larger evaluator or harness upgrades occurring one to three times per week. A six-hour interval with no candidate reaching paired evaluation should be treated as a pipeline incident, not as normal idea selectivity.

Work completed

Measure the post-repair funnel

Completed
  • The observation window began with event 3607 at 2026-08-23T01:07:50.325035+00:00 and ended during inspection near 2026-08-23T04:06:50+00:00.
  • Event totals were 36 investigation_started, 34 potential_confirmed, 30 candidate_revision_scheduled, three artifact_generation_blocked, two idea_not_reproduced, and four new idea_queued events.
  • No candidate_tests_passed, canary checkpoint, paired-evaluation, or implementation event occurred in that window.
  • The queue remained live throughout the observation and immediately selected subsequent candidates after bounded failures.

Separate scheduler health from promotion health

Completed
  • Hiro's task service, benchmark service, and Qwen health endpoint each returned HTTP 200.
  • The active candidate at inspection was incident-87abbb2d08f47155714c5a1f, a causally complete correction-recovery replay in candidate construction.
  • The queue held five implemented, 71 queued, 45 rejected, 270 superseded, three artifact-blocked, and one active candidate record.
  • The missing-prompt defect did not recur. The current dominant failure is that generated changes do not pass their harness-owned isolated replay, so they never earn the right to consume paired-evaluation capacity.

Establish an operational expectation

Diagnostic recommendation
  • Promotion frequency cannot be guaranteed because correct rejection is necessary, but the system can be held to funnel-health expectations.
  • A healthy near-term target is at least one candidate reaching paired evaluation within six hours or approximately every 8 to 12 serious construction attempts.
  • Given a backlog of reproducible failures, a reasonable result target is one small promotion per 12 to 24 active hours. Larger research, evaluator, or harness changes should be expected less often, approximately one to three per week.
  • A 24-to-72-hour promotion dry spell can be legitimate only when candidates are reaching evaluation and losing for measured reasons. A dry spell where zero candidates reach evaluation is a technical throughput failure.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Post-repair event audit diagnostic finding Thirty distinct ideas were investigated and 28 confirmed; 27 received construction retries, while zero reached paired evaluation or promotion.
Retry-reason classification diagnostic finding All 30 candidate_revision_scheduled events reported candidate validation failure, baseline-already-passes, or both. Candidate validation failure appeared in 28 events and baseline-already-passes appeared in three.
Live service health passed The task endpoint, benchmark endpoint, and Qwen health endpoint each returned HTTP 200.
Replay completeness passed for active candidate The active audit candidate retained both its full prompt and query; the historical missing-evidence defect was not the cause of the current construction failures.
Repository tests not run This session performed a read-only live diagnosis and made no Hiro source change. The previously qualified revision remained active.

Current state

Next steps