{
  "schemaVersion": 2,
  "date": "2026.08.22",
  "publishedAt": "2026-08-22T21:07:30-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "A realistic promotion cadence after the queue repair",
  "publicationStatus": "Live throughput diagnosis complete; no source change in this session",
  "executiveSummary": [
    "The stale-replay repair restored active, causally complete candidate construction, but it has not yet restored a healthy promotion funnel. In approximately three hours after restart, the queue began 36 investigations, confirmed 34 potentials, scheduled 30 candidate revisions, marked three artifacts blocked, and reproduced two cases as already fixed. No candidate reached paired evaluation or promotion.",
    "Counting distinct ideas in the same interval gives 30 investigated and 28 confirmed. Twenty-seven distinct ideas received at least one construction retry. Candidate validation failure appeared in 28 retry events; three also reported that targeted tests already passed on the untouched baseline.",
    "The honest current promotion expectation is therefore effectively zero until the construction-to-evaluation transition improves. The system is busy and Qwen is healthy, but activity at five-to-six-minute intervals is not equivalent to improvement throughput.",
    "For a functioning version of this architecture, a reasonable operating target is one small, attributable promotion per 12 to 24 hours of active runtime, with larger evaluator or harness upgrades occurring one to three times per week. A six-hour interval with no candidate reaching paired evaluation should be treated as a pipeline incident, not as normal idea selectivity."
  ],
  "workstreams": [
    {
      "title": "Measure the post-repair funnel",
      "status": "Completed",
      "details": [
        "The observation window began with event 3607 at 2026-08-23T01:07:50.325035+00:00 and ended during inspection near 2026-08-23T04:06:50+00:00.",
        "Event totals were 36 investigation_started, 34 potential_confirmed, 30 candidate_revision_scheduled, three artifact_generation_blocked, two idea_not_reproduced, and four new idea_queued events.",
        "No candidate_tests_passed, canary checkpoint, paired-evaluation, or implementation event occurred in that window.",
        "The queue remained live throughout the observation and immediately selected subsequent candidates after bounded failures."
      ]
    },
    {
      "title": "Separate scheduler health from promotion health",
      "status": "Completed",
      "details": [
        "Hiro's task service, benchmark service, and Qwen health endpoint each returned HTTP 200.",
        "The active candidate at inspection was incident-87abbb2d08f47155714c5a1f, a causally complete correction-recovery replay in candidate construction.",
        "The queue held five implemented, 71 queued, 45 rejected, 270 superseded, three artifact-blocked, and one active candidate record.",
        "The missing-prompt defect did not recur. The current dominant failure is that generated changes do not pass their harness-owned isolated replay, so they never earn the right to consume paired-evaluation capacity."
      ]
    },
    {
      "title": "Establish an operational expectation",
      "status": "Diagnostic recommendation",
      "details": [
        "Promotion frequency cannot be guaranteed because correct rejection is necessary, but the system can be held to funnel-health expectations.",
        "A healthy near-term target is at least one candidate reaching paired evaluation within six hours or approximately every 8 to 12 serious construction attempts.",
        "Given a backlog of reproducible failures, a reasonable result target is one small promotion per 12 to 24 active hours. Larger research, evaluator, or harness changes should be expected less often, approximately one to three per week.",
        "A 24-to-72-hour promotion dry spell can be legitimate only when candidates are reaching evaluation and losing for measured reasons. A dry spell where zero candidates reach evaluation is a technical throughput failure."
      ]
    }
  ],
  "decisions": [
    "Do not use queue activity, model utilization, investigation counts, or construction attempts as a proxy for successful improvement throughput.",
    "Treat paired-evaluation entry as the minimum useful funnel-health signal. Promotion selectivity is credible only after candidates consistently reach that stage.",
    "Use separate expectations for bounded corrective patches and larger research or harness upgrades; they have materially different construction and attribution difficulty.",
    "Escalate automatically when six active hours or 12 serious construction attempts produce zero candidates that pass isolated validation.",
    "Do not lower evaluator standards merely to manufacture a promotion. The immediate target is to make candidate construction satisfy existing reproducible contracts often enough to test real baseline contrasts."
  ],
  "validation": [
    {
      "check": "Post-repair event audit",
      "status": "diagnostic finding",
      "result": "Thirty distinct ideas were investigated and 28 confirmed; 27 received construction retries, while zero reached paired evaluation or promotion."
    },
    {
      "check": "Retry-reason classification",
      "status": "diagnostic finding",
      "result": "All 30 candidate_revision_scheduled events reported candidate validation failure, baseline-already-passes, or both. Candidate validation failure appeared in 28 events and baseline-already-passes appeared in three."
    },
    {
      "check": "Live service health",
      "status": "passed",
      "result": "The task endpoint, benchmark endpoint, and Qwen health endpoint each returned HTTP 200."
    },
    {
      "check": "Replay completeness",
      "status": "passed for active candidate",
      "result": "The active audit candidate retained both its full prompt and query; the historical missing-evidence defect was not the cause of the current construction failures."
    },
    {
      "check": "Repository tests",
      "status": "not run",
      "result": "This session performed a read-only live diagnosis and made no Hiro source change. The previously qualified revision remained active."
    }
  ],
  "currentState": [
    "The live queue is active and Qwen 3.8 27B is healthy, but the observed promotion rate since the repair is zero.",
    "Five records remain implemented; no sixth implementation has occurred.",
    "The bottleneck is now candidate construction quality and isolated validation, not queue starvation, stale replay evidence, model availability, paired-evaluator variance, or promotion safety policy.",
    "The six scheduled-research mechanisms remain queued behind higher-priority reproduced interaction failures."
  ],
  "limitations": [
    "The post-repair observation window is approximately three hours, so it is enough to identify a zero-throughput funnel stage but not enough to estimate a stable long-run promotion probability.",
    "The proposed 12-to-24-hour cadence is an operating target for a healthy funnel, not a guarantee that every day contains a genuinely beneficial patch.",
    "Historical promotions occurred under several different pipeline revisions and periods of direct supervision, so averaging them into a single long-run autonomous rate would be misleading."
  ],
  "nextSteps": [
    "Instrument and repair the construction-to-isolated-validation boundary so the candidate builder receives the exact failing assertion, actual baseline behavior, allowed edit surface, and a concise patch strategy before generating code.",
    "Add a funnel watchdog that opens a platform diagnosis after six active hours or 12 serious attempts without one candidate_tests_passed event.",
    "Track construction pass rate, paired-evaluation entry rate, evaluation win rate, and promotion rate separately on the benchmark page.",
    "Re-estimate promotion cadence after at least 30 candidates have passed isolated validation; until then, do not report model activity as evidence of expected promotions."
  ]
}
