{
  "schemaVersion": 2,
  "date": "2026.08.22",
  "publishedAt": "2026-08-22T21:12:20-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Why two supervised promotions did not become autonomous throughput",
  "publicationStatus": "Supervision-boundary diagnosis complete; no Hiro source change in this session",
  "executiveSummary": [
    "The two recent promotions were valid code improvements, but they were not evidence that Hiro had achieved a self-sustaining autonomous improvement loop. Both emerged from long-lived candidates that were repeatedly made-next, rebuilt after platform changes, and carried through invalid harness or governor outcomes during direct supervision.",
    "The arithmetic candidate had been in the queue since August 17. It was explicitly made-next three times, crossed multiple builder and harness revisions, passed candidate tests several times, and was finally promoted on August 22 after a governor full-suite failure was diagnosed as infrastructure-related and retried.",
    "The prompt-injection candidate was also made-next three times. Its successful path required a sequence of platform corrections for editing, missing task context, a double-applied response boundary, production reachability, an invalid regression expectation, and a governor manifest mismatch. Twelve repository commits occurred during the concentrated supervision interval before its final promotion.",
    "After supervision stopped, Hiro reverted to breadth-first queue churn. In the current post-repair window, 28 ideas received one investigation each and only three ideas reached three investigations. None reached candidate_tests_passed. The missing component is therefore an encoded supervisory control loop that holds focus, clusters systemic failures, repairs the builder or harness, and resumes the same candidate with accumulated causal feedback."
  ],
  "workstreams": [
    {
      "title": "Reconstruct the arithmetic promotion",
      "status": "Completed",
      "details": [
        "Incident d2074a2645707df25a06f74e entered the queue on August 17 and received three explicit make-next events.",
        "Before promotion it encountered malformed patch plans, dirty-repository infrastructure blocks, candidate validation failures, an evaluation regression, a failed canary, several builder-change requeues, and a governor full-suite rejection.",
        "Candidate tests passed at events 2636, 2940, 3011, and 3057 across successive revisions. The final candidate completed canary checkpoints and was promoted at event 3085 as candidate_implemented_after_infrastructure_retry.",
        "This history demonstrates genuine iterative improvement, but the iteration was sustained by external prioritization and repeated platform interventions rather than autonomous task persistence."
      ]
    },
    {
      "title": "Reconstruct the prompt-injection promotion",
      "status": "Completed",
      "details": [
        "Incident fe44ffbe49335e3aa0d5f30d originated on August 12 and was made-next three times during the successful recovery period.",
        "Two attempts were explicitly voided after editor defects, and another was voided because candidate construction lacked the complete audit prompt and task-specific guidance.",
        "Later attempts exposed a double-applied response boundary, evaluator-only reachability, a platform regression test that required defective behavior, and a governor manifest mismatch. Each was diagnosed and corrected before the same improvement thread resumed.",
        "The final candidate changed core/response_envelope.py and its autonomous contract test, passed the full suite with 715 tests, and was promoted at event 3285 after the manifest defect was voided."
      ]
    },
    {
      "title": "Compare autonomous behavior after supervision",
      "status": "Completed",
      "details": [
        "Since the replay migration began at event 3607, 28 distinct ideas received exactly one investigation and three ideas received three investigations.",
        "The three repeatedly attempted ideas exhausted bounded construction and became artifact-blocked; the remaining scheduler capacity moved across new high-priority interaction incidents.",
        "No post-repair idea reached candidate_tests_passed, canary, governor, or implementation.",
        "The system retained complete replay evidence and healthy Qwen service, so the discontinuity is task persistence and supervisory adaptation rather than model availability or the earlier stale-prompt defect."
      ]
    }
  ],
  "decisions": [
    "Classify the two promotions as supervised autonomous-candidate successes: Qwen generated the retained patches, but human-supervised platform diagnosis and queue focus were material dependencies of the successful trajectories.",
    "Do not cite those two outcomes as proof of steady-state autonomous promotion throughput.",
    "Treat repeated construction failures across different incidents as a shared builder-level signal. The current queue incorrectly handles them primarily as unrelated candidate-level failures.",
    "The next structural requirement is a persistent supervisor loop, not another isolated rejector adjustment. It must own a candidate thread across attempts, accumulate exact failure deltas, distinguish candidate defects from platform defects, and escalate systemic clusters into builder repairs.",
    "Promotion standards should remain intact. The objective is to recreate the productive parts of supervision inside the system, not to bypass validation."
  ],
  "validation": [
    {
      "check": "Successful-candidate event reconstruction",
      "status": "passed",
      "result": "Both promotion histories were reconstructed from append-only queue events, including make-next actions, candidate revisions, canary checkpoints, governor outcomes, platform voids, and final implementation events."
    },
    {
      "check": "Post-supervision attempt distribution",
      "status": "diagnostic finding",
      "result": "Twenty-eight ideas received one investigation and three ideas received three; none reached candidate_tests_passed."
    },
    {
      "check": "Supervision-era repository activity",
      "status": "diagnostic finding",
      "result": "Twelve repository commits occurred during the concentrated August 22 supervision interval, including causal candidate guidance, full audit context, end-to-end contract validation, production reachability, Qwen context, and governor manifest corrections."
    },
    {
      "check": "Repository tests",
      "status": "not run",
      "result": "This session performed a read-only historical and live diagnosis and made no Hiro source change."
    }
  ],
  "currentState": [
    "The two retained promotions are real improvements, but their completion depended on active supervisory diagnosis and repeated re-entry of the same candidate thread.",
    "The autonomous scheduler is currently optimized for continuously selecting work, not for ensuring that one difficult improvement thread receives enough reflective iterations to cross validation.",
    "Systemic construction failures do not yet trigger an autonomous platform-repair episode; they are consumed as bounded candidate failures and the scheduler moves on.",
    "This explains the apparent cliff when supervision ended: the missing supervisory function was never fully encoded into Hiro."
  ],
  "limitations": [
    "The queue event log establishes what interventions occurred and when, but it cannot measure the counterfactual probability that either candidate would eventually have succeeded without them.",
    "Some intervention commits improved the platform generally and remain valuable; classifying the trajectory as supervision-dependent does not mean the resulting promotions were invalid.",
    "A persistent supervisor loop must avoid infinite fixation on one candidate. It needs explicit progress evidence, bounded reflection rounds, and a criterion for escalating from candidate repair to builder repair or abandoning the thread."
  ],
  "nextSteps": [
    "Add a sticky improvement episode that retains one reproduced problem across several construction and reflection rounds while progress is measurable.",
    "Feed each failed assertion, observed candidate output, diff summary, and unchanged baseline result into the next revision rather than restarting from a generic task plan.",
    "Cluster repeated validation failures across ideas. When a common editor, prompt, test, or reachability pattern appears, pause candidate churn and open a qualified builder-repair task automatically.",
    "Resume the originating candidate after the builder repair and require it to pass isolated validation before declaring the supervisory episode successful.",
    "Expose autonomous versus supervised contribution in promotion lineage so future throughput claims cannot conflate model-generated patches with externally sustained recovery."
  ]
}
