{
  "schemaVersion": 2,
  "date": "2026.08.22",
  "publishedAt": "2026-08-22T21:33:05-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Designing the missing autonomous improvement supervisor",
  "publicationStatus": "Implementation design complete; no Hiro source change in this session",
  "executiveSummary": [
    "The promotion drought should be fixed by adding a distinct upstream ImprovementSupervisor while leaving the independent ContinuousGovernor intact. The supervisor owns progress toward a valid candidate; the governor continues to own promotion authority.",
    "The current engine performs bounded retries but does not retain an improvement episode. After a failed candidate it stores a truncated previous reason, schedules a retry, and returns the record to the globally ranked queue. This permits many fresh incidents to displace reflective work and prevents failure evidence from accumulating into a coherent strategy.",
    "The repair needs four capabilities: sticky candidate ownership, a structured attempt ledger, cross-candidate failure clustering with builder-repair escalation, and automatic resume of the originating candidate after the builder repair qualifies.",
    "The change must be proven with end-to-end autonomous qualification. A scripted candidate must fail its first construction attempt, improve from exact causal feedback, pass isolated validation, traverse canary, and reach the real governor without a human make-next action or manual state edit. A second scenario must demonstrate detection and repair of a shared builder defect."
  ],
  "workstreams": [
    {
      "title": "Add persistent improvement episodes",
      "status": "Designed",
      "details": [
        "Create hiro/improvement/supervisor.py with an ImprovementSupervisor that maintains one active episode independently of ordinary queue priority.",
        "Persist episode_id, originating idea, objective, current strategy, attempt number, progress evidence, failure fingerprint, builder version, status, and next allowed action in SQLite and the append-only event ledger.",
        "When an episode is active, run its due reflective attempt before selecting a fresh queue record. Release the lock only after candidate_tests_passed, reproduced failure disproven, explicit non-progress, systemic escalation, or a bounded terminal decision.",
        "Do not use make-next as the implementation. Human priority remains a separate user control; episode ownership must be an auditable autonomous state transition."
      ]
    },
    {
      "title": "Replace generic retries with causal reflection packets",
      "status": "Designed",
      "details": [
        "Record each attempt's exact failing assertion, baseline output, candidate output, changed-file diff summary, validator stage, test command, return code, and bounded stderr/stdout excerpt.",
        "Before a revision attempt, create a compact code-owned reflection packet containing what changed, what improved, what remained wrong, and which constraint the next strategy must satisfy.",
        "Require successive strategies to declare whether they are correcting an edit, changing the mechanism, narrowing scope, or expanding the permitted implementation surface. Identical retries without new evidence do not count as progress.",
        "Qwen 3.8 27B receives the complete replay plus the structured reflection packet within the 16K context budget; raw external research text remains outside candidate authority."
      ]
    },
    {
      "title": "Escalate systemic failures into builder repairs",
      "status": "Designed",
      "details": [
        "Fingerprint construction failures by stage and causal class, including malformed edit, missing symbol preservation, incorrect contract invocation, encoding corruption, baseline-already-passes, and production-reachability mismatch.",
        "When three independent ideas exhibit the same fingerprint, or twelve serious construction attempts yield no candidate_tests_passed event, pause fresh candidate churn and open one platform-repair episode.",
        "Evaluate the builder repair on a frozen diagnostic set drawn from the triggering failures plus held-out construction cases. It must increase construction success without changing evaluator thresholds, governor policy, or expected task answers.",
        "Advance the builder compatibility version only after the repair passes focused tests, the full suite, and five clean pipeline qualification cycles. Then resume the original candidate episode automatically."
      ]
    },
    {
      "title": "Expose and enforce funnel health",
      "status": "Designed",
      "details": [
        "Add supervisor state, construction-pass rate, paired-evaluation entry rate, canary entry rate, and governor decision rate to the benchmark page.",
        "A watchdog opens platform diagnosis after six active hours or twelve serious construction attempts without candidate_tests_passed, whichever occurs first.",
        "A candidate episode may use three ordinary reflection rounds and one mechanism revision while measurable progress continues. Repeating the same failure without progress escalates or terminates rather than looping indefinitely.",
        "Promotion cadence is reported only from governor-completed implementations; investigation and model-activity counts are never presented as upgrades."
      ]
    }
  ],
  "decisions": [
    "Keep ContinuousGovernor protected and unchanged as the final promotion authority. The supervisor cannot alter its policy, manufacture canary evidence, or promote directly.",
    "Implement autonomous episode ownership as first-class persisted state, not as repeated queue boosts or human-style make-next events.",
    "Classify construction failures structurally and aggregate them across ideas so common platform defects are repaired once rather than charged repeatedly to candidate merit.",
    "Preserve boundedness through progress-aware reflection limits, explicit escalation, independent qualification, and automatic return to the original task.",
    "Do not lower isolated validation, evaluator, canary, or governor standards to increase promotion count."
  ],
  "validation": [
    {
      "check": "Current queue-selection inspection",
      "status": "diagnostic finding",
      "result": "Failed candidates return to the globally ranked queued state with a retry time and only a bounded previous-failure field; no persisted episode forces the reflective retry to retain ownership."
    },
    {
      "check": "Current retry-context inspection",
      "status": "diagnostic finding",
      "result": "The next candidate receives a previous failure reason and revision counter, but not a normalized attempt ledger containing exact assertion, baseline/candidate contrast, diff summary, progress delta, and failure fingerprint."
    },
    {
      "check": "Governor separation",
      "status": "passed by design",
      "result": "The proposed supervisor sits upstream and has no path to ContinuousGovernor.promote; only a canary-complete frozen candidate can enter the existing governor call path."
    },
    {
      "check": "Repository tests",
      "status": "not run",
      "result": "This session produced an implementation design and made no Hiro source change."
    }
  ],
  "currentState": [
    "Hiro has queue ranking, bounded candidate retries, isolated validation, canary checks, and a final governor, but no component owns an improvement objective end to end.",
    "The retry mechanism preserves too little structured causal history and permits fresh high-priority records to dominate reflective work.",
    "Repeated failures are not automatically aggregated into a builder-repair hypothesis, leaving direct supervision as the only current mechanism for platform adaptation.",
    "The design is ready for implementation, but no claim is made that the supervisor exists or that autonomous throughput has been restored."
  ],
  "limitations": [
    "Sticky ownership can become unproductive fixation unless progress is measured and reflection rounds remain bounded.",
    "Failure fingerprints must be deterministic and code-owned; asking the candidate model alone to label its own failure would make escalation unreliable.",
    "A builder repair that succeeds on only the triggering examples may overfit. Qualification requires frozen held-out construction cases and the full existing pipeline suite.",
    "The supervisor can improve the probability that valid candidates reach evaluation, but it cannot guarantee that genuinely beneficial patches exist at a fixed daily rate."
  ],
  "nextSteps": [
    "Implement ImprovementSupervisor, episode persistence, attempt-ledger events, deterministic failure fingerprints, and episode-first queue selection.",
    "Add systemic-failure thresholds and a separately qualified builder-repair episode that resumes the original candidate after activation.",
    "Add benchmark-page supervisor and funnel metrics plus explicit unhealthy-funnel status.",
    "Run focused unit tests, the full repository suite, five exact-revision qualification cycles, and two end-to-end autonomous scenarios: reflective candidate recovery and shared builder-defect recovery.",
    "Restart Hiro only after qualification, then require one unsupervised candidate to reach the real governor before declaring the repair operational."
  ]
}
