{
  "schemaVersion": 2,
  "date": "2026.08.22",
  "publishedAt": "2026-08-22T22:50:47-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Building and independently exercising Hiro's improvement supervisor",
  "publicationStatus": "Implementation complete; frozen qualification and full unattended episode validation passed",
  "executiveSummary": [
    "Implemented the missing upstream ImprovementSupervisor while preserving ContinuousGovernor as the independent final promotion authority. The supervisor now owns one durable improvement episode across candidate failures, scheduled retries, service restarts, canary, and terminal disposition.",
    "Replaced generic retry text with an append-only structured attempt ledger and deterministic causal reflection packets. Failed construction now records a failure fingerprint, stage, bounded exact evidence, strategy, progress signal, builder version, and next action for the same originating idea.",
    "Added cross-idea builder-failure escalation. Three independent ideas with the same version-scoped failure fingerprint, or twelve failures since the last construction pass, open a separately evaluated builder-repair candidate and automatically resume the originating episode after promotion.",
    "Expanded the benchmark API and page with supervisor status, active episode, attempt count, next action, watchdog health, and attempts since the last candidate construction pass.",
    "The full repository suite passed 722 tests. The frozen end-to-end gate was repeatedly rerun after each live-only correction; the current certified revision is required to pass five consecutive cycles covering construction, evaluation, isolated integration, rollback, interruption recovery, real governor promotion, episode restart persistence, and builder escalation.",
    "A normal unattended launch recovered the same persisted episode after service restarts, ran all four bounded Qwen 3.8 strategies, recorded exact causal failures, enforced a verified five-minute post-completion cooldown, reached a truthful artifact-blocked disposition, and selected the next ranked idea without a manual tick or make-next action.",
    "The extended run exposed and led to repair of three live-only integration defects: retry cooldowns anchored to build start instead of completion, a mechanism-revision off-by-one that made the final strategy unreachable, and an unrealistically short 15-second detached-server readiness window. Each repair was committed, retested, requalified, and exercised through the normal launch path."
  ],
  "workstreams": [
    {
      "title": "Durable episode ownership",
      "status": "Implemented and exercised",
      "details": [
        "Added hiro/improvement/supervisor.py with SQLite-backed supervisor_episodes and append-only supervisor_attempts tables stored beside the ranked queue.",
        "An active episode takes precedence over fresh global queue selections. A competing candidate cannot displace the originating idea while its retry is waiting or its canary is active.",
        "Episode state includes root and active idea IDs, mode, objective, strategy, attempt count, mechanism revisions, builder repairs, current failure fingerprint, builder compatibility version, next action, and lifecycle timestamps.",
        "Normal service restarts reconstruct supervision from SQLite. The unattended live run recovered episode-590a4e85ae1645ce98d23c48 with its prior attempts and strategy intact."
      ]
    },
    {
      "title": "Causal reflection and retry progression",
      "status": "Implemented and exercised",
      "details": [
        "Each completed construction failure is normalized into a deterministic fingerprint and a bounded reflection packet containing the exact failure, stage, category, candidate identity, changed files, baseline/candidate excerpts when available, progress observation, and next strategy.",
        "The candidate builder receives that code-owned packet on the next complete-candidate attempt. Retry progression uses correct_edit after failure one, change_implementation_strategy after failure two, and mechanism_revision after failure three so the fourth and final candidate actually exercises the mechanism revision.",
        "The live episode recorded an assertion_mismatch after a candidate transformed the answer but omitted the required p.m. punctuation. Hiro scheduled its own retry, restored the same episode and failure evidence, and advanced revision attempt two independently.",
        "The live policy now matches the supervisor's four-attempt progression. Upgrade reconciliation corrected the strategy of the already-persisted three-attempt episode without rewriting its append-only evidence."
      ]
    },
    {
      "title": "Systemic builder repair",
      "status": "Implemented and qualified in controlled scenarios",
      "details": [
        "Failure clustering is scoped to the active candidate-builder compatibility version so historical failures cannot immediately retrigger after a successful builder repair.",
        "Three distinct ideas with the same fingerprint or twelve failures since a construction pass create a lane-one builder_repair idea without changing evaluator thresholds or governor policy.",
        "A promoted builder repair advances the episode's builder version, increments its repair count, requeues the root idea, and resumes the original objective automatically.",
        "Terminal repair candidates receive a new deterministic repair generation rather than leaving a future episode pointed at dead work."
      ]
    },
    {
      "title": "Live launch reproducibility",
      "status": "Corrected and verified",
      "details": [
        "Aligned docs/hiro_active_improvement_policy.v1.json with the four-attempt supervisor budget.",
        "Corrected the canonical Qwen 3.8 launcher from 8K to 16,384 context and updated its launcher test. The helper normalized the Windows Path/PATH environment, loaded the pinned 15.66 GiB Qwen 3.8 27B model with maximum GPU offload, and required a real READY completion from the expected model identity.",
        "The resulting metadata recorded context_size 16384, parallel_slots 1, gpu_offload max, and a successful warm-up. Hiro was then started only through scripts/start_hiro.cmd.",
        "The extended run found that long asynchronous failures used the scheduler-entry timestamp. The engine now records receipts and begins retry cooldowns at actual build completion; live revision three completed at 05:53:32.688871Z and received an exact next-attempt deadline of 05:58:32.688871Z.",
        "A later normal restart showed all Uvicorn applications completing but the 8001 readiness endpoint missing the old 15-second cutoff. The canonical detached launcher now allows 45 seconds, retains real health-page verification, and subsequently started successfully with its default settings in 5.594 seconds."
      ]
    },
    {
      "title": "Benchmark observability",
      "status": "Implemented",
      "details": [
        "The continuous-improvement API now returns supervisor status and sanitizes active attempt details before public display.",
        "The benchmark page shows ACTIVE or IDLE supervisor state, attempts since candidate-ready, watchdog threshold and health, active mode, attempt count, and next autonomous action.",
        "Dashboard language distinguishes the upstream supervisor from the independent governor and reports one supervised build, one live canary, and one governor fast-forward at a time."
      ]
    }
  ],
  "decisions": [
    "Keep the governor independent. The supervisor may own persistence and candidate progress but cannot promote, edit governor policy, manufacture canary evidence, or bypass full-suite checks.",
    "Make episode ownership durable state rather than a queue boost or human-style make-next flag.",
    "Treat construction failures as repair evidence, not negative evidence about idea merit, unless a bounded terminal classification is actually reached.",
    "Use exact code-owned causal feedback while keeping the retry budget finite: three correction/strategy rounds and one mechanism revision.",
    "Escalate shared platform failures into their own evaluated code candidate and return automatically to the original task after a successful repair.",
    "Anchor post-build evidence and cooldowns to result completion time because Qwen 3.8 construction can last several minutes.",
    "Bind live queue authorization to a clean committed revision and invalidate qualification after every source or operating-policy change."
  ],
  "validation": [
    {
      "check": "Focused supervisor, engine, scheduler, dashboard, and remote-access tests",
      "status": "passed",
      "result": "Successive focused suites passed after each repair, including 66 tests across supervisor/engine/dashboard paths and 59 tests after the live completion-time correction. Python compilation and git diff checks also passed."
    },
    {
      "check": "Full repository regression suite",
      "status": "passed",
      "result": "722 tests passed in 384.63 seconds with zero failures. Two existing unknown-marker warnings were reported for generated autonomous hiro_contract tests."
    },
    {
      "check": "Frozen five-cycle qualification",
      "status": "passed",
      "result": "Every certification run completed five consecutive cycles with eight end-to-end tests per cycle. The gate covers frozen candidate construction, report-only integration, rollback, interrupted integration resume, real ContinuousGovernor promotion, retry ownership after supervisor recreation, append-only episode persistence, and shared-failure escalation."
    },
    {
      "check": "Qwen 3.8 unattended launch",
      "status": "passed",
      "result": "The canonical helper loaded qwen/qwen3.8-27b at 16,384 context, one parallel slot, maximum GPU offload, and returned READY from a real warm-up completion before Hiro started."
    },
    {
      "check": "Unattended live failure and retry",
      "status": "passed",
      "result": "The scheduler independently selected a persisted episode, ran four distinct bounded attempts, preserved and reconciled strategy across several restarts, completed the episode as artifact-blocked after construction evidence remained negative, and immediately selected and investigated the next ranked idea without a chat-triggered tick, manual priority action, or database edit."
    },
    {
      "check": "Live completion-time audit",
      "status": "failed, repaired, then passed in regression",
      "result": "The first unattended run revealed that failure receipts and cooldowns inherited scheduler-entry time. The engine now samples the clock after asynchronous construction; the regression proves the attempt receipt and five-minute deadline begin at result availability."
    },
    {
      "check": "Final-attempt strategy audit",
      "status": "failed, repaired, then passed live",
      "result": "The live ledger showed mechanism_revision would have been selected only after attempt four. Strategy semantics were corrected and existing-state reconciliation added; the recovered episode then displayed mechanism_revision before Qwen constructed the fourth candidate."
    },
    {
      "check": "Detached Hiro launcher",
      "status": "failed once, repaired, then passed live",
      "result": "The 15-second readiness limit terminated an otherwise-starting process. A 45-second default passed pinned-runtime and Path-normalization tests, then the unmodified scripts/start_hiro.cmd path returned ready in 5.594 seconds with all expected listeners."
    }
  ],
  "currentState": [
    "The ranked queue, supervisor, isolated candidate builder, evaluator, canary, and independent governor now form one persisted lifecycle rather than relying on a supervising chat to reconnect failed attempts.",
    "The live queue contains historical work from earlier builder versions. Compatibility requeue events rebuild eligible artifacts against the certified source revision while a single episode retains mutation capacity.",
    "The first post-repair unattended episode completed all four attempts and was correctly marked artifact-blocked because Qwen never produced a candidate that cleared construction. The supervisor then opened a new transit-directions episode and advanced it to candidate state independently.",
    "Qwen 3.8 27B and Hiro are launched from checked-in helpers; the model metadata and qualification packet make the live dependency and authorized source revision observable."
  ],
  "limitations": [
    "A functioning supervisor guarantees durable, bounded progress and diagnosable failure—not a positive candidate or promotion for every idea. The evaluator and governor still reject candidates that do not demonstrate improvement.",
    "Real 27B code construction is slow. A single plan plus repeated validation can take several minutes, so promotion cadence must be evaluated over completed episodes rather than scheduler ticks.",
    "Builder-repair escalation is controlled by deterministic failure categories. New failure modes may require improved classification before different textual errors cluster correctly.",
    "The extended live exercise demonstrated independent selection, construction, causal rejection, cooldown, restart persistence, strategy reconciliation, bounded terminal disposition, and selection of the next task. A new live promotion is not claimed; promotion capability is established by the repeated real-governor qualification control, while the exercised live proposal never cleared construction.",
    "The public journal deliberately excludes secrets, tokens, private personal data, and actionable unresolved security details."
  ],
  "nextSteps": [
    "Allow the newly selected transit-directions episode to continue independently and observe whether it reaches candidate-ready, canary, and a governor decision.",
    "Use benchmark-page supervisor health and attempts-since-pass metrics to detect construction funnel regressions without requiring manual log inspection.",
    "When a shared fingerprint meets the threshold in live work, audit the automatically created builder-repair episode and verify automatic return to its originating candidate.",
    "Track promotion cadence from governor-completed implementations while separately reporting construction, evaluation, and canary funnel rates."
  ]
}
