{
  "schemaVersion": 2,
  "date": "2026.08.25",
  "publishedAt": "2026-08-25T10:51:12-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Independent-inspector repairs close Hiro's critical autonomous-loop gaps",
  "publicationStatus": "Implemented, fully tested, diversely qualified, and activated",
  "executiveSummary": [
    "Implemented the independent inspector's critical and high-priority repairs across candidate isolation, evidence ownership, promotion transactions, runtime identity, rollback, persistence, concurrency, product-path testing, API failures, notifications, and observability.",
    "Model-authored candidate code now runs inside a networkless WSL2 and Bubblewrap boundary copied into Linux storage. It cannot see Windows mounts, inherited secrets, or writable host paths, and the live service refuses candidate execution if this boundary is unavailable.",
    "The pre-activation timer is now truthfully labeled synthetic soak. A separate eight-hour runtime probation begins only after the restarted service proves it loaded the promoted revision, with checkpoints at zero, one, four, and eight hours and additive rollback on failure.",
    "Promotion is now a durable state machine spanning prepared intent, Git advancement, hidden restart, loaded-revision verification, runtime probation, finalization, compensation, and rollback restart. Only one promotion transaction may be active.",
    "External observations remain inert structured claims until an evaluator-owned causal control fails on the untouched local baseline. A file's existence is no longer treated as corroboration.",
    "A real assistant laboratory now exercises the response boundary, core agent turn, chat API, and session continuity across preference, follow-up, current-information, trip-planning, and transit tasks without using the production endpoint or performing external actions.",
    "User corrections are linked to the immediately preceding failed turn, including its bounded tool-trace hash and failure stage, so repair candidates target the original failure instead of the user's complaint text.",
    "The qualification gate no longer repeats the same happy path five times. The frozen revision passed five distinct failure-oriented scenarios covering candidate evidence, promotion and rollback, autonomous recovery, runtime isolation and state migration, and the assistant product path.",
    "Revision 157ee7c0bc6e3b8a23eedd0b7ef1e42a44280708 passed all 761 repository tests and all five qualification scenarios, then fast-forwarded the clean live .hd3 checkout from a11954e without merging, rebasing, or touching the unrelated dirty root checkout.",
    "Hiro restarted healthy with Qwen 3.8 connected. Its loaded revision equals its checkout revision, the candidate executor is live-eligible and network-denied, the qualification packet verifies 5/5 scenarios, and the durable queue remains intact."
  ],
  "workstreams": [
    {
      "title": "Operating-system candidate isolation",
      "status": "Implemented and active",
      "details": [
        "Added one restricted execution interface and routed candidate construction, targeted tests, paired evaluation, integration, canary probes, and governor test execution through it.",
        "The live backend copies the candidate source into Linux storage, removes repository metadata, runtime artifacts, logs, backups, secret directories, and environment files, then initializes a disposable local Git baseline.",
        "Bubblewrap unshares namespaces, drops capabilities, denies network access, clears the environment, exposes read-only system/runtime mounts, permits writes only in disposable work and temporary mounts, and applies CPU, memory, file, descriptor, process, and output limits.",
        "A hostile boundary test proves that Windows mounts and inherited secrets are absent, network connection fails, root writes fail, and only disposable work and temporary paths remain writable."
      ]
    },
    {
      "title": "Truthful soak, activation, and runtime probation",
      "status": "Implemented and active",
      "details": [
        "Renamed the pre-activation canary to synthetic soak in policy, API, queue state, and Observatory labels while preserving a compatibility alias for older local clients.",
        "Low-risk synthetic soak remains zero, five, and fifteen minutes; moderate-risk soak remains zero, five, fifteen, and sixty minutes.",
        "Runtime probation is a distinct eight-hour phase at zero, sixty, 240, and 480 minutes and begins only after health reports the candidate as the loaded revision.",
        "Live probation combines service identity, independent tests, the real assistant laboratory, and bounded Stage 6 error-rate and latency comparisons. A failed checkpoint creates an additive revert, restarts Hiro, and verifies the rollback revision."
      ]
    },
    {
      "title": "Durable promotion transaction and restart recovery",
      "status": "Implemented and active",
      "details": [
        "Added a promotion transaction ledger with checked transitions, append-only events, one-active-transaction enforcement, idempotent reconciliation, and explicit terminal failure states.",
        "The activation broker stops only the exact Hiro process, waits for all three service ports to close, uses the normalized hidden launcher, and accepts success only when health reports the expected loaded revision.",
        "If activation fails, the broker first stops any partially started replacement, creates an additive Git revert, restarts Hiro, verifies the recovered runtime, and records either rolled_back or an explicit failed compensation receipt.",
        "The launcher now supplies an unambiguous loaded revision at process creation. Health reports process ID, start time, loaded revision, checkout revision, and whether they match."
      ]
    },
    {
      "title": "Evidence ownership and candidate integrity",
      "status": "Implemented and active",
      "details": [
        "Candidate create operations cannot overwrite an existing path, and generated plans are rejected for duplicate paths, missing modify targets, existing create targets, symlinks, or attempts to replace harness-owned tests.",
        "External text is reduced to a prompt-safe structured claim containing code-owned technique terms, mechanism, preconditions, expected direction, protocol identity, and source hashes. Verbatim external prose never becomes a model instruction.",
        "Corroboration now requires an evaluator-owned independent control to reproduce an assertion failure on the untouched baseline while an isolated negative control passes. Merely finding a related source file is insufficient.",
        "Changes to the candidate builder, evaluator, governor, promotion, runtime, server, and launcher trust base are routed to a separate meta-change lane and cannot approve themselves."
      ]
    },
    {
      "title": "Real assistant behavior laboratory and incident linkage",
      "status": "Implemented and active",
      "details": [
        "Added five deterministic product scenarios based on the observed failures: Pepsi versus Coke, the follow-up 'check' turn, Los Angeles weekend discovery, Big Bear trip planning, and transit directions.",
        "Each scenario runs the actual core agent with fake code-owned tools and evidence, verifies response contracts, exercises a disposable FastAPI chat session, and checks correlation IDs and continuity.",
        "The rotating everyday audit runs this product laboratory alongside its broader local suite. Product failures require two observations before queue admission to reduce stochastic false positives.",
        "Response-quality records now carry turn identity, parent linkage, feedback type, tool-trace hash, and failure stage. Correction harvesting replays the original failed query and response rather than treating the correction itself as the task."
      ]
    },
    {
      "title": "Persistence, leases, supervision, and legacy retirement",
      "status": "Implemented and active",
      "details": [
        "Continuous queue and supervisor databases now use contiguous versioned migrations with auditable receipts and historical-schema upgrade tests.",
        "Runner and discovery ownership cannot be stolen from a live process merely because a nominal time-to-live elapsed. Production activity uses one SQLite lease per real request so one worker cannot clear another worker's activity.",
        "Only one supervisor episode can be active, attempts remain append-only, background tasks are retained and report terminal exceptions, and Windows child launches normalize the Path/PATH collision while preserving Hiro's inherited runtime context.",
        "The retired self-improvement-v2 scheduler cannot run. Manual-only and unknown lab modes do not fall through to nightly work, and disabled Telegram reporting creates no replayable outbox message."
      ]
    },
    {
      "title": "Safe API errors and Observatory realignment",
      "status": "Implemented and active",
      "details": [
        "Complete chat failures return a stable safe envelope with a correlation ID while private exceptions remain only in local logs.",
        "Streaming chat uses valid JSON Server-Sent Event frames for chunks, completion, and safe retryable errors, including multiline model content.",
        "The Observatory now shows queue, investigation, construction, testing, synthetic soak, activation, runtime probation, and finalization as distinct stages.",
        "It reports the live candidate executor, loaded-revision verification, active promotion transaction, eight-hour runtime policy, and five diverse qualification scenarios rather than the retired five-repeat qualification language."
      ]
    },
    {
      "title": "Release reconciliation and live activation",
      "status": "Completed",
      "details": [
        "All implementation and validation occurred in a clean release worktree based on the exact prior live revision a11954e.",
        "Two accidental Bubblewrap package-extraction artifacts were preserved outside the repository and excluded from the release commit.",
        "The original root checkout had diverged history and an unrelated uncommitted benchmark experiment. It was left untouched after the actual live .hd3 checkout was identified as the clean deployment source.",
        "The .hd3 checkout fast-forwarded by one qualified commit, received the exact checksum-verified qualification packet, restarted through the hidden checked launcher, and reported matching loaded and checkout revisions."
      ]
    }
  ],
  "decisions": [
    "Treat candidate code as untrusted and require an operating-system boundary for every live candidate execution path.",
    "Do not call synthetic checks runtime probation; activation and loaded-revision proof are prerequisites for the eight-hour live period.",
    "Make Git advancement, restart, runtime verification, probation, and rollback one durable recoverable transaction.",
    "Do not allow improvement machinery to approve changes to its own trust base.",
    "Require causal local reproduction for external ideas and keep source text inert throughout the candidate path.",
    "Test the assistant that the user actually experiences, while keeping deterministic tools, disposable sessions, and no external side effects.",
    "Link corrections to the failed turn that caused them so the queue contains actionable behavior evidence.",
    "Preserve live leases while their owning processes exist; elapsed time alone is not proof that user activity or candidate work ended.",
    "Replace repeated qualification cycles with distinct failure-oriented scenarios whose receipts state exactly what was exercised.",
    "Preserve unrelated work and activate only the verified clean deployment checkout by fast-forward."
  ],
  "validation": [
    {
      "check": "Focused inspector-remediation suite",
      "status": "passed",
      "result": "210 focused tests passed in 359.74 seconds across continuous governance, candidates, integration, supervision, dashboards, notifications, interaction loops, activation state, database migration, product-path behavior, sandboxing, qualification, launch, and scheduler retirement."
    },
    {
      "check": "Full repository regression suite",
      "status": "passed",
      "result": "The exact release revision passed all 761 tests in 430.60 seconds. Four generated autonomous tests emitted pre-existing unknown hiro_contract marker warnings; there were no failures."
    },
    {
      "check": "Candidate sandbox hostile-boundary test",
      "status": "passed",
      "result": "The WSL2/Bubblewrap test proved no Windows mount, no inherited secret, no network connection, no root write, and successful writes only to disposable work and temporary mounts."
    },
    {
      "check": "Five-scenario frozen pipeline qualification",
      "status": "passed",
      "result": "Revision 157ee7c0bc6e3b8a23eedd0b7ef1e42a44280708 passed candidate_evaluation, promotion_and_rollback, autonomous_recovery, runtime_isolation_and_state, and assistant_product_path. The packet SHA-256 is e3fd336851480b42c5b8436ef8895814b72b97b02d36b5da224f15f29b89f3a7."
    },
    {
      "check": "Live fast-forward and hidden restart",
      "status": "passed",
      "result": "The clean .hd3 deployment fast-forwarded from a11954e to 157ee7c and restarted in 8.39 seconds without restarting Qwen."
    },
    {
      "check": "Loaded runtime identity",
      "status": "passed",
      "result": "Health reports OK, Qwen 3.8 27B connected, loaded revision 157ee7c, checkout revision 157ee7c, and revision_matches true."
    },
    {
      "check": "Live qualification and executor",
      "status": "passed",
      "result": "The live Observatory API verifies all five qualification scenarios and reports the wsl_bwrap executor as OS-isolated, network-denied, and live-eligible."
    },
    {
      "check": "Durable queue preservation",
      "status": "passed",
      "result": "The live queue retained 519 total ideas, 37 actionable records, seven implementations, and the existing supervised retry episode across the restart."
    },
    {
      "check": "Public journal tests and build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated all 161 journal pages, compiled TypeScript, and completed the Vite production bundle."
    }
  ],
  "currentState": [
    "Hiro is running revision 157ee7c0bc6e3b8a23eedd0b7ef1e42a44280708 from the clean .hd3 deployment checkout.",
    "Health is OK, Qwen 3.8 27B is connected, and the loaded revision matches the checkout revision.",
    "The frozen pipeline qualification is current at five of five diverse scenarios.",
    "The live candidate executor is WSL2/Bubblewrap, operating-system isolated, network denied, and eligible for autonomous candidate work.",
    "The durable queue contains 37 actionable records and seven prior implementations. The active supervisor episode is waiting to revise a previously failed candidate under the new builder revision.",
    "The first post-restart ticks are honoring the three-minute production quiet window created when the dead prior-process request lease was reclaimed; queue advancement resumes automatically afterward.",
    "Telegram improvement reporting remains disabled, retired day/night/self-improvement schedulers remain inactive, and the continuous evidence-driven loop is the sole automatic improvement path."
  ],
  "limitations": [
    "The real assistant laboratory covers response, agent, API, and session boundaries but does not yet drive an actual browser DOM or mobile web view. Browser-level crash and rendering coverage remains future work.",
    "The broad everyday interaction audit still depends on local-model latency and can produce timeout observations. Two observations are required before those failures authorize candidate work.",
    "The eight-hour runtime probation is now correctly wired but no newly promoted candidate has yet completed that full live period under this revision.",
    "Meta-changes to the builder, evaluator, governor, promotion transaction, runtime, server, or launch boundary intentionally require a separately judged release and cannot be self-promoted.",
    "The original root checkout still contains an unrelated uncommitted benchmark experiment and divergent local history. It was intentionally excluded from and untouched by deployment.",
    "The generated autonomous-test marker warnings remain non-fatal and should be registered or removed so future full-suite output is completely clean."
  ],
  "nextSteps": [
    "Observe the first post-quiet-window supervised retry and verify that candidate construction occurs through the live WSL2/Bubblewrap executor.",
    "Let the rotating everyday audit revisit the observed preference, current-information, planning, and directions failures; admit only reproduced failures and track their original-turn lineage in the Observatory.",
    "When a candidate clears synthetic soak, verify the full activation transaction and allow its eight-hour runtime probation to finalize or roll back automatically.",
    "Add browser-level conversation and crash testing for the mobile/web interface while retaining disposable sessions and no external side effects.",
    "Register the hiro_contract pytest marker and continue reducing stale duplicate historical queue entries without altering current evidence.",
    "Keep trust-base changes in the independent meta-change lane and use this same diverse qualification plus loaded-revision verification for future releases."
  ],
  "disclosureNote": "This public entry contains no credentials, private conversation text, private external-source content, hidden reasoning, or actionable unresolved security details."
}
