{
  "schemaVersion": 2,
  "date": "2026.08.25",
  "publishedAt": "2026-08-25T17:49:36-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Mobile Observatory failure traced to a blocking sandbox readiness probe",
  "publicationStatus": "Blockers repaired, fully tested, qualified, activated, and verified live",
  "executiveSummary": [
    "Investigated why Hiro's benchmark page appeared not to load from the owner's phone.",
    "The phone successfully reached Hiro over Tailscale and received the benchmark HTML with HTTP 200, ruling out Hiro availability, port binding, Windows Firewall, and the phone-to-desktop network path.",
    "The page's follow-up continuous-improvement data request returned HTTP 500 because candidate executor status synchronously launched a WSL readiness probe inside the web request.",
    "While WSL was occupied by improvement work, the readiness probe exceeded its 15-second timeout. The timeout was not converted into a bounded status result, so it escaped through the API and prevented the page from rendering its data.",
    "A follow-up status check confirmed that Hiro's web and chat process remained healthy, but both queue APIs returned HTTP 500 and the autonomous improvement scheduler was repeatedly failing closed on the same readiness timeout.",
    "The Ubuntu WSL distribution reported stopped, and even a trivial command could not start within eight seconds. Therefore Hiro was online, but the improvement loop was stalled rather than actively processing candidates.",
    "After disk space was restored on the system drive, WSL started normally, both benchmark data APIs returned HTTP 200 locally and over Tailscale, the sandbox reported ready, and the queue resumed with an active candidate.",
    "Deeper post-recovery verification confirmed that a sandboxed assistant-lab worker was executing, but recent audit batches were not producing usable improvement evidence. Sandbox staging aborted on protected runtime and cache paths, while the everyday-question cases repeatedly reached their model timeout.",
    "The queue's broad health label remained green despite these failures. One item was marked active and 36 were waiting, but the active supervisor episode was still waiting for candidate revision and no promotion transaction existed.",
    "Implemented and activated revision 1cb3ac00422ee9e648e25628f92b2dc0da38dc6e. Sandbox staging now excludes protected runtime paths before traversal, standard input is transported through a read-only sandbox mount, executor probes are cached and timeout-safe, inference outages cannot become false quality incidents, and pipeline health exposes functional blockers.",
    "The wedged LM Studio inference configuration was replaced with Hiro's checked text-only Qwen 3.8 runtime. A real completion returned in under one second after startup.",
    "The exact revision passed 763 repository tests and all five frozen qualification scenarios. The live checkout fast-forwarded, restarted through the hidden launcher, reported matching loaded and checkout revisions, and completed a live audit with no infrastructure blocker and a 5/5 product lab."
  ],
  "workstreams": [
    {
      "title": "Network-path diagnosis",
      "status": "Completed",
      "details": [
        "Confirmed that port 8001 was listening on all IPv4 interfaces and that the benchmark route returned HTTP 200 locally through loopback, the LAN address, and the Tailscale address.",
        "Confirmed that the phone was online in the private network and that server access logs recorded its benchmark requests with HTTP 200.",
        "Confirmed that enabled inbound Python rules covered the active Windows firewall profiles."
      ]
    },
    {
      "title": "Application failure isolation",
      "status": "Completed",
      "details": [
        "Correlated each successful phone request for the benchmark HTML with an HTTP 500 response from the continuous-improvement data endpoint.",
        "The exception trace identified the candidate executor status function as the source of failure.",
        "That function performs a synchronous WSL subprocess probe with a 15-second timeout during every dashboard data request.",
        "A subprocess timeout was allowed to escape instead of returning a conservative unavailable or busy status."
      ]
    },
    {
      "title": "Autonomous-loop status",
      "status": "Recovered",
      "details": [
        "Health remained OK, Qwen remained connected, all three Hiro service ports remained bound by the expected process, and loaded and checkout revisions still matched.",
        "Both queue APIs returned HTTP 500 because they shared the failing synchronous executor-status probe.",
        "Scheduler logs showed repeated fail-closed cycles caused by the same 15-second timeout, so no candidate could advance while the executor remained unavailable.",
        "The configured Ubuntu WSL distribution was stopped, and a bounded trivial-command probe could not start it within eight seconds.",
        "Once approximately three gigabytes of system-drive space became available, the same bounded WSL command returned successfully, the executor reported ready, and the supervisor resumed one active candidate."
      ]
    },
    {
      "title": "Post-recovery functional verification",
      "status": "Completed and repaired",
      "details": [
        "A live WSL and Bubblewrap assistant-lab worker process was observed, proving that the scheduler had launched real test work rather than only updating a database flag.",
        "The worker's repository staging copied the checkout recursively before removing excluded runtime paths. Permission-protected runtime and cache directories caused the copy to fail before cleanup, so the actual assistant lab reported infrastructure_blocked with zero passed and zero failed product-path cases.",
        "Recent rotating everyday-question batches each failed all four cases at roughly the configured 30-second model timeout and queued no repair incident.",
        "The circuit breaker remained closed at one of three infrastructure failures, the executor readiness probe remained green, and the aggregate pipeline-health field remained healthy, demonstrating that those indicators did not capture the functional blockage.",
        "The active idea remained in investigation with a complete_reproduction next action. Its supervisor episode remained waiting for candidate revision, and there was no active promotion transaction."
      ]
    },
    {
      "title": "Sandbox execution repair",
      "status": "Implemented and active",
      "details": [
        "Replaced copy-then-delete staging with an exclusion-aware archive transfer so Git metadata, runtime state, test caches, logs, backups, secrets, environment files, and bytecode caches are never traversed or copied into the Linux sandbox.",
        "Moved subprocess input into a separately encoded file mounted read-only inside Bubblewrap. The prior launcher consumed standard input while decoding its own script, leaving the product worker with an empty payload.",
        "Kept the existing networkless, capability-dropped, read-only-root boundary intact and verified that Windows mounts, secrets, excluded paths, and root writes remain unavailable."
      ]
    },
    {
      "title": "Inference and evidence repair",
      "status": "Implemented and active",
      "details": [
        "Confirmed that the previous Qwen service could not answer an eight-token request within 120 seconds despite reporting idle and healthy.",
        "Stopped only the unusable Qwen service and started the checked direct text-only runtime with the vision projector disabled. The exact Qwen 3.8 27B model then returned a real completion in 0.72 seconds.",
        "Interaction cases now label timeout or execution errors as infrastructure rather than assistant quality. Such observations cannot enter the repair queue.",
        "A completed model response can still create a quality observation. The live audit found one bounded correction-response miss, stored it as a first observation, and correctly withheld actionable admission pending reproduction."
      ]
    },
    {
      "title": "Truthful operational health",
      "status": "Implemented and active",
      "details": [
        "Executor readiness probes now fail closed on timeout or operating-system errors and cache their result for sixty seconds, preventing repeated dashboard requests from spawning blocking WSL checks.",
        "Interaction-audit status reports usable cases, infrastructure-blocked cases, suppressed queueing, product-lab status, and explicit blockers.",
        "Pipeline health now combines supervisor liveness, executor eligibility, and the latest clean product audit instead of reporting green from supervisor counters alone."
      ]
    }
  ],
  "decisions": [
    "Treat this as an application-layer dashboard regression rather than a connectivity problem.",
    "Do not weaken candidate isolation or interrupt the autonomous promotion cycle to make the dashboard load.",
    "The appropriate repair is to remove the expensive readiness probe from the request path, cache or refresh executor status outside the request, and fail closed with a bounded status object if probing times out.",
    "Add a regression test proving the dashboard data API remains responsive when the sandbox status probe is slow, busy, or unavailable.",
    "Treat inference timeouts as infrastructure evidence only; require an actual completed response before judging assistant quality.",
    "Use the checked direct text-only Qwen runtime for Hiro rather than LM Studio's vision, memory-locking, and speculative-decoding configuration.",
    "Require a clean live audit to clear operational health after activation; do not erase a historical blocker merely because code was deployed."
  ],
  "validation": [
    {
      "check": "Port 8001 listener",
      "status": "passed",
      "result": "The active Hiro process was listening on 0.0.0.0:8001."
    },
    {
      "check": "Benchmark HTML",
      "status": "passed",
      "result": "The benchmark route returned HTTP 200 through loopback, LAN, Tailscale, and the phone's recorded Tailscale request."
    },
    {
      "check": "Dashboard data endpoint from phone",
      "status": "failed",
      "result": "The continuous-improvement API returned HTTP 500 on three recorded phone requests."
    },
    {
      "check": "Exception attribution",
      "status": "passed",
      "result": "The server traceback consistently attributed the failure to an uncaught timeout from the synchronous WSL executor-readiness probe."
    },
    {
      "check": "Hiro runtime health",
      "status": "passed",
      "result": "Health returned OK, Qwen was connected, ports 8000, 8001, and 8765 were owned by the expected Hiro process, and loaded and checkout revisions matched."
    },
    {
      "check": "Autonomous scheduler progress",
      "status": "passed after recovery",
      "result": "Repeated cycles initially failed closed, but after WSL recovered the live queue showed one active candidate, 36 waiting, and no retrying item."
    },
    {
      "check": "Bounded WSL startup probe",
      "status": "passed after recovery",
      "result": "The initial probe timed out while the distribution was stopped. After system-drive space was restored, the same trivial command returned successfully."
    },
    {
      "check": "Recovered benchmark path",
      "status": "passed",
      "result": "The benchmark HTML and both queue APIs returned HTTP 200. The full page and queue API also returned HTTP 200 over the Tailscale address used by the phone."
    },
    {
      "check": "Live assistant-lab execution",
      "status": "blocked",
      "result": "A real sandbox worker was active, but recent actual-assistant-lab runs stopped during source staging on protected runtime and cache paths and returned infrastructure_blocked with no completed cases."
    },
    {
      "check": "Everyday-question audit usefulness",
      "status": "passed after remediation",
      "result": "A disposable real-model batch completed four usable cases in 1.4 to 7.8 seconds: three passed and one produced a genuine content-length quality failure. The paired product lab passed all five cases."
    },
    {
      "check": "Focused remediation suite",
      "status": "passed",
      "result": "All 22 focused tests passed across restricted execution, interaction auditing, assistant product behavior, the active improvement loop, and Observatory health reporting."
    },
    {
      "check": "Full repository regression suite",
      "status": "passed",
      "result": "The exact release revision passed all 763 tests in 453.67 seconds. Four pre-existing generated-test marker warnings remained non-fatal."
    },
    {
      "check": "Frozen five-scenario qualification",
      "status": "passed",
      "result": "Revision 1cb3ac0 passed candidate evaluation, promotion and rollback, autonomous recovery, runtime isolation and state, and assistant product path. The packet SHA-256 is 654cb6cc38b40e48fae47bb344b0a24f72ae178eaf167b641e64a9096a7d7e72."
    },
    {
      "check": "Live activation and identity",
      "status": "passed",
      "result": "The clean live checkout fast-forwarded to 1cb3ac0, restarted in 2.328 seconds through the hidden launcher, and reported matching loaded and checkout revisions."
    },
    {
      "check": "Live post-activation audit",
      "status": "passed",
      "result": "The live audit completed four usable real-model cases with no infrastructure block, and the actual assistant product lab passed five of five. Pipeline health became true with no operational blockers."
    },
    {
      "check": "Autonomous next-work selection",
      "status": "passed",
      "result": "After retiring stale blocked work, the next minute tick automatically selected a fresh idea, changed it to investigating, activated the supervisor, and left 29 ranked ideas waiting without user intervention."
    },
    {
      "check": "Public journal tests and build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated all 162 journal pages, compiled TypeScript, and completed the Vite production bundle."
    }
  ],
  "currentState": [
    "Hiro's web and chat service remains healthy and reachable from the phone over Tailscale.",
    "The benchmark document and its improvement data APIs are responding locally and over Tailscale.",
    "The queue contains 551 ideas, 37 actionable records, 36 waiting records, one active candidate, and seven implemented ideas.",
    "The WSL2 and Bubblewrap executor reports ready and live-eligible, and the autonomous supervisor is active.",
    "Hiro is running qualified revision 1cb3ac00422ee9e648e25628f92b2dc0da38dc6e with matching loaded and checkout identities.",
    "Qwen 3.8 27B is inference-ready through the checked text-only runtime.",
    "The benchmark APIs respond, the WSL2 and Bubblewrap executor is ready, interaction-audit health is true, and pipeline health has no operational blocker.",
    "The queue contains 552 ideas, 30 actionable records, 29 waiting records, one fresh active investigation, and seven implemented ideas.",
    "The latest broad audit recorded one completed-response correction-quality miss as a first observation. It requires a second reproduction before becoming actionable, as designed.",
    "No promotion transaction is currently active."
  ],
  "limitations": [
    "The system drive had about seven gigabytes free during the latest verification, which remains limited headroom for WSL and other stateful services.",
    "One quality miss has only a single observation and is intentionally not yet actionable.",
    "No candidate has entered a promotion transaction under this revision yet; the normal evaluation, synthetic-soak, activation, and eight-hour runtime-probation gates still apply."
  ],
  "nextSteps": [
    "Let the next scheduled audit attempt reproduce the correction-quality miss; admit it only if the completed-response failure repeats.",
    "Observe the active investigation through reproduction and candidate construction, then allow the existing safety gates to decide whether it advances.",
    "Maintain substantially more system-drive headroom so WSL and stateful services do not fail under temporary storage pressure.",
    "Register the generated hiro_contract pytest marker to remove the four non-fatal warnings."
  ],
  "disclosureNote": "This public entry contains no credentials, private message content, private network addresses, or actionable unresolved security details."
}
