Executive summary
Investigated why Hiro's benchmark page appeared not to load from the owner's phone.
The phone successfully reached Hiro over Tailscale and received the benchmark HTML with HTTP 200, ruling out Hiro availability, port binding, Windows Firewall, and the phone-to-desktop network path.
The page's follow-up continuous-improvement data request returned HTTP 500 because candidate executor status synchronously launched a WSL readiness probe inside the web request.
While WSL was occupied by improvement work, the readiness probe exceeded its 15-second timeout. The timeout was not converted into a bounded status result, so it escaped through the API and prevented the page from rendering its data.
A follow-up status check confirmed that Hiro's web and chat process remained healthy, but both queue APIs returned HTTP 500 and the autonomous improvement scheduler was repeatedly failing closed on the same readiness timeout.
The Ubuntu WSL distribution reported stopped, and even a trivial command could not start within eight seconds. Therefore Hiro was online, but the improvement loop was stalled rather than actively processing candidates.
After disk space was restored on the system drive, WSL started normally, both benchmark data APIs returned HTTP 200 locally and over Tailscale, the sandbox reported ready, and the queue resumed with an active candidate.
Deeper post-recovery verification confirmed that a sandboxed assistant-lab worker was executing, but recent audit batches were not producing usable improvement evidence. Sandbox staging aborted on protected runtime and cache paths, while the everyday-question cases repeatedly reached their model timeout.
The queue's broad health label remained green despite these failures. One item was marked active and 36 were waiting, but the active supervisor episode was still waiting for candidate revision and no promotion transaction existed.
Implemented and activated revision 1cb3ac00422ee9e648e25628f92b2dc0da38dc6e. Sandbox staging now excludes protected runtime paths before traversal, standard input is transported through a read-only sandbox mount, executor probes are cached and timeout-safe, inference outages cannot become false quality incidents, and pipeline health exposes functional blockers.
The wedged LM Studio inference configuration was replaced with Hiro's checked text-only Qwen 3.8 runtime. A real completion returned in under one second after startup.
The exact revision passed 763 repository tests and all five frozen qualification scenarios. The live checkout fast-forwarded, restarted through the hidden launcher, reported matching loaded and checkout revisions, and completed a live audit with no infrastructure blocker and a 5/5 product lab.