Executive summary
Phase 3F-RT replayed only the exact two-source reproducer preserved by Phase 3F-CE. It performed no discovery, full-corpus extraction, claim-semantic changes, reproduction, candidate construction, governor action, or promotion.
The baseline was frozen before changes: LM Studio 0.4.21+2, LM Studio CLI 1.3.3, Qwen3.8-27B Q4_K_M, model SHA-256 e00082f779fa385cee8c68a3ec8833a75778cc87272240b942f74e0b8243e520, a 16,384-token context, full GPU offload, one slot, flash attention on, and llama.cpp CUDA runtime 2.31.2.
The current/latest 2.31.2 direct server, prior supported 2.29.1 direct server, LM Studio-managed 2.31.2 path, and 2.31.2 with flash attention disabled all produced a hard generation hang. No configuration qualified for the requested 10-consecutive-cycle reliability campaign.
Removing server-side JSON-schema constrained decoding did not remove the hang. The long unconstrained request stopped after 916 decoded tokens and left the slot occupied. The shorter unconstrained output completed but was correctly rejected by the unchanged validator as non-JSON.
Ordinary streamed cancellation recovered cleanly and a subsequent structured inference succeeded. Hard-hung generations did not recover after client disconnect and required process replacement.
The managed LM Studio path could report the model IDLE while bounded inference timed out. Model unload reported success, but reload stalled because the backend process survived. Forced replacement of the exact backend process and a clean reload restored inference.
A qualification-only harness, exact request builder, injectable flash-attention control with unchanged production default, tests, evidence documentation, and normalized runtime telemetry were committed on the isolated qualification branch at f7f704c53ae3b5656833f951ef4f02b0d09ea0e2. They were not merged or activated.
Production was restored to the original direct 2.31.2, flash-attention-on, single-slot configuration. A tiny strict structured inference returned valid JSON, the slot was idle, and Hiro health showed the expected model with matching checkout and loaded revision 19ff77a5b09cf3b76712a4b5eef3d6617089fdd9.
The final disposition is PHASE 3F-RT NOT DEMONSTRATED — MODEL/RUNTIME INCOMPATIBILITY. The frozen 20-source Phase 3F-CE corpus must not be rerun until a separately qualified supervisor can detect, replace, reload, verify, and boundedly retry a hard-hung runtime.