Hiro development journal

Phase 3F-RT finds no reliable Qwen3.8 serving configuration

Runtime comparison, repository tests, public journal tests, production journal build, and production restoration complete Machine-readable JSON

Executive summary

Phase 3F-RT replayed only the exact two-source reproducer preserved by Phase 3F-CE. It performed no discovery, full-corpus extraction, claim-semantic changes, reproduction, candidate construction, governor action, or promotion.

The baseline was frozen before changes: LM Studio 0.4.21+2, LM Studio CLI 1.3.3, Qwen3.8-27B Q4_K_M, model SHA-256 e00082f779fa385cee8c68a3ec8833a75778cc87272240b942f74e0b8243e520, a 16,384-token context, full GPU offload, one slot, flash attention on, and llama.cpp CUDA runtime 2.31.2.

The current/latest 2.31.2 direct server, prior supported 2.29.1 direct server, LM Studio-managed 2.31.2 path, and 2.31.2 with flash attention disabled all produced a hard generation hang. No configuration qualified for the requested 10-consecutive-cycle reliability campaign.

Removing server-side JSON-schema constrained decoding did not remove the hang. The long unconstrained request stopped after 916 decoded tokens and left the slot occupied. The shorter unconstrained output completed but was correctly rejected by the unchanged validator as non-JSON.

Ordinary streamed cancellation recovered cleanly and a subsequent structured inference succeeded. Hard-hung generations did not recover after client disconnect and required process replacement.

The managed LM Studio path could report the model IDLE while bounded inference timed out. Model unload reported success, but reload stalled because the backend process survived. Forced replacement of the exact backend process and a clean reload restored inference.

A qualification-only harness, exact request builder, injectable flash-attention control with unchanged production default, tests, evidence documentation, and normalized runtime telemetry were committed on the isolated qualification branch at f7f704c53ae3b5656833f951ef4f02b0d09ea0e2. They were not merged or activated.

Production was restored to the original direct 2.31.2, flash-attention-on, single-slot configuration. A tiny strict structured inference returned valid JSON, the slot was idle, and Hiro health showed the expected model with matching checkout and loaded revision 19ff77a5b09cf3b76712a4b5eef3d6617089fdd9.

The final disposition is PHASE 3F-RT NOT DEMONSTRATED — MODEL/RUNTIME INCOMPATIBILITY. The frozen 20-source Phase 3F-CE corpus must not be rerun until a separately qualified supervisor can detect, replace, reload, verify, and boundedly retry a hard-hung runtime.

Work completed

Frozen reproducer and stack inventory

Completed
  • Frozen the exact two source IDs, source-content hashes, source-text hashes, extraction system prompt, user payload, JSON schema, sampling settings, model identity, model hash, context, GPU settings, request hashes, and existing deterministic validator.
  • Confirmed LM Studio reported 2.31.2 as the latest stable supported CUDA runtime, making the current and latest-stable rows the same build.
  • Confirmed the strict request path reports peg-native chat formatting, reasoning disabled, temperature zero, and a 3,072-token output ceiling.
  • Captured GPU and driver identity, runtime manifests, selected backend preferences, server load arguments, model metadata, slot state, and immutable per-run evidence before runtime changes.

Current and prior direct-runtime comparison

No reliable build
  • Direct 2.31.2 produced three valid strict completions across the accepted baseline and clean repeat, but also produced the same unrecovered hard hang. Valid latencies ranged from approximately 5.16 to 63.39 seconds.
  • Direct 2.29.1 hard-hung on the first strict frozen request after ordinary cancellation had recovered normally.
  • Disabling flash attention on direct 2.31.2 did not repair the issue; the first strict request again failed to produce recoverable progress.
  • Every hard-hang attempt stopped immediately at the first unresolved divergence. Later sources were not queued behind the unavailable slot.

Structured-output diagnostic

Strict-schema-only cause rejected
  • The unconstrained diagnostic kept the exact semantic prompt and output shape request but removed only server-side JSON-schema constrained decoding.
  • The short source completed in about 8.24 seconds but returned content that the unchanged deterministic validator rejected as invalid JSON.
  • The long source stopped after 916 decoded tokens, failed cancellation recovery, and left the single slot processing.
  • This demonstrates that removing response_format does not remove the hard hang and is not a production-ready fallback.

LM Studio managed serving path

Hard hang reproduced
  • The latest 2.31.2 runtime was loaded through LM Studio's supported server and model-management commands with the frozen context, GPU offload, parallelism, and model identity.
  • Three strict requests completed with valid output before the fourth hard-hung, so moving from the direct launcher to the managed serving layer did not eliminate the failure.
  • The managed control plane later reported the model IDLE while a tiny inference timed out, proving dashboard process state is not inference health.
  • Unload reported success, but reload stalled because the backend process remained alive. Stopping the HTTP server did not terminate that backend. Exact backend process replacement, state cleanup, and reload restored bounded inference.

Health and recovery contract

Defined; unattended supervisor not yet qualified
  • PROCESS_HEALTHY now means the expected process/model identity plus a real health signal where one exists; HTTP 200 with an error body is not healthy.
  • INFERENCE_HEALTHY requires the expected loaded model, a tiny bounded structured inference whose JSON passes deterministic validation, and an idle direct-server slot afterward.
  • An ordinary two-second streamed cancellation released the direct slot and passed subsequent inference health.
  • A future supervisor must preserve the failed attempt, detect stopped progress, attempt cancellation, verify inference health, replace the exact serving process when necessary, reload the same frozen configuration, reverify inference, and permit at most one separately versioned retry.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused Phase 3F-RT, request-runtime, source-claim, and launcher tests passed Sixteen focused tests passed in 1.04 seconds. They cover exact request construction, frozen-source enforcement, unchanged validation in unconstrained diagnostics, hard-hang classification, bounded request behavior, and unchanged flash-attention defaults.
Complete Hiro repository suite passed 897 tests passed, one expected test skipped, and zero tests failed in 515.11 seconds. Six existing unknown-mark warnings were reported.
Direct 2.31.2 strict control mixed The frozen pair completed cleanly in one cycle, but the same runtime also reproduced an unrecovered strict hard hang. A single successful cycle therefore did not qualify the runtime.
Direct 2.31.2 unconstrained diagnostic failed Two requests were attempted: one schema-invalid completion and one hard generation hang after 916 decoded tokens.
Direct prior 2.29.1 strict runtime failed The first frozen request hard-hung and did not recover after client cancellation.
LM Studio-managed 2.31.2 strict runtime failed Three of four attempted requests completed validly; the fourth hard-hung. Managed unload/reload did not restore service until the exact backend process was replaced.
Direct 2.31.2 with flash attention off failed The first strict frozen request hard-hung, so flash attention was not the isolating configuration variable.
Production restoration passed The original 2.31.2 flash-on runtime returned valid strict JSON from a tiny bounded inference, exposed an idle slot, and Hiro health reported connected Qwen with matching live revisions.
Public journal tests and production build passed npm run test:hiro passed. npm run build generated and validated 187 timestamped journal entries, compiled TypeScript, and completed the Vite production bundle.

Current state

Next steps