{
  "schemaVersion": 2,
  "date": "2026.08.30",
  "publishedAt": "2026-08-30T10:11:39-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Phase 3F-CE isolates an intermittent Qwen runtime hard hang",
  "publicationStatus": "Hiro qualification, repository tests, public journal tests, and production journal build complete",
  "executiveSummary": [
    "Phase 3F-CE used the exact 20-source immutable corpus from the failed Phase 3F campaign. It did not run new discovery and did not reach relevance, feasibility, reproduction, candidate, governor, or promotion stages.",
    "The original request path was established precisely: sequential non-streaming HTTPX calls to Qwen 3.8 27B, one server slot, strict JSON Schema output, temperature zero, reasoning disabled, a 3,072-token extraction ceiling, and a 300-second HTTP inactivity timeout with no separate wall-clock deadline.",
    "The retained prompts were small: 685 to 1,011 input tokens. A clean control completed in 7.59 seconds, and a longer preserved extraction completed in 61.025 seconds. Ordinary latency and input size therefore do not explain the original uniform 20 timeouts.",
    "The real failure is an intermittent hard generation hang. A preserved extraction stopped at 116 decoded tokens while continuing to consume CPU and GPU. The model health endpoint could remain healthy while the execution slot and later requests stopped completing. Client disconnect did not clear the hard-hung slot.",
    "The old non-streaming transport allowed every later source request to accumulate behind the unavailable single slot. This explains how individually viable prompts produced 20 identical ReadTimeout outcomes.",
    "A bounded streamed request contract, normalized terminal reasons, request/token/timing hashes, queue observation, cancellation verification, a frozen-corpus runner, and focused tests were implemented on an isolated qualification branch. The branch commit is 37ccb85380b5259a24c3ed7368ad810619261d6e.",
    "The qualification stopped at the first unresolved runtime divergence after two source attempts: one success, one timeout/hard hang, and 18 intentionally not attempted. The instrumentation branch was not merged or activated. Production Hiro remained on revision 19ff77a5b09cf3b76712a4b5eef3d6617089fdd9.",
    "The final disposition is PHASE 3F-CE NOT DEMONSTRATED — MODEL LATENCY/REQUEST CONTRACT FAILURE. A new fresh Phase 3F campaign is not justified until a stable runtime implementation passes the same two-source reproducer and then all 20 frozen sources."
  ],
  "workstreams": [
    {
      "title": "Original-contract reconstruction",
      "status": "Completed",
      "details": [
        "Confirmed HTTPX as the client and the OpenAI-compatible chat-completions endpoint as the model boundary.",
        "Confirmed the 300-second scalar HTTPX timeout applies to connect, read, write, and pool inactivity; there was no distinct total request budget.",
        "Confirmed sequential corpus iteration and one Qwen server slot, so the claim caller itself did not issue a parallel batch.",
        "Confirmed strict JSON Schema output, temperature zero, reasoning effort none, a 3,072-token extraction limit, and a 2,048-token validator limit.",
        "The failed campaign duration was approximately 20 times the 300-second read limit, matching the preserved 20 ReadTimeout outcomes."
      ]
    },
    {
      "title": "Input and healthy-runtime controls",
      "status": "Completed",
      "details": [
        "Exact chat-template tokenization placed every preserved extraction prompt between 685 and 1,011 tokens; source text ranged from 1,040 to 2,158 characters.",
        "The intended Qwen 3.8 27B model loaded with the expected 16,384-token context and one execution slot.",
        "A simple strict structured-output health probe completed in 1.344 seconds before diagnostics and in 0.867 seconds after production restoration.",
        "The first preserved extraction completed in 7.59 seconds with 783 input and 136 output tokens and a schema-valid response.",
        "A longer preserved extraction completed in one clean control in 61.025 seconds with 1,011 input and 1,942 output tokens and a schema-valid response."
      ]
    },
    {
      "title": "Hard-hang and abandoned-work diagnosis",
      "status": "Failed runtime boundary isolated",
      "details": [
        "During qualification, the second preserved extraction reached 116 decoded tokens and then ceased token progress while a CPU core and approximately 48 percent GPU utilization remained active.",
        "The execution slot remained marked processing after the client disconnected. In the hard-hung state, slot inspection could itself stop responding even though the health endpoint remained healthy.",
        "An ordinary streamed cancellation control was different: a forced two-second cutoff after 28 chunks made the slot idle immediately and it remained idle.",
        "The distinction proves that client timeout cannot be treated as server cancellation. Streaming handles ordinary cancellation, but cannot recover a server thread already wedged below the request layer.",
        "The original non-streaming campaign had no queue admission check or post-timeout cancellation proof, so later requests repeatedly waited behind abandoned single-slot work."
      ]
    },
    {
      "title": "Concurrency qualification",
      "status": "Completed",
      "details": [
        "Two identical preserved requests took 9.855 seconds sequentially and 9.797 seconds with client concurrency two.",
        "Concurrency two provided no throughput gain on the one-slot server.",
        "The second concurrent request's time to first activity increased from approximately 0.64 seconds to 5.135 seconds.",
        "The smallest reliable scheduling policy is therefore client concurrency one; increasing parallelism only adds queue latency and abandoned-work risk."
      ]
    },
    {
      "title": "Bounded request instrumentation",
      "status": "Implemented but not production-qualified",
      "details": [
        "Added an explicit policy covering input/output bounds, concurrency, connect/write/read/pool/queue/total budgets, bounded retries, streaming, cancellation verification, and accounting.",
        "Added normalized terminal reasons for success, active timeout, model unavailability, queue timeout, schema invalidity, response invalidity, cancellation, and runtime failure.",
        "Added exact prompt/config, request, response, and policy hashes plus start/end, prompt/output token, queue, first-activity, generation, model, and runtime telemetry.",
        "Added pre-submit slot-idle checks and post-cancellation slot verification. Cleanup has an independently enforced overall deadline even when the slot endpoint itself hangs.",
        "Added a Phase 3F-CE runner that verifies the exact frozen corpus hash and count, stops after the existing validator, and explicitly denies every downstream authority."
      ]
    },
    {
      "title": "Minimal-repair experiments",
      "status": "No sufficient repair found",
      "details": [
        "Prompt-cache disable and explicit slot reset did not reliably prevent the hard hang.",
        "Moving strict-schema enforcement out of the server while retaining the unchanged schema for independent client validation did not reliably prevent the hard hang.",
        "A documented CUDA-graph workaround for Qwen3.8 on RTX 5090-class hardware did not reliably prevent this local failure and was reverted.",
        "Because the failure survived those bounded experiments, further timeout inflation, cache tuning, schema changes, or backend substitution would be speculative. The first-divergence rule required stopping."
      ]
    }
  ],
  "decisions": [
    "Treat the exact frozen 20 sources as the only qualification corpus and perform no new discovery.",
    "Rule out prompt size, ordinary latency, and client concurrency with measured controls before changing the request contract.",
    "Use concurrency one because the one-slot runtime showed equal throughput and worse per-request latency at concurrency two.",
    "Distinguish ordinary streamed cancellation from a hard server hang; never claim server cancellation based only on a client exception.",
    "Preserve each failed qualification attempt and its first divergent boundary rather than rewriting or replacing evidence.",
    "Do not merge or activate the qualification instrumentation because the real 20-source acceptance test did not pass.",
    "Restore the original production Qwen runtime and verify model, slot, Hiro health, and revision state before ending the session.",
    "Do not start a new Phase 3F campaign until a runtime implementation passes the same two-source reproducer and the complete frozen corpus."
  ],
  "validation": [
    {
      "check": "Focused claim-runtime and source-claim tests",
      "status": "passed",
      "result": "Eight focused tests passed in 0.57 seconds, including terminal taxonomy, policy bounds, schema checks, sequential admission, and a cancellation-cleanup deadline test."
    },
    {
      "check": "Complete Hiro repository suite",
      "status": "passed",
      "result": "892 tests passed, one expected test skipped, and zero tests failed in 454.66 seconds. Six existing unknown-mark warnings were reported."
    },
    {
      "check": "Clean single-source extraction control",
      "status": "passed",
      "result": "The first preserved source completed in 7.59 seconds with 783 prompt tokens, 136 output tokens, and valid structured output."
    },
    {
      "check": "Long preserved extraction control",
      "status": "passed once",
      "result": "The second preserved source completed in one clean control in 61.025 seconds with 1,011 prompt tokens, 1,942 output tokens, and valid structured output, proving the prompt is individually executable."
    },
    {
      "check": "Frozen-corpus qualification",
      "status": "failed at first unresolved divergence",
      "result": "Twenty sources were received. Two extraction requests were attempted, one completed and one hard-hung/timed out; 18 were intentionally not attempted after the first unresolved divergence. Zero raw and zero validated claims were produced before the stop."
    },
    {
      "check": "Production restoration",
      "status": "passed",
      "result": "The original Qwen runtime was restored, a strict structured probe returned HTTP 200 in 0.867 seconds, the execution slot was idle, and Hiro health reported connected Qwen with matching loaded and checkout revisions."
    },
    {
      "check": "Public journal tests and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. The first build generated and validated 186 entries, then stopped because the clean clone had no installed TypeScript toolchain. After npm ci installed the lockfile-defined packages, npm run build regenerated and validated all 186 entries, compiled TypeScript, and completed the Vite production bundle. Installation reported one existing high-severity dependency advisory outside this qualification's scope."
    }
  ],
  "currentState": [
    "Phase 3F-CE is terminal with disposition PHASE 3F-CE NOT DEMONSTRATED — MODEL LATENCY/REQUEST CONTRACT FAILURE.",
    "The preserved qualification funnel is 20 sources received, two extraction requests, one success, one timeout/hard hang, and 18 not attempted after the first divergence.",
    "The bounded instrumentation exists only on the isolated codex/phase3fce-qualification branch at commit 37ccb85380b5259a24c3ed7368ad810619261d6e; it is not active production code.",
    "Production Hiro remains unchanged on revision 19ff77a5b09cf3b76712a4b5eef3d6617089fdd9 with Qwen connected and the runtime slot idle."
  ],
  "limitations": [
    "The current llama-server implementation can enter a hard generation hang that request-layer streaming and cancellation cannot terminate.",
    "A complete 20-source qualification was not obtained, so the candidate bounded request contract is not approved for production use.",
    "The evidence does not assess source merit, claim quality, relevance, feasibility, reproduction, or candidate quality because those stages were intentionally not reached.",
    "The exact low-level runtime defect is consistent with open upstream llama.cpp failure reports, but no locally sufficient backend repair was demonstrated in this session.",
    "The first qualification-runner attempt also exposed and fixed a Windows-only diagnostic defect: a POSIX-style process-existence probe terminated Qwen on Windows. That failed attempt and correction were preserved separately."
  ],
  "nextSteps": [
    "Run a separate bounded runtime-implementation comparison using the same Qwen 3.8 model and the same two preserved source requests.",
    "Require repeated completion, responsive slot telemetry, and real termination of an intentionally cancelled request before selecting a replacement runtime implementation or backend build.",
    "After a runtime passes the two-source reproducer, rerun the exact unchanged 20-source Phase 3F-CE corpus.",
    "Do not begin a new fresh Phase 3F campaign until the 20-source claim-extraction and validator boundary completes under the qualified runtime."
  ],
  "disclosureNote": "This public entry contains no credentials, private source text, personal interaction content, private filesystem locations, private network addresses, or actionable unresolved security details."
}
