{
  "schemaVersion": 2,
  "date": "2026.08.31",
  "publishedAt": "2026-08-31T15:35:26-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Mistral Q6_K startup reaches a conclusive GPU-memory failure",
  "publicationStatus": "900-second startup/backend qualification, immutable telemetry, production restoration, and repository validation complete",
  "executiveSummary": [
    "The earlier 300-second Mistral startup cutoff was correctly treated as inconclusive for runtime reliability because no generation request had run and large local models can take substantially longer to load.",
    "A narrowly scoped startup/backend qualification retested only the exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact. The llama.cpp command, model identity, context, slot count, GPU-offload request, flash attention, extraction semantics, inference bounds, and downstream policies were unchanged; only the diagnostic startup ceiling increased to 900 seconds.",
    "The runner captured process, CPU, device-memory, log, port, and API state every ten seconds. All forty-five checkpoints showed changing progress signals, so loading was not stalled.",
    "Mistral acquired its port and returned HTTP 503 while loading. Device use rose to approximately 29.4 GiB and CPU time continued increasing.",
    "At approximately 836.5 seconds, llama.cpp attempted a final 228.01 MiB CUDA allocation for context compute buffers. The allocation failed because device memory was exhausted, and the process exited. The monitor observed exit at 849.676 seconds.",
    "The normalized cause is CUDA_OUT_OF_MEMORY_DURING_CONTEXT_INITIALIZATION. The final disposition is MISTRAL STARTUP NOT QUALIFIED — LOAD FAILURE, not load stall and not post-load inference failure.",
    "API readiness was never reached, so model-identity verification through the API, tiny inference, the frozen structured-output control, extraction quality, and the twenty-source corpus did not run.",
    "The candidate was confirmed absent, central Qwen was restored and passed real inference health, production routing remained disabled, and no other model was tested."
  ],
  "workstreams": [
    {
      "title": "Reproducible startup/backend diagnostic",
      "status": "Completed",
      "details": [
        "The diagnostic accepts only the frozen Mistral Q6_K file and SHA-256 and independently verifies the complete artifact before launch.",
        "It changes the candidate startup ceiling from 300 to 900 seconds without changing the generated llama.cpp command or model/configuration identity.",
        "It polls active progress every ten seconds and declares a load stall only after 180 seconds with no process, CPU, device-memory, log, port, or API change.",
        "If readiness occurs, the same runner requires exact model identity, the existing tiny structured health inference, one frozen structured-output control, and a final idle slot. Those post-load checks were correctly skipped because readiness never occurred."
      ]
    },
    {
      "title": "Observed Mistral load",
      "status": "Failed conclusively",
      "details": [
        "The exact 19,345,944,704-byte Q6_K artifact retained SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.",
        "Forty-five immutable checkpoints were captured through 849.676 seconds, with progress present at every checkpoint.",
        "The worker stayed alive and accumulated CPU time, acquired its configured port, returned HTTP 503 while loading, and used approximately 29.1 to 29.4 GiB of device memory near the end.",
        "llama.cpp then failed a 228.01 MiB CUDA allocation needed for compute buffers while initializing the context.",
        "The server emitted a concrete model-load error and exited. No request entered generation."
      ]
    },
    {
      "title": "Failure classification and evidence",
      "status": "Completed",
      "details": [
        "The first conclusive boundary was ACTIVE_MODEL_LOADING to CONTEXT_COMPUTE_BUFFER_ALLOCATION to CUDA_OUT_OF_MEMORY to PROCESS_EXITED_DURING_LOAD.",
        "The terminal disposition is MISTRAL STARTUP NOT QUALIFIED — LOAD FAILURE.",
        "The result is not LOAD STALL because measurable progress continued until the server emitted an allocation error.",
        "The result is not POST-LOAD INFERENCE FAILURE because API readiness and inference were never reached.",
        "A normalized sidecar binds CUDA_OUT_OF_MEMORY_DURING_CONTEXT_INITIALIZATION to the immutable qualification report and the hashed server log without publishing raw operational details."
      ]
    },
    {
      "title": "Cleanup and production restoration",
      "status": "Completed",
      "details": [
        "The candidate had exited by the time exact pre-listen cleanup ran. Cleanup verified that the process was gone and the candidate port was unowned.",
        "Central Qwen was restored with its identical model and command.",
        "Central health passed process ownership, port ownership, API response, model identity, real structured inference, and final idle-slot checks before admission reopened.",
        "No production extraction route, discovery component, candidate path, governor, promotion mechanism, or corpus state changed."
      ]
    }
  ],
  "decisions": [
    "Do not interpret the 300-second cutoff as model unreliability; replace it with direct startup evidence.",
    "Allow active loading to continue within the 900-second bound while CPU, memory, port, API, or log evidence demonstrates progress.",
    "Stop at the concrete CUDA allocation failure rather than extending time further, because additional time cannot create missing device memory.",
    "Do not reduce context, batch allocation, GPU offload, or quantization within this task because the existing serving command was part of the controlled variable set.",
    "Do not run inference, extraction quality, the twenty-source corpus, another model, or production routing after startup failed."
  ],
  "validation": [
    {
      "check": "Exact artifact identity",
      "status": "passed",
      "result": "The full 19,345,944,704-byte artifact matched SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874 before execution."
    },
    {
      "check": "Active-load progress monitoring",
      "status": "passed",
      "result": "Forty-five immutable checkpoints were recorded; all forty-five contained a changed progress signal and no 180-second stall occurred."
    },
    {
      "check": "API readiness and post-load inference",
      "status": "failed prerequisite",
      "result": "The process exited during context initialization before healthy API readiness; no inference or frozen structured control ran."
    },
    {
      "check": "Focused startup, qualification, supervisor, runtime, and source-claim tests",
      "status": "passed",
      "result": "Twenty-four focused tests passed after the final exception-consumption and load-error classification changes."
    },
    {
      "check": "Complete Hiro repository suite",
      "status": "passed",
      "result": "914 tests passed, one expected test skipped, and zero tests failed in 451.95 seconds. Six existing unknown-mark warnings were reported."
    },
    {
      "check": "Final production runtime",
      "status": "passed",
      "result": "Central Qwen is READY with admission open after exact identity, real inference, and idle-slot verification; the Mistral process is absent."
    }
  ],
  "currentState": [
    "Mistral Small 3.2 24B Q6_K is not startup-qualified under the unchanged full-GPU-offload serving configuration.",
    "The conclusive blocker is insufficient GPU memory for the final context compute-buffer allocation, not elapsed startup time.",
    "No Mistral inference, structured-control, extraction-quality, or repeated runtime result exists.",
    "Production routing remains disabled and the frozen twenty-source corpus remains untouched.",
    "Central Qwen remains Hiro's active, healthy central model."
  ],
  "limitations": [
    "This result does not evaluate Mistral's response quality or generation reliability because the backend never reached usable API readiness.",
    "The diagnostic intentionally did not test a smaller quantization, partial CPU offload, smaller context or batch allocation, or another backend.",
    "Device-wide memory telemetry was available, but Windows WDDM did not provide reliable per-process attribution."
  ],
  "nextSteps": [
    "Do not rerun the same command with a longer timeout; the observed failure is resource exhaustion, not a timeout.",
    "If Mistral remains desired, separately qualify one bounded resource/backend change such as partial CPU offload, a smaller quantization, reduced context/batch allocation, or another inference backend.",
    "Require startup, exact identity, tiny inference, structured control, and idle-slot success before resuming repeated claim-extraction qualification.",
    "Continue blocking quality, production routing, and the twenty-source corpus until the startup/backend prerequisite passes."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, private source text, personal data, or actionable unresolved security details."
}
