{
  "schemaVersion": 2,
  "date": "2026.08.30",
  "publishedAt": "2026-08-30T21:22:38-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Mistral Q6_K extractor fails the frozen startup boundary",
  "publicationStatus": "Exact artifact qualification, failure evidence, cleanup repair, repository validation, and central-runtime restoration complete",
  "executiveSummary": [
    "The exact Bartowski Mistral Small 3.2 24B Instruct 2506 Q6_K GGUF became available and was qualified as the sole newly authorized claim-extraction candidate.",
    "The 19,345,944,704-byte artifact was independently hashed as 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874 before execution and again by the qualification runner before its immutable manifest was created.",
    "Mistral used the unchanged llama.cpp 2.31.2 configuration, full-GPU-offload request, flash attention, 8,192-token context, one slot, frozen extraction contract, supervisor policy, and startup boundary.",
    "The candidate never became API-ready within the frozen 300-second startup policy. It therefore failed at MODEL_START_TO_API_READY with normalized reason MODEL_RELOAD_FAILED before any extraction request was submitted.",
    "Because no request ran, this result contains zero requests, zero attempts, zero generation hangs, zero retries, and zero completed cycles. It is a startup failure and is not mislabeled as the hard-generation-hang pattern seen in prior llama.cpp candidates.",
    "Quality testing was not authorized. No production routing changed, the frozen twenty-source corpus was not run, and no other model was tested.",
    "The failed startup exposed a cleanup gap for processes still loading before API readiness. Exact pre-listen cleanup was added and used to terminate only the verified candidate process, after which central Qwen was restored and passed real inference health.",
    "Final disposition for this bounded candidate: NO QUALIFIED EXTRACTION MODEL — RUNTIME BOTTLENECK."
  ],
  "workstreams": [
    {
      "title": "Exact artifact identity",
      "status": "Completed",
      "details": [
        "The accepted artifact identity is Mistral Small 3.2 24B Instruct 2506, Bartowski Q6_K GGUF.",
        "The exact file size is 19,345,944,704 bytes and its SHA-256 is 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.",
        "The previously found Q4_K_S artifact was not substituted, retested, or treated as equivalent.",
        "The runner independently reverified the model hash before stopping the central worker or creating runtime evidence."
      ]
    },
    {
      "title": "Frozen runtime qualification",
      "status": "Rejected at startup",
      "details": [
        "The runtime command retained llama.cpp 2.31.2, one slot, an 8,192-token context, full GPU-layer offload request, flash attention, memory mapping, no multimodal projector, and the existing batch configuration.",
        "The exactly owned central model was verified and temporarily released before candidate launch.",
        "Mistral remained in model loading and did not expose a healthy API within the unchanged 300-second startup boundary.",
        "The first failed boundary was MODEL_START_TO_API_READY and the normalized terminal reason was MODEL_RELOAD_FAILED.",
        "No reliability source was submitted: requests 0, attempts 0, valid completions 0, hard hangs 0, retries 0, successful retries 0, and best consecutive cycles 0 of 10.",
        "Observed device use while the loading candidate was resident was 28,123 MiB. Per-process attribution was unavailable under Windows WDDM."
      ]
    },
    {
      "title": "Startup-failure cleanup",
      "status": "Completed",
      "details": [
        "The runner previously armed abort cleanup only after model startup returned successfully. A process that timed out while still loading could therefore remain resident and block central restoration.",
        "The minimal repair extends qualification-only cleanup across the launch-attempt interval without changing startup duration or serving policy.",
        "Pre-listen cleanup requires the frozen configuration hash, model hash, exact PID command, unique matching process, and either no port owner or the same PID as port owner. It fails closed on ambiguity.",
        "The Mistral process was terminated with exact evidence, its process was confirmed gone, and its candidate port was confirmed unowned."
      ]
    },
    {
      "title": "Production restoration and evidence",
      "status": "Completed",
      "details": [
        "The first automatic central reload could not allocate GPU memory because the timed-out candidate was still resident. After exact candidate cleanup, the identical central Qwen model was launched again.",
        "Central Qwen passed process ownership, port ownership, API health, model identity, structured inference, and idle-slot checks before admission reopened.",
        "The immutable Mistral report records that the candidate is absent, quality was unauthorized, production routing is disabled, and the twenty-source corpus is not cleared."
      ]
    }
  ],
  "decisions": [
    "Treat startup readiness as part of runtime qualification and retain the frozen 300-second policy rather than granting a model-specific extension.",
    "Do not call the result a generation hang because generation was never reached.",
    "Do not run the six-source quality workload for a runtime-failed model.",
    "Do not retest prior models, activate routing, run the frozen twenty-source corpus, or start another model automatically.",
    "Repair only the exact cleanup gap demonstrated by the failed launch."
  ],
  "validation": [
    {
      "check": "Exact Q6_K file and hash",
      "status": "passed",
      "result": "The 19,345,944,704-byte artifact matched SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874 in both the preliminary and runner-owned verification passes."
    },
    {
      "check": "Mistral runtime startup",
      "status": "failed acceptance",
      "result": "The worker did not become API-ready within the frozen startup policy; MODEL_RELOAD_FAILED occurred before any request."
    },
    {
      "check": "Frozen extraction quality",
      "status": "not authorized",
      "result": "Runtime did not qualify, so no six-source extraction or independent validation request was run."
    },
    {
      "check": "Focused qualification and supervisor tests",
      "status": "passed",
      "result": "Twenty-one focused tests passed after the exact pre-listen cleanup repair."
    },
    {
      "check": "Complete Hiro repository suite",
      "status": "passed",
      "result": "911 tests passed, one expected test skipped, and zero tests failed in 449.43 seconds. Six existing unknown-mark warnings were reported."
    },
    {
      "check": "Final production runtime",
      "status": "passed",
      "result": "Central Qwen is READY with admission open after exact identity and real inference verification; the Mistral worker is absent."
    }
  ],
  "currentState": [
    "Mistral Small 3.2 24B Q6_K is not qualified for claim extraction under the frozen runtime contract.",
    "No Mistral quality result exists because quality was correctly blocked.",
    "Production extraction routing remains disabled and unchanged.",
    "The frozen twenty-source corpus is not cleared for requalification.",
    "Central Qwen remains Hiro's active, healthy central model."
  ],
  "limitations": [
    "This run establishes failure to meet the current startup policy; it does not provide evidence about Mistral's extraction fidelity or generation-hang incidence.",
    "A different inference backend or separately authorized startup-policy qualification could investigate whether Mistral can load reliably, but that would be a new task and cannot be inferred from this result.",
    "Per-process GPU-memory usage was unavailable, so only the observed device total is reported."
  ],
  "nextSteps": [
    "Preserve the immutable startup-failure result and do not repeat this exact comparison without a specifically authorized runtime or backend change.",
    "If Mistral remains a desired candidate, qualify inference-backend/startup behavior separately before resuming extraction-model selection.",
    "Do not run the frozen quality or twenty-source workloads until a candidate passes runtime qualification."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, private source text, personal data, or actionable unresolved security details."
}
