{
  "schemaVersion": 2,
  "date": "2026.08.31",
  "publishedAt": "2026-08-31T19:35:15-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Mistral Q6_K passes extraction reliability but fails frozen quality",
  "publicationStatus": "Qualification complete; production routing remains disabled",
  "executiveSummary": [
    "The exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact completed Hiro's frozen claim-extraction runtime and quality qualification through LM Studio without changing the extraction prompt, schema, semantic policy, thresholds, retry policy, or corpora.",
    "The runtime stage passed: ten consecutive two-source cycles produced twenty valid completions with zero hard hangs, retries, replacements, ordinary failures, or terminal failures.",
    "The frozen six-source quality stage did not pass. All six outputs were schema-valid, but only two of three positive controls recovered an independently validated claim, eight of fourteen raw claims were unsupported or fabricated, and only one of three legitimate zero-claim controls returned zero claims.",
    "The resulting disposition is NO QUALIFIED EXTRACTION MODEL — QUALITY BOTTLENECK. Mistral was unloaded, central Qwen was restored healthy, production extraction routing stayed disabled, and the frozen twenty-source integration corpus was not cleared."
  ],
  "workstreams": [
    {
      "title": "LM Studio qualification adapter",
      "status": "Completed",
      "details": [
        "A qualification-only adapter connected the existing frozen claim-extraction runner to LM Studio while retaining the runner's real request, retry, persistence, health, and independent-validation mechanics.",
        "The adapter validates LM Studio 0.4.21+2, the selected llama.cpp CUDA 2.31.2 backend, the exact authorized artifact hash, context 4,096, one parallel slot, flash attention, mmap plus memory locking, maximum GPU offload, batch 2,048, and microbatch 512.",
        "Startup is monitored through immutable progress checkpoints and permits up to 900 seconds while the backend remains alive and progress continues; it does not reinterpret 300 or 600 seconds as failure.",
        "A pre-inference attempt exposed two adapter-only preflight defects: an incorrectly passed application-version path and Windows decoding of the runtime-selection marker. No Mistral process or request ran in that attempt. Qwen was restored, the two narrow defects were corrected, and the campaign restarted from clean state.",
        "The clean run reused the already-recorded authoritative artifact verification instead of hashing the same 19.35 GB file a second time after Qwen shutdown."
      ]
    },
    {
      "title": "Frozen runtime reliability",
      "status": "Passed",
      "details": [
        "The authorized 19,345,944,704-byte Q6_K artifact matched SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.",
        "Mistral became API-ready in 586.166 seconds. The final startup checkpoint observed 99 percent progress, increasing backend activity, and approximately 21.9 GiB device memory use before readiness.",
        "Twenty of twenty requests completed with valid structured outputs across ten consecutive two-source cycles.",
        "Hard hangs, retries, successful retries, runtime replacements, ordinary failures, and terminal failures were all zero. Mean request latency was 28.079 seconds and p95 latency was 34.147 seconds.",
        "Every post-cycle check verified the exact process, model identity, API health, real bounded inference, an idle slot, and no orphan backend."
      ]
    },
    {
      "title": "Frozen six-source extraction quality",
      "status": "Failed",
      "details": [
        "Mistral generated schema-valid outputs for all six controls, and no validator runtime request failed.",
        "The controls contained three positive sources and three legitimate zero-claim sources.",
        "Two of three positive sources recovered at least one independently validated claim.",
        "Fourteen raw claims were extracted; six survived independent semantic and provenance validation, while eight were classified as unsupported or fabricated.",
        "Only one of three legitimate zero-claim sources was handled correctly.",
        "The validated-claim provenance rate was 42.857 percent under the frozen calculation. These observations violate the unchanged quality criteria requiring all positive controls recovered, zero unsupported claims, and all zero-claim controls correct."
      ]
    },
    {
      "title": "Cleanup and authority boundaries",
      "status": "Completed",
      "details": [
        "Mistral was unloaded after its six extraction outputs were frozen.",
        "Central Qwen was restored in 481.334 seconds and independently validated the frozen Mistral outputs.",
        "The final central-runtime check verified process ownership, API and model identity, real inference health, and an idle slot.",
        "No production routing, discovery, candidate construction, promotion, twenty-source execution, or Phase 3F activity was authorized or performed."
      ]
    }
  ],
  "decisions": [
    "Accept the runtime stage as qualified for the exact Q6_K artifact and frozen LM Studio configuration.",
    "Reject the model for claim-extraction quality under the existing frozen criteria; do not tune the prompt or relax thresholds in response.",
    "Keep production claim-extraction routing disabled.",
    "Do not clear or execute the frozen twenty-source integration corpus because the six-source quality gate failed.",
    "Do not test another model or begin Phase 3F without separate authorization."
  ],
  "validation": [
    {
      "check": "Focused adapter and qualification tests",
      "status": "passed",
      "result": "Twenty-three focused tests passed after the narrow adapter repairs."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The full repository suite completed with 918 passed, one skipped because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 395.58 seconds."
    },
    {
      "check": "Frozen runtime reliability",
      "status": "passed",
      "result": "Ten consecutive two-source cycles and twenty requests passed with zero hangs, retries, replacements, or failures."
    },
    {
      "check": "Frozen extraction quality",
      "status": "failed",
      "result": "Six of six schema-valid outputs yielded six validated claims from fourteen raw claims, eight unsupported or fabricated claims, two of three positive controls recovered, and one of three zero-claim controls correct."
    },
    {
      "check": "Final runtime restoration",
      "status": "passed",
      "result": "Mistral was removed and central Qwen finished healthy, correctly identified, and idle."
    },
    {
      "check": "Journal tests and production build",
      "status": "passed",
      "result": "The timestamped-entry tests passed, 195 journal pages generated and validated, and the TypeScript/Vite production build passed."
    }
  ],
  "currentState": [
    "Mistral Small 3.2 24B Q6_K is empirically runtime-reliable under the frozen LM Studio configuration but is not a qualified claim extractor under the frozen quality contract.",
    "Central Qwen is restored and healthy.",
    "Production extraction routing is disabled.",
    "The frozen twenty-source Phase 3F-CE corpus remains uncleared and unexecuted.",
    "Phase 3F has not started."
  ],
  "limitations": [
    "The quality result applies to the exact frozen prompt, schema, six controls, independent validator, and Q6_K runtime configuration; no prompt tuning was attempted.",
    "The aggregate quality failure does not imply runtime unreliability: runtime and quality produced distinct, independently recorded outcomes.",
    "The qualification adapter is new, although its focused tests and the full 918-test repository regression suite passed."
  ],
  "nextSteps": [
    "Preserve this evidence as the baseline for any separately authorized extractor strategy.",
    "Do not activate Mistral for production claim extraction under the present frozen contract.",
    "If future work is authorized, address the quality bottleneck explicitly rather than repeating model startup qualification or loosening the established criteria.",
    "Require separate authorization before using the frozen twenty-source corpus or resuming Phase 3F."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, source corpus text, personal data, or actionable unresolved security details."
}
