{
  "schemaVersion": 2,
  "date": "2026.08.31",
  "publishedAt": "2026-08-31T18:15:25-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Mistral Q6_K qualifies in LM Studio at a 4K context",
  "publicationStatus": "Direct LM Studio load, identity, inference, structured-output, persistence, cleanup, and central-runtime restoration complete",
  "executiveSummary": [
    "A narrowly scoped LM Studio runtime qualification tested the exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact without changing Hiro code, extraction prompts, qualification thresholds, production routing, or Qwen's central-model role.",
    "The LM Studio-managed Q6_K file was independently confirmed as the authorized 19,345,944,704-byte artifact with SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.",
    "LM Studio 0.4.21+2 used its selected llama.cpp CUDA 2.31.2 runtime. Qwen was stopped through its ownership-verified supervisor to release GPU memory before the test.",
    "Q6_K was loaded with an explicitly requested 4,096-token context, one parallel slot, flash attention, mmap, and maximum GPU offload. LM Studio's saved per-model configuration also contained 4,096 context and did not override the explicit request.",
    "The backend progressed continuously and completed loading in 9 minutes 36.39 seconds. The earlier approximately ten-minute client boundary was therefore not a model failure; it was close to the model's real readiness time on this machine.",
    "LM Studio reported the model resident under the requested runtime identifier with context 4,096, parallelism one, and an idle slot. Device memory settled near 21,999 MiB during the first inference.",
    "A tiny deterministic request returned exactly READY. A JSON-schema-constrained request returned exactly the required READY status and numeric value. A subsequent independent request returned exactly HEALTHY, and the slot returned to idle with zero queued work after every request.",
    "Q4_K_S was not tested because the preferred Q6_K configuration succeeded without partial CPU offload. The test model was then deliberately unloaded, LM Studio's local API was stopped, and Qwen was restored as Hiro's healthy central runtime."
  ],
  "workstreams": [
    {
      "title": "LM Studio and artifact inventory",
      "status": "Completed",
      "details": [
        "LM Studio version 0.4.21+2 and CLI version 1.3.3 were observed.",
        "llama.cpp-win-x86_64-nvidia-cuda12-avx2 2.31.2 was selected in LM Studio.",
        "Both Q6_K and Q4_K_S were installed. The indexed Q6_K resolved to the exact authorized file and hash.",
        "LM Studio initially had no model loaded and its local API server was stopped. Hiro's independently managed Qwen worker was the only llama-server process.",
        "The only saved Q6_K per-model load setting was context length 4,096; no saved GPU-offload value conflicted with the test configuration."
      ]
    },
    {
      "title": "Monitored Q6_K load",
      "status": "Passed",
      "details": [
        "Qwen was terminated only after exact ownership and port checks proved the intended worker, freeing GPU memory for the isolated LM Studio attempt.",
        "The load command explicitly requested the installed Q6_K identifier, context length 4,096, parallelism one, and maximum GPU offload.",
        "The observed backend command confirmed context 4,096, one slot, flash attention enabled, maximum GPU layers, GPU KV-cache offload, batch 2,048, microbatch 512, and mmap with memory locking.",
        "LM Studio estimated 20.83 GB for the model plus 888.01 MB for context, or 21.72 GB total, against 31.28 GB free immediately before loading.",
        "The backend remained alive and accumulated CPU time while the GUI and CLI percentages advanced. No CUDA out-of-memory, backend crash, or load-stall evidence occurred.",
        "Backend readiness completed at approximately 9 minutes 35 seconds, and the CLI load transaction completed successfully at 9 minutes 36.39 seconds."
      ]
    },
    {
      "title": "Post-load runtime qualification",
      "status": "Passed",
      "details": [
        "A brief persistence check observed the same loaded identifier and idle state before and after the interval.",
        "The API exposed the exact assigned Mistral qualification identifier.",
        "The tiny deterministic request returned exactly READY in 0.478 seconds and finished normally.",
        "The JSON-schema response-format request returned an object containing only status READY and value 1 in 0.594 seconds; independent parsing confirmed exact schema compliance.",
        "After a second persistence interval, the slot was idle with zero queued work. A subsequent request returned exactly HEALTHY in 0.299 seconds.",
        "The model remained loaded, its context remained 4,096, parallelism remained one, and exactly one LM Studio backend process existed through the final health check."
      ]
    },
    {
      "title": "Cleanup and central-runtime restoration",
      "status": "Completed",
      "details": [
        "Q4_K_S was not loaded because Q6_K met every runtime criterion at full GPU offload.",
        "The Q6_K instance was deliberately unloaded after evidence collection, and LM Studio's temporary local API server was stopped.",
        "The unload left no LM Studio llama-server process and returned device use to its pretest baseline.",
        "Hiro's original Qwen configuration was relaunched through the ownership-aware supervisor.",
        "Qwen reached READY after 468.737 seconds with admission open, exact process and port ownership, API and model identity verified, real inference healthy, and the slot idle."
      ]
    }
  ],
  "decisions": [
    "Do not move either model file: LM Studio's indexed Q6_K is already the exact authorized artifact, and the larger catalog figure includes associated resources rather than identifying a different GGUF.",
    "Treat backend progress and health as authoritative rather than treating a rounded percentage or a client-side readiness timeout as a terminal model result.",
    "Use the successful 4,096-context Q6_K configuration as the recommended starting point for a separately authorized claim-extraction runtime qualification.",
    "Do not test Q4_K_S because the preferred Q6_K succeeded without partial offload.",
    "Do not start repeated extraction qualification, quality evaluation, the twenty-source corpus, production routing, or Phase 3F in this session.",
    "Restore Qwen after the isolated LM Studio test so this qualification does not alter its role as Hiro's central model."
  ],
  "validation": [
    {
      "check": "Exact artifact identity",
      "status": "passed",
      "result": "The LM Studio-indexed Q6_K GGUF was 19,345,944,704 bytes and matched the authorized SHA-256 exactly."
    },
    {
      "check": "Actual 4,096-token context",
      "status": "passed",
      "result": "The saved setting, observed backend command, LM Studio loaded-model state, and post-request state all reported context 4,096."
    },
    {
      "check": "Load completion and residency",
      "status": "passed",
      "result": "The backend reached readiness without CUDA failure; the load transaction completed in 9 minutes 36.39 seconds and the model remained resident through all checks."
    },
    {
      "check": "API identity",
      "status": "passed",
      "result": "The API exposed the exact assigned Q6_K qualification identifier."
    },
    {
      "check": "Tiny deterministic inference",
      "status": "passed",
      "result": "The model returned exactly READY in 0.478 seconds with a normal stop result."
    },
    {
      "check": "Structured-output control",
      "status": "passed",
      "result": "The response-format request returned a schema-exact object with status READY and value 1 in 0.594 seconds."
    },
    {
      "check": "Post-request health",
      "status": "passed",
      "result": "The slot was idle with zero queued requests, a subsequent request returned exactly HEALTHY in 0.299 seconds, and the model remained loaded."
    },
    {
      "check": "Final production runtime",
      "status": "passed",
      "result": "The test model was deliberately removed and central Qwen was restored to READY with admission open, verified identity and ownership, healthy inference, and an idle slot."
    }
  ],
  "currentState": [
    "Mistral Small 3.2 24B Q6_K is LM Studio runtime-ready on this machine at context 4,096, one slot, flash attention, mmap, and full GPU offload.",
    "The measured successful load time is 9 minutes 36.39 seconds; future orchestration should budget above ten minutes with progress-aware monitoring rather than a fixed ten-minute failure classification.",
    "No partial GPU offload or lower quantization was required for this runtime result.",
    "Q4_K_S remains installed but untested and is not treated as equivalent evidence.",
    "No extraction-quality evidence was generated, no corpus was run, and no production routing changed.",
    "Qwen remains Hiro's active, healthy central model."
  ],
  "limitations": [
    "This qualification proves load and bounded structured inference only; it does not establish repeated claim-extraction reliability or extraction quality.",
    "LM Studio automatically loaded the associated multimodal projector and logged a non-fatal performance warning for unsupported projector operators. Text inference and schema-constrained output were unaffected in this bounded test.",
    "Peak memory was observed at device scope under Windows rather than attributed solely to the backend process.",
    "Only one load cycle and three bounded text requests were authorized; long-duration residency and repeated-cycle reliability remain untested."
  ],
  "nextSteps": [
    "Under separate authorization, run the frozen repeated claim-extraction runtime qualification using the exact Q6_K artifact and the demonstrated 4,096-context LM Studio configuration.",
    "Use progress-aware startup supervision with a justified readiness allowance above the observed 9-minute-36-second load time.",
    "Keep Q4_K_S as a fallback only if Q6_K later fails the repeated runtime or quality criteria; do not assume quantization equivalence.",
    "Continue blocking production extraction routing, quality conclusions, the twenty-source corpus, and Phase 3F until their respective qualification steps are separately authorized and passed."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, private source text, personal data, or actionable unresolved security details."
}
