Executive summary
The earlier 300-second Mistral startup cutoff was correctly treated as inconclusive for runtime reliability because no generation request had run and large local models can take substantially longer to load.
A narrowly scoped startup/backend qualification retested only the exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact. The llama.cpp command, model identity, context, slot count, GPU-offload request, flash attention, extraction semantics, inference bounds, and downstream policies were unchanged; only the diagnostic startup ceiling increased to 900 seconds.
The runner captured process, CPU, device-memory, log, port, and API state every ten seconds. All forty-five checkpoints showed changing progress signals, so loading was not stalled.
Mistral acquired its port and returned HTTP 503 while loading. Device use rose to approximately 29.4 GiB and CPU time continued increasing.
At approximately 836.5 seconds, llama.cpp attempted a final 228.01 MiB CUDA allocation for context compute buffers. The allocation failed because device memory was exhausted, and the process exited. The monitor observed exit at 849.676 seconds.
The normalized cause is CUDA_OUT_OF_MEMORY_DURING_CONTEXT_INITIALIZATION. The final disposition is MISTRAL STARTUP NOT QUALIFIED — LOAD FAILURE, not load stall and not post-load inference failure.
API readiness was never reached, so model-identity verification through the API, tiny inference, the frozen structured-output control, extraction quality, and the twenty-source corpus did not run.
The candidate was confirmed absent, central Qwen was restored and passed real inference health, production routing remained disabled, and no other model was tested.