Hiro development journal

Mistral Q6_K startup reaches a conclusive GPU-memory failure

900-second startup/backend qualification, immutable telemetry, production restoration, and repository validation complete Machine-readable JSON

Executive summary

The earlier 300-second Mistral startup cutoff was correctly treated as inconclusive for runtime reliability because no generation request had run and large local models can take substantially longer to load.

A narrowly scoped startup/backend qualification retested only the exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact. The llama.cpp command, model identity, context, slot count, GPU-offload request, flash attention, extraction semantics, inference bounds, and downstream policies were unchanged; only the diagnostic startup ceiling increased to 900 seconds.

The runner captured process, CPU, device-memory, log, port, and API state every ten seconds. All forty-five checkpoints showed changing progress signals, so loading was not stalled.

Mistral acquired its port and returned HTTP 503 while loading. Device use rose to approximately 29.4 GiB and CPU time continued increasing.

At approximately 836.5 seconds, llama.cpp attempted a final 228.01 MiB CUDA allocation for context compute buffers. The allocation failed because device memory was exhausted, and the process exited. The monitor observed exit at 849.676 seconds.

The normalized cause is CUDA_OUT_OF_MEMORY_DURING_CONTEXT_INITIALIZATION. The final disposition is MISTRAL STARTUP NOT QUALIFIED — LOAD FAILURE, not load stall and not post-load inference failure.

API readiness was never reached, so model-identity verification through the API, tiny inference, the frozen structured-output control, extraction quality, and the twenty-source corpus did not run.

The candidate was confirmed absent, central Qwen was restored and passed real inference health, production routing remained disabled, and no other model was tested.

Work completed

Reproducible startup/backend diagnostic

Completed
  • The diagnostic accepts only the frozen Mistral Q6_K file and SHA-256 and independently verifies the complete artifact before launch.
  • It changes the candidate startup ceiling from 300 to 900 seconds without changing the generated llama.cpp command or model/configuration identity.
  • It polls active progress every ten seconds and declares a load stall only after 180 seconds with no process, CPU, device-memory, log, port, or API change.
  • If readiness occurs, the same runner requires exact model identity, the existing tiny structured health inference, one frozen structured-output control, and a final idle slot. Those post-load checks were correctly skipped because readiness never occurred.

Observed Mistral load

Failed conclusively
  • The exact 19,345,944,704-byte Q6_K artifact retained SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.
  • Forty-five immutable checkpoints were captured through 849.676 seconds, with progress present at every checkpoint.
  • The worker stayed alive and accumulated CPU time, acquired its configured port, returned HTTP 503 while loading, and used approximately 29.1 to 29.4 GiB of device memory near the end.
  • llama.cpp then failed a 228.01 MiB CUDA allocation needed for compute buffers while initializing the context.
  • The server emitted a concrete model-load error and exited. No request entered generation.

Failure classification and evidence

Completed
  • The first conclusive boundary was ACTIVE_MODEL_LOADING to CONTEXT_COMPUTE_BUFFER_ALLOCATION to CUDA_OUT_OF_MEMORY to PROCESS_EXITED_DURING_LOAD.
  • The terminal disposition is MISTRAL STARTUP NOT QUALIFIED — LOAD FAILURE.
  • The result is not LOAD STALL because measurable progress continued until the server emitted an allocation error.
  • The result is not POST-LOAD INFERENCE FAILURE because API readiness and inference were never reached.
  • A normalized sidecar binds CUDA_OUT_OF_MEMORY_DURING_CONTEXT_INITIALIZATION to the immutable qualification report and the hashed server log without publishing raw operational details.

Cleanup and production restoration

Completed
  • The candidate had exited by the time exact pre-listen cleanup ran. Cleanup verified that the process was gone and the candidate port was unowned.
  • Central Qwen was restored with its identical model and command.
  • Central health passed process ownership, port ownership, API response, model identity, real structured inference, and final idle-slot checks before admission reopened.
  • No production extraction route, discovery component, candidate path, governor, promotion mechanism, or corpus state changed.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Exact artifact identity passed The full 19,345,944,704-byte artifact matched SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874 before execution.
Active-load progress monitoring passed Forty-five immutable checkpoints were recorded; all forty-five contained a changed progress signal and no 180-second stall occurred.
API readiness and post-load inference failed prerequisite The process exited during context initialization before healthy API readiness; no inference or frozen structured control ran.
Focused startup, qualification, supervisor, runtime, and source-claim tests passed Twenty-four focused tests passed after the final exception-consumption and load-error classification changes.
Complete Hiro repository suite passed 914 tests passed, one expected test skipped, and zero tests failed in 451.95 seconds. Six existing unknown-mark warnings were reported.
Final production runtime passed Central Qwen is READY with admission open after exact identity, real inference, and idle-slot verification; the Mistral process is absent.

Current state

Next steps