Hiro development journal

Mistral Q6_K passes extraction reliability but fails frozen quality

Qualification complete; production routing remains disabled Machine-readable JSON

Executive summary

The exact Mistral Small 3.2 24B Instruct 2506 Q6_K artifact completed Hiro's frozen claim-extraction runtime and quality qualification through LM Studio without changing the extraction prompt, schema, semantic policy, thresholds, retry policy, or corpora.

The runtime stage passed: ten consecutive two-source cycles produced twenty valid completions with zero hard hangs, retries, replacements, ordinary failures, or terminal failures.

The frozen six-source quality stage did not pass. All six outputs were schema-valid, but only two of three positive controls recovered an independently validated claim, eight of fourteen raw claims were unsupported or fabricated, and only one of three legitimate zero-claim controls returned zero claims.

The resulting disposition is NO QUALIFIED EXTRACTION MODEL — QUALITY BOTTLENECK. Mistral was unloaded, central Qwen was restored healthy, production extraction routing stayed disabled, and the frozen twenty-source integration corpus was not cleared.

Work completed

LM Studio qualification adapter

Completed
  • A qualification-only adapter connected the existing frozen claim-extraction runner to LM Studio while retaining the runner's real request, retry, persistence, health, and independent-validation mechanics.
  • The adapter validates LM Studio 0.4.21+2, the selected llama.cpp CUDA 2.31.2 backend, the exact authorized artifact hash, context 4,096, one parallel slot, flash attention, mmap plus memory locking, maximum GPU offload, batch 2,048, and microbatch 512.
  • Startup is monitored through immutable progress checkpoints and permits up to 900 seconds while the backend remains alive and progress continues; it does not reinterpret 300 or 600 seconds as failure.
  • A pre-inference attempt exposed two adapter-only preflight defects: an incorrectly passed application-version path and Windows decoding of the runtime-selection marker. No Mistral process or request ran in that attempt. Qwen was restored, the two narrow defects were corrected, and the campaign restarted from clean state.
  • The clean run reused the already-recorded authoritative artifact verification instead of hashing the same 19.35 GB file a second time after Qwen shutdown.

Frozen runtime reliability

Passed
  • The authorized 19,345,944,704-byte Q6_K artifact matched SHA-256 3b1f9516b3446859f145f114152b260388253b4f911528bfe7545a79a09a8874.
  • Mistral became API-ready in 586.166 seconds. The final startup checkpoint observed 99 percent progress, increasing backend activity, and approximately 21.9 GiB device memory use before readiness.
  • Twenty of twenty requests completed with valid structured outputs across ten consecutive two-source cycles.
  • Hard hangs, retries, successful retries, runtime replacements, ordinary failures, and terminal failures were all zero. Mean request latency was 28.079 seconds and p95 latency was 34.147 seconds.
  • Every post-cycle check verified the exact process, model identity, API health, real bounded inference, an idle slot, and no orphan backend.

Frozen six-source extraction quality

Failed
  • Mistral generated schema-valid outputs for all six controls, and no validator runtime request failed.
  • The controls contained three positive sources and three legitimate zero-claim sources.
  • Two of three positive sources recovered at least one independently validated claim.
  • Fourteen raw claims were extracted; six survived independent semantic and provenance validation, while eight were classified as unsupported or fabricated.
  • Only one of three legitimate zero-claim sources was handled correctly.
  • The validated-claim provenance rate was 42.857 percent under the frozen calculation. These observations violate the unchanged quality criteria requiring all positive controls recovered, zero unsupported claims, and all zero-claim controls correct.

Cleanup and authority boundaries

Completed
  • Mistral was unloaded after its six extraction outputs were frozen.
  • Central Qwen was restored in 481.334 seconds and independently validated the frozen Mistral outputs.
  • The final central-runtime check verified process ownership, API and model identity, real inference health, and an idle slot.
  • No production routing, discovery, candidate construction, promotion, twenty-source execution, or Phase 3F activity was authorized or performed.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused adapter and qualification tests passed Twenty-three focused tests passed after the narrow adapter repairs.
Repository regression suite passed The full repository suite completed with 918 passed, one skipped because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 395.58 seconds.
Frozen runtime reliability passed Ten consecutive two-source cycles and twenty requests passed with zero hangs, retries, replacements, or failures.
Frozen extraction quality failed Six of six schema-valid outputs yielded six validated claims from fourteen raw claims, eight unsupported or fabricated claims, two of three positive controls recovered, and one of three zero-claim controls correct.
Final runtime restoration passed Mistral was removed and central Qwen finished healthy, correctly identified, and idle.
Journal tests and production build passed The timestamped-entry tests passed, 195 journal pages generated and validated, and the TypeScript/Vite production build passed.

Current state

Next steps