Executive summary
Hiro's evidence-first claim extraction was decomposed into independent supporting-span selection and gold-span structuring tests using the same three frozen positive controls and Mistral Small 3.2 24B Q6_K.
Six previously independently validated historical claims were recovered from immutable qualification artifacts. Their exact supporting excerpts were located uniquely in the frozen sources, assigned exact offsets, frozen before inference, and preserved under one control-set hash.
Stage A was exonerated: all three positive sources achieved SPAN_RECALL_PASS, and every historical supporting span was completely covered by at least one selected span.
The original Stage B contract failed all six gold claims: five claim-bearing outputs were contract-invalid because the required nullable rejection field was emitted as an empty string, and one known-good span was rejected as incomplete. The resulting component diagnosis was STRUCTURING_BOTTLENECK.
A minimal tagged-union repair made claim and reject responses structurally exclusive. On the versioned component rerun, Stage B improved to four of six independently validated passes; the remaining two chose the claim branch but omitted both the required outcome and intervention or comparison.
A further prompt-only clarification could not receive a semantic result: one model load exited at 99 percent and the single retry reached API readiness but failed bounded post-load inference health before Stage A. Those infrastructure failures do not alter Mistral's prior 20-of-20 runtime qualification or the last completed semantic result.
The frozen six-source corpus was not rerun because Stage B did not pass all gold controls. Production routing, the twenty-source corpus, Phase 3F, model selection, candidates, and promotion remained untouched.