Executive summary
This session continued the existing claim-extraction model qualification without rerunning models already tested. It evaluated the two remaining locally installed candidates: Qwen 3.5 9B and Gemma 4 12B instruction QAT.
Both candidates used the exact frozen two-source reliability workload, extraction prompt, response schema, 30-second progress-stall rule, exact-process recovery, one-retry maximum, and 5% hard-hang ceiling used by the prior qualification.
Qwen 3.5 9B failed in cycle 2 when one source request hard-hung on both its original attempt and its sole retry. The accepted evidence window contains four requests, five attempts, three valid completions, two hard hangs, one terminal failure, and a best streak of one cycle.
Gemma 4 12B completed four successful two-source cycles, but two of ten attempts hard-hung. Even six perfect remaining cycles would have produced a 9.09% hard-hang rate at the required ten-cycle boundary, so the frozen 5% target had become mathematically unreachable.
Quality testing was not authorized for either model. The frozen 20-source corpus was not executed, no production extraction route was changed, and no discovery, candidate, governor, promotion, or Phase 3F work occurred.
The production Qwen 3.8 worker was restored with its exact model and configuration. Its first cold-start inference check hung; the existing exact-process replacement mechanism then restored a healthy worker that passed ownership, API, model identity, real inference, and idle-slot verification.
The final cross-model disposition is MIXED MODEL BOTTLENECK: one previously tested candidate passed runtime but failed extraction quality, while the remaining installed candidates failed runtime reliability. No local model is qualified for claim extraction.