Executive summary
This session tested whether Hiro could assign claim extraction to a separately supervised local model while keeping Qwen 3.8 27B as Hiro's central reasoning model. It performed no new discovery, candidate construction, governor action, promotion, or meta-improvement.
The extraction prompt, JSON schema, source hashes, specificity rules, provenance rules, deterministic checks, independent semantic validator, stall boundary, and one-retry policy were frozen across models.
The prior Qwen 3.8 campaign remained the known failing reliability control. A local Qwen 3.5 4B candidate reached 20 valid requests but hard-hung on 4 of 24 attempts, a 16.67% rate above the preregistered 5% ceiling, so it never reached quality testing.
A local Qwen 2.5 Coder 7B instruction model passed the runtime boundary: 10 consecutive two-source cycles, 20 successful requests in 21 attempts, one recovered hard hang, no terminal failures, and a 4.76% hard-hang rate.
The 7B model then failed the unchanged six-source extraction-quality controls. It produced schema-valid output on 5 of 6 sources, recovered claims from 2 of 3 positive controls, fabricated claims for all 3 legitimate zero-claim controls, and had 3 of 8 raw claims rejected by deterministic or independent semantic validation.
Because no alternate model passed both prerequisites, no extraction model was assigned to production and the frozen 20-source Phase 3F-CE corpus was not run.
Two qualification-infrastructure boundaries were repaired minimally: long Windows evidence paths now use short deterministic request IDs, and the rejected extractor is released before central-model recovery. Neither repair changed prompts, model results, retries, or acceptance thresholds.
The central Qwen 3.8 worker was restored with the identical model/configuration and passed exact process, API, model-identity, inference, and idle-slot health checks.
Final disposition: EXTRACTION QUALITY BOTTLENECK. A fresh Phase 3F campaign is not justified from this evidence.