Executive summary
Hiro's next dedicated claim-extraction comparison tested only fastino/gliner2-large-v1, pinned to immutable revision 6a498b5a28ec3908bbc5277aeb47d22bcfc02f33. Production routing, discovery, candidate construction, the frozen twenty-source corpus, and Phase 3F remained disabled.
The official GLiNER2 1.3.1 local interface ran in a fresh Python 3.11 CPU-only environment with PyTorch 2.13.0+cpu and Transformers 5.13.1. The verified twelve-file model snapshot totals 1,962,226,926 bytes and has aggregate manifest SHA-256 67f4e2f14add7ae904076b1b014d71e568395da77706998537cf4d916545d6fa.
The native smoke test passed, including exact entity offsets and identical repeat inference. A full workload reliability campaign then completed ten consecutive cycles over all sixteen frozen spans from the six-source qualification set: 160/160 requests completed, with zero hangs, crashes, or within-process output changes.
Semantic qualification nevertheless failed at GOLD_SPAN_INPUT_TO_CANONICAL_CLAIM. GLiNER2 produced seven native structures for six known claim spans, but only two claims satisfied Hiro's unchanged canonical and independent semantic validation contract. Four expected claims were missed, and five structures were unusable because required outcome or intervention/comparison fields were absent.
Native span integrity was strong: zero fabricated fields and no offset mismatch were observed. The result is therefore an extraction recall/field-completeness limitation, not a provenance fabrication failure. The six-source quality stage was correctly withheld because its gold-span prerequisite failed.