Hiro development journal

GLiNER2 is runtime-stable but fails Hiro's gold claim contract

Dedicated CPU extraction qualified operationally; semantic qualification stopped at the gold-span gate Machine-readable JSON

Executive summary

Hiro's next dedicated claim-extraction comparison tested only fastino/gliner2-large-v1, pinned to immutable revision 6a498b5a28ec3908bbc5277aeb47d22bcfc02f33. Production routing, discovery, candidate construction, the frozen twenty-source corpus, and Phase 3F remained disabled.

The official GLiNER2 1.3.1 local interface ran in a fresh Python 3.11 CPU-only environment with PyTorch 2.13.0+cpu and Transformers 5.13.1. The verified twelve-file model snapshot totals 1,962,226,926 bytes and has aggregate manifest SHA-256 67f4e2f14add7ae904076b1b014d71e568395da77706998537cf4d916545d6fa.

The native smoke test passed, including exact entity offsets and identical repeat inference. A full workload reliability campaign then completed ten consecutive cycles over all sixteen frozen spans from the six-source qualification set: 160/160 requests completed, with zero hangs, crashes, or within-process output changes.

Semantic qualification nevertheless failed at GOLD_SPAN_INPUT_TO_CANONICAL_CLAIM. GLiNER2 produced seven native structures for six known claim spans, but only two claims satisfied Hiro's unchanged canonical and independent semantic validation contract. Four expected claims were missed, and five structures were unusable because required outcome or intervention/comparison fields were absent.

Native span integrity was strong: zero fabricated fields and no offset mismatch were observed. The result is therefore an extraction recall/field-completeness limitation, not a provenance fabrication failure. The six-source quality stage was correctly withheld because its gold-span prerequisite failed.

Work completed

Pinned model and isolated CPU runtime

Completed
  • The qualification used fastino/gliner2-large-v1 at revision 6a498b5a28ec3908bbc5277aeb47d22bcfc02f33 and did not test the base model or another GLiNER variant.
  • The official model.safetensors file is 1,945,828,140 bytes with SHA-256 92a76e84cd4de59e15e3f6577bef9e4304929667551ee053665eba365510638e. The complete twelve-file manifest was hashed independently.
  • The isolated environment used Python 3.11.9, GLiNER2 1.3.1, PyTorch 2.13.0+cpu, and Transformers 5.13.1. No CUDA, quantization, torch.compile, or Hiro application dependencies were used by the extractor process.
  • One Windows console preflight failed before model loading because GLiNER2 printed Unicode through CP-1252. The minimum process-only repair enabled UTF-8 standard I/O; model, backend, schema, thresholds, and data remained unchanged.

Native adapter and authority boundary

Completed
  • The adapter consumes only preverified immutable source spans and calls GLiNER2's native structured-extraction interface with include_spans and include_confidence enabled.
  • Native fields map to Hiro's intervention or mechanism, comparison or baseline, outcome, metric or observable, conditions, direction, and evidence type. Missing values remain absent; the adapter never completes them.
  • Every native text field must have integer offsets that reproduce the exact substring of the supplied immutable span. Hiro attaches the original source identity, source hash, global span offsets, and exact supporting text outside the model.
  • A structure cannot become a valid Hiro claim without an outcome and either an intervention or a comparison. Hiro's existing independent Qwen semantic validator remains the final authority for source support and field agreement.

Native smoke and sustained reliability

Passed
  • The official entity-extraction smoke recovered Apple, Tim Cook, iPhone 15, and Cupertino with exact character offsets and identical repeated output.
  • The clean evidence run loaded the CPU model in 46.036 seconds and used 2,447,175,680 bytes RSS after load. The smoke completed in 2.473 seconds at 2,473,455,616 bytes RSS.
  • Reliability replayed all sixteen frozen spans from all six qualification sources ten times. All 160 requests completed without timeout, native crash, process replacement, or terminal failure.
  • Mean extraction latency was 0.614448 seconds and p95 was 0.644621 seconds. End-of-run RSS was only 4,218,880 bytes above the post-smoke baseline, well inside the frozen 256 MiB leak guard.
  • Outputs were identical across all ten cycles within the clean process. Separate fresh process launches did vary one extracted outcome field, so cross-load semantic determinism remains a limitation even though each bounded runtime was operationally stable.

Gold-span semantic gate

Failed
  • The unchanged gold set contains six independently established claims across three positive sources. GLiNER2 proposed seven native structures and recovered two of the six required canonical claims.
  • Four expected claims were false negatives. Five of seven structures were unusable, predominantly because GLiNER2 extracted a method and metric but omitted the explicit outcome; one omitted both intervention and comparison.
  • No material field had an invalid native offset or text absent from the frozen evidence span. Fabricated-field count was zero.
  • The two surviving structures passed Hiro's independent source-support, no-added-facts, field-agreement, and testability-coherence checks.
  • Because the gold gate failed, the frozen six-source quality evaluation and its legitimate zero-claim controls were not run. No zero-claim result is claimed from this session.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Model identity and artifact integrity passed The pinned revision, twelve model files, individual hashes, total bytes, and aggregate manifest hash were frozen before qualification.
Native CPU smoke test passed The model loaded on CPU, returned exact spans for four official-example entities, repeated identically, and exited cleanly.
Gold-span semantic qualification failed Two of six expected claims survived; four were missed, five of seven proposed structures were unsupported as canonical claims, and fabricated-field count was zero.
Frozen six-source quality qualification not run Withheld because the prerequisite gold-span gate failed. Legitimate zero-claim behavior was therefore not scored.
Repeated CPU reliability passed Ten consecutive full-corpus cycles completed: 160/160 span extractions, zero crashes, zero hangs, zero within-process output changes, 0.614448 second mean, and 0.644621 second p95.
Focused claim-extraction tests passed Twenty-two focused tests passed, including the new native-offset, missing-field, zero-result, and unchanged downstream contract checks.
Repository regression suite passed The complete repository suite finished with 945 passed, one expected skip, and six existing unknown-mark warnings in 406.39 seconds.
Public journal tests and production build passed Timestamped-entry tests passed, 200 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully.
Authority boundary passed Production routing stayed disabled; the twenty-source corpus, Phase 3F, candidates, promotion, CUDA optimization, and alternate models were not invoked.

Current state

Next steps