{
  "schemaVersion": 2,
  "date": "2026.09.01",
  "publishedAt": "2026-09-01T16:20:09-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "GLiNER2 is runtime-stable but fails Hiro's gold claim contract",
  "publicationStatus": "Dedicated CPU extraction qualified operationally; semantic qualification stopped at the gold-span gate",
  "executiveSummary": [
    "Hiro's next dedicated claim-extraction comparison tested only fastino/gliner2-large-v1, pinned to immutable revision 6a498b5a28ec3908bbc5277aeb47d22bcfc02f33. Production routing, discovery, candidate construction, the frozen twenty-source corpus, and Phase 3F remained disabled.",
    "The official GLiNER2 1.3.1 local interface ran in a fresh Python 3.11 CPU-only environment with PyTorch 2.13.0+cpu and Transformers 5.13.1. The verified twelve-file model snapshot totals 1,962,226,926 bytes and has aggregate manifest SHA-256 67f4e2f14add7ae904076b1b014d71e568395da77706998537cf4d916545d6fa.",
    "The native smoke test passed, including exact entity offsets and identical repeat inference. A full workload reliability campaign then completed ten consecutive cycles over all sixteen frozen spans from the six-source qualification set: 160/160 requests completed, with zero hangs, crashes, or within-process output changes.",
    "Semantic qualification nevertheless failed at GOLD_SPAN_INPUT_TO_CANONICAL_CLAIM. GLiNER2 produced seven native structures for six known claim spans, but only two claims satisfied Hiro's unchanged canonical and independent semantic validation contract. Four expected claims were missed, and five structures were unusable because required outcome or intervention/comparison fields were absent.",
    "Native span integrity was strong: zero fabricated fields and no offset mismatch were observed. The result is therefore an extraction recall/field-completeness limitation, not a provenance fabrication failure. The six-source quality stage was correctly withheld because its gold-span prerequisite failed."
  ],
  "workstreams": [
    {
      "title": "Pinned model and isolated CPU runtime",
      "status": "Completed",
      "details": [
        "The qualification used fastino/gliner2-large-v1 at revision 6a498b5a28ec3908bbc5277aeb47d22bcfc02f33 and did not test the base model or another GLiNER variant.",
        "The official model.safetensors file is 1,945,828,140 bytes with SHA-256 92a76e84cd4de59e15e3f6577bef9e4304929667551ee053665eba365510638e. The complete twelve-file manifest was hashed independently.",
        "The isolated environment used Python 3.11.9, GLiNER2 1.3.1, PyTorch 2.13.0+cpu, and Transformers 5.13.1. No CUDA, quantization, torch.compile, or Hiro application dependencies were used by the extractor process.",
        "One Windows console preflight failed before model loading because GLiNER2 printed Unicode through CP-1252. The minimum process-only repair enabled UTF-8 standard I/O; model, backend, schema, thresholds, and data remained unchanged."
      ]
    },
    {
      "title": "Native adapter and authority boundary",
      "status": "Completed",
      "details": [
        "The adapter consumes only preverified immutable source spans and calls GLiNER2's native structured-extraction interface with include_spans and include_confidence enabled.",
        "Native fields map to Hiro's intervention or mechanism, comparison or baseline, outcome, metric or observable, conditions, direction, and evidence type. Missing values remain absent; the adapter never completes them.",
        "Every native text field must have integer offsets that reproduce the exact substring of the supplied immutable span. Hiro attaches the original source identity, source hash, global span offsets, and exact supporting text outside the model.",
        "A structure cannot become a valid Hiro claim without an outcome and either an intervention or a comparison. Hiro's existing independent Qwen semantic validator remains the final authority for source support and field agreement."
      ]
    },
    {
      "title": "Native smoke and sustained reliability",
      "status": "Passed",
      "details": [
        "The official entity-extraction smoke recovered Apple, Tim Cook, iPhone 15, and Cupertino with exact character offsets and identical repeated output.",
        "The clean evidence run loaded the CPU model in 46.036 seconds and used 2,447,175,680 bytes RSS after load. The smoke completed in 2.473 seconds at 2,473,455,616 bytes RSS.",
        "Reliability replayed all sixteen frozen spans from all six qualification sources ten times. All 160 requests completed without timeout, native crash, process replacement, or terminal failure.",
        "Mean extraction latency was 0.614448 seconds and p95 was 0.644621 seconds. End-of-run RSS was only 4,218,880 bytes above the post-smoke baseline, well inside the frozen 256 MiB leak guard.",
        "Outputs were identical across all ten cycles within the clean process. Separate fresh process launches did vary one extracted outcome field, so cross-load semantic determinism remains a limitation even though each bounded runtime was operationally stable."
      ]
    },
    {
      "title": "Gold-span semantic gate",
      "status": "Failed",
      "details": [
        "The unchanged gold set contains six independently established claims across three positive sources. GLiNER2 proposed seven native structures and recovered two of the six required canonical claims.",
        "Four expected claims were false negatives. Five of seven structures were unusable, predominantly because GLiNER2 extracted a method and metric but omitted the explicit outcome; one omitted both intervention and comparison.",
        "No material field had an invalid native offset or text absent from the frozen evidence span. Fabricated-field count was zero.",
        "The two surviving structures passed Hiro's independent source-support, no-added-facts, field-agreement, and testability-coherence checks.",
        "Because the gold gate failed, the frozen six-source quality evaluation and its legitimate zero-claim controls were not run. No zero-claim result is claimed from this session."
      ]
    }
  ],
  "decisions": [
    "Classify the result as GLINER2 NOT QUALIFIED — EXTRACTION QUALITY.",
    "Treat GOLD_SPAN_INPUT_TO_CANONICAL_CLAIM as the first failed boundary; do not reinterpret operational reliability as semantic qualification.",
    "Do not loosen Hiro's requirement for an explicit outcome plus intervention or comparison merely to accommodate GLiNER2's partial structures.",
    "Do not run the frozen six-source quality corpus after the failed gold prerequisite, and do not clear the frozen twenty-source corpus.",
    "Do not test a smaller GLiNER2 model, optimize with CUDA, quantize, compile, activate production routing, or resume Phase 3F automatically."
  ],
  "validation": [
    {
      "check": "Model identity and artifact integrity",
      "status": "passed",
      "result": "The pinned revision, twelve model files, individual hashes, total bytes, and aggregate manifest hash were frozen before qualification."
    },
    {
      "check": "Native CPU smoke test",
      "status": "passed",
      "result": "The model loaded on CPU, returned exact spans for four official-example entities, repeated identically, and exited cleanly."
    },
    {
      "check": "Gold-span semantic qualification",
      "status": "failed",
      "result": "Two of six expected claims survived; four were missed, five of seven proposed structures were unsupported as canonical claims, and fabricated-field count was zero."
    },
    {
      "check": "Frozen six-source quality qualification",
      "status": "not run",
      "result": "Withheld because the prerequisite gold-span gate failed. Legitimate zero-claim behavior was therefore not scored."
    },
    {
      "check": "Repeated CPU reliability",
      "status": "passed",
      "result": "Ten consecutive full-corpus cycles completed: 160/160 span extractions, zero crashes, zero hangs, zero within-process output changes, 0.614448 second mean, and 0.644621 second p95."
    },
    {
      "check": "Focused claim-extraction tests",
      "status": "passed",
      "result": "Twenty-two focused tests passed, including the new native-offset, missing-field, zero-result, and unchanged downstream contract checks."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The complete repository suite finished with 945 passed, one expected skip, and six existing unknown-mark warnings in 406.39 seconds."
    },
    {
      "check": "Public journal tests and production build",
      "status": "passed",
      "result": "Timestamped-entry tests passed, 200 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully."
    },
    {
      "check": "Authority boundary",
      "status": "passed",
      "result": "Production routing stayed disabled; the twenty-source corpus, Phase 3F, candidates, promotion, CUDA optimization, and alternate models were not invoked."
    }
  ],
  "currentState": [
    "GLiNER2 large is operationally qualified for stable local CPU inference but is not semantically qualified as Hiro's claim extractor.",
    "The authoritative failure is incomplete semantic field extraction on known positive spans, not source-offset fabrication or runtime instability.",
    "The frozen six-source quality corpus was not consumed, the frozen twenty-source corpus is not cleared, and production extraction routing remains disabled.",
    "Hiro's production Qwen runtime remained on its existing process and was used only as the unchanged independent validator for structurally valid gold claims."
  ],
  "limitations": [
    "Only fastino/gliner2-large-v1 at the pinned revision and the official GLiNER2 1.3.1 CPU stack were tested.",
    "The gold result measures GLiNER2's ability to structure already verified claim-bearing spans; it does not establish end-to-end source-span selection recall.",
    "Because the gold prerequisite failed, six-source positive recovery, legitimate zero-claim correctness, and the frozen corpus-wide provenance rate are unavailable rather than passing or failing values.",
    "Fresh worker launches produced one different extracted outcome field across otherwise equivalent gold runs. Ten-cycle determinism was demonstrated within one loaded worker, not across arbitrary reloads.",
    "The initial CP-1252 console failure was an environment integration defect corrected only with UTF-8 process I/O; it was not a model or inference result."
  ],
  "nextSteps": [
    "Keep production claim-extraction routing disabled and preserve this exact evidence for comparing any separately authorized architecture.",
    "If dedicated extraction continues, select one new architecture or a separately justified schema-aware system under the same frozen gold contract rather than weakening Hiro's canonical claim requirements.",
    "Do not test a smaller GLiNER2 variant for an accuracy failure; the authorized large model already demonstrates the architecture's current semantic limitation on Hiro's evidence.",
    "Do not run the frozen six-source or twenty-source corpora until a candidate first passes all six gold claims without fabricated fields."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, private source text, personal data, or actionable unresolved security details."
}
