Hiro development journal

Evidence-first extraction improves abstention but loses positive claims

Protocol qualification complete; production routing remains disabled Machine-readable JSON

Executive summary

Hiro's generation-first claim-extraction contract was replaced in a qualification-only pathway with a two-stage evidence-first protocol: exact supporting-span selection followed by independent span-bound structuring.

The clean frozen six-source campaign completed using the already runtime-qualified Mistral Small 3.2 24B Q6_K configuration. Seventeen spans were proposed, sixteen passed exact-offset verification, and all Stage A and Stage B model responses were schema-valid.

Grounding improved in one important respect: all three legitimate zero-claim controls returned zero validated claims, compared with one of three under the preserved generation-first result.

The protocol nevertheless failed the unchanged quality gate. It recovered zero of three positive sources, produced only two raw structured claims, and both violated the span-to-structure contract, leaving zero independently validated claims. The evidence supports a mixed bottleneck involving positive-claim recall and span-to-structure contract failure.

Mistral was unloaded, central Qwen was restored healthy and idle, and no production routing, twenty-source corpus execution, Phase 3F activity, candidate construction, or promotion occurred.

Work completed

Evidence-first protocol

Completed
  • Stage A presents the complete immutable source and bounded sentence-like candidates, then permits explicit zero-span output or exact supporting spans with source-relative character offsets.
  • A deterministic verifier rejects invalid offsets, empty or overlong spans, text mismatches, duplicate spans, and any proposed text absent from the source.
  • Stage B processes each verified span independently. Its nullable contract records intervention or mechanism, comparison, outcome, metric, conditions, direction, and evidence type without requiring absent semantics to be invented.
  • Every material non-null field must be an exact substring of its supporting span. Claim-bearing outputs retain the exact span, offsets, source hash, and immutable request evidence.
  • Zero claims arise naturally when Stage A selects nothing, every span is rejected, or Stage B marks or validates no span as claim-bearing.

Frozen six-source qualification

Failed
  • The campaign used the same three positive and three legitimate zero-claim controls and the same independent semantic expectations used by the prior Qwen 2.5 Coder and Mistral generation-first qualifications.
  • Stage A proposed seventeen spans; sixteen were deterministically valid, for a 94.118 percent span-verification rate. One proposal was rejected for mutated text and invalid offsets.
  • All six Stage A responses and all sixteen Stage B responses were schema-valid, with no independent-validator runtime failures.
  • The protocol produced two raw claim-bearing structures. Both carried a rejection reason while also declaring a claim, violating the frozen span-structure contract; both were deterministically rejected before semantic acceptance.
  • Zero of three positive controls recovered an independently validated claim, all three positive controls were false negatives, and zero claims survived independent validation.
  • All three legitimate zero-claim controls were handled correctly, but selected candidate spans in each negative control show that Stage A precision was not independently improved even though downstream abstention prevented false claims.

Baseline comparison and diagnosis

Completed
  • The preserved generation-first baseline recovered two of three positive sources, generated fourteen raw claims, retained six independently validated claims, produced eight unsupported or fabricated claims, achieved 42.857 percent provenance validity, and handled one of three zero-claim sources correctly.
  • Evidence-first extraction reduced accepted unsupported claims by rejecting all raw structures, and improved zero-claim correctness from one of three to three of three, but this was achieved with unacceptable suppression: positive-source recovery fell from two of three to zero of three.
  • The final disposition is EVIDENCE-FIRST CLAIM EXTRACTION NOT QUALIFIED — MIXED. The observed signals are span-selection or structuring recall loss and span-to-structure contract failure, not model startup, hard-hang, schema, independent-validator, or production-runtime failure.

First-divergence repair and clean rerun

Completed
  • The first attempt reached a successful Stage A request and a returned Stage B response, then exposed that the shared local JSON-schema checker could not evaluate a nullable type union.
  • The minimum infrastructure repair added standards-compatible type-union and null handling to that checker without changing extraction semantics, prompt content, corpus content, thresholds, retry policy, or validation policy.
  • The campaign was restarted from a clean run directory. Only the clean rerun is used for the reported qualification metrics.
  • Mistral loaded in 596.517 seconds under progress-aware startup supervision, completed the six-source protocol, unloaded cleanly, and central Qwen was restored in 467.299 seconds.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused protocol and runtime tests passed Fifteen focused tests passed, covering exact offsets, mutated-span rejection, natural zero-span behavior, nullable fields, material-field grounding, schema type unions, and strict qualification truth tables.
Repository regression suite passed The full suite completed with 928 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 387.15 seconds.
Frozen evidence-first quality campaign failed Seventeen spans proposed, sixteen offset-valid, two raw structured claims, zero validated claims, zero of three positive controls recovered, three of three zero-claim controls correct, and two unsupported contract-invalid raw claims.
Runtime restoration passed Mistral was unloaded and central Qwen was restored with verified process ownership, model identity, API health, bounded inference, and an idle slot.
Journal tests and production build passed The timestamped-entry tests passed, 196 journal pages generated and validated, and the TypeScript/Vite production build passed.

Current state

Next steps