Hiro development journal

Grounded Qwen extraction removes hangs but merges a gold claim

Runtime qualified; semantic qualification stopped at the one-claim-per-span gold gate Machine-readable JSON

Executive summary

Hiro's final bounded extractor attempt tested the already-running qwen/qwen3.8-27b only. No model was installed, no inference backend was changed, and production routing, the frozen twenty-source corpus, and Phase 3F remained disabled.

The attempt replaced server-constrained generation with a two-stage plain-text evidence contract. Stage 1 could emit only exact delimited source excerpts or an explicit zero-span token. Stage 2 structured one deterministically verified excerpt into seven bounded fields, with material values accepted only when they were exact substrings of that excerpt.

Runtime reliability improved decisively. Ten consecutive cycles completed 60 Stage 1 and 60 Stage 2 requests with zero hard hangs, zero generation stalls, zero retries, and zero terminal failures. Stage 1 averaged 1.350156 seconds and Stage 2 averaged 1.432633 seconds.

Semantic qualification nevertheless failed at VERIFIED_SPANS_TO_ONE_CLAIM_PER_SPAN_GOLD_GATE. Qwen covered all six known gold evidence fragments but returned five verified spans because one excerpt merged two distinct claims. The resulting five canonical claims were provenance-valid and independently validated, with zero unsupported claims and zero fabricated fields, but the frozen contract requires six separately recoverable claims.

The frozen six-source quality qualification was correctly withheld after the failed gold prerequisite. Legitimate zero-claim behavior was therefore not scored, the twenty-source corpus is not cleared, and no production behavior changed.

Work completed

Plain-text evidence contracts

Completed
  • Stage 1 instructs the model to copy exact contiguous claim-bearing excerpts between CLAIM_SPAN_START and CLAIM_SPAN_END delimiters, one claim per span, or return exactly NO_CLAIM_SPANS. Its frozen contract hash is 2a51511ad2e2dbaedb6a27b2da49c204af0fd50933d1373e6668efbb1f6fa3.
  • Stage 2 receives exactly one verified span and emits seven single-line fields: intervention, comparison, outcome, metric, conditions, direction, and evidence type. Missing information must remain blank. Its frozen contract hash is 3545d87d9bc2869a32ce5112532fd4ee34ed52d372d8c11c699c90f1f0e797e0.
  • No server-side constrained JSON or grammar was used in either model stage. The existing supervised streaming transport was minimally extended to return plain text while retaining progress observation, cancellation, bounded retry, ownership checks, and immutable attempt evidence.
  • Output budgets were 384 tokens for Stage 1 and 192 tokens for Stage 2. Stage 1 allowed at most eight excerpts per source and 512 characters per excerpt.

Deterministic provenance and field grounding

Completed
  • Hiro independently recomputes the immutable source hash, requires every proposed excerpt to be nonempty, bounded, unique, and present exactly once in the source, and derives offsets locally rather than accepting model-provided offsets.
  • Stage 2 accepts exactly seven ordered lines. Each nonempty material field must be an exact contiguous substring of the verified excerpt; direction and evidence type are limited to frozen enumerations. Missing fields remain null and no defaults are supplied.
  • The unchanged canonical claim assembler and independent semantic validator remain downstream authorities. A claim cannot survive if the structured evidence violates source support, adds material facts, or fails field agreement.
  • Focused tests cover plain-text transport, exact parsing, explicit zero-span output, missing-field preservation, invalid grounded fields, and the distinction between overbroad containment and multi-claim merging.

Ten-cycle runtime qualification

Passed
  • Ten consecutive cycles each exercised all six frozen full-source Stage 1 requests and all six frozen gold-span Stage 2 requests, for 120 model requests total.
  • Stage 1 completed 60/60 valid requests in 60 attempts: zero hard hangs, zero stalls, zero retries, zero terminal failures, 1.350156-second mean latency, and 1.782458-second p95 latency.
  • Stage 2 completed 60/60 valid requests in 60 attempts: zero hard hangs, zero stalls, zero retries, zero terminal failures, 1.432633-second mean latency, and 1.776331-second p95 latency.
  • The same owned Qwen process, PID 15316, was healthy and idle before and after qualification. No runtime recovery or process replacement occurred.
  • The preserved one-pass Qwen baseline had 80 requests, 105 attempts, 38 hard hangs, 25 retries, 13 terminal failures, 22.77-second mean latency, and 54.54-second p95 latency. The grounded protocol therefore resolved the observed runtime fragility under this bounded workload.

Gold semantic gate

Failed
  • The immutable gold set contains six canonical claims across three positive sources. Qwen's Stage 1 output contained five deterministically valid excerpts covering all six gold evidence fragments.
  • Three excerpts were exact gold boundaries and two were overbroad containment spans. One of the overbroad spans combined two separately expected directional findings into one excerpt, violating the one-claim-per-span contract.
  • Five canonical claims were recovered from the five excerpts. All five had exact-span provenance and passed independent semantic validation. Unsupported-claim count and fabricated-field count were both zero.
  • There were no false spans, missed gold evidence fragments, Stage 1 request failures, or Stage 2 request failures. The failure is claim-segmentation granularity and canonical recall: five separately structured claims instead of six.
  • Because the gold gate failed, the frozen six-source quality corpus was not run. Positive-source recovery and legitimate zero-claim correctness for that corpus are unavailable, not implicitly passing or failing.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Contract immutability and authority passed Stage contracts, hashes, budgets, frozen corpus hashes, model identity, disabled production authority, and unchanged quality gates were recorded before execution.
Ten-cycle runtime qualification passed 120/120 plain-text requests completed in 120 attempts with zero hard hangs, stalls, retries, recoveries, or terminal failures.
Gold-span semantic qualification failed Five verified excerpts yielded five valid canonical claims while the frozen gold contract requires six; one excerpt merged two claims. Fabricated and unsupported counts were zero.
Frozen six-source quality qualification not run Withheld because the prerequisite gold semantic gate failed. Zero-claim correctness was not scored.
Focused grounded-extraction tests passed Twenty-eight focused extraction, supervisor, evidence-first, and model-qualification tests passed in 0.70 seconds.
Repository regression suite passed The complete repository suite finished with 951 passed, one expected skip, and six existing unknown-mark warnings in 387.64 seconds.
Public journal tests and production build passed Timestamped-entry tests passed, 201 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully.

Current state

Next steps