Hiro development journal

Bounded multi-claim parsing fixes the merge but Qwen remains five of six

Final bounded Qwen repair stopped at the unchanged gold gate Machine-readable JSON

Executive summary

Hiro's last authorized repair to the grounded Qwen extractor changed only Stage 2 cardinality. Stage 1, the model, llama.cpp backend, frozen corpora, canonical claim schema, validators, runtime thresholds, and production routing remained unchanged.

Stage 2 can now return zero, one, two, or three delimited claim blocks from one verified evidence span. Each block retains the same seven grounded fields, missing information remains absent, duplicate assertions are deterministically removed, and no server-side constrained JSON is used.

The repair solved the exact prior defect: the formerly merged evidence span produced two separately structured, provenance-valid claims. However, Qwen returned NO_VALID_CLAIMS for a different known-positive span about logical planning and constrained execution.

The gold result therefore remained five of six canonical claims. All five produced claims were provenance-valid, with zero unsupported claims, zero fabricated fields, zero request failures, and zero parser failures.

The critical stop rule was applied. The frozen six-source quality corpus, repeated runtime campaign, frozen twenty-source corpus, production routing, and Phase 3F were not run.

Work completed

Bounded Stage 2 multi-claim contract

Completed
  • Stage 2 now returns exactly NO_VALID_CLAIMS or one to three CLAIM_START and CLAIM_END blocks. Each block contains intervention, comparison, outcome, metric, conditions, direction, and evidence type in a fixed order.
  • The maximum is three candidate claims per verified span. The output budget increased from 192 to 384 tokens, the smallest frozen allowance selected for up to three short grounded blocks without returning to the old 3,072-token generation pattern.
  • Stage 1 remained byte-for-byte frozen: its contract hash is 2a51511ad2e2dbaedb6a27b2da49c204af0fd50933d1373e6668efbb1f6fa3ae, with a 384-token budget, eight-span maximum, 512-character span limit, exact matching, immutable source hashes, and locally derived offsets.
  • The new Stage 2 contract hash is d1f82b95e9abf44dcf8fb267b7c595ea090ebf70b27dc79496e81aa56c144ee4.

Parser, grounding, and deduplication

Completed
  • The parser accepts only the explicit zero token or complete claim blocks containing all seven labels in the frozen order. Responses exceeding three blocks or containing malformed delimiters are invalid.
  • Every nonempty material field must still be an exact substring of the parent verified span. Direction and evidence type retain their frozen enumerations; missing fields remain null and no comparison, metric, outcome, or condition is inferred.
  • Assertions are deduplicated using normalized canonical material fields rather than generated wording or a model judgment. A duplicate block is retained in diagnostics but cannot become a canonical claim.
  • Each surviving claim is sent independently through Hiro's unchanged deterministic claim assembler and independent semantic validator.

Frozen gold qualification

Failed
  • All three positive gold sources and all six known evidence fragments were recovered through five verified Stage 1 spans. Stage 1 had no request or parser failures.
  • The prior merged span was successfully repaired: Qwen emitted two distinct claim blocks, one for reducing low-capability trajectory information and one for removing high-capability trajectory information. Both passed deterministic grounding and independent validation.
  • A different verified positive span concerning logical planning and safe or efficient execution under resource constraints returned NO_VALID_CLAIMS. This left the aggregate at five canonical claims instead of the required six.
  • The final gold metrics were five raw claim blocks, five canonical claims, five provenance-valid claims, zero unsupported claims, zero fabricated fields, zero duplicate claims, zero request failures, zero parser failures, and zero validator runtime failures.
  • The first divergence is VERIFIED_SPAN_TO_BOUNDED_MULTI_CLAIM_GOLD_GATE. This is a remaining Stage 2 false negative, not the previous cardinality defect.

Observed runtime and stop enforcement

Passed within the stopped gold run
  • The gold run made three Stage 1 requests and five Stage 2 requests. All eight completed on their first attempt with zero hard hangs, stalls, retries, or terminal failures.
  • Stage 1 mean latency was 2.097960 seconds with a 2.268775-second p95. Stage 2 mean latency was 1.976355 seconds with a 2.789420-second p95.
  • Because the gold semantic gate failed, the six-source quality corpus and ten-cycle repeated runtime campaign were not authorized. Their results are unavailable rather than inferred.
  • The frozen twenty-source corpus was not cleared, production extraction routing stayed disabled, and Phase 3F was not started.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Stage 1 immutability passed The Stage 1 contract hash remained exactly 2a51511ad2e2dbaedb6a27b2da49c204af0fd50933d1373e6668efbb1f6fa3ae.
Focused multi-claim tests passed Thirty-three focused grounding, parser, supervisor, validator, and model-qualification tests passed, including zero, multiple, duplicate, over-limit, and malformed-field cases.
Frozen gold semantic gate failed Five of six canonical gold claims were recovered. The prior merged claim split correctly, but another known-positive span produced NO_VALID_CLAIMS.
Frozen six-source quality qualification not run Withheld under the critical stop rule because the gold prerequisite failed. Zero-claim correctness was not rescored.
Repository regression suite passed The complete repository suite finished with 954 passed, one expected skip, and six existing unknown-mark warnings in 429.06 seconds.
Public journal tests and production build passed Timestamped-entry tests passed, 202 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully.

Current state

Next steps