Hiro development journal

The single hosted atomic-segmentation repair did not clear gold

Qualification stopped at the frozen gold gate Machine-readable JSON

Executive summary

Hiro performed the one authorized atomic-claim segmentation repair for the hosted GPT-5.6 Sol extraction path. The work began with a frozen post-hoc mapping of all ten prior emitted claims to the six gold claims, then added exactly one source-level consolidation pass between initial grounded extraction and canonical acceptance.

The diagnosis established that all six gold propositions were present in the prior run. The four actual surplus outputs were three over-split components of compound empirical findings and one non-claim background statement. A previously reported merged span contained two valid propositions but had not itself created a surplus claim.

The consolidation contract was generic: it received no gold identities, expected spans, source-specific exceptions, or expected claim count. It could merge decomposed pieces, suppress background or duplicate proposals, or split genuinely merged proposals, while deterministic Hiro checks continued to enforce exact source text, offsets, field grounding, canonical validity, provenance, and independent validation.

The authoritative frozen-gold rerun did not qualify. One source's initial extraction used blank-line-separated blocks that the unchanged strict span parser treated as an invalid response, preventing two gold propositions from reaching consolidation. Two consolidation responses also violated the new accounting contract by both using and suppressing the same proposal; one response additionally used blank separators rejected by the strict parser. The third source consolidated correctly.

The mandatory stop rule was applied. No second semantic repair was attempted, the six-source and twenty-source corpora were not run, production routing stayed disabled, and Phase 3F did not start. Final disposition: HOSTED EXTRACTION NOT QUALIFIED — ATOMIC CLAIM SEGMENTATION.

Work completed

Frozen ten-to-six diagnosis

Completed before protocol modification
  • Each of the ten prior emitted claim IDs was assigned at most one best gold correspondence, and each gold claim received at most one GOLD_EQUIVALENT output.
  • Six outputs were the best gold equivalents. Three surplus outputs were OVER_SPLIT_CLAIM components: a secondary cost component, a coordinated resource-utilization component, and a second benchmark measurement of the same intervention, comparison, and outcome.
  • The fourth surplus output was a FALSE_SPAN: a qualitative transition statement promoted despite lacking a sufficiently explicit empirical relationship.
  • No gold proposition was semantically missing in the prior run. The two propositions sharing the trajectory-interface span were represented separately, so the span-level merge did not translate into a claim-level merge.

Generic atomicity and minimal-span contract

Implemented and frozen
  • An atomic claim is one independently testable relationship linking an intervention or mechanism and applicable comparison to one primary asserted finding under material conditions.
  • Metrics, numerical values, methods, implementation details, conditions, mechanism explanations, and supporting evidence do not become separate claims merely because they occupy separate fields or sentences.
  • Coordinated outcome components expressed as one effect or tradeoff, and measurements of the same relationship on multiple benchmarks, remain one claim unless the tested intervention, comparison, primary finding, or material conditions differ.
  • Each surviving claim must use the smallest contiguous verified source span that preserves its complete assertion. Character count alone is not used as a semantic rule.

One-pass hosted consolidation

Implemented for qualification only
  • Exactly one consolidation request is allowed per source with surviving initial claims. There is no recursive revision and no third semantic pass.
  • The consolidator sees only proposed claim IDs, verified source spans, and canonical material fields. It does not receive the full gold controls, expected claim count, expected spans, or source-specific outcomes.
  • Every input claim ID must be represented in a surviving atomic claim or explicitly suppressed with a normalized reason. Unknown, unaccounted, duplicated, or simultaneously used-and-suppressed identifiers invalidate the response.
  • Hiro deterministically verifies that every final evidence span is verbatim, uniquely located, bounded, and contained in a referenced verified parent span. Material fields must remain exact substrings, duplicate signatures are rejected, and the existing local independent validator makes the final provenance decision.

Authoritative frozen-gold result

Failed
  • All ten hosted requests completed successfully on their first attempt. There were no API errors, timeouts, retries, or terminal transport failures.
  • The final gate observed one accepted canonical claim covering one of six expected evidence units. Five expected units remained uncovered, two consolidation responses were parser-invalid, and one initial Stage 1 response was parser-invalid.
  • The failed run produced three consolidated claim blocks, of which one passed all deterministic and independent validation. Two were rejected because the source-level accounting contract was invalid.
  • Unsupported accepted claims and fabricated fields remained zero. The failure was contract adherence and segmentation-path robustness, not hosted service availability or invented source content.
  • The third source successfully merged two benchmark-specific decompositions into one provenance-valid claim using a minimal exact span.

Runtime and cost accounting

Completed for the permitted gold rerun
  • The local Qwen validator was initially offline, so the first attempt stopped before any hosted inference. The existing hidden launcher restored the exact expected Qwen runtime; the authoritative attempt then began from a new immutable directory.
  • The authoritative run made three Stage 1 requests, four Stage 2 requests, and three consolidation requests. All ten completed successfully.
  • Total usage was 5,269 input tokens and 1,301 output tokens. Observed qualification cost was $0.047096, including $0.023480 for consolidation.
  • Mean hosted latency was 2.710549 seconds and p95 was 4.260472 seconds. The three-source projection was approximately $1.569867 per one hundred similarly positive-heavy sources.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen diagnostic mapping passed The ten prior claim IDs, their one-to-one best gold correspondences, surplus status, and generic classifications were persisted before protocol code changed and protected by a frozen hash.
Focused protocol tests passed Thirty-nine atomic-consolidation, hosted-adapter, grounded-extraction, model-qualification, runtime-supervisor, and evidence-first tests passed.
Repository regression suite passed The full repository suite completed with 961 passed, two expected skips, and six existing unknown-marker warnings in 552.75 seconds.
Hosted service execution passed operationally Ten of ten requests completed on their first attempt with zero API errors, ordinary timeouts, retries, or terminal failures.
Frozen gold semantic gate failed Only one of six expected evidence units reached a final accepted claim. Five were missed after one initial parser-invalid source and two consolidation-invalid sources.
Frozen six-source and twenty-source corpora not run Both were withheld under the mandatory gold-first stop rule. Zero-claim generalization was therefore not evaluated.

Current state

Next steps