Hiro development journal

Boundary diagnosis isolates evidence-first structuring failures

Boundary diagnosis complete; evidence-first extraction remains unqualified Machine-readable JSON

Executive summary

Hiro's evidence-first claim extraction was decomposed into independent supporting-span selection and gold-span structuring tests using the same three frozen positive controls and Mistral Small 3.2 24B Q6_K.

Six previously independently validated historical claims were recovered from immutable qualification artifacts. Their exact supporting excerpts were located uniquely in the frozen sources, assigned exact offsets, frozen before inference, and preserved under one control-set hash.

Stage A was exonerated: all three positive sources achieved SPAN_RECALL_PASS, and every historical supporting span was completely covered by at least one selected span.

The original Stage B contract failed all six gold claims: five claim-bearing outputs were contract-invalid because the required nullable rejection field was emitted as an empty string, and one known-good span was rejected as incomplete. The resulting component diagnosis was STRUCTURING_BOTTLENECK.

A minimal tagged-union repair made claim and reject responses structurally exclusive. On the versioned component rerun, Stage B improved to four of six independently validated passes; the remaining two chose the claim branch but omitted both the required outcome and intervention or comparison.

A further prompt-only clarification could not receive a semantic result: one model load exited at 99 percent and the single retry reached API readiness but failed bounded post-load inference health before Stage A. Those infrastructure failures do not alter Mistral's prior 20-of-20 runtime qualification or the last completed semantic result.

The frozen six-source corpus was not rerun because Stage B did not pass all gold controls. Production routing, the twenty-source corpus, Phase 3F, model selection, candidates, and promotion remained untouched.

Work completed

Immutable positive controls

Completed
  • The controls were derived solely from the historical independently validated qualification report and immutable six-source corpus; current Mistral outputs were not used to define gold claims.
  • The three positive sources contained six validated claims: three about trajectory handoff behavior, two about resource-constrained execution, and one benchmark comparison.
  • Each historical supporting excerpt occurred exactly once in its source. Exact character offsets, the original structured fields, historical semantic validation, source hashes, and provenance decisions were frozen before testing.
  • The frozen control artifact retained the same SHA-256 across every versioned attempt, preventing prompt or schema changes from modifying the qualification targets.

Independent Stage A diagnosis

Passed
  • Stage A ran alone on the three positive sources; Stage B was not invoked as part of these measurements.
  • All three sources were classified SPAN_RECALL_PASS. There were no partial, missed, or false-positive-only source outcomes.
  • Every one of the six frozen supporting spans had complete deterministic coverage by a selected exact-offset candidate.
  • This evidence corrects the prior aggregate diagnosis: positive claims were not lost because Stage A failed to surface their source evidence.

Independent Stage B diagnosis

Failed
  • Stage B bypassed Stage A and received each of the six frozen known-good supporting spans independently.
  • Under the original schema, zero of six gold claims passed. Five outputs declared themselves claim-bearing while also supplying an empty rejection field, and one rejected a known-good span as incomplete.
  • Inspection established that the strict schema required one nullable rejection field in both accepted and rejected output states. Although deterministic validation correctly rejected the contradiction, the schema itself permitted it and encouraged empty-string null substitution.
  • The component disposition was STRUCTURING_BOTTLENECK rather than span-selection, validator, or boundary-handoff failure.

Minimal tagged-result repair

Partially successful
  • Stage B now uses a mutually exclusive tagged JSON result: a claim branch containing the canonical span-bound structure or a reject branch containing a reason. The schema cannot represent both states simultaneously.
  • The shared deterministic schema checker gained oneOf support so local validation enforces the same exclusivity as generation-time structured output.
  • On the clean versioned rerun, all six outputs selected the claim branch with no claim/rejection contradiction. Four became independently validated STRUCTURING_PASS results.
  • Two remained STRUCTURING_CONTRACT_INVALID because outcome and intervention or comparison were both null despite selecting the claim branch. No span-to-structure hallucination or independent-validator rejection was observed in the four passing claims.
  • A topic-neutral prompt clarification made the already-required claim fields explicit; it did not add source context, source-specific examples, relaxed validation, or new claim semantics.

Prompt-clarification attempts and cleanup

Infrastructure-blocked before semantic evaluation
  • The first prompt-clarification attempt stopped at MODEL_START_TO_API_READY when the backend exited at 99 percent after approximately 604 seconds. No extraction request ran.
  • The one permitted identical retry reached API readiness in 603.587 seconds but failed the supervisor's bounded inference health check before Stage A. Again, no component request ran.
  • Each failure was preserved in a separate immutable run directory with its own model state, logs, manifest, frozen controls, cleanup evidence, and central-runtime restoration evidence.
  • After each failure, the temporary Mistral instance was removed and central Qwen was restored. The final Qwen restoration completed healthy with verified identity, real inference, and an idle slot.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen positive-control integrity passed Six historical validated claims across three positive sources were frozen with unique exact source offsets before inference; the identical control-set hash was retained across attempts.
Stage A positive-control recall passed Three of three sources were SPAN_RECALL_PASS, with complete coverage of every frozen supporting span.
Original Stage B gold-span structuring failed Zero of six passed: five contract-invalid claim/rejection combinations and one false-negative rejection.
Tagged Stage B gold-span structuring failed Four of six passed independent validation; two were contract-invalid for missing the required outcome and intervention or comparison. Mutual exclusivity itself worked for all six.
Prompt-only clarification not evaluated The initial attempt exited during load and the single retry failed post-load inference health before Stage A, so no semantic conclusion was drawn.
Focused boundary and extraction tests passed Twenty-one focused tests passed for gold derivation, exact coverage, boundary classifications, tagged-union exclusivity, nullable fields, provenance checks, and qualification truth tables.
Repository regression suite passed The full suite completed with 934 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 393.67 seconds.
Final runtime restoration passed Mistral was removed and central Qwen finished healthy, correctly identified, inference-verified, and idle.
Journal tests and production build passed The timestamped-entry tests passed, 197 journal pages generated and validated, and the TypeScript/Vite production build passed.

Current state

Next steps