Hiro development journal

Tightened Stage B fails semantic-completion controls

Stage B qualification complete; end-to-end extraction remains blocked Machine-readable JSON

Executive summary

Hiro's Stage B claim structurer was qualified independently on the exact six gold spans frozen by the accepted boundary diagnosis. Stage A was not executed.

The mutually exclusive claim-versus-reject result remained in place. The claim branch was tightened so it cannot validate without a nonempty outcome and at least one nonempty source-grounded antecedent field.

Mistral Small 3.2 24B Q6_K loaded successfully under its frozen runtime configuration, passed post-load health, and executed all six requests on one runtime instance.

Stage B failed all six controls. Four responses were schema-invalid, one was invalid JSON, and the only schema-valid response explicitly rejected a historically validated claim-bearing span. The corrected normalized outcome is five STRUCTURING_CONTRACT_INVALID and one STRUCTURING_FALSE_NEGATIVE, with zero passes and zero accepted hallucinations.

Because Stage B did not reach six of six, the six-source end-to-end corpus was not run. Production routing remained disabled, and the twenty-source corpus, Phase 3F, model selection, candidates, and promotion were untouched.

Work completed

Frozen Stage B-only contract

Completed
  • The runner consumes the previously frozen six-span artifact under its exact hash and verifies source order, source hashes, character offsets, exact text, historical provenance, and the six-claim count before model work.
  • No Stage A request path exists in the runner, ensuring the semantic-completion result cannot be confounded by span selection.
  • The claim and reject branches remain structurally exclusive. The claim branch has no rejection field.
  • A claim branch now requires a nonempty outcome and at least one nonempty intervention, mechanism, subject behavior, comparison, or other antecedent represented through the canonical fields. Metrics, conditions, and comparisons remain nullable when absent from the span.
  • Deterministic exact-substring and provenance rules remain unchanged; the model is still forbidden to manufacture missing empirical details.

Six gold-span semantic qualification

Failed
  • The exact authorized model loaded in 595.820 seconds under progress-aware supervision and completed post-load identity, bounded inference, and idle-slot health checks.
  • All six gold spans were submitted in their frozen order to one healthy runtime instance.
  • Four responses failed the tightened JSON schema, one response was not valid JSON, and one response validly selected the reject branch.
  • The explicit reject asserted that its supplied span lacked a supported proposition, contrary to the frozen historical validation; it is therefore a false negative.
  • No response produced a structurally valid claim, so no claim could proceed to independent semantic acceptance or exact-provenance success. There were no accepted unsupported additions because nothing reached claim acceptance.

Normalized outcome correction

Completed
  • The initial report grouped every absent claim as a false negative, obscuring the distinction between an explicit reject and an invalid response.
  • The preserved request terminal reasons showed four RESPONSE_SCHEMA_INVALID outcomes, one RESPONSE_JSON_INVALID outcome, and one valid reject response.
  • A separate immutable correction artifact was written without overwriting the original report. It normalizes missing or invalid Stage B responses as STRUCTURING_CONTRACT_INVALID and reserves STRUCTURING_FALSE_NEGATIVE for an explicit reject branch.
  • Corrected counts are zero STRUCTURING_PASS, one STRUCTURING_FALSE_NEGATIVE, zero STRUCTURING_HALLUCINATION, and five STRUCTURING_CONTRACT_INVALID. The qualification disposition is unchanged.

Runtime restoration and authority boundaries

Completed
  • Mistral was unloaded after the six requests with no orphan model instance.
  • Central Qwen was restored in 469.286 seconds and finished healthy, correctly identified, inference-verified, and idle.
  • The frozen six-source end-to-end campaign was withheld because the prerequisite Stage B component qualification failed.
  • No production extraction routing, twenty-source execution, Phase 3F activity, model substitution, candidate construction, or promotion was performed.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen control integrity passed The exact six-span hash, source order, source hashes, offsets, exact span text, historical provenance, and six-claim count were verified before runtime work.
Tagged claim-branch contract passed Focused tests prove claim and reject states are exclusive and that empty outcomes or absent antecedent fields cannot satisfy the claim branch.
Stage B semantic completion failed Corrected outcomes: zero passes, five contract-invalid responses, one explicit false-negative reject, and zero accepted hallucinations.
Focused qualification tests passed Twenty-four focused tests passed for frozen controls, schema exclusivity, semantic field requirements, classification normalization, provenance, and qualification truth tables.
Repository regression suite passed The full suite completed with 937 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 407.05 seconds.
Final runtime restoration passed Mistral was removed and central Qwen finished healthy, correctly identified, inference-verified, and idle.
Journal tests and production build passed The timestamped-entry tests passed, 198 journal pages generated and validated, and the TypeScript/Vite production build passed.

Current state

Next steps