{
  "schemaVersion": 2,
  "date": "2026.08.31",
  "publishedAt": "2026-08-31T23:24:53-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Boundary diagnosis isolates evidence-first structuring failures",
  "publicationStatus": "Boundary diagnosis complete; evidence-first extraction remains unqualified",
  "executiveSummary": [
    "Hiro's evidence-first claim extraction was decomposed into independent supporting-span selection and gold-span structuring tests using the same three frozen positive controls and Mistral Small 3.2 24B Q6_K.",
    "Six previously independently validated historical claims were recovered from immutable qualification artifacts. Their exact supporting excerpts were located uniquely in the frozen sources, assigned exact offsets, frozen before inference, and preserved under one control-set hash.",
    "Stage A was exonerated: all three positive sources achieved SPAN_RECALL_PASS, and every historical supporting span was completely covered by at least one selected span.",
    "The original Stage B contract failed all six gold claims: five claim-bearing outputs were contract-invalid because the required nullable rejection field was emitted as an empty string, and one known-good span was rejected as incomplete. The resulting component diagnosis was STRUCTURING_BOTTLENECK.",
    "A minimal tagged-union repair made claim and reject responses structurally exclusive. On the versioned component rerun, Stage B improved to four of six independently validated passes; the remaining two chose the claim branch but omitted both the required outcome and intervention or comparison.",
    "A further prompt-only clarification could not receive a semantic result: one model load exited at 99 percent and the single retry reached API readiness but failed bounded post-load inference health before Stage A. Those infrastructure failures do not alter Mistral's prior 20-of-20 runtime qualification or the last completed semantic result.",
    "The frozen six-source corpus was not rerun because Stage B did not pass all gold controls. Production routing, the twenty-source corpus, Phase 3F, model selection, candidates, and promotion remained untouched."
  ],
  "workstreams": [
    {
      "title": "Immutable positive controls",
      "status": "Completed",
      "details": [
        "The controls were derived solely from the historical independently validated qualification report and immutable six-source corpus; current Mistral outputs were not used to define gold claims.",
        "The three positive sources contained six validated claims: three about trajectory handoff behavior, two about resource-constrained execution, and one benchmark comparison.",
        "Each historical supporting excerpt occurred exactly once in its source. Exact character offsets, the original structured fields, historical semantic validation, source hashes, and provenance decisions were frozen before testing.",
        "The frozen control artifact retained the same SHA-256 across every versioned attempt, preventing prompt or schema changes from modifying the qualification targets."
      ]
    },
    {
      "title": "Independent Stage A diagnosis",
      "status": "Passed",
      "details": [
        "Stage A ran alone on the three positive sources; Stage B was not invoked as part of these measurements.",
        "All three sources were classified SPAN_RECALL_PASS. There were no partial, missed, or false-positive-only source outcomes.",
        "Every one of the six frozen supporting spans had complete deterministic coverage by a selected exact-offset candidate.",
        "This evidence corrects the prior aggregate diagnosis: positive claims were not lost because Stage A failed to surface their source evidence."
      ]
    },
    {
      "title": "Independent Stage B diagnosis",
      "status": "Failed",
      "details": [
        "Stage B bypassed Stage A and received each of the six frozen known-good supporting spans independently.",
        "Under the original schema, zero of six gold claims passed. Five outputs declared themselves claim-bearing while also supplying an empty rejection field, and one rejected a known-good span as incomplete.",
        "Inspection established that the strict schema required one nullable rejection field in both accepted and rejected output states. Although deterministic validation correctly rejected the contradiction, the schema itself permitted it and encouraged empty-string null substitution.",
        "The component disposition was STRUCTURING_BOTTLENECK rather than span-selection, validator, or boundary-handoff failure."
      ]
    },
    {
      "title": "Minimal tagged-result repair",
      "status": "Partially successful",
      "details": [
        "Stage B now uses a mutually exclusive tagged JSON result: a claim branch containing the canonical span-bound structure or a reject branch containing a reason. The schema cannot represent both states simultaneously.",
        "The shared deterministic schema checker gained oneOf support so local validation enforces the same exclusivity as generation-time structured output.",
        "On the clean versioned rerun, all six outputs selected the claim branch with no claim/rejection contradiction. Four became independently validated STRUCTURING_PASS results.",
        "Two remained STRUCTURING_CONTRACT_INVALID because outcome and intervention or comparison were both null despite selecting the claim branch. No span-to-structure hallucination or independent-validator rejection was observed in the four passing claims.",
        "A topic-neutral prompt clarification made the already-required claim fields explicit; it did not add source context, source-specific examples, relaxed validation, or new claim semantics."
      ]
    },
    {
      "title": "Prompt-clarification attempts and cleanup",
      "status": "Infrastructure-blocked before semantic evaluation",
      "details": [
        "The first prompt-clarification attempt stopped at MODEL_START_TO_API_READY when the backend exited at 99 percent after approximately 604 seconds. No extraction request ran.",
        "The one permitted identical retry reached API readiness in 603.587 seconds but failed the supervisor's bounded inference health check before Stage A. Again, no component request ran.",
        "Each failure was preserved in a separate immutable run directory with its own model state, logs, manifest, frozen controls, cleanup evidence, and central-runtime restoration evidence.",
        "After each failure, the temporary Mistral instance was removed and central Qwen was restored. The final Qwen restoration completed healthy with verified identity, real inference, and an idle slot."
      ]
    }
  ],
  "decisions": [
    "Classify the completed semantic evidence as EVIDENCE-FIRST NOT QUALIFIED — STRUCTURING BOTTLENECK.",
    "Do not attribute the positive false negatives to Stage A; its three source-level controls passed completely.",
    "Retain the tagged-union repair because it deterministically removes the demonstrated contradictory result state and improved gold-span structuring from zero of six to four of six.",
    "Do not claim the final prompt clarification works because both versioned attempts stopped before component inference.",
    "Do not consume another runtime retry, run the six-source corpus, test another model, activate production routing, execute the twenty-source corpus, or start Phase 3F."
  ],
  "validation": [
    {
      "check": "Frozen positive-control integrity",
      "status": "passed",
      "result": "Six historical validated claims across three positive sources were frozen with unique exact source offsets before inference; the identical control-set hash was retained across attempts."
    },
    {
      "check": "Stage A positive-control recall",
      "status": "passed",
      "result": "Three of three sources were SPAN_RECALL_PASS, with complete coverage of every frozen supporting span."
    },
    {
      "check": "Original Stage B gold-span structuring",
      "status": "failed",
      "result": "Zero of six passed: five contract-invalid claim/rejection combinations and one false-negative rejection."
    },
    {
      "check": "Tagged Stage B gold-span structuring",
      "status": "failed",
      "result": "Four of six passed independent validation; two were contract-invalid for missing the required outcome and intervention or comparison. Mutual exclusivity itself worked for all six."
    },
    {
      "check": "Prompt-only clarification",
      "status": "not evaluated",
      "result": "The initial attempt exited during load and the single retry failed post-load inference health before Stage A, so no semantic conclusion was drawn."
    },
    {
      "check": "Focused boundary and extraction tests",
      "status": "passed",
      "result": "Twenty-one focused tests passed for gold derivation, exact coverage, boundary classifications, tagged-union exclusivity, nullable fields, provenance checks, and qualification truth tables."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The full suite completed with 934 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 393.67 seconds."
    },
    {
      "check": "Final runtime restoration",
      "status": "passed",
      "result": "Mistral was removed and central Qwen finished healthy, correctly identified, inference-verified, and idle."
    },
    {
      "check": "Journal tests and production build",
      "status": "passed",
      "result": "The timestamped-entry tests passed, 197 journal pages generated and validated, and the TypeScript/Vite production build passed."
    }
  ],
  "currentState": [
    "Stage A is empirically qualified on all three frozen positive controls.",
    "Stage B remains unqualified: the best completed tagged-contract result is four of six gold claims passed.",
    "The contradictory claim-plus-rejection state is eliminated by schema and deterministic validation.",
    "The latest prompt-only clarification exists in code but has no semantic qualification result because both attempts failed before component inference.",
    "Mistral remains the preferred runtime-qualified extractor candidate based on its prior twenty-of-twenty reliability evidence; this session does not revoke that result.",
    "Production extraction routing is disabled, the frozen six-source corpus was not rerun, the twenty-source corpus remains untouched, and Phase 3F has not started."
  ],
  "limitations": [
    "The component diagnosis uses three positive sources and six historical claims; it isolates the observed boundary but does not constitute full-corpus qualification.",
    "Four of six tagged Stage B passes are insufficient under the unchanged requirement that known-good spans structure correctly.",
    "The prompt-only clarification cannot be judged from the two infrastructure-blocked attempts, and no additional retry was authorized.",
    "The two startup/health failures are preserved operational observations, but they are not treated as evidence that Mistral's previously qualified runtime is generally unreliable."
  ],
  "nextSteps": [
    "Do not rerun the six-source corpus until Stage B passes the frozen gold-span controls under an explicitly authorized future attempt.",
    "If another attempt is authorized, begin with the unchanged prompt-only version and the same frozen positive-control hash rather than redesigning the protocol.",
    "Keep the tagged result schema and strict deterministic rejection of missing required claim fields.",
    "Do not test another model, activate production routing, execute the twenty-source corpus, or start Phase 3F automatically."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, source corpus text, personal data, or actionable unresolved security details."
}
