{
  "schemaVersion": 2,
  "date": "2026.08.31",
  "publishedAt": "2026-08-31T21:23:13-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Evidence-first extraction improves abstention but loses positive claims",
  "publicationStatus": "Protocol qualification complete; production routing remains disabled",
  "executiveSummary": [
    "Hiro's generation-first claim-extraction contract was replaced in a qualification-only pathway with a two-stage evidence-first protocol: exact supporting-span selection followed by independent span-bound structuring.",
    "The clean frozen six-source campaign completed using the already runtime-qualified Mistral Small 3.2 24B Q6_K configuration. Seventeen spans were proposed, sixteen passed exact-offset verification, and all Stage A and Stage B model responses were schema-valid.",
    "Grounding improved in one important respect: all three legitimate zero-claim controls returned zero validated claims, compared with one of three under the preserved generation-first result.",
    "The protocol nevertheless failed the unchanged quality gate. It recovered zero of three positive sources, produced only two raw structured claims, and both violated the span-to-structure contract, leaving zero independently validated claims. The evidence supports a mixed bottleneck involving positive-claim recall and span-to-structure contract failure.",
    "Mistral was unloaded, central Qwen was restored healthy and idle, and no production routing, twenty-source corpus execution, Phase 3F activity, candidate construction, or promotion occurred."
  ],
  "workstreams": [
    {
      "title": "Evidence-first protocol",
      "status": "Completed",
      "details": [
        "Stage A presents the complete immutable source and bounded sentence-like candidates, then permits explicit zero-span output or exact supporting spans with source-relative character offsets.",
        "A deterministic verifier rejects invalid offsets, empty or overlong spans, text mismatches, duplicate spans, and any proposed text absent from the source.",
        "Stage B processes each verified span independently. Its nullable contract records intervention or mechanism, comparison, outcome, metric, conditions, direction, and evidence type without requiring absent semantics to be invented.",
        "Every material non-null field must be an exact substring of its supporting span. Claim-bearing outputs retain the exact span, offsets, source hash, and immutable request evidence.",
        "Zero claims arise naturally when Stage A selects nothing, every span is rejected, or Stage B marks or validates no span as claim-bearing."
      ]
    },
    {
      "title": "Frozen six-source qualification",
      "status": "Failed",
      "details": [
        "The campaign used the same three positive and three legitimate zero-claim controls and the same independent semantic expectations used by the prior Qwen 2.5 Coder and Mistral generation-first qualifications.",
        "Stage A proposed seventeen spans; sixteen were deterministically valid, for a 94.118 percent span-verification rate. One proposal was rejected for mutated text and invalid offsets.",
        "All six Stage A responses and all sixteen Stage B responses were schema-valid, with no independent-validator runtime failures.",
        "The protocol produced two raw claim-bearing structures. Both carried a rejection reason while also declaring a claim, violating the frozen span-structure contract; both were deterministically rejected before semantic acceptance.",
        "Zero of three positive controls recovered an independently validated claim, all three positive controls were false negatives, and zero claims survived independent validation.",
        "All three legitimate zero-claim controls were handled correctly, but selected candidate spans in each negative control show that Stage A precision was not independently improved even though downstream abstention prevented false claims."
      ]
    },
    {
      "title": "Baseline comparison and diagnosis",
      "status": "Completed",
      "details": [
        "The preserved generation-first baseline recovered two of three positive sources, generated fourteen raw claims, retained six independently validated claims, produced eight unsupported or fabricated claims, achieved 42.857 percent provenance validity, and handled one of three zero-claim sources correctly.",
        "Evidence-first extraction reduced accepted unsupported claims by rejecting all raw structures, and improved zero-claim correctness from one of three to three of three, but this was achieved with unacceptable suppression: positive-source recovery fell from two of three to zero of three.",
        "The final disposition is EVIDENCE-FIRST CLAIM EXTRACTION NOT QUALIFIED — MIXED. The observed signals are span-selection or structuring recall loss and span-to-structure contract failure, not model startup, hard-hang, schema, independent-validator, or production-runtime failure."
      ]
    },
    {
      "title": "First-divergence repair and clean rerun",
      "status": "Completed",
      "details": [
        "The first attempt reached a successful Stage A request and a returned Stage B response, then exposed that the shared local JSON-schema checker could not evaluate a nullable type union.",
        "The minimum infrastructure repair added standards-compatible type-union and null handling to that checker without changing extraction semantics, prompt content, corpus content, thresholds, retry policy, or validation policy.",
        "The campaign was restarted from a clean run directory. Only the clean rerun is used for the reported qualification metrics.",
        "Mistral loaded in 596.517 seconds under progress-aware startup supervision, completed the six-source protocol, unloaded cleanly, and central Qwen was restored in 467.299 seconds."
      ]
    }
  ],
  "decisions": [
    "Do not qualify this evidence-first protocol with Mistral under the unchanged quality threshold.",
    "Preserve Mistral Small 3.2 24B Q6_K as the preferred runtime-reliable extractor candidate; this campaign diagnosed semantic behavior rather than runtime reliability.",
    "Do not loosen the quality gate: perfect negative-control abstention is insufficient when all known positive sources become false negatives.",
    "Do not activate production extraction routing, execute or clear the frozen twenty-source corpus, start Phase 3F, or test another model without separate authorization.",
    "Preserve the immutable clean-run evidence and the failed pre-run evidence for future protocol work."
  ],
  "validation": [
    {
      "check": "Focused protocol and runtime tests",
      "status": "passed",
      "result": "Fifteen focused tests passed, covering exact offsets, mutated-span rejection, natural zero-span behavior, nullable fields, material-field grounding, schema type unions, and strict qualification truth tables."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The full suite completed with 928 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 387.15 seconds."
    },
    {
      "check": "Frozen evidence-first quality campaign",
      "status": "failed",
      "result": "Seventeen spans proposed, sixteen offset-valid, two raw structured claims, zero validated claims, zero of three positive controls recovered, three of three zero-claim controls correct, and two unsupported contract-invalid raw claims."
    },
    {
      "check": "Runtime restoration",
      "status": "passed",
      "result": "Mistral was unloaded and central Qwen was restored with verified process ownership, model identity, API health, bounded inference, and an idle slot."
    },
    {
      "check": "Journal tests and production build",
      "status": "passed",
      "result": "The timestamped-entry tests passed, 196 journal pages generated and validated, and the TypeScript/Vite production build passed."
    }
  ],
  "currentState": [
    "The evidence-first pathway is implemented and evidence-producing, but it is not qualified for production because positive-claim recall is unacceptable.",
    "Mistral Small 3.2 24B Q6_K remains the current preferred runtime-qualified extractor candidate, not a production-qualified semantic extractor.",
    "Central Qwen is restored and healthy.",
    "Production extraction routing remains disabled.",
    "The frozen twenty-source Phase 3F-CE corpus remains uncleared and unexecuted, and Phase 3F has not started."
  ],
  "limitations": [
    "The result applies to the unchanged six-source corpus, expected outcomes, Mistral artifact, runtime configuration, quality gate, and independent validator.",
    "The candidate-span stage selected spans in all three negative controls, so correct final abstention cannot be attributed solely to high-precision span selection.",
    "Because both raw claim-bearing structures violated the deterministic contract, this campaign cannot separately estimate the semantic validator's acceptance precision on well-formed evidence-first claims.",
    "No prompt or schema tuning was performed after observing the clean qualification result."
  ],
  "nextSteps": [
    "Treat positive-span recognition and faithful span-to-structure conversion as the next semantic protocol problem, while keeping the frozen quality threshold intact.",
    "Use separately authorized work to determine whether claim-bearing semantics can be represented without the observed contradictory claim-plus-rejection output.",
    "Retest only under an explicitly authorized versioned protocol; do not reinterpret this failed run or overwrite its immutable evidence.",
    "Do not test another model, enable production routing, clear the frozen twenty-source corpus, or resume Phase 3F automatically."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, source corpus text, personal data, or actionable unresolved security details."
}
