{
  "schemaVersion": 2,
  "date": "2026.09.01",
  "publishedAt": "2026-09-01T16:44:00-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Grounded Qwen extraction removes hangs but merges a gold claim",
  "publicationStatus": "Runtime qualified; semantic qualification stopped at the one-claim-per-span gold gate",
  "executiveSummary": [
    "Hiro's final bounded extractor attempt tested the already-running qwen/qwen3.8-27b only. No model was installed, no inference backend was changed, and production routing, the frozen twenty-source corpus, and Phase 3F remained disabled.",
    "The attempt replaced server-constrained generation with a two-stage plain-text evidence contract. Stage 1 could emit only exact delimited source excerpts or an explicit zero-span token. Stage 2 structured one deterministically verified excerpt into seven bounded fields, with material values accepted only when they were exact substrings of that excerpt.",
    "Runtime reliability improved decisively. Ten consecutive cycles completed 60 Stage 1 and 60 Stage 2 requests with zero hard hangs, zero generation stalls, zero retries, and zero terminal failures. Stage 1 averaged 1.350156 seconds and Stage 2 averaged 1.432633 seconds.",
    "Semantic qualification nevertheless failed at VERIFIED_SPANS_TO_ONE_CLAIM_PER_SPAN_GOLD_GATE. Qwen covered all six known gold evidence fragments but returned five verified spans because one excerpt merged two distinct claims. The resulting five canonical claims were provenance-valid and independently validated, with zero unsupported claims and zero fabricated fields, but the frozen contract requires six separately recoverable claims.",
    "The frozen six-source quality qualification was correctly withheld after the failed gold prerequisite. Legitimate zero-claim behavior was therefore not scored, the twenty-source corpus is not cleared, and no production behavior changed."
  ],
  "workstreams": [
    {
      "title": "Plain-text evidence contracts",
      "status": "Completed",
      "details": [
        "Stage 1 instructs the model to copy exact contiguous claim-bearing excerpts between CLAIM_SPAN_START and CLAIM_SPAN_END delimiters, one claim per span, or return exactly NO_CLAIM_SPANS. Its frozen contract hash is 2a51511ad2e2dbaedb6a27b2da49c204af0fd50933d1373e6668efbb1f6fa3.",
        "Stage 2 receives exactly one verified span and emits seven single-line fields: intervention, comparison, outcome, metric, conditions, direction, and evidence type. Missing information must remain blank. Its frozen contract hash is 3545d87d9bc2869a32ce5112532fd4ee34ed52d372d8c11c699c90f1f0e797e0.",
        "No server-side constrained JSON or grammar was used in either model stage. The existing supervised streaming transport was minimally extended to return plain text while retaining progress observation, cancellation, bounded retry, ownership checks, and immutable attempt evidence.",
        "Output budgets were 384 tokens for Stage 1 and 192 tokens for Stage 2. Stage 1 allowed at most eight excerpts per source and 512 characters per excerpt."
      ]
    },
    {
      "title": "Deterministic provenance and field grounding",
      "status": "Completed",
      "details": [
        "Hiro independently recomputes the immutable source hash, requires every proposed excerpt to be nonempty, bounded, unique, and present exactly once in the source, and derives offsets locally rather than accepting model-provided offsets.",
        "Stage 2 accepts exactly seven ordered lines. Each nonempty material field must be an exact contiguous substring of the verified excerpt; direction and evidence type are limited to frozen enumerations. Missing fields remain null and no defaults are supplied.",
        "The unchanged canonical claim assembler and independent semantic validator remain downstream authorities. A claim cannot survive if the structured evidence violates source support, adds material facts, or fails field agreement.",
        "Focused tests cover plain-text transport, exact parsing, explicit zero-span output, missing-field preservation, invalid grounded fields, and the distinction between overbroad containment and multi-claim merging."
      ]
    },
    {
      "title": "Ten-cycle runtime qualification",
      "status": "Passed",
      "details": [
        "Ten consecutive cycles each exercised all six frozen full-source Stage 1 requests and all six frozen gold-span Stage 2 requests, for 120 model requests total.",
        "Stage 1 completed 60/60 valid requests in 60 attempts: zero hard hangs, zero stalls, zero retries, zero terminal failures, 1.350156-second mean latency, and 1.782458-second p95 latency.",
        "Stage 2 completed 60/60 valid requests in 60 attempts: zero hard hangs, zero stalls, zero retries, zero terminal failures, 1.432633-second mean latency, and 1.776331-second p95 latency.",
        "The same owned Qwen process, PID 15316, was healthy and idle before and after qualification. No runtime recovery or process replacement occurred.",
        "The preserved one-pass Qwen baseline had 80 requests, 105 attempts, 38 hard hangs, 25 retries, 13 terminal failures, 22.77-second mean latency, and 54.54-second p95 latency. The grounded protocol therefore resolved the observed runtime fragility under this bounded workload."
      ]
    },
    {
      "title": "Gold semantic gate",
      "status": "Failed",
      "details": [
        "The immutable gold set contains six canonical claims across three positive sources. Qwen's Stage 1 output contained five deterministically valid excerpts covering all six gold evidence fragments.",
        "Three excerpts were exact gold boundaries and two were overbroad containment spans. One of the overbroad spans combined two separately expected directional findings into one excerpt, violating the one-claim-per-span contract.",
        "Five canonical claims were recovered from the five excerpts. All five had exact-span provenance and passed independent semantic validation. Unsupported-claim count and fabricated-field count were both zero.",
        "There were no false spans, missed gold evidence fragments, Stage 1 request failures, or Stage 2 request failures. The failure is claim-segmentation granularity and canonical recall: five separately structured claims instead of six.",
        "Because the gold gate failed, the frozen six-source quality corpus was not run. Positive-source recovery and legitimate zero-claim correctness for that corpus are unavailable, not implicitly passing or failing."
      ]
    }
  ],
  "decisions": [
    "Classify the result as QWEN GROUNDED EXTRACTION NOT QUALIFIED — QUALITY.",
    "Treat VERIFIED_SPANS_TO_ONE_CLAIM_PER_SPAN_GOLD_GATE as the first semantic divergence; do not reinterpret full evidence coverage as six separately recovered claims.",
    "Preserve the material runtime improvement as evidence about protocol design without qualifying Qwen as Hiro's extractor.",
    "Do not loosen the six-claim gold threshold, split the model's merged excerpt after observing the result, or alter the frozen source, prompt, validator, and quality criteria.",
    "Do not run the frozen six-source or twenty-source corpora, activate production extraction routing, test another model, install vLLM, change llama.cpp, or resume Phase 3F."
  ],
  "validation": [
    {
      "check": "Contract immutability and authority",
      "status": "passed",
      "result": "Stage contracts, hashes, budgets, frozen corpus hashes, model identity, disabled production authority, and unchanged quality gates were recorded before execution."
    },
    {
      "check": "Ten-cycle runtime qualification",
      "status": "passed",
      "result": "120/120 plain-text requests completed in 120 attempts with zero hard hangs, stalls, retries, recoveries, or terminal failures."
    },
    {
      "check": "Gold-span semantic qualification",
      "status": "failed",
      "result": "Five verified excerpts yielded five valid canonical claims while the frozen gold contract requires six; one excerpt merged two claims. Fabricated and unsupported counts were zero."
    },
    {
      "check": "Frozen six-source quality qualification",
      "status": "not run",
      "result": "Withheld because the prerequisite gold semantic gate failed. Zero-claim correctness was not scored."
    },
    {
      "check": "Focused grounded-extraction tests",
      "status": "passed",
      "result": "Twenty-eight focused extraction, supervisor, evidence-first, and model-qualification tests passed in 0.70 seconds."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The complete repository suite finished with 951 passed, one expected skip, and six existing unknown-mark warnings in 387.64 seconds."
    },
    {
      "check": "Public journal tests and production build",
      "status": "passed",
      "result": "Timestamped-entry tests passed, 201 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully."
    }
  ],
  "currentState": [
    "Qwen 3.8 27B is runtime-stable under the new two-stage plain-text protocol but is not semantically qualified as Hiro's claim extractor.",
    "The first failed boundary is Stage 1 claim segmentation: a single valid excerpt merged two distinct gold claims and left the canonical pipeline at five of six required claims.",
    "All five resulting claims were exactly grounded and independently valid; the failure is not fabrication, provenance, request reliability, or the independent validator.",
    "Production extraction routing remains disabled. The six-source quality corpus was not consumed, the frozen twenty-source corpus is not cleared, and Phase 3F was not started."
  ],
  "limitations": [
    "The runtime conclusion is bounded to the exact 120-request qualification workload and the already-running Qwen 3.8 27B runtime.",
    "Gold evidence coverage does not equal canonical recall when one excerpt contains multiple separately expected claims; the frozen contract requires independent claim units.",
    "The six-source corpus was deliberately not run, so this session cannot report corpus-level positive-source recovery, zero-claim correctness, or provenance rate.",
    "An initial launcher attempt failed before producing a classifiable model result because plain-text response handling was placed in the wrong transport scope. The minimum transport fix was applied and the entire authoritative qualification restarted from clean evidence.",
    "The initial gold scorer counted overbroad exact containment as false. The scorer was minimally corrected to preserve both containment coverage and multi-claim-merge diagnostics, and the entire authoritative campaign was rerun without changing model output or thresholds."
  ],
  "nextSteps": [
    "Keep production extraction routing disabled and preserve the immutable report and attempt evidence as the final result of this authorized Qwen attempt.",
    "Any future extraction architecture should be evaluated first on whether it emits one independently structureable claim per evidence unit without sacrificing the zero-fabrication result.",
    "Do not proceed to another model, backend, corpus, or Phase 3F without separate authorization."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem locations, private source text, personal data, or actionable unresolved security details."
}
