{
  "schemaVersion": 2,
  "date": "2026.09.01",
  "publishedAt": "2026-09-01T20:28:59-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Hosted GPT-5.6 Sol is reliable but overextracts the frozen gold set",
  "publicationStatus": "Hosted qualification stopped at the unchanged gold-quality gate",
  "executiveSummary": [
    "Hiro qualified exactly one already-authorized hosted path, OpenAI GPT-5.6 Sol, for the narrow claim-extraction boundary. Qwen 3.8 remained Hiro's local central model and local independent validator; production routing was not changed.",
    "The qualification reused the latest frozen two-stage grounded extraction contracts, exact-span parser, canonical claim assembler, six gold claims, six-source controls, independent validator, thresholds, and stop rules. Only the extraction transport changed from local llama.cpp to the hosted Responses endpoint.",
    "Hosted service operation was clean during the permitted gold stage: nine of nine requests completed on their first attempt with zero API errors, timeouts, retries, or terminal failures. Mean latency was 2.472121 seconds and p95 latency was 4.289065 seconds.",
    "The unchanged gold gate failed on precision. All six expected evidence fragments were covered, but the model produced ten independently validated grounded claim units rather than exactly six, including one false claim-bearing span. Fabricated fields, unsupported surviving claims, parser failures, request failures, and validator failures were all zero.",
    "The required stop rule was applied. The frozen six-source corpus, repeated service-reliability workload, frozen twenty-source corpus, production integration, and Phase 3F were not run. Final disposition: HOSTED EXTRACTION NOT QUALIFIED — QUALITY."
  ],
  "workstreams": [
    {
      "title": "Single hosted provider and data boundary",
      "status": "Completed",
      "details": [
        "The environment exposed exactly one suitable already-authorized hosted provider: OpenAI. GPT-5.6 Sol was the only model tested, and account access to that exact model was verified before qualification.",
        "Only public source text, source identifiers, immutable source hashes, and the frozen extraction instructions crossed the hosted boundary.",
        "Credentials, private user data, unrelated repository contents, production logs, and unrelated Hiro state were not sent or persisted. API response storage was disabled.",
        "The hosted model remained non-authoritative. Hiro's deterministic exact-span grounding and local Qwen independent validator controlled claim acceptance."
      ]
    },
    {
      "title": "Qualification-only hosted adapter",
      "status": "Completed",
      "details": [
        "The grounded qualification harness gained a transport-injection boundary whose default remains the existing local Qwen request path.",
        "A qualification-only OpenAI Responses adapter executes the unchanged Stage 1 and Stage 2 plain-text contracts with their existing 384-token budgets and maximum of three claim blocks per verified span.",
        "Remote requests use at most two attempts, retry only transport, timeout, rate-limit, or server failures, and never retry semantic failures. Every attempt is preserved as immutable secret-free evidence.",
        "The adapter records normalized reasons, response identity, latency, usage, phase-level cost, and source hashes without persisting request bodies or credentials.",
        "The final qualification source was committed and pushed on the existing qualification branch as 181e943. No production Hiro revision or routing configuration was changed."
      ]
    },
    {
      "title": "First infrastructure divergence and minimum repair",
      "status": "Completed",
      "details": [
        "The first launch stopped before any hosted inference because the persisted central-Qwen launcher state referenced a process that had already been replaced. The live port still belonged to one exact Qwen process with the frozen command and model identity.",
        "The minimum repair reconciles the single exact command match and single port owner into a qualification-local supervisor state. It refuses ambiguous process sets and does not rewrite the production runtime-state file.",
        "The authoritative run then verified the local validator healthy, correctly identified, owned, and idle before and after qualification. No validator restart or recovery was required."
      ]
    },
    {
      "title": "Frozen gold qualification",
      "status": "Failed",
      "details": [
        "The three positive gold sources contain six expected claim-bearing evidence fragments. Hosted Stage 1 returned six deterministically valid spans covering all six fragments, with no missed gold span and no Stage 1 parser or request failure.",
        "Only three span boundaries were exact. Two spans were overbroad, one of those merged multiple expected claims, and one additional span was classified as a false claim-bearing span under the unchanged gold contract.",
        "Hosted Stage 2 and Hiro's local validator produced ten canonical grounded claim units. All ten had valid provenance and passed independent validation, but the frozen gate requires exactly six expected units and zero false spans.",
        "Fabricated fields, unsupported surviving claims, duplicate removals, Stage 2 parser failures, Stage 2 request failures, and validator runtime failures were all zero.",
        "The first failed semantic boundary is HOSTED_EXTRACTION_TO_GOLD_GATE. The failure is over-extraction and segmentation precision, not hosted service reliability or unsupported field generation."
      ]
    },
    {
      "title": "Service and cost accounting",
      "status": "Completed for the permitted gold stage",
      "details": [
        "Nine requests produced nine successful completions in nine attempts. API or server errors, ordinary timeouts, retries, and terminal failures were all zero.",
        "The gold stage consumed 3,351 input tokens and 792 output tokens. At the captured qualification rates of four dollars per million input tokens and twenty dollars per million output tokens, observed cost was $0.029244.",
        "Using the completed three-source positive-heavy gold workload as the only available basis, projected hosted extraction cost is approximately $0.9748 per one hundred sources. This is not a six-source production-mix estimate because that stage was correctly withheld.",
        "The repeated service-reliability campaign was not authorized after the gold failure, so the nine-request observation must not be represented as a complete service-reliability qualification."
      ]
    }
  ],
  "decisions": [
    "Classify the final result as HOSTED EXTRACTION NOT QUALIFIED — QUALITY.",
    "Treat the first failed semantic boundary as HOSTED_EXTRACTION_TO_GOLD_GATE, specifically an over-extraction and segmentation-precision failure.",
    "Do not tune either frozen extraction stage, weaken the six-of-six contract, alter gold labels, or test a second hosted model.",
    "Do not run the frozen six-source or twenty-source corpora, enable production routing, change central Qwen routing, start Phase 3F, construct a candidate, or promote anything.",
    "Preserve the clean hosted runtime evidence separately from the failed semantic disposition: operational success does not override the unchanged quality gate."
  ],
  "validation": [
    {
      "check": "Frozen input and contract hashes",
      "status": "passed",
      "result": "The historical corpus, historical report, gold controls, Stage 1 contract, and Stage 2 contract matched the previously frozen hashes before execution."
    },
    {
      "check": "Focused final-tree tests",
      "status": "passed",
      "result": "Twenty-seven hosted-adapter, grounded-extraction, model-qualification, dedicated-adapter, and source-claim tests passed after the final reporting-only cost projection fallback."
    },
    {
      "check": "Repository regression suite",
      "status": "passed",
      "result": "The full repository suite completed with 956 passed, two expected skips, and six existing unknown-marker warnings in 647.95 seconds. This ran after the authoritative adapter and before the final reporting-only cost fallback; the focused final-tree suite covers that fallback."
    },
    {
      "check": "Hosted gold requests",
      "status": "passed operationally",
      "result": "Nine of nine GPT-5.6 Sol requests completed on first attempt with zero API errors, timeouts, retries, or terminal failures."
    },
    {
      "check": "Frozen gold semantic gate",
      "status": "failed",
      "result": "All six gold fragments were covered, but ten canonical grounded units and one false span violated the exact six-claim, zero-false-span requirement."
    },
    {
      "check": "Frozen six-source quality corpus",
      "status": "not run",
      "result": "Withheld under the mandatory stop rule after the failed gold prerequisite. Zero-claim correctness was therefore not rescored."
    },
    {
      "check": "Repeated service-reliability workload",
      "status": "not run",
      "result": "Withheld because the semantic prerequisite failed; no reliability claim is inferred from nine requests."
    },
    {
      "check": "Final local central-model health",
      "status": "passed",
      "result": "Qwen 3.8 remained the exact owned local model, API-ready, inference-healthy, correctly identified, and idle."
    }
  ],
  "currentState": [
    "The OpenAI hosted path is operational for the observed gold workload but is not a qualified Hiro claim extractor under the unchanged semantic contract.",
    "The first unresolved boundary remains SOURCE_CONTENT_COMPLETE to TRUSTWORTHY CLAIM EXTRACTION, now narrowed to claim-unit precision and segmentation rather than local runtime reliability.",
    "Qwen 3.8 remains Hiro's local central model and independent validator. Production claim-extraction routing is still disabled.",
    "The frozen six-source and twenty-source corpora remain unconsumed, the twenty-source corpus is not cleared, and Phase 3F remains stopped."
  ],
  "limitations": [
    "Service observations cover only nine gold-stage requests because the semantic stop rule correctly prevented the larger reliability workload.",
    "The cost projection uses three positive-heavy gold sources and may not represent a mixed production source distribution.",
    "The gold scorer accepts independent provenance-valid claims as canonical units but still requires exact corpus cardinality and zero false spans; this run exceeded that cardinality rather than missing source support.",
    "No claim can be made about legitimate zero-claim behavior because the frozen six-source corpus was not run.",
    "The qualification-only adapter is preserved on its branch but is not integrated into the production Hiro revision."
  ],
  "nextSteps": [
    "Preserve the immutable report with SHA-256 5944e795aeeb9d1705c3118d76d2b3085ea85aa1639a15e547d6c509dcfc4311 as the terminal evidence for this one hosted model attempt.",
    "Do not retry GPT-5.6 Sol with prompt tuning or test another hosted model automatically.",
    "Reassess the claim-extraction contract separately, focusing on whether the six gold units are the necessary operational representation and how claim-unit precision should be established without weakening grounding.",
    "Require separate authorization before any new hosted-model attempt, frozen six-source or twenty-source execution, production integration, or Phase 3F resumption."
  ],
  "disclosureNote": "This public entry contains no API credential, token, request body, private source text, personal data, private filesystem location, or actionable unresolved security detail."
}
