Hiro development journal

Hosted GPT-5.6 Sol is reliable but overextracts the frozen gold set

Hosted qualification stopped at the unchanged gold-quality gate Machine-readable JSON

Executive summary

Hiro qualified exactly one already-authorized hosted path, OpenAI GPT-5.6 Sol, for the narrow claim-extraction boundary. Qwen 3.8 remained Hiro's local central model and local independent validator; production routing was not changed.

The qualification reused the latest frozen two-stage grounded extraction contracts, exact-span parser, canonical claim assembler, six gold claims, six-source controls, independent validator, thresholds, and stop rules. Only the extraction transport changed from local llama.cpp to the hosted Responses endpoint.

Hosted service operation was clean during the permitted gold stage: nine of nine requests completed on their first attempt with zero API errors, timeouts, retries, or terminal failures. Mean latency was 2.472121 seconds and p95 latency was 4.289065 seconds.

The unchanged gold gate failed on precision. All six expected evidence fragments were covered, but the model produced ten independently validated grounded claim units rather than exactly six, including one false claim-bearing span. Fabricated fields, unsupported surviving claims, parser failures, request failures, and validator failures were all zero.

The required stop rule was applied. The frozen six-source corpus, repeated service-reliability workload, frozen twenty-source corpus, production integration, and Phase 3F were not run. Final disposition: HOSTED EXTRACTION NOT QUALIFIED — QUALITY.

Work completed

Single hosted provider and data boundary

Completed
  • The environment exposed exactly one suitable already-authorized hosted provider: OpenAI. GPT-5.6 Sol was the only model tested, and account access to that exact model was verified before qualification.
  • Only public source text, source identifiers, immutable source hashes, and the frozen extraction instructions crossed the hosted boundary.
  • Credentials, private user data, unrelated repository contents, production logs, and unrelated Hiro state were not sent or persisted. API response storage was disabled.
  • The hosted model remained non-authoritative. Hiro's deterministic exact-span grounding and local Qwen independent validator controlled claim acceptance.

Qualification-only hosted adapter

Completed
  • The grounded qualification harness gained a transport-injection boundary whose default remains the existing local Qwen request path.
  • A qualification-only OpenAI Responses adapter executes the unchanged Stage 1 and Stage 2 plain-text contracts with their existing 384-token budgets and maximum of three claim blocks per verified span.
  • Remote requests use at most two attempts, retry only transport, timeout, rate-limit, or server failures, and never retry semantic failures. Every attempt is preserved as immutable secret-free evidence.
  • The adapter records normalized reasons, response identity, latency, usage, phase-level cost, and source hashes without persisting request bodies or credentials.
  • The final qualification source was committed and pushed on the existing qualification branch as 181e943. No production Hiro revision or routing configuration was changed.

First infrastructure divergence and minimum repair

Completed
  • The first launch stopped before any hosted inference because the persisted central-Qwen launcher state referenced a process that had already been replaced. The live port still belonged to one exact Qwen process with the frozen command and model identity.
  • The minimum repair reconciles the single exact command match and single port owner into a qualification-local supervisor state. It refuses ambiguous process sets and does not rewrite the production runtime-state file.
  • The authoritative run then verified the local validator healthy, correctly identified, owned, and idle before and after qualification. No validator restart or recovery was required.

Frozen gold qualification

Failed
  • The three positive gold sources contain six expected claim-bearing evidence fragments. Hosted Stage 1 returned six deterministically valid spans covering all six fragments, with no missed gold span and no Stage 1 parser or request failure.
  • Only three span boundaries were exact. Two spans were overbroad, one of those merged multiple expected claims, and one additional span was classified as a false claim-bearing span under the unchanged gold contract.
  • Hosted Stage 2 and Hiro's local validator produced ten canonical grounded claim units. All ten had valid provenance and passed independent validation, but the frozen gate requires exactly six expected units and zero false spans.
  • Fabricated fields, unsupported surviving claims, duplicate removals, Stage 2 parser failures, Stage 2 request failures, and validator runtime failures were all zero.
  • The first failed semantic boundary is HOSTED_EXTRACTION_TO_GOLD_GATE. The failure is over-extraction and segmentation precision, not hosted service reliability or unsupported field generation.

Service and cost accounting

Completed for the permitted gold stage
  • Nine requests produced nine successful completions in nine attempts. API or server errors, ordinary timeouts, retries, and terminal failures were all zero.
  • The gold stage consumed 3,351 input tokens and 792 output tokens. At the captured qualification rates of four dollars per million input tokens and twenty dollars per million output tokens, observed cost was $0.029244.
  • Using the completed three-source positive-heavy gold workload as the only available basis, projected hosted extraction cost is approximately $0.9748 per one hundred sources. This is not a six-source production-mix estimate because that stage was correctly withheld.
  • The repeated service-reliability campaign was not authorized after the gold failure, so the nine-request observation must not be represented as a complete service-reliability qualification.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen input and contract hashes passed The historical corpus, historical report, gold controls, Stage 1 contract, and Stage 2 contract matched the previously frozen hashes before execution.
Focused final-tree tests passed Twenty-seven hosted-adapter, grounded-extraction, model-qualification, dedicated-adapter, and source-claim tests passed after the final reporting-only cost projection fallback.
Repository regression suite passed The full repository suite completed with 956 passed, two expected skips, and six existing unknown-marker warnings in 647.95 seconds. This ran after the authoritative adapter and before the final reporting-only cost fallback; the focused final-tree suite covers that fallback.
Hosted gold requests passed operationally Nine of nine GPT-5.6 Sol requests completed on first attempt with zero API errors, timeouts, retries, or terminal failures.
Frozen gold semantic gate failed All six gold fragments were covered, but ten canonical grounded units and one false span violated the exact six-claim, zero-false-span requirement.
Frozen six-source quality corpus not run Withheld under the mandatory stop rule after the failed gold prerequisite. Zero-claim correctness was therefore not rescored.
Repeated service-reliability workload not run Withheld because the semantic prerequisite failed; no reliability claim is inferred from nine requests.
Final local central-model health passed Qwen 3.8 remained the exact owned local model, API-ready, inference-healthy, correctly identified, and idle.

Current state

Next steps