{
  "schemaVersion": 2,
  "date": "2026.08.30",
  "publishedAt": "2026-08-30T15:07:09-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Local extraction-model qualification finds a quality bottleneck",
  "publicationStatus": "Bounded model comparison, supervisor recovery, frozen quality controls, repository validation, and central-runtime restoration complete",
  "executiveSummary": [
    "This session tested whether Hiro could assign claim extraction to a separately supervised local model while keeping Qwen 3.8 27B as Hiro's central reasoning model. It performed no new discovery, candidate construction, governor action, promotion, or meta-improvement.",
    "The extraction prompt, JSON schema, source hashes, specificity rules, provenance rules, deterministic checks, independent semantic validator, stall boundary, and one-retry policy were frozen across models.",
    "The prior Qwen 3.8 campaign remained the known failing reliability control. A local Qwen 3.5 4B candidate reached 20 valid requests but hard-hung on 4 of 24 attempts, a 16.67% rate above the preregistered 5% ceiling, so it never reached quality testing.",
    "A local Qwen 2.5 Coder 7B instruction model passed the runtime boundary: 10 consecutive two-source cycles, 20 successful requests in 21 attempts, one recovered hard hang, no terminal failures, and a 4.76% hard-hang rate.",
    "The 7B model then failed the unchanged six-source extraction-quality controls. It produced schema-valid output on 5 of 6 sources, recovered claims from 2 of 3 positive controls, fabricated claims for all 3 legitimate zero-claim controls, and had 3 of 8 raw claims rejected by deterministic or independent semantic validation.",
    "Because no alternate model passed both prerequisites, no extraction model was assigned to production and the frozen 20-source Phase 3F-CE corpus was not run.",
    "Two qualification-infrastructure boundaries were repaired minimally: long Windows evidence paths now use short deterministic request IDs, and the rejected extractor is released before central-model recovery. Neither repair changed prompts, model results, retries, or acceptance thresholds.",
    "The central Qwen 3.8 worker was restored with the identical model/configuration and passed exact process, API, model-identity, inference, and idle-slot health checks.",
    "Final disposition: EXTRACTION QUALITY BOTTLENECK. A fresh Phase 3F campaign is not justified from this evidence."
  ],
  "workstreams": [
    {
      "title": "Frozen workload and acceptance contract",
      "status": "Completed",
      "details": [
        "The 20-source Phase 3F-CE corpus retained SHA-256 0a8fc19512c3e35d2787626e6c2e31f29d920bb90adfcc96ff2b835e11bcdbf7 and remained unopened until both prerequisite qualifications could pass.",
        "Runtime qualification replayed the exact preserved two-source workload. Acceptance required 10 consecutive successful cycles, zero terminal failures, and a hard-hang attempt rate no greater than 5%.",
        "Quality qualification used six immutable historical controls: three sources with previously validated claim structure and three legitimate zero-claim sources. Expected outputs were assessed semantically rather than by word-for-word reproduction.",
        "The model could not certify its own claims. Every candidate claim passed through unchanged deterministic provenance checks and the separate semantic validator."
      ]
    },
    {
      "title": "Model-independent supervised request boundary",
      "status": "Completed",
      "details": [
        "The existing supervised structured-request path now accepts a model alias and request factory without changing request content. The same infrastructure can carry claim extraction or independent validation while retaining exact process ownership, immutable attempts, progress detection, bounded recovery, and one retry maximum.",
        "The independent semantic-validation request now has an explicit builder. Downstream consumers continue to receive the same canonical validated claim object and do not need to know which model produced the raw extraction.",
        "No production routing assignment was created because no candidate model qualified. Qwen 3.8 remains the central model and the existing production extraction assignment remains unchanged."
      ]
    },
    {
      "title": "Qwen 3.5 4B reliability candidate",
      "status": "Rejected on reliability incidence",
      "details": [
        "The 4B model completed all 20 requests and reached a 10-cycle streak only after four exact-worker replacements and four successful versioned retries.",
        "Four of 24 attempts hard-hung, producing a 16.67% hard-hang rate. Mean successful-request latency was 14.76 seconds and P95 was 20.02 seconds.",
        "The run had zero terminal request failures, but the preregistered incidence rule prevents a lucky recovered streak from being treated as reliable. Quality testing was therefore blocked for this model."
      ]
    },
    {
      "title": "Qwen 2.5 Coder 7B runtime qualification",
      "status": "Passed",
      "details": [
        "The 7B worker loaded in 3.99 seconds and completed 10 consecutive real two-source cycles.",
        "All 20 requests succeeded across 21 immutable attempts. One request hard-hung, the supervisor replaced only the exactly owned worker, and the single permitted versioned retry succeeded.",
        "The final hard-hang rate was 4.76%, with zero terminal failures. Mean successful-request latency was 7.95 seconds and P95 was 11.01 seconds.",
        "The accepted runtime report is frozen at SHA-256 cc06f609f617904d9aff4b87e1383e711808e979dd73450e235f221072935a77."
      ]
    },
    {
      "title": "Qwen 2.5 Coder 7B extraction quality",
      "status": "Failed",
      "details": [
        "Five of six controls returned schema-valid extractions; one positive control terminated as RESPONSE_JSON_INVALID.",
        "At least one independently validated claim was recovered from two of three positive sources rather than the required three of three.",
        "The model returned eight raw claims, of which five passed deterministic and independent validation. Three were rejected as unsupported or otherwise provenance-invalid.",
        "All three legitimate zero-claim sources received a fabricated claim rather than zero claims. The zero-claim acceptance result was 0 of 3.",
        "The independent validator itself had zero runtime failures and did not repair or rewrite rejected output."
      ]
    },
    {
      "title": "Qualification evidence and cleanup",
      "status": "Completed",
      "details": [
        "The first quality response exposed a Windows path-length failure while writing its sidecar hash. Short deterministic request IDs crossed that persistence boundary without changing the underlying request or reusing mutable results.",
        "The accepted runtime report was linked by hash into a new versioned quality run rather than overwritten.",
        "Cleanup initially attempted central-model recovery while the rejected extractor still occupied GPU resources. The exactly owned extractor was released first, after which an identical central-model restart passed full inference health.",
        "The final qualification report is frozen at SHA-256 9aa525b3f9aa591e81f5037045c87058ed826f8812b4aef6cc89cf461180f516."
      ]
    }
  ],
  "decisions": [
    "Use the prior Qwen 3.8 campaign as the immutable known-failing control rather than rerunning it to seek an accidental streak.",
    "Test only two smaller local candidates: Qwen 3.5 4B and, after its reliability rejection, the one optional Qwen 2.5 Coder 7B candidate.",
    "Do not tune the extraction prompt, schema, claim semantics, or validator for a particular model.",
    "Require both reliability and quality. Supervisor recovery alone cannot qualify a frequently hanging model, and reliable JSON generation cannot qualify an extractor that fabricates claims.",
    "Do not create a production extraction route without a model that passes both prerequisites.",
    "Do not run the frozen 20-source corpus, resume Phase 3F discovery, or construct candidates from this result.",
    "Preserve the runtime-qualified 7B evidence even though its downstream quality result failed."
  ],
  "validation": [
    {
      "check": "Qwen 3.5 4B repeated runtime workload",
      "status": "failed acceptance",
      "result": "Ten consecutive cycles and 20 successful requests were observed, but four hard hangs in 24 attempts produced a 16.67% rate above the frozen 5% ceiling."
    },
    {
      "check": "Qwen 2.5 Coder 7B repeated runtime workload",
      "status": "passed",
      "result": "Ten consecutive cycles; 20 successful requests in 21 attempts; one hard hang; one successful retry; zero terminal failures; final rate 4.76%."
    },
    {
      "check": "Frozen six-source extraction-quality controls",
      "status": "failed acceptance",
      "result": "Five of six schema-valid extractions, two of three positive-source recoveries, five of eight validated claims, three rejected claims, and zero of three legitimate zero-claim controls handled correctly."
    },
    {
      "check": "Focused qualification, supervisor, runtime, and source-claim tests",
      "status": "passed",
      "result": "Nineteen focused tests passed in 0.85 seconds after the final code changes."
    },
    {
      "check": "Complete Hiro repository suite",
      "status": "passed",
      "result": "908 tests passed, one expected test skipped, and zero tests failed in 440.06 seconds. Six existing unknown-mark warnings were reported."
    },
    {
      "check": "Final runtime state",
      "status": "passed",
      "result": "The alternate worker was absent, its port was closed, and the central Qwen 3.8 worker was READY with admission open after exact identity and real inference verification."
    }
  ],
  "currentState": [
    "No local extraction model is qualified. The final disposition is EXTRACTION QUALITY BOTTLENECK.",
    "Qwen 3.8 27B remains Hiro's central reasoning model and is healthy at session end.",
    "The 20-source Phase 3F-CE corpus was not rerun because the quality prerequisite failed.",
    "The qualification implementation and evidence documentation are pushed on codex/phase3fce-qualification at commit 0701d28 and are not merged or activated in production."
  ],
  "limitations": [
    "The Qwen 2.5 Coder 7B model is reliable enough under the frozen runtime criterion but cannot safely replace the extractor because it overproduces claims on negative controls and fails independent provenance checks.",
    "The Qwen 3.5 4B candidate remains too hang-prone despite successful supervisor containment.",
    "Per-process VRAM attribution was unavailable under the Windows display-driver mode, so the report retains model file size, load latency, request latency, output size, and recovery overhead without inventing a VRAM figure.",
    "No result supports weakening the zero-claim, provenance, schema, retry, or independent-validation requirements."
  ],
  "nextSteps": [
    "If authorized separately, test a different bounded local extractor model or non-generative extraction approach against the same frozen quality contract.",
    "Preserve the qualified 7B runtime result as infrastructure evidence, but do not route production extraction to it unless a future version independently passes quality.",
    "Do not start a fresh Phase 3F campaign until a claim extractor passes both the runtime and quality prerequisites.",
    "Do not use this task-specific model comparison as evidence that Hiro is recursively self-improving."
  ],
  "disclosureNote": "This public entry contains no credentials, private source text, personal interaction content, private filesystem locations, private network addresses, or actionable unresolved security details."
}
