Hiro development journal

Remaining local claim extractors fail the frozen runtime boundary

Bounded model qualification, immutable evidence sealing, repository validation, and production-runtime restoration complete Machine-readable JSON

Executive summary

This session continued the existing claim-extraction model qualification without rerunning models already tested. It evaluated the two remaining locally installed candidates: Qwen 3.5 9B and Gemma 4 12B instruction QAT.

Both candidates used the exact frozen two-source reliability workload, extraction prompt, response schema, 30-second progress-stall rule, exact-process recovery, one-retry maximum, and 5% hard-hang ceiling used by the prior qualification.

Qwen 3.5 9B failed in cycle 2 when one source request hard-hung on both its original attempt and its sole retry. The accepted evidence window contains four requests, five attempts, three valid completions, two hard hangs, one terminal failure, and a best streak of one cycle.

Gemma 4 12B completed four successful two-source cycles, but two of ten attempts hard-hung. Even six perfect remaining cycles would have produced a 9.09% hard-hang rate at the required ten-cycle boundary, so the frozen 5% target had become mathematically unreachable.

Quality testing was not authorized for either model. The frozen 20-source corpus was not executed, no production extraction route was changed, and no discovery, candidate, governor, promotion, or Phase 3F work occurred.

The production Qwen 3.8 worker was restored with its exact model and configuration. Its first cold-start inference check hung; the existing exact-process replacement mechanism then restored a healthy worker that passed ownership, API, model identity, real inference, and idle-slot verification.

The final cross-model disposition is MIXED MODEL BOTTLENECK: one previously tested candidate passed runtime but failed extraction quality, while the remaining installed candidates failed runtime reliability. No local model is qualified for claim extraction.

Work completed

Exclusive-GPU qualification scheduling

Completed
  • The remaining candidates could not coexist in GPU memory with central Qwen 3.8, so qualification changed resource scheduling without changing evaluation semantics.
  • The process verified the exactly owned central worker, stopped it, loaded one candidate, ran the frozen runtime workload, stopped the candidate, and restored the identical central model.
  • Had a candidate passed runtime, its six frozen quality extractions would have been preserved before releasing it and independently validated only after central Qwen was restored. Neither candidate reached that stage.
  • Production defaults and routing remained unchanged throughout.

Qwen 3.5 9B runtime qualification

Rejected
  • The exact model artifact was verified by hash before launch.
  • Cycle 1 passed. In cycle 2, one source completed and the second source hard-hung twice: once on the original request and once on its single versioned retry.
  • The first conclusive boundary was a terminal retry failure. Accepted metrics were four requests, five attempts, three valid completions, two hard hangs, one retry, zero successful retries, and one terminal failure.
  • An older runner continued briefly beyond the conclusive boundary. Those later immutable artifacts were preserved for audit but explicitly excluded from qualification scoring.
  • Quality testing and the 20-source corpus were blocked for this model.

Gemma 4 12B runtime qualification

Rejected
  • The exact instruction/QAT model artifact was verified by hash before launch.
  • The model completed eight source requests across four successful cycles. Two requests required exact-worker replacement and succeeded on their one permitted retry.
  • The observed incidence was two hard hangs in ten attempts. At a ten-cycle target with six remaining clean cycles, the best possible final incidence was two hangs in twenty-two attempts, or 9.09%.
  • Because that optimistic rate still exceeded the frozen 5% ceiling, the run ended at normalized boundary RUNTIME_HARD_HANG_RATE_UNRECOVERABLE rather than consuming additional work that could not change the outcome.
  • Quality testing and the 20-source corpus were blocked for this model.

Qualification-runner boundary fixes

Completed
  • Long evidence request identifiers were shortened to deterministic hashes to remain safe under Windows path limits without changing request semantics.
  • The runtime loop now stops immediately on the first terminal request failure instead of beginning later source requests or cycles.
  • The runtime loop also calculates the optimistic minimum hard-hang rate at the required streak and stops with an explicit reason when the frozen criterion cannot be recovered.
  • A preserved-evidence finalizer records the first conclusive failure boundary, excludes later accidental work from metrics, verifies the candidate is absent, and requires healthy central inference before sealing a report.
  • None of these repairs changed prompts, schemas, thresholds, retry counts, source selection, claim semantics, or model outcomes.

Production-runtime restoration

Completed
  • The first post-Gemma cold reload produced a healthy API and exact model identity, but its real inference-health request hung and left the single slot occupied. Production was not declared healthy on API status alone.
  • The existing supervisor terminated only the exactly owned stalled process and launched the identical model and configuration.
  • The replacement passed process ownership, port ownership, API health, model identity, real structured inference, and idle-slot checks. Admission reopened only after those checks passed.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Qwen 3.5 9B frozen reliability workload failed acceptance First conclusive failure in cycle 2: four requests, five attempts, three valid completions, two hard hangs, one terminal retry failure, best streak one of ten.
Gemma 4 12B frozen reliability workload failed acceptance Four successful cycles and eight valid requests, but two hard hangs in ten attempts made the best attainable ten-cycle rate 9.09%, above the 5% ceiling.
Focused qualification, supervisor, runtime, and source-claim tests passed Twenty focused tests passed after the final stop-condition and evidence-finalization changes.
Complete Hiro repository suite passed 910 tests passed, one expected test skipped, and zero tests failed in 480.24 seconds. Six existing unknown-mark warnings were reported.
Final production runtime passed Central Qwen 3.8 is READY with admission open after exact process, port, model identity, API, real inference, and idle-slot verification. Both candidate workers are absent.

Current state

Next steps