Hiro development journal

Local extraction-model qualification finds a quality bottleneck

Bounded model comparison, supervisor recovery, frozen quality controls, repository validation, and central-runtime restoration complete Machine-readable JSON

Executive summary

This session tested whether Hiro could assign claim extraction to a separately supervised local model while keeping Qwen 3.8 27B as Hiro's central reasoning model. It performed no new discovery, candidate construction, governor action, promotion, or meta-improvement.

The extraction prompt, JSON schema, source hashes, specificity rules, provenance rules, deterministic checks, independent semantic validator, stall boundary, and one-retry policy were frozen across models.

The prior Qwen 3.8 campaign remained the known failing reliability control. A local Qwen 3.5 4B candidate reached 20 valid requests but hard-hung on 4 of 24 attempts, a 16.67% rate above the preregistered 5% ceiling, so it never reached quality testing.

A local Qwen 2.5 Coder 7B instruction model passed the runtime boundary: 10 consecutive two-source cycles, 20 successful requests in 21 attempts, one recovered hard hang, no terminal failures, and a 4.76% hard-hang rate.

The 7B model then failed the unchanged six-source extraction-quality controls. It produced schema-valid output on 5 of 6 sources, recovered claims from 2 of 3 positive controls, fabricated claims for all 3 legitimate zero-claim controls, and had 3 of 8 raw claims rejected by deterministic or independent semantic validation.

Because no alternate model passed both prerequisites, no extraction model was assigned to production and the frozen 20-source Phase 3F-CE corpus was not run.

Two qualification-infrastructure boundaries were repaired minimally: long Windows evidence paths now use short deterministic request IDs, and the rejected extractor is released before central-model recovery. Neither repair changed prompts, model results, retries, or acceptance thresholds.

The central Qwen 3.8 worker was restored with the identical model/configuration and passed exact process, API, model-identity, inference, and idle-slot health checks.

Final disposition: EXTRACTION QUALITY BOTTLENECK. A fresh Phase 3F campaign is not justified from this evidence.

Work completed

Frozen workload and acceptance contract

Completed
  • The 20-source Phase 3F-CE corpus retained SHA-256 0a8fc19512c3e35d2787626e6c2e31f29d920bb90adfcc96ff2b835e11bcdbf7 and remained unopened until both prerequisite qualifications could pass.
  • Runtime qualification replayed the exact preserved two-source workload. Acceptance required 10 consecutive successful cycles, zero terminal failures, and a hard-hang attempt rate no greater than 5%.
  • Quality qualification used six immutable historical controls: three sources with previously validated claim structure and three legitimate zero-claim sources. Expected outputs were assessed semantically rather than by word-for-word reproduction.
  • The model could not certify its own claims. Every candidate claim passed through unchanged deterministic provenance checks and the separate semantic validator.

Model-independent supervised request boundary

Completed
  • The existing supervised structured-request path now accepts a model alias and request factory without changing request content. The same infrastructure can carry claim extraction or independent validation while retaining exact process ownership, immutable attempts, progress detection, bounded recovery, and one retry maximum.
  • The independent semantic-validation request now has an explicit builder. Downstream consumers continue to receive the same canonical validated claim object and do not need to know which model produced the raw extraction.
  • No production routing assignment was created because no candidate model qualified. Qwen 3.8 remains the central model and the existing production extraction assignment remains unchanged.

Qwen 3.5 4B reliability candidate

Rejected on reliability incidence
  • The 4B model completed all 20 requests and reached a 10-cycle streak only after four exact-worker replacements and four successful versioned retries.
  • Four of 24 attempts hard-hung, producing a 16.67% hard-hang rate. Mean successful-request latency was 14.76 seconds and P95 was 20.02 seconds.
  • The run had zero terminal request failures, but the preregistered incidence rule prevents a lucky recovered streak from being treated as reliable. Quality testing was therefore blocked for this model.

Qwen 2.5 Coder 7B runtime qualification

Passed
  • The 7B worker loaded in 3.99 seconds and completed 10 consecutive real two-source cycles.
  • All 20 requests succeeded across 21 immutable attempts. One request hard-hung, the supervisor replaced only the exactly owned worker, and the single permitted versioned retry succeeded.
  • The final hard-hang rate was 4.76%, with zero terminal failures. Mean successful-request latency was 7.95 seconds and P95 was 11.01 seconds.
  • The accepted runtime report is frozen at SHA-256 cc06f609f617904d9aff4b87e1383e711808e979dd73450e235f221072935a77.

Qwen 2.5 Coder 7B extraction quality

Failed
  • Five of six controls returned schema-valid extractions; one positive control terminated as RESPONSE_JSON_INVALID.
  • At least one independently validated claim was recovered from two of three positive sources rather than the required three of three.
  • The model returned eight raw claims, of which five passed deterministic and independent validation. Three were rejected as unsupported or otherwise provenance-invalid.
  • All three legitimate zero-claim sources received a fabricated claim rather than zero claims. The zero-claim acceptance result was 0 of 3.
  • The independent validator itself had zero runtime failures and did not repair or rewrite rejected output.

Qualification evidence and cleanup

Completed
  • The first quality response exposed a Windows path-length failure while writing its sidecar hash. Short deterministic request IDs crossed that persistence boundary without changing the underlying request or reusing mutable results.
  • The accepted runtime report was linked by hash into a new versioned quality run rather than overwritten.
  • Cleanup initially attempted central-model recovery while the rejected extractor still occupied GPU resources. The exactly owned extractor was released first, after which an identical central-model restart passed full inference health.
  • The final qualification report is frozen at SHA-256 9aa525b3f9aa591e81f5037045c87058ed826f8812b4aef6cc89cf461180f516.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Qwen 3.5 4B repeated runtime workload failed acceptance Ten consecutive cycles and 20 successful requests were observed, but four hard hangs in 24 attempts produced a 16.67% rate above the frozen 5% ceiling.
Qwen 2.5 Coder 7B repeated runtime workload passed Ten consecutive cycles; 20 successful requests in 21 attempts; one hard hang; one successful retry; zero terminal failures; final rate 4.76%.
Frozen six-source extraction-quality controls failed acceptance Five of six schema-valid extractions, two of three positive-source recoveries, five of eight validated claims, three rejected claims, and zero of three legitimate zero-claim controls handled correctly.
Focused qualification, supervisor, runtime, and source-claim tests passed Nineteen focused tests passed in 0.85 seconds after the final code changes.
Complete Hiro repository suite passed 908 tests passed, one expected test skipped, and zero tests failed in 440.06 seconds. Six existing unknown-mark warnings were reported.
Final runtime state passed The alternate worker was absent, its port was closed, and the central Qwen 3.8 worker was READY with admission open after exact identity and real inference verification.

Current state

Next steps