Hiro development journal

Dedicated NuExtract qualification stops at native backend load

Runtime qualification failed before inference; semantic testing withheld Machine-readable JSON

Executive summary

Hiro's accepted Stage B diagnosis was followed with a dedicated information-extraction qualification for NuMind NuExtract 2.0 8B. Stage A, Mistral prompting, canonical quality thresholds, the twenty-source corpus, production routing, and Phase 3F were not changed or resumed.

The exact official BF16 model revision was downloaded and verified as a 15-file, 16.600 GB snapshot. A NuExtract-native adapter was built around the model's typed extraction template, with verbatim material fields, deterministic span provenance, clean abstention, and Hiro's existing independent semantic validator retained as final authority.

An isolated CUDA runtime was installed and verified on the RTX 5090. The final runtime attempt failed during model weight loading, before readiness or any inference request: the worker reached 57 of 729 tensors and then Windows recorded an access violation in PyTorch's torch_cpu.dll.

The failure boundary is MODEL_ARTIFACT_VALID_TO_NATIVE_WORKER_READY with normalized reason NATIVE_BACKEND_LOAD_ACCESS_VIOLATION. Runtime requests attempted: zero. Gold-span and six-source semantic qualification were correctly withheld.

Hiro's central Qwen worker was restored through its canonical supervisor and finished READY and inference-healthy. Production extraction routing remains disabled and the frozen twenty-source corpus is not cleared.

Work completed

Exact model and isolated native runtime

Completed
  • The selected model is numind/NuExtract-2.0-8B at immutable revision a470f0b5dd0b42fa7182cdbe7c6113a232f671e4, the authorized 8B release rather than the smaller 2B or 4B variants.
  • The verified snapshot contains 15 files totaling 16,600,363,222 bytes. Its canonical local manifest hash is 74d9f125e25392068f6f51023f88f3c60d6e96c0a3ef72afed1a59da45801f32.
  • The four BF16 weight shards matched the official expected sizes and SHA-256 values before GPU handoff.
  • A qualification-only Python environment was installed with PyTorch 2.11.0+cu128, torchvision 0.26.0+cu128, Transformers 5.16.1, Accelerate 1.14.0, Hugging Face Hub 1.20.1, safetensors 0.8.0, Pillow 12.1.1, and psutil 7.2.2.
  • CUDA was directly verified on the RTX 5090 with compute capability 12.0. The official NuExtract Qwen2.5-VL processor and empty model architecture constructed successfully before the bounded full-weight attempt.

NuExtract-native canonical adapter

Completed
  • NuExtract receives only one immutable evidence span plus its native JSON extraction template; it is not asked to imitate Mistral's generation-first protocol.
  • Intervention or mechanism, comparison or baseline, outcome, metric, and conditions use NuExtract's verbatim-string type. Direction and evidence type use bounded enums, and absent values remain null.
  • Claim-bearing versus no-claim is explicit. A no-claim result containing substantive assertions is contract-invalid rather than silently accepted.
  • Hiro attaches source identity, exact offsets, immutable source hashes, and the supporting span outside the model. No model output can alter those provenance anchors.
  • The existing deterministic claim checks and independent Qwen semantic validator remain authoritative. The adapter does not repair missing fields or complete semantics.

Process-bounded runtime qualification

Failed before inference
  • The native BF16 model runs in a dedicated one-slot child process with offline model resolution, hidden child-process creation, a 900-second startup ceiling, a 120-second request ceiling, and at most one infrastructure retry per request.
  • The first startup attempt exposed a supervisor integration defect: module-style launch imported Hiro's broad package initializer and failed on an irrelevant application dependency before model loading.
  • The minimal repair launched the dependency-isolated native worker file directly and made startup supervision poll child liveness every 250 milliseconds, preserve return code and stderr, and stop immediately when a child exits.
  • After a clean restart, the direct worker began loading BF16 weights and allocated the model on the GPU. It reached 57 of 729 tensors, approximately eight percent, before the Python process crashed.
  • Windows Error Reporting recorded exception 0xc0000005 in torch_cpu.dll. The worker return code was 3221225477. No API-ready event and no inference result existed.

Stop rule, recovery, and evidence

Completed
  • Because runtime qualification failed, the six gold-span semantic stage was not run and the frozen six-source quality stage was not authorized.
  • The final immutable failure report records zero requests, zero successful completions, zero hard generation hangs, one terminal startup failure, no retry, no fabrication result, and no semantic or provenance result.
  • The failure report SHA-256 is 6ab299f2749a0f633fa9167c4754e21d3802545315e86eaf67036c79ac032ed2.
  • Hiro's central Qwen worker was restored by the canonical runtime supervisor and verified READY, correctly owned, model-identified, inference-healthy, and idle.
  • No alternate NuExtract backend, quantization, model variant, general model, quality prompt, corpus, candidate, routing change, or promotion was attempted.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Official model identity and integrity passed The exact 8B revision, snapshot contents, aggregate manifest, four shard sizes, and four shard SHA-256 values were verified before runtime work.
Native runtime preflight passed CUDA PyTorch recognized the RTX 5090; the NuExtract processor, Qwen2.5-VL class, configuration, and empty architecture constructed successfully.
Runtime qualification failed The native worker crashed in torch_cpu.dll at 57/729 weight tensors before READY. Requests attempted: 0; hard generation hangs: 0; terminal startup failures: 1.
Six gold-span semantic qualification not run Withheld because runtime qualification did not pass.
Frozen six-source quality qualification not run Withheld because the prerequisite gold-span stage was not authorized.
Focused adapter and boundary tests passed Twenty-two focused tests passed for NuExtract request construction, exact-span canonical mapping, abstention, hallucinated field rejection, and the unchanged Stage B/evidence-first contracts.
Repository regression suite passed The full suite completed with 942 passed, one expected skip because its historical memory candidate is absent from this revision, and six existing unknown-mark warnings in 396.90 seconds.
Final runtime restoration passed Central Qwen finished READY and inference-healthy through the canonical supervisor after the NuExtract process crash.
Journal tests and production build passed Timestamped-entry tests passed, 199 journal pages generated and validated, and the TypeScript/Vite production build passed.

Current state

Next steps