Hiro development journal

Alternate vLLM extraction backend stops before inference

Bounded backend qualification stopped at engine startup reliability Machine-readable JSON

Executive summary

A narrowly bounded qualification tested whether vLLM under WSL2 could provide a reliable alternate backend for one stronger claim-extraction model. Hiro's production routing, extraction prompts, schemas, validators, frozen corpora, and downstream RSI policy were not changed.

The isolated vLLM environment and model storage were placed on a secondary drive after relocating the Ubuntu WSL2 distribution away from the constrained system drive. The RTX 5090 remained visible and usable after relocation.

vLLM 0.28.0 installed successfully and PyTorch 2.13.0 with CUDA 12.9 detected the RTX 5090. One Mistral Small 3.2 24B AWQ artifact was pinned to immutable upstream commit 49ec31ec0975e46a1aeef24186a45f11be3042b8 and its four local weight shards were hashed before execution.

No inference request executed. The first launch failed because vLLM's V2 runner requires unified virtual addressing unavailable under WSL2. The documented V1-runner fallback passed that boundary, but Marlin repacking failed after all weights loaded. A single bounded Triton-kernel retry also loaded all weights, then the engine process terminated before API readiness.

The alternate backend is therefore not qualified for runtime reliability in this configuration. No semantic quality test, six-source corpus, twenty-source corpus, production activation, or Phase 3F campaign ran. Hiro's original Qwen 3.8 runtime was restored and verified healthy with a real bounded response.

Work completed

WSL2 and storage prerequisites

Completed
  • Ubuntu 24.04 on WSL2 reported kernel 6.6.87.2 and exposed an NVIDIA GeForce RTX 5090 with the Windows 576.88 driver and CUDA 12.9 compatibility.
  • The original system drive had insufficient headroom for multi-gigabyte backend dependencies. A partial isolated install was stopped, its temporary package cache was removed, and the managed Ubuntu distribution was relocated to a secondary drive using WSL's supported move operation.
  • After relocation, the Linux root filesystem, package cache, isolated environment, and downloaded model consumed secondary-drive storage. The distribution restarted successfully and retained GPU visibility.

Isolated vLLM backend

Installed but not qualified
  • A fresh Python 3.12 virtual environment installed vLLM 0.28.0 from the supported prebuilt-wheel path without modifying Hiro's Python environment.
  • The resulting stack reported Python 3.12.3, vLLM 0.28.0, PyTorch 2.13.0+cu129, CUDA 12.9, one visible CUDA device, and the expected RTX 5090 identity.
  • Because the current wheel's stable extension depends on bundled CUDA runtime libraries, the serve command explicitly supplied only the wheel-local library directories. No system CUDA toolkit or Linux display driver was installed.

Single pinned extraction model

Downloaded and verified
  • The only model selected was jeffcookio/Mistral-Small-3.2-24B-Instruct-2506-awq-sym, the successful AWQ artifact cited in the official Mistral model discussion and traceable to the official Mistral Small 3.2 base.
  • The model repository was pinned to commit 49ec31ec0975e46a1aeef24186a45f11be3042b8. The local checkpoint contained four safetensor shards totaling approximately 15 GB, with per-shard SHA-256 values retained in the private qualification evidence.
  • No alternate model was downloaded or tested.

Backend startup qualification

Failed before inference
  • The first launch resolved the expected Mistral3 architecture and text-only 4,096-token configuration, then stopped at EngineCore initialization with UVA unavailable. This exactly matches an upstream WSL2 V2-runner limitation.
  • The minimal documented repair set VLLM_USE_V2_MODEL_RUNNER=0. The V1 runner then loaded all four checkpoint shards in 93.66 seconds but failed during the Marlin gptq_marlin_repack transformation before API readiness.
  • One bounded kernel retry selected vLLM's supported Triton W4A16 linear backend. It loaded all four shards in 16.18 seconds and reported 12.0 GiB of model memory, then the engine process terminated before the API became ready.
  • Because no server reached API readiness, the READY control, structured-output control, twenty sequential requests, and extraction semantic gates were not authorized to run.

Production restoration

Passed
  • The temporary qualification required exclusive GPU memory, so Hiro's existing Qwen server was stopped without changing its routing or configuration.
  • After the vLLM stop condition, the checked-in hidden Qwen launcher restored the exact Qwen 3.8 27B artifact, llama.cpp 2.31.2 backend, 16,384-token context, one parallel slot, full GPU offload, flash attention, and port 8080.
  • Restoration was verified by HTTP health 200, the expected qwen/qwen3.8-27b model identity, and a real bounded chat completion returning READY.

Decisions and reasoning

Validation and evidence

CheckStatusResult
GPU prerequisite passed WSL2 and the isolated PyTorch environment both identified the RTX 5090 and reported CUDA availability.
Storage isolation passed The managed Ubuntu distribution, vLLM environment, caches, and selected model now reside on secondary-drive storage rather than the constrained system drive.
vLLM API readiness failed The V2 runner failed at unavailable UVA; the V1 Marlin path failed during weight repacking; the V1 Triton path loaded weights but its engine terminated before API readiness.
Inference and semantic quality not run No vLLM inference request executed, so runtime request reliability and frozen semantic quality remain untested.
Hiro Qwen restoration passed The production Qwen endpoint returned health 200, the expected model identity, and READY from a bounded real inference request.
Public journal tests and production build passed Timestamped-entry tests passed, 203 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully.

Current state

Next steps