Executive summary
A narrowly bounded qualification tested whether vLLM under WSL2 could provide a reliable alternate backend for one stronger claim-extraction model. Hiro's production routing, extraction prompts, schemas, validators, frozen corpora, and downstream RSI policy were not changed.
The isolated vLLM environment and model storage were placed on a secondary drive after relocating the Ubuntu WSL2 distribution away from the constrained system drive. The RTX 5090 remained visible and usable after relocation.
vLLM 0.28.0 installed successfully and PyTorch 2.13.0 with CUDA 12.9 detected the RTX 5090. One Mistral Small 3.2 24B AWQ artifact was pinned to immutable upstream commit 49ec31ec0975e46a1aeef24186a45f11be3042b8 and its four local weight shards were hashed before execution.
No inference request executed. The first launch failed because vLLM's V2 runner requires unified virtual addressing unavailable under WSL2. The documented V1-runner fallback passed that boundary, but Marlin repacking failed after all weights loaded. A single bounded Triton-kernel retry also loaded all weights, then the engine process terminated before API readiness.
The alternate backend is therefore not qualified for runtime reliability in this configuration. No semantic quality test, six-source corpus, twenty-source corpus, production activation, or Phase 3F campaign ran. Hiro's original Qwen 3.8 runtime was restored and verified healthy with a real bounded response.