{
  "schemaVersion": 2,
  "date": "2026.09.01",
  "publishedAt": "2026-09-01T19:18:29-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Alternate vLLM extraction backend stops before inference",
  "publicationStatus": "Bounded backend qualification stopped at engine startup reliability",
  "executiveSummary": [
    "A narrowly bounded qualification tested whether vLLM under WSL2 could provide a reliable alternate backend for one stronger claim-extraction model. Hiro's production routing, extraction prompts, schemas, validators, frozen corpora, and downstream RSI policy were not changed.",
    "The isolated vLLM environment and model storage were placed on a secondary drive after relocating the Ubuntu WSL2 distribution away from the constrained system drive. The RTX 5090 remained visible and usable after relocation.",
    "vLLM 0.28.0 installed successfully and PyTorch 2.13.0 with CUDA 12.9 detected the RTX 5090. One Mistral Small 3.2 24B AWQ artifact was pinned to immutable upstream commit 49ec31ec0975e46a1aeef24186a45f11be3042b8 and its four local weight shards were hashed before execution.",
    "No inference request executed. The first launch failed because vLLM's V2 runner requires unified virtual addressing unavailable under WSL2. The documented V1-runner fallback passed that boundary, but Marlin repacking failed after all weights loaded. A single bounded Triton-kernel retry also loaded all weights, then the engine process terminated before API readiness.",
    "The alternate backend is therefore not qualified for runtime reliability in this configuration. No semantic quality test, six-source corpus, twenty-source corpus, production activation, or Phase 3F campaign ran. Hiro's original Qwen 3.8 runtime was restored and verified healthy with a real bounded response."
  ],
  "workstreams": [
    {
      "title": "WSL2 and storage prerequisites",
      "status": "Completed",
      "details": [
        "Ubuntu 24.04 on WSL2 reported kernel 6.6.87.2 and exposed an NVIDIA GeForce RTX 5090 with the Windows 576.88 driver and CUDA 12.9 compatibility.",
        "The original system drive had insufficient headroom for multi-gigabyte backend dependencies. A partial isolated install was stopped, its temporary package cache was removed, and the managed Ubuntu distribution was relocated to a secondary drive using WSL's supported move operation.",
        "After relocation, the Linux root filesystem, package cache, isolated environment, and downloaded model consumed secondary-drive storage. The distribution restarted successfully and retained GPU visibility."
      ]
    },
    {
      "title": "Isolated vLLM backend",
      "status": "Installed but not qualified",
      "details": [
        "A fresh Python 3.12 virtual environment installed vLLM 0.28.0 from the supported prebuilt-wheel path without modifying Hiro's Python environment.",
        "The resulting stack reported Python 3.12.3, vLLM 0.28.0, PyTorch 2.13.0+cu129, CUDA 12.9, one visible CUDA device, and the expected RTX 5090 identity.",
        "Because the current wheel's stable extension depends on bundled CUDA runtime libraries, the serve command explicitly supplied only the wheel-local library directories. No system CUDA toolkit or Linux display driver was installed."
      ]
    },
    {
      "title": "Single pinned extraction model",
      "status": "Downloaded and verified",
      "details": [
        "The only model selected was jeffcookio/Mistral-Small-3.2-24B-Instruct-2506-awq-sym, the successful AWQ artifact cited in the official Mistral model discussion and traceable to the official Mistral Small 3.2 base.",
        "The model repository was pinned to commit 49ec31ec0975e46a1aeef24186a45f11be3042b8. The local checkpoint contained four safetensor shards totaling approximately 15 GB, with per-shard SHA-256 values retained in the private qualification evidence.",
        "No alternate model was downloaded or tested."
      ]
    },
    {
      "title": "Backend startup qualification",
      "status": "Failed before inference",
      "details": [
        "The first launch resolved the expected Mistral3 architecture and text-only 4,096-token configuration, then stopped at EngineCore initialization with UVA unavailable. This exactly matches an upstream WSL2 V2-runner limitation.",
        "The minimal documented repair set VLLM_USE_V2_MODEL_RUNNER=0. The V1 runner then loaded all four checkpoint shards in 93.66 seconds but failed during the Marlin gptq_marlin_repack transformation before API readiness.",
        "One bounded kernel retry selected vLLM's supported Triton W4A16 linear backend. It loaded all four shards in 16.18 seconds and reported 12.0 GiB of model memory, then the engine process terminated before the API became ready.",
        "Because no server reached API readiness, the READY control, structured-output control, twenty sequential requests, and extraction semantic gates were not authorized to run."
      ]
    },
    {
      "title": "Production restoration",
      "status": "Passed",
      "details": [
        "The temporary qualification required exclusive GPU memory, so Hiro's existing Qwen server was stopped without changing its routing or configuration.",
        "After the vLLM stop condition, the checked-in hidden Qwen launcher restored the exact Qwen 3.8 27B artifact, llama.cpp 2.31.2 backend, 16,384-token context, one parallel slot, full GPU offload, flash attention, and port 8080.",
        "Restoration was verified by HTTP health 200, the expected qwen/qwen3.8-27b model identity, and a real bounded chat completion returning READY."
      ]
    }
  ],
  "decisions": [
    "Classify the alternate backend attempt as VLLM BACKEND NOT QUALIFIED — RUNTIME RELIABILITY because the selected model never reached API readiness or executed inference.",
    "Do not interpret the result as evidence about Mistral extraction semantics; the semantic gate was never reached.",
    "Stop after the documented V1 fallback and one supported Triton-kernel retry rather than beginning a backend or model safari.",
    "Keep production extraction routing disabled and preserve the frozen six-source and twenty-source corpora for a later explicitly authorized path."
  ],
  "validation": [
    {
      "check": "GPU prerequisite",
      "status": "passed",
      "result": "WSL2 and the isolated PyTorch environment both identified the RTX 5090 and reported CUDA availability."
    },
    {
      "check": "Storage isolation",
      "status": "passed",
      "result": "The managed Ubuntu distribution, vLLM environment, caches, and selected model now reside on secondary-drive storage rather than the constrained system drive."
    },
    {
      "check": "vLLM API readiness",
      "status": "failed",
      "result": "The V2 runner failed at unavailable UVA; the V1 Marlin path failed during weight repacking; the V1 Triton path loaded weights but its engine terminated before API readiness."
    },
    {
      "check": "Inference and semantic quality",
      "status": "not run",
      "result": "No vLLM inference request executed, so runtime request reliability and frozen semantic quality remain untested."
    },
    {
      "check": "Hiro Qwen restoration",
      "status": "passed",
      "result": "The production Qwen endpoint returned health 200, the expected model identity, and READY from a bounded real inference request."
    },
    {
      "check": "Public journal tests and production build",
      "status": "passed",
      "result": "Timestamped-entry tests passed, 203 journal pages and aliases were generated and validated, and the TypeScript/Vite production build completed successfully."
    }
  ],
  "currentState": [
    "Hiro is back on its established Qwen 3.8 llama.cpp runtime and is healthy.",
    "The isolated vLLM stack and pinned Mistral AWQ artifact remain available on secondary-drive storage for evidence, but no vLLM server is running.",
    "The trustworthy-claim-extraction boundary remains unresolved because the alternate backend failed before inference.",
    "Production extraction routing remains disabled; frozen corpora and Phase 3F remain untouched."
  ],
  "limitations": [
    "The terminal Triton attempt ended before API readiness without a complete engine-side reason in the parent process output, so its precise post-load termination mechanism remains unresolved.",
    "This qualification used one reputable community quantization of the official Mistral model, not the full-precision official checkpoint, because the latter requires substantially more GPU memory than the single RTX 5090 provides.",
    "No claim-extraction quality conclusion can be drawn because no generation request ran."
  ],
  "nextSteps": [
    "Evaluate the preserved backend evidence before authorizing any different inference backend, model artifact, or runtime architecture.",
    "Do not resume the frozen extraction corpora or Phase 3F until a backend/model combination first reaches stable API inference and passes the unchanged semantic gates.",
    "Continue using the restored Qwen runtime for Hiro's existing central reasoning role."
  ],
  "disclosureNote": "This public entry contains no credentials, private filesystem paths, private source corpus text, personal data, or actionable unresolved security details."
}
