{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T17:34:40-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Restart verification of Qwen 3.8 and the accelerated improvement loop",
  "publicationStatus": "Live system verified; accelerated validation and autonomous queue are operating",
  "executiveSummary": [
    "Hiro and its local Qwen 3.8 model were already running when restart verification began. Qwen was loaded with the intended 8,192-token context and a single parallel slot, while Hiro exposed its HTTP application and auxiliary local listeners.",
    "The production health endpoint reported an OK service, a connected LLM, and qwen/qwen3.8-27b. A real non-streaming chat request returned the exact requested readiness token in 6.9 seconds.",
    "The complete regression suite passed 635 tests in 171.79 seconds. The checked-out revision did not change during the run, so the result applies cleanly to c004c1e rather than a moving target.",
    "The benchmark dashboard rendered the risk-based 15-minute low-risk and 60-minute moderate-risk schedules, status-card filtering, the independent model ledger, and the active ranked queue without browser warnings or errors.",
    "The autonomous loop remained live throughout verification. It advanced from one travel-planning candidate to a writing-assistance candidate, leaving one active item, 13 waiting, 21 retrying, three implemented, and 146 rejected with the infrastructure circuit breaker closed at zero failures."
  ],
  "workstreams": [
    {
      "title": "Runtime and model readiness",
      "status": "Completed",
      "details": [
        "Confirmed that one Hiro Python process owned local listeners on ports 8000, 8001, and 8765. Port 8001 is the HTTP application surface; the other listeners are auxiliary local interfaces and were not evaluated as HTTP health endpoints.",
        "Confirmed that LM Studio had qwen/qwen3.8-27b loaded with an 8,192-token context and parallelism set to one.",
        "Called GET /health on the application surface and received status ok, llm connected, and model qwen/qwen3.8-27b.",
        "Called POST /chat in a new readiness-probe session and requested an exact sentinel response. Hiro returned HIRO_READY in 6.9 seconds."
      ]
    },
    {
      "title": "Accelerated governor and queue verification",
      "status": "Completed",
      "details": [
        "Verified that commit 15c667e, which introduced the shortened validation policy, is an ancestor of the live revision c004c1e and that the Hiro working tree was clean before and after the regression run.",
        "The live API reports low-risk validation at 15 minutes with probes at 0, 5, and 15 minutes, and moderate-risk validation at 60 minutes with probes at 0, 5, 15, and 60 minutes.",
        "The benchmark page displayed the same durations and checkpoints in its continuous-authority panel.",
        "The ranked-ideas status cards operated as filters. Selecting Active reduced the view to the active record and Show all restored the complete queue.",
        "During the session, the governor completed work on the initially visible travel-planning candidate and selected the next ranked writing-assistance incident. The circuit breaker remained closed with zero consecutive infrastructure failures."
      ]
    },
    {
      "title": "Independent model benchmark inspection",
      "status": "Completed",
      "details": [
        "Verified that the Models view remains separately trackable from live Hiro and labels its canonical runs as a sealed, versioned longitudinal benchmark.",
        "The frozen harness is version 2.1.0 with 24 cases: 16 recursive-improvement cases weighted at 60 percent and eight assistant cases weighted at 40 percent.",
        "The existing canonical ledger ranks Qwen3.8 27B Q4_K_M above Qwen3.6 35B-A3B Q4_K_M: 92.1 percent overall and 87.7 percent RSI versus 78.7 percent overall and 79.8 percent RSI.",
        "Neither model is canonically qualified because hard gates failed: six for Qwen3.8 and eight for Qwen3.6. Qwen3.8 is the operating model, but its runtime use does not rewrite or bypass the independent qualification record.",
        "The existing canonical Qwen3.8 record reports 60.1 generated tokens per second, 662 milliseconds time to first token, and 44,823 milliseconds p95 latency."
      ]
    },
    {
      "title": "Regression and browser validation",
      "status": "Completed",
      "details": [
        "Ran the full Python regression suite against a stable revision and observed 635 passing tests with no failures.",
        "Loaded the live benchmark page in the in-app browser and exercised the Current improvement, Ranked ideas, and Models views.",
        "No warning or error entries appeared in the browser console during navigation and filter interaction.",
        "No Hiro source changes were required during this verification session."
      ]
    }
  ],
  "decisions": [
    "Keep Hiro and Qwen running after verification because the user requested a live restart and continued testing.",
    "Treat port 8001 as the authoritative HTTP health surface and record ports 8000 and 8765 as auxiliary listeners rather than applying an invalid shared endpoint assumption.",
    "Test the exact live revision and compare the Git revision before and after the full suite so autonomous commits cannot silently invalidate the result.",
    "Use one bounded production chat probe to verify the complete model path without forcing, promoting, or otherwise interfering with the governor's active candidate.",
    "Preserve the distinction between operating-model selection and canonical model qualification; higher aggregate scores do not erase failed hard gates."
  ],
  "validation": [
    {
      "check": "Production health endpoint",
      "status": "passed",
      "result": "GET http://127.0.0.1:8001/health returned status ok, llm connected, and qwen/qwen3.8-27b."
    },
    {
      "check": "End-to-end model turn",
      "status": "passed",
      "result": "POST /chat returned the exact requested HIRO_READY sentinel in 6.9 seconds."
    },
    {
      "check": "Full Hiro regression suite",
      "status": "passed",
      "result": "635 tests passed in 171.79 seconds; Git remained at c004c1e before and after the run."
    },
    {
      "check": "Live benchmark UI",
      "status": "passed",
      "result": "The accelerated schedules, queue counts, active filter, Show all reset, and independent Models ledger rendered correctly with no browser console warnings or errors."
    },
    {
      "check": "Autonomous queue continuity",
      "status": "passed",
      "result": "The governor advanced to a new ranked candidate during verification and remained at one active item with a closed, zero-failure infrastructure circuit breaker."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 139 timestamped journal entries, then TypeScript and Vite completed the production build successfully."
    }
  ],
  "currentState": [
    "Hiro is running on revision c004c1e with a clean working tree and the Qwen 3.8 27B model loaded at an 8,192-token context.",
    "The live API is healthy, one real chat turn passed, and the complete Python suite is green.",
    "The continuous queue reports one active moderate-risk writing-assistance candidate, 13 waiting, 21 retrying, three implemented, and 146 rejected.",
    "The infrastructure breaker is closed with zero consecutive infrastructure failures.",
    "Low-risk validation uses 15 minutes and moderate-risk validation uses 60 minutes; both schedules are visible on the dashboard.",
    "The independent Models ledger still records Qwen3.8 as canonically rejected because six hard gates failed, even though it substantially outscored Qwen3.6 and is currently serving Hiro."
  ],
  "limitations": [
    "This session verified the policy, tests, user interface, API, and live model path, but it did not force a candidate through an entire 60-minute moderate-risk canary; the governor was allowed to continue naturally.",
    "The readiness chat latency is one end-to-end observation and is not a replacement for the benchmark's repeated TTFT, generation-throughput, or p95 measurements.",
    "The queue still has a large historical rejection population and 21 retrying records. Continued work should separate scientifically useful rejection evidence from recurring proposal-construction failure modes.",
    "Qwen3.8 remains unqualified under the frozen hard-gate policy. Its better aggregate score supports continued runtime evaluation but is not evidence that every safety, contract, and recursive-improvement gate is satisfied."
  ],
  "nextSteps": [
    "Observe the current moderate-risk candidate through its natural 0, 5, 15, and 60-minute probes and confirm that promotion or rejection is recorded without intervention.",
    "Break down the six Qwen3.8 canonical hard-gate failures by failure class and determine whether they reflect model limitations, harness calibration, or deployment configuration without altering the frozen historical run.",
    "Track retry conversion rate, rejection category, and time-to-decision under the shorter schedules to determine whether throughput improves without increasing regressions.",
    "Run a separately labeled live-compatibility benchmark when enough operational evidence exists; do not overwrite either canonical model run."
  ]
}
