Hiro development journal

Restart verification of Qwen 3.8 and the accelerated improvement loop

Live system verified; accelerated validation and autonomous queue are operating Machine-readable JSON

Executive summary

Hiro and its local Qwen 3.8 model were already running when restart verification began. Qwen was loaded with the intended 8,192-token context and a single parallel slot, while Hiro exposed its HTTP application and auxiliary local listeners.

The production health endpoint reported an OK service, a connected LLM, and qwen/qwen3.8-27b. A real non-streaming chat request returned the exact requested readiness token in 6.9 seconds.

The complete regression suite passed 635 tests in 171.79 seconds. The checked-out revision did not change during the run, so the result applies cleanly to c004c1e rather than a moving target.

The benchmark dashboard rendered the risk-based 15-minute low-risk and 60-minute moderate-risk schedules, status-card filtering, the independent model ledger, and the active ranked queue without browser warnings or errors.

The autonomous loop remained live throughout verification. It advanced from one travel-planning candidate to a writing-assistance candidate, leaving one active item, 13 waiting, 21 retrying, three implemented, and 146 rejected with the infrastructure circuit breaker closed at zero failures.

Work completed

Runtime and model readiness

Completed
  • Confirmed that one Hiro Python process owned local listeners on ports 8000, 8001, and 8765. Port 8001 is the HTTP application surface; the other listeners are auxiliary local interfaces and were not evaluated as HTTP health endpoints.
  • Confirmed that LM Studio had qwen/qwen3.8-27b loaded with an 8,192-token context and parallelism set to one.
  • Called GET /health on the application surface and received status ok, llm connected, and model qwen/qwen3.8-27b.
  • Called POST /chat in a new readiness-probe session and requested an exact sentinel response. Hiro returned HIRO_READY in 6.9 seconds.

Accelerated governor and queue verification

Completed
  • Verified that commit 15c667e, which introduced the shortened validation policy, is an ancestor of the live revision c004c1e and that the Hiro working tree was clean before and after the regression run.
  • The live API reports low-risk validation at 15 minutes with probes at 0, 5, and 15 minutes, and moderate-risk validation at 60 minutes with probes at 0, 5, 15, and 60 minutes.
  • The benchmark page displayed the same durations and checkpoints in its continuous-authority panel.
  • The ranked-ideas status cards operated as filters. Selecting Active reduced the view to the active record and Show all restored the complete queue.
  • During the session, the governor completed work on the initially visible travel-planning candidate and selected the next ranked writing-assistance incident. The circuit breaker remained closed with zero consecutive infrastructure failures.

Independent model benchmark inspection

Completed
  • Verified that the Models view remains separately trackable from live Hiro and labels its canonical runs as a sealed, versioned longitudinal benchmark.
  • The frozen harness is version 2.1.0 with 24 cases: 16 recursive-improvement cases weighted at 60 percent and eight assistant cases weighted at 40 percent.
  • The existing canonical ledger ranks Qwen3.8 27B Q4_K_M above Qwen3.6 35B-A3B Q4_K_M: 92.1 percent overall and 87.7 percent RSI versus 78.7 percent overall and 79.8 percent RSI.
  • Neither model is canonically qualified because hard gates failed: six for Qwen3.8 and eight for Qwen3.6. Qwen3.8 is the operating model, but its runtime use does not rewrite or bypass the independent qualification record.
  • The existing canonical Qwen3.8 record reports 60.1 generated tokens per second, 662 milliseconds time to first token, and 44,823 milliseconds p95 latency.

Regression and browser validation

Completed
  • Ran the full Python regression suite against a stable revision and observed 635 passing tests with no failures.
  • Loaded the live benchmark page in the in-app browser and exercised the Current improvement, Ranked ideas, and Models views.
  • No warning or error entries appeared in the browser console during navigation and filter interaction.
  • No Hiro source changes were required during this verification session.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Production health endpoint passed GET http://127.0.0.1:8001/health returned status ok, llm connected, and qwen/qwen3.8-27b.
End-to-end model turn passed POST /chat returned the exact requested HIRO_READY sentinel in 6.9 seconds.
Full Hiro regression suite passed 635 tests passed in 171.79 seconds; Git remained at c004c1e before and after the run.
Live benchmark UI passed The accelerated schedules, queue counts, active filter, Show all reset, and independent Models ledger rendered correctly with no browser console warnings or errors.
Autonomous queue continuity passed The governor advanced to a new ranked candidate during verification and remained at one active item with a closed, zero-failure infrastructure circuit breaker.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 139 timestamped journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps