Hiro development journal

Qwen3.8 becomes Hiro's active runtime and continuous improvement resumes

Implemented, validated, and running Machine-readable JSON

Executive summary

Hiro's production inference profile now selects the locally installed Qwen3.8 27B Q4_K_M model by exact identifier, with an 8,192-token context, single-request parallelism, and full GPU offload.

The model qualification system remains independent of Hiro's production profile. Qwen3.6 and Qwen3.8 retain their immutable matched benchmark evidence on the Models page even though the production model changed.

Hiro and its continuous evidence-driven improvement scheduler were restarted after removing the explicit operator stop marker. The first post-restart rotating interaction audit completed four cases on Qwen3.8 and passed all four.

The previously open infrastructure circuit breaker was traced to the intentionally dirty repository containing the completed benchmark and model-migration work. After all 576 repository tests passed, the known changes were committed as e5c4897 and the breaker was reset to zero against a clean baseline.

Stage 6C runtime-fragment mutation remains inactive. Restarting the existing promotion system did not silently grant the separate exact-policy activation that its governance contract requires.

Work completed

First-class Qwen3.8 runtime profile

Completed
  • Added a qwen38 router profile for both fast and standard local inference using the exact qwen/qwen3.8-27b identifier on the loopback LM Studio endpoint.
  • Changed persistent configuration defaults, health-check startup selection, operator documentation, and nightly API preflight to Qwen3.8 while retaining qwen36 and qwen9 as reversible fallback profiles.
  • Added a pinned PowerShell launcher and background batch wrapper for the installed Qwen3.8 artifact. The launcher normalizes the Windows Path/PATH collision, requests an 8,192-token context, one parallel request, maximum GPU offload, verifies the returned model identifier, performs a warmup, and writes startup metadata.

Runtime restart and live verification

Completed
  • Validated the real Qwen3.8 launcher against the already loaded model. It reported the exact model key, 8,192-token context, one parallel request, maximum GPU offload, and a successful READY warmup.
  • Removed the explicit servers-intentionally-stopped marker and launched Hiro through the checked-in Windows-safe process helper.
  • Verified the task API, HTTPS and HTTP listeners, Evaluation Observatory, continuous-improvement endpoint, and improvement-pipeline endpoint were live.
  • Startup telemetry identified Qwen3.8 as both the fast and standard model, and LM Studio accepted inference from the restarted service.

Continuous-improvement recovery

Completed
  • Confirmed the active policy remains enabled in continuous_evidence_driven mode with a sixty-second scheduler poll, three maximum transitions per tick, automatic low-risk candidate handling under the stable governor, held-out evaluation, isolated worktrees, and append-only evidence.
  • The first post-restart everyday-interaction audit completed at 2026-08-15T15:41:32+00:00. Travel planning, transit directions, current information, and follow-up continuity all passed; no production endpoint, user memory, or external action was used.
  • The queue retained eighteen actionable records, including internal reliability and external innovation work, but candidate advancement initially failed closed because the repository was dirty.
  • Committed the completed model benchmark and Qwen3.8 runtime work as e5c4897 after full validation, then recorded an append-only circuit-breaker reset. The breaker reports open false and zero consecutive infrastructure failures.
  • At 2026-08-15T15:47:32+00:00, the next clean-baseline tick opened an investigation on the highest-ranked arithmetic-response incident, reproduced its failure, recorded potential_confirmed against e5c4897, and advanced it to candidate construction.

Independent model qualification preservation

Completed
  • Committed both frozen benchmark generations, their hash-verified assets, append-only ledger code, command-line runners, Models observatory APIs and interface, and qualification documentation with the runtime migration.
  • The frozen benchmark packages do not use the production router or active-model configuration, so changing Hiro to Qwen3.8 does not rewrite the instrument or historical Qwen3.6 baseline.
  • Qwen3.8's production selection is an operator runtime decision, not a retroactive benchmark qualification: the canonical ledger still records its hard-gate failures separately.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused Qwen3.8 runtime tests passed Twenty-six launcher, profile, health-selection, nightly-preflight, and operational-boundary tests passed.
Real Windows child-process launcher passed The pinned launcher started and verified qwen/qwen3.8-27b with context length 8192, parallelism 1, maximum GPU offload, and a successful warmup response.
Full Hiro repository suite passed All 576 tests passed in 177.94 seconds after the final persistent-default and documentation changes.
Live service health passed Hiro's task API, both API listeners, Evaluation Observatory, and continuous-improvement reporting endpoints responded successfully after restart.
First Qwen3.8 autonomous audit passed Four of four rotating interaction cases passed in 30.222 seconds, with zero external actions, no production endpoint use, and no user-memory use.
Promotion safety recovery passed The known work was committed as e5c4897, git status became clean, and the continuous-improvement circuit breaker was reset from three blocked-dirty-repository failures to open false with zero consecutive failures.
First clean-baseline promotion transition passed The scheduler recorded investigation_started and potential_confirmed for incident-fbdbc95b249d11b648ba7d00, then moved it into candidate state with construct_isolated_candidate as the next action. Qwen3.8 was actively generating for that work while Hiro remained healthy.

Current state

Next steps