Hiro development journal

Hiro returns online while Stage 6 remains fail-closed

Core runtime healthy; replacement approval pending Machine-readable JSON

Executive summary

Hiro was restarted after the user's planned maintenance interval, with local Qwen brought online and verified before Hiro's API was allowed to start.

Qwen's checked-in launcher twice left its command-line load client waiting indefinitely. The first attempt also accepted a prompt but did not complete within four minutes. Both attempts were contained while Hiro remained offline.

After one clean supported unload and server restart, LM Studio reported the exact Qwen 3.6 model loaded and idle. The stuck command-line wrapper was terminated without unloading the model, and a minimal real inference completed in 24.5 seconds with the expected model identity.

Hiro then started successfully and reported healthy API status, connected local inference, and the exact expected Qwen model. Stage 6 metrics matched the frozen replacement baseline.

Daylab and Telegram are running again. All Hiro scheduled tasks and the benchmark tunnel remain disabled.

The prior Stage 6 candidate was rejected before fast-forward, so there is no probation to resume. The replacement candidate remains frozen and requires a new exact approval before a fresh 72-hour, one-promotion window can begin.

Work completed

Qwen cold restart

Healthy after bounded recovery
  • Started the pinned Qwen 3.6 profile through the checked-in launcher and preserved its model-identity and inference requirements.
  • The first launcher did not return from the LM Studio load client, and a direct warm-up remained in prompt processing beyond its 240-second request timeout.
  • Performed one supported model unload and local-server restart, then retried the exact profile with an extended readiness allowance.
  • The second load client also failed to return even though LM Studio reported the model loaded and idle. Only the verified stuck client wrapper was terminated.
  • A one-token inference probe completed in 24.5 seconds and returned the exact expected model identity, proving that local generation was operational.

Hiro core restart

Healthy
  • A first background launch failed before startup because PowerShell split the repository path at its space. The process exited without opening a listener.
  • Restarted with the entry-point path explicitly quoted and the inherited Windows search path normalized.
  • Hiro's local health endpoint returned status ok, local inference connected, and model identity qwen_qwen3.6-35b-a3b.
  • The Stage 6 metrics endpoint returned 100 observations, error rate 0.04, and p95 latency 10,022 ms.

Auxiliary services

Partially restored by design
  • Restarted Daylab at its configured 30-minute cadence.
  • The first Telegram start exited on a Windows console-encoding error. Restarting with UTF-8 mode produced a running bot and enabled-notifications marker.
  • Telegram emitted a nonfatal logging-rollover error because its existing log file was in use, but the bot process remained running.
  • The externally accessible benchmark tunnel was intentionally not restarted, and no scheduled task was re-enabled.

Stage 6 boundary

Disabled
  • Verified the replacement candidate remains the only packet in the Stage 6 inbox and the persistent DISABLED sentinel remains present.
  • The Stage 6 scheduled task remains disabled and no promotion or probation is active.
  • The prior approval cannot authorize the replacement because the corrected revision and packet hash differ from the rejected identity.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Qwen model identity passed The local model endpoint and real generation response both identified qwen_qwen3.6-35b-a3b.
Qwen generation passed A bounded one-token inference request completed in 24.5 seconds after recovery.
Hiro health passed The local health endpoint returned status ok with connected local inference and the expected model.
Stage 6 monitoring baseline passed The restarted service reported 100 observations, error rate 0.04, and p95 latency 10,022 ms, matching the frozen replacement evidence.
Daylab runtime passed The managed Daylab parent and child processes were running at the configured cadence.
Telegram runtime passed with warning The bot process and notifications marker were present after a UTF-8 restart; a nonfatal log-rollover sharing error remains.
Stage 6 controls passed The scheduler remained disabled and the persistent sentinel remained present throughout restart.

Current state

Next steps