Hiro development journal

Hiro is offline with fast risk-based candidate validation ready

Implemented and fully tested; intentionally offline Machine-readable JSON

Executive summary

Hiro, its continuous scheduler, all local HTTP listeners on ports 8000, 8001, and 8765, and the loaded Qwen3.8 model were stopped before the validation policy changed. No queue work could advance during implementation or testing.

The inherited eight-hour pre-promotion canary was replaced by short risk-based repeated-probe schedules: low-risk candidates validate at 0, 5, and 15 minutes; moderate-risk candidates validate at 0, 5, 15, and 60 minutes.

Targeted baseline contrast, candidate-authored tests, public and held-out evaluation, invariant checks, category non-regression, repeated cross-suite latency checks, the final full repository suite, protected-path enforcement, clean-tree validation, fast-forward-only promotion, and rollback controls remain unchanged.

The pre-change active candidate was cancelled, its candidate and canary payloads were cleared, and it was returned to the queue with an append-only policy-change event. All 599 Hiro tests passed after the final implementation. Hiro and the model remain stopped.

Work completed

Controlled shutdown

Completed
  • The running Hiro process was verified by exact command line before termination.
  • After shutdown, no listeners remained on ports 8000, 8001, or 8765.
  • LM Studio unloaded qwen/qwen3.8-27b and reported no loaded models.
  • Hiro was not restarted during validation and will remain offline after this session.

Fast risk-based validation

Completed
  • The protected continuous-governor policy now authorizes a 15-minute low-risk schedule with checkpoints at 0, 5, and 15 minutes.
  • Moderate-risk candidates now use a 60-minute schedule with checkpoints at 0, 5, 15, and 60 minutes.
  • The queue derives checkpoint schedules from the candidate's frozen risk classification and calculates its deadline from the final required checkpoint plus five minutes.
  • The governor and queue share one code-owned schedule definition, preventing policy, execution, and promotion validation from drifting.
  • Unknown and high-risk candidates remain outside autonomous authority.

Visibility and recovery

Completed
  • The continuous-improvement API now exposes both low- and moderate-risk validation schedules.
  • The benchmark dashboard describes the low 15-minute and moderate 60-minute schedules instead of the obsolete eight-hour checkpoints.
  • Investigation records now persist their exact base revision inside candidate metadata so a later platform revision can reliably invalidate unfinished work.
  • The active pre-change prompt-injection-resistance candidate was returned to queued, all provisional candidate/canary/outcome data was cleared, and the event validation_policy_changed_requeued preserved the reason without counting the policy cancellation as a behavioral rejection.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused governor, queue, API, dashboard, and interaction suite passed All 61 focused tests passed after both the validation-schedule change and base-revision recovery fix.
First authoritative full Hiro suite passed All 599 tests passed in 151.89 seconds after the risk-based schedule implementation.
Final authoritative full Hiro suite passed All 599 tests passed again in 150.63 seconds after adding investigation base-revision persistence.
Old-run cancellation passed The exact active candidate changed from candidate to queued; candidate and canary payloads are null; next action is reproduce_locally; the audited policy-change event is the latest lineage event.
Offline state passed Hiro listeners were zero after shutdown and LM Studio reported no loaded models.

Current state

Next steps