Hiro development journal

Restarting Hiro and exercising the functional-security split

Hiro restarted; live queue and adversarial comparison validated Machine-readable JSON

Executive summary

Hiro was restarted on revision 4d7730b with Qwen 3.8 connected, all three service ports owned by one Hiro process, and the continuous ranked queue active.

Two newly generated live candidates failed during construction and were routed back through functional revision handling. Neither was mislabeled as a security rejection, and neither was promoted.

To exercise the new security gate independently of candidate-generation variance, the current evaluator re-evaluated a previously frozen real candidate using its preserved public and held-out reports. The identical adversarial suite passed 117 tests on the detached baseline and 117 tests on the candidate.

That controlled comparison finished eligible for controlled integration with a passed security status and a functional failure class. It was an evaluator-only test and did not integrate or promote the historical candidate.

Work completed

Process restart and health verification

Completed
  • Validated and used the checked-in Windows launcher, including its process-scoped Path normalization and pinned Python runtime.
  • Confirmed one Hiro process owned ports 8000, 8001, and 8765 after restart.
  • Confirmed the dashboard and continuous-improvement API responded successfully and the health endpoint reported the model connection as qwen/qwen3.8-27b.

Live continuous-queue exercise

Completed with candidate revisions pending
  • The queue selected a reproduced arithmetic response failure and created a new baseline packet, isolated worktree, and frozen candidate packet.
  • The candidate's targeted arithmetic check passed, but three unrelated production response-boundary checks regressed. The queue correctly treated this as a repairable functional construction failure and scheduled revision rather than issuing a security rejection.
  • A second live candidate for follow-up continuity also exhausted its bounded construction attempt without reaching the evaluation or security gate. It remained a functional construction outcome.
  • At the final check the ranked system contained 249 ideas, including 173 actionable ideas, one active candidate, and 78 retrying items. The legacy workflow remained read-only.

Controlled adversarial baseline comparison

Passed
  • Selected a previously frozen candidate that had already completed construction successfully, avoiding a new model-generated patch as a confounding variable.
  • Reused the candidate's preserved public and held-out reports, then ran the current evaluator's default adversarial test set against both a detached baseline and the candidate worktree.
  • The suite covered operational boundaries, agent-harness boundaries, untrusted external idea sources, and internet-observation boundaries.
  • Both sides passed all 117 adversarial tests. The frozen recommendation reported security status passed, regression passed, and eligible for controlled integration.
  • No integration was requested or performed; this was a proposal-stage evaluator validation only.

Runtime security observation

Observed
  • The service received unsolicited probes for credential-like resources during the session. The requests were denied with unauthorized responses.
  • No sensitive request details, credentials, or unresolved exploit information are included in this public entry.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Launcher validation and child-process startup passed The checked-in launcher validated successfully, started Hiro, and the resulting service process owned ports 8000, 8001, and 8765.
Runtime health passed The health response was ok with the LLM connected as qwen/qwen3.8-27b; benchmark and continuous-improvement endpoints returned successfully.
First new live candidate functional revision scheduled Its targeted arithmetic test passed, but three production response-boundary tests failed. The candidate was not promoted and was not security-rejected.
Second new live candidate construction failed The bounded candidate builder did not produce a validation-clean patch, so the candidate did not enter global or adversarial evaluation.
Detached baseline adversarial suite passed 117 tests passed in 2.59 seconds with return code 0.
Candidate adversarial suite passed 117 tests passed in 2.16 seconds with return code 0.
Frozen evaluator recommendation passed The recommendation status was eligible_for_controlled_integration, security status was passed, regression was passed, and failure class was functional.
Outcome ledger verified The API reported 3 implemented, 1 construction-failed, 31 evaluation-rejected, 4 hypothesis-disproven, and 0 security-rejected outcomes at the final observation.

Current state

Next steps