Hiro development journal

Giving Hiro broad autonomous sandbox evaluation with controlled promotion

Validated and published Machine-readable JSON

Executive summary

Implemented the user-selected operating model for Hiro: broad autonomous candidate scope, broad autonomous evaluation gates, broad retained evidence, and a controlled promotion boundary.

The nightly self-improvement path can now hand complete code-changing specifications into Hiro's existing Stage 2 through Stage 4 pipeline, where candidates are built and evaluated in external Git worktrees.

Every autonomously constructed candidate is preserved for review, including failed and rejected attempts, and the benchmark dashboard now has a dedicated Autonomous changes view.

This workflow has no integration or deployment authority. It does not invoke Stage 5, mutate Hiro's active branch, restart services, or enable Stage 6 automatic promotion.

Work completed

Autonomous sandbox orchestration

Implemented and focused tests passed
  • Added a hands-off runner that composes the existing baseline coordinator, external-worktree candidate builder, and proposal-only candidate evaluator.
  • Each candidate retains one falsifiable hypothesis, exact allowed paths, a required targeted test, immutable baseline evidence, and external public and held-out evaluation identities.
  • The runner processes multiple eligible specifications sequentially, preserves every result, and ranks retained candidates using eligibility and public/held-out score deltas.
  • The completed run packet and its SHA-256 companion are written outside Hiro's repository.

Fail-closed operating boundary

Implemented and focused tests passed
  • The sandbox refuses to construct candidates from a dirty Hiro repository because that would make the evaluated baseline ambiguous.
  • An external held-out vault remains mandatory. Missing held-out evidence blocks the run rather than weakening the evaluation gate.
  • Candidate construction continues to use the existing scope audit, protected-path rules, targeted test requirement, bounded repair attempts, and frozen packet verification.
  • The new runner never imports or invokes the Stage 5 integrator and records promotion, deployment, restart, and live-repository mutation authority as false.

Nightly hands-off execution

Configured; awaiting the next clean eligible run
  • Enabled the existing nightly self-improvement schedule and connected its completed summary to the autonomous sandbox runner.
  • The handoff considers complete code-changing specifications only; diagnostic-only ideas remain proposals until their causal hypothesis and executable scope are complete.
  • The current Hiro working tree contains pre-existing development changes, so the new runner will pause safely until the repository presents a clean, unambiguous baseline.
  • No scheduler process or service was restarted during this session, and no autonomous candidate was launched.

Benchmark-page candidate review

Implemented and focused tests passed
  • Added a read-only sandbox-candidates API that joins autonomous experiment metadata, append-only events, frozen packet locations, and proposal decisions.
  • Added an Autonomous changes dashboard view showing every autonomously tested candidate, including eligible queued, rejected, incomplete, and construction-stage states.
  • Each expanded record displays the hypothesis, exact allowed paths, decision evidence, packet location, and gate/event history.
  • The page intentionally has no merge, deploy, restart, or promotion action.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Autonomous sandbox, scheduler, dashboard, candidate pipeline, and Stage 6 boundary tests passed 42 focused tests passed in 42.08 seconds. Coverage included sandbox policy preparation, disabled and no-candidate behavior, scheduler idempotency, dashboard candidate enumeration, coordinator handoffs, candidate construction/evaluation, and unchanged Stage 6 disabled-policy assertions.
Python syntax compilation passed The autonomous runner, configuration, scheduler, and evaluation API modules compiled successfully.
Repository-wide regression suite passed 319 tests passed in 133.88 seconds after the focused validation completed.
Initial test-runtime attempt invalid environment attempt The bundled base Python runtime did not include pytest, so it executed no tests. Validation was rerun with Hiro's existing dashboard virtual environment, where all 42 focused tests passed.
Hiro journal generation and frontend build passed Timestamped-entry unit tests passed; the generator produced and validated 65 journal pages, and the TypeScript and Vite production build completed successfully.

Current state

Next steps