Hiro development journal

RSI obstacles become recoverable, testable evidence

Implemented, validated, and live Machine-readable JSON

Executive summary

Hiro no longer treats every failed candidate transaction as evidence that an idea was bad. The queue now distinguishes construction failure, evaluation rejection, safety rejection, infrastructure blockage, disproven hypotheses, legacy outcomes, supersession, and implementation.

Candidate repair attempts now restart from a clean isolated baseline, receive the prior failure evidence, and must produce a complete alternative candidate. Python candidates can use AST-bounded top-level symbol replacement instead of fragile text matching.

Continuous candidates now prove targeted improvement by running the candidate-authored test against an untouched baseline: it must execute and fail assertions there, then pass on the candidate. Public and held-out suites enforce global non-regression, invariants, categories, and repeated cross-suite latency evidence.

A shadow replay found eleven retained candidates that satisfy the corrected targeted and global evidence contract. Four were rejected because baseline already passed, and twenty-four remained inconclusive rather than receiving false credit.

A real Qwen3.8 retry of a previously failed memory candidate produced a candidate-ready packet after one clean repair. All 597 repository tests passed and the live circuit breaker remains closed.

Work completed

Persistent candidate recovery

Completed
  • Each non-terminal builder repair now performs a validated hard reset and untracked-file cleanup inside the exact external candidate worktree before asking for a new strategy.
  • Repair prompts state that the workspace is clean and require a complete candidate rather than an incremental patch against discarded code.
  • The builder supplies file-existence state, compact prior validation evidence, exact relevant symbol context, and instructions that distinguish create, modify, append, and symbol replacement operations.
  • AST-bounded replace_symbol changes only one explicitly named top-level Python function or class inside an allowlisted file and then passes through the existing patch-safety and scope checks.

Targeted improvement with global non-regression

Completed
  • Continuous RSI candidates use a dedicated evidence mode rather than the general laboratory's global-superiority policy.
  • The evaluator creates a temporary detached baseline worktree, copies only the candidate's targeted tests into it, and requires an executable assertion failure. A passing baseline proves no improvement; collection or usage failures are inconclusive.
  • The same tests must pass on the candidate alongside the stable regression suite.
  • Public and held-out score may vary by at most one weighted case, category regression remains capped at 0.03, candidate invariant failures remain forbidden, and latency rejection requires the regression to repeat in both suites.
  • The general promotion policy remains global superiority by default; the continuous policy and all thresholds are explicit in the active policy and frozen proposal evidence.

Failure taxonomy and observability

Completed
  • Frozen builder attempts are summarized into specific obstacle classes including stale anchors, duplicate creation, empty plans, patch-application failures, scope rejection, and safety rejection.
  • Queue API records now expose outcome_category and aggregate outcome_counts without rewriting historical evidence.
  • The Ranked Ideas Rejected view displays category totals and each card carries its derived outcome badge.
  • The live 120 rejected records resolve to 86 construction failures, 28 evaluation rejections, and six legacy rejections; none are silently relabeled as successful.

Safe historical shadow replay

Completed
  • Added a reusable command that replays retained sandbox evidence against the corrected global policy and optionally runs targeted baseline contrasts.
  • Shadow replay never promotes, changes queue state, or rewrites old decisions.
  • Of 74 candidate-ready historical packets, 39 passed corrected global non-regression. Eleven also demonstrated executable targeted baseline failures, four already passed on baseline, and 24 produced inconclusive baseline collection results.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused recovery suite passed Eighty-seven focused tests passed for clean repairs, AST symbol replacement, builder diagnostics, queue taxonomy, continuous policy, targeted baseline contrast, latency confirmation, shadow replay, dashboard rendering, and active policy loading.
Historical shadow replay passed Replayed 74 candidate-ready packets without promotion or queue mutation: 39 cleared global non-regression, eleven proved targeted improvement, four already passed on baseline, 24 were inconclusive, and zero contrast executions raised infrastructure errors.
Real Qwen3.8 construction retry passed A previously failed typed-memory hypothesis was rebuilt from the same 8e19fd0 baseline. The first candidate failed validation; one clean repair returned a complete alternative whose syntax, collection, and targeted tests passed, producing a candidate-ready frozen packet.
Authoritative full Hiro suite passed All 597 tests passed in 186.75 seconds in Hiro's real pinned environment.
Live restart and API passed Hiro restarted at commit a0b456f, ports 8001 and 8765 returned on one validated process, eight actionable ideas were admitted, and the circuit breaker remained closed with zero failures.
Live dashboard passed The Rejected filter visibly displayed construction, evaluation, safety, and disproven-hypothesis totals plus classified historical cards.

Current state

Next steps