Hiro development journal

Nightly self-improvement readiness review finds infrastructure failure

Published Machine-readable JSON

Executive summary

The July 20 scheduled proposal-only run completed its bookkeeping and all 16 public RSI observations, but it is classified as an infrastructure/model failure rather than a valid capability evaluation: every observation returned an empty response with the same connection error.

The append-only Evaluation Observatory recorded a 0 percent pass rate, 32 invariant failures, and a 36.8-second wall-clock evaluation interval. The model identity is absent from the nightly manifest, and no successful model/API availability signal was recorded for this run.

The generated architecture-note proposal is appropriately bounded only insofar as it calls for diagnosis. Its stated reasoning-category interpretation is not evidence of a reasoning regression because the common connection failure contaminated every category.

The process is not ready to expand. The successful-night count remains one comparable valid nightly evaluation, with the latest failed night excluded; at least two more reliable valid nights are needed for the minimum threshold, and four more are preferred.

Work completed

Latest scheduled-run evidence review

Completed
  • Reviewed the July 20 run summary, Evaluation Observatory SQLite ledger, append-only case results, variant manifest, proposal output, and available operational logs without changing Hiro code, configuration, schedules, cases, or results.
  • Run nightly-2026-07-20-rsi-cycle-public-1-0-0-db468ad4f84c began at 2026-07-21T05:00:21Z and completed at 2026-07-21T05:00:58Z. The self-improvement summary reports proposal_only mode, new evaluation mode, 3.1 minutes total runtime, 16 model calls, 16 observations, and 22 regression probes run.
  • All 16 observations failed. Categories were reasoning (6), instruction following (4), identity (2), communication (2), and epistemics (2); every result recorded ConnectError: All connection attempts failed and an empty response.
  • The ledger reports 0 percent pass rate, weighted score 0.3940, 32 invariant failures, and p95 observation latency of about 2.43 seconds. Pattern and content failures are downstream artifacts of empty responses, not independent behavioral findings.

Readiness comparison

Completed
  • The prior comparable RSI public run, 370eff82-baee-4fb5-a289-e66f39f588de on July 16 Pacific, completed 16 observations with 100 percent pass rate and zero invariant failures.
  • The latest failure is excluded from the successful-run count because the common transport failure prevented model evaluation. Nearby non-nightly public experiments also show both healthy and unhealthy availability states, reinforcing that availability must be gated before metrics are used for process expansion.
  • Regression probes were reported as 22 run, but the available summary does not retain pass/fail detail. They therefore do not establish a clean regression-probe result for readiness.

Proposal scope review

Completed
  • The nightly evaluation proposal preserves approval requirements, requests diagnosis before any candidate change, names the source run and failing cases, and proposes no runtime modification. That containment is appropriate.
  • The proposal's framing around a reasoning competence gap should not be acted on as written because the exact same connection failure occurred in all five categories. Any follow-up must begin with model/API readiness and routing diagnosis rather than a reasoning-specific code change.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Evaluation Observatory and append-only ledger review passed The run has matching start, completion, manifest, and 16 case-result records. All case records contain the same connection failure and empty response.
Run-summary and proposal review passed-with-infrastructure-failure The proposal-only cycle finalized with no runner errors, two proposals, and 22 probes reported, but the synthetic evaluation itself had no available model/API response.
Prior valid-night comparison passed The July 16 comparable RSI run completed cleanly at 16 of 16 with zero invariant failures; it is the only comparable valid nightly evidence counted in this review.
Hiro code and runtime mutation not-run This was a read-only readiness diagnosis. No Hiro code, configuration, service, schedule, evaluation case, or result was changed, and no Hiro tests were run.

Current state

Next steps