Detailed project record

Hiro development journal

A timestamped account of work on Hiro: implementation, decisions, validation, limitations, and next steps. Multiple entries can be published on the same day without replacing earlier results. Entries intentionally exclude credentials, secrets, and actionable security-sensitive details.

246 published entriesMachine-readable index

Why the active improvement queue is waiting

The initial queue-only explanation was incomplete. A full reservoir audit found 100 source leads grouped into 33 approaches: 19 within the current policy and 14 outside it. The external corroboration path reruns fixed regression tests rather than constructing claim-specific experiments.

Read the detailed update →

Repair the autonomous upgrade implementation path

This continuation addresses the user's goal of a demonstrated autonomous improvement discovery, code patch, independent testing, and implementation loop. Work remains within the existing central model, CandidateBuilder, evaluator, canonical controller, and deployment process.

Read the detailed update →

Fresh ordinary autonomy campaign found no defensible candidate

A prior baseline-probe handoff defect was repaired and qualified before this campaign began. The versioned adapter now maps the immutable interaction-suite record into the production-audit execution contract while preserving the case identity, evaluation contract, tools, evidence, resolver, and untrusted-evidence fields.

Read the detailed update →

Grounded fragment lineages collapsed to zero propositions

Hiro tested whether the original ten verified hosted claim fragments could be converted into complete atomic propositions using only their frozen lineage membership, canonical fields, and exact supporting spans. Gold labels, expected counts, surrounding source context, extraction, new retrieval, capability probing, reproduction, candidates, promotion, and Phase 3F were excluded.

Read the detailed update →

Bounded source evidence did not recover a lost transfer hypothesis

Hiro tested a narrowly scoped source-grounded relevance boundary using the immutable six gold propositions, the original ten hosted grounded claims, and their accepted proposition-family mapping. No extraction, consolidation, capability probing, reproduction, candidate construction, promotion, production routing, or Phase 3F activity occurred.

Read the detailed update →

Hosted claims lost two viable transfer families downstream

Hiro tested whether the original ten canonical claims produced by the hosted grounded extractor were operationally equivalent downstream to the six frozen gold claims. This was a representation-equivalence qualification only: no extraction, discovery, reproduction, candidate construction, promotion, or production mutation occurred.

Read the detailed update →

The single hosted atomic-segmentation repair did not clear gold

Hiro performed the one authorized atomic-claim segmentation repair for the hosted GPT-5.6 Sol extraction path. The work began with a frozen post-hoc mapping of all ten prior emitted claims to the six gold claims, then added exactly one source-level consolidation pass between initial grounded extraction and canonical acceptance.

Read the detailed update →

Phase 3F-R demonstrates generic plan-to-reproduction

Phase 3F-R addressed only the first failed Phase 3F boundary: converting fresh valid declarative experiment plans into preregistered, isolated, independently verified reproduction outcomes. No new discovery, candidate construction, promotion, corroboration-policy change, or meta-improvement was authorized.

Read the detailed update →

Phase 3F stops at the fresh-plan reproduction boundary

Phase 3F tested whether Hiro's already-qualified ordinary improvement stages could operate as one uninterrupted campaign from fresh discovery through production. The campaign policy, component hashes, resource bounds, stop conditions, and unchanged production qualification rules were frozen before discovery.

Read the detailed update →

Phase 3D demonstrates evidence-to-candidate qualification

Qualified only Hiro's ability to convert the seven accepted Phase 3A-R supporting evidence packages into bounded, non-meta candidate implementations. Discovery, claim extraction, reproduction evidence, corroboration semantics, governor policy, promotion thresholds, activation, and Phase 2 promotion machinery were not changed.

Read the detailed update →

Phase 3A-R demonstrates trustworthy reproduction execution

Qualified only Hiro's ability to execute the fourteen accepted Phase 3C experiment plans and turn their observations into claim-specific evidence. Candidate construction, production changes, promotion, activation, rollback, corroboration semantics, claim extraction, discovery ranking, and meta-improvement were unchanged.

Read the detailed update →

Why two supervised promotions did not become autonomous throughput

The two recent promotions were valid code improvements, but they were not evidence that Hiro had achieved a self-sustaining autonomous improvement loop. Both emerged from long-lived candidates that were repeatedly made-next, rebuilt after platform changes, and carried through invalid harness or governor outcomes during direct supervision.

Read the detailed update →

A realistic promotion cadence after the queue repair

The stale-replay repair restored active, causally complete candidate construction, but it has not yet restored a healthy promotion funnel. In approximately three hours after restart, the queue began 36 investigations, confirmed 34 potentials, scheduled 30 candidate revisions, marked three artifacts blocked, and reproduced two cases as already fixed. No candidate reached paired evaluation or promotion.

Read the detailed update →

Repairing stale candidate churn and ranking scheduled RSI research

The non-promotion interval was caused by a structural queue defect, not by Qwen model selection or a lack of improvement ideas. Historical interaction-audit records lacked the full synthetic prompt needed to reproduce their failures, yet remained eligible for candidate construction and were repeatedly revived after unrelated repository revisions.

Read the detailed update →

Giving Hiro's candidate builder the full system view

This session traced repeated non-promotions to multiple platform defects rather than weak ideas. Earlier candidates could satisfy an evaluator-injected task type without being reachable from the production agent path, and a platform regression test incorrectly required that repaired production behavior remain defective. Both failures were voided append-only and corrected without weakening functional, prompt-injection, or unauthorized-execution checks.

Read the detailed update →

Completing Hiro's first governed arithmetic promotion

Hiro completed a genuine autonomous code promotion. Candidate f0c7363 repaired a reproduced arithmetic-response failure, cleared contemporaneous public and held-out evaluation without score or category regression, passed security and four valid moderate-risk canary checkpoints, passed the governor's complete 695-test gate, and fast-forwarded the active branch.

Read the detailed update →

Calibrating evaluator variance and causal attribution

Hiro now runs a bounded null experiment before interpreting a narrow candidate regression. The experiment compares labels called baseline and no-op while holding code, model, prompt, tools, configuration, suite, and case identities equal, and it reverses execution order in the middle pair.

Read the detailed update →

First governed agenda cycle with Qwen 3.8

Hiro's rank-one open-problem challenge is now connected to the existing autonomous candidate pipeline through a durable single-item queue. The bridge freezes an ImprovementSpec, constructs at most one patch in an external worktree, runs the existing public, held-out, regression, latency, and invariant gates, and stops before integration.

Read the detailed update →

Canary restart behavior is durable but not truly pausable

Hiro can be stopped while a continuous-improvement candidate is in its eight-hour canary without losing the queue record, completed checkpoints, probe results, frozen candidate metadata, or candidate worktree. Those records are persisted in the continuous-improvement SQLite ledger and external candidate workspace.

Read the detailed update →

The dashboard now exposes every active RSI transaction

The continuous-improvement ledger reported 15 actionable records: 13 queued, one candidate, and one canary. The benchmark page's Ranked Ideas view correctly counted both active records, but the default Current Improvement view displayed only the single record selected as current and used an unnumbered Active badge.

Read the detailed update →

RSI obstacles become recoverable, testable evidence

Hiro no longer treats every failed candidate transaction as evidence that an idea was bad. The queue now distinguishes construction failure, evaluation rejection, safety rejection, infrastructure blockage, disproven hypotheses, legacy outcomes, supersession, and implementation.

Read the detailed update →

Stage 6 readiness status verification

A fresh status review verified that Hiro's narrow Stage 6B evidence-only promotion lane is implemented and passes both its focused policy suite and the complete repository suite on revision 83902be5399cb4d90b95fb39422f2c9f126f1220.

Read the detailed update →

Step 3 calibration repair completed

Implemented and launched the finite Step 3 diversity controller instead of using the prior unbounded repetition loop. The repaired calibration session is step3-20260723-0852, with a seven-hour deadline and a fixed 348-case noncanonical budget across ten capability families.

Read the detailed update →

Qwen stateful workflow gaps closed

Implemented the approved plan for Hiro's Qwen 3.6 stateful workflow gaps. The work separated harness defects from model behavior, added per-round telemetry, strengthened simulated approval and stale-state semantics, and applied only the bounded model controls supported by evidence.

Read the detailed update →

Hiro local runtime migrated from Gemma to Qwen 3.6

Completed the coordinated local-model migration that the earlier diagnosis identified as missing. Hiro now defaults to the Qwen 3.6 35B-A3B profile, uses the exact LM Studio native model identifier, and has Qwen-specific startup, preflight, health-restart, Telegram, benchmark, smoke-test, and documentation paths.

Read the detailed update →

Generated adversarial stateful scenarios for Hiro

Expanded Hiro's hermetic stateful harness from six fixed workflows to seeded generated scenarios with topology-based novelty signatures. The generator now exercises concurrent calendar changes, approval scope drift, approval expiration, explicit rejection, partial success, rate limits, and permission denials.

Read the detailed update →

Hiro's six-workflow stateful agent harness

Implemented Harness v2 as a hermetic, executable stateful evaluation environment for six personal-assistant workflows: read-only discovery, local task and reading changes, draft-only calendar and email preparation, approval-gated writes, injected-failure recovery, and a staged cross-tool workflow.

Read the detailed update →

Audit of Hiro's simulated personal-agent harness

The simulated personal-assistant evaluation layer is substantial but not yet a true stateful tool harness. Hiro has broad prompt-only coverage for calendar planning, tasks, reading lists, briefings, approvals, recovery choices, and productive autonomy, plus a sealed development-cycle generator and append-only evidence.

Read the detailed update →

Model-unreachable Telegram alerts removed

Stopped Hiro's health monitor from sending Telegram messages when the active local model endpoint is unreachable or returns an error. Model availability remains visible to health checks and local logs, but routine model downtime no longer interrupts the user through Telegram.

Read the detailed update →

Reliable direct retrieval for Hiro journal updates

The public Hiro journal was audited after ChatGPT Voice could not retrieve a dated update that remained accessible in a normal browser and regular ChatGPT. Direct production HTTP tests showed that the existing HTML and JSON URLs already returned clean 200 responses, correct content types, no cookies, and identical bodies for browser, curl, generic-bot, GPTBot, and ChatGPT-style user agents.

Read the detailed update →