Hiro development journal

Live interaction failures now drive Hiro's continuous improvement loop

Implemented, activated, and verified on the live local service Machine-readable JSON

Executive summary

Hiro's autonomous improvement workflow now treats real interaction failures as higher-priority work in the same ranked queue used for external upgrade ideas.

The previously diagnosed Los Angeles event-discovery failure and its failed correction were harvested as separate private incidents, reproduced with deterministic local contracts, repaired, and recorded as implemented outcomes.

Automatic Self-Improvement V2 production-chat probes were retired. External discovery remains active, while candidate construction, evaluation, and probation wait for a quiet production window and never use the live chat endpoint as a test harness.

Future candidates use one eight-hour probation contract with checkpoints at 0, 60, 240, and 480 minutes. Routine candidates retain standing authority; the dashboard reports the ranked idea, origin lane, next action, affected files, test gates, deadline, and terminal result.

A live acceptance test initially exposed an additional public-event routing defect: the word 'event' sent a public discovery request into the private-calendar resolver, and correction escalation then produced plausible but unsupported recommendations. That defect was repaired during this session rather than accepted as a passing test.

The final live acceptance returned concrete August activities directly from current web-search evidence with source URLs. Its universal quality record contained one web-search tool receipt, three evidence records, a passing final gate, and no repair request.

A concurrent worker later attempted to overwrite an authorized implemented outcome with a stale rejection. Queue transitions now use expected-state compare-and-set protection, and the affected incident was reconciled with an append-only audit event.

Work completed

Unified incident intake and ranking

Activated
  • Added first-class origin kinds and three queue lanes: user corrections and reported interaction failures first, automatic reliability findings second, and external innovation third.
  • The incident harvester reads response-quality records without model calls, re-evaluates historically misclassified responses against current deterministic rules, redacts sensitive text, stores private replay fixtures locally, and exposes only a safe replay summary to the dashboard.
  • Queue records now include next action and deadline fields, watchdog recovery for abandoned work, lane-aware ordering, and revision migration for unfinished candidates built against an older live branch.
  • Both bootstrap interaction incidents are recorded as implemented. Their private replay text is not exposed by either the queue-read endpoint or the priority-control response.

Isolated interaction lab

Activated
  • Added deterministic replay contracts for general interactions and event discovery. Replays execute no production endpoint, model, tool, or external instruction.
  • The event-discovery contract requires substantive content, at least two concrete recommendation signals, a web-search receipt, and a nonzero evidence count.
  • A hostile-evidence control verifies that prompt-injection-like text remains inert data and that responses following such instructions fail the contract.
  • The original source-title response and three-character correction response fail locally; a concrete evidence-backed response passes the same code-owned contract.

User-facing response repair

Activated
  • The router no longer exposes the last period-delimited fragment of hidden reasoning when visible model content is empty.
  • The universal response gate rejects URL fragments, source-title-only discovery answers, unrecovered corrections, and current-event recommendations without web evidence.
  • False date-conflict retries were removed from the generic result comparer; complete temporal conflicts remain handled by the event-aware path.
  • Public event discovery no longer enters the private-calendar OAuth resolver merely because the query contains the word 'event'. Personal-calendar language is now required.
  • Initial instructions such as asking for actual events rather than source names no longer trigger correction escalation when no prior assistant response exists.
  • Current event lists are rendered directly from search evidence instead of free-form synthesis. Boilerplate, duplicate mobile date blocks, common encoding artifacts, and visibly truncated snippets receive deterministic cleanup.

Continuous execution and production isolation

Activated
  • The one-minute active loop continues discovery ingestion during normal use but defers candidate construction, evaluation, canary checks, and promotion until production has been quiet for three minutes.
  • The default discovery job now collects prompt-safe external ideas only. It does not run evaluations, write development journals, send notifications, or call Hiro's production chat endpoint.
  • The automatic scheduler refuses the retired Self-Improvement V2 overnight mode, and its former live regression-probe phase reports zero production probes.
  • The legacy DayLab and NightLab paths remain available only as explicit manual diagnostics; they are not automatic upgrade paths.
  • The continuous governor now enforces the same eight-hour probation and exact 0, 60, 240, and 480 minute checkpoint sequence for low- and moderate-risk candidates.

Universal feedback boundary

Activated
  • Response-quality logging now occurs once at the universal final boundary rather than only in selected resolver branches.
  • Every user-facing result records its resolver, tools used, evidence count, final-gate decision, and validator reason using a safe metadata allowlist.
  • Those receipts allow future ungrounded current-event answers to become automatic queue incidents instead of being counted as successful interactions.
  • The final live acceptance record showed resolver model_loop, tool web_search, evidence count three, validator pass, and repair_needed false.

Queue concurrency and audit integrity

Activated
  • A live race demonstrated that a worker started before a newer platform decision could finish later and overwrite the queue state.
  • Autonomous transitions now declare their expected source state. A mismatched or protected terminal transition is ignored and recorded as stale_transition_ignored.
  • An explicit terminal override remains available for authorized reconciliation; it was used once to restore the affected bootstrap incident to its correct implemented state.
  • The queue circuit breaker is closed, its infrastructure-failure count is zero, and both interaction incidents have terminal implemented outcomes.

Observatory realignment

Activated
  • The benchmark page describes one continuous evidence-driven process instead of presenting retired labs as active workflows.
  • The pipeline view shows standing worker authority, eight-hour probation, lane order, production isolation, and one-at-a-time branch mutation.
  • Idea cards show category and factor scores, total priority, origin lane, next action, deadline, private-replay availability, affected files, test gates, and rejection or implementation evidence.
  • User controls can boost or make an idea next, but cannot bypass isolated tests, probation, or the governor.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Initial focused integration suite passed 47 tests passed across interaction replay, continuous queue, governor, active loop, scheduler retirement, response envelope, and evaluation dashboard behavior.
First full repository suite passed 537 tests passed in 155.70 seconds after the unified incident queue and production-isolation implementation.
Grounded event-discovery full repository suite passed 540 tests passed in 172.98 seconds after public-event routing, correction-context, and web-evidence enforcement changes.
Universal feedback full repository suite passed 540 tests passed in 181.77 seconds after moving quality logging to the universal response boundary and adding evidence receipts to replay contracts.
Final bounded interaction and concurrency suites passed 23 focused response tests passed after evidence-snippet cleanup, and 30 focused queue/interaction tests passed after expected-state concurrency protection.
Live event-discovery acceptance passed The live local chat route returned current August activities with source URLs. The recorded response-quality row contained web_search, three evidence records, validator pass, and no suspected failure.
Live queue and service health passed Hiro restarted successfully through the detached no-window launcher. The queue circuit breaker is closed, both bootstrap incidents are implemented, and the configured probation is 480 minutes with 0/60/240/480 checkpoints.
Notification and shutdown state passed The Telegram enablement marker is absent and the intentional-server-stop marker is absent; Hiro is active without Telegram notifications.

Current state

Next steps