Hiro development journal

Autonomous controller campaign stops at corroboration before build

Phase 3 not demonstrated; 16 bounded non-meta candidates reached explicit corroboration failures Machine-readable JSON

Executive summary

Qualified the autonomous controller upstream of the already-proven promotion actuator without changing promotion, activation, probation, finalization, or rollback behavior.

Added one authoritative funnel projection over the existing continuous-queue rows and append-only event ledger. Each candidate now exposes discovery source, proposed change, expected benefit, build, evaluation, regression, corroboration, governor, promotion request, final outcome, and a normalized terminal reason.

Fixed the first observed controller divergence: a stale qualification status containing multiple reason clauses did not trigger the existing automatic refresh because the code required an exact whole-string match.

Fixed a second evidence-driven discovery boundary: excluded meta candidates were still counted toward the ordinary non-meta backlog, preventing new in-scope candidates from being admitted.

Ran a clean, bounded campaign on one qualified revision with no synthetic candidate, no manual candidate selection, and the production controller's ranking and transition logic.

The complete frozen ordinary source reservoir contained sixteen distinct non-meta lineages, so twenty terminal outcomes were not available without fetching or fabricating more source material.

All sixteen candidates were discovered, selected, retried within the bounded policy, and terminally rejected at external corroboration. None reached build, evaluation, governor, promotion request, activation, promotion, or rollback.

The corroboration truth table behaved as intended. Eleven candidates lacked a complete structured claim; the other five had structured identifiers but no candidate-specific executable local reproduction. Healthy generic controls therefore did not establish a causal local gap.

Phase 3 is not demonstrated. The evidence supports 'insufficient viable candidates; no downstream controller or actuator defect established,' not a claim of autonomous or recursive self-improvement.

Work completed

Authoritative autonomous funnel

Completed
  • The existing continuous-improvement SQLite queue remains the only candidate state authority. The new funnel is computed from its ideas table and append-only events table rather than maintained as a second workflow database.
  • Funnel milestones are DISCOVERED, SELECTED, BUILT, EVALUATION_PASS, REGRESSION_PASS, CORROBORATION_PASS or N/A, GOVERNOR_PASS, PROMOTION_REQUESTED, ACTIVATED, PROMOTED, ROLLED_BACK, and TERMINAL_REJECTED.
  • Per-candidate records expose source, origin kind, proposed change, expected behavior and metrics, current state, gate results, final outcome, normalized terminal reason, event count, and last event.
  • Normalized terminal categories distinguish discovery rejection, duplicate, build failure, no measurable gain, regression failure, corroboration failure, governor rejection, stale candidate, builder-version mismatch, verification failure, and promotion failure where the underlying canonical evidence supports them.

Bounded campaign evidence capture

Completed
  • Added a campaign marker that records an immutable event-ledger boundary and source revision. It has no authority to create, seed, rank, select, build, approve, or promote candidates.
  • Campaign membership is derived from ideas queued or selected after the boundary. Reports reuse the canonical funnel projection and expose target progress and remaining in-scope actionable work.
  • A campaign now ends explicitly as source_exhausted when the target is unmet and no ordinary non-meta candidate remains actionable, avoiding an ambiguous permanent running state.

First divergence: automatic qualification refresh

Repaired and verified
  • The live controller had thirty queued records and zero active work because its required clean-revision qualification was stale.
  • The status correctly reported both a revision change and an earlier incomplete qualification, but refresh logic compared the combined prose string against a set of single exact reasons.
  • The minimum repair parses the status into explicit reason clauses and refreshes when any recognized refresh reason is present.
  • The real active-loop entrypoint then automatically ran all five qualification scenario groups and authorized queue advancement on the current clean revision.

Non-meta campaign scope

Completed
  • The active policy now limits this phase to ordinary non-meta candidates whose approved implementation surfaces do not touch Hiro's improvement subsystem.
  • Already-active lifecycle work would still be allowed to finish safely, while new queued meta records are excluded from selection.
  • No candidate was allowed to modify the controller, evaluator, governor, promotion machinery, or improvement machinery during the campaign.

First divergence: backlog admission scope

Repaired and verified
  • After the first two ordinary candidates terminated, discovery admitted no replacements because twenty-eight excluded meta records still counted toward the global external backlog target.
  • The minimum repair applies the same ordinary non-meta scope to reservoir filtering, active-backlog counting, represented-lineage accounting, and admissions.
  • After requalification, the live controller admitted the remaining ordinary reservoir lineages without manually selecting any candidate.

Final bounded autonomous campaign

Completed without promotions
  • The final campaign ran on revision 06b86c65832826144beacb71b11b94763706d772 with zero seeded candidates and zero manual selections.
  • Sixteen candidates were discovered and sixteen were selected. Sources were fourteen prompt-safe Moltbook-derived leads and two arXiv-derived leads.
  • Mechanisms represented were six tool-routing, three context-memory, three latency-freshness, two retrieval-freshness, one decision-traceability, and one assistant-interaction candidate.
  • All sixteen candidates completed three persisted defer events followed by a fourth corroboration attempt and an explicit terminal rejection: forty-eight defer events and sixteen terminal events in total.
  • The campaign ended source_exhausted with zero candidates left silently active or waiting.

Corroboration semantics

Policy correct; source evidence insufficient
  • The explicit truth table remains: supporting evidence plus healthy regression is PASS; contradictory evidence is FAIL; absent evidence plus healthy regression is DEFER; supporting evidence plus a new regression is FAIL; unavailable evidence is DEFER.
  • The regression suite being healthy was not interpreted as contradictory evidence. It was correctly treated as absence of a reproduced local gap.
  • Eleven candidates lacked one or more of claim identity, reproduction protocol, expected direction, and source fingerprint. Five carried those labels, but their generic control suites passed and did not execute a candidate-specific causal reproduction.
  • Forcing those candidates past corroboration would have weakened the gate. No threshold or truth-table change was made.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused controller and funnel suite passed Seventy-six focused continuous-engine, active-loop, and campaign tests passed after the scope-aware admission repair. Additional campaign closeout tests passed after the reporting-only change.
Corroboration truth table passed The five intended PASS, FAIL, and DEFER combinations are covered and matched implementation behavior.
Final clean-revision pipeline qualification passed All five required qualification scenario groups passed on revision 06b86c65832826144beacb71b11b94763706d772; the packet authorized live queue advancement.
Final autonomous campaign completed_no_promotion 16 discovered, 16 selected, 0 built, 0 evaluation passes, 0 regression passes, 0 corroboration passes, 0 governor passes, 0 promotion requests, 0 activations, 0 promotions, 0 rollbacks, and 16 terminal corroboration failures.
Retry termination and history passed All sixteen candidates retained three explicit defer events and one terminal event; no candidate remained actionable or entered an unbounded requeue loop.
Active runtime identity passed Hiro restarted healthy with the model connected and loaded revision equal to checkout revision 06b86c65832826144beacb71b11b94763706d772.
Public journal tests and build passed npm run test:hiro passed. npm run build generated and validated 169 journal pages, compiled TypeScript, and completed the Vite production bundle.

Current state

Next steps