Hiro development journal

Strategy: make open harness problems a durable research program

Strategy designed; implementation intentionally deferred Machine-readable JSON

Executive summary

Before restarting Hiro or its upgrade process, the research strategy was broadened from repeatedly improving recent behavior to maintaining a durable, evidence-backed program of open problems in AI harness construction.

Hiro already contains important execution machinery: reactive opportunity detection from interaction records, prompt-safe external idea lineage including arXiv and other sources, bounded falsifiable experiments, isolated candidate worktrees, evaluation gates, and governed capability lanes.

The missing strategic layer is durable problem identity and agenda governance across many observations and experiments. Without it, individual fixes and external ideas can be processed competently while the system still loses sight of the largest unresolved capability bottlenecks.

The proposed integration is an Open Problems Registry connected to, but distinct from, the experiment queue. Problems persist across failed hypotheses, accumulate evidence and contradictions, and are periodically reprioritized or reframed as Hiro and the external research frontier change.

Work completed

Current-system fit assessment

Completed as a read-only design review
  • The current opportunity detector prioritizes recent high-severity failures, surprising successes, related weak signals, novel observations, and repeated low confidence. This is useful operational feedback but does not by itself maintain long-lived research questions.
  • The external-idea system accepts prompt-safe reductions from supported sources, preserves lineage, requires local corroboration, and prevents external content from authorizing candidate construction or promotion.
  • The experiment planner and variant laboratory already provide much of the downstream machinery needed once a problem has yielded a falsifiable hypothesis.
  • The current capability-lane policy supplies strong authority and evaluation gates, but its lanes are narrow implementation routes rather than a complete, revisable map of open harness problems.

Open Problems Registry

Architecture proposed
  • Create a durable registry in which each problem has a stable identity, precise capability gap, operational evidence, strategic significance, current hypotheses, unresolved contradictions, relevant external findings, evaluation gaps, and a history of attempted experiments.
  • Treat problems as longer-lived than experiments. A failed experiment should update the parent problem's evidence and hypothesis space rather than close or duplicate the problem.
  • Keep observed symptoms separate from proposed causes and proposed interventions. This prevents an early guess about implementation from becoming the permanent definition of the problem.
  • Allow explicit states such as observing, characterized, experiment-ready, active, blocked-on-measurement, monitoring, provisionally resolved, superseded, and retired, with recorded reasons for every transition.

Research-agenda loop

Operating model proposed
  • Ingest signals from real interactions, evaluation failures, surprising successes, instrumentation, external research, repositories, and human strategic observations into a candidate-problem inbox.
  • Merge evidence into existing problems when it concerns the same underlying capability gap; create a new problem only when the distinction is meaningful and testable.
  • Select a portfolio rather than a single highest score: urgent reliability work, high-leverage capability bottlenecks, measurement and infrastructure work, and bounded exploratory research should each receive deliberate capacity.
  • Translate selected problems into competing hypotheses and discriminating experiments, then return all outcomes, including negative and inconclusive results, to the problem record.
  • Run periodic agenda reviews that can reprioritize, split, merge, reframe, provisionally resolve, or reopen problems based on accumulated local and external evidence.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Hiro architecture inspection completed Read-only inspection confirmed existing reactive opportunity detection, external-source lineage, falsifiable experiment planning, isolated variant records, and governed capability lanes that can serve as downstream components of the proposed registry.
Hiro code or runtime tests not run No Hiro implementation or runtime state was changed in this strategy session, and Hiro was not restarted.
Journal test and production build passed npm run test:hiro passed. After installing the clean checkout's locked dependencies, npm run build passed; generation validated 132 journal entries and Vite completed the production build. The initial pre-install build attempt had stopped because tsc was not yet installed in the clean checkout.

Current state

Next steps