Hiro development journal

Hiro now has a durable Open Problems Registry and dashboard

Implemented and fully tested; Hiro remains offline Machine-readable JSON

Executive summary

Hiro now has a durable strategic research layer above its reactive opportunity and experiment queues. Open harness problems retain stable identities across observations, competing hypotheses, negative results, agenda reviews, and multiple experiments.

A new Open problems tab was added to the Evaluation Observatory. It shows the ten initial foundational harness problems, research-portfolio targets, priority dimensions, measurement readiness, dependencies, hypotheses, evidence, experiment lineage, reopen conditions, and agenda-governance boundaries.

The registry is deliberately unable to construct or promote candidates. External material has no instruction authority, a passed experiment cannot automatically close its parent problem, and Hiro's existing isolation, evaluation, invariant, latency, stable-governor, and rollback controls remain unchanged.

All 606 Hiro tests passed after implementation. The dashboard JavaScript passed an independent syntax check, and 19 focused open-problem, dashboard, and experiment-planner tests passed. Hiro was not restarted during this work.

Work completed

Durable registry and audit model

Completed
  • Added a dedicated SQLite registry under Hiro's runtime directory with separate tables for open problems, evidence, hypotheses, linked experiments, and append-only problem events.
  • Problem records include a stable identifier, capability gap, strategic importance, lifecycle status, research portfolio, multidimensional priority, measurement readiness, review timing, dependencies, and explicit reopen conditions.
  • Evidence is append-only and typed as human strategy, local observation, evaluation, external research, or experiment result. Each record retains a bounded summary, source reference, observation time, and metadata.
  • Agenda reviews can reframe a problem, update measurement readiness or priority, change state, revise dependencies, and set review information only with a recorded reason and complete audit delta.
  • Experiments and hypotheses cannot be silently reassigned across problem boundaries. Negative or inconclusive experiments remain linked evidence and do not resolve their parent problem.

Initial governed research agenda

Completed
  • Added a versioned policy containing ten manually reviewed foundational harness problems: agenda governance, evaluator validity, real-world transfer, research continuity, discriminating experiment design, causal attribution, safe self-modification, capability discovery, external research synthesis, and model-versus-harness attribution.
  • The policy defines a portfolio across capability bottlenecks, reliability, measurement infrastructure, bounded exploration, and rapid frontier response. Target shares are validated to total exactly one.
  • Priority is computed from strategic leverage, evidence strength, learning value, transfer potential, urgency, and measurement readiness rather than a single short-term failure score.
  • Policy loading validates unique problem and hypothesis identities, known portfolios, dependency integrity, bounded scores, and the absence of candidate-construction or promotion authority.
  • Seeding is idempotent: the policy creates missing durable identities without overwriting later runtime evidence or agenda-review decisions.

Experiment lineage integration

Completed
  • Improvement opportunities and planned experiments now carry optional parent-problem and hypothesis identifiers.
  • When a planned experiment declares a parent problem, the planner records that experiment in the registry with its hypothesis linkage and an initial planned outcome.
  • Existing opportunities without a strategic parent remain backward compatible, allowing deliberate triage rather than automatically forcing every local incident into a foundational problem.
  • The experiment planner now ensures its persistence directory exists at save time, preserving safe operation in isolated test and candidate environments.

Evaluation Observatory Open problems tab

Completed
  • Added a read-only API endpoint that initializes missing seed identities and returns safe aggregate and detailed registry state.
  • The new dashboard tab displays problem counts, active or experiment-ready work, measurement blockers, evidence and experiment totals, portfolio targets, review cadence, and explicit authority boundaries.
  • Each expandable problem card displays its capability gap, strategic importance, measurement path, priority dimensions, dependencies, competing hypotheses, recent evidence, and reopen conditions.
  • The navigation now scrolls horizontally on constrained screens, and the problem and portfolio layouts adapt to narrower displays.
  • The README now documents the registry's purpose and its relationship to the existing continuous-improvement and promotion controls.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused open-problem, dashboard, and planner suite passed All 19 focused tests passed in 2.07 seconds. Coverage included idempotent seeding, evidence and event persistence, negative-result behavior, agenda reviews, boundary enforcement, portfolio summaries, API output, dashboard presence, and parent-problem experiment linkage.
Dashboard JavaScript syntax passed Node successfully parsed the complete inline Evaluation Observatory script after the new tab and rendering functions were added.
Full Hiro repository suite passed All 606 tests passed in 148.93 seconds.
Policy structure passed The policy parses as JSON, contains ten unique seed problems, and its five portfolio target shares total 1.00. Runtime loading also validates identifiers, dependencies, priorities, and authority boundaries.
Live Hiro and browser verification not run Hiro remained intentionally offline. The API and HTML were tested through isolated application clients and syntax checks; live dashboard interaction will be verified after the next explicitly authorized restart.

Current state

Next steps