Hiro development journal

Implementation: a source-grounded open-problem agenda and task protocol

Implementation and local validation completed Machine-readable JSON

Executive summary

Hiro now has a versioned, ranked research agenda above its durable Open Problems Registry. The agenda freezes the primary-source observations behind each priority and assigns every seeded problem a bounded first challenge rather than treating a broad research label as an executable task.

Evaluator validity is ranked first because a self-improvement loop cannot trust apparent gains until it can detect evaluator defects, instability, and reward-shaped strategies. Causal attribution and experiment design follow because Hiro must distinguish a harness improvement from model, prompt, environment, or measurement variance.

A six-tier interaction protocol now moves work from deterministic diagnostics through isolated public tasks, external benchmark adapters, held-out discrimination, cross-family transfer, and production shadow evaluation. Advancement requires explicit gates; passing a task updates evidence but cannot by itself close a problem, construct an unrestricted candidate, or authorize promotion.

The Evaluation Observatory's Open problems tab now presents the grounded agenda, source count, admission rule, escalation ladder, rank rationale, and first challenge packet for every problem. External research is attached idempotently as evidence with preserved source lineage.

Work completed

Versioned research agenda

Completed
  • Added a schema-validated agenda containing twelve frozen primary sources and ten ranked problem entries. Each source records its role, a conservative observed claim, affected problem identities, and the observation timestamp.
  • Ranked evaluator validity first, causal attribution second, experiment design third, and real-world transfer fourth. Research continuity, safe self-modification, model-harness interaction, external synthesis, capability discovery, and agenda governance complete the initial ordering.
  • Assigned each problem one first challenge packet with an objective, environment, baseline, permitted candidate scope, public and held-out case counts, repetition count, wall-clock and model-call budgets, primary metric, minimum effect, required gates, transfer gate, and halt conditions.
  • Separated the stable problem registry from the changing agenda so problem identities and longitudinal evidence survive reprioritization.

Agenda validation and evidence integration

Completed
  • Added a strict agenda loader that rejects unknown problem references, duplicate ranks or challenge identities, non-HTTPS or credential-bearing source URLs, invalid decisions and task kinds, missing held-out evaluation, malformed budgets, empty gates, authority expansion, and an incomplete tier sequence.
  • Added idempotent evidence attachment to the SQLite-backed Open Problems Registry. Repeated dashboard reads do not duplicate the same source observation.
  • Preserved a hard authority boundary: the agenda explicitly denies candidate-construction and promotion authority, and all external text is classified as untrusted evidence.
  • The read-only evaluation API merges current agenda metadata into durable problem records while allowing deployments without an agenda file to remain operational.

Task-interaction protocol

Completed
  • Defined a portable task-family contract inspired by current agent evaluation practice: environment identity, instructions, scoring contract, case split, budgets, baselines, repetitions, failure taxonomy, artifacts, and provenance.
  • Defined paired baseline and candidate run packets that hold model, prompt, tools, environment, budget, and evaluator constant except for the harness change under test.
  • Required public development cases and blind held-out cases before a change can advance, followed by cross-family transfer and production-shadow evidence for higher-risk claims.
  • Defined explicit outcomes: supported, contradicted, inconclusive, evaluator failure, environment failure, and safety halt. Every outcome appends evidence to the parent problem; none automatically resolves it.
  • Specified initial execution order: adversarially audit the evaluator, run replicated causal variants, test competing hypotheses, add external task adapters, then evaluate continuity and bounded self-modification. Novel strategy exploration is deliberately later, after measurement is trustworthy.

Evaluation Observatory presentation

Completed
  • Extended the Open problems tab with agenda source count, admission criteria, the six-tier escalation ladder, agenda ranks, ranking rationales, first-challenge summaries, evaluation sizes, budgets, metrics, minimum effects, transfer requirements, and mandatory gates.
  • Kept the interface read-only and made its lack of promotion authority visible so strategic planning cannot be mistaken for an execution bypass.
  • Updated repository documentation to link the registry, ranked agenda, and task protocol as separate governed artifacts.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused agenda, registry, API, dashboard, and patch-generation tests passed 22 tests passed. Coverage included repository agenda validation, rejection of unknown problems and authority expansion, API exposure, idempotent source-evidence attachment, dashboard delivery, and existing candidate-generation boundaries.
Evaluation Observatory JavaScript syntax passed The complete inline dashboard script was compiled with Node.js using the Function constructor without syntax errors.
Full Hiro regression suite passed 609 tests passed in 153.38 seconds.
Live Hiro runtime and browser interaction not run Hiro remains offline by design. No restart, live model call, queued experiment, candidate construction, or promotion was performed.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 135 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps