Executive summary
Hiro now has a versioned, ranked research agenda above its durable Open Problems Registry. The agenda freezes the primary-source observations behind each priority and assigns every seeded problem a bounded first challenge rather than treating a broad research label as an executable task.
Evaluator validity is ranked first because a self-improvement loop cannot trust apparent gains until it can detect evaluator defects, instability, and reward-shaped strategies. Causal attribution and experiment design follow because Hiro must distinguish a harness improvement from model, prompt, environment, or measurement variance.
A six-tier interaction protocol now moves work from deterministic diagnostics through isolated public tasks, external benchmark adapters, held-out discrimination, cross-family transfer, and production shadow evaluation. Advancement requires explicit gates; passing a task updates evidence but cannot by itself close a problem, construct an unrestricted candidate, or authorize promotion.
The Evaluation Observatory's Open problems tab now presents the grounded agenda, source count, admission rule, escalation ladder, rank rationale, and first challenge packet for every problem. External research is attached idempotently as evidence with preserved source lineage.