Hiro development journal

Research: external systems for discovering Hiro's real open problems

Research and sourcing strategy completed; implementation deferred Machine-readable JSON

Executive summary

The new Open Problems Registry supplies durable structure, but it still needs a disciplined external discovery system that finds consequential problems rather than merely accumulating papers, issues, or benchmark scores.

No single public platform currently serves as a complete registry of open problems in self-improving AI harness construction. The strongest design is a composite: METR's portable task standard, ARC Prize's benchmark-and-open-challenge model, PaperBench's hierarchical rubrics, real task benchmarks such as SWE-bench, OSWorld, and MLE-bench, and issue ecosystems from active agent frameworks.

Hugging Face is valuable as a discovery, dataset, leaderboard, collection, and competition substrate. It should feed Hiro's candidate-problem inbox and later host bounded challenges when useful, but Hugging Face should not define the research agenda by popularity or leaderboard rank alone.

The recommended next layer is a Problem Discovery Observatory that separates raw leads from corroborated problems, clusters signals by underlying mechanism, requires reproducible evidence, and converts validated gaps into versioned challenge packets with public and held-out evaluation.

Work completed

External system survey

Completed
  • METR's Task Standard is the closest reusable technical foundation for packaging real agent problems. It defines an environment, instructions, optional scoring, and task families, and provides adapters for existing suites including SWE-bench, GAIA, and AgentBench.
  • ARC Prize provides the strongest open-research campaign pattern: a difficult benchmark, public and private evaluation, milestone structure, reproducible open-source submissions, research papers, and incentives for genuinely different approaches. Its current agentic work explicitly distinguishes harness innovation from model-only results.
  • Hugging Face Competitions supports public or private challenges, code submissions, hidden test data, custom metrics, resource limits, and leaderboards. Daily Papers, Collections, community leaderboards, datasets, and agent-framework discussions provide useful discovery feeds.
  • PaperBench provides a reusable rubric pattern: decompose a complex research outcome into a weighted tree of individually gradable requirements and separately evaluate judge quality.
  • SWE-bench, OSWorld, MLE-bench, and METR public tasks provide real or semi-real task distributions. Their failure cases, benchmark revisions, infrastructure issues, and unsaturated categories are stronger problem signals than paper titles alone.
  • OpenHands and Hugging Face smolagents issue and discussion systems provide bottom-up production evidence about planning, tool use, recovery, timeouts, memory, interface contracts, and stability-versus-flexibility tradeoffs.

Composite sourcing model

Designed
  • Use research feeds for proposed mechanisms, benchmark results for repeatable capability gaps, issue trackers for real operational failures, and Hiro's own interactions for local relevance.
  • Ingest all incoming material as untrusted problem leads rather than open problems. Preserve source lineage and safe summaries while withholding candidate-construction and promotion authority.
  • Cluster leads by underlying harness mechanism instead of keywords or product names. A timeout report, failed long-horizon task, and stalled tool workflow may all support one recovery-and-state-management problem.
  • Promote a lead into the durable registry only after independent corroboration or one high-quality reproducible benchmark gap plus local relevance to Hiro.
  • Convert validated problems into challenge packets containing an operational capability gap, current baselines, human calibration where possible, environment identity, public development cases, held-out evaluation, resource limits, failure taxonomy, rubric, and reopen conditions.

Challenge and hackathon role

Clarified
  • General hackathons are weak sources of durable research problems because prompts are often broad, novelty-oriented, and poorly evaluated.
  • Challenge sprints become valuable after a problem is validated. Hiro can freeze one problem packet, give several competing strategies identical budgets and environments, score them against blind held-out cases, and require a postmortem that updates the parent problem.
  • Hugging Face script competitions or a similar private challenge runner could later supply execution infrastructure, but Hiro should first run the format locally and prove that the metric cannot be trivially gamed.
  • The purpose of a challenge is strategy diversity and discriminating evidence, not producing a winner that bypasses normal integration or promotion gates.

Decisions and reasoning

Validation and evidence

CheckStatusResult
External source verification passed Current primary sources were reviewed for Hugging Face Competitions and evaluation facilities, METR task standards and public tasks, ARC Prize competitions and research, PaperBench, MLE-bench, OSWorld, SWE-bench, OpenHands, and smolagents.
Hiro implementation or runtime tests not run This was a research and system-design session. No Hiro code, configuration, queue state, model state, or runtime process was changed.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 134 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps