{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T10:55:20-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Research: external systems for discovering Hiro's real open problems",
  "publicationStatus": "Research and sourcing strategy completed; implementation deferred",
  "executiveSummary": [
    "The new Open Problems Registry supplies durable structure, but it still needs a disciplined external discovery system that finds consequential problems rather than merely accumulating papers, issues, or benchmark scores.",
    "No single public platform currently serves as a complete registry of open problems in self-improving AI harness construction. The strongest design is a composite: METR's portable task standard, ARC Prize's benchmark-and-open-challenge model, PaperBench's hierarchical rubrics, real task benchmarks such as SWE-bench, OSWorld, and MLE-bench, and issue ecosystems from active agent frameworks.",
    "Hugging Face is valuable as a discovery, dataset, leaderboard, collection, and competition substrate. It should feed Hiro's candidate-problem inbox and later host bounded challenges when useful, but Hugging Face should not define the research agenda by popularity or leaderboard rank alone.",
    "The recommended next layer is a Problem Discovery Observatory that separates raw leads from corroborated problems, clusters signals by underlying mechanism, requires reproducible evidence, and converts validated gaps into versioned challenge packets with public and held-out evaluation."
  ],
  "workstreams": [
    {
      "title": "External system survey",
      "status": "Completed",
      "details": [
        "METR's Task Standard is the closest reusable technical foundation for packaging real agent problems. It defines an environment, instructions, optional scoring, and task families, and provides adapters for existing suites including SWE-bench, GAIA, and AgentBench.",
        "ARC Prize provides the strongest open-research campaign pattern: a difficult benchmark, public and private evaluation, milestone structure, reproducible open-source submissions, research papers, and incentives for genuinely different approaches. Its current agentic work explicitly distinguishes harness innovation from model-only results.",
        "Hugging Face Competitions supports public or private challenges, code submissions, hidden test data, custom metrics, resource limits, and leaderboards. Daily Papers, Collections, community leaderboards, datasets, and agent-framework discussions provide useful discovery feeds.",
        "PaperBench provides a reusable rubric pattern: decompose a complex research outcome into a weighted tree of individually gradable requirements and separately evaluate judge quality.",
        "SWE-bench, OSWorld, MLE-bench, and METR public tasks provide real or semi-real task distributions. Their failure cases, benchmark revisions, infrastructure issues, and unsaturated categories are stronger problem signals than paper titles alone.",
        "OpenHands and Hugging Face smolagents issue and discussion systems provide bottom-up production evidence about planning, tool use, recovery, timeouts, memory, interface contracts, and stability-versus-flexibility tradeoffs."
      ]
    },
    {
      "title": "Composite sourcing model",
      "status": "Designed",
      "details": [
        "Use research feeds for proposed mechanisms, benchmark results for repeatable capability gaps, issue trackers for real operational failures, and Hiro's own interactions for local relevance.",
        "Ingest all incoming material as untrusted problem leads rather than open problems. Preserve source lineage and safe summaries while withholding candidate-construction and promotion authority.",
        "Cluster leads by underlying harness mechanism instead of keywords or product names. A timeout report, failed long-horizon task, and stalled tool workflow may all support one recovery-and-state-management problem.",
        "Promote a lead into the durable registry only after independent corroboration or one high-quality reproducible benchmark gap plus local relevance to Hiro.",
        "Convert validated problems into challenge packets containing an operational capability gap, current baselines, human calibration where possible, environment identity, public development cases, held-out evaluation, resource limits, failure taxonomy, rubric, and reopen conditions."
      ]
    },
    {
      "title": "Challenge and hackathon role",
      "status": "Clarified",
      "details": [
        "General hackathons are weak sources of durable research problems because prompts are often broad, novelty-oriented, and poorly evaluated.",
        "Challenge sprints become valuable after a problem is validated. Hiro can freeze one problem packet, give several competing strategies identical budgets and environments, score them against blind held-out cases, and require a postmortem that updates the parent problem.",
        "Hugging Face script competitions or a similar private challenge runner could later supply execution infrastructure, but Hiro should first run the format locally and prove that the metric cannot be trivially gamed.",
        "The purpose of a challenge is strategy diversity and discriminating evidence, not producing a winner that bypasses normal integration or promotion gates."
      ]
    }
  ],
  "decisions": [
    "Borrow the METR task-family abstraction for portable problem environments and scoring contracts.",
    "Borrow the ARC Prize campaign model for milestone-driven, reproducible, open strategy exploration with private evaluation.",
    "Borrow PaperBench's hierarchical rubric method for complex research outcomes and separately test evaluator reliability.",
    "Use Hugging Face as infrastructure and a source ecosystem, not as the authority that determines which problems matter.",
    "Treat active benchmark repositories and agent-framework issue trackers as evidence streams whose recurring failure clusters can reveal open harness problems.",
    "Keep raw leads, validated problems, hypotheses, challenge packets, experiments, and promoted changes as separate governed objects.",
    "Require cross-source corroboration, reproducibility, local relevance, and measurement readiness before a discovery consumes substantial experiment capacity."
  ],
  "validation": [
    {
      "check": "External source verification",
      "status": "passed",
      "result": "Current primary sources were reviewed for Hugging Face Competitions and evaluation facilities, METR task standards and public tasks, ARC Prize competitions and research, PaperBench, MLE-bench, OSWorld, SWE-bench, OpenHands, and smolagents."
    },
    {
      "check": "Hiro implementation or runtime tests",
      "status": "not run",
      "result": "This was a research and system-design session. No Hiro code, configuration, queue state, model state, or runtime process was changed."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 134 journal entries, then TypeScript and Vite completed the production build successfully."
    }
  ],
  "currentState": [
    "Hiro's registry can represent the output of this discovery process, but no automated external source adapters or candidate-problem inbox have been added in this session.",
    "The initial ten problems remain a manually reviewed seed agenda rather than a complete externally derived problem map.",
    "Hiro remains offline and no existing improvement candidate or queue item was advanced."
  ],
  "limitations": [
    "Benchmark gaps can reflect evaluation defects, environment brittleness, leakage, or infrastructure problems rather than true harness limitations; each signal requires validation.",
    "GitHub issues and discussions contain duplicates, user configuration mistakes, speculative proposals, and popularity bias, so issue counts are not evidence strength by themselves.",
    "Paper popularity and trending feeds favor recent and attention-grabbing work and can miss negative results or foundational engineering constraints.",
    "Public challenge leaderboards invite hill-climbing and contamination; hidden tests, versioned environments, submission budgets, replication, and evaluator audits remain necessary.",
    "Human baselines and realistic long-horizon environments can be expensive to create, which means some important problems will initially remain blocked on measurement.",
    "The proposed source list is a starting portfolio and must be revised as benchmarks saturate, projects change, and new research venues emerge."
  ],
  "nextSteps": [
    "Implement a read-only Problem Discovery Inbox with adapters for selected primary research feeds, benchmark releases and failures, and agent-framework issue trackers.",
    "Define a safe lead schema containing source identity, claim type, mechanism, reproducibility, benchmark evidence, local relevance, and proposed parent-problem matches.",
    "Create corroboration rules that require multiple independent signals or one reproducible high-quality benchmark gap before registry admission.",
    "Adopt a portable task-family and challenge-packet format compatible in spirit with METR's standard while retaining Hiro's stricter authority and privacy boundaries.",
    "Run a first manual discovery cycle over METR public tasks, ARC-AGI-3 harness results, SWE-bench, OSWorld, MLE-bench, PaperBench, OpenHands, and smolagents, then compare the resulting problem clusters with the ten seeded problems.",
    "Choose one validated, measurable problem for an internal multi-strategy challenge sprint before considering external competition hosting."
  ]
}
