{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T10:39:54-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Strategy: make open harness problems a durable research program",
  "publicationStatus": "Strategy designed; implementation intentionally deferred",
  "executiveSummary": [
    "Before restarting Hiro or its upgrade process, the research strategy was broadened from repeatedly improving recent behavior to maintaining a durable, evidence-backed program of open problems in AI harness construction.",
    "Hiro already contains important execution machinery: reactive opportunity detection from interaction records, prompt-safe external idea lineage including arXiv and other sources, bounded falsifiable experiments, isolated candidate worktrees, evaluation gates, and governed capability lanes.",
    "The missing strategic layer is durable problem identity and agenda governance across many observations and experiments. Without it, individual fixes and external ideas can be processed competently while the system still loses sight of the largest unresolved capability bottlenecks.",
    "The proposed integration is an Open Problems Registry connected to, but distinct from, the experiment queue. Problems persist across failed hypotheses, accumulate evidence and contradictions, and are periodically reprioritized or reframed as Hiro and the external research frontier change."
  ],
  "workstreams": [
    {
      "title": "Current-system fit assessment",
      "status": "Completed as a read-only design review",
      "details": [
        "The current opportunity detector prioritizes recent high-severity failures, surprising successes, related weak signals, novel observations, and repeated low confidence. This is useful operational feedback but does not by itself maintain long-lived research questions.",
        "The external-idea system accepts prompt-safe reductions from supported sources, preserves lineage, requires local corroboration, and prevents external content from authorizing candidate construction or promotion.",
        "The experiment planner and variant laboratory already provide much of the downstream machinery needed once a problem has yielded a falsifiable hypothesis.",
        "The current capability-lane policy supplies strong authority and evaluation gates, but its lanes are narrow implementation routes rather than a complete, revisable map of open harness problems."
      ]
    },
    {
      "title": "Open Problems Registry",
      "status": "Architecture proposed",
      "details": [
        "Create a durable registry in which each problem has a stable identity, precise capability gap, operational evidence, strategic significance, current hypotheses, unresolved contradictions, relevant external findings, evaluation gaps, and a history of attempted experiments.",
        "Treat problems as longer-lived than experiments. A failed experiment should update the parent problem's evidence and hypothesis space rather than close or duplicate the problem.",
        "Keep observed symptoms separate from proposed causes and proposed interventions. This prevents an early guess about implementation from becoming the permanent definition of the problem.",
        "Allow explicit states such as observing, characterized, experiment-ready, active, blocked-on-measurement, monitoring, provisionally resolved, superseded, and retired, with recorded reasons for every transition."
      ]
    },
    {
      "title": "Research-agenda loop",
      "status": "Operating model proposed",
      "details": [
        "Ingest signals from real interactions, evaluation failures, surprising successes, instrumentation, external research, repositories, and human strategic observations into a candidate-problem inbox.",
        "Merge evidence into existing problems when it concerns the same underlying capability gap; create a new problem only when the distinction is meaningful and testable.",
        "Select a portfolio rather than a single highest score: urgent reliability work, high-leverage capability bottlenecks, measurement and infrastructure work, and bounded exploratory research should each receive deliberate capacity.",
        "Translate selected problems into competing hypotheses and discriminating experiments, then return all outcomes, including negative and inconclusive results, to the problem record.",
        "Run periodic agenda reviews that can reprioritize, split, merge, reframe, provisionally resolve, or reopen problems based on accumulated local and external evidence."
      ]
    }
  ],
  "decisions": [
    "Make the open-problem registry a strategic control layer above the existing opportunity, experiment, and promotion pipelines rather than replacing those governed mechanisms.",
    "Use a problem portfolio so immediately measurable bug fixes do not consume all research capacity at the expense of foundational harness questions and exploratory capability discovery.",
    "Require evidence of external or real-world relevance, a clear statement of what remains unknown, and a path toward measurement before a problem receives substantial experiment resources.",
    "Preserve negative results and contradictions as first-class research evidence; do not treat failed variants as disposable queue events.",
    "Do not let external research directly authorize code changes. External discoveries can update evidence, hypotheses, and priority, while existing local corroboration, isolation, evaluation, and promotion gates remain authoritative.",
    "Separate agenda review from experiment execution so the component generating or building a candidate cannot silently redefine the research objective or declare the parent problem solved."
  ],
  "validation": [
    {
      "check": "Hiro architecture inspection",
      "status": "completed",
      "result": "Read-only inspection confirmed existing reactive opportunity detection, external-source lineage, falsifiable experiment planning, isolated variant records, and governed capability lanes that can serve as downstream components of the proposed registry."
    },
    {
      "check": "Hiro code or runtime tests",
      "status": "not run",
      "result": "No Hiro implementation or runtime state was changed in this strategy session, and Hiro was not restarted."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. After installing the clean checkout's locked dependencies, npm run build passed; generation validated 132 journal entries and Vite completed the production build. The initial pre-install build attempt had stopped because tsc was not yet installed in the clean checkout."
    }
  ],
  "currentState": [
    "The discussion has produced a concrete integration direction but has not changed Hiro's implementation, configuration, queue, model, or runtime state.",
    "The existing improvement loop remains suited to executing bounded experiments once a strategic problem has been characterized.",
    "The proposed registry, agenda-review process, portfolio allocation policy, and problem-to-experiment linkage still need an implementation design and explicit schemas."
  ],
  "limitations": [
    "No initial canonical list of open harness problems has yet been agreed, so the proposed categories and allocation percentages should not be frozen prematurely.",
    "Automated priority scoring can create false precision and can be gamed by easily measured work; human strategic review should remain part of early agenda governance.",
    "Research literature can contain weak, non-reproducible, or context-specific claims. Source recency or popularity must not substitute for local testing and transfer evidence.",
    "A provisionally resolved problem can recur under new tasks, models, tools, or environments, so resolution requires monitoring and explicit reopen conditions.",
    "The design has not yet established storage format, API surfaces, user interface, review cadence, resource budgets, or migration behavior for existing opportunities and experiments."
  ],
  "nextSteps": [
    "Before restart, write the minimum schema for problem identity, evidence, hypotheses, experiments, status transitions, priority history, and reopen conditions.",
    "Seed the registry with a small manually reviewed set of foundational harness problems derived from Hiro's observed limitations and current external research, rather than bulk-importing a large undifferentiated backlog.",
    "Define the agenda-review cadence and a portfolio policy that balances reliability, capability expansion, measurement infrastructure, and exploration.",
    "Link every new experiment to a parent problem and require its result packet to state what evidence changed, which hypothesis was weakened or strengthened, and what the next discriminating action should be.",
    "Add views that distinguish problem progress from experiment throughput, because a high experiment count is not evidence that a hard capability bottleneck is being resolved.",
    "Only after the strategic layer and restart acceptance criteria are agreed should Hiro's upgrade process resume."
  ]
}
