{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T11:11:07-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Implementation: a source-grounded open-problem agenda and task protocol",
  "publicationStatus": "Implementation and local validation completed",
  "executiveSummary": [
    "Hiro now has a versioned, ranked research agenda above its durable Open Problems Registry. The agenda freezes the primary-source observations behind each priority and assigns every seeded problem a bounded first challenge rather than treating a broad research label as an executable task.",
    "Evaluator validity is ranked first because a self-improvement loop cannot trust apparent gains until it can detect evaluator defects, instability, and reward-shaped strategies. Causal attribution and experiment design follow because Hiro must distinguish a harness improvement from model, prompt, environment, or measurement variance.",
    "A six-tier interaction protocol now moves work from deterministic diagnostics through isolated public tasks, external benchmark adapters, held-out discrimination, cross-family transfer, and production shadow evaluation. Advancement requires explicit gates; passing a task updates evidence but cannot by itself close a problem, construct an unrestricted candidate, or authorize promotion.",
    "The Evaluation Observatory's Open problems tab now presents the grounded agenda, source count, admission rule, escalation ladder, rank rationale, and first challenge packet for every problem. External research is attached idempotently as evidence with preserved source lineage."
  ],
  "workstreams": [
    {
      "title": "Versioned research agenda",
      "status": "Completed",
      "details": [
        "Added a schema-validated agenda containing twelve frozen primary sources and ten ranked problem entries. Each source records its role, a conservative observed claim, affected problem identities, and the observation timestamp.",
        "Ranked evaluator validity first, causal attribution second, experiment design third, and real-world transfer fourth. Research continuity, safe self-modification, model-harness interaction, external synthesis, capability discovery, and agenda governance complete the initial ordering.",
        "Assigned each problem one first challenge packet with an objective, environment, baseline, permitted candidate scope, public and held-out case counts, repetition count, wall-clock and model-call budgets, primary metric, minimum effect, required gates, transfer gate, and halt conditions.",
        "Separated the stable problem registry from the changing agenda so problem identities and longitudinal evidence survive reprioritization."
      ]
    },
    {
      "title": "Agenda validation and evidence integration",
      "status": "Completed",
      "details": [
        "Added a strict agenda loader that rejects unknown problem references, duplicate ranks or challenge identities, non-HTTPS or credential-bearing source URLs, invalid decisions and task kinds, missing held-out evaluation, malformed budgets, empty gates, authority expansion, and an incomplete tier sequence.",
        "Added idempotent evidence attachment to the SQLite-backed Open Problems Registry. Repeated dashboard reads do not duplicate the same source observation.",
        "Preserved a hard authority boundary: the agenda explicitly denies candidate-construction and promotion authority, and all external text is classified as untrusted evidence.",
        "The read-only evaluation API merges current agenda metadata into durable problem records while allowing deployments without an agenda file to remain operational."
      ]
    },
    {
      "title": "Task-interaction protocol",
      "status": "Completed",
      "details": [
        "Defined a portable task-family contract inspired by current agent evaluation practice: environment identity, instructions, scoring contract, case split, budgets, baselines, repetitions, failure taxonomy, artifacts, and provenance.",
        "Defined paired baseline and candidate run packets that hold model, prompt, tools, environment, budget, and evaluator constant except for the harness change under test.",
        "Required public development cases and blind held-out cases before a change can advance, followed by cross-family transfer and production-shadow evidence for higher-risk claims.",
        "Defined explicit outcomes: supported, contradicted, inconclusive, evaluator failure, environment failure, and safety halt. Every outcome appends evidence to the parent problem; none automatically resolves it.",
        "Specified initial execution order: adversarially audit the evaluator, run replicated causal variants, test competing hypotheses, add external task adapters, then evaluate continuity and bounded self-modification. Novel strategy exploration is deliberately later, after measurement is trustworthy."
      ]
    },
    {
      "title": "Evaluation Observatory presentation",
      "status": "Completed",
      "details": [
        "Extended the Open problems tab with agenda source count, admission criteria, the six-tier escalation ladder, agenda ranks, ranking rationales, first-challenge summaries, evaluation sizes, budgets, metrics, minimum effects, transfer requirements, and mandatory gates.",
        "Kept the interface read-only and made its lack of promotion authority visible so strategic planning cannot be mistaken for an execution bypass.",
        "Updated repository documentation to link the registry, ranked agenda, and task protocol as separate governed artifacts."
      ]
    }
  ],
  "decisions": [
    "Treat open problems as durable research objects and agendas as replaceable, versioned views over those objects.",
    "Rank evaluator validity ahead of capability expansion because unreliable measurement can manufacture false progress and select reward-shaped behavior.",
    "Require tasks to produce discriminating evidence about a named hypothesis, not merely a leaderboard score.",
    "Escalate task realism only after cheaper tiers pass, while retaining a direct safety halt at every tier.",
    "Use external benchmarks through pinned adapters and provenance records rather than granting external task text operational authority.",
    "Require held-out discrimination, repeated trials, paired baselines, negative-result retention, and cross-family transfer before making broad improvement claims.",
    "Do not allow a passed challenge to resolve its parent problem or promote a change; those remain separate review and governor decisions.",
    "Keep Hiro offline during this session and defer runtime interaction until the planned restart and upgrade process."
  ],
  "validation": [
    {
      "check": "Focused agenda, registry, API, dashboard, and patch-generation tests",
      "status": "passed",
      "result": "22 tests passed. Coverage included repository agenda validation, rejection of unknown problems and authority expansion, API exposure, idempotent source-evidence attachment, dashboard delivery, and existing candidate-generation boundaries."
    },
    {
      "check": "Evaluation Observatory JavaScript syntax",
      "status": "passed",
      "result": "The complete inline dashboard script was compiled with Node.js using the Function constructor without syntax errors."
    },
    {
      "check": "Full Hiro regression suite",
      "status": "passed",
      "result": "609 tests passed in 153.38 seconds."
    },
    {
      "check": "Live Hiro runtime and browser interaction",
      "status": "not run",
      "result": "Hiro remains offline by design. No restart, live model call, queued experiment, candidate construction, or promotion was performed."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 135 journal entries, then TypeScript and Vite completed the production build successfully."
    }
  ],
  "currentState": [
    "The working tree contains the Open Problems Registry, ranked source-grounded agenda, agenda validator, idempotent evidence integration, read-only evaluation API and dashboard presentation, and task-interaction protocol.",
    "The agenda contains a concrete first challenge for all ten seeded problems. Evaluator adversarial audit is the only run-next item; several later items remain prepare or measure-first decisions.",
    "The implementation is validated locally but is not exercising a live Hiro process because the user intentionally deferred restart and upgrade.",
    "External source observations are static and reviewable in this version; automated discovery adapters and recurring source refresh are not yet implemented."
  ],
  "limitations": [
    "The twelve-source catalog is a curated initial snapshot rather than an exhaustive or continuously refreshed research map.",
    "Challenge packets define contracts and gates, but their concrete task instances, pinned environments, scoring implementations, and human calibration data still need to be built before execution.",
    "Several minimum-effect statements are semantic acceptance criteria rather than fully calibrated numeric thresholds. Initial baseline runs must establish variance before final thresholds are frozen.",
    "External benchmark adapters can inherit evaluator bugs, infrastructure brittleness, contamination, and licensing constraints; adapter validation remains part of the required work.",
    "A six-tier process reduces risk but cannot prove that a change will generalize beyond the tested distributions or remain safe under all future model and environment changes.",
    "The dashboard is a planning and evidence surface, not a workflow control plane."
  ],
  "nextSteps": [
    "Implement the rank-one evaluator adversarial-audit task family with deterministic fixtures, known-good and known-bad trajectories, perturbation tests, and an explicit defect taxonomy.",
    "Run an unchanged baseline repeatedly to estimate evaluator and environment variance before allowing any candidate comparison.",
    "Build paired baseline and candidate packet serialization with immutable model, prompt, tool, environment, budget, evaluator, and seed identities.",
    "Add a held-out case vault whose contents are unavailable to candidate construction and whose access is auditable.",
    "Create the replicated causal-variant challenge and hypothesis-discrimination challenge after the evaluator audit passes.",
    "Implement one pinned METR-style external adapter as the first tier-two integration and prove that local scoring agrees with the upstream task contract.",
    "Add recurring agenda review that records source refreshes, rank changes, blocked measurement, negative results, splits, merges, resolution decisions, and reopen triggers.",
    "After restart and upgrade, execute tasks only through the existing isolation and stable-governor path; do not let the agenda or external sources bypass promotion controls."
  ]
}
