{
  "schemaVersion": 2,
  "date": "2026.09.19",
  "publishedAt": "2026-09-19T21:27:13-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Research Map and HRO design reconciled with the existing registry",
  "publicationStatus": "Design and critical review complete; structural integration awaits approval",
  "executiveSummary": [
    "Inspected the repository, existing architecture and research policies, experiment and promotion code, read-only registry and queue state, historical qualifications, relevant project discussion and primary literature before designing a research-management extension.",
    "The central finding is that Hiro already has a durable Open Problems Registry, ranked agenda, experiment linkage, evidence return and append-only review events. The proposal enriches that system rather than creating another research queue.",
    "Prepared a populated machine-readable Research Map, an exact JSON Schema and a realistic unexecuted HRO example about discovering and independently validating a new evaluative dimension. The review packet remains a proposal, not a deployed subsystem.",
    "Recovered the established five-level RSI labels: AI-R&D autonomy, autonomous experiment loops, persistent self-modification, meta-improvement and compounding R&D velocity. These remain distinct from experiment tiers and Phase 2/3 qualification names.",
    "Preserved negative evidence, including null evaluator calibration, frozen-corpus mismatch, unjustified production transfer, viability false positives and the latest capability candidate failure after an earlier narrow holdout pass.",
    "No Hiro source, live registry data, queue, model configuration, process or promotion authority was changed. Historical tests were not rerun or relabeled as new validation. The originating design brief calls for review before structural implementation."
  ],
  "workstreams": [
    {
      "title": "Existing-state reconciliation",
      "status": "Completed",
      "details": [
        "Keep the system knowledge-base JSON as the architecture and terminology index; extend it with established objective, sourced capability claims and a small reviewed research frontier.",
        "Extend OpenProblemRegistry rather than adding an HRO database. Reuse problem_id as HRO identity, existing hypotheses and experiment links, and the existing event ledger.",
        "Keep the ranked agenda and task-interaction protocol as the existing priority and experimental rigor mechanisms. The map must not own a competing priority score.",
        "Keep discovery, source claims, feasibility, baseline probes, reproduction, production-surface binding, candidate construction, restricted execution and independent evaluation. Link their immutable evidence rather than copying control state.",
        "Keep governor, promotion transactions, activation and probation as the sole existing production path. A research decision cannot grant promotion authority.",
        "Retain historical Stage 6 and lab evidence with explicit historical/manual status. Import selected attributed conversation observations, never bulk conversations as executable work."
      ]
    },
    {
      "title": "Verified snapshot and historical scope",
      "status": "Completed",
      "details": [
        "Inspected Hiro revision 387a4b00c38ca5ab9358a4e69c2980c7fa719b86. Working tree remained clean after the design work.",
        "Read-only research registry snapshot: 10 problems, 33 evidence records, 2 hypotheses, 1 linked experiment and 55 events. The linked experiment is failed because null calibration found the evaluator insufficiently calibrated.",
        "Read-only continuous queue snapshot: 9 historical implemented labels; promotion transactions include 4 finalized and 3 failed records. Historical label counts do not establish current-design retained promotions.",
        "The existing September 18 frozen pipeline qualification report passed all five scenarios at the inspected revision. This session inspected the report and hashed it; it did not rerun those tests.",
        "A bounded health read did not establish current service availability. Persisted operating history is not evidence that the service is running now; no restart was attempted.",
        "Phase 2 is historically reported as demonstrated after golden/rollback reliability qualification. Phase 3C demonstrated source-to-plan feasibility; Phase 3D reached promotion eligibility; Phase 3E reported one retained non-meta production cycle. These distinct scopes remain explicit.",
        "The latest capability candidate was not retained: the governor full suite recorded 1 failure, 1102 passes and 12 skips. The transaction failed and the candidate was later superseded. Earlier validation/holdout success does not override that result.",
        "Historical report hashes verify the documents inspected, not fresh reproduction of the raw campaigns. Some old campaign artifacts were unavailable in the current runtime tree."
      ]
    },
    {
      "title": "Populated canonical map proposal",
      "status": "Completed",
      "details": [
        "The objective remains an always-on local agent progressively taking over both theorizing/research direction and implementation/experimentation under demonstrated competence and real resources.",
        "Consciousness remains a separate unproven long-range hypothesis, never an architectural assumption or success criterion.",
        "The map preserves existing architecture, terminology, all ten registry problem identities, both persisted hypotheses, sourced capability claims, limitations, negative outcomes, prior art, deprioritized ideas and three candidate frontier questions.",
        "Capability labels distinguish proposed, implemented, tested, demonstrated, qualified and production/persistent. The existing queue meaning of implemented remains unchanged; code existence does not imply runtime retention.",
        "Model, hardware/resource, architecture, evaluator, reliability and knowledge/research limitations are separate. The user-reported local envelope is 32 GiB VRAM, 32 GiB RAM and an i7-10700K-class host; usable headroom is not freshly measured.",
        "Recursive abstraction and recursive reorganization remain named historical concepts with operational definitions explicitly unrecovered, rather than newly invented replacements.",
        "The proposed small frontier is evaluator validity, research continuity through one evidence handoff, and then independent validation of dimension discovery. This is a recommendation, not a silent agenda rewrite."
      ]
    },
    {
      "title": "Negative evidence and reopen conditions",
      "status": "Completed",
      "details": [
        "The Phase 3F frozen-corpus mismatch was an infrastructure/preregistration boundary failure. It is not a scientific rejection of the source hypotheses.",
        "The resumed Phase 3F campaign had six supported reproductions: three mechanisms were already present and three lacked justified production transfer. It produced zero candidates; reopening requires a measured gap and causal production surface.",
        "Phase 3F-VC had two reproduced mechanisms with measured gaps but no valid production transfer. Phase 3F-VS subsequently strengthened production binding; all eight preserved hypotheses were nonviable under that contract.",
        "The persisted evaluator null-calibration failure should be revisited only with independent null/positive/negative controls and a reproducible frozen baseline.",
        "The latest capability candidate needs a genuinely corrected, freshly eligible candidate and independent regression evidence. Consumed holdout results and retired candidate identity must not be revived as success.",
        "Earlier discussion deprioritized large Hiro populations and a multi-domain research institution before a discriminating single-agent experiment. Neuro-fuzzy scheduling remains only a competing mechanism, not a research thesis."
      ]
    },
    {
      "title": "HRO contract and example",
      "status": "Proposed; not executed",
      "details": [
        "The exact proposed export includes identity/originator/provenance, project relationship, precise question, prior art and novelty confidence, falsifiable hypotheses and alternatives, experiment design, controls, success criteria, resource envelope, implementation/rollback plan, evidence, interpretation, decision, learning, deduplication and contribution attribution.",
        "Early observations do not require fabricated experiments or results. Missing/irrelevant sections remain absent or explicitly not applicable. Strict experimental completeness is required only before readiness under the existing admission policy.",
        "Example question: at equal local inference budget, can Hiro identify an omitted failure dimension, construct a valid evaluator and improve independently scored held-out utility beyond fixed-taxonomy and goal-generation controls?",
        "The example is a proposed child of existing capability discovery, linked to evaluator validity, experiment design, causal attribution and model-harness interaction. It remains observing, with no experiment IDs or execution authority.",
        "Controls include equal model/tool/token budgets, null and known-positive dossiers, and a relabel-only sham or equivalence audit. A curator-owned hidden scorer judges the generated evaluator; the generated evaluator cannot certify its own improvement.",
        "Draft criteria include a 10-percentage-point independent utility advantage over each main control with paired uncertainty excluding zero, valid metric calibration, transfer and no invariant regression. These thresholds are not preregistered; a separate pilot must establish variance, sample size and feasible budget.",
        "The bounded local pilot estimate uses one existing model, sequential arms, no downloads/training/external services, at most 12 calls per main arm and 48k generated tokens total, 28 GiB VRAM, 24 GiB RAM, 1 GiB scratch and a 15-minute stop limit. These are proposed caps, not measured feasibility or authorization.",
        "One successful dimension-discovery experiment would not prove compounding RSI. Meta-improvement requires a later equal-compute comparison of subsequent research yield using frozen original and modified processes."
      ]
    },
    {
      "title": "Prior-art review",
      "status": "Targeted primary-source review completed",
      "details": [
        "DGM (https://arxiv.org/abs/2505.22954), STOP (https://arxiv.org/abs/2310.02304), HyperAgents (https://arxiv.org/abs/2603.19461) and ADAS (https://arxiv.org/abs/2408.08435) establish direct conceptual neighbors for agent self-modification, improvement of an improver and automatic agent design.",
        "OMNI (https://arxiv.org/abs/2306.01711) and the autotelic-agent survey (https://arxiv.org/abs/2012.09830) connect goal selection and interestingness to established open-ended learning research.",
        "LMA3 (https://arxiv.org/abs/2305.12487) already combines goal representation, generation and reward functions. This directly challenges a novelty claim based merely on inventing a goal or metric.",
        "The AI Scientist (https://sakana.ai/ai-scientist/) provides prior art for an automated research workflow in bounded settings.",
        "Novelty of the narrower taxonomy-inadequacy question remains unknown pending literature review. Primary abstract/project-page inspection is not exhaustive prior-art clearance or a local replication."
      ]
    },
    {
      "title": "Minimal integration and critical review",
      "status": "Proposed; awaiting approval",
      "details": [
        "Add one optional research-context JSON column to the existing problem table. Reuse the latest event ID for concurrency revision; do not add another revision ledger or state machine.",
        "Complete existing append-only events with before/after hypothesis statements and experiment learning summaries. Current replacement updates can otherwise lose the earlier reasoning even when a status event remains.",
        "Use validated local import/export through Work and the existing read-only Observatory. No new network write service is required. Exports carry snapshot time, repository revision and event watermark.",
        "Deduplicate exact IDs/aliases first, then normalized token overlap over active and retired questions, hypotheses, prior art and negative outcomes. Record semantic review because lexical absence is not novelty proof; no vector infrastructure is justified yet.",
        "Bind admitted experiments to immutable question/hypothesis revision and predeclared criteria. Return conceptual flaws for revision instead of silently changing intent.",
        "Evidence receipt ingestion must be idempotent and separate from interpretation. Stale edits conflict; corrections reference prior evidence; the map is generated after commits rather than independently edited.",
        "The critical review rejected a second queue, global HRO numbering, giant mandatory templates, automatic conversation-to-experiment conversion and unsourced capability/novelty claims.",
        "The acceptance demonstration should be one ordinary question-to-experiment-to-negative-result-to-revised-map handoff across a context reset, with migration/conflict/history/privacy tests. Dimension discovery is not part of that first integration."
      ]
    }
  ],
  "decisions": [
    "Reuse problem identity and the existing registry; use HRO as its enriched research representation.",
    "Preserve the original five RSI labels and separate them from qualification phases and experiment tiers.",
    "Preserve failures, infrastructure outcomes and inconclusive results, with explicit reopen conditions.",
    "Keep exact research questions and raw evidence separate from engineering plans and interpretations.",
    "Require structural implementation approval at the recovered design-first stop point; the example is not authorized research execution.",
    "Publish this detailed session record under the existing journal authorization, retaining noindex metadata and the unlisted navigation arrangement."
  ],
  "validation": [
    {
      "check": "Draft 2020-12 schema and populated packet",
      "status": "passed",
      "result": "Schema self-validation and packet validation passed; all source references and example problem lineage checked; 17 inspected repository source hashes matched."
    },
    {
      "check": "Invalid packet controls",
      "status": "passed",
      "result": "Nine invalid variants rejected: invalid status, novelty category, confidence, execution authority, timestamp, missing provenance, resource cap, consciousness milestone and unknown field. The first run exposed missing optional date-time validation support; after adding it only to artifact-local validation dependencies, all nine checks passed."
    },
    {
      "check": "Hiro working tree / runtime scope",
      "status": "passed",
      "result": "Working tree remained clean. Database reads used SQLite read-only mode. No Hiro structural implementation or model experiment ran."
    },
    {
      "check": "Hiro full test suite",
      "status": "not_run",
      "result": "Design-only session. Historical qualification results are attributed to their existing reports and were not rerun."
    },
    {
      "check": "Journal timestamped-entry tests",
      "status": "passed",
      "result": "npm run test:hiro passed before the journal build."
    },
    {
      "check": "Journal frontend build",
      "status": "passed",
      "result": "npm run build passed: generated and validated 229 entries with timestamp, aliases and noindex checks, then TypeScript and Vite completed. Final publication regenerated from this same source entry."
    }
  ],
  "currentState": [
    "Review artifacts are complete outside the Hiro working tree; no proposed schema/migration is live.",
    "Current local service availability remains unverified. No liveness or hardware-cause claim is inferred from the failed health read.",
    "Research autonomy metrics are not yet established: origin attribution, accepted hypotheses, valid designs, reproducibility and comparable gains per compute need accountable stage records.",
    "Exact recursive abstraction/reorganization definitions and exhaustive novelty review remain incomplete; historical raw campaign reproduction was outside this design session."
  ],
  "nextSteps": [
    "Review and approve or revise the minimal registry extension and generated map/HRO view.",
    "After approval, implement migration, complete history, conflict-safe local edits, deduplication and existing executor evidence linkage with focused tests.",
    "Demonstrate one real research-memory handoff and context reset before adding autonomous research responsibility.",
    "Keep dimension discovery as a candidate pending independent calibration, literature review, resource measurement and explicit experiment admission."
  ],
  "disclosureNote": "Public design summary only. Private conversations, personal data, credentials, held-out cases and sensitive runtime details are excluded. Scientific hypotheses and historical reports are not presented as current demonstrated capability."
}
