Hiro development journal

Phase 3C finds fourteen actionable local experiments

Actionable discovery works; Phase 3A now has sufficient planned inputs but remains stopped Machine-readable JSON

Executive summary

Qualified experimental feasibility and actionable discovery only. Promotion, activation, rollback, corroboration, source-claim extraction, autonomous implementation, and meta-improvement were not changed.

Reclassified all twenty-five previously validated Phase 3B claims using explicit relevance, asset, compute, and feasibility fields. Twenty-four remain not Hiro-relevant and one cross-domain pre-commitment gating mechanism has a defensible light local experiment path.

Defined a conservative research-asset policy with bounded download, storage, CPU, and GPU tiers; approved public source classes; provenance and license requirements; isolation; supply-chain precautions; and explicit denial of production-write, candidate, and promotion authority.

Added an experiment-feasibility planner after claim validation. It produces exact asset inventories, local hypotheses, observables, falsification criteria, resource budgets, isolation requirements, relevance classifications, feasibility tiers, an independent semantic audit, and deterministic policy results.

Added bounded actionable discovery ranking for empirical, Hiro-relevant, and reproducible research. Every considered source receives a recorded retention or exclusion decision, and the retained corpus is frozen before claim extraction.

The new campaign considered seventy-one sources, retained twenty, extracted and validated thirty-four claims, classified sixteen as directly relevant or explicitly transferable, and produced fourteen valid LOCAL_READY experiment plans.

The first demonstrated feasibility divergence was deterministic rather than semantic: LOCAL_READY correctly prohibited downloads but also prohibited generated scratch output. The minimal correction preserved zero downloads while allowing at most one GiB of disposable scratch storage.

The final disposition is ACTIONABLE DISCOVERY WORKS. Phase 3A now has more than the requested five legitimate experiment plans, but no asset was acquired and no reproduction experiment was executed or resumed.

Work completed

Existing Phase 3B claim feasibility

Completed
  • All twenty-five frozen validated claims received exact required-asset records, available and missing inventories, resource estimates, relevance decisions, feasibility tiers, local hypotheses where defensible, observables, falsification criteria, and independent semantic audits.
  • Twenty-four claims terminate as NOT_HIRO_RELEVANT. Their missing assets remain explicit even though relevance overrides an apparent cost tier.
  • One driving-system claim transfers defensibly to Hiro: a gate that acts before commitment can be tested as an uncertainty-abstention control in routing. Its bounded plan requires only a small locally constructed adapter and no external download.
  • No other old claim could responsibly move to a light or moderate acquisition tier from the retained evidence. Common blockers were unavailable paper-specific implementations, benchmarks, models, simulators, annotations, and proprietary agent scaffolds.

Bounded research-asset policy

Implemented
  • Automatic acquisition stops at ACQUIRABLE_LIGHT: at most 250 MiB download, one GiB storage, sixty CPU minutes, and fifteen GPU minutes.
  • LOCAL_READY permits no missing-asset download and at most one GiB of disposable generated fixtures and results, thirty CPU minutes, and fifteen GPU minutes.
  • Moderate and heavy tiers are classified but cannot be acquired automatically. Credentials and paid APIs require explicit authorization and are unavailable to autonomous mode.
  • Approved source classes are HTTPS GitHub, Hugging Face, Zenodo, PyPI, and arXiv locations. Licenses, immutable provenance, and SHA-256 receipts are mandatory.
  • Downloaded content remains untrusted input and may execute only in a disposable, network-disabled environment with read-only source, writable scratch, resource limits, dependency locking, archive and malware checks, and setup hooks disabled until review.
  • Research assets have no authority to modify production, construct or approve a candidate, request promotion, or change the governor.

Experiment-feasibility assessment

Implemented
  • Added relevance states HIRO_RELEVANT, TRANSFERABLE_WITH_EXPLICIT_HYPOTHESIS, and NOT_HIRO_RELEVANT.
  • Added feasibility tiers LOCAL_READY, ACQUIRABLE_LIGHT, ACQUIRABLE_MODERATE, ACQUIRABLE_HEAVY, UNAVAILABLE, and NOT_HIRO_RELEVANT.
  • The planner preserves the exact validated claim and may transfer only a concrete shared mechanism with an explicit Hiro-local hypothesis; generic topical similarity is insufficient.
  • Each plan records every required asset and its availability, URL, license, byte estimates, requirement status, notes, total resource estimates, isolation controls, effort, observable, falsification criterion, and tier reason.
  • A separate model operation audits source preservation, transfer defensibility, missing-asset completeness, source support, resource conservatism, policy coherence, and meaningful falsification. Deterministic validation independently enforces schemas, approved hosts, resource limits, and inventory consistency.
  • Immutable reports and companion hashes preserve both model judgments and deterministic results. A revalidation operation can apply a corrected deterministic policy without rerunning or rewriting either frozen model judgment.

Actionable discovery

Implemented
  • The new bounded strategy searches empirical work in tool use and selection, retrieval and memory, self-correction and verification, and agent planning and evaluation.
  • Ranking records distinct Hiro-relevance, empirical-claim, numeric-result, and asset-or-reproduction signals. It prefers actionability but does not select for a predetermined conclusion or claimed improvement.
  • Seventy-one sources were considered: fifty-three arXiv records, fifteen sanitized Moltbook records, and three GitHub records. Twenty were retained and frozen before claim extraction: fifteen arXiv and five Moltbook.
  • Fifty-one sources were explicitly excluded: one overall capacity limit, nineteen without asset or reproduction signal, eleven without an empirical-claim signal, one without a Hiro-relevance signal, seventeen per-source capacity limits, and two prompt-injection risks.
  • The unchanged Phase 3B claim extractor then produced thirty-four claims from the frozen twenty-source corpus. All thirty-four passed claim validation; there were no extraction failures or unavailable-source outcomes.

Observed first divergence and minimal repair

Completed
  • The first new feasibility run classified sixteen plans LOCAL_READY, and the independent semantic auditor approved their resource bounds, but deterministic validation rejected fourteen because any nonzero generated storage exceeded a zero-byte LOCAL_READY limit.
  • That limit confused asset acquisition with scratch output. A local experiment can require no download while still generating kilobytes of fixtures and measurements.
  • The only policy correction was to allow up to one GiB of bounded disposable scratch storage for LOCAL_READY while retaining the zero-download rule and all isolation controls.
  • Revalidation reused the same frozen planner and auditor outputs. It produced fourteen valid LOCAL_READY experiment plans; no semantic threshold, source claim, discovery result, or model output was changed.
  • Two assessments remain invalid: one optional PeakBench asset was omitted from its exact missing-assets inventory, and one scientifically irrelevant materials claim had internally inconsistent semantic fields. Neither represents a silently lost valid experiment.

Final qualification funnel

Actionable discovery works
  • SOURCES_CONSIDERED 71; SOURCES_RETAINED 20; CLAIMS_EXTRACTED 34; CLAIMS_VALIDATED 34.
  • HIRO_RELEVANT or explicitly transferable 16; NOT_HIRO_RELEVANT 18.
  • LOCAL_READY 16; ACQUIRABLE_LIGHT 0; ACQUIRABLE_MODERATE 0; ACQUIRABLE_HEAVY 0; UNAVAILABLE 0.
  • EXPERIMENT_PLANS_VALID 14. The plans cover planner/executor handoffs, graph-based failure attribution, resource-aware scheduling, multi-step planning degradation, failure attribution, verification rules, misleading-premise handling, answer selection, judge reliability, and repository-context test generation.
  • No asset acquisition, experiment execution, reproduction, candidate construction, governor decision, promotion request, activation, or rollback occurred.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen actionable discovery corpus passed 71 sources considered, 20 retained, and 51 explicitly excluded. The frozen corpus SHA-256 is 23006b78f3a86024bf93cbb43b7e22f1dfa968439ef27540788584484f64b19f.
Unchanged source-claim qualification passed 20 frozen source records produced 34 extracted and 34 validated claims with zero extraction failures. The immutable report SHA-256 is 7d45b4a5bed2b4bc4964df5054e88913b434726f4fc1fdccf0aedadeb2747915.
New experiment-feasibility qualification passed 34 claims were assessed; 16 were LOCAL_READY, 18 NOT_HIRO_RELEVANT, and 14 produced complete valid light-or-local plans. The corrected deterministic revalidation SHA-256 is cf574283c1f6b1144ff3bc43315815b6eda3abd6d8fc4fd2181ca6ad159db093.
Existing twenty-five-claim reclassification passed All 25 prior claims received explicit asset and feasibility assessments: 24 NOT_HIRO_RELEVANT and one ACQUIRABLE_LIGHT valid plan. The report SHA-256 is fd9bb6a29d7552414555b25d51a531f2e147ce57545ea9cd871d63aa9085e1ee.
Focused discovery, feasibility, and source-claim tests passed 12 tests passed in 0.54 seconds. One non-failing warning concerned the inaccessible pytest cache directory.
Full Hiro repository regression suite passed 815 tests passed in 443.66 seconds. Six non-failing warnings concerned five existing unregistered test marks and the inaccessible pytest cache directory.
Public journal tests and build passed npm run test:hiro passed. npm run build generated and validated 172 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; the install reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session.

Current state

Next steps