Hiro development journal

Designing a bounded adaptive-discovery comparison

Research design recorded; execution awaits approval Machine-readable JSON

Executive summary

Designed a comparison that distinguishes fixed-query repair from value added by adaptive discovery. No production discovery policy, ontology or observer safety state changed.

Extended the existing external-research-synthesis problem with one proposed hypothesis; no new top-level problem, child HRO or experiment was created. Status is blocked_on_measurement.

Prior-art review classifies the proposal as an application of established retrieval and exploration methods, with its value for Hiro unmeasured.

Historical packets cannot reconstruct counterfactual GitHub queries; any historical metadata experiment must be labeled a simulation and pass evidence-availability gates.

Work completed

Existing-state reconciliation

Completed
  • Reviewed all 11 registry problem identities. Acquisition fits the existing external synthesis problem and relates to capability discovery, agenda governance and evaluator validity.
  • Kept the prior improvement-dimension discovery question separate. Fixed retrieval evidence does not establish ontology inadequacy.
  • Attached the existing forensic provenance and a hash-bound design artifact as observations, not experimental results. Preserved the parent question and existing experiment links.
  • Canonical research edit attributed to Work, originating request attributed to Alex. No independently generated Hiro research is claimed.

Prior art and critical review

Completed with stated limits
  • Reviewed relevance-based language modeling, focused crawling, active information acquisition, novelty search, curiosity, FLARE adaptive retrieval, ResearchAgent and Voyager.
  • The methods establish useful precedents, not evidence that adaptive discovery beats static breadth under Hiro constraints. No algorithmic novelty claimed.
  • Identified baseline weakness, hindsight case selection, incomplete historical indexes, model contamination, evaluator circularity and hidden compute as principal confounds.

Smallest discriminating design

Proposed only
  • A preserves the frozen narrow repository query; B uses a credible broader static schedule; C adapts evidence-grounded vocabulary within the same source; D changes allocation without expanding vocabulary.
  • Target 24 case clusters with six development and 18 sealed cases, independent relevance judgments, negative controls and origin-source strata. Named cases and historical corpus have not been assembled.
  • The motivating Jev announcement family remains outside the pilot. No hindsight-derived search terms or arbitrary source expansion are permitted.
  • Compare encounter recall, recognition losses, retained precision, novelty, time, diversity, actionable yield, unedited downstream experiment specifications, cost and redundancy separately.
  • Proposed useful-effect target: 10 percentage points of encounter-recall gain over broad static, at least 60 percent retained precision and no more than five points precision loss. Small-sample uncertainty can leave the pilot inconclusive.
  • All planning costs count against the same resource envelope. Sequential offline arms, bounded branch lifetimes, fixed endpoint authority and full stage lineage are required.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Research schema and canonical state passed Validated the research edit/context before application and exported HRO after application. Registry remains at 11 problems, parent question and experiment links unchanged; hypothesis proposed and execution_authorized false.
Protected source and safety files passed Before/after hashes match for inspected source policy, adapter, configuration, knowledge-base and observer sentinel files.
Empirical benchmark / hardware throughput / Hiro test suite not_run Design and state-recording work only. No discovery rerun, source replay, model benchmark or production code change.
Journal tests/build passed npm run test:hiro followed by npm run build passed; 234 entries validated with timestamp, alias and noindex checks, plus TypeScript and Vite build. No Hiro benchmark ran.

Current state

Next steps