Hiro development journal

Hiro's autonomous idea pool expands beyond Reddit to GitHub, arXiv, and official technology feeds

Source expansion active; grounded idea representation repaired and live candidate processing resumed Machine-readable JSON

Executive summary

Reddit is no longer an active dependency for Hiro's improvement discovery. Its adapter remains disabled and fail-closed for possible future use.

GitHub's public repository-search API, arXiv's official Atom API, BBC Technology RSS, WIRED AI RSS, and Ars Technica Technology Lab RSS now feed the same governed idea-lineage pipeline as Moltbook.

The idea pool can retain up to 100 ranked leads per discovery packet, with a 20-item per-source ceiling and round-robin source diversity so one high-engagement source cannot consume the pool.

A live discovery reviewed 91 items and persisted 17 new, deduplicated ideas from five new source families; Hiro rejected the first non-improving candidate and automatically advanced the next idea without operator intervention.

External text remains untrusted evidence: it is screened, deterministically reduced to a theme and opaque reference, and never treated as instruction or authority.

Post-activation review found that the original reduction was too lossy for useful autonomous implementation: distinct leads collapsed into identical themes, scores, and generic candidate prompts, so the strict evaluation gates correctly rejected mostly ungrounded patches.

That defect is now repaired: external text is converted into one of ten code-owned technical mechanisms, supporting observations are clustered, scores incorporate mechanism and corroboration, and candidate construction receives a specific intent, file boundary, metrics, and generated test path.

A fresh live collection reviewed 91 items and produced seven distinct clustered proposals. The queue is advancing autonomously; the first grounded proposal closed safely during construction, later proposals advanced automatically, and a moderate-risk latency/freshness proposal was in construction at the final verification checkpoint.

Work completed

First-class source adapters

Activated
  • Added bounded, read-only adapters for the official GitHub REST search endpoint and the official arXiv Atom query endpoint.
  • Added feed-only adapters for BBC Technology, WIRED AI, and Ars Technica Technology Lab; linked article pages are not fetched.
  • Kept every source name, endpoint, query, response budget, and permitted article-link origin in a code-validated allowlist rather than permitting arbitrary configured domains.
  • Disabled Reddit in the active policy so missing Reddit approval cannot delay the discovery cycle.

Capacity and source diversity

Implemented
  • Raised discovery packet capacity from 3 ideas to 100 and raised downstream external-idea intake to the same ceiling.
  • Limited each source to at most 20 accepted ideas in a packet.
  • Changed final selection to take one ranked idea from each productive source before taking a second from any source, repeating by rounds until the pool is full.
  • Added a consumer-experience theme and a lower news-feed relevance threshold so product and usability signals can compete without lowering the threshold for technical sources.

Rate and format controls

Implemented
  • GitHub, Moltbook, and news feeds are fetched no more than once every two hours; arXiv is fetched no more than once per day.
  • Successful fetch times are persisted separately from content deduplication state, allowing independent source cadence without losing idea lineage.
  • All network calls are single-connection GET requests with redirects disabled, strict response-size budgets, accepted-content-type checks, and no credentials for the new sources.
  • XML document types and entity declarations are rejected before parsing, and article links must match exact approved origins.

Live activation

Processing
  • Restarted Hiro through the hidden Windows launcher after committing the implementation.
  • Confirmed Telegram remained disabled during restart and that the self-improvement run skipped notification delivery.
  • The first expanded live packet contained 8 arXiv ideas, 4 Ars Technica ideas, 3 BBC Technology ideas, 1 WIRED idea, and 1 GitHub idea after deduplication.
  • The continuous queue reached 28 total records. The first arXiv-derived candidate failed its independent benefit gates, after which the next scheduler tick automatically moved a BBC-derived idea into candidate construction.

Post-activation candidate-quality diagnosis

Structural defect confirmed
  • The visible count of 17 rejections combines 6 older terminal Moltbook records, the deliberately bad negative control, and 10 organic candidate failures; it is not 17 newly tested organic failures.
  • The remaining queue has 7 evaluation-reliability entries with identical impact 5.38 and priority 10.28, plus 3 memory-and-retrieval entries with identical impact 5.36 and priority 10.16.
  • The deterministic reducer preserves only a broad theme, generic question, and opaque source hash. It discards the concrete mechanism, proposed behavior, affected capability, expected metric, and distinguishing evidence needed for ranking or implementation.
  • Candidate construction receives that generic summary and a hard-coded theme-to-file mapping, so candidates have included generic authority directives, an unrelated model profile, sports-routing regex changes, and in one case no substantive product change.
  • Most evaluated candidates produced zero or near-zero public and held-out score deltas; several also worsened latency or held-out categories. The gates therefore rejected them for valid evidence-based reasons.

Prompt-safe idea grounding repair

Implemented and activated
  • Replaced generic theme-plus-hash queue summaries with code-owned mechanism records covering fault injection, benchmark coverage, retrieval freshness, context memory, latency and freshness, rollback resilience, provenance replay, tool routing, decision traceability, and assistant interaction.
  • External titles and summaries remain outside candidate prompts and persistent safe summaries. The reducer emits only approved mechanism names, declarative proposal templates, intended behavior, bounded metrics, opaque references, and source counts.
  • Clustered observations by mechanism before queue insertion. Multiple papers or stories now increase one proposal's support and confidence rather than occupying repeated queue positions.
  • Added a queue schema migration for structured details, mechanism-aware impact and priority scoring, and a separate terminal state for superseded theme-only records so representation migration is not reported as experimental failure.
  • Replaced broad theme-to-file routing with validated mechanism plans containing exact candidate paths, change intent, risk class, metrics, and generated test paths.
  • Added a boundary test that verifies every mechanism path exists and clears both the isolated candidate builder's protected-path rules and the stable governor's promotion policy.
  • Updated the benchmark queue cards to show the mechanism, supporting observation count and sources, and the metrics that will decide the candidate.

Live migration and reactivation

Autonomous processing active
  • Paused queue transitions during the repair while leaving discovery available, then restarted Hiro through the hidden launcher after tests passed.
  • A fresh mechanism-v2 collection reviewed 91 source items and yielded seven distinct proposals spanning benchmark coverage, routing, observability, provenance, latency, context memory, and retrieval freshness.
  • Moved ten generic waiting records into a separate superseded state. The queue now reports 17 historical rejections, 10 superseded representations, seven actionable proposals, and one earlier implementation.
  • Found and corrected a stale circuit breaker caused by three earlier protected-path failures. The relevant mechanism plans now target permitted worker surfaces, the breaker is closed, and its failure counter is zero.
  • Inspected the first live mapping before releasing the loop: seven arXiv and Ars observations map to benchmark coverage, hiro/improvement/benchmark_generator.py, one generated targeted test, and case-coverage, score-reproducibility, and held-out pass-rate metrics.
  • The first grounded benchmark proposal produced a relevant benchmark-generation plan, but its generated test used a recursive cleanup primitive prohibited by the patch-safety layer. The complete candidate was discarded without touching the active branch.
  • Kept the deletion protection intact and added explicit construction constraints requiring pytest temporary paths and prohibiting cleanup, process, network, credential, and dynamic-execution primitives in generated tests.
  • Added an explicit vocabulary translation because the candidate-spec contract calls moderate risk 'medium' while the stable governor calls it 'moderate'. The affected routing item was deferred for retry rather than lost, and the circuit breaker remained closed.
  • At the final checkpoint the queue had advanced past additional closed candidates and was constructing a moderate-risk latency/freshness proposal against core/router.py. No promotion is claimed.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused external-source security and parsing suite passed 19 tests passed, including XML entity rejection, approved-origin enforcement, injection screening, lineage loading, deduplication, and source-diverse ranking.
Continuous engine, governor, active loop, workflow, and dashboard integration suite passed 49 tests passed.
Full Hiro repository suite passed 520 tests passed in 158.57 seconds.
Live endpoint preview passed All six enabled source adapters returned valid bounded data; Reddit remained disabled. Ninety-one items were reviewed and 24 were relevant before persistent deduplication.
Live persisted discovery and queue ingestion passed Seventeen new safe idea records were frozen with checksums and ingested; the top-ranked record advanced, failed its benefit gates, and the next eligible record was selected automatically on the following scheduler tick.
Rejection and candidate-diff audit failed design expectation Ten organic candidates were evaluated and none cleared the gates. Their diffs showed weak or absent grounding in the originating idea, while queue scores were identical within each broad theme. This identifies upstream representation and construction as the primary defect rather than excessive promotion strictness.
Grounded representation, clustering, migration, dashboard, and path-boundary tests passed 38 focused tests passed, including mechanism distinction, semantic clustering, legacy lineage compatibility, mechanism-specific scoring and candidate mapping, superseded-record handling, dashboard API behavior, and validation of every candidate path against both protection layers.
Full Hiro repository suite after activation repair passed 526 tests passed in 162.94 seconds against commit c098469 after the path, risk-vocabulary, and generated-test safety repairs.
Fresh mechanism-v2 live discovery passed Ninety-one items were reviewed and reduced to seven distinct clustered proposals; Reddit remained disabled and no raw external text was propagated.
First live mechanism-to-candidate mapping mapping passed; candidate closed safely The top proposal mapped to a permitted benchmark-generation file, a generated targeted test, explicit evaluation metrics, seven opaque supporting references, and two supporting source families. Construction later discarded the candidate because its generated test included a prohibited recursive-cleanup primitive; the active branch was unchanged.
Moderate-risk contract translation passed Focused tests confirm that the queue retains the governor's 'moderate' classification while emitting the candidate-spec contract's required 'medium' value.
Generated-test safety guidance passed Candidate-builder tests confirm that every model request now requires nonempty patches, all targeted tests, pytest temporary paths, and no cleanup, process, network, credential, or dynamic-execution primitives.

Current state

Next steps