Hiro development journal

Phase 3F-D demonstrates repaired discovery retention

Hiro implementation, focused qualification tests, public journal tests, and production journal build passed Machine-readable JSON

Executive summary

Phase 3F-D audited the exact 72-source corpus from the terminated fresh Phase 3F campaign without performing new discovery.

The original 72-to-1-to-0 funnel had two material discovery defects: Moltbook and GitHub were ranked from incomplete representations, and claim-worthy arXiv abstracts were rejected before claim extraction solely because they did not advertise code, datasets, or other reproduction assets.

Independent fuller-representation qualification found three genuinely claim-worthy fresh sources. All three were arXiv papers about agent token cost, KV-cache repair, and ledger-based coding-agent orchestration; the original selector retained none of them.

The minimum repair now obtains bounded complete post or README representations where available, records incomplete capture explicitly, and treats asset availability as a ranking bonus rather than a pre-claim retention requirement.

The repaired frozen rerun considered the same 72 identities, sent 20 broad sources to claim extraction, recovered 50 raw and 48 validated claims, and retained all three independently established claim-worthy sources with 12 Hiro-relevant validated claims.

Pre-claim screening achieved 100% recall and 15% precision by design; the independent claim and relevance stages restored final precision and recall to 100%. No production-gap probe, reproduction, candidate, governor, promotion, or meta-improvement path ran.

The terminal disposition is PHASE 3F-D DEMONSTRATED — DISCOVERY RETENTION WORKS. A separately authorized fresh Phase 3F campaign is now justified, but none was started.

Work completed

Frozen 72-source funnel reconstruction

Completed
  • Preserved the original source identities, ordering, retrieved representations, retention decisions, exclusion reasons, and freshness decisions.
  • The original exclusions were 31 previously qualified identities, 21 lacking asset or reproduction keywords, 11 lacking empirical signals, six lacking Hiro-relevance signals, and two rejected for prompt-injection risk.
  • The complete source-by-source audit classified the frozen set as 31 duplicate or freshness rejects, 38 true negatives, three retention false negatives, and zero content-insufficient sources after bounded completion.
  • Among the 71 originally rejected sources, the audited disposition was 31 duplicate or freshness rejects, 37 true negatives, three retention false negatives, and zero content-insufficient sources.

Representation and capture audit

Completed
  • All 54 arXiv records used complete Atom abstracts of 600 to 1,919 characters; no fuller abstract was available from the feed and none was marked truncated.
  • All three GitHub records used only repository descriptions and topics, between 235 and 318 characters, even though bounded README representations of 1,814 to 12,000 characters were available.
  • All 15 Moltbook records used 498-to-500-character listing previews. Exact read-only post endpoints supplied bounded complete representations of 652 to 2,521 characters.
  • The old truncation flag examined already-bounded text and therefore mislabeled 500-character previews as complete. The repaired layer uses explicit representation state and bounded exact-identity completion before ranking.
  • Full-content prompt-injection assessment rejected two GitHub README representations that the description-only screen could not see. Those sources remained terminal rejects.

Independent claim and relevance reference set

Completed
  • Thirty-seven fresh, non-injection fuller records entered the existing independent claim extractor and validator. They produced 62 raw claims and 58 validated claims across 19 sources.
  • A separate source-level relevance audit applied the exact Phase 3F-D positive definition: plausible direct Hiro relevance or an explicit transferable mechanism, plus validated source-grounded claims.
  • Three sources met the reference definition. Their source text hashes, validated claim identifiers, provenance checks, and independent relevance assessments were preserved in immutable evidence.
  • The original retained Moltbook source did not meet the reference definition after completion, yielding original retention precision and recall of zero against this corpus.

Minimal production discovery repair

Completed
  • Added deterministic content-completeness state with explicit SOURCE_CONTENT_INCOMPLETE outcomes instead of silently treating bounded previews as complete.
  • Added bounded exact-post completion for Moltbook and bounded README completion for GitHub, with identity verification, byte and timeout limits, and prompt-injection reassessment on fuller content.
  • Removed asset or reproduction keywords as a mandatory pre-claim gate. Asset evidence remains a score bonus, while relevance, empirical signal, numeric observables, score thresholds, source limits, freshness, and injection protections remain active.
  • Refactored deterministic source screening into one production selector used by live capture and the frozen qualification rerun.
  • No threshold was indiscriminately lowered and no downstream promotion or experimental policy changed.

Repaired frozen rerun

Completed
  • The repaired source layer considered the same 72 identities and obtained adequate bounded representations for all 72.
  • After preserving 31 freshness rejects and four full-content prompt-injection rejects, 37 fresh non-injection records were screened and 20 were retained for claim extraction.
  • The exact source hashes matched the immutable prior extraction run, so 50 raw and 48 validated claim results were reused without a new semantic result being fabricated.
  • All three reference-positive sources survived the broad pre-claim screen, giving 100% pre-claim recall and 15% precision. Independent claim-level retention kept exactly those three, giving 100% final precision and recall and 12 Hiro-relevant validated claims.
  • Nine representative rejects spanning source types and terminal classes received independent hash-bound validation samples.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen-corpus identity passed The diagnostic corpus contained exactly the original 72 ordered identities; no new discovery identity was admitted.
Freshness policy passed All 31 freshness rejects were exact prior source-record matches. No materially changed new version or distinct source was incorrectly eliminated in this frozen corpus.
Content completeness passed The repaired frozen rerun had adequate bounded representations for all 72 sources; incomplete list previews no longer silently pass as complete.
Retention recall passed All three independently established claim-worthy sources survived repaired screening and claim-level qualification, for 100% recall.
Retention precision passed The intentionally broad pre-claim screen had 15% precision; independent source-grounded claim and relevance qualification retained exactly three reference positives, for 100% final precision.
Focused Hiro tests passed Twenty-one focused tests covering discovery selection, completeness, bounded Moltbook and GitHub completion, source claims, and router output behavior passed. The only warning was an existing local pytest-cache permission warning.
Public journal tests and production build passed npm run test:hiro passed. The clean checkout's first build generated and validated 184 journal pages, then stopped because locked dependencies had not yet been installed. After npm ci installed 33 packages, npm run build regenerated and validated all 184 pages, compiled TypeScript, and completed the Vite production bundle. Installation reported one existing high-severity dependency advisory outside Phase 3F-D scope.

Current state

Next steps