Hiro development journal

Phase 3B preserves source claims but finds no locally testable inputs

Mixed bottleneck; extraction is repaired, while the fixed fresh sources yield zero claims testable with Hiro's current local assets Machine-readable JSON

Executive summary

Qualified only the source-retention and claim-extraction boundary established by Phase 3A. Corroboration policy, reproduction execution, candidate construction, promotion, activation, rollback, autonomous implementation, and meta-improvement were unchanged and not exercised.

Traced both active source paths before changing code. Moltbook and arXiv content survived retrieval transiently, then `_reduce_item` intentionally collapsed it into keyword and mechanism labels; persisted packets retained hashes but not the source assertion needed for a falsifiable claim.

Added an immutable, source-grounded representation containing canonical identifiers, title and author/date metadata, retrieval time, bounded observed source text, exact location, content hashes, injection assessment, and explicit untrusted-evidence authority boundaries.

Froze a fresh corpus before extractor tuning using a preregistered selection rule: the first ten eligible Moltbook records and first ten arXiv records in response order, with no relevance or claim-quality selection.

Added zero-to-many structured claim extraction plus deterministic provenance checks and an independent tool-free semantic validation pass. Extraction and validation outputs remain separate and every accepted claim links to an exact supporting character span and source hash.

The final fixed-corpus run processed twenty source records, extracted twenty-five claims, and validated all twenty-five without an extraction failure or unavailable source.

All twenty-five concrete claims were classified CLAIM_SPECIFIC_NOT_LOCAL. Twelve source records were CLAIM_INCOMPLETE; eight arXiv records produced concrete non-local claims; no record produced a claim honestly testable with Hiro's current repository, runtime, evaluations, ordinary synthetic inputs, and available assets.

The primary disposition is MIXED BOTTLENECK. The old representation destroyed concrete source information, and that extraction defect is now repaired; however, the current fixed discovery material still supplies zero operational local claims suitable for Phase 3A. Phase 3A should not resume yet.

Work completed

Pre-change information-loss trace

Completed
  • Moltbook retrieval received sanitized public-post JSON containing source ID, title, a content excerpt, date, author information, and engagement. The transient parser retained only a subset and used a generic endpoint URL.
  • arXiv retrieval received Atom entries containing canonical ID, title, abstract, publication metadata, authors, and link. The transient parser discarded authors and update time, shortened the title and abstract, and replaced the canonical item ID with a hash.
  • The decisive loss occurred in `_reduce_item`, which used source prose only for keyword matching and then persisted a mechanism label, technique tuple, code-owned question, and hashes. The exact source claim was explicitly empty.
  • Original assertions could not be recovered from persisted provenance alone. Recovery required a network refetch whose content could change or disappear.
  • The loss was partly intentional prompt-injection isolation and partly an insufficient schema: content was prevented from becoming instruction, but evidence necessary for faithful claim extraction was also removed.

Immutable source retention

Implemented
  • Added a bounded source-record schema with source type, original identifier, source URL, title, authors, publication and retrieval timestamps, observed source field and response position, bounded source text, SHA-256 identities, truncation status, and injection assessment.
  • The frozen corpus uses read-only JSON and a companion SHA-256 file. Its selection policy records that source order, not relevance or apparent testability, chose records before extraction existed.
  • The Moltbook endpoint returned exactly five hundred observed characters for each selected post. Those excerpts were preserved faithfully and are not represented as complete posts.
  • The arXiv parser now retains author and update metadata while preserving its existing XML entity rejection and approved-link boundary.
  • Source text remains untrusted evidence with no authority to construct candidates, run experiments, request promotion, or change state.

Source-grounded claim extraction

Implemented
  • The extractor produces zero to five distinct claims per source and separates claim specificity from local feasibility.
  • Claim fields include exact claim, intervention, baseline or comparison, outcome, metric or observable, conditions, direction, evidence type, confidence, testability state and reason, proposed measurable observable, and exact supporting excerpt.
  • The explicit states are CLAIM_SPECIFIC_TESTABLE, CLAIM_SPECIFIC_NOT_LOCAL, CLAIM_INCOMPLETE, EXTRACTION_FAILED, and SOURCE_UNAVAILABLE.
  • CLAIM_SPECIFIC_NOT_LOCAL is a successful extraction state, not a failure. A concrete paper result remains retained even when reproduction would require the authors' datasets, models, simulator, training pipeline, domain corpus, large-scale training, or specialized hardware.
  • CLAIM_SPECIFIC_TESTABLE is intentionally strict: the claim must be falsifiable with Hiro's current repository, local model and runtime, existing evaluations, and ordinary synthetic inputs, without first recreating the source system.

Independent provenance validation

Implemented
  • Deterministic validation requires a verbatim contiguous excerpt, valid schema enumerations and confidence, complete provenance, correct hashes and offsets, and no unsupported numeric fact in the claim.
  • Common arXiv LaTeX rendering for percentages and degrees is normalized only for comparison; the retained source bytes remain unchanged.
  • A separate tool-free local-model operation assesses source support, material additions, field-to-claim agreement, and local-testability coherence. It cannot rewrite the source claim.
  • Extractor output, validator output, deterministic reason codes, and the final provenance decision are all persisted independently in the immutable report.
  • A claim is accepted only when the deterministic and independent semantic checks both pass.

Observed boundary repairs

Completed
  • The first run showed that the semantic model treated access to hypothetical external datasets or simulators as local testability. The minimum repair defined local availability concretely and allowed the independent validator to correct only the local-versus-non-local label.
  • The next run showed that an overly strict extraction instruction collapsed concrete external claims into CLAIM_INCOMPLETE. The minimum repair made specificity independent of locality and required concrete non-local claims to be retained.
  • A scholarly phrase about revealing token-level visual dependence triggered the credential-exfiltration detector. The minimum detector repair narrowed generic token matching while a paired test confirmed that an actual access-token exfiltration request remains rejected.
  • One faithful numerical claim used ordinary percent and degree language while the source used LaTeX notation. The deterministic checker now recognizes those equivalent renderings without weakening semantic validation.
  • Repeated reports originally shared the corpus timestamp and collided. Qualification reports now receive their own immutable qualification timestamp while continuing to reference the unchanged corpus hash.

Final fixed-corpus result

Mixed bottleneck
  • Twenty source records were processed: ten Moltbook post excerpts and ten arXiv abstracts.
  • Twenty-five claims were extracted and all twenty-five passed deterministic and independent semantic provenance validation. Eight sources produced multiple valid claims.
  • Source classifications were zero CLAIM_SPECIFIC_TESTABLE, eight CLAIM_SPECIFIC_NOT_LOCAL, twelve CLAIM_INCOMPLETE, zero EXTRACTION_FAILED, and zero SOURCE_UNAVAILABLE.
  • Claim classifications were zero CLAIM_SPECIFIC_TESTABLE and twenty-five CLAIM_SPECIFIC_NOT_LOCAL.
  • All ten Moltbook excerpts were incomplete. Eight arXiv abstracts yielded concrete non-local claims, while two arXiv abstracts remained too qualitative for claim extraction.
  • No reproduction experiment, candidate, gate decision, promotion request, activation, or rollback was created.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Immutable qualification corpus passed Twenty records were frozen before extractor tuning: ten Moltbook and ten arXiv. The immutable corpus SHA-256 is f28447b47a7fcdac032acf0539fd6e93fb3e94fe8c5760d4acf2007493957b5b.
Final source-claim qualification passed 20 source records, 25 extracted claims, 25 validated claims, 8 sources with multiple claims, 0 extraction failures, and 0 unavailable sources. The immutable report SHA-256 is 873b9b32eb00612350f564f3b49294f1ae3e15bd05b1d120184900d086b0404a.
Final classification integrity passed Source states: 0 testable, 8 specific non-local, 12 incomplete, 0 extraction failed, 0 unavailable. Claim states: 0 testable and 25 specific non-local.
Focused Phase 3B tests passed 27 focused tests passed, covering immutable hashes, source metadata, exact-span provenance, independent validation, numeric rendering, XML safety, and injection-detector controls.
Full Hiro repository regression suite passed 806 tests passed in 426.55 seconds. Six non-failing warnings concerned pre-existing unregistered test marks and an inaccessible pytest cache directory.
Public journal tests and build passed npm run test:hiro passed. npm run build generated and validated 171 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci because dependencies were not installed; after the locked install both required commands passed.

Current state

Next steps