Hiro development journal

Bounded source evidence did not recover a lost transfer hypothesis

Qualification stopped at the first material false negative Machine-readable JSON

Executive summary

Hiro tested a narrowly scoped source-grounded relevance boundary using the immutable six gold propositions, the original ten hosted grounded claims, and their accepted proposition-family mapping. No extraction, consolidation, capability probing, reproduction, candidate construction, promotion, production routing, or Phase 3F activity occurred.

The qualification froze a small deterministic context window: the verified supporting sentence plus at most its immediately preceding sentence, no following sentence, capped at 1,200 characters. Every window was checked against exact source offsets and the immutable source-text hash before the planner could see it.

The control arm completed all six relevance and feasibility assessments. The hosted arm was stopped at the first required recovery failure. The downshift trajectory-context family remained NOT_HIRO_RELEVANT instead of recovering the control's TRANSFERABLE_WITH_EXPLICIT_HYPOTHESIS and LOCAL_READY decision.

The false-span control remained correctly rejected and did not become an independent opportunity. Provenance-based accounting also showed that the ten hosted fragments can be reduced to seven lineage units, representing six substantive families plus the false-span control.

Final disposition: SOURCE-GROUNDED RELEVANCE RECOVERY NOT DEMONSTRATED — BOUNDED_SOURCE_EVIDENCE_TO_TRANSFER_HYPOTHESIS. Exact atomic representation remains the first unresolved operational blocker under the current planner semantics, and the frozen twenty-source corpus remains uncleared.

Work completed

Frozen evidence boundary

Implemented and verified
  • The relevance planner gained an optional evidence input that is absent by default, preserving existing production behavior.
  • Structured claim, immutable source evidence, and transfer interpretation are recorded as separate objects. Source evidence cannot rewrite or silently repair the canonical claim.
  • The context rule includes the sentence containing the verified span and at most one immediately preceding sentence from the same immutable source record. It includes no following sentence or unrelated source and rejects invalid offsets, hash mismatches, exact-span mismatches, and oversized evidence.

Qualification preflight and lineage

Passed
  • All ten original hosted spans matched their source offsets and source-text hashes. The resulting context windows ranged from 112 to 432 characters, below the frozen 1,200-character ceiling.
  • A deterministic lineage key based on source record, exact supporting offsets, and normalized intervention collapsed the ten hosted claim IDs to seven lineage units without merging the two distinct propositions that share one compound source sentence.
  • A lineage can create at most one later opportunity. This prevents duplicate fragments from independently consuming bounded downstream opportunity capacity or masquerading as independent corroboration.

Control relevance baseline

Completed
  • All six frozen gold propositions completed the unchanged relevance and feasibility planner plus its existing independent validator.
  • Four controls reached valid TRANSFERABLE_WITH_EXPLICIT_HYPOTHESIS and LOCAL_READY decisions. Two correctly terminated as NOT_HIRO_RELEVANT.
  • The downshift trajectory-context control produced a valid hypothesis that downstream execution quality depends on retaining the upstream planning trajectory.

Hosted source-aware recovery

Stopped at first divergence
  • The false-span control remained NOT_HIRO_RELEVANT with no valid transfer plan, establishing that the bounded evidence did not automatically turn an irrelevant fragment into an opportunity.
  • The hosted downshift claim received verified evidence covering the complete compound source sentence. Nevertheless, the planner characterized it as specific to external model interactions and produced no transferable Hiro mechanism or local hypothesis.
  • This differed from the gold control at BOUNDED_SOURCE_EVIDENCE_TO_TRANSFER_HYPOTHESIS and was classified MATERIAL_FALSE_NEGATIVE. The remaining required hosted families were not run after that boundary, in accordance with the stop rule.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen-input integrity passed The historical corpus, historical claim report, six gold controls, original hosted report, and accepted proposition-family mapping matched their frozen hashes.
Context provenance passed Ten of ten hosted claims produced exact, same-source evidence windows with valid offsets and source-text hashes.
Focused qualification tests passed Fifteen source-grounded relevance, feasibility, and prior representation-equivalence tests passed. Python compilation and repository diff checks also passed.
False-positive protection passed for the exercised control The known false-span control remained NOT_HIRO_RELEVANT and produced no valid transfer plan.
Required trajectory-context recovery failed The gold control remained transferable and local-ready, while the hosted source-aware claim remained not Hiro-relevant. The first divergent boundary was BOUNDED_SOURCE_EVIDENCE_TO_TRANSFER_HYPOTHESIS.
Authority containment passed No extraction, capability probe, reproduction, candidate, promotion, production routing, twenty-source qualification, or Phase 3F action occurred.

Current state

Next steps