Hiro development journal

Hosted claims lost two viable transfer families downstream

Representation-equivalence qualification completed and stopped Machine-readable JSON

Executive summary

Hiro tested whether the original ten canonical claims produced by the hosted grounded extractor were operationally equivalent downstream to the six frozen gold claims. This was a representation-equivalence qualification only: no extraction, discovery, reproduction, candidate construction, promotion, or production mutation occurred.

Both immutable claim sets passed their already-recorded deterministic, semantic, and provenance validation. Each set then traversed the same current experiment-feasibility planner, independent plan validator, current-capability assessor, and independent capability auditor against one repository revision and one shared telemetry cutoff.

The representations were not operationally equivalent. Three proposition families were decision-equivalent, two were benign duplications, and two were material false negatives. In both false-negative cases, the gold claim produced a valid local Hiro-transfer experiment plan while its hosted counterpart or counterparts were rejected as not Hiro-relevant before current-capability assessment.

The hosted false span terminated safely as not Hiro-relevant and created no capability assessment, reproduction, corroboration, or candidate lineage. The hosted set nevertheless expanded six proposition families into ten independent claim IDs, causing four additional feasibility planning and validation units and creating a capacity-amplification risk because no semantic family deduplication occurs before feasibility.

Final disposition: HOSTED GROUNDED EXTRACTION NOT OPERATIONALLY SUFFICIENT — CLAIM_REPRESENTATION_TO_HIRO_RELEVANCE. The frozen twenty-source qualification remains uncleared, and Phase 3F remains stopped.

Work completed

Immutable control and hosted inputs

Verified
  • The control consisted of exactly six frozen gold canonical claims from three positive sources. The hosted arm consisted of the exact original ten hosted canonical claims from the pre-consolidation qualification.
  • The accepted post-hoc mapping remained unchanged: six GOLD_EQUIVALENT outputs, three OVER_SPLIT_CLAIM outputs, and one FALSE_SPAN output, organized into six proposition families plus the false-span control.
  • All five source artifacts were checked against their frozen SHA-256 values before the qualification began. No claim, source text, mapping, expected result, or semantic field was regenerated or edited.

Canonical validation and identity behavior

Completed
  • All six control claims and all ten hosted claims retained valid deterministic checks, independent semantic validation, and source provenance from their original immutable evidence.
  • The authoritative downstream identity is claim_id. Feasibility iterates every validated claim independently, and current-capability qualification consumes those claim-keyed assessments up to its configured maximum.
  • No semantic proposition-family deduplication or lineage collapse occurs before feasibility or capability eligibility. Distinct claim IDs retain common source provenance, so over-splitting does not manufacture independent sources, but it does manufacture independent downstream work units.

Control downstream funnel

Completed
  • Six control claims produced six valid feasibility assessments, four valid LOCAL_READY experiment plans, and two NOT_HIRO_RELEVANT outcomes.
  • The four relevant control claims reached current-capability assessment. All four were assessed as CAPABILITY_PRESENT_WITH_GAP with an additive transfer surface, then terminated as GAP_UNPROVEN under the independent evidence audit.
  • No control claim became viable for reproduction, and no reproduction or later authority was exercised.

Hosted downstream funnel

Completed
  • Ten hosted claims produced ten valid feasibility assessments, three valid LOCAL_READY experiment plans, and seven NOT_HIRO_RELEVANT outcomes.
  • The three relevant hosted claims reached current-capability assessment. All three were assessed as CAPABILITY_PRESENT_WITH_GAP with an additive transfer surface, then terminated as GAP_UNPROVEN under the independent evidence audit.
  • Compared with control, the hosted representation lost two proposition families at the Hiro-relevance and experiment-feasibility decision, while one resource-awareness family expanded into two downstream assessments.

Per-family operational result

Two material divergences
  • The escalation cost-quality family was a benign duplication: both hosted pieces and the gold control terminated as not Hiro-relevant.
  • The escalation-interface, resource-planning failure, and false-span cases were decision-equivalent. The false span was safely filtered before capability assessment.
  • The resource-information family was a benign duplication: both hosted pieces and the gold control passed feasibility, reached the same capability state and transfer surface, and ended GAP_UNPROVEN. The two hosted pieces still consumed twice the downstream work.
  • The downshift-context family materially diverged. Its gold claim produced a transferable LOCAL_READY hypothesis, but its one-to-one hosted representation was rejected as model-specific and not Hiro-relevant.
  • The structured multi-agent code-generation family materially diverged. Its combined gold claim produced a transferable LOCAL_READY structured-handoff hypothesis, while both benchmark-specific hosted splits were rejected as system-specific and not Hiro-relevant.

Amplification and segmentation decision

Measured and bounded
  • The hosted representation amplified six claim units to ten, a 1.666667 ratio. It caused four additional feasibility planner calls, four additional independent feasibility validations, and exposed four additional potential claim-keyed slots at later probe, reproduction, and candidate-lineage boundaries.
  • No false independent corroboration was created because all hosted fragments retained their original source-record and source-content identities.
  • Under the qualification's explicit decision rule, exact atomic segmentation is operationally necessary because fragmentation produced material false negatives that existing downstream controls did not recover.
  • Necessary atomicity does not mean uniformly shorter spans. The required unit must preserve the complete transferable mechanism and material conditions: one failed family was a one-to-one GOLD_EQUIVALENT mapping with lost context, while the structured-handoff control succeeded by retaining its related benchmark observations as one complete mechanism-level claim rather than isolated numeric fragments.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen-input integrity passed The historical corpus, historical claim report, six gold controls, original hosted report, and accepted ten-to-six diagnostic mapping all matched their frozen hashes.
Real downstream replay passed operationally Both representations ran through the same current feasibility planner, plan validator, capability assessor, and capability auditor with Qwen 3.8 at temperature zero and one shared telemetry cutoff.
Qualification guard tests passed Twenty-two focused representation-equivalence, feasibility, and viability tests passed. The comparator distinguishes gate-visible divergence from harmless relevance-label vocabulary differences and classifies lost advancing families as material false negatives.
Immutable result evidence passed The corrected v3 report is read-only and has SHA-256 5b27ad61549399e94f0bf3bc27cd799c95a7f252adf99509292d9a4ef3f8143d. It reuses the immutable downstream artifacts and records that no model calls were repeated during analysis correction.
Production health and authority passed Hiro remained healthy with connected Qwen and matching checkout and loaded revision 19ff77a5b09cf3b76712a4b5eef3d6617089fdd9. No production code, routing, state, candidate, or promotion was changed.

Current state

Next steps