Hiro development journal

Finite diversity campaign for Hiro Step 3

Published Machine-readable JSON

Executive summary

Designed a finite execution-grade roadmap Step 3 campaign that a lighter supervising model can initiate, monitor, resume, and close through deterministic controller commands and a fixed evidence decision table.

The plan directly addresses the July 21 repetition failure. That session recorded 44,592 observations but only 2,820 structural signatures, a 93.68% structural-reuse rate. Seven of eight families had only four to eight effective topologies.

The replacement campaign uses ten capability families, eight finite phases, a maximum of 360 planned cases including opening and closing canonical checks, global historical novelty rejection, family and topology caps, mastery retirement, sealed blind validation, and mandatory worker closeout.

Work completed

Repetition diagnosis

Completed
  • Inspected the existing development-cycle runbook, v1 runner, generated catalogs, adaptive Daylab rotation, and stateful adversarial generator.
  • Confirmed that the v1 runner extends beyond its nominal cycle count until the timebox and that signatures are checked only within each generated batch rather than against historical catalogs.
  • Recomputed the continuous-session corpus: 929 cycles, 44,592 observations, 2,820 unique signatures, and 93.68% structural reuse.

Finite capability portfolio

Designed
  • Defined ten families spanning temporal calendar reasoning, task replanning, email grounding, cross-tool coordination, approval lifecycle, concurrency, recovery, memory continuity, source judgment, and productive autonomy.
  • Defined a finite sequence from calibration through breadth, evolving state, cross-family composition, adaptive micro-waves, blind frontier, and closeout.
  • Set the campaign target at 320-420 valid observations, with a manifest maximum of 348 noncanonical cases plus twelve opening and closing canonical workflow checks.

Enforceable novelty and mastery

Designed
  • Structural signatures exclude superficial names, prose, dates, and labels and instead encode constraint graphs, event ordering, tool classes, approval transitions, resource versions, fault sequences, temporal relations, oracle class, and family composition.
  • The plan requires exact and near-duplicate rejection against repository suites and historical development and adversarial catalogs, at least 85% noncanary novelty, a fifteen-percent family cap, and a two-percent topology cap.
  • Mastered family-band combinations retire after two fresh statistically supported waves and return only as small rotating canaries.

Weaker-model supervision contract

Designed
  • The supervisor follows machine-readable checkpoints and allowed-next-actions instead of reconstructing state from conversation history.
  • A fixed table determines whether to increase difficulty, hold a band, create a bounded diagnostic wave, permit one evidence-based proposal, reject a batch, or close the campaign.
  • Terminal handling covers verified duplicate workers, wrapper timeouts, endpoint failures, incomplete reports, exhausted generators, failed candidates, blind leakage, and hard-stop closeout.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Historical structural-reuse audit completed The July 21 continuous session contained 44,592 cataloged cases and 2,820 unique structural signatures, yielding 93.68% reuse. Seven families had only four to eight signatures; long-context attention supplied 2,778.
Plan structure validation passed The runbook contains all ten required families, all eight phases, the finite manifest, novelty/family/topology limits, mastery retirement, supervisor decision table, closeout contract, evidence packet, and lighter-model handoff prompt. Planned maximum is 360 cases including canonical open and close checks.
Step 3 controller availability not-ready The specified hiro.benchmarks.development_cycle.step3 interface does not yet exist. The runbook therefore makes controller implementation and validation a bounded prerequisite and prohibits fallback to v1.
Capability evaluation not-run-by-design No Step 3 evaluation campaign, model run, candidate change, production connector, service, or schedule was started during this planning session.

Current state

Next steps