Hiro development journal

Giving Hiro's candidate builder the full system view

Published after successful promotion and post-promotion qualification Machine-readable JSON

Executive summary

This session traced repeated non-promotions to multiple platform defects rather than weak ideas. Earlier candidates could satisfy an evaluator-injected task type without being reachable from the production agent path, and a platform regression test incorrectly required that repaired production behavior remain defective. Both failures were voided append-only and corrected without weakening functional, prompt-injection, or unauthorized-execution checks.

The final construction failure exposed the central model-context problem. Qwen 3.8 27B was running with a 16,384-token context window, but the candidate builder supplied only 2,400 characters of repository source and 1,600 characters of symbol context. Qwen consequently guessed at an existing classifier and a permissive fuzzy editor replaced nearby module lines, deleting unrelated public exports.

Framework revision 67c8794 increased repository context to 16,000 characters, added 8,000 characters of exact symbol context, prioritized explicitly named entrypoints, preserved existing classifier branches in guidance, and made complete named-symbol edits incapable of overwriting neighboring exports. The full repository passed 714 tests and five exact-revision pipeline qualification cycles.

On the qualified framework, Qwen used a 13,272-token candidate prompt and produced a narrow production-reachable patch on its first attempt. Construction, paired public and held-out evaluation, 117-test security comparison, isolated Stage 5 integration, and all four live canary checkpoints passed. The governor then passed 715 tests and promoted candidate e3353a1. A final manifest-plumbing repair was added on top, bringing the qualified live head to af62760.

Work completed

Make the originating contract executable end to end

Completed
  • Raw isolated model inference was separated from boundary-validated evaluation so interaction audits apply the universal response gate exactly once with the full task contract.
  • CandidateBuildRequest now carries the code-owned replay fixture. Stage 3 independently replays the complete contract, including grounding classification, required phrases, forbidden phrases, and production reachability without an evaluator-only task-type keyword.
  • Captured and fresh canaries each exercise both contract-aware and production-reachable paths, plus independent platform tests.

Require production-reachable evidence

Completed
  • Candidate sandbox-20260822184609-ce0986ac-01 passed contract-aware checks but depended on an explicit task_type that core.agent.run_turn never supplies. It was voided with a platform_defect_production_reachability_voided event.
  • Framework revision 6327eaf added captured and fresh production-path replays to construction and canary evidence.
  • A subsequent candidate passed all four behavioral paths, but an independent test asserted that production replay must remain broken. That platform-test rejection was voided, the test was rewritten to inject a defective classifier explicitly, and revision 958efa4 passed 712 tests plus five qualification cycles.

Use Qwen's 16K context window structurally

Completed
  • The live model was correctly Qwen 3.8 27B with --ctx-size 16384, but candidate source context was capped at 2,400 characters and frequently omitted final_gate and its requested insertion anchor.
  • The builder now supplies up to 16,000 characters of balanced repository source and 8,000 characters of exact symbol excerpts. Task-specific guidance and contract identifiers participate in symbol ranking, so explicitly named entrypoints outrank large generic classes.
  • A complete function or class supplied as a named modify edit is now applied only to that AST symbol. Stale fuzzy text can no longer replace adjacent public definitions such as begin_response_trace.
  • Prompt-injection guidance explicitly requires preserving every existing inference branch, including arithmetic_reasoning, and forbids inventing unrelated weather, news, or general classifications.

Qualify the corrected controller

Completed
  • Forty-six focused candidate-builder and editor tests passed, including new regressions for explicit-symbol prioritization and adjacent-export preservation.
  • The full repository passed 714 tests in 403.88 seconds with one existing unregistered custom-mark warning.
  • Five of five frozen pipeline cycles passed construction, contemporaneous paired evaluation, isolated integration, canary and promotion simulation, restart idempotency, and rollback. Qualification packet qualification-20260822T204403Z.json is pinned to revision 67c87941521a66ab28fd307edea3897e7986ac94.

Build and evaluate the production candidate

Completed and promoted
  • Candidate sandbox-20260822204455-45a4f37b-01 used 13,272 prompt tokens and completed on Qwen's first attempt. It preserved the arithmetic classifier, added exact untrusted-snapshot routing, inserted a logged prompt-injection-resistance pass-through at Rule 3b, and added a marked regression test.
  • Stage 3 passed syntax, production guards, targeted candidate execution, assertion-based untouched-baseline contrast, captured contract replay, and captured production replay. Frozen packet content hash begins d92686b.
  • Across three alternating baseline/candidate pairs, public weighted score was 0.9291667 on both variants and held-out weighted score was 0.9041667 on both variants. Invariant failures and category regressions were zero.
  • Public candidate p95 latency was 12,571.94 ms versus 13,098.98 ms baseline. Held-out candidate p95 was 19,066.75 ms versus 18,975.01 ms baseline, within the 1.10 ratio limit and not a repeated cross-suite regression.
  • Baseline and candidate each passed all 117 configured security tests. Isolated Stage 5 revision e3353a1 passed smoke, canary, and three repeated 118-test monitoring runs.
  • Live canary checkpoints 0, 5, 15, and 60 passed captured and fresh contract-aware replays, captured and fresh production replays, and independent platform tests.
  • The first governor call exposed a platform manifest defect: the queue carried the four-file allowlist as changed_files even though Stage 5 had verified a two-file commit. The rejection was voided append-only, checkpoint 60 was rerun with the authoritative Stage 5 manifest, and the governor passed 715 tests before promoting revision e3353a1.

Repair promotion-manifest plumbing

Completed
  • The continuous engine now reads changed_files from the frozen Stage 5 candidate verification receipt. It no longer substitutes the broader candidate allowlist.
  • The positive full-path acceptance test now deliberately uses a two-file allowlist and a one-file verified commit, then asserts that the governor receives only the verified file.
  • All 44 continuous-engine tests passed. The complete post-promotion repository passed 715 tests in 408.68 seconds.
  • Five of five frozen pipeline qualification cycles passed on final revision af627600bfa3d497d6d9b27ab12426057cda093f. Packet qualification-20260822T223350Z.json has SHA-256 2dd795416a7c3f290a25b0f9323354baa990155dc5a7e9f80ee84ac8f084f438.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Framework repository suite passed 714 tests passed in 403.88 seconds on revision 67c8794; zero failures; one known unregistered hiro_contract mark warning.
Exact-revision pipeline qualification passed Five of five end-to-end cycles passed. Packet qualification-20260822T204403Z.json; SHA-256 9ff7216373a191bab25fa1f9d6b9abfad3ac5560365a92bda44a156b364d2dfe.
Candidate construction passed First attempt passed targeted execution, untouched-baseline assertion contrast, production guards, full contract replay, and production-path replay without an injected task type.
Stage 4 paired evaluation passed Three alternating pairs produced score parity on public and held-out suites, zero invariant and category regression, acceptable latency, and 117 of 117 security tests on both variants.
Stage 5 isolated integration passed Candidate revision e3353a123072b22689ddf780852c890ff44c2d8e passed smoke, 117-test canary, and three 118-test monitoring repetitions.
Live moderate-risk canary passed Checkpoints 0, 5, 15, and 60 passed all five evidence paths. The corrected governor run then passed 715 tests in 405.09 seconds and promoted candidate e3353a1.
Post-promotion manifest repair passed The full repository passed 715 tests in 408.68 seconds, followed by five of five frozen pipeline cycles on final head af62760.

Current state

Next steps