{
  "schemaVersion": 2,
  "date": "2026.08.22",
  "publishedAt": "2026-08-22T12:11:07-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Giving Hiro's candidate builder the full system view",
  "publicationStatus": "Published after successful promotion and post-promotion qualification",
  "executiveSummary": [
    "This session traced repeated non-promotions to multiple platform defects rather than weak ideas. Earlier candidates could satisfy an evaluator-injected task type without being reachable from the production agent path, and a platform regression test incorrectly required that repaired production behavior remain defective. Both failures were voided append-only and corrected without weakening functional, prompt-injection, or unauthorized-execution checks.",
    "The final construction failure exposed the central model-context problem. Qwen 3.8 27B was running with a 16,384-token context window, but the candidate builder supplied only 2,400 characters of repository source and 1,600 characters of symbol context. Qwen consequently guessed at an existing classifier and a permissive fuzzy editor replaced nearby module lines, deleting unrelated public exports.",
    "Framework revision 67c8794 increased repository context to 16,000 characters, added 8,000 characters of exact symbol context, prioritized explicitly named entrypoints, preserved existing classifier branches in guidance, and made complete named-symbol edits incapable of overwriting neighboring exports. The full repository passed 714 tests and five exact-revision pipeline qualification cycles.",
    "On the qualified framework, Qwen used a 13,272-token candidate prompt and produced a narrow production-reachable patch on its first attempt. Construction, paired public and held-out evaluation, 117-test security comparison, isolated Stage 5 integration, and all four live canary checkpoints passed. The governor then passed 715 tests and promoted candidate e3353a1. A final manifest-plumbing repair was added on top, bringing the qualified live head to af62760."
  ],
  "workstreams": [
    {
      "title": "Make the originating contract executable end to end",
      "status": "Completed",
      "details": [
        "Raw isolated model inference was separated from boundary-validated evaluation so interaction audits apply the universal response gate exactly once with the full task contract.",
        "CandidateBuildRequest now carries the code-owned replay fixture. Stage 3 independently replays the complete contract, including grounding classification, required phrases, forbidden phrases, and production reachability without an evaluator-only task-type keyword.",
        "Captured and fresh canaries each exercise both contract-aware and production-reachable paths, plus independent platform tests."
      ]
    },
    {
      "title": "Require production-reachable evidence",
      "status": "Completed",
      "details": [
        "Candidate sandbox-20260822184609-ce0986ac-01 passed contract-aware checks but depended on an explicit task_type that core.agent.run_turn never supplies. It was voided with a platform_defect_production_reachability_voided event.",
        "Framework revision 6327eaf added captured and fresh production-path replays to construction and canary evidence.",
        "A subsequent candidate passed all four behavioral paths, but an independent test asserted that production replay must remain broken. That platform-test rejection was voided, the test was rewritten to inject a defective classifier explicitly, and revision 958efa4 passed 712 tests plus five qualification cycles."
      ]
    },
    {
      "title": "Use Qwen's 16K context window structurally",
      "status": "Completed",
      "details": [
        "The live model was correctly Qwen 3.8 27B with --ctx-size 16384, but candidate source context was capped at 2,400 characters and frequently omitted final_gate and its requested insertion anchor.",
        "The builder now supplies up to 16,000 characters of balanced repository source and 8,000 characters of exact symbol excerpts. Task-specific guidance and contract identifiers participate in symbol ranking, so explicitly named entrypoints outrank large generic classes.",
        "A complete function or class supplied as a named modify edit is now applied only to that AST symbol. Stale fuzzy text can no longer replace adjacent public definitions such as begin_response_trace.",
        "Prompt-injection guidance explicitly requires preserving every existing inference branch, including arithmetic_reasoning, and forbids inventing unrelated weather, news, or general classifications."
      ]
    },
    {
      "title": "Qualify the corrected controller",
      "status": "Completed",
      "details": [
        "Forty-six focused candidate-builder and editor tests passed, including new regressions for explicit-symbol prioritization and adjacent-export preservation.",
        "The full repository passed 714 tests in 403.88 seconds with one existing unregistered custom-mark warning.",
        "Five of five frozen pipeline cycles passed construction, contemporaneous paired evaluation, isolated integration, canary and promotion simulation, restart idempotency, and rollback. Qualification packet qualification-20260822T204403Z.json is pinned to revision 67c87941521a66ab28fd307edea3897e7986ac94."
      ]
    },
    {
      "title": "Build and evaluate the production candidate",
      "status": "Completed and promoted",
      "details": [
        "Candidate sandbox-20260822204455-45a4f37b-01 used 13,272 prompt tokens and completed on Qwen's first attempt. It preserved the arithmetic classifier, added exact untrusted-snapshot routing, inserted a logged prompt-injection-resistance pass-through at Rule 3b, and added a marked regression test.",
        "Stage 3 passed syntax, production guards, targeted candidate execution, assertion-based untouched-baseline contrast, captured contract replay, and captured production replay. Frozen packet content hash begins d92686b.",
        "Across three alternating baseline/candidate pairs, public weighted score was 0.9291667 on both variants and held-out weighted score was 0.9041667 on both variants. Invariant failures and category regressions were zero.",
        "Public candidate p95 latency was 12,571.94 ms versus 13,098.98 ms baseline. Held-out candidate p95 was 19,066.75 ms versus 18,975.01 ms baseline, within the 1.10 ratio limit and not a repeated cross-suite regression.",
        "Baseline and candidate each passed all 117 configured security tests. Isolated Stage 5 revision e3353a1 passed smoke, canary, and three repeated 118-test monitoring runs.",
        "Live canary checkpoints 0, 5, 15, and 60 passed captured and fresh contract-aware replays, captured and fresh production replays, and independent platform tests.",
        "The first governor call exposed a platform manifest defect: the queue carried the four-file allowlist as changed_files even though Stage 5 had verified a two-file commit. The rejection was voided append-only, checkpoint 60 was rerun with the authoritative Stage 5 manifest, and the governor passed 715 tests before promoting revision e3353a1."
      ]
    },
    {
      "title": "Repair promotion-manifest plumbing",
      "status": "Completed",
      "details": [
        "The continuous engine now reads changed_files from the frozen Stage 5 candidate verification receipt. It no longer substitutes the broader candidate allowlist.",
        "The positive full-path acceptance test now deliberately uses a two-file allowlist and a one-file verified commit, then asserts that the governor receives only the verified file.",
        "All 44 continuous-engine tests passed. The complete post-promotion repository passed 715 tests in 408.68 seconds.",
        "Five of five frozen pipeline qualification cycles passed on final revision af627600bfa3d497d6d9b27ab12426057cda093f. Packet qualification-20260822T223350Z.json has SHA-256 2dd795416a7c3f290a25b0f9323354baa990155dc5a7e9f80ee84ac8f084f438."
      ]
    }
  ],
  "decisions": [
    "Treat complete code-owned contracts and production reachability as harness-owned approval evidence rather than trusting candidate-authored tests alone.",
    "Void demonstrated platform-caused rejections append-only while preserving the original idea, evidence, and bounded candidate lifecycle.",
    "Use the model's configured context window to expose exact production code; context size is a system property, not merely a server launch flag.",
    "Scope complete named Python edits by AST symbol so permissive text recovery cannot corrupt neighboring module interfaces.",
    "Keep approval focused on functional improvement, global non-regression, prompt injection, unauthorized execution, and hacking boundaries rather than broad text-edit prohibitions.",
    "Require exact-revision full-suite and five-cycle qualification before restarting the live queue after framework changes."
  ],
  "validation": [
    {
      "check": "Framework repository suite",
      "status": "passed",
      "result": "714 tests passed in 403.88 seconds on revision 67c8794; zero failures; one known unregistered hiro_contract mark warning."
    },
    {
      "check": "Exact-revision pipeline qualification",
      "status": "passed",
      "result": "Five of five end-to-end cycles passed. Packet qualification-20260822T204403Z.json; SHA-256 9ff7216373a191bab25fa1f9d6b9abfad3ac5560365a92bda44a156b364d2dfe."
    },
    {
      "check": "Candidate construction",
      "status": "passed",
      "result": "First attempt passed targeted execution, untouched-baseline assertion contrast, production guards, full contract replay, and production-path replay without an injected task type."
    },
    {
      "check": "Stage 4 paired evaluation",
      "status": "passed",
      "result": "Three alternating pairs produced score parity on public and held-out suites, zero invariant and category regression, acceptable latency, and 117 of 117 security tests on both variants."
    },
    {
      "check": "Stage 5 isolated integration",
      "status": "passed",
      "result": "Candidate revision e3353a123072b22689ddf780852c890ff44c2d8e passed smoke, 117-test canary, and three 118-test monitoring repetitions."
    },
    {
      "check": "Live moderate-risk canary",
      "status": "passed",
      "result": "Checkpoints 0, 5, 15, and 60 passed all five evidence paths. The corrected governor run then passed 715 tests in 405.09 seconds and promoted candidate e3353a1."
    },
    {
      "check": "Post-promotion manifest repair",
      "status": "passed",
      "result": "The full repository passed 715 tests in 408.68 seconds, followed by five of five frozen pipeline cycles on final head af62760."
    }
  ],
  "currentState": [
    "Hiro is serving from qualified final revision af627600bfa3d497d6d9b27ab12426057cda093f, which contains promoted candidate e3353a1 and the manifest-plumbing repair.",
    "Qwen 3.8 27B is resident on port 8080 with a 16,384-token context window and serves both fast and standard local tiers.",
    "The prompt-injection-resistance candidate is implemented; the queue's implemented count is now five.",
    "The active production-path replay infers prompt_injection_resistance from the complete snapshot markers and passes the originating contract with zero failures."
  ],
  "limitations": [
    "The candidate is intentionally narrow to complete UNTRUSTED_WEB_SNAPSHOT marker pairs and does not claim general prompt-injection detection.",
    "One candidate test uses an unregistered hiro_contract pytest mark, producing a warning without changing pass/fail evidence.",
    "The strict JSON request advertised a 3,500-token maximum completion after a 13,272-token prompt; the actual completion was 755 tokens and fit the 16K runtime, but future budget estimation should use tokenizer telemetry rather than byte heuristics."
  ],
  "nextSteps": [
    "Use the verified Stage 5 changed-file manifest for every future governor request and retain the broader allowlist only as a scope boundary.",
    "Use tokenizer telemetry rather than byte heuristics for future candidate completion-budget calibration.",
    "Continue the ranked queue from the qualified promoted head and apply the same production-reachability evidence to subsequent interaction incidents.",
    "Publish this finalized entry after journal tests and the production build pass."
  ]
}
