{
  "schemaVersion": 2,
  "date": "2026.08.21",
  "publishedAt": "2026-08-21T17:45:55-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Paired evaluation, pipeline qualification, and live builder recovery",
  "publicationStatus": "Implementation complete, qualified, and exercised through controlled integration and independent canary",
  "executiveSummary": [
    "Hiro's Stage 4 evaluator now measures every candidate against a fresh checkout of its exact base commit in three alternating baseline/candidate pairs for both public and held-out suites. Historical pinned runs remain identity and suite-provenance evidence but can no longer supply causal latency measurements.",
    "A fail-closed pipeline qualification gate now binds live queue advancement to a checksum-verified packet for the exact active Git revision. Five repeated cycles exercise candidate construction, contemporaneous paired evaluation, isolated integration, governor canary/promotion, interruption-resume idempotency, and failed-canary rollback.",
    "Revision 2cb578d passed all 683 repository tests and five qualification cycles, each with five passing checks. Hiro restarted, exposed the qualified state on the benchmark page, recovered stale queue work, and selected a new item based on the qualified revision.",
    "The first real live candidate then exposed a gap the deterministic qualification fixture had not covered: Qwen returned malformed or empty structured plans and later exceeded its pinned 8K context, producing HTTP 400. The service was stopped and the exact preserved arithmetic failure was replayed repeatedly through the real CandidateBuilder while the builder was repaired.",
    "The revised builder uses context-aware completion budgeting, split production/test recovery for malformed or empty plans, a symbol-bounded additive insertion operator, task-contract-derived guidance, and harness-owned pytest contract markers. The same preserved failure that previously exhausted five attempts now reached candidate_ready on its first attempt with improvement_demonstrated baseline contrast.",
    "The final builder revision dfda200 passes all 687 repository tests and all five exact-revision qualification cycles. After restart, the preserved live item produced a frozen candidate, passed every paired Stage 4 gate, entered controlled integration, and reached the independent canary. The canary exposed a remaining functional miss, so Hiro scheduled a bounded candidate revision instead of promoting or terminally rejecting the idea, then continued with the queue."
  ],
  "workstreams": [
    {
      "title": "Contemporaneous paired Stage 4 evaluation",
      "status": "Implemented and regression-tested",
      "details": [
        "Stage 4 verifies the pinned Step 2 baseline reports for immutable variant identity, suite version, and coverage, but does not reuse their timing for a candidate decision.",
        "For each candidate, the evaluator creates or verifies a detached clean worktree at the candidate's exact base commit and starts both baseline and candidate through the same child-process entrypoint and runtime boundary.",
        "Public and held-out suites each run in three pairs with alternating order: baseline/candidate, candidate/baseline, baseline/candidate.",
        "The evaluator aggregates all three full-suite reports per role and suite, persists the aggregates in the append-only ledger, and sends those fresh aggregates to the existing functional, invariant, category, statistical, and latency gates.",
        "Frozen recommendation evidence records the pair order, all source run IDs, the base commit, aggregate run IDs, and an explicit historical_latency_used false field.",
        "Report override fields remain backward-compatible at the request boundary but cannot bypass paired measurement; a regression test verifies that supplied old/fresh override IDs are ignored for the decision."
      ]
    },
    {
      "title": "Revision-bound pipeline qualification circuit",
      "status": "Implemented; final revision qualified",
      "details": [
        "Added hiro.improvement.pipeline_qualification with a frozen checksum packet, exact source revision identity, five required cycles, and fail-closed status verification.",
        "Each cycle runs five tests spanning a real CandidateBuilder/paired-CandidateEvaluator path, standing-authority isolated integration, failed-canary rollback, interruption/resume idempotency, and stable-governor canary promotion in a disposable repository.",
        "The active improvement scheduler permits discovery and queue ingestion while unqualified but passes advance_allowed false to construction, evaluation, canary, and promotion transitions.",
        "The benchmark API and Improvement Control Center display qualification status, cycles passed, required cycles, revision identity, and the reason a queue is blocked.",
        "Revision 2cb578d passed five of five cycles; each cycle reported five passed checks in approximately 82 to 84 seconds. The packet SHA-256 was 0819cef6f7b4a352867881162635f83a981c6271de5b9fa293c102ec01370e86.",
        "After the live builder repair, revision dfda200 passed a new five-cycle campaign in 78.06 to 79.39 seconds per cycle. Its packet SHA-256 is 0aa873f98d11c3753903dd99f9e4f4cf86d359d07410b53f45246aa96bea96ed."
      ]
    },
    {
      "title": "Live restart and queue exercise",
      "status": "Completed through first real construction failure",
      "details": [
        "Hiro restarted through scripts/start_hiro.py using the checked-in Windows Path/PATH normalization and pinned runtime. Startup completed in 11.172 seconds.",
        "Ports 8000, 8001, and 8765 responded; the tasks and benchmark pages returned HTTP 200; Qwen 3.8 remained on port 8080.",
        "The continuous API reported qualification qualified, five cycles passed, one active item, and 199 actionable items.",
        "The scheduler recovered stale pre-change candidate state, requeued an artifact after the builder revision changed, selected a fresh candidate, and pinned it to the then-qualified revision 2cb578d.",
        "The server established live inference connections to Qwen and created an external candidate worktree. The frozen candidate packet then recorded five construction failures: two non-JSON responses followed by three HTTP 400 responses.",
        "Hiro correctly classified this as repairable construction failure, scheduled a bounded retry, and continued ranking work. The service was deliberately stopped before source repair so dirty code could not contaminate automatic candidates."
      ]
    },
    {
      "title": "Local candidate construction recovery",
      "status": "Implemented and validated against the preserved live failure",
      "details": [
        "Measured the exact failed prompt at approximately 13.4K characters. The old fixed 4,096-token completion request could not reliably coexist with the prompt and strict schema inside Qwen's pinned 8,192-token context, especially after repair history was added.",
        "Added context-aware completion budgeting and reduced duplicated full-file/symbol excerpts while retaining source, signatures, validation history, security constraints, and the task contract.",
        "Raw diagnostic output showed Qwen degenerating into repeated text inside one large escaped new_content JSON string. Malformed, HTTP-400, and valid-but-empty full plans now fall back to two bounded structured calls: production patch first, targeted test second.",
        "Added contract-derived arithmetic guidance that requires operands to come from user_text when the captured malformed output omits them. This prevents a candidate from searching the known-bad response for information it cannot contain.",
        "Added an insert_before change type that operates inside one named Python symbol, resolves a close anchor only within that symbol, inserts additively, and preserves the anchor and surrounding control flow. It replaces destructive approximate replacement for bounded branch insertion.",
        "Generated targeted tests are mechanically annotated with pytest.mark.hiro_contract when the model omits the marker. The harness changes no assertion or expected outcome; it only records which failure is the declared candidate contract.",
        "The preserved arithmetic failure progressed from five transport/parse failures, through syntax and locator diagnostics, to candidate_ready on the first attempt. Its frozen packet records both changed files, passing candidate tests, a failing marked assertion on the untouched baseline, and improvement_demonstrated."
      ]
    },
    {
      "title": "Final qualified restart and real end-to-end cycle",
      "status": "Completed through independent canary decision",
      "details": [
        "Hiro restarted on revision dfda200 in 12.422 seconds through the normalized pinned launcher. The tasks and benchmark endpoints returned HTTP 200, and the API reported qualified with five of five cycles for the exact current revision.",
        "The scheduler detected that the active arithmetic item had been built against obsolete revision 2cb578d, emitted stale_revision_requeued, and repinned the work to dfda200 before allowing mutation.",
        "The revised live builder produced frozen candidate sandbox-20260822013004-40bccb7a-01 with status candidate_ready. The packet contains changes only to core/response_envelope.py and its generated targeted contract test.",
        "The queue then entered the real CandidateEvaluator. It completed 96 baseline and 96 candidate cases across three alternating public/held-out pairs; no diagnostic packet was substituted for this live candidate.",
        "Stage 4 marked the candidate eligible for controlled integration. Public weighted score was 0.9291667 for both roles, held-out weighted score was 0.9041667 for both, invariant failures were zero, category regressions were empty, and candidate p95 latency remained inside the 1.10 ratio policy.",
        "The targeted candidate contract passed while the untouched baseline failed the single marked assertion. The same 117 security and operational-boundary tests passed on both baseline and candidate.",
        "After controlled integration, the fresh independent canary returned only $408 and still omitted the requested expression. Hiro emitted canary_revision_scheduled with that redacted failure evidence and a five-minute retry time rather than promoting or terminally rejecting the idea.",
        "The scheduler immediately continued to investigate the next ranked queue item while the arithmetic revision waited, proving that one canary miss does not block queue throughput."
      ]
    }
  ],
  "decisions": [
    "Make fresh contemporaneous paired evidence mandatory inside CandidateEvaluator rather than maintaining it as a separate agenda-replication utility.",
    "Use historical pinned baselines only for immutable code/suite identity; do not let baseline age or runtime drift determine causal latency rejection.",
    "Fail closed on source-revision changes: any commit after qualification requires a new checksum-verified five-cycle campaign before live queue advancement.",
    "Keep discovery and queue ingestion operating when unqualified so evidence is not lost, while preventing construction, evaluation, canary, or promotion from advancing.",
    "Treat the live Qwen construction failure as a qualification-coverage gap, stop Hiro, and repair the exact preserved request rather than dismissing it as a bad idea.",
    "Retain all functional, regression, prompt-injection/security, baseline-contrast, and held-out gates. The builder recovery changes candidate generation and safe application semantics, not approval thresholds.",
    "Use harness-owned contract markers because attribution metadata is evaluator infrastructure; the candidate model remains responsible for the actual assertion and implementation behavior.",
    "Prefer additive, symbol-bounded insertion over destructive fuzzy replacement when a candidate adds one local branch to a large shared function."
  ],
  "validation": [
    {
      "check": "Affected evaluator, integrator, governor, active-loop, API, and benchmark tests",
      "status": "passed",
      "result": "72 tests passed in 270.82 seconds after paired evaluation and the qualification circuit were implemented."
    },
    {
      "check": "Complete suite for revision 2cb578d",
      "status": "passed",
      "result": "683 tests passed in 377.92 seconds."
    },
    {
      "check": "Five-cycle qualification for revision 2cb578d",
      "status": "passed",
      "result": "Five cycles passed; each cycle ran five checks covering construction/evaluation, integration, rollback, restart idempotency, and governor promotion."
    },
    {
      "check": "Normalized Windows child startup",
      "status": "passed",
      "result": "Hiro started through the checked-in detached launcher in 11.172 seconds; tasks and benchmarks returned HTTP 200 and all expected ports were owned by the Hiro process."
    },
    {
      "check": "Live qualified queue advancement",
      "status": "passed through construction entry",
      "result": "The scheduler recovered stale state, selected work pinned to 2cb578d, created an external worktree, and called Qwen. The attempt then exposed the builder transport/format gap described in this entry."
    },
    {
      "check": "Builder focused tests after recovery changes",
      "status": "passed",
      "result": "Thirty CandidateBuilder tests passed, including split recovery, dynamic budgeting, additive insertion preservation, and mechanical contract annotation."
    },
    {
      "check": "Preserved live arithmetic candidate replay",
      "status": "passed",
      "result": "The exact previously failing request reached candidate_ready on attempt one with improvement_demonstrated baseline contrast."
    },
    {
      "check": "Complete suite for revision dfda200",
      "status": "passed",
      "result": "687 tests passed in 371.87 seconds."
    },
    {
      "check": "Five-cycle qualification for revision dfda200",
      "status": "passed",
      "result": "Five of five cycles passed. Cycle durations were 79.07, 78.77, 78.20, 78.06, and 79.39 seconds; every cycle passed all five required checks."
    },
    {
      "check": "Final normalized restart and exact-revision gate",
      "status": "passed",
      "result": "Hiro started in 12.422 seconds, tasks and benchmarks returned HTTP 200, and the live API verified qualified revision dfda200 with packet SHA-256 0aa873f98d11c3753903dd99f9e4f4cf86d359d07410b53f45246aa96bea96ed."
    },
    {
      "check": "Real live candidate construction after restart",
      "status": "passed",
      "result": "The preserved active queue item was invalidated as stale, rebuilt against dfda200, and frozen as candidate_ready in sandbox-20260822013004-40bccb7a-01 before entering paired Stage 4."
    },
    {
      "check": "Real live mandatory paired Stage 4",
      "status": "passed",
      "result": "All 192 case executions completed. The frozen recommendation was eligible_for_controlled_integration with equal public and held-out scores, no invariant or category regression, acceptable paired latency, passing targeted regression, and 117/117 security tests on both roles."
    },
    {
      "check": "Controlled integration and fresh independent canary",
      "status": "revision scheduled",
      "result": "Controlled integration completed, but the fresh canary still omitted the arithmetic expression. Hiro preserved the example, scheduled a bounded revision for 18:55 Pacific, did not promote, and continued processing the queue."
    }
  ],
  "currentState": [
    "Paired Stage 4 evaluation and the revision-bound qualification circuit are committed.",
    "The initial implementation commit is 2cb578d and its qualification packet passed five of five cycles.",
    "The builder recovery commit is dfda200; all 687 repository tests and all five qualification cycles pass.",
    "The preserved real failure now produces a first-attempt candidate_ready packet with demonstrated baseline contrast.",
    "Hiro and Qwen 3.8 are running. The live API reports dfda200 qualified, and the repaired queue continues processing candidates.",
    "Live candidate sandbox-20260822013004-40bccb7a-01 passed mandatory paired Stage 4 and reached controlled integration; its independent canary scheduled a revision because the fresh response remained incomplete.",
    "The arithmetic idea remains retryable rather than rejected. Another ranked interaction-audit item entered investigation while its retry timer ran.",
    "No diagnostic candidate was integrated, promoted, deployed, or merged into Hiro."
  ],
  "limitations": [
    "The deterministic qualification fixture did not initially exercise the live local-model structured generation path; that coverage gap was discovered only after the first qualified restart.",
    "The preserved replay establishes that one formerly blocked arithmetic candidate now reaches Stage 3. It does not establish that arbitrary Qwen candidates will be correct or pass Stage 4.",
    "Split recovery uses two model calls and therefore costs more inference time than a successful compact full-plan response.",
    "The additive insert_before operator reduces accidental deletion but can still insert semantically wrong code; syntax, targeted, regression, security, held-out, latency, and canary tests remain mandatory.",
    "Task-specific guidance currently includes an arithmetic contract path. Other recurring task contracts may need their own evidence-derived guidance if preserved failures demonstrate a stable construction obstacle.",
    "The first live candidate passed Stage 4 but did not pass the independent canary. This is a remaining candidate-behavior limitation, not an evaluator transport or attribution failure.",
    "The independent canary showed that repairing one captured malformed form does not guarantee the same patch handles a shorter fresh output such as $408; the revision prompt must incorporate this new evidence without hard-coding one answer."
  ],
  "nextSteps": [
    "Allow the scheduled arithmetic revision to use the fresh canary failure and repeat targeted, paired, integration, and independent-canary validation without weakening thresholds.",
    "Inspect whether the next revision generalizes expression reconstruction from user_text across both captured malformed outputs; promote only if the independent canary and normal governor checkpoints pass.",
    "Replay the seven previously guided Qwen candidates through the mandatory paired evaluator without rebuilding their patches.",
    "Extend the qualification campaign with a deterministic surrogate for malformed/empty/HTTP-400 local candidate responses so the split recovery path is always covered without making qualification dependent on nondeterministic live inference."
  ]
}
