{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T19:05:46-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Requiring attributable regression tests before candidate evaluation",
  "publicationStatus": "Early baseline-contrast gate implemented, tested, and loaded into the live queue",
  "executiveSummary": [
    "Hiro's candidate builder previously required a generated targeted test to pass on the candidate but deferred the untouched-baseline counterfactual until Stage 4. This allowed expensive evaluations to begin before discovering that a test already passed on baseline or could not execute there.",
    "Stage 3 now copies the exact candidate-authored tests into a detached checkout of the candidate's immutable base commit. A candidate can be frozen as ready only when those tests collect, execute, and produce a JUnit-recorded assertion failure on baseline while passing on the candidate.",
    "Import errors, collection errors, setup or runtime errors, absent reports, zero collected tests, and skip-only runs are explicitly inconclusive rather than evidence of improvement. Tests that pass on baseline are classified separately as failing to demonstrate a change.",
    "The bounded repair loop now receives the contrast classification, outcome counts, and compact collection and execution output, together with a repair focus tailored to the failure. Hiro was restarted on the new commit, its implementation-time dirty-repository breaker was audit-reset, and the scheduler selected a live candidate under the new code.",
    "After a live packet showed that Qwen twice imported newly invented candidate-only router functions, the builder was strengthened again: it now supplies an exact baseline module-to-public-symbol allowlist and explicitly requires every production import in a targeted test to come from that list."
  ],
  "workstreams": [
    {
      "title": "Early counterfactual gate",
      "status": "Completed",
      "details": [
        "Extended CandidateBuilder validation after candidate syntax, collection, and targeted tests pass.",
        "Creates a uniquely named detached Git worktree at the exact base commit under the external candidate-worktree root.",
        "Copies only the exact targeted tests from the candidate workspace into that untouched baseline and validates collection before execution.",
        "Runs pytest with a bounded JUnit artifact and removes the detached baseline worktree in a finally block.",
        "Candidate readiness now requires both local candidate success and an improvement_demonstrated baseline contrast."
      ]
    },
    {
      "title": "Assertion-aware evidence classification",
      "status": "Completed",
      "details": [
        "Added a shared targeted-contrast module used by both Stage 3 construction and Stage 4 evaluation.",
        "Improvement is demonstrated only when at least one non-skipped case executed, JUnit recorded one or more failures, JUnit recorded zero errors, and pytest returned the ordinary test-failure code.",
        "A clean baseline pass is classified as already_passes_baseline rather than inconclusive.",
        "Missing or malformed JUnit evidence, no executable cases, all-skipped cases, test errors, and inconsistent return-code/report combinations are inconclusive.",
        "Stage 4 now repeats this stronger JUnit-aware check as defense in depth for frozen packets."
      ]
    },
    {
      "title": "Bounded test repair",
      "status": "Completed",
      "details": [
        "Preserved the existing maximum of two repairs after the initial complete candidate attempt.",
        "Each failed attempt still resets to the clean base commit; repairs must return a complete revised patch and tests rather than editing a contaminated failed attempt.",
        "Repair context now contains the baseline status, reason, outcome counts, and compact stdout and stderr from collection and execution.",
        "The local candidate agent receives a specific focus: create an assertion contrast when baseline already passes, repair imports to use baseline-existing entrypoints after collection failure, or remove errors and skips for other inconclusive executions.",
        "The prompt now explicitly forbids treating candidate-only imports, setup failures, collection failures, skipped tests, or empty collections as improvement evidence.",
        "The full code-owned task contract is now carried from the incident spec into CandidateBuildRequest, the Qwen request, and the frozen packet as the authoritative behavioral oracle.",
        "The request also contains baseline_test_imports, a bounded AST-derived map of public functions and classes that exist in the untouched baseline production files; private helpers and test-local symbols are excluded."
      ]
    },
    {
      "title": "Live service transition",
      "status": "Completed",
      "details": [
        "Committed the implementation before service activation so new candidates pin a stable base revision.",
        "The old scheduler opened its infrastructure breaker after three blocked_dirty_repository events during the edit window, correctly refusing to work against changing source.",
        "Stopped only the existing Hiro API process tree, preserved the loaded Qwen model server, and restarted Hiro through the checked-in path-normalizing detached launcher.",
        "Reset the resolved dirty-repository breaker through ContinuousQueue.clear_infrastructure_failures, producing an append-only circuit_breaker_reset event.",
        "The restarted scheduler selected a preference-comparison incident and began candidate construction on the new code with Qwen connected.",
        "That live frozen packet was pinned to 3a86749, preserved the full preference-comparison task contract, and classified attempts 0 and 2 as baseline collection/import failures because their tests imported candidate-only router functions. No invalid test was credited.",
        "The live evidence motivated the final 0b2a784 baseline_test_imports constraint. Hiro was restarted again, the interrupted lease was released only after its exact worker process had been stopped, and stale candidate state is rebuilt on the active revision by the queue's existing upgrade recovery."
      ]
    }
  ],
  "decisions": [
    "Move targeted baseline contrast into Stage 3 so invalid tests are repaired before public and protected held-out evaluation costs are incurred.",
    "Require structured JUnit evidence instead of inferring an assertion failure from pytest return code 1 alone.",
    "Treat setup and runtime errors as inconclusive even when pytest returns the general test-failure code.",
    "Require targeted tests to use entrypoints that exist on the immutable baseline; candidate-only imports cannot establish a counterfactual.",
    "Give the candidate model an exact AST-derived baseline import allowlist instead of relying on prose alone.",
    "Carry the complete code-owned task contract into construction and freeze it in the candidate packet for auditability.",
    "Retain the Stage 4 contrast as an independent verification of the frozen candidate rather than trusting Stage 3's declaration.",
    "Keep repair attempts bounded and reset each attempt to the clean baseline.",
    "Do not weaken global public, held-out, latency, invariant, causal-attribution, or promotion gates after the targeted test succeeds."
  ],
  "validation": [
    {
      "check": "Candidate builder and evaluator tests",
      "status": "passed",
      "result": "The final focused suite contains 60 passing builder, evaluator, sandbox, continuous-queue, and agenda tests. Coverage includes real detached baseline assertion contrast, tests that already pass on baseline, repair-context diagnostics, strict return-code/JUnit classification, errors, skipped tests, absent reports, task-contract propagation, and exclusion of private or test-local symbols from the baseline import allowlist."
    },
    {
      "check": "Sandbox, continuous queue, and agenda integration tests",
      "status": "passed",
      "result": "The integration-focused subset is included in the 60 focused passes and confirms that stronger candidate packet requirements do not break queue revision handling, sandbox orchestration, or agenda execution."
    },
    {
      "check": "Full Hiro regression suite",
      "status": "passed",
      "result": "638 tests passed in 183.28 seconds after the final baseline-symbol constraint."
    },
    {
      "check": "Live service restart and scheduler",
      "status": "passed",
      "result": "Hiro restarted through scripts/start_hiro.py, reported healthy with qwen/qwen3.8-27b connected, reset the resolved breaker to zero failures, and selected a scheduler-owned candidate on the new commit."
    },
    {
      "check": "Live candidate contrast packet",
      "status": "passed",
      "result": "The scheduler produced sandbox-20260817021750-570da205-01 on base 3a86749 with the complete task contract frozen. Attempts 0 and 2 were explicitly inconclusive because baseline collection returned 2 for candidate-only router imports; attempt 1 failed local validation. After adding the exact-import constraint, sandbox-20260817022645-55020c0e-01 was pinned to final base 0b2a784 with its writing-assistance contract; it failed earlier at empty-plan, ambiguous-edit, and candidate-local validation gates and therefore never reached baseline contrast. Neither packet advanced."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 140 journal entries, then TypeScript and Vite completed the production build successfully."
    }
  ],
  "currentState": [
    "The implementation is committed as 39ffa14, 3a86749, and 0b2a784 on codex/rsi-first-cycle and the Hiro source tree is clean.",
    "Hiro is healthy, Qwen 3.8 27B remains connected, Daylab remains absent, and the continuous queue breaker is closed with zero infrastructure failures.",
    "The scheduler has exercised the new contrast classification with a real preference-comparison candidate, rejected its candidate-only imports, rebuilt work on the final revision, and remains operative with another candidate active.",
    "No candidate has been credited or integrated merely because its generated test errored or failed to collect."
  ],
  "limitations": [
    "A well-formed assertion contrast proves the patch changes the targeted behavior, but it does not by itself prove the behavior generalizes; public, held-out, and non-regression gates remain necessary.",
    "The repair agent still has only two bounded retries and may fail to produce a robust test within that budget.",
    "The first final-revision packet failed before it reached baseline contrast, so the explicit baseline symbol allowlist's effect on a locally passing candidate still needs longitudinal measurement.",
    "JUnit distinguishes test failures from errors but cannot determine that every assertion encodes the correct product requirement; success criteria and broader suites remain part of review.",
    "Existing retrying candidates were created from earlier incidents and will receive the stronger gate only when the live queue constructs their next candidate packet.",
    "The separate open-problem agenda remains calibration-gated and is not advanced by this change."
  ],
  "nextSteps": [
    "Measure the live rate of improvement_demonstrated, already_passes_baseline, collection/import failure, execution error, and skip-only outcomes.",
    "Compare Stage 3 rejection rates and Stage 4 inconclusive rates before and after this change to verify that failures move earlier and become more repairable.",
    "Build reusable contract-level test templates for common interaction categories so generated tests exercise stable public entrypoints.",
    "Add mutation checks where valuable to detect assertions that pass for irrelevant reasons despite producing a baseline failure.",
    "Continue evaluator calibration and dedicated-inference-lane work independently of this deterministic targeted-test improvement."
  ]
}
