{
  "schemaVersion": 2,
  "date": "2026.08.17",
  "publishedAt": "2026-08-17T21:45:08-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Controlled patch-guidance capability experiment",
  "publicationStatus": "Experiment completed, evaluator confound isolated, no promotion performed",
  "executiveSummary": [
    "A frozen ten-case experiment tested whether Hiro's repeated construction failures primarily reflected weak patch-generation capability or inadequate task guidance.",
    "The preserved candidates formed the control and reached Stage 4 in zero of ten cases. Guided Qwen received full production source, a fixed independent targeted test, production guards, and the previous failure output; seven of ten one-shot candidates reached Stage 4, exceeding the predeclared three-of-ten threshold.",
    "All seven evaluated Qwen candidates passed their targeted regression, the configured security comparison, and public and held-out quality non-regression. None became eligible because both suites reported latency ratios between 1.2799 and 1.3820 against a 1.1000 limit.",
    "A current alternating three-pair null calibration then compared identical code back-to-back. Its ratios were 1.0038, 0.9866, and 1.0382 with zero false rejections, while the candidate comparisons used baseline timings recorded roughly two and a half hours earlier. The evidence therefore isolates non-contemporaneous latency attribution as the remaining approval-system flaw rather than inability by Qwen to construct useful patches.",
    "GPT-4o reached Stage 4 in zero of ten cases with the identical evidence-rich prompt. Its proposed edits were usually fixture-specific canned answers or unsafe status-field substitutions and failed the fixed targeted tests."
  ],
  "workstreams": [
    {
      "title": "Frozen ten-case capability design",
      "status": "Completed",
      "details": [
        "Selected ten preserved live failures spanning arithmetic reasoning, follow-up continuity, current information, and transit directions.",
        "Used each preserved candidate packet as the control, including its five failed construction attempts and frozen targeted-test artifact.",
        "Ran guided arms with exactly one construction attempt, fixed targeted tests appended by the harness, unchanged CandidateBuilder validation, and unchanged CandidateEvaluator gates.",
        "Copied the evaluation SQLite database into an external experiment directory so experiment events could not contaminate the production ledger.",
        "Disabled integration and promotion by design and kept the production queue frozen throughout the experiment."
      ]
    },
    {
      "title": "Guided Qwen arm",
      "status": "Threshold passed",
      "details": [
        "Seven of ten guided Qwen candidates became candidate-ready and reached Stage 4 in one attempt; the predeclared success threshold was at least three.",
        "The control reached Stage 4 in zero of ten cases despite five attempts per preserved candidate.",
        "Median changed production-and-test lines fell from 160.0 in the controls to 62.5 in the guided Qwen arm.",
        "The three construction failures were one arithmetic, one current-information, and one transit case; each failed its fixed targeted test rather than an editor, transport, or text-anchor check.",
        "The seven evaluated candidates all retained zero public and held-out score delta, no category regression, passing targeted regression, and passing configured security comparison."
      ]
    },
    {
      "title": "GPT-4o comparison arm",
      "status": "Completed without a Stage 4 candidate",
      "details": [
        "A valid ten-case GPT-4o rerun used the same evidence-rich task prompt and the same one-attempt construction policy.",
        "Zero of ten GPT-4o candidates passed the fixed construction tests or reached Stage 4.",
        "The median changed-line count was 52.0, but smaller patches did not make them correct.",
        "Representative failures changed validator status from skipped to pass, hard-coded one benchmark answer into shared production behavior, or substituted a clarification tied to a single fixture.",
        "The comparison demonstrates that stronger general model branding was not sufficient for these repository-specific tasks; Qwen used the supplied evidence more effectively."
      ]
    },
    {
      "title": "Latency attribution diagnosis",
      "status": "Approval-system confound isolated",
      "details": [
        "All seven functionally evaluated Qwen candidates were rejected only for public and held-out p95 latency ratios above 1.1000.",
        "Public ratios were 1.2799, 1.3698, 1.3820, 1.3552, 1.3568, 1.3284, and 1.3316. Held-out ratios were 1.3110, 1.3170, 1.3417, 1.3689, 1.3375, 1.3418, and 1.3386.",
        "The pinned baseline evidence was created near 18:32 UTC for the first examined case, while its candidate was evaluated near 21:07 UTC. The evaluator re-used the earlier latency instead of timing an unchanged control beside the candidate.",
        "A fresh three-pair alternating null calibration of code-equivalent baseline and no-op variants produced ratios of 1.0038, 0.9866, and 1.0382, zero false rejections, stable case outcomes, and no category variance.",
        "Causal reachability confirmed the candidate production edit was on the static and runtime-imported evaluation path, so it was genuinely testable. The uniform multi-candidate slowdown against old baselines, combined with stable contemporaneous null pairs, makes the old timing comparison unsuitable for causal rejection."
      ]
    },
    {
      "title": "Reusable experiment harness",
      "status": "Implemented and committed",
      "details": [
        "Added scripts/run_patch_guidance_experiment.py with explicit preserved cases, arm selection, model telemetry, isolated worktrees, isolated ledger backup, progress artifacts, and a frozen summary.",
        "Added deterministic normalization for strict patch schemas and one-edit shorthand without allowing models to modify the harness-supplied tests.",
        "Added unit coverage for schema normalization, path aliases, exclusion of test files from the production-source prompt, and the predeclared Stage 4 threshold.",
        "Committed the harness and tests as Hiro revision 2011a75."
      ]
    }
  ],
  "decisions": [
    "Accept the capability hypothesis for guided Qwen: reaching Stage 4 in seven of ten cases materially exceeds both the zero-of-ten control and the predeclared three-of-ten threshold.",
    "Reject the claim that Qwen is categorically unable to code Hiro patches. Guidance quality was the dominant construction variable in this experiment.",
    "Do not treat these seven latency rejections as demonstrated patch regressions because candidate and baseline timing were not measured contemporaneously.",
    "Do not weaken functional, targeted, security, or quality gates. Repair latency attribution by pairing a fresh unchanged baseline with each candidate under matched runtime conditions.",
    "Keep the live queue frozen until the latency-comparison path is corrected; restarting the existing evaluator would continue generating misleading rejections.",
    "Do not promote any experiment candidate. This session measured capability and evaluator behavior only."
  ],
  "validation": [
    {
      "check": "Predeclared guided-Qwen construction threshold",
      "status": "passed",
      "result": "Seven of ten candidates reached Stage 4; at least three were required."
    },
    {
      "check": "Preserved control",
      "status": "completed",
      "result": "Zero of ten preserved candidates reached Stage 4 after five recorded attempts each."
    },
    {
      "check": "Guided-Qwen Stage 4 functional evidence",
      "status": "passed before latency gate",
      "result": "All seven Stage 4 candidates passed their fixed targeted regression and unchanged security comparison, with zero public and held-out weighted-score delta."
    },
    {
      "check": "GPT-4o identical-prompt comparison",
      "status": "completed",
      "result": "Zero of ten candidates reached Stage 4; all failed fixed candidate validation."
    },
    {
      "check": "Current paired evaluator null calibration",
      "status": "passed",
      "result": "Three alternating code-equivalent pairs produced zero false rejections and latency ratios from 0.9866 to 1.0382."
    },
    {
      "check": "Candidate causal reachability",
      "status": "passed",
      "result": "The examined Qwen candidate edit was static-reachable and runtime-imported by the evaluation entrypoint and classified directly attributable."
    },
    {
      "check": "Experiment harness unit tests",
      "status": "passed",
      "result": "Four tests passed in 0.44 seconds."
    },
    {
      "check": "Complete Hiro suite",
      "status": "passed",
      "result": "683 tests passed in 230.43 seconds."
    },
    {
      "check": "Preliminary harness runs",
      "status": "excluded from capability evidence",
      "result": "Early setup runs exposed a local-model context overflow, invalid uppercase candidate IDs, Windows long-path failure, and GPT shorthand-schema differences. The harness was corrected and all reported arm results come from subsequent valid complete runs."
    }
  ],
  "currentState": [
    "The controlled experiment is complete and no candidate was promoted or integrated.",
    "Guided Qwen demonstrated seven-of-ten Stage 4 construction capability on preserved failures.",
    "The live Hiro service and production queue remain stopped intentionally while the non-contemporaneous latency comparison remains in the approval path.",
    "The Qwen inference service remains available for further controlled work.",
    "The experiment harness is committed locally at revision 2011a75 and all 683 Hiro tests pass."
  ],
  "limitations": [
    "Ten preserved cases are sufficient to falsify the zero-capability hypothesis but not to estimate general patch quality across the full open-problem agenda.",
    "Reaching Stage 4 proves the patch passed construction and targeted guards; it does not establish that every candidate is desirable enough to integrate.",
    "The current null calibration establishes that back-to-back timing is stable today. It does not recover the runtime conditions of the hours-old baselines used for the seven rejected candidates.",
    "The evidence strongly identifies baseline age and unmatched runtime state as the latency confound, but each candidate must still be compared with a fresh paired baseline before an eligibility decision.",
    "GPT-4o returned a deterministic one-edit shorthand instead of the requested strict schema. The valid rerun normalized that format, but model behavior remained fixture-specific and incorrect.",
    "No live queue throughput or promotion result was measured after the experiment because restarting the known-confounded evaluator would produce invalid approval data."
  ],
  "nextSteps": [
    "Change Stage 4 latency evaluation to run a fresh unchanged baseline and candidate as alternating or randomized matched pairs during the same evaluation window.",
    "Require latency rejection to be based on paired evidence rather than a historical pinned timing metric; keep the frozen baseline for quality identity and suite provenance.",
    "Replay the seven guided Qwen candidates through the corrected paired evaluator without rebuilding their patches and inspect whether any become eligible.",
    "Integrate the evidence-rich guidance packet into the normal candidate constructor only after the paired evaluator fix is validated.",
    "Restart Hiro and the continuously ranked queue only after the corrected evaluator passes null calibration and one preserved-candidate end-to-end replay."
  ]
}
