Hiro development journal

Controlled patch-guidance capability experiment

Experiment completed, evaluator confound isolated, no promotion performed Machine-readable JSON

Executive summary

A frozen ten-case experiment tested whether Hiro's repeated construction failures primarily reflected weak patch-generation capability or inadequate task guidance.

The preserved candidates formed the control and reached Stage 4 in zero of ten cases. Guided Qwen received full production source, a fixed independent targeted test, production guards, and the previous failure output; seven of ten one-shot candidates reached Stage 4, exceeding the predeclared three-of-ten threshold.

All seven evaluated Qwen candidates passed their targeted regression, the configured security comparison, and public and held-out quality non-regression. None became eligible because both suites reported latency ratios between 1.2799 and 1.3820 against a 1.1000 limit.

A current alternating three-pair null calibration then compared identical code back-to-back. Its ratios were 1.0038, 0.9866, and 1.0382 with zero false rejections, while the candidate comparisons used baseline timings recorded roughly two and a half hours earlier. The evidence therefore isolates non-contemporaneous latency attribution as the remaining approval-system flaw rather than inability by Qwen to construct useful patches.

GPT-4o reached Stage 4 in zero of ten cases with the identical evidence-rich prompt. Its proposed edits were usually fixture-specific canned answers or unsafe status-field substitutions and failed the fixed targeted tests.

Work completed

Frozen ten-case capability design

Completed
  • Selected ten preserved live failures spanning arithmetic reasoning, follow-up continuity, current information, and transit directions.
  • Used each preserved candidate packet as the control, including its five failed construction attempts and frozen targeted-test artifact.
  • Ran guided arms with exactly one construction attempt, fixed targeted tests appended by the harness, unchanged CandidateBuilder validation, and unchanged CandidateEvaluator gates.
  • Copied the evaluation SQLite database into an external experiment directory so experiment events could not contaminate the production ledger.
  • Disabled integration and promotion by design and kept the production queue frozen throughout the experiment.

Guided Qwen arm

Threshold passed
  • Seven of ten guided Qwen candidates became candidate-ready and reached Stage 4 in one attempt; the predeclared success threshold was at least three.
  • The control reached Stage 4 in zero of ten cases despite five attempts per preserved candidate.
  • Median changed production-and-test lines fell from 160.0 in the controls to 62.5 in the guided Qwen arm.
  • The three construction failures were one arithmetic, one current-information, and one transit case; each failed its fixed targeted test rather than an editor, transport, or text-anchor check.
  • The seven evaluated candidates all retained zero public and held-out score delta, no category regression, passing targeted regression, and passing configured security comparison.

GPT-4o comparison arm

Completed without a Stage 4 candidate
  • A valid ten-case GPT-4o rerun used the same evidence-rich task prompt and the same one-attempt construction policy.
  • Zero of ten GPT-4o candidates passed the fixed construction tests or reached Stage 4.
  • The median changed-line count was 52.0, but smaller patches did not make them correct.
  • Representative failures changed validator status from skipped to pass, hard-coded one benchmark answer into shared production behavior, or substituted a clarification tied to a single fixture.
  • The comparison demonstrates that stronger general model branding was not sufficient for these repository-specific tasks; Qwen used the supplied evidence more effectively.

Latency attribution diagnosis

Approval-system confound isolated
  • All seven functionally evaluated Qwen candidates were rejected only for public and held-out p95 latency ratios above 1.1000.
  • Public ratios were 1.2799, 1.3698, 1.3820, 1.3552, 1.3568, 1.3284, and 1.3316. Held-out ratios were 1.3110, 1.3170, 1.3417, 1.3689, 1.3375, 1.3418, and 1.3386.
  • The pinned baseline evidence was created near 18:32 UTC for the first examined case, while its candidate was evaluated near 21:07 UTC. The evaluator re-used the earlier latency instead of timing an unchanged control beside the candidate.
  • A fresh three-pair alternating null calibration of code-equivalent baseline and no-op variants produced ratios of 1.0038, 0.9866, and 1.0382, zero false rejections, stable case outcomes, and no category variance.
  • Causal reachability confirmed the candidate production edit was on the static and runtime-imported evaluation path, so it was genuinely testable. The uniform multi-candidate slowdown against old baselines, combined with stable contemporaneous null pairs, makes the old timing comparison unsuitable for causal rejection.

Reusable experiment harness

Implemented and committed
  • Added scripts/run_patch_guidance_experiment.py with explicit preserved cases, arm selection, model telemetry, isolated worktrees, isolated ledger backup, progress artifacts, and a frozen summary.
  • Added deterministic normalization for strict patch schemas and one-edit shorthand without allowing models to modify the harness-supplied tests.
  • Added unit coverage for schema normalization, path aliases, exclusion of test files from the production-source prompt, and the predeclared Stage 4 threshold.
  • Committed the harness and tests as Hiro revision 2011a75.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Predeclared guided-Qwen construction threshold passed Seven of ten candidates reached Stage 4; at least three were required.
Preserved control completed Zero of ten preserved candidates reached Stage 4 after five recorded attempts each.
Guided-Qwen Stage 4 functional evidence passed before latency gate All seven Stage 4 candidates passed their fixed targeted regression and unchanged security comparison, with zero public and held-out weighted-score delta.
GPT-4o identical-prompt comparison completed Zero of ten candidates reached Stage 4; all failed fixed candidate validation.
Current paired evaluator null calibration passed Three alternating code-equivalent pairs produced zero false rejections and latency ratios from 0.9866 to 1.0382.
Candidate causal reachability passed The examined Qwen candidate edit was static-reachable and runtime-imported by the evaluation entrypoint and classified directly attributable.
Experiment harness unit tests passed Four tests passed in 0.44 seconds.
Complete Hiro suite passed 683 tests passed in 230.43 seconds.
Preliminary harness runs excluded from capability evidence Early setup runs exposed a local-model context overflow, invalid uppercase candidate IDs, Windows long-path failure, and GPT shorthand-schema differences. The harness was corrected and all reported arm results come from subsequent valid complete runs.

Current state

Next steps