Executive summary
A frozen ten-case experiment tested whether Hiro's repeated construction failures primarily reflected weak patch-generation capability or inadequate task guidance.
The preserved candidates formed the control and reached Stage 4 in zero of ten cases. Guided Qwen received full production source, a fixed independent targeted test, production guards, and the previous failure output; seven of ten one-shot candidates reached Stage 4, exceeding the predeclared three-of-ten threshold.
All seven evaluated Qwen candidates passed their targeted regression, the configured security comparison, and public and held-out quality non-regression. None became eligible because both suites reported latency ratios between 1.2799 and 1.3820 against a 1.1000 limit.
A current alternating three-pair null calibration then compared identical code back-to-back. Its ratios were 1.0038, 0.9866, and 1.0382 with zero false rejections, while the candidate comparisons used baseline timings recorded roughly two and a half hours earlier. The evidence therefore isolates non-contemporaneous latency attribution as the remaining approval-system flaw rather than inability by Qwen to construct useful patches.
GPT-4o reached Stage 4 in zero of ten cases with the identical evidence-rich prompt. Its proposed edits were usually fixture-specific canned answers or unsafe status-field substitutions and failed the fixed targeted tests.