Executive summary
Hiro now runs a bounded null experiment before interpreting a narrow candidate regression. The experiment compares labels called baseline and no-op while holding code, model, prompt, tools, configuration, suite, and case identities equal, and it reverses execution order in the middle pair.
The first live Qwen 3.8 calibration found that identical code would be rejected by the existing gate in two of three pairs, a 66.7% false-rejection rate. Both false rejections came from p95 latency ratios on a shared single-slot model server. The public uncertainty case also flipped outcome across repetitions and had a score standard deviation of 0.046.
Causal attribution classified the stopped evaluator-diagnostics patch as unreached by the global isolated-agent evaluation: no static dependency or runtime import connected the changed implementation file to the entrypoint, whose normalized source was identical. This evidence is diagnostic only and did not override the rejection.
Future agenda candidates are now evaluated with three fresh baseline/candidate pairs in alternating order and aggregate scoring across all repetitions. Latency can pass only when the evaluation explicitly records a dedicated inference slot; shared-slot measurements produce an inconclusive result rather than acceptance or candidate blame.