Executive summary
Hiro completed a sealed Stage 4 gate campaign using Qwen 3.6 35B-A3B to construct bounded frozen candidates for one eligible path and six distinct rejection paths.
All seven valid candidate decisions matched their sealed expectations: one candidate was eligible for human review, while held-out transfer, public minimum improvement, regression, latency, invariant, and category-regression candidates were rejected for the intended evidence.
The valid evidence covered 260 baseline-and-candidate observations, one exact interrupted-run recovery, zero false accepts, zero false rejects, zero integrity failures, and no merge, promotion, or deployment.
One initial category probe was correctly excluded as an invalid benchmark because its case-only output distinction was incompatible with the scorer's documented case-insensitive regex semantics. A separately sealed semantic-token replacement produced the intended category regression and was rejected.
The campaign exposed and corrected a separate process-logging issue: standalone candidate builders now establish a process-specific namespace before loading the local model/router logging stack.