Executive summary
A forensic review of the live continuous-improvement queue confirms that the extreme rejection rate is substantially caused by approval-system design, not evidence that nearly every underlying idea is harmful.
The durable queue currently contains 173 rejected and 3 implemented ideas, a 98.3 percent rejection share among terminal rejected-or-implemented records. Since recovery event 1395, the latest classified outcomes for 34 ideas comprise 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.
The engine currently converts a repairable candidate-construction failure into a terminal idea rejection after the bounded attempt budget is exhausted. This conflates failure to manufacture an evaluable patch/test artifact with evidence that the idea itself lacks merit.
The single recent candidate-gate pass exposed the opposite defect. Its untouched-baseline run failed with TypeError because the candidate-authored test called the baseline chat function with a keyword its baseline signature did not accept. Pytest represented this call-phase exception as a JUnit failure, and Hiro's count-only parser classified any such failure as an attributable assertion contrast.
The live canary cannot correct that false positive because its default probe simply reruns the same single candidate-authored test at synthetic checkpoints. The approval stages therefore reuse correlated evidence instead of adding an independent behavioral check.