Executive summary
A completed seven-cycle Daylab evidence set revealed that synthetic personal-assistant prompts were entering the normal interactive resolver path, making those benchmark scores unsuitable for capability assessment.
Daylab was paused immediately. The evaluation harness now uses a dedicated local-only inference path with no tools, resolver access, user history, session persistence, or remote-model fallback.
A clean replacement evaluation completed 16 observations at an 87.5% pass rate with zero tool, source, resolver, or execution-error observations. Its only measured gap was the reading-list tag formatting case.
No external write action was enabled or performed. The contaminated observations are retained in the append-only ledger for audit but excluded from readiness conclusions.