Executive summary
Daylab now selects public evaluation suites adaptively: a stable regression canary, synthetic personal-assistant workflows, and suites ranked by observed failure rate.
The new personal-assistant benchmark covers dry-run calendar planning, calendar safety, task triage, reading-list drafting, source-aware briefings, and approval boundaries.
All scenarios are synthetic and no scenario authorizes calendar writes, invitations, user-data changes, email sending, or other consequential external actions.
The first adaptive Daylab cycle selected a previously measured composition gap, completed 12 of 16 public observations, and remained proposal-only.