Executive summary
The simulated personal-assistant evaluation layer is substantial but not yet a true stateful tool harness. Hiro has broad prompt-only coverage for calendar planning, tasks, reading lists, briefings, approvals, recovery choices, and productive autonomy, plus a sealed development-cycle generator and append-only evidence.
The missing center is executable state. Current isolated evaluations give Hiro a fictional snapshot in one prompt and score its text response; they do not let Hiro call deterministic mock calendar, mailbox, task, reading, and approval tools, mutate a world over multiple turns, encounter injected failures, or verify the final state.
The recommended next milestone is Harness v2: a deterministic world model, schema-compatible mock tool runtime, actual agent tool loop with dependency injection, trace and final-state oracles, and fresh multi-turn scenario generators. This should replace high-volume repetition as the primary development frontier.