Executive summary
Hiro now tests its everyday conversational behavior proactively instead of relying only on user-reported failures and a small set of known failure patterns.
A versioned suite contains 32 ordinary questions across 16 categories, including preferences, explanations, arithmetic, practical instructions, writing, summarization, recommendations, trip planning, directions, current information, follow-up continuity, ambiguity, action safety, correction recovery, and prompt-injection resistance.
Four cases run every fifteen minutes after a quiet production window. The complete suite rotates in approximately two hours without calling the production chat endpoint, reading user memory, accessing personal accounts, or performing external actions.
Every response passes through Hiro's real universal response boundary and a deterministic task contract. A failed case becomes a deduplicated lane-two item in the same ranked continuous-improvement queue used by other reliability findings.
The first live scheduled batch passed the ordinary comparison and general-knowledge cases and found two additional usability failures. Both entered the live queue automatically; one immediately advanced to candidate construction.
The full repository suite passed all 557 tests, Hiro restarted successfully through the hidden launcher, and the Observatory now displays audit coverage and the latest batch result.