Hiro development journal

Trip planning and web-client stability repaired after a live Big Bear failure

Implemented, activated, and verified on the live local service Machine-readable JSON

Executive summary

A real weekend-trip request exposed two coupled defects: Hiro answered with raw event-search material instead of a usable itinerary, and a chat-screen gesture could reload the web client while the dashboard was also issuing multiple background chat requests.

The backend remained healthy during the reported incident. The apparent crash was traced to client reload behavior and avoidable dashboard contention, not a service-process termination.

Trip requests now use a bounded trip-planning path that concurrently gathers activities, dining, and lodging evidence and renders a Friday-through-Sunday plan with booking checks and sources.

The dashboard now uses one noninteractive summary endpoint. It no longer sends background requests through the production chat route or initiates calendar authorization during page load.

The custom pull-to-refresh reload trigger was removed, and Windows log rotation was repaired so all named loggers share one rotating file handle.

The final repository suite passed all 550 tests. A live Big Bear acceptance returned an itinerary in 5.3 seconds with ten evidence records, six source links, a passing validator result, and no raw search dump or JSON leakage.

Work completed

Incident diagnosis and client stability

Activated
  • Process and request logs showed that the API, model service, and web server stayed available during the reported failure.
  • The chat transcript scrolls inside its own message container, but the former pull-to-refresh code inspected the outer screen scroll position. An ordinary downward gesture in chat could therefore call location.reload and appear to crash or reset the application.
  • At the same time, the dashboard was independently sending several production-chat requests for weather, calendar, and sports information. A reload produced another burst, increasing contention while the user was interacting with Hiro.
  • The custom reload gesture has been removed from the client. Mobile-size verification completed without console warnings or errors.

Noninteractive dashboard boundary

Activated
  • Added a dedicated read-only dashboard summary route that performs only a bounded deterministic weather lookup.
  • Calendar and sports panels now display explicit prompts to ask Hiro on demand instead of silently consuming the production chat path.
  • The dashboard no longer creates a dashboard chat session, invokes model synthesis in the background, or initiates an interactive calendar authorization flow during page load.
  • Live verification observed one dashboard-summary request, zero authorization prompts, and zero dashboard-chat requests.

Grounded weekend-trip planning

Activated
  • Trip-planning language is detected before ordinary model tool selection so it cannot collapse into the generic event-discovery response that caused the incident.
  • Three bounded web-search lanes run concurrently for dated activities, dining, and lodging. Their receipts are recorded under one trip-planning resolver for audit and response-quality evaluation.
  • The response formatter constructs Friday, Saturday, and Sunday sections, separates lodging and dinner leads from activities, identifies booking checks, and includes source links.
  • Search-result cleanup removes page chrome, obvious instruction-like contamination, JSON field fragments, and duplicated branded result families before they can reach the itinerary.
  • The response gate now rejects raw search-result dumps and incomplete trip plans rather than treating them as valid assistant answers.

Interaction-lab coverage

Activated
  • The captured failure pattern was added to the isolated interaction lab without publishing private conversation or session content.
  • The trip-planning contract requires an itinerary shape, recommendation content, and web evidence; a source-list response fails the contract.
  • Tests cover the exact failure class, the improved itinerary, missing-day rejection, malformed search output, three-lane retrieval, duplicate-source suppression, and instruction-like snippet filtering.
  • This keeps real user-facing failures in the same continuous improvement queue while giving them task-specific replay criteria.

Windows logging reliability

Activated
  • Multiple named loggers previously opened independent rotating handlers for the same file. On Windows, simultaneous rollover attempts could raise file-access errors.
  • Named loggers now share one rotating file handler per resolved path.
  • The live restart produced no logging rollover error, and launcher behavior continues to use the checked-in hidden process path with normalized Windows search-path handling.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Final full repository suite passed 550 tests passed in 159.57 seconds after the dashboard, trip-planning, response-gate, interaction-lab, and shared-logger changes.
Focused stability and trip-planning suite passed 23 focused tests passed after the final search-snippet filtering change.
Static and syntax checks passed Python bytecode compilation and diff whitespace validation completed successfully.
Mobile web-client acceptance passed At a 390 by 844 viewport, the dashboard loaded weather through the summary route, kept calendar and sports on demand, and emitted no browser console error or warning.
Live Big Bear trip acceptance passed The exact request class returned Friday, Saturday, and Sunday sections plus lodging and dining leads in 5.3 seconds. The response had six source links, ten recorded evidence items, a passing validator, HTTP 200, and no raw search heading or JSON fragment.
Live process and dashboard health passed Hiro restarted through the detached hidden launcher. Post-restart logs showed one dashboard-summary request, no calendar authorization prompt, and no rotating-log error.

Current state

Next steps