Moving the Hiro journal to a dedicated subdomain
The Hiro development journal has been reworked to use https://hiro.bballstatistics.com as its canonical public origin, with the journal index at the subdomain root.
Read the detailed update →Detailed project record
A timestamped account of work on Hiro: implementation, decisions, validation, limitations, and next steps. Multiple entries can be published on the same day without replacing earlier results. Entries intentionally exclude credentials, secrets, and actionable security-sensitive details.
The Hiro development journal has been reworked to use https://hiro.bballstatistics.com as its canonical public origin, with the journal index at the subdomain root.
Read the detailed update →Hiro now applies two deterministic mailbox rules before local Qwen sees a conversation: USPS-originated mail is always excluded, and automated sign-in or login notices are excluded from task and reading creation.
Read the detailed update →The Gmail productivity rules were refined again: Important remains the normal outer gate, while email sent from the owner to the owner is an explicit exception that must always become exactly one task or reading item.
Read the detailed update →The Gmail productivity workflow now uses Gmail importance as a hard outer gate. Non-important conversations never reach local Qwen or participate in task or reading creation.
Read the detailed update →The earlier pagination repair exposed the next twelve unanalyzed conversations, but it still required the user to click Analyze once per twelve-conversation group.
Read the detailed update →The task dashboard correctly reported that all eleven conversations on the first Gmail result page had already been analyzed, but the importer stopped there instead of continuing to older matching mail.
Read the detailed update →A screenshot from the private task dashboard showed that the corrected label-free Gmail importer still failed when the default Analyze action processed a larger set of full conversations.
Read the detailed update →The first Gmail importer incorrectly assumed that the mailbox owner would organize messages under dedicated task and reading labels.
Read the detailed update →A reported loading problem was clarified to affect Gmail import rather than the task page itself.
Read the detailed update →The private task page was reported as not loading shortly after Hiro's detached startup fix.
Read the detailed update →Hiro's prior manual startup command successfully launched the service but remained attached to the long-running process tree, making the command appear stuck even after the dashboard was available.
Read the detailed update →Hiro was started after the task dashboard was found offline.
Read the detailed update →Hiro now recognizes rejected or malformed Google credentials as a recoverable authorization condition instead of exposing a generic Gmail RefreshError.
Read the detailed update →The personal task dashboard's Gmail import returned a RefreshError during its first live authorization attempt.
Read the detailed update →Hiro was restarted so the newly implemented personal task and reading-list dashboard could be evaluated through its normal local and private-network listeners.
Read the detailed update →A work email task-manager design was reviewed as the basis for a personal system organized by areas of life such as Hiro, self-care, and family.
Read the detailed update →The scheduled July 24 Pacific nightly evaluation did not run. The canonical ledger contains no nightly record for that date, while the checked-in schedule remains disabled.
Read the detailed update →The overnight scheduler is disabled and no Hiro-related Windows Scheduled Task is currently registered, but an independent Daylab process was still running continuously.
Read the detailed update →Added a phased geolocation-trigger workstream to Hiro's public product direction.
Read the detailed update →Audited Hiro's current calendar, mailbox, task, reading-list, reminder, memory, and web-update capabilities to distinguish simulated evidence from production readiness.
Read the detailed update →Added a Performance tab to Hiro's private Evaluation Observatory so routed development tasks can be reviewed by time, task label, selected model, route, status, duration, and token usage.
Read the detailed update →Implemented a separate Hiro development-task router that selects among local Qwen, Luna Low, Terra Medium, Sol Medium, and explicitly approved Sol High without changing Hiro's conversational router.
Read the detailed update →Implemented and launched the finite Step 3 diversity controller instead of using the prior unbounded repetition loop. The repaired calibration session is step3-20260723-0852, with a seven-hour deadline and a fixed 348-case noncanonical budget across ten capability families.
Read the detailed update →Designed a finite execution-grade roadmap Step 3 campaign that a lighter supervising model can initiate, monitor, resume, and close through deterministic controller commands and a fixed evidence decision table.
Read the detailed update →Implemented the approved plan for Hiro's Qwen 3.6 stateful workflow gaps. The work separated harness defects from model behavior, added per-round telemetry, strengthened simulated approval and stale-state semantics, and applied only the bounded model controls supported by evidence.
Read the detailed update →Reviewed both Qwen 3.6 canonical stateful reports and the exact contracts for workflows three, five, and six. Both runs passed three of six; increasing the harness output budget did not change the score.
Read the detailed update →Completed the coordinated local-model migration that the earlier diagnosis identified as missing. Hiro now defaults to the Qwen 3.6 35B-A3B profile, uses the exact LM Studio native model identifier, and has Qwen-specific startup, preflight, health-restart, Telegram, benchmark, smoke-test, and documentation paths.
Read the detailed update →Diagnosed unexpectedly high memory use while Hiro appeared idle. The local inference server was still active with the Gemma 12B quantized model loaded in LM Studio, even though its status was idle and no evaluation was currently running.
Read the detailed update →Expanded Hiro's hermetic stateful harness from six fixed workflows to seeded generated scenarios with topology-based novelty signatures. The generator now exercises concurrent calendar changes, approval scope drift, approval expiration, explicit rejection, partial success, rate limits, and permission denials.
Read the detailed update →Implemented Harness v2 as a hermetic, executable stateful evaluation environment for six personal-assistant workflows: read-only discovery, local task and reading changes, draft-only calendar and email preparation, approval-gated writes, injected-failure recovery, and a staged cross-tool workflow.
Read the detailed update →The simulated personal-assistant evaluation layer is substantial but not yet a true stateful tool harness. Hiro has broad prompt-only coverage for calendar planning, tasks, reading lists, briefings, approvals, recovery choices, and productive autonomy, plus a sealed development-cycle generator and append-only evidence.
Read the detailed update →The approved first two phases of Hiro's governed memory plan were implemented: an additive typed SQLite operational store and a pinned, local-only GBrain recall bridge.
Read the detailed update →A two-phase implementation plan was designed for user review; no Hiro code, database, service, dependency, configuration, memory, or evaluation data was changed.
Read the detailed update →A review of the curated AI second-brain landscape found several useful memory layers, but no product should become Hiro's sole source of truth for calendar, task, reading-list, or approval state.
Read the detailed update →A new isolated development-cycle runner generated calibration, development, blind, and canary batches and recorded all observations append-only.
Read the detailed update →The repeated perfect scores on Hiro's current active suite rotation were diagnosed as evaluation saturation rather than proof that the broader assistant goals are solved.
Read the detailed update →A supervised evaluation, proposal review, bounded implementation, and re-evaluation loop ran across calibration, composed reasoning, personal-assistant planning, grounding, and operational regression coverage.
Read the detailed update →The user approved two bounded Hiro improvements: proactive grounding for fresh factual questions and a candidate response to repeated ordered-transformation failures.
Read the detailed update →A completed seven-cycle Daylab evidence set revealed that synthetic personal-assistant prompts were entering the normal interactive resolver path, making those benchmark scores unsuitable for capability assessment.
Read the detailed update →Daylab now selects public evaluation suites adaptively: a stable regression canary, synthetic personal-assistant workflows, and suites ranked by observed failure rate.
Read the detailed update →Hiro Daylab was upgraded from a fixed 30-minute cadence into an on-demand, bounded evidence-burst capability.
Read the detailed update →The protected remote Evaluation Observatory was diagnosed and corrected so an authorized viewer needs only their Cloudflare identity, not a separately carried Hiro API key.
Read the detailed update →This session added Hiro Daylab: a bounded, proposal-only development lane that runs immediately and then every 30 minutes while preserving the canonical nightly lane for readiness evidence.
Read the detailed update →A user-requested manual run of Hiro's guarded proposal-only nightly workflow completed successfully after the normal readiness preflight confirmed that the local model and dashboard API were healthy.
Read the detailed update →This session completed a private, browser-only route to Hiro's live benchmark dashboard without requiring a VPN client or remote desktop on the work computer.
Read the detailed update →The nightly evaluation connection failure was traced to a dependency gap: the preflight kept Hiro's API available but did not ensure the local Gemma inference server was available on port 8080, which the evaluator calls directly.
Read the detailed update →The July 20 scheduled proposal-only run completed its bookkeeping and all 16 public RSI observations, but it is classified as an infrastructure/model failure rather than a valid capability evaluation: every observation returned an empty response with the same connection error.
Read the detailed update →Recorded the recurring Windows Path/PATH case-collision as a standing Hiro workspace constraint so future planning accounts for it before introducing process launchers.
Read the detailed update →Stopped Hiro's health monitor from sending Telegram messages when the active local model endpoint is unreachable or returns an error. Model availability remains visible to health checks and local logs, but routine model downtime no longer interrupts the user through Telegram.
Read the detailed update →Reworked the Hiro development journal so a calendar date is no longer the unique identity of an entry. New sessions use schema-version-2 source files named with local date and time, allowing multiple independent entries on the same day without rewriting earlier results.
Read the detailed update →Hiro's two local-model libraries were relocated from the space-constrained system volume to dedicated directories on E:. LM Studio now downloads to E:\AI\Models\LMStudio, while Hiro's legacy GGUF library resides at E:\AI\Models\llama.cpp.
Read the detailed update →Hiro's host was inventoried after its GPU upgrade. It now has an NVIDIA GeForce RTX 5090 with 32,607 MiB of VRAM, an Intel Core i7-10700K with 16 logical processors, 31.84 GiB of physical RAM, and several terabytes of free space on E:.
Read the detailed update →Hiro's local dashboard was unreachable because no process was listening on its web ports. The browser's connection-refused message was accurate and was unrelated to the dashboard API key.
Read the detailed update →The launcher now uses the supported pinned LM Studio CLI, verifies Gemma 4 12B model identity and readiness, warms the model, records telemetry, and enables Telegram only after success.
Read the detailed update →The public Hiro journal was audited after ChatGPT Voice could not retrieve a dated update that remained accessible in a normal browser and regular ChatGPT. Direct production HTTP tests showed that the existing HTML and JSON URLs already returned clean 200 responses, correct content types, no cookies, and identical bodies for browser, curl, generic-bot, GPTBot, and ChatGPT-style user agents.
Read the detailed update →Hiro's local operating workflow was tightened so the model server and Telegram integration behave as one intentional lifecycle instead of unrelated background processes. Starting Hiro now enables Telegram notifications and starts the bot, while the new shutdown path disables notifications before stopping services.
Read the detailed update →