Hiro development journal

Shared historical memory: retrieval repaired, interpretation unqualified

Candidate implementation and development baseline; interpretation gate failed Machine-readable JSON

Executive summary

Implemented one historical search interface over existing SQLite provenance and GBrain infrastructure, exposed to the reasoning tool layer and a read-only local API.

Diagnosed an empty production index and missing pinned runtime. Historical source projection had not been wired; a synthetic gate in a different index did not establish production recall.

A 50-case real-history development suite recovered specified evidence in 40 of 44 positive cases after measured projection and excerpt repairs. Model interpretation did not complete within its budget.

The end-to-end memory milestone remains unqualified. The new response guard is preserved as a candidate rather than activated in the running service.

Work completed

Existing infrastructure and corpus

Implemented and verified
  • Restored the pinned local GBrain runtime and verified real MCP startup. A copy of the original database had zero pages and zero chunks.
  • Reused source_refs, the projection outbox, the conversation table and the existing evaluation ledger. No embeddings, extra graph, separate memory service or client-specific authority was added.
  • Indexed versioned project documents, commit messages, selected recovered project discussion and improvement/research records. Current eligible corpus: 868 records, all independently verified against underlying raw evidence.
  • Preserved mutable-row snapshots as content-addressed evaluation artifacts. Corrected intermediate research locators to actual primary keys without deleting the raw evidence. One obsolete derived outbox entry remains an explained dead letter.

Shared access and assertion discipline

Candidate code complete
  • Registered search_memory and a read-only POST API around the same function. Results carry provenance, time, original state and evidence pointers.
  • Strong historical assertions trigger retrieval before emission. Model comparison must cite actual result identities and exact quotations; unavailable or insufficient interpretation returns UNKNOWN.
  • Generated a compact repository-readable view from existing knowledge records and the current research registry. Its source hashes and dates distinguish historical review claims from current registry snapshots.
  • Designed scoped authenticated external read access; no public endpoint or write connector was deployed.
  • Verified the generated export hashes against Git blob bytes, correcting Windows newline normalization before delivery.

Measurement and improvement integration

Retrieval measured; interpretation failed
  • Initial development baseline: 16/44 specified positive evidence recovered. Record-sized projections over the same underlying knowledge material: 39/44. Query-bearing excerpts: 40/44.
  • The interpreted run recovered 40/44 evidence targets, with a 0.3045 precision lower bound and roughly 742 ms mean retrieval latency.
  • No valid model interpretations completed in the 180-second budget. All labels were UNKNOWN; 30 percent label agreement reflects abstention, not working status reasoning. Token usage was unavailable.
  • The suite uses the existing append-only experiment ledger, and invariant tests are included in existing pipeline qualification. Research-continuity evidence receipt 39 records the outcome; no ready experiment or promotion was created.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused and integration regressions passed 71 tests passed across shared memory, typed memory, GBrain launcher/bridge, agent/product paths, pipeline qualification and evaluation ledger.
Raw evidence integrity passed 868 of 868 eligible records verified; zero hash failures. Tampering rejection is covered by a regression test.
Historical benchmark partial 50 public development cases; 44 positive targets and 6 negative controls. 40 positive targets recovered. Gold annotations require independent review and relevance judgments are incomplete.
Interpretation gate failed Zero valid local-model assessments within 180 seconds; end-to-end milestone not demonstrated.
Journal tests and build passed npm run test:hiro and npm run build passed for the timestamped entry, including generated-page validation. Final publication metadata is regenerated below.

Current state

Next steps