Hiro development journal

Repairing stale candidate churn and ranking scheduled RSI research

Repair implemented, qualified, and running live Machine-readable JSON

Executive summary

The non-promotion interval was caused by a structural queue defect, not by Qwen model selection or a lack of improvement ideas. Historical interaction-audit records lacked the full synthetic prompt needed to reproduce their failures, yet remained eligible for candidate construction and were repeatedly revived after unrelated repository revisions.

The repair adds an explicit replay-completeness invariant and an append-only migration. It superseded 239 prompt-incomplete historical audit records and left zero such records actionable. Artifact-blocked work is now reconsidered only after a candidate-builder compatibility change, rather than after every Git revision.

Twelve research signals from the user's scheduled ChatGPT research conversation were reduced into six locally defined, testable mechanisms and admitted to the ranked queue: evaluator blind spots, measured-null calibration, strategy reconsideration, recursive iteration, security drift, and algorithmic invention.

The implementation passed 77 focused tests, all 718 repository tests, and five exact-revision pipeline qualification cycles. Hiro restarted on the qualified revision with Qwen 3.8 27B at a 16,384-token context window. Its first post-repair audit candidate used the complete synthetic prompt, produced a real patch, ran the isolated replay, and received a bounded retry for one concrete failed assertion rather than failing from absent evidence.

Work completed

Remove invalid historical work from the active queue

Completed
  • Added a replay-completeness predicate for interaction-audit records. An audit replay is constructible only when it carries both the synthetic prompt and the user query needed to reproduce the observed failure.
  • Added an idempotent, append-only queue migration that supersedes prompt-incomplete automatic quality records, clears their scheduling fields, records an idea_superseded event, and preserves their prior evidence and event history.
  • Applied the migration to the live queue. It superseded 239 stale records; a direct database audit found zero prompt-incomplete records remaining in actionable states.
  • After migration and research admission, the queue contained five implemented, 72 queued, 43 rejected, and 270 superseded records. The former stale cohort can no longer consume candidate-construction capacity.

Make artifact reactivation causally relevant

Completed
  • Introduced candidate-builder-2026-08-22-context-v2 as an explicit builder compatibility version.
  • Artifact blocks now retain blocked_on_revision for provenance while using blocked_on_builder_version as the reactivation boundary.
  • A new source promotion or unrelated code commit no longer revives every old artifact. Reconsideration requires a builder compatibility change and a complete replay fixture.
  • Historical artifact records receive a legacy builder marker when reclassified, preserving bounded migration behavior without pretending that their original evidence was complete.

Convert scheduled research into ranked improvement work

Completed
  • Read the scheduled ChatGPT conversation as an external research source and extracted twelve cited arXiv signals published from August 15 through August 21, 2026.
  • Reduced the signals into six code-owned mechanisms with local hypotheses, intended behaviors, metrics, construction plans, and independent tests. No raw ChatGPT prose was sent to the candidate model or treated as an instruction.
  • Admitted six normal ranked records without bypass or make-next authority. Evaluator blind spots ranked at 14.96, measured-null calibration at 14.95, strategy reconsideration at 9.66, recursive iteration at 9.53, security drift at 6.07, and algorithmic invention at 5.76.
  • The source packet is content-addressed with SHA-256 be1519e21aacb72e74794ab112bcb05330872881937b6ace4e1f857bc2d07f85 and marks all imported material as untrusted external evidence requiring local corroboration.

Qualify and restart the live system

Completed
  • Committed the repair as revision 0200eab0681aa51f125c976388037c08c17c3898.
  • Ran five clean pipeline qualification cycles against that exact revision. All five passed, authorizing live queue advance.
  • Restarted Hiro with the checked-in Windows launcher, which preserves the inherited runtime while normalizing the Path/PATH collision before spawning child processes.
  • Hiro returned HTTP 200 on the task and benchmark services. Qwen remained healthy as qwen/qwen3.8-27b on port 8080 with a 16,384-token context window.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused repair and integration suite passed 77 tests passed in 6.61 seconds across continuous-engine, external-source, and evaluation-dashboard coverage.
Full Hiro repository suite passed 718 tests passed in 384.75 seconds. Two warnings reported existing unregistered hiro_contract marks; no test failed.
Exact-revision pipeline qualification passed Five of five cycles passed for revision 0200eab0681aa51f125c976388037c08c17c3898. The qualification packet is qualification-20260823T010623Z.json with SHA-256 63fa1b0afda299cc06a24074431b6843872fbb353a5d77b1195fc2e8833527ba.
Live historical migration passed 239 prompt-incomplete audit records were superseded append-only, and zero prompt-incomplete records remained actionable.
Scheduled research admission passed Twelve research signals were consolidated into six ranked queue records with ordinary authority, measurable metrics, and local corroboration requirements.
Live service and model health passed Hiro's task and benchmark endpoints returned HTTP 200; Qwen 3.8 27B returned healthy at a 16K context window.
First post-repair candidate cycle passed The first selected audit candidate stored the complete 170-character synthetic schedule prompt, query, task contract, required facts, and malformed baseline output. Qwen generated a patch and ran its isolated test; the test caught that the proposed answer lost 'p.m.' after '4:00'. The queue scheduled a bounded retry and advanced to the next causally complete candidate.

Current state

Next steps