{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T09:57:16-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Canary restart behavior is durable but not truly pausable",
  "publicationStatus": "Diagnosed and validated; no runtime change",
  "executiveSummary": [
    "Hiro can be stopped while a continuous-improvement candidate is in its eight-hour canary without losing the queue record, completed checkpoints, probe results, frozen candidate metadata, or candidate worktree. Those records are persisted in the continuous-improvement SQLite ledger and external candidate workspace.",
    "A restart on the same active Git revision resumes the canary. The scheduler computes checkpoint eligibility from the original started_at timestamp, so shutdown time counts as elapsed canary time. This is restart recovery, not a true pause: overdue checkpoints run after restart rather than moving the canary end time forward.",
    "If Hiro's active revision changes while it is stopped, unfinished candidate, testing, and canary work built against the old revision is cleared and audibly requeued for reproduction instead of being promoted against a stale base.",
    "The active canary is pre-promotion isolated work. Stopping Hiro during its waiting period does not require rolling back the active branch because the candidate has not yet been fast-forwarded into Hiro."
  ],
  "workstreams": [
    {
      "title": "Durable canary state",
      "status": "Confirmed",
      "details": [
        "The canary payload stores its original start time, required checkpoints at 0, 60, 240, and 480 minutes, completed checkpoints, probe results, and frozen candidate metadata in the queue ledger.",
        "Canary records are deliberately excluded from the watchdog that requeues ordinary investigating, candidate, and testing states after 120 minutes of inactivity.",
        "On every scheduler tick the next incomplete checkpoint is selected from persisted state and compared with wall-clock elapsed time since the original start."
      ]
    },
    {
      "title": "Stop and restart behavior",
      "status": "Confirmed",
      "details": [
        "The scheduler normally ticks once per minute. While the process is stopped, no probes, evaluations, discovery, interaction audits, queue transitions, or promotions run.",
        "After restart on the same revision, the next due canary probe runs and the canary continues from its previously completed checkpoint list.",
        "If more than one checkpoint became due during downtime, the engine handles the next incomplete due checkpoint on a tick and preserves the rest for later ticks.",
        "Candidate construction interrupted outside canary is recoverable: work idle for more than 120 minutes is returned to the queue with an incremented attempt and an auditable stalled-state event."
      ]
    },
    {
      "title": "Revision and transaction boundaries",
      "status": "Confirmed with one narrow caveat",
      "details": [
        "At the start of a queue tick, Hiro compares unfinished candidate metadata with the active branch revision. Mismatched work is requeued, its candidate and canary payloads are cleared, and a stale_revision_requeued event is appended.",
        "The candidate remains outside the live Hiro branch throughout the canary. Only after all checkpoints pass does the stable governor rerun the full suite and perform a fast-forward promotion.",
        "There is a narrow non-atomic interruption window between the governor's Git fast-forward and the subsequent queue transition to implemented. The current continuous path does not expose a coordinated operator pause that waits for transaction boundaries."
      ]
    }
  ],
  "decisions": [
    "Describe current behavior as restart-resumable, not pause-resumable, because downtime is included in elapsed canary time.",
    "Do not recommend an abrupt kill while a queue transaction is actively constructing, probing, or promoting; stop during a visible canary-wait interval whenever possible.",
    "Treat an active revision change as invalidating unfinished candidate evidence and require a rebuild from the new base.",
    "Do not claim continuous observation during downtime: no probes execute while Hiro is stopped."
  ],
  "validation": [
    {
      "check": "Canary wait persistence behavior",
      "status": "passed",
      "result": "The focused test confirmed that a persisted canary returns canary_waiting with a finite next checkpoint derived from its original start timestamp."
    },
    {
      "check": "Concurrent queue behavior during canary wait",
      "status": "passed",
      "result": "The focused test confirmed that the waiting canary remains active while the queue may investigate another item."
    },
    {
      "check": "Stale revision recovery",
      "status": "passed",
      "result": "The focused test confirmed that a canary built from a different active revision is requeued, has stale candidate data cleared, and receives an auditable stale_revision_requeued event."
    },
    {
      "check": "Focused restart-behavior suite",
      "status": "passed",
      "result": "All three selected continuous-engine tests passed in 0.51 seconds. No Hiro process or queue state was changed during diagnosis."
    }
  ],
  "currentState": [
    "The continuous canary mechanism survives process restarts when the active revision is unchanged.",
    "There is no explicit operator pause command or paused-duration accounting in the current policy or engine.",
    "Stopping Hiro pauses execution but not the canary's wall clock.",
    "No implementation changes were made in this diagnostic session."
  ],
  "limitations": [
    "The deadline_at field records an eight-hour-five-minute boundary for visibility, but the continuous canary advancement path does not currently reject or extend a canary solely because a checkpoint was late after downtime.",
    "A long shutdown can therefore allow late probes to satisfy checkpoints even though Hiro was not continuously available for observation during the full window.",
    "The continuous promotion's Git fast-forward and queue implemented record are sequential rather than one atomic cross-system transaction, leaving a narrow abrupt-termination edge case.",
    "A controlled pause-at-boundary operation has not yet been implemented."
  ],
  "nextSteps": [
    "Design an explicit pause request that stops new transactions, waits for the current atomic boundary, persists paused_at, and shifts all remaining checkpoint deadlines by the accumulated paused duration.",
    "Add startup reconciliation for the case where the active branch already equals the canary candidate revision but the queue did not record implementation before shutdown.",
    "Decide whether long downtime should extend the canary, restart it from checkpoint zero, or reject it as insufficient continuous evidence; encode that choice in policy and tests.",
    "Expose paused, draining, safe-to-stop, and resumed states on the benchmark dashboard before presenting pause as an operator feature."
  ]
}
