{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T09:59:46-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "The eight-hour canary is inherited policy, not useful current evidence",
  "publicationStatus": "Historical and architectural diagnosis; no policy change",
  "executiveSummary": [
    "The continuous governor originally authorized a 30-minute low-risk canary and a 120-minute moderate-risk canary. Commit 1d2964b on August 11 changed both classes to 480 minutes with checkpoints at 0, 60, 240, and 480 minutes while unifying interaction failures with continuous improvement.",
    "The eight-hour value came from the older Stage 6 post-promotion probation design, where a change was already live and the system needed time to collect service availability, error-rate, latency, and at least 40 compatible observations before retention.",
    "The current continuous engine uses the same duration before promotion while the candidate remains in an external worktree. Each checkpoint reruns a deterministic probe against that isolated worktree. Wall-clock waiting between identical probes supplies little additional evidence and blocks candidate construction and promotion throughput.",
    "The user's expectation of faster improvement is correct: the present eight-hour pre-promotion canary is a policy carryover that no longer matches the evidence being gathered."
  ],
  "workstreams": [
    {
      "title": "Policy provenance",
      "status": "Confirmed",
      "details": [
        "Git history shows the initial continuous governor used low-risk checkpoints at 0, 5, 15, and 30 minutes and a 30-minute duration.",
        "The same initial policy used a 120-minute moderate-risk duration.",
        "The interaction-failure unification changed both risk classes to one eight-hour schedule without adding a distinct evidence rationale for isolated pre-promotion candidates."
      ]
    },
    {
      "title": "Old versus current canary semantics",
      "status": "Diagnosed",
      "details": [
        "The older Stage 6 probation design fast-forwarded a verified change before probation and monitored the live revision for eight hours.",
        "Its design called for repository identity checks, service and model availability, error and p95 latency comparisons, and at least 40 compatible observations, making elapsed time relevant.",
        "The current continuous queue performs all targeted, public, held-out, invariant, category, and latency evaluation before canary, leaves the candidate outside the active branch, and reruns a stable worktree probe at scheduled times.",
        "Because the candidate is not serving production traffic, the current eight-hour wait does not accumulate real user exposure or live candidate telemetry."
      ]
    },
    {
      "title": "RSI throughput effect",
      "status": "Quantified",
      "details": [
        "The governor allows only one active canary. An eight-hour minimum therefore caps theoretical promotion throughput at three changes per day before construction and evaluation time.",
        "While a canary is present, the queue can investigate another idea but parks candidate construction and testing, so the wall-clock gate also slows experimentation upstream.",
        "The original 30-minute low-risk policy allowed a theoretical 48 low-risk decisions per day; the original two-hour moderate policy allowed 12, before other constraints."
      ]
    }
  ],
  "decisions": [
    "Do not defend the eight-hour duration merely because it is conservative; require each delay to identify the new evidence it produces.",
    "Separate isolated pre-promotion soak from live post-promotion observation rather than applying one inherited duration to both.",
    "Preserve independent tests, held-out evaluation, invariants, scope controls, rollback, and governor checks while removing evidence-free waiting.",
    "Make future canary duration risk- and evidence-based rather than a single constant for every candidate."
  ],
  "validation": [
    {
      "check": "Git provenance",
      "status": "passed",
      "result": "Blame and commit inspection identified 1d2964b as the change from 30/120 minutes to 480/480 minutes and showed the exact policy diff."
    },
    {
      "check": "Current probe semantics",
      "status": "passed",
      "result": "Code inspection confirmed that each continuous canary checkpoint runs the candidate worktree probe before governor promotion and records its result in the queue ledger."
    },
    {
      "check": "Legacy probation rationale",
      "status": "passed",
      "result": "The Stage 6 design explicitly associates eight-hour post-promotion probation with live health, error-rate, latency, identity, and observation-count evidence."
    }
  ],
  "currentState": [
    "The active continuous-governor policy still requires eight hours for both low- and moderate-risk candidates.",
    "The current canary remains pre-promotion and isolated from live user traffic.",
    "No policy, queue, canary, or running process was changed during this diagnosis."
  ],
  "limitations": [
    "Shortening the pre-promotion canary alone would not create post-promotion live monitoring or automatically restart Python services after a promoted code change.",
    "A faster policy needs explicit checkpoint evidence and risk classes; replacing eight hours with an arbitrary shorter constant would repeat the same design mistake.",
    "Theoretical daily throughput ignores model time, evaluation time, quiet-window constraints, and failures."
  ],
  "nextSteps": [
    "Design a fast isolated gate based on repeated probes rather than idle wall time, with a candidate target such as 15 to 30 minutes for low-risk changes and a longer evidence-triggered path for moderate risk.",
    "Add real post-promotion monitoring bound to the promoted revision, compatible observations, service restart state, and automatic rollback without unnecessarily blocking isolated evaluation of later candidates.",
    "Measure false-pass and rollback rates by risk class, then adapt checkpoint count and duration from observed evidence quality.",
    "Expose which evidence each waiting checkpoint is expected to add on the benchmark dashboard."
  ]
}
