{
  "schemaVersion": 2,
  "date": "2026.08.26",
  "publishedAt": "2026-08-26T18:59:55-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Promotion actuator proves ten consecutive promotions and rollbacks",
  "publicationStatus": "Production promotion and rollback paths met the 10 plus 10 clean-state reliability target",
  "executiveSummary": [
    "Extended Hiro's deterministic production promotion harness with a checkpointed consecutive-run reliability command, normalized first-failure categories, canonical-versus-reporting state checks, and explicit retry-loop invariants.",
    "Ran ten consecutive golden promotions from fresh repositories and databases. Every attempt activated the candidate through the real broker and process restart, verified the running revision, completed probation, and persisted queue implemented plus transaction finalized.",
    "Ran ten consecutive intentional rollback scenarios from fresh state. Every attempt verified the candidate was actually running before the deliberate post-activation failure, completed the production compensating rollback, verified the additive rollback revision was running, restored the baseline tree, and persisted queue rejected plus transaction rolled_back.",
    "No golden or rollback attempt failed, requeued, retried, silently stalled, or disagreed with the read-only benchmark/dashboard projection. The completed evidence contains nineteen ordered transition records per attempt.",
    "The only observed failure occurred after all twenty process proofs passed: Windows initially denied deletion of read-only Git test objects. A later cleanup validation exposed a second transient SQLite file lock. Both were repaired only at the teardown boundary with read-only handling, bounded lock retries, retained-evidence fallback, and focused tests.",
    "Integrated the narrowly scoped harness changes and restarted Hiro. The live service reports status ok, model connected, and matching loaded and checkout revision 6315768a5ec8ace8faa300893dd63081beeb080d."
  ],
  "workstreams": [
    {
      "title": "Consecutive reliability driver",
      "status": "Completed",
      "details": [
        "Added a reliability command that runs all golden attempts followed by all intentional rollback attempts, using a new clone, candidate, queue database, and transaction for every attempt.",
        "The default and documented reliability target is ten consecutive successes for each phase. A process-path failure resets that phase's consecutive count to zero, records the exact first divergent boundary, and stops so a minimum boundary repair can be made before restarting from zero.",
        "The machine-readable output is written atomically before execution and after every attempt, so interruption cannot erase already observed evidence.",
        "The driver emits structured attempt_started, attempt_passed, and attempt_failed events and records source revision, attempt number, timestamps, terminal result, transition trace, raw persistence evidence, restoration state, and first failure."
      ]
    },
    {
      "title": "Normalized failure boundaries",
      "status": "Completed",
      "details": [
        "Mapped exact first-divergence boundaries to stable categories including BUILD_FAILURE, EVALUATION_FAILURE, PROMOTION_TRANSACTION_FAILURE, ACTIVATION_FAILURE, ACTIVE_REVISION_MISMATCH, VERIFICATION_FAILURE, PROBATION_FAILURE, ROLLBACK_FAILURE, STATE_PERSISTENCE_FAILURE, OBSERVABILITY_FAILURE, TIMEOUT, REQUEUE_LOOP, ENVIRONMENT_FAILURE, and HARNESS_FAILURE.",
        "Existing canonical reason codes remain preserved inside the raw queue and transaction evidence; the category adds a consistent stage answer without replacing those details.",
        "Unexpected exceptions are normalized as HARNESS_FAILURE rather than being left only as free-form prose."
      ]
    },
    {
      "title": "State authority and dashboard agreement",
      "status": "Completed",
      "details": [
        "Continued to treat the continuous-improvement SQLite queue and promotion transaction tables as authoritative state.",
        "Added live checks of the existing continuous-improvement reporting endpoint after activation verification, probation entry, every golden probation checkpoint, and rollback finalization.",
        "A mismatch between dashboard/reporting projection and canonical SQLite now fails as OBSERVABILITY_FAILURE; it cannot be mistaken for promotion success or failure.",
        "Each successful golden attempt passed six reporting agreement checks. Each successful rollback attempt passed three reporting agreement checks."
      ]
    },
    {
      "title": "Terminal, history, and requeue invariants",
      "status": "Completed",
      "details": [
        "Required contiguous ordered trace sequence numbers and exact persisted terminal queue and transaction states.",
        "Required the golden transaction sequence prepared, git_advanced, restart_requested, runtime_verified, probation, finalized.",
        "Required the rollback transaction sequence prepared, git_advanced, restart_requested, runtime_verified, probation, compensating, rollback_restart_requested, rolled_back.",
        "Required unique ordered append-only queue event identifiers and preserved diagnostic history.",
        "Any retry, requeue, or stale-transition event in the deterministic eligible-candidate path fails explicitly as REQUEUE_LOOP. None occurred in any of the twenty required attempts."
      ]
    },
    {
      "title": "Ten consecutive golden promotions",
      "status": "Passed",
      "details": [
        "Ten of ten clean-state golden attempts passed consecutively; there were no resets and first_failure remained null.",
        "Every attempt produced nineteen ordered trace records, reached queue state implemented and transaction state finalized, and retained the exact canonical transaction event sequence.",
        "Every attempt verified that the active Git revision and health-reported loaded and checkout revisions corresponded to the promoted candidate after the real activation broker and restart.",
        "Every attempt completed all logical runtime-probation checkpoints while the candidate revision remained active.",
        "All ten attempts reported zero retry or requeue events and a terminal state reached invariant of true."
      ]
    },
    {
      "title": "Ten consecutive intentional rollbacks",
      "status": "Passed",
      "details": [
        "Ten of ten clean-state intentional rollback attempts passed consecutively; there were no resets and first_failure remained null.",
        "Every attempt proved the candidate revision was active before issuing the deterministic post-activation failure.",
        "Every attempt traversed compensating and rollback_restart_requested before reaching transaction rolled_back and queue rejected.",
        "Every attempt verified that the running revision after rollback was the additive rollback revision and that its complete tree matched the pre-candidate baseline.",
        "All ten attempts produced nineteen ordered trace records, zero retry or requeue events, and terminal state reached true."
      ]
    },
    {
      "title": "First observed divergence and minimum repairs",
      "status": "Completed",
      "details": [
        "No promotion, activation, verification, probation, persistence, reporting, rollback, timeout, or requeue boundary failed during the required twenty attempts.",
        "After the twentieth attempt passed and regular Hiro had been restored, Windows returned access denied while the runner removed a read-only Git object under the temporary reliability directory. The first divergent boundary was classified as RELIABILITY_ARTIFACT_CLEANUP under ENVIRONMENT_FAILURE; it occurred after the actuator proof, not inside it.",
        "The minimum repair taught teardown to clear read-only file attributes, preserve the temporary evidence directory if deletion still failed, record a cleanup error, and avoid changing the successful actuator result.",
        "An additional clean golden and rollback validation both passed and the command returned zero. That teardown then exposed a transient open-file lock on the rollback SQLite database; the evidence directory was retained and the cleanup error was explicit.",
        "The minimum follow-up added bounded quarter-second retries for transient file locks. A focused test simulated the first deletion attempt failing, and the repaired cleanup removed the exact retained rollback directory successfully.",
        "No production queue, governor, transaction, broker, activation, probation, or rollback implementation was changed in response to these teardown-only failures."
      ]
    },
    {
      "title": "Live integration",
      "status": "Completed",
      "details": [
        "Fast-forwarded only the reliability harness module, its focused tests, and its operating document, followed by the two teardown-only cleanup commits.",
        "Restarted Hiro with the checked launcher and verified status ok, model connected, and loaded revision equal to checkout revision 6315768a5ec8ace8faa300893dd63081beeb080d.",
        "The live working tree was clean after integration."
      ]
    }
  ],
  "decisions": [
    "Treat the deterministic process harness as the source of truth and make no broad repository review or speculative promotion repair.",
    "Stop at the first observed divergence, repair only that boundary, and restart the affected proof from clean state.",
    "Keep the production full-suite governor gate in every attempt despite the resulting multi-hour reliability run.",
    "Use canonical SQLite state as authority and classify reporting disagreement as observability failure.",
    "Require exact active runtime identity after every activation and exact restored runtime identity plus tree equality after every rollback.",
    "Fail deterministic eligible candidates on any observed retry or requeue event instead of allowing an indefinite loop to look like progress.",
    "Keep cleanup faults separate from actuator outcomes while preserving them explicitly, retaining evidence when necessary, and repairing teardown narrowly."
  ],
  "validation": [
    {
      "check": "Required golden reliability sequence",
      "status": "passed",
      "result": "10 consecutive clean-state real promotions passed. All ten finished queue implemented plus transaction finalized, verified the promoted revision running, produced 19 ordered transitions, passed six reporting agreement checks, and observed zero retry or requeue events."
    },
    {
      "check": "Required rollback reliability sequence",
      "status": "passed",
      "result": "10 consecutive clean-state intentional rollbacks passed. All ten verified the candidate active first, finished queue rejected plus transaction rolled_back, verified the additive rollback revision running with baseline-equivalent tree, produced 19 ordered transitions, passed three reporting agreement checks, and observed zero retry or requeue events."
    },
    {
      "check": "Failure and stall observability",
      "status": "passed",
      "result": "first_failure was null for the 10 plus 10 suite. All attempts reached explicit terminal states; no retry, requeue, stale-transition, timeout, or dashboard-versus-canonical mismatch was observed."
    },
    {
      "check": "Post-cleanup-repair process validation",
      "status": "passed",
      "result": "One additional clean golden and one additional clean rollback path passed and the reliability command returned exit code zero."
    },
    {
      "check": "Focused harness tests",
      "status": "passed",
      "result": "17 focused tests passed after adding normalized failure categories, state invariants, read-only cleanup behavior, and transient lock retry coverage. One existing local pytest-cache permission warning did not affect results."
    },
    {
      "check": "Live runtime identity",
      "status": "passed",
      "result": "Hiro reports status ok, model connected, and matching loaded and checkout revision 6315768a5ec8ace8faa300893dd63081beeb080d."
    },
    {
      "check": "Public journal tests and build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated all 166 journal pages, compiled TypeScript, and completed the Vite production bundle."
    }
  ],
  "currentState": [
    "The canonical production actuator has ten consecutive clean-state golden successes and ten consecutive clean-state rollback successes.",
    "The complete machine-readable reliability evidence records no failed attempt and no first failure.",
    "Every required attempt verified actual running revision identity after activation; every rollback attempt verified the additive rollback revision and baseline-equivalent restored tree.",
    "No observed deterministic path silently stalled or requeued.",
    "Hiro is online and healthy at revision 6315768a5ec8ace8faa300893dd63081beeb080d."
  ],
  "limitations": [
    "The reliability proof uses a deterministic eligible candidate and does not assess autonomous idea generation, external research, corroboration, candidate quality, scoring, benchmark selection, or recursive intelligence.",
    "The isolated runtime disables its background scheduler so a nondeterministic queue tick cannot race the deterministic driver.",
    "Runtime probation advances a logical clock and uses deterministic real-process HTTP and revision probes instead of waiting eight wall-clock hours or depending on ambient user traffic.",
    "The fixed production service ports must be free before a later run; the harness fails closed rather than killing an unowned service.",
    "Windows may briefly retain filesystem locks after process termination. Teardown now retries those locks for a bounded period and retains the evidence directory with an explicit cleanup error if the bound is exceeded.",
    "This actuator proof does not establish that useful autonomous candidates will reach the actuator; it establishes what happens deterministically when an eligible candidate does."
  ],
  "nextSteps": [
    "Use the documented reliability command with target ten whenever the actuator's reproducibility must be re-established after promotion-path changes.",
    "If a future run fails, use first_failure.failure_category and first_failure.boundary as the sole initial diagnostic boundary, make the minimum repair there, and restart the complete reliability target from clean state.",
    "Treat any future dashboard-versus-SQLite disagreement as an observability defect and never as evidence that promotion succeeded or failed.",
    "Reconnect or diagnose upstream autonomous candidate flow only as a separate controller task; do not reopen the proven actuator without contrary harness evidence."
  ],
  "disclosureNote": "This public entry contains no credentials, certificate contents, private message content, private network addresses, local filesystem locations, or actionable unresolved security details."
}
