{
  "schemaVersion": 2,
  "date": "2026.08.04",
  "publishedAt": "2026-08-04T11:26:01-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Implementing and benchmarking Hiro Stage 6A shadow decisions",
  "publicationStatus": "Validated and published",
  "executiveSummary": [
    "Implemented Hiro's Stage 6A shadow-only decision engine and ran the required 30-decision allow/deny benchmark without enabling automatic mutation.",
    "The engine reads the disabled draft policy, classifies hypothetical evidence-only promotions, and writes only process-safe append-only SQLite decisions and simulated phase events. It has no Git, service, scheduler, deployment, candidate-materialization, or active-branch operation.",
    "The frozen campaign completed 30 valid externally labeled decisions with 30 correct classifications: five hypothetical allows and 25 denials, for 100 percent accuracy with zero false allows and zero false denies.",
    "Five controlled interruption boundaries resumed idempotently without duplicate events, and four competing child processes converged on one durable decision and five events.",
    "Stage 6 remains disabled. The host-level persistent disable sentinel is present, policy status remains draft_disabled with enabled false, Git HEAD was unchanged, and automatic promotion readiness remains false."
  ],
  "workstreams": [
    {
      "title": "Shadow-only policy evaluator",
      "status": "Implemented and passed",
      "details": [
        "Added a standalone evaluator that accepts normalized replay manifests and computes would-allow, would-deny, or invalid-infrastructure outcomes without consulting the externally supplied expected label.",
        "The evaluator refuses any policy that is not a disabled draft or that grants active-branch or runtime mutation authority.",
        "It validates user-bound policy identity, enablement timing, repository and branch identity, exact base revision, worktree state, leases, rate limits, cooldowns, both kill switches, evidence completeness, packet integrity, and restart/external-action declarations.",
        "It enforces literal add-only paths, regular nonexecutable UTF-8 files, count/size/line limits, opportunity-key binding, Python AST restrictions for regression tests, deterministic JSONL structure, and safe documentation content."
      ]
    },
    {
      "title": "Append-only decision and transition ledger",
      "status": "Implemented and passed",
      "details": [
        "Added SQLite decision and phase-event tables protected by no-update and no-delete triggers.",
        "Decision identity is bound to a canonical input SHA-256. Exact retries return the existing record; reuse with different evidence is rejected.",
        "Five shadow phase events represent future preflight, staging, pre-fast-forward, post-fast-forward, and probation boundaries. Every event is explicitly marked simulation-only with zero mutation, restart, or external action.",
        "A real four-process contention test showed all writers returning the same classification while the ledger retained one decision and five unique events."
      ]
    },
    {
      "title": "Externally labeled 30-decision benchmark",
      "status": "Passed",
      "details": [
        "Created exactly 30 deterministic replay cases: five allowed evidence-only shapes and 25 denied operation, path, content, repository-state, authorization, concurrency, rate, or integrity cases.",
        "Denied probes covered modification, deletion, rename, mode change, symlink, binary content, traversal, policy self-modification, stale base, dirty worktree, expired enablement, lease collision, rate limit, both kill switches, packet tampering, submodule, file-count overflow, runtime path, restart, external action, unsafe test code, invalid JSONL, dependent judging evidence, and active probation.",
        "The evaluator classified all 30 correctly: five would-allow and 25 would-deny, with zero false allows, zero false denies, and no infrastructure exclusions.",
        "The durable ledger contains 30 decisions and 150 unique decision/event pairs, exactly five events per decision."
      ]
    },
    {
      "title": "Interruption and idempotency exercise",
      "status": "Passed",
      "details": [
        "Independently interrupted one otherwise eligible decision after each of the five shadow phase boundaries.",
        "Every retry resumed from append-only evidence, completed with the same would-allow result, and retained exactly one event per phase.",
        "Completed-decision reuse returned the frozen decision without adding events, while changed evidence under an existing decision identifier was rejected."
      ]
    },
    {
      "title": "Disabled-state preservation",
      "status": "Passed",
      "details": [
        "Created the persistent Stage 6 DISABLED sentinel with a plain statement that only shadow evaluation is allowed.",
        "The policy remains draft_disabled and enabled false; both shadow and evidence-only active-branch mutation fields remain false.",
        "The benchmark recorded zero active-branch mutations, promotions, service restarts, schedule changes, or external actions, and the repository HEAD stayed unchanged."
      ]
    }
  ],
  "decisions": [
    "Treat would-allow as a hypothetical classification only. It is not authorization and never invokes a mutation operation.",
    "Keep expected benchmark labels outside the evaluator input so classification accuracy measures policy behavior rather than label leakage.",
    "Use a separate Stage 6 shadow ledger instead of the Stage 5 integrator so the new engine cannot accidentally inherit Git or candidate-materialization authority.",
    "Classify model or infrastructure outage as invalid-infrastructure and exclude it from policy accuracy rather than counting it as a correct behavioral denial or a successful decision.",
    "Preserve the first frozen benchmark result and its SHA-256 companion instead of rewriting campaign history.",
    "Pass only the classification portion of the activation benchmark. Do not claim activation readiness while real rollback drills and mutation safety components are absent."
  ],
  "validation": [
    {
      "check": "Focused Stage 6 policy and shadow suite",
      "status": "passed",
      "result": "18 tests passed in 3.79 seconds on the final rerun."
    },
    {
      "check": "Repository-wide regression suite",
      "status": "passed",
      "result": "303 tests passed in 149.54 seconds."
    },
    {
      "check": "Thirty-decision classification campaign",
      "status": "passed",
      "result": "30 of 30 valid decisions were correct: five hypothetical allows and 25 denials, with 100 percent accuracy, zero false allows, zero false denies, and zero infrastructure exclusions."
    },
    {
      "check": "Append-only ledger audit",
      "status": "passed",
      "result": "The durable ledger contained 30 decisions, 150 events, and 150 unique decision/event pairs; no update or delete is permitted by its triggers."
    },
    {
      "check": "Five interruption boundaries",
      "status": "passed",
      "result": "Preflight, staging, pre-fast-forward, post-fast-forward, and probation-boundary interruptions all resumed exactly without duplicate events."
    },
    {
      "check": "Competing-process logging",
      "status": "passed",
      "result": "Four child processes concurrently evaluated the same input and converged on one durable decision and five phase events."
    },
    {
      "check": "Frozen packet integrity",
      "status": "passed",
      "result": "Benchmark packet SHA-256 00400ec385a98a8c5392b57b23db7ddf23278085128adc9da80036721db3c5e1 verified."
    },
    {
      "check": "Mutation and activation boundary",
      "status": "passed",
      "result": "Policy and host sentinel remained disabled; Git HEAD was unchanged; active mutations, promotions, restarts, schedule changes, and external actions were all zero."
    },
    {
      "check": "Hiro journal generation and frontend build",
      "status": "passed",
      "result": "Timestamped-entry unit tests passed; the generator produced and validated 62 journal pages, and the TypeScript and Vite production build completed successfully."
    }
  ],
  "currentState": [
    "Stage 6A shadow classification is implemented and its 30-decision accuracy gate has passed.",
    "The engine can durably explain hypothetical allow/deny results, distinguish infrastructure invalidity, resume five simulated boundaries, and tolerate competing writers.",
    "Stage 6 automatic promotion remains disabled in the policy and by the persistent host sentinel.",
    "Stage 6B activation readiness is false because the real mutation, canary/probation, and rollback components and drills do not yet exist."
  ],
  "limitations": [
    "The benchmark cases are deterministic historical/adversarial replays, not live active-branch promotions.",
    "The five interruption points simulate future transaction boundaries; no Git fast-forward or rollback occurred.",
    "No production canary or 24-hour probation monitor exists yet.",
    "The required three real rollback drills, including overlapping-state emergency stop, have not run.",
    "The current policy is a disabled draft and has no user approval bound to a final activation hash.",
    "This result does not establish consciousness, unrestricted autonomy, or permission for self-deployment."
  ],
  "nextSteps": [
    "Implement the future 6B transaction, canary/probation monitor, and history-preserving rollback executor behind an adapter that is structurally forced to shadow mode.",
    "Use sealed fixture repositories to run at least three real rollback drills, including one clean revert, one post-fast-forward interruption, and one overlapping-state emergency stop.",
    "Verify that every drill preserves append-only evidence, never uses reset or force update, and leaves the real Hiro active branch untouched.",
    "Review the resulting false-deny behavior and operational evidence, then freeze a final policy version and hash.",
    "Request a separate explicit user approval bound to that final policy hash before enabling any evidence-only active-branch mutation."
  ],
  "disclosureNote": "This public entry contains no credentials, tokens, private held-out cases or expected answers, personal data, or actionable details about unresolved security weaknesses."
}
