{
  "schemaVersion": 2,
  "date": "2026.08.11",
  "publishedAt": "2026-08-11T10:49:03-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Hiro's restart-safe autonomous improvement workflow enters live Stage 6C probation",
  "publicationStatus": "Implementation, live terminal-path validation, two timing corrections, final exact-policy reactivation, and clean service restart completed",
  "executiveSummary": [
    "Hiro now has one durable autonomous workflow that advances safe external ideas through local corroboration, frozen specification, bounded candidate construction, public and held-out evaluation, Stage 6C probation, and an explicit terminal outcome.",
    "The workflow is append-only, idempotent, lease-protected, restart-safe, retry-aware, and monitored for nonterminal records without an executable next action.",
    "The exact Stage 6C behavior-fragment policy was revalidated on the final implementation revision and activated for one eight-hour window. Per-candidate user approval is disabled inside that narrow policy, while post-action reporting remains required.",
    "Live execution exposed a final-checkpoint boundary defect before retention. The active candidate was failed closed, exact prechange bytes were restored, and the lineage reached a real rolled-back terminal outcome. A second concurrent candidate was also rolled back during maintenance quiescence.",
    "The next authority-and-rollback candidate exposed a second timing defect: a long evaluation allowed the workflow's original tick timestamp to precede actual candidate application, making a checkpoint early. That candidate was also rolled back with exact restoration, and transition timestamps are now refreshed after long-running work.",
    "The final revision passed 498 repository tests and a fresh exact-revision campaign, the eight-hour policy was reactivated, and Hiro restarted cleanly. All 10 known lineages are now terminal with zero stuck records; the scheduler is ready to advance newly discovered safe ideas."
  ],
  "workstreams": [
    {
      "title": "Durable lineage state machine",
      "status": "Implemented and running",
      "details": [
        "A SQLite event ledger records append-only workflow transitions and rejects event updates or deletions.",
        "Each external or local improvement has one lineage identity, current state, terminal flag, next action, event history, retry metadata, affected files, candidate revision, and gate evidence binding.",
        "The transition runner uses a singleton lease and idempotency keys so repeated scheduler ticks and restarts cannot duplicate candidates or decisions.",
        "Transition failures are preserved with a bounded error description, retry count, and next retry time; successful recovery resumes from the incomplete state."
      ]
    },
    {
      "title": "Prompt-safe local corroboration",
      "status": "Implemented and exercised",
      "details": [
        "External material remains curiosity-only input. The autonomous workflow receives prompt-safe themes and references rather than raw external prose.",
        "Local corroboration uses bounded metadata derived from public evaluation failures and local response envelopes. It does not persist response bodies or held-out prompts in lineage receipts.",
        "A theme advances with one reproducible public evaluation failure or two independent local signals. Unsupported themes end as explicit rejections instead of silently disappearing.",
        "Corroboration receipts, improvement specifications, candidate fragments, and gate evidence are frozen with SHA-256 bindings for later audit."
      ]
    },
    {
      "title": "Candidate construction and evaluation",
      "status": "Implemented and exercised across passing and failing live candidates",
      "details": [
        "Autonomous construction is restricted to the schema-validated assistant-behavior fragment governed by Stage 6C. Arbitrary code, new filesystem scope, network actions, and external actions remain outside the authorization.",
        "Candidates are compared with the active baseline on both public and locally held-out suites using deterministic offline model calls.",
        "Admission requires public, held-out, invariant, behavior-regression, and latency gates, plus zero false allows.",
        "Only one fragment candidate may be active at a time. Passing candidates enter probation; a live latency-and-freshness candidate failed its gates and closed as rejected without changing the fragment."
      ]
    },
    {
      "title": "Stage 6C activation and probation",
      "status": "Active probation",
      "details": [
        "The activation helper accepts only the authorized policy identifier and exact policy checksum, a ready campaign bound to the current Git revision, and an absent disable sentinel.",
        "The final campaign completed 30 correct shadow decisions out of 30, zero false allows, 10 hot-reload rehearsals, 3 rollback drills, 5 interruption-boundary rehearsals, and a passing full repository suite.",
        "The enablement packet expires after eight hours, permits at most one active candidate, removes per-candidate approval, and requires post-action reporting.",
        "Three live probation attempts were deliberately failed closed and rolled back with exact byte restoration while integration timing defects were diagnosed. No candidate was falsely represented as retained.",
        "The final checkpoint now validates the still-probationary controller and fragment hash binding directly while the regular prompt loader continues to stop applying an expired probationary fragment. The existing five-minute lateness guard still fails closed."
      ]
    },
    {
      "title": "Scheduler, recovery, and observability",
      "status": "Implemented and live",
      "details": [
        "Workflow transitions and probation checkpoints run on every one-minute scheduler poll; the longer discovery interval no longer delays already-known work.",
        "The normalized hidden launcher supplies the Stage 6C runtime flag while preserving the existing packet and controller checks, so the flag alone grants no authority.",
        "Hiro was stopped and restarted through the hidden launcher after probation began. Maintenance later used a disable sentinel and a stopped service to prevent scheduler mutation during shared-repository tests.",
        "After the application-time correction and final campaign, Hiro restarted successfully again; task, benchmark, lineage, and pipeline endpoints returned successfully with all known lineages terminal and no stuck record.",
        "The benchmark lineage view now shows workflow state, next action, retry timing, affected files, candidate revision, test gates, evidence binding, and event history. A recovered failure remains in history but no longer appears as the current blocker."
      ]
    }
  ],
  "decisions": [
    "Treat retained, rejected, rolled back, or explicitly blocked as the only completed lineage states.",
    "Run state transitions independently of discovery cadence so reporting or a quiet discovery cycle cannot strand active work.",
    "Allow autonomous changes only within the exact Stage 6C assistant-behavior fragment policy; terminate broader code candidates with an explicit policy-scope decision rather than implying they were implemented.",
    "Use one exact-revision activation packet for the eight-hour window and continue requiring all candidate gates, atomic replacement, frozen rollback bytes, and scheduled checkpoints.",
    "Preserve historical failures and recovery events in the ledger so successful retries do not erase operational evidence.",
    "Quiesce the live scheduler with a fail-closed sentinel before running tests that share the active fragment path.",
    "Anchor probation and retry timing to the completion of long-running transitions, especially model evaluation and actual controller application, rather than to the scheduler tick's initial timestamp."
  ],
  "validation": [
    {
      "check": "Focused workflow, activation, launcher, dashboard, controller, and validation tests",
      "status": "passed",
      "result": "The focused sets passed before activation, including a regression test for the live controller gate adapter."
    },
    {
      "check": "Full repository regression suite",
      "status": "passed",
      "result": "498 tests passed on the final implementation revision with Hiro quiesced to prevent concurrent fragment mutation."
    },
    {
      "check": "Exact-revision Stage 6C validation campaign",
      "status": "passed",
      "result": "30 of 30 shadow decisions were correct with zero false allows; all hot-reload, rollback, interruption, and repository-test requirements passed."
    },
    {
      "check": "Live terminal canary",
      "status": "passed",
      "result": "A synthetic unauthorized/unsupported lineage reached a terminal rejected outcome without attempting promotion."
    },
    {
      "check": "Live lineage terminal outcomes and corrected probation",
      "status": "passed",
      "result": "Three live candidates reached explicit rolled-back terminal outcomes with exact byte restoration during defect discovery and maintenance. A latency-and-freshness candidate failed gates and closed as rejected. Unsupported bounded themes and the rejection canary also ended explicitly rather than becoming stuck."
    },
    {
      "check": "Restart recovery and service health",
      "status": "passed",
      "result": "After the final hidden restart, the task, benchmark, lineage, and pipeline endpoints returned successfully; the workflow reported 10 terminal records, zero nonterminal records, and zero stuck records."
    }
  ],
  "currentState": [
    "The exact Stage 6C policy is active for a fresh bounded eight-hour window on the final implementation revision.",
    "All 10 known lineages are terminal: rolled back after safe live diagnosis, rejected by candidate gates, rejected for unsupported bounded mappings, or completed as the synthetic rejection canary.",
    "There are zero nonterminal records, zero stuck records, and no active candidate; the tracked fragment is restored to its baseline bytes.",
    "Newly discovered safe ideas can enter the same autonomous path without per-candidate approval. Future probation checkpoints remain scheduled at 0, 15, 60, 360, and 480 minutes; a failed or excessively late checkpoint restores exact prechange bytes."
  ],
  "limitations": [
    "No live candidate from this session was retained: each probationary candidate was rolled back when the audit found a process defect or maintenance required quiescence.",
    "Final retention requires at least 40 live response observations and no critical response errors during probation.",
    "Autonomous arbitrary-code promotion remains intentionally outside the active Stage 6C policy; such candidates receive an explicit scope rejection pending a separately authorized broader policy.",
    "The process requires explicit maintenance quiescence before repository-wide tests that share the live fragment file; this session demonstrated and documented that operating constraint."
  ],
  "nextSteps": [
    "Allow the scheduler to discover and advance the next safely reduced and locally corroborated idea without manual intervention.",
    "For the next passing candidate, run the 0-, 15-, 60-, 360-, and 480-minute checkpoints from the actual controller application time.",
    "Retain a future fragment only if the final checkpoint has sufficient live observations and no critical failures; otherwise restore the exact prechange bytes and record rollback as the terminal outcome.",
    "Continue post-action reporting for every retained, rejected, rolled-back, or explicitly blocked lineage."
  ],
  "disclosureNote": "This public entry omits credentials, raw external posts, private prompts and responses, held-out test content, machine-local paths, and actionable details about unresolved security weaknesses."
}
