{
  "schemaVersion": 2,
  "date": "2026.08.25",
  "publishedAt": "2026-08-25T08:43:43-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Hiro's improvement factory now separates pipeline health from candidate merit",
  "publicationStatus": "Core process implemented, qualified, and activated; independent inspection found critical completion gaps",
  "executiveSummary": [
    "Completed and activated a simplified continuous-improvement process centered on reproducible failures, harness-owned tests, isolated candidates, paired evaluation, timed canary, independent promotion, and automatic rollback.",
    "Candidate construction failures no longer become negative verdicts about the underlying idea. Exhausted build attempts move to artifact_blocked with idea_merit_evaluated false and become eligible again only after a relevant builder change.",
    "Valid comparative evaluation failures remain rejected evidence. Security regressions remain terminal, infrastructure failures receive cooldown retries, and canary or governor failures retain their own evidence and revision paths.",
    "The candidate builder now mechanically normalizes harmless generated whitespace and common edit-shape mistakes, uses the exact harness-owned production replay instead of allowing a candidate to author its own approval test, and rejects repeated failed patch semantics.",
    "A durable supervisor watches the improvement factory itself. Recursive model-authored builder-repair tasks were retired; a bounded run of construction failures now stops queue manufacture and exposes a diagnostic incident without consuming idea merit.",
    "Corrected a watchdog defect discovered during implementation: legitimate evaluation rejections no longer count as candidate-builder failures or move the pipeline toward a false global shutdown.",
    "Corrected a Windows linked-worktree defect discovered by release qualification. Qualification pytest roots now anchor at the short common Git checkout instead of inheriting an arbitrarily long release-worktree path.",
    "Preserved all eight previously promoted commits on the live branch. The final release is a fast-forward from the live revision rather than a replacement or history rewrite.",
    "The exact activated revision a11954ef646e7f43404af7909678073af701edae passed 750 repository tests and five consecutive ten-test end-to-end qualification cycles. Hiro restarted healthy with Qwen 3.8 connected and the existing durable queue intact.",
    "A separate read-only whole-code inspection then found that the system is not yet a complete autonomous production loop: candidate execution needs stronger operating-system isolation, the current timed checks are synthetic rather than product-traffic probation, promotion is not transactionally joined to runtime restart verification, and post-promotion rollback monitoring is not wired."
  ],
  "workstreams": [
    {
      "title": "Explicit outcome separation",
      "status": "Implemented and active",
      "details": [
        "The live queue distinguishes implemented, evaluation-rejected, security-rejected, hypothesis-disproven, infrastructure-blocked, superseded, and artifact-blocked outcomes.",
        "Artifact-blocked means patch or test manufacture failed and idea merit was not evaluated. These records are not included in the ordinary rejected count.",
        "Evaluation rejection requires a valid candidate and evidence from public and held-out comparisons. Construction text, patch errors, inconclusive baseline contrast, and malformed test evidence cannot masquerade as an idea verdict.",
        "Infrastructure faults use a separate retry circuit breaker and do not consume the candidate's bounded engineering-attempt budget."
      ]
    },
    {
      "title": "Reproducible behavior before implementation",
      "status": "Implemented and active",
      "details": [
        "The rotating everyday interaction audit exercises thirty-two questions across sixteen assistant categories and retains a deterministic replay fixture and task contract for every admitted failure.",
        "Stochastic one-off audit results enter an observing state. A second independent observation is required before the automatic audit admits a patch-authorizing queue item.",
        "Legacy single-sample audit records are superseded rather than allowed to consume current candidate capacity.",
        "Markdown emphasis is removed for semantic phrase checks so presentation punctuation does not create false improvement incidents."
      ]
    },
    {
      "title": "Harness-owned candidate evidence",
      "status": "Implemented and active",
      "details": [
        "For captured interaction failures, the model proposes only production changes. The candidate builder replaces model-authored replay tests with a deterministic test generated from the pre-candidate code-owned fixture.",
        "The identical target test must collect on the untouched baseline, reach a marked assertion, fail there by AssertionError, and pass only when the production-reachable boundary changes.",
        "Tests that merely classify or log the bad response do not qualify as a response improvement.",
        "Repairs start from a clean baseline and receive the exact prior failure plus a bounded excerpt of the failed implementation, while semantic duplicates are mechanically rejected."
      ]
    },
    {
      "title": "Candidate construction reliability",
      "status": "Implemented and active",
      "details": [
        "Generated text files are normalized mechanically before validation so trailing whitespace does not consume an autonomous repair attempt.",
        "A common multiline insert-before mistake is converted into an exact modify operation while still requiring the old block to resolve in the isolated workspace.",
        "Every pytest phase receives an isolated, bounded basetemp outside deep candidate paths, preventing cross-run cache contamination and Windows nested-path failures.",
        "Candidate, evaluator, integration, governor, canary, and qualification test roots use purpose-specific short runtime locations."
      ]
    },
    {
      "title": "Independent evaluation and promotion",
      "status": "Implemented and active",
      "details": [
        "Candidates run paired public and held-out evaluation, independent regression tests, security comparison, and targeted baseline contrast before Stage 5 integration.",
        "Only the exact changed-file manifest verified by frozen Stage 5 evidence reaches the promotion governor; an allowlist is not treated as proof that every permitted file changed.",
        "Low-risk and moderate-risk canaries retain their authorized risk-based checkpoints. The stable governor rechecks the complete suite and permits only a fast-forward from the unchanged candidate baseline.",
        "Rollback creates an additive Git revert and preserves the failed post-promotion evidence."
      ]
    },
    {
      "title": "Improvement-factory health controller",
      "status": "Implemented and active",
      "details": [
        "The durable supervisor records candidate episodes, strategies, failure fingerprints, exact evidence, and candidate-ready receipts in append-only tables.",
        "A bounded no-canary watchdog now counts construction failures for one builder compatibility revision. Evaluation rejections are explicitly excluded because they prove the evaluator is working, not that candidate construction is broken.",
        "When the construction budget is exhausted, the active artifact is parked without a merit verdict and queue manufacture stops with diagnose_candidate_pipeline as the required action.",
        "The Observatory now displays pipeline health, the construction-failure count and threshold, and whether diagnosis is required."
      ]
    },
    {
      "title": "Release reconciliation and activation",
      "status": "Completed",
      "details": [
        "The original workspace contained an unrelated uncommitted benchmark experiment. It was preserved and excluded from the process release rather than accidentally committed.",
        "The live .hd3 checkout contained eight promoted commits beyond the initial workspace revision. A clean integration branch was based on that live head so none of those improvements was discarded.",
        "The final live update was a fast-forward from 4c15149 to a11954e. The durable queue database, local configuration, credentials, model runtime, and historical artifacts remained in place.",
        "The exact live checkout received the frozen qualification packet for a11954e and an exact safe-directory entry required by its existing Windows ownership boundary."
      ]
    }
  ],
  "decisions": [
    "Use the simple measurable cycle: real failure or bounded idea, reproducible contract, isolated candidate, comparative evaluation, timed canary, governor retain-or-revert.",
    "Do not interpret inability to manufacture a valid patch as evidence that the idea is bad.",
    "Do not weaken comparative evaluation, security, canary, full-suite, protected-path, or rollback boundaries to increase throughput.",
    "Let promising ideas receive materially different bounded implementations, but stop recursive repair when the factory itself is unhealthy.",
    "Keep candidate-authored production code separate from harness-owned approval evidence.",
    "Count only construction-stage failures in the global builder-health watchdog; valid negative evaluations must not halt the factory.",
    "Preserve live Git lineage and runtime state during activation.",
    "Expose pipeline health and failure taxonomy directly in the Observatory instead of collapsing them into a generic rejection count."
  ],
  "validation": [
    {
      "check": "Focused state-machine and release tests",
      "status": "passed",
      "result": "The combined release passed 79 focused supervisor, queue, Observatory, API, and linked-worktree qualification tests."
    },
    {
      "check": "Full repository regression suite",
      "status": "passed",
      "result": "The exact combined revision passed 750 tests in 461.77 seconds. Four generated autonomous tests emitted pre-existing unknown hiro_contract marker warnings; there were no failures."
    },
    {
      "check": "Frozen five-cycle pipeline qualification",
      "status": "passed",
      "result": "Revision a11954ef646e7f43404af7909678073af701edae passed five consecutive cycles of ten end-to-end checks, for 50 successful lifecycle checks. The frozen packet is qualification-20260825T154151Z.json with SHA-256 7157a0153cf4f1dbfc59cfa7ae50f4bb8f17aa4429fec7d4b9b9ecfb0fd0eb84."
    },
    {
      "check": "Windows linked-worktree qualification",
      "status": "passed after repair",
      "result": "The first detached release qualification exposed Git's administrative-path limit. A new regression test proves qualification now anchors pytest under the primary common checkout, and all five subsequent cycles passed."
    },
    {
      "check": "Live fast-forward and restart",
      "status": "passed",
      "result": "The clean live checkout fast-forwarded from 4c15149 to a11954e. The checked hidden launcher restarted Hiro in 3.406 seconds without restarting Qwen."
    },
    {
      "check": "Live health and qualification",
      "status": "passed",
      "result": "Ports 8000, 8001, and 8765 are owned by the new Hiro process. Health is OK, Qwen 3.8 27B is connected, the live revision matches the qualified revision, and the pipeline watchdog is not blocked."
    },
    {
      "check": "Durable queue preservation",
      "status": "passed",
      "result": "The post-restart ledger retained 487 ideas, 35 actionable records, seven implementations, nine artifact-blocked records, and the existing scheduled infrastructure retry."
    },
    {
      "check": "Independent whole-code inspection",
      "status": "completed with critical findings",
      "result": "A separate read-only inspector reviewed the exact activated revision across architecture, security boundaries, candidate construction, evaluation, promotion, runtime lifecycle, concurrency, persistence, observability, legacy code, and tests. It produced a severity-ordered remediation plan with five critical system-boundary findings and eleven high-priority findings."
    },
    {
      "check": "Public journal timestamped-entry tests",
      "status": "passed",
      "result": "npm run test:hiro passed in the clean journal checkout."
    },
    {
      "check": "Public journal production build",
      "status": "passed after installing locked dependencies",
      "result": "The initial clean-checkout build generated and validated all 160 journal pages but could not find the uninstalled local TypeScript compiler. After npm ci installed the package-lock dependencies, npm run build completed the generator, timestamp and alias validation, TypeScript build, and Vite production bundle successfully."
    }
  ],
  "currentState": [
    "Hiro is running the qualified revision a11954ef646e7f43404af7909678073af701edae from the clean .hd3 deployment checkout.",
    "Hiro reports healthy and Qwen 3.8 27B is connected.",
    "The continuous pipeline is qualified and its construction watchdog is healthy at nine construction failures against a threshold of twelve.",
    "The final live check found 37 actionable records: 36 waiting and one infrastructure retry. Source ingestion continued adding bounded ideas during publication, and no item was forced ahead of its recorded cooldown.",
    "The Observatory separately displays implemented, evidence-rejected, artifact-blocked, superseded, active, waiting, retrying, and pipeline-health information.",
    "The independent inspection is complete. The queue and candidate-evidence process are materially improved, but the system should not yet be described as a complete production self-improvement loop until execution isolation, transactional runtime activation, and automatic recovery are implemented."
  ],
  "limitations": [
    "This architecture can autonomously measure and integrate bounded code improvements; it cannot guarantee that the local model will invent a successful implementation for every worthwhile idea.",
    "The construction watchdog is a circuit breaker, not an automatic meta-repair mechanism. When it opens, the correct action is to diagnose the factory with independent engineering evidence.",
    "Two independent audit observations reduce stochastic false positives but can delay automatic admission of an intermittent user-visible failure.",
    "The generated autonomous-test marker warnings remain non-fatal but should be cleaned up so qualification output stays high-signal.",
    "The unrelated experimental benchmark edit remains uncommitted in the original workspace and was intentionally excluded from the live release.",
    "Candidate worktrees isolate source revisions but do not yet provide a complete operating-system security sandbox for generated code execution.",
    "The current timed candidate checks are synthetic worktree probes. They do not yet prove behavior under shadow or live product traffic.",
    "Git promotion, hidden service restart, loaded-revision verification, post-promotion monitoring, and automatic rollback are not yet one durable idempotent transaction."
  ],
  "nextSteps": [
    "First make status reporting truthful: label current timed checks as synthetic soak, distinguish every infrastructure and transaction failure from merit rejection, and expose stable versus loaded runtime revision.",
    "Build a restricted candidate execution boundary with explicit environment, network, filesystem, process, memory, and time limits before expanding autonomous promotion.",
    "Create an idempotent promotion transaction covering prepared intent, Git fast-forward, hidden restart, runtime revision verification, smoke tests, finalization, post-promotion monitoring, and compensating rollback.",
    "Strengthen state and concurrency with versioned database migrations, durable expiring activity leases, heartbeating runner ownership, one active supervisor episode, and supervised background tasks.",
    "Move causal test ownership to the evaluator for external ideas and replace file-existence corroboration with mechanism-specific local experiments.",
    "Add a four-level real-assistant laboratory covering output boundaries, core agent turns with fake tools, disposable API sessions, and browser conversation/crash behavior.",
    "After the system-boundary work, improve source-claim diversity, retire overlapping legacy schedulers and hidden dashboard sections, split monolithic modules, and add progressive cached evaluation."
  ],
  "disclosureNote": "This public entry contains no credentials, private conversation text, private source content, hidden reasoning, or actionable unresolved security details."
}
