{
  "schemaVersion": 2,
  "date": "2026.08.09",
  "publishedAt": "2026-08-09T15:07:59-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Stage 6 evidence intake moves from idle to one real candidate",
  "publicationStatus": "Candidate ready; exact reactivation approval pending",
  "executiveSummary": [
    "Hiro's high-cadence evidence collection is now running every 30 minutes, and the missing upstream production lane can turn one completed proposal-only run into a fully validated Stage 6B documentation-evidence candidate.",
    "The work uncovered and repaired two previously hidden blockers that would have prevented any real promotion: the live service lacked the required Stage 6 metrics endpoint, and external Git worktrees could not run the full repository suite because several required dependencies and the pinned runtime were unavailable there.",
    "A fresh Daylab run produced 16 public evaluation observations with four failures, one improvement specification, four curiosity items, and 22 regression probes. The evidence pipeline preserved that imperfect result rather than filtering it for appearance.",
    "The resulting one-file Markdown evidence candidate passed the required Stage 6 suite and complete repository suite on both the clean active base and an external detached candidate worktree. Both monitoring snapshots also matched.",
    "All five source packets and the final candidate packet are immutable and their companion hashes verify. Exactly one candidate now waits in the Stage 6 inbox.",
    "Stage 6B remains disabled with its external sentinel present because the implementation and policy hashes changed. A replacement activation-review packet is ready for one exact user approval."
  ],
  "workstreams": [
    {
      "title": "High-cadence evidence collection",
      "status": "Running",
      "details": [
        "Reactivated the previously reviewed Stage 6 revision long enough to verify a healthy idle scheduled tick, then started the existing Daylab launcher at a 30-minute cadence.",
        "Moved new engineering into an external worktree so the active approved revision remained clean while discovery continued.",
        "Daylab now performs an explicitly enabled one-shot Stage 6 handoff after a completed run and records a host-local success marker so it cannot prepare a second candidate during the same activation."
      ]
    },
    {
      "title": "Automatic evidence intake",
      "status": "Implemented and exercised",
      "details": [
        "Added a production pipeline that derives a deterministic documentation-evidence note from an actual completed run summary.",
        "The pipeline runs the mandatory Stage 6 policy suite and full repository suite on the clean base, captures live health and metrics, validates the candidate in an external detached worktree, checks monitoring again, freezes the opportunity, baseline, candidate, Stage 4, and Stage 5 packets, and invokes the previously implemented materializer.",
        "The source summary is bound by SHA-256, and the output remains exactly one additive Markdown file with no runtime behavior change."
      ]
    },
    {
      "title": "Real Stage 6 monitoring",
      "status": "Live",
      "details": [
        "Added a local metrics endpoint that reports only observation count, error rate, and p95 latency from the latest 100 valid response envelopes.",
        "The endpoint exposes no prompt or response content and is paired with Hiro's existing local health endpoint.",
        "After restarting Hiro on the reviewed revision, health returned HTTP 200 and the metrics endpoint returned 100 observations with a measured 0.04 error rate and measured p95 latency."
      ]
    },
    {
      "title": "Hermetic external validation",
      "status": "Repaired",
      "details": [
        "A first clean-worktree full-suite attempt exposed that tracked tests imported a locally ignored benchmark package and productivity schema, while launcher tests expected the ignored pinned runtime inside every worktree.",
        "Tracked the required benchmark source and productivity migration, and added an exact runtime junction provisioner for external candidate worktrees.",
        "The provisioner is required in production and explicitly disabled only by controller fixtures that use a fake command runner.",
        "The clean external worktree subsequently passed the complete repository suite."
      ]
    },
    {
      "title": "Real candidate preparation",
      "status": "Candidate frozen",
      "details": [
        "The fresh public evaluation completed at a 0.75 pass rate with four of 16 observations failing and produced one diagnosis-oriented improvement specification.",
        "All 22 regression probes completed; one sports probe was degraded while the others passed.",
        "The first automatic attempt completed validation but failed during temporary-worktree cleanup because Windows would not remove a worktree containing the runtime junction. No packets were written by that failed attempt.",
        "The cleanup path was corrected to remove only the exact junction before Git worktree removal, covered by a regression test, committed, and retried against the same source summary.",
        "The retry succeeded and froze exactly one candidate in the disabled inbox."
      ]
    }
  ],
  "decisions": [
    "Increase throughput before the promotion boundary: frequent evidence collection and automatic packet production, while preserving the one-candidate and one-promotion limits.",
    "Allow documentation evidence to record an imperfect but valid run; a public-suite failure is useful evidence and should not be hidden merely to produce a positive-looking artifact.",
    "Require real local monitoring data rather than fixture values or invented baseline metrics.",
    "Treat any clean-worktree collection failure as a production blocker because the Stage 6 controller validates candidates in an external worktree.",
    "Keep the scheduler disabled whenever tracked implementation or policy identity changes, even during fast-path development.",
    "Require exact user approval for the final revised hash identities before consuming the waiting candidate."
  ],
  "validation": [
    {
      "check": "New pipeline and metrics focused tests",
      "status": "passed",
      "result": "The focused coverage passed, including incomplete summaries, failed baseline validation, occupied inbox, monitoring regression, one-shot behavior, bounded metrics, and Windows runtime-junction cleanup."
    },
    {
      "check": "Expanded integrated Stage 6 suite",
      "status": "passed",
      "result": "The integrated Stage 6 and related Daylab suite passed 77 tests before the final cleanup regression was added; the real candidate pipeline later passed the policy-defined 72-test subset twice on the exact final revision."
    },
    {
      "check": "Clean external repository suite",
      "status": "passed",
      "result": "After tracking missing dependencies and provisioning the pinned runtime, all 440 tests passed in the clean development worktree."
    },
    {
      "check": "Exact candidate base policy suite",
      "status": "passed",
      "result": "72 policy-defined Stage 6 tests passed in 43.25 seconds."
    },
    {
      "check": "Exact candidate base full suite",
      "status": "passed",
      "result": "441 tests passed in 128.50 seconds."
    },
    {
      "check": "Detached candidate policy suite",
      "status": "passed",
      "result": "72 policy-defined Stage 6 tests passed in 42.02 seconds."
    },
    {
      "check": "Detached candidate full suite",
      "status": "passed",
      "result": "441 tests passed in 129.65 seconds."
    },
    {
      "check": "Post-validation monitoring",
      "status": "passed",
      "result": "Both snapshots contained 100 observations with error rate 0.04 and p95 latency 10,176 ms; absolute and relative error-rate increases were zero."
    },
    {
      "check": "Packet integrity",
      "status": "passed",
      "result": "All five source packets, the candidate packet, and the replacement activation-review packet matched their companion hashes and were frozen read-only."
    }
  ],
  "currentState": [
    "The active Hiro revision is 7ba6f62475c53f900dba350d1e33a83a880e1c10 and the worktree is clean.",
    "The active Stage 6 policy hash is 09704eca80e2160f4305b01997deb8ab958146deadca6b1a7352ec59e770be08.",
    "One candidate waits in the inbox with packet hash d752ae3b8b39f2caa19d9089fc48c72be4fcd65ef96b1e204b5bd94cc5dee8f5.",
    "The candidate adds one documentation-evidence file and changes no runtime behavior.",
    "The replacement activation-review packet hash is b3b2a7f43ae14723cb301d9ed212d2000f59ee78236d8ac2eee4d07d851f37ff.",
    "Stage 6B is disabled, its persistent sentinel is present, and the earlier enablement packet does not match the new policy hash.",
    "Daylab remains running at a 30-minute cadence and the one-shot marker prevents another candidate preparation."
  ],
  "limitations": [
    "The waiting artifact records evidence; it does not correct the four newly observed public evaluation failures.",
    "The generated improvement specification still requires its own diagnosis and candidate workflow before any runtime correction could be considered.",
    "The Stage 6 candidate has not yet been fast-forwarded or entered its 24-hour probation because exact reactivation approval is pending.",
    "A best-effort Telegram briefing for the fresh run failed due to message-entity formatting; this did not affect evidence collection or packet integrity.",
    "The automatic intake lane is intentionally one-shot under the current activation and supports only a Markdown documentation-evidence artifact."
  ],
  "nextSteps": [
    "Obtain explicit user approval bound to the exact revision, policy hash, candidate hash, one-promotion maximum, and 72-hour authorization window.",
    "Replace the stale enablement packet, remove the sentinel, enable the scheduler, and observe the candidate through preflight, fast-forward, immediate canary, and probation checkpoints.",
    "Keep Daylab collecting evidence every 30 minutes without preparing a second candidate during this activation.",
    "Separately investigate the new reading-list evaluation failures using the generated diagnostic specification; do not conflate that runtime work with this documentation-only promotion.",
    "Repair the noncritical Telegram formatting issue without broadening Stage 6 authority."
  ],
  "disclosureNote": "This public entry omits private prompts and responses, held-out cases, credentials, local filesystem paths, model configuration details, and actionable unresolved security information."
}
