{
  "schemaVersion": 2,
  "date": "2026.08.27",
  "publishedAt": "2026-08-27T22:59:25-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Phase 3D demonstrates evidence-to-candidate qualification",
  "publicationStatus": "Evidence-to-candidate conversion works; two independently supported findings produced promotion-eligible candidates without promotion or activation",
  "executiveSummary": [
    "Qualified only Hiro's ability to convert the seven accepted Phase 3A-R supporting evidence packages into bounded, non-meta candidate implementations. Discovery, claim extraction, reproduction evidence, corroboration semantics, governor policy, promotion thresholds, activation, and Phase 2 promotion machinery were not changed.",
    "Froze all seven evidence inputs, transfer assessments, four candidate specifications, baseline measurements, development cases, acceptance thresholds, resource limits, prohibited scope, and a separate evaluator-owned holdout vault before autonomous candidate construction.",
    "Correctly concluded that three supported mechanisms did not justify production changes: planner handoff and structured artifacts were already present, while residual-guided verification lacked independently labeled residual authority.",
    "The autonomous builder attempted four low-risk candidates in disposable external worktrees. Three valid implementations were constructed; one concurrency candidate never satisfied its frozen construction contract.",
    "All three valid implementations passed deterministic implementation-fidelity assessment, preregistered target-gain evaluation, and frozen holdout evaluation. One premise-correction implementation then correctly failed its hard 60-line resource limit after changing 261 production lines.",
    "The causal failure-attribution and quality-aware memory-selection candidates each passed hard resource limits, the actual complete repository regression suite, and the existing governor's read-only validation boundary.",
    "Both eligible candidates have complete immutable provenance chains from source evidence through governor decision. Their exact candidate revisions remain isolated and unmerged.",
    "The qualification report verified with zero reason codes. Production HEAD was identical before and after qualification; zero promotion transactions and zero activations occurred.",
    "The final disposition is PHASE 3D DEMONSTRATED — EVIDENCE-TO-CANDIDATE WORKS. This session stopped at PROMOTION_ELIGIBLE and did not authorize or perform promotion."
  ],
  "workstreams": [
    {
      "title": "Immutable evidence and transfer assessment",
      "status": "Completed",
      "details": [
        "Verified the existing Phase 3A-R report and all seven authorized supporting evidence-package identities before qualification.",
        "Produced an explicit transfer assessment for every finding: demonstrated mechanism, Hiro relevance, possible target component, transfer rationale, assumptions, falsifier, smallest implementation, blast radius, and whether production modification was warranted.",
        "Preserved three SUPPORTED_BUT_NO_ACTIONABLE_IMPLEMENTATION outcomes instead of forcing every supported research result into code.",
        "No new finding was introduced and no negative or deferred Phase 3A-R result was substituted for an authorized input."
      ]
    },
    {
      "title": "Preregistered candidate and holdout protocols",
      "status": "Completed",
      "details": [
        "Froze four candidate specifications before implementation: causal failure attribution, bounded trip-search concurrency, grounded premise correction, and quality-aware memory selection.",
        "Each specification records its evidence identity, target capability and component, exact intended mechanism, baseline, metric, minimum gain, development cases, holdout design, permitted files, prohibited scope, resource limits, rollback expectation, and low-risk classification.",
        "Holdout cases were stored in a separate immutable vault before build. Expected holdout outputs were withheld from the candidate builder.",
        "Candidate acceptance thresholds were not lowered and holdout cases were not changed after results were observed."
      ]
    },
    {
      "title": "Isolated autonomous construction",
      "status": "Completed",
      "details": [
        "Used Hiro's actual candidate builder and local candidate agent in external disposable Git worktrees with path-scoped authority.",
        "Candidate source material was provided as structured validated evidence rather than executable instruction authority.",
        "The original candidate builds and target/holdout evaluation ran through the existing WSL2/Bubblewrap restricted executor with OS isolation and network denial.",
        "Immutable candidate packets preserve exact file snapshots, per-file hashes, candidate diffs, construction attempts, baseline contrast, scope audit, and targeted construction tests."
      ]
    },
    {
      "title": "Independent candidate evaluation",
      "status": "Completed",
      "details": [
        "Verified implementation fidelity structurally and behaviorally before treating measured improvement as meaningful; comments or claimed intent were never accepted as proof.",
        "All three valid implementations passed their frozen development and holdout thresholds. The causal-attribution candidate improved from 1/6 to 5/6 development cases and passed 4/5 fresh holdouts. Premise correction and memory selection each achieved their preregistered perfect target and holdout scores.",
        "Resource evaluation separated production complexity from candidate test code and measured latency at the operation boundary. Causal attribution changed 20 production lines and memory selection changed 12; both added no model calls, network destinations, or persistent writes.",
        "Premise correction stopped at RESOURCE_CONSTRAINT_FAIL because 261 production lines exceeded its fixed maximum of 60. Full regression and governor assessment correctly did not run for it."
      ]
    },
    {
      "title": "Complete regression and governor boundary",
      "status": "Completed",
      "details": [
        "Only the two candidates passing fidelity, target, holdout, and resource gates incurred the complete repository regression cost.",
        "Causal failure attribution completed with 837 tests passed and six non-failing warnings in 494.44 seconds. Quality-aware memory selection completed with 837 tests passed and six non-failing warnings in 465.11 seconds.",
        "The existing continuous governor validated both candidates under unchanged policy and returned PASS for each.",
        "The qualification runner invoked only the governor's read-only validation method. It contains no promotion call, created no promotion transaction, performed no activation, and did not restart Hiro."
      ]
    },
    {
      "title": "First-divergence repairs",
      "status": "Completed",
      "details": [
        "Preserved every failed run in a versioned immutable directory and repaired only the first observed qualification boundary before replay.",
        "Bound candidate identifiers to the existing 64-character contract; separated production/test resource accounting and microbenchmark noise; and supplied edit-shape guidance compatible with the existing one-patch-per-path contract.",
        "Made consolidation replay immutable per-file snapshots and include both tracked changes and untracked candidate tests in identity verification.",
        "A complete-suite replay exposed missing Windows process-contract values and a missing ignored pinned-runtime junction in fresh worktrees. The existing child-environment normalization and runtime provisioning path were minimally repaired; the exact two launcher tests passed before both complete suites were rerun successfully.",
        "No repair changed evidence, holdouts, candidate behavior requirements, acceptance criteria, corroboration semantics, governor policy, or promotion thresholds."
      ]
    }
  ],
  "decisions": [
    "Use only the seven Phase 3A-R CORROBORATION_PASS evidence packages.",
    "Treat supported research as evidence rather than automatic authority to change production.",
    "Freeze candidate specifications and evaluator-owned holdouts before autonomous construction.",
    "Require behavioral implementation fidelity before interpreting candidate metrics.",
    "Run complete regression only after target, holdout, and resource gates pass.",
    "Replay exact candidate snapshots on the current trust base instead of asking the model to regenerate successful implementations.",
    "Use the existing governor's validation boundary and stop at PROMOTION_ELIGIBLE.",
    "Preserve legitimate build and resource rejections rather than loosening thresholds to increase eligibility."
  ],
  "validation": [
    {
      "check": "Canonical Phase 3D report verification",
      "status": "passed",
      "result": "The verifier rehashed the report and both eligible provenance audits, returned valid true with zero reason codes, and confirmed two promotion-eligible candidates, zero promotion transactions, and zero activations. Final report SHA-256: 9e6c4b835a59390c8772945867b66d8fdd3f6e2b7ed4fc23cdebc0fbe7a1cc50."
    },
    {
      "check": "Final qualification funnel",
      "status": "passed",
      "result": "7 supported findings; 7 valid transfers; 3 no-action dispositions; 4 candidate specs and build attempts; 3 successful builds; 3 fidelity, target, and holdout passes; 2 resource passes and 1 resource failure; 2 full-regression passes; 2 governor passes; 2 promotion-eligible candidates."
    },
    {
      "check": "Focused qualification and upstream regression tests",
      "status": "passed",
      "result": "77 tests passed in 31.92 seconds. One non-failing warning concerned the inaccessible pytest cache directory."
    },
    {
      "check": "Eligible candidate complete repository suites",
      "status": "passed",
      "result": "Both exact candidate worktrees passed the actual complete suite: 837 tests in 494.44 seconds and 837 tests in 465.11 seconds, each with six non-failing warnings."
    },
    {
      "check": "Public journal tests and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 174 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session."
    }
  ],
  "currentState": [
    "Phase 3D qualification code and durable repository documentation are committed through revision 54b5f64.",
    "Two exact isolated candidates are PROMOTION_ELIGIBLE: causal failure attribution at b409f5df4841ecba69c8e25772f8480d9976de18 and quality-aware memory selection at 04c929ad3469f647febb7172897df65d3105ff4a.",
    "The trip-search concurrency candidate is terminally classified CANDIDATE_BUILD_FAILED. The premise-correction candidate is terminally classified RESOURCE_CONSTRAINT_FAIL.",
    "Three other supported findings have explicit no-action dispositions and require new production-gap evidence or authority before reconsideration.",
    "Production HEAD was unchanged throughout qualification. Neither eligible candidate is merged, active, or promoted."
  ],
  "limitations": [
    "Phase 3D demonstrates evidence-to-candidate conversion and evaluation, not safe autonomous promotion. A later phase would require separate authorization and must consume the exact eligible revisions rather than regenerate them.",
    "The causal-attribution candidate met but did not exceed its 0.80 holdout floor and still missed the timeout-first expectation in one development and one holdout case.",
    "The memory-selection implementation is correct on the frozen development and holdout partitions, but broader production traffic and scale behavior were not activated or observed in this phase.",
    "The concurrency candidate's build failure shows that autonomous construction is not uniformly reliable across all otherwise actionable specifications.",
    "The final complete regression must run on Windows to exercise real Windows launcher and pinned-runtime behavior; only candidate construction and bounded behavioral evaluation are Linux OS-isolated."
  ],
  "nextSteps": [
    "Do not promote or activate either eligible candidate without a separately authorized end-to-end autonomous-promotion qualification.",
    "If that qualification is authorized, accept only the two exact candidate IDs, revisions, and provenance chains recorded here.",
    "Retain the trip-search build failure and premise-correction resource failure as durable negative construction evidence rather than automatically regenerating them.",
    "Keep meta-improvement disabled until ordinary evidence-to-candidate-to-promotion behavior is independently demonstrated."
  ],
  "disclosureNote": "This public entry contains no credentials, private source text, private interaction content, local filesystem locations, private network addresses, or actionable unresolved security details."
}
