{
  "schemaVersion": 2,
  "date": "2026.08.27",
  "publishedAt": "2026-08-27T21:00:03-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Phase 3A-R demonstrates trustworthy reproduction execution",
  "publicationStatus": "Evidence acquisition works; thirteen experiments completed validly and seven produced admissible supporting evidence",
  "executiveSummary": [
    "Qualified only Hiro's ability to execute the fourteen accepted Phase 3C experiment plans and turn their observations into claim-specific evidence. Candidate construction, production changes, promotion, activation, rollback, corroboration semantics, claim extraction, discovery ranking, and meta-improvement were unchanged.",
    "Froze all fourteen source claims, provenance spans and hashes, Hiro-local hypotheses, baselines, treatments, measurements, acceptance and falsification criteria, asset plans, resource estimates, sample strategies, and complete experiment specifications before observing any result.",
    "Built a declarative reproduction runner that separates experiment code from production candidates. Synthetic fixtures and hidden ground truth are code-owned; local-model subjects receive neither ground truth nor the other experimental condition.",
    "Executed numerical calculation inside disposable WSL2/Bubblewrap workspaces with OS isolation and network denial. A typed loopback bridge supplied synthetic cases to the fixed local model where a plan required model behavior, without exposing source text, credentials, arbitrary code, or production-write authority.",
    "Thirteen plans were experimentally executable and all thirteen completed with valid, independently recalculated measurements. Seven satisfied their preregistered local criteria and six validly did not. Negative results were preserved as successful evidence acquisition rather than experiment failures.",
    "One plan was classified EXPERIMENT_INVALID before build because it attempted to attribute failures away from model size while providing only one model and no independent model-size comparator. The plan and reason remain preserved rather than being weakened.",
    "The first execution attempt for the planner-handoff experiment failed before measurement because a batched model response did not preserve the frozen case identities. The immutable v1 failure remains recorded; a minimal preregistered v2 transport repair requested one response per unchanged case and then completed validly.",
    "The unchanged corroboration truth table produced seven PASS, zero FAIL, and seven DEFER decisions. Every PASS and two valid DEFER results have complete source-to-decision provenance audits.",
    "The final disposition is PHASE 3A-R DEMONSTRATED — EVIDENCE ACQUISITION WORKS. Seven evidence packages are eligible as inputs to a later separately authorized candidate-construction phase, but no candidate was constructed or promoted."
  ],
  "workstreams": [
    {
      "title": "Immutable preregistration",
      "status": "Completed",
      "details": [
        "The accepted fourteen-plan Phase 3C feasibility report was used as the only qualification corpus; no replacement plan or new discovery result was selected.",
        "Each preregistration preserves the exact source claim, supporting source location and hashes, local transfer hypothesis, baseline, treatment, measurement, acceptance criterion, falsifier, required assets, expected resource use, isolation controls, and sample policy.",
        "Every preregistration passed the existing evidence-package contract, which rejects missing fields and any result written before execution.",
        "The frozen manifest SHA-256 is 944f9942daaadc2715832f6ad7fb6684526c0d8da4e654ba582a09029c81ac5a."
      ]
    },
    {
      "title": "Experiment-only construction",
      "status": "Implemented",
      "details": [
        "Added bounded code-owned protocols for planner handoff, failure attribution, resource scheduling, planning complexity, structured artifacts, residual verification, misleading-premise behavior, answer selection, judge rarity, repository-context test generation, and dense-context reliability.",
        "Protocols use declarative fixtures and measurements. They cannot write Hiro production code, create a candidate, request promotion, or inherit authority from source material.",
        "The experiment worker is copied into a disposable research workspace and consumes only the frozen specification and typed subject observations.",
        "Safe assertion experiments accept only a narrow Python assertion grammar and run in the OS sandbox. Other answer-quality experiments compare against hidden deterministic expected values rather than an unrestricted LLM evaluator."
      ]
    },
    {
      "title": "Execution and integrity validation",
      "status": "Completed",
      "details": [
        "Thirteen experiment builds were attempted and all thirteen built successfully. All thirteen executable plans completed and passed integrity validation.",
        "Validation checked baseline and treatment execution, exact case counts, experiment and attempt identity, preregistration and specification hashes, unchanged acceptance criteria, raw-artifact identity, and deterministic result recalculation.",
        "Every numerical result was calculated in the isolated worker and independently recalculated from raw rows by the host validator. All calculations matched.",
        "The local model was treated only as a subject. Synthetic prompts excluded hidden answers and the other condition; model identity and telemetry were retained.",
        "No experiment downloaded an asset, required a resource reclassification, or touched production."
      ]
    },
    {
      "title": "Observed plan defect",
      "status": "Preserved as invalid",
      "details": [
        "One failure-cause plan proposed distinguishing evidence, tool-use, and constraint failures from model-size or reasoning-length effects.",
        "The supplied assets contained one local model and no independently varied model-size condition, making the causal comparison unidentifiable.",
        "The plan was frozen, classified EXPERIMENT_INVALID with reason PLAN_DEFECT_MODEL_SIZE_COMPARATOR_UNAVAILABLE, and sent to corroboration as unavailable evidence, which correctly deferred it.",
        "No weaker hypothesis, substitute metric, or post-result plan was created."
      ]
    },
    {
      "title": "First execution divergence and versioned repair",
      "status": "Repaired",
      "details": [
        "Planner-handoff attempt v1 failed before experiment execution because one batched treatment response omitted or duplicated a case identity.",
        "The original failure is immutable in the first qualification report, whose SHA-256 is 548c2d51b17220fdc4212240b8b48500c44db20d6f26108e74b840eff6ded5ea.",
        "The minimum v2 repair changed only response transport from one batch to one schema-bound response per frozen case. Cases, sample count, prompts, hidden ground truth, model configuration, acceptance criterion, falsifier, and preregistration hash remained unchanged.",
        "The repair record was frozen before v2 execution, explicitly denies outcome-based rerunning, and has SHA-256 5181092c9b6b806f551dbe1e5da6d2e58f5b618bf7a9d7eae9784b1ab99dcb6e.",
        "Attempt v2 completed validly and met the preregistered criterion."
      ]
    },
    {
      "title": "Reproduction outcomes",
      "status": "Completed",
      "details": [
        "Seven experiments supported their local hypotheses: planner/executor handoff, causal graph attribution, resource-aware scheduling, structured artifact handoff, residual-guided verification, misleading-premise answer correction, and combined frequency/quality selection.",
        "Six valid experiments did not support their local hypotheses: graph versus flat model attribution, simple versus multi-step planning, misleading-premise routing, judge reliability by candidate rarity, repository-context unit-test generation, and dense-context reliability.",
        "Several non-supporting experiments produced 100 percent accuracy in both conditions. The system retained the zero-delta result and deferred corroboration instead of manufacturing an improvement.",
        "There were no opposite-direction effects crossing a preregistered threshold, so REPRODUCTION_CONTRADICTED remained zero. There were no inconclusive, build-failed, final runtime-failed, or resource-reclassification outcomes."
      ]
    },
    {
      "title": "Existing corroboration integration",
      "status": "Completed without semantic changes",
      "details": [
        "Valid supporting evidence plus the independently established healthy regression context mapped to PASS/CORROBORATED.",
        "Valid non-supporting evidence mapped to DEFER/INSUFFICIENT_EVIDENCE rather than being misreported as contradiction.",
        "The invalid plan mapped to DEFER/EVIDENCE_UNAVAILABLE. No result mapped to FAIL because no valid opposite-direction evidence or attributable regression occurred.",
        "Final counts were seven PASS, zero FAIL, and seven DEFER.",
        "Nine complete provenance audits cover every PASS and two valid DEFER results. Each chain records source, claim, hypothesis, preregistration, experiment implementation, raw observations, calculated result, criterion, interpretation, integrity status, and corroboration decision."
      ]
    }
  ],
  "decisions": [
    "Use only the fourteen frozen Phase 3C plans and never stop after finding five positive outcomes.",
    "Treat a valid negative result as evidence-acquisition success and distinguish it from experiment failure.",
    "Reject unidentifiable causal comparisons before execution rather than weakening the frozen hypothesis.",
    "Keep model subjects separate from deterministic ground truth and never use the tested judge as the sole evaluator of its own output.",
    "Run calculations in the existing WSL2/Bubblewrap isolation mechanism and use only a typed local-model bridge for synthetic subject responses.",
    "Preserve the first malformed-response attempt and implement a new versioned transport attempt without changing methodology or criteria.",
    "Apply the existing corroboration truth table exactly; do not modify its vocabulary or thresholds.",
    "Stop at admissible evidence. Do not construct or promote production candidates."
  ],
  "validation": [
    {
      "check": "Focused Phase 3A-R and upstream evidence tests",
      "status": "passed",
      "result": "21 tests passed. One non-failing warning concerned the inaccessible pytest cache directory."
    },
    {
      "check": "Pre-execution regression context",
      "status": "passed",
      "result": "820 repository tests passed in 421.99 seconds before experiment execution. Six non-failing warnings concerned five existing unregistered marks and the inaccessible pytest cache directory."
    },
    {
      "check": "Final full Hiro regression suite",
      "status": "passed",
      "result": "821 tests passed in 418.60 seconds after the transport repair and evidence verifier were complete, with the same six non-failing warnings."
    },
    {
      "check": "Frozen evidence-package verification",
      "status": "passed",
      "result": "The verifier rehashed all thirteen valid evidence packages, raw measurement artifacts, calculated result artifacts, the source manifest, and all seven PASS audits. It returned zero reason codes. The final report SHA-256 is 423a41156b6f15c1cba465df6bda904bc6c1813bea734eaec3f04ab1c4c2820d."
    },
    {
      "check": "Final reproduction funnel",
      "status": "passed",
      "result": "14 plans frozen; 13 builds attempted and successful; 13 executed and valid; 7 supported; 6 not supported; 0 contradicted; 0 inconclusive; 1 invalid plan; 0 final build/runtime failures; 0 resource reclassifications; corroboration 7 PASS, 0 FAIL, 7 DEFER."
    },
    {
      "check": "Public journal tests and build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 173 journal pages, compiled TypeScript, and completed the Vite production bundle. The clean checkout first required npm ci; installation reported one existing high-severity dependency advisory, which was not automatically modified because dependency maintenance was outside this session."
    }
  ],
  "currentState": [
    "Phase 3A-R implementation is committed at revision 0708ab550aa99b007371ebd314f66bf1eb274904.",
    "Thirteen of fourteen accepted plans now have valid executed evidence; the remaining plan has an explicit preserved methodological defect.",
    "Seven evidence packages are admissible under the unchanged corroboration semantics and may become inputs to a later separately authorized candidate-construction phase.",
    "Six valid non-supporting packages remain durable negative evidence and seven total outcomes are deferred.",
    "No candidate exists from these results, and no production code, promotion state, or active runtime was changed by any experiment."
  ],
  "limitations": [
    "The experiments validate bounded Hiro-local transfer hypotheses, not the source papers' full external benchmarks or exact published effect sizes.",
    "Model-based conditions used one fixed local model at temperature zero and one preregistered observation per case; broader model-family generalization was not tested.",
    "The typed local-model bridge uses the existing loopback model service outside the network-denied calculation sandbox. It carries only synthetic prompts and answers and has no arbitrary-code or production-write authority.",
    "Some deterministic synthetic protocols create deliberately discriminating cases. Subsequent candidate qualification must establish benefit on independent Hiro evaluations before any production consideration.",
    "A PASS decision makes evidence eligible for later candidate construction; it does not prove that a production implementation will improve Hiro or pass promotion gates."
  ],
  "nextSteps": [
    "Do not automatically construct candidates. Obtain separate authorization for the next phase.",
    "If candidate construction is authorized, consume only the seven audited PASS evidence packages and preserve their exact claim, hypothesis, criterion, and scope.",
    "Keep all six valid non-supporting results available to prevent rediscovering and retesting the same unsupported local hypotheses without new evidence.",
    "Require any future corrected version of the invalid failure-cause experiment to add a genuine independent comparator and create a new preregistration rather than modifying this record."
  ],
  "disclosureNote": "This public entry contains no credentials, private source text, private interaction content, local filesystem locations, private network addresses, or actionable unresolved security details."
}
