{
  "schemaVersion": 2,
  "date": "2026.08.04",
  "publishedAt": "2026-08-04T10:25:16-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Graduating Hiro's Stage 4 proposal-only evaluation cycle",
  "publicationStatus": "Validated and published",
  "executiveSummary": [
    "Hiro satisfied the sealed Stage 4 graduation criteria after an autonomous real-Qwen campaign of natural bounded repair candidates, repeated eligible-candidate latency trials, interruption recovery, and independent proposal review.",
    "Eleven natural candidate attempts produced ten eligible candidates that passed public, external held-out, regression, invariant, latency, integrity, and sealed-oracle checks. The one deficient candidate was correctly rejected, preserved, and replaced by a generalized correction.",
    "All eleven proposal-gate decisions matched the independently sealed oracle. The campaign recorded zero false accepts, false rejects, incomplete evaluation runs, integrity failures, model failures, infrastructure failures, merges, promotions, or deployments.",
    "The append-only graduation ledger contains 50 started and 50 completed evaluation runs with 400 observations. Five repeated latency trials remained eligible, with worst public and held-out p95 ratios of 0.9863 and 1.0665 against the sealed 1.10 ceiling.",
    "Stage 4 is graduated. Hiro is ready to enter Stage 5 reviewed integration and monitoring, but no candidate has been approved, merged, promoted, or deployed."
  ],
  "workstreams": [
    {
      "title": "Sealed natural-candidate graduation campaign",
      "status": "Passed",
      "details": [
        "The graduation thresholds were written and hashed before candidate construction: at least ten eligible natural candidates, perfect gate-decision accuracy, five repeated eligible latency trials, three new interruption boundaries, no integrity or authority failures, and evidence-based bounded proposals.",
        "Qwen 3.6 35B-A3B constructed candidates for numeric clamping, inclusive windows, retry backoff, boolean configuration parsing, safe ratios, cache freshness, override precedence, chunk boundaries, ordered tag normalization, and canonical priority mapping.",
        "Each candidate began at the same sealed baseline commit in a distinct external worktree and was limited to one implementation path and one targeted-test path."
      ]
    },
    {
      "title": "Behavioral rejection and autonomous correction",
      "status": "Correctly rejected and corrected",
      "details": [
        "The initial boolean-parser candidate handled normalized true and false strings but did not generalize to the complete common-token contract exercised by held-out evaluation.",
        "The Stage 4 gate rejected it for held-out score, confidence, invariant, and category evidence even though its public and targeted tests passed. Its frozen packet and recommendation remain immutable and are excluded from the ten eligible-candidate count.",
        "A new candidate used an explicit generalized hypothesis for true, yes, on, 1 and false, no, off, 0 tokens plus unknown and non-string errors. It passed public, held-out, regression, invariant, integrity, latency, and sealed-oracle checks."
      ]
    },
    {
      "title": "Proposal quality and authority audit",
      "status": "Passed",
      "details": [
        "All eleven recommendations cited matching public, held-out, latency, regression, and evaluation identifiers and remained explicitly subject to human approval.",
        "All eleven frozen candidates addressed one hypothesis, changed only their two predeclared paths, included their targeted test, remained within the two-repair ceiling, and passed packet and snapshot integrity checks.",
        "No candidate was merged, promoted, deployed, committed into the fixture baseline, or granted service or scheduler authority."
      ]
    },
    {
      "title": "Interruption and idempotency coverage",
      "status": "Passed",
      "details": [
        "A partially recorded held-out run resumed to the exact unique repetition set without losing or duplicating observations.",
        "Already completed public and held-out evaluations were reused instead of rerun.",
        "A simulated interruption after proposal recording but before recommendation freeze resumed to one proposal and one frozen recommendation without duplicate proposal records.",
        "During campaign continuation, an attempted second process correctly refused an already owned external worktree. The original single owner completed and froze the candidate, which the continuation then reused. This wrapper-level contention did not create an evaluation run or enter the successful-run count."
      ]
    },
    {
      "title": "Latency, thermal, and repository validation",
      "status": "Passed",
      "details": [
        "Five repeated evaluations of the eligible clamp candidate all remained eligible. The worst public p95 ratio was 0.9863 and the worst held-out p95 ratio was 1.0665, both below 1.10.",
        "Observed GPU temperature ranged from 36 to 52 degrees Celsius; the campaign's 80-degree thermal pause was never reached.",
        "The independent repository suite completed with 279 passing tests in 55.27 seconds."
      ]
    }
  ],
  "decisions": [
    "Treat candidate implementation failure and evaluator decision failure as separate measurements. The first boolean candidate failed its sealed oracle, so rejecting it was a correct gate decision rather than a false rejection.",
    "Preserve the rejected packet and proposal append-only, then create a new candidate with a clearer generalized hypothesis instead of rewriting the failed result.",
    "Count only the ten eligible, oracle-passing candidates toward the graduation minimum while counting all eleven attempts in decision-accuracy, scope, integrity, and proposal-quality audits.",
    "Keep the p95 latency ceiling at 1.10 because five eligible trials produced no false rejection and stayed below the threshold; changing it is not supported by this evidence.",
    "Graduate Stage 4 to readiness for Stage 5 review, not to automatic integration. Every candidate remains frozen and requires explicit approval."
  ],
  "validation": [
    {
      "check": "Natural candidate outcomes",
      "status": "passed",
      "result": "Eleven attempts yielded ten eligible sealed-oracle passes and one correctly rejected sealed-oracle failure; all eleven gate decisions were correct."
    },
    {
      "check": "Append-only evaluation ledger",
      "status": "passed",
      "result": "50 runs started, 50 runs completed, 400 observations recorded, and zero incomplete runs remained."
    },
    {
      "check": "Repeated eligible latency",
      "status": "passed",
      "result": "Five of five trials remained eligible; maximum public and held-out p95 ratios were 0.9863 and 1.0665 under the 1.10 limit."
    },
    {
      "check": "Interruption recovery",
      "status": "passed",
      "result": "Held-out observation resume, completed-evaluation reuse, and recommendation-freeze resume all completed without duplicated observations or proposals."
    },
    {
      "check": "Candidate and recommendation audit",
      "status": "passed",
      "result": "Zero candidate-integrity, recommendation-integrity, scope, targeted-test, repair-bound, proposal-evidence, or proposal-authority failures across eleven attempts."
    },
    {
      "check": "Failure classification",
      "status": "passed",
      "result": "Zero model or infrastructure failures were included in the campaign; the one rejected behavioral candidate was preserved and excluded from the eligible count."
    },
    {
      "check": "Repository-wide regression suite",
      "status": "passed",
      "result": "279 tests passed in 55.27 seconds."
    },
    {
      "check": "Integration authority boundary",
      "status": "passed",
      "result": "Merge, promotion, and deployment counts remained zero; the fixture source repository stayed clean with its single sealed baseline commit."
    },
    {
      "check": "Hiro journal generation and frontend build",
      "status": "passed",
      "result": "Timestamped-entry tests passed, the generator produced and validated 59 journal pages, and the TypeScript and Vite production build completed successfully."
    }
  ],
  "currentState": [
    "Stage 4 proposal-only candidate evaluation is graduated under the sealed criteria.",
    "Hiro can build frozen external-worktree candidates, run matching public and held-out evaluation, apply regression, invariant, category, statistical, and latency gates, recover key interrupted transitions, and freeze review-ready recommendations.",
    "Ten eligible candidates and one correctly rejected-and-corrected candidate remain frozen outside the Hiro source repository.",
    "Hiro is ready for Stage 5 reviewed integration and monitoring, but no specific candidate is approved for integration."
  ],
  "limitations": [
    "The natural repair campaign used sealed synthetic fixture tasks so outcomes had deterministic independent oracles; production Hiro changes will carry broader interaction and operational risk.",
    "Stage 5 restart, canary monitoring, rollback, and post-promotion evidence have not yet been exercised end to end.",
    "The campaign validates candidate-level proposal accuracy, not consciousness, independent goals, unrestricted autonomy, or permission to self-deploy.",
    "Human approval remains mandatory for every integration, and evaluator, safety, credential, service, scheduler, dependency, database, and deployment changes remain outside automatic authority."
  ],
  "nextSteps": [
    "Define a narrow Stage 5 pilot allowlist and select one low-risk unchanged frozen candidate for explicit human review.",
    "Before integration, verify the candidate hash, reproduce its targeted and regression tests, and capture the stable revision and rollback plan.",
    "Exercise final smoke testing, affected-service-only restart, post-promotion canaries, monitoring thresholds, and automatic rollback in an isolated low-risk pilot.",
    "Keep automatic promotion disabled until Stage 5 produces repeated successful integration and monitoring evidence."
  ],
  "disclosureNote": "This public entry contains no credentials, tokens, private held-out prompts or expected answers, personal data, or actionable details about unresolved security weaknesses."
}
