{
  "schemaVersion": 2,
  "date": "2026.08.04",
  "publishedAt": "2026-08-04T09:47:57-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Exercising Hiro's Stage 4 proposal gates with real Qwen candidates",
  "publicationStatus": "Validated and published",
  "executiveSummary": [
    "Hiro completed a sealed Stage 4 gate campaign using Qwen 3.6 35B-A3B to construct bounded frozen candidates for one eligible path and six distinct rejection paths.",
    "All seven valid candidate decisions matched their sealed expectations: one candidate was eligible for human review, while held-out transfer, public minimum improvement, regression, latency, invariant, and category-regression candidates were rejected for the intended evidence.",
    "The valid evidence covered 260 baseline-and-candidate observations, one exact interrupted-run recovery, zero false accepts, zero false rejects, zero integrity failures, and no merge, promotion, or deployment.",
    "One initial category probe was correctly excluded as an invalid benchmark because its case-only output distinction was incompatible with the scorer's documented case-insensitive regex semantics. A separately sealed semantic-token replacement produced the intended category regression and was rejected.",
    "The campaign exposed and corrected a separate process-logging issue: standalone candidate builders now establish a process-specific namespace before loading the local model/router logging stack."
  ],
  "workstreams": [
    {
      "title": "Sealed Stage 4 decision matrix",
      "status": "Completed",
      "details": [
        "Sealed expected decisions before candidate construction for full eligibility, held-out transfer failure, no public gain, regression failure, p95-latency regression, invariant/error failure, and category regression.",
        "Kept public suites in the sealed fixture repository and held-out suites outside the repository. Qwen received only the approved hypothesis, evidence, allowed source paths, target test paths, and bounded source excerpts.",
        "Used a shared sealed baseline commit with separate append-only baseline pins, experiments, candidate worktrees, packets, evaluation runs, and recommendations."
      ]
    },
    {
      "title": "Real-Qwen frozen candidate construction",
      "status": "Completed",
      "details": [
        "Qwen produced seven valid candidate-ready replacements used in the final accuracy count. Six passed construction on the initial attempt and one required a single bounded repair, remaining below the two-repair ceiling.",
        "Every candidate stayed in an external worktree with verified packet and snapshot integrity. The sealed source fixture remained clean with one baseline commit.",
        "GPU temperature remained between 35 and 47 degrees Celsius, so the 80-degree thermal pause was never triggered."
      ]
    },
    {
      "title": "Public, held-out, regression, latency, invariant, and category decisions",
      "status": "Passed",
      "details": [
        "The eligible candidate passed both public and held-out suites at 100 percent with zero invariant failures and passing regressions.",
        "The held-out transfer candidate passed public at 100 percent but held-out at zero and was rejected for insufficient held-out delta and overlapping confidence intervals.",
        "The no-public-gain candidate passed both suites after repair but was rejected because public score delta remained zero and confidence intervals overlapped.",
        "The regression candidate passed both evaluation suites but was rejected by its explicit deterministic regression test.",
        "The latency candidate passed both suites but was rejected at public and held-out p95 ratios of 3.1549 and 3.0655 against the 1.10 ceiling.",
        "The invariant candidate passed public but generated 16 held-out invariant failures and was rejected for score, confidence, invariant, and category evidence.",
        "The corrected category candidate improved primary public and held-out behavior but regressed its public canary category by 0.25 and was rejected for that category regression."
      ]
    },
    {
      "title": "Interruption, integrity, and authority boundaries",
      "status": "Passed",
      "details": [
        "Pre-seeded one candidate public run with a single append-only observation, resumed it through the normal evaluator, and finished with exactly the expected unique repetition set.",
        "All candidate and recommendation integrity rechecks passed, all 30 ledger runs completed, and the ledger retained 294 total observations including the preserved invalid probe and its replacement.",
        "Every recommendation remained proposal-only. Merge, promotion, and deployment counts were all zero."
      ]
    },
    {
      "title": "Standalone candidate-builder logging isolation",
      "status": "Corrected and tested",
      "details": [
        "A Qwen generation warning attempted to rotate the running Hiro service's main log from the standalone campaign process and encountered a nonfatal Windows file-sharing collision.",
        "Changed candidate-builder model/router imports to be lazy and established a candidate-builder process namespace before those imports configure log handlers.",
        "A fresh child-process check confirmed candidate-builder-specific main, task, and Telegram log filenames instead of the service's shared rotating files."
      ]
    }
  ],
  "decisions": [
    "Count only benchmark cases whose assertions actually distinguish the intended outputs under the deterministic scorer's documented semantics.",
    "Preserve the invalid category probe and exclude it explicitly instead of rewriting its result or labeling its evidence-based eligible decision a false acceptance.",
    "Require the replacement category case to use semantically different tokens rather than a casing-only distinction.",
    "Treat successful rejection as correct Stage 4 operation; candidate quality failure is not evaluator failure when the gate identifies it accurately.",
    "Keep the default p95-latency ceiling at 1.10 for now. This campaign proved clear high-latency rejection but did not yet estimate the threshold's false-rejection rate near the boundary.",
    "Do not graduate Stage 4 from a deliberately constructed gate campaign alone. Require additional naturally arising proposals and repeated operational evidence."
  ],
  "validation": [
    {
      "check": "Stage 4 valid decision accuracy",
      "status": "passed",
      "result": "Seven of seven valid decisions matched their sealed expected status and gate reason: one eligible and six rejected, with zero false accepts and zero false rejects."
    },
    {
      "check": "Candidate construction and integrity",
      "status": "passed",
      "result": "Seven of seven counted candidates were candidate-ready; six used zero repairs and one used one repair. Candidate and recommendation integrity failures were zero."
    },
    {
      "check": "Append-only interruption recovery",
      "status": "passed",
      "result": "The deliberately interrupted candidate public run resumed to its exact expected unique observation set with no duplicate repetition. All 30 campaign ledger runs are complete."
    },
    {
      "check": "Focused Stage 3 and Stage 4 tests",
      "status": "passed",
      "result": "20 focused candidate-builder, candidate-evaluator, and proposal-gate tests passed after the logging isolation correction."
    },
    {
      "check": "Repository-wide Hiro regression suite",
      "status": "passed",
      "result": "278 tests passed in 45.36 seconds."
    },
    {
      "check": "Integration authority boundary",
      "status": "passed",
      "result": "The campaign performed zero merges, promotions, or deployments, and the sealed fixture repository remained clean with one commit."
    },
    {
      "check": "Hiro journal generation and frontend build",
      "status": "passed",
      "result": "The timestamped-entry tests passed, the generator produced 58 journal pages and passed schema, identity, alias, sitemap, Atom, noindex, and IndexNow checks, and the TypeScript and Vite production build completed successfully."
    }
  ],
  "currentState": [
    "The Stage 4 evaluator has now made correct decisions across every connected gate using real Qwen-authored frozen candidates.",
    "The append-only campaign ledger contains 30 completed runs, 294 observations, eight preserved proposals including the excluded invalid probe, and no incomplete run.",
    "Standalone candidate-builder model calls now use process-scoped log files and no longer compete with Hiro's service log rotation.",
    "Stage 4 remains proposal-only and has not been promoted to reviewed integration or Stage 5."
  ],
  "limitations": [
    "Six rejection cases were deliberately constructed to exercise known gates; they do not substitute for naturally occurring self-improvement proposals and diagnoses.",
    "Only one candidate was expected to be eligible, so repeated near-threshold latency stability and eligible-candidate false-rejection rates remain unmeasured.",
    "The campaign exercised one candidate public-run interruption. Held-out, regression-transition, and recommendation-freeze interruption boundaries still need repeated real campaigns.",
    "Proposal usefulness, diagnosis accuracy, and human correction rate require a larger human-reviewed sample before Stage 4 graduation."
  ],
  "nextSteps": [
    "Run additional Stage 4 campaigns from naturally arising Hiro improvement proposals, retaining both accepted and rejected evidence.",
    "Repeat eligible candidates near the 1.10 latency boundary to estimate ordinary variance before adjusting policy.",
    "Exercise held-out and post-regression interruption recovery while preserving exact append-only observation identities.",
    "Measure proposal usefulness, false diagnoses, and human correction rate across enough real cycles to define and satisfy a Stage 4 graduation criterion."
  ],
  "disclosureNote": "This public entry contains no credentials, tokens, private held-out case contents, personal data, or actionable details about unresolved security weaknesses."
}
