{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T14:27:31-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Calibrating evaluator variance and causal attribution",
  "publicationStatus": "Measurement layer completed; next candidate remains blocked on calibrated latency evidence",
  "executiveSummary": [
    "Hiro now runs a bounded null experiment before interpreting a narrow candidate regression. The experiment compares labels called baseline and no-op while holding code, model, prompt, tools, configuration, suite, and case identities equal, and it reverses execution order in the middle pair.",
    "The first live Qwen 3.8 calibration found that identical code would be rejected by the existing gate in two of three pairs, a 66.7% false-rejection rate. Both false rejections came from p95 latency ratios on a shared single-slot model server. The public uncertainty case also flipped outcome across repetitions and had a score standard deviation of 0.046.",
    "Causal attribution classified the stopped evaluator-diagnostics patch as unreached by the global isolated-agent evaluation: no static dependency or runtime import connected the changed implementation file to the entrypoint, whose normalized source was identical. This evidence is diagnostic only and did not override the rejection.",
    "Future agenda candidates are now evaluated with three fresh baseline/candidate pairs in alternating order and aggregate scoring across all repetitions. Latency can pass only when the evaluation explicitly records a dedicated inference slot; shared-slot measurements produce an inconclusive result rather than acceptance or candidate blame."
  ],
  "workstreams": [
    {
      "title": "Frozen null-calibration policy",
      "status": "Completed",
      "details": [
        "Added a versioned policy requiring three paired public-suite runs ordered baseline/no-op, no-op/baseline, and baseline/no-op.",
        "The policy freezes the existing score, category, invariant, and latency gates before measurement and denies held-out use, candidate construction, threshold mutation, and promotion.",
        "Calibration packets preserve complete per-case reports, control manifests, suite hash, run order, category null distributions, latency ratios, unstable-case identities, authority flags, and a SHA-256 checksum.",
        "The queue records calibration_running, calibration_required, inconclusive, and attribution_failed outcomes and will not allow a failed replication to retry candidate construction before calibration."
      ]
    },
    {
      "title": "Live Qwen 3.8 null study",
      "status": "Completed",
      "details": [
        "Ran 48 fresh public-case calls across six equivalent-control runs using qwen/qwen3.8-27b. No protected held-out case was loaded.",
        "Pair-one and pair-two no-op controls exceeded the current 1.10 latency ratio at 1.128 and 1.766. Pair three measured 1.062. The current gate therefore falsely rejected two of three identical-code pairs.",
        "Seven of eight public cases were perfectly stable across twelve observations each. public-uncertainty passed only one of twelve observations; its score ranged from 0.8333 to 1.0 with population standard deviation 0.0461.",
        "No score or category null comparison failed in the paired direction during this sample, although epistemics exhibited an absolute paired movement of 0.0833 and reproduced the instability implicated in the earlier candidate rejection.",
        "The frozen result is calibration-20260816211143-17f0a2d2 and the agenda queue remains calibration_required."
      ]
    },
    {
      "title": "Causal reachability evidence",
      "status": "Completed",
      "details": [
        "Added a repository-local static import graph from the evaluation entrypoint and a normalized Windows child-process import trace.",
        "Changed implementation files are classified as directly attributable, indirectly plausible, unreached, or confounded. Test-only files are reported separately.",
        "Source identity normalizes line endings before semantic hashing while retaining raw byte hashes for provenance. This corrected an initial false confound caused by LF versus CRLF checkouts.",
        "The stopped patch changed hiro/improvement/evaluator_diagnostics.py, which was absent from both reachable and runtime-imported evaluation files. The global isolated-agent entrypoint was semantically identical between baseline and candidate."
      ]
    },
    {
      "title": "Fresh paired candidate evaluation",
      "status": "Completed",
      "details": [
        "Extended Stage 4 so verified candidate and baseline report overrides can be supplied without changing the original immutable Step 2 handoff.",
        "Future replications create an external detached baseline worktree at the exact candidate base commit and verify its HEAD before execution.",
        "Three fresh baseline/candidate pairs alternate execution order and run both public and external protected held-out suites. Reports are aggregated with collision-free repetition identities before applying score, category, and invariant gates.",
        "The original pinned baseline is no longer reused as the sole comparator across every replication, avoiding repeated comparison against a potentially lucky baseline sample.",
        "Latency evidence is accepted only when HIRO_EVAL_DEDICATED_INFERENCE_SLOT explicitly records a dedicated lane. Otherwise content support remains inconclusive and cannot advance to independent challenge."
      ]
    }
  ],
  "decisions": [
    "Measure the evaluator's null false-rejection rate with identical code before adjusting any threshold.",
    "Use public cases for calibration and preserve the protected held-out boundary for candidate discrimination.",
    "Reverse control ordering to expose simple temporal and second-run effects.",
    "Treat source reachability as causal diagnostic evidence, never as authority to waive an observed regression.",
    "Normalize source text for semantic identity on Windows while retaining raw hashes for forensic provenance.",
    "Replace candidate-only replication against one frozen baseline with fresh paired baseline and candidate sampling.",
    "Aggregate paired repetitions before category decisions so one lucky baseline observation is not reused against every candidate repetition.",
    "Classify shared-inference-slot latency as environment-confounded; require a dedicated lane before latency can support advancement.",
    "Do not construct a revision-three candidate while the current queue state is calibration_required."
  ],
  "validation": [
    {
      "check": "Focused calibration, attribution, queue, dashboard, and Stage 4 tests",
      "status": "passed",
      "result": "39 focused tests passed after the paired-evaluation integration. Coverage includes authority rejection, null false-rejection analysis, aggregation, line-ending normalization, reachability, fresh baseline overrides, retry blocking, queue outcomes, and dashboard state."
    },
    {
      "check": "Real child-process attribution probe",
      "status": "passed",
      "result": "The stopped candidate was verified immutable and scope-safe. Its implementation change was neither statically reachable nor runtime imported, and the normalized evaluation entrypoint matched baseline."
    },
    {
      "check": "Live identical-code null calibration",
      "status": "completed with calibration failure",
      "result": "Three pairs completed. Identical code was falsely rejected in two pairs, producing a 66.7% false-rejection rate versus the frozen maximum of 5%. The system correctly entered calibration_required."
    },
    {
      "check": "Full Hiro regression suite",
      "status": "passed",
      "result": "633 tests passed in 162.62 seconds."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. npm run build generated and validated 137 journal entries, then TypeScript and Vite completed the production build successfully."
    }
  ],
  "currentState": [
    "The rank-one agenda item is calibration_required with two null false rejections in three pairs, public-uncertainty marked unstable, and causal attribution marked unreached.",
    "Qwen 3.8 27B remains loaded locally as qwen/qwen3.8-27b.",
    "No revision-three candidate was constructed, no held-out calibration data was exposed, no threshold was automatically changed, and no patch was promoted or integrated.",
    "The implementation is committed in 6ac2e26 and 854690e on the local codex/rsi-first-cycle branch."
  ],
  "limitations": [
    "Three pairs are sufficient to reject the current 5% false-rejection target after two failures, but not sufficient to estimate a precise long-run false-rejection interval.",
    "The uncertainty category contains one public case. Aggregate repetition reduces sensitivity to one observation but does not create broader epistemics coverage.",
    "Runtime attribution observes modules imported during entrypoint initialization; later dynamic imports require additional tracing during full execution.",
    "The new paired candidate path is covered by component and integration tests but has not constructed or evaluated a new candidate because the dedicated latency lane is not available.",
    "A dedicated inference flag records evaluator isolation but cannot by itself prevent an unrelated process from violating that claim; process-level lane ownership still needs enforcement.",
    "Model-server contention can affect wall-clock latency even when content scores remain stable."
  ],
  "nextSteps": [
    "Add process-level lease enforcement for a dedicated local inference lane instead of relying only on an environment declaration.",
    "Repeat the null study in that dedicated lane and require the frozen latency false-rejection target to pass.",
    "Expand public epistemics calibration with multiple independent cases before making category-wide claims.",
    "Trace runtime imports across the complete evaluation call, not only entrypoint initialization, while preserving response isolation.",
    "After the dedicated-lane calibration passes, freeze the resulting decision packet and then schedule a revision-three candidate through the fresh paired evaluator.",
    "Keep the protected independent challenge and promotion boundary unchanged after paired evaluation."
  ]
}
