Hiro development journal

Calibrating evaluator variance and causal attribution

Measurement layer completed; next candidate remains blocked on calibrated latency evidence Machine-readable JSON

Executive summary

Hiro now runs a bounded null experiment before interpreting a narrow candidate regression. The experiment compares labels called baseline and no-op while holding code, model, prompt, tools, configuration, suite, and case identities equal, and it reverses execution order in the middle pair.

The first live Qwen 3.8 calibration found that identical code would be rejected by the existing gate in two of three pairs, a 66.7% false-rejection rate. Both false rejections came from p95 latency ratios on a shared single-slot model server. The public uncertainty case also flipped outcome across repetitions and had a score standard deviation of 0.046.

Causal attribution classified the stopped evaluator-diagnostics patch as unreached by the global isolated-agent evaluation: no static dependency or runtime import connected the changed implementation file to the entrypoint, whose normalized source was identical. This evidence is diagnostic only and did not override the rejection.

Future agenda candidates are now evaluated with three fresh baseline/candidate pairs in alternating order and aggregate scoring across all repetitions. Latency can pass only when the evaluation explicitly records a dedicated inference slot; shared-slot measurements produce an inconclusive result rather than acceptance or candidate blame.

Work completed

Frozen null-calibration policy

Completed
  • Added a versioned policy requiring three paired public-suite runs ordered baseline/no-op, no-op/baseline, and baseline/no-op.
  • The policy freezes the existing score, category, invariant, and latency gates before measurement and denies held-out use, candidate construction, threshold mutation, and promotion.
  • Calibration packets preserve complete per-case reports, control manifests, suite hash, run order, category null distributions, latency ratios, unstable-case identities, authority flags, and a SHA-256 checksum.
  • The queue records calibration_running, calibration_required, inconclusive, and attribution_failed outcomes and will not allow a failed replication to retry candidate construction before calibration.

Live Qwen 3.8 null study

Completed
  • Ran 48 fresh public-case calls across six equivalent-control runs using qwen/qwen3.8-27b. No protected held-out case was loaded.
  • Pair-one and pair-two no-op controls exceeded the current 1.10 latency ratio at 1.128 and 1.766. Pair three measured 1.062. The current gate therefore falsely rejected two of three identical-code pairs.
  • Seven of eight public cases were perfectly stable across twelve observations each. public-uncertainty passed only one of twelve observations; its score ranged from 0.8333 to 1.0 with population standard deviation 0.0461.
  • No score or category null comparison failed in the paired direction during this sample, although epistemics exhibited an absolute paired movement of 0.0833 and reproduced the instability implicated in the earlier candidate rejection.
  • The frozen result is calibration-20260816211143-17f0a2d2 and the agenda queue remains calibration_required.

Causal reachability evidence

Completed
  • Added a repository-local static import graph from the evaluation entrypoint and a normalized Windows child-process import trace.
  • Changed implementation files are classified as directly attributable, indirectly plausible, unreached, or confounded. Test-only files are reported separately.
  • Source identity normalizes line endings before semantic hashing while retaining raw byte hashes for provenance. This corrected an initial false confound caused by LF versus CRLF checkouts.
  • The stopped patch changed hiro/improvement/evaluator_diagnostics.py, which was absent from both reachable and runtime-imported evaluation files. The global isolated-agent entrypoint was semantically identical between baseline and candidate.

Fresh paired candidate evaluation

Completed
  • Extended Stage 4 so verified candidate and baseline report overrides can be supplied without changing the original immutable Step 2 handoff.
  • Future replications create an external detached baseline worktree at the exact candidate base commit and verify its HEAD before execution.
  • Three fresh baseline/candidate pairs alternate execution order and run both public and external protected held-out suites. Reports are aggregated with collision-free repetition identities before applying score, category, and invariant gates.
  • The original pinned baseline is no longer reused as the sole comparator across every replication, avoiding repeated comparison against a potentially lucky baseline sample.
  • Latency evidence is accepted only when HIRO_EVAL_DEDICATED_INFERENCE_SLOT explicitly records a dedicated lane. Otherwise content support remains inconclusive and cannot advance to independent challenge.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused calibration, attribution, queue, dashboard, and Stage 4 tests passed 39 focused tests passed after the paired-evaluation integration. Coverage includes authority rejection, null false-rejection analysis, aggregation, line-ending normalization, reachability, fresh baseline overrides, retry blocking, queue outcomes, and dashboard state.
Real child-process attribution probe passed The stopped candidate was verified immutable and scope-safe. Its implementation change was neither statically reachable nor runtime imported, and the normalized evaluation entrypoint matched baseline.
Live identical-code null calibration completed with calibration failure Three pairs completed. Identical code was falsely rejected in two pairs, producing a 66.7% false-rejection rate versus the frozen maximum of 5%. The system correctly entered calibration_required.
Full Hiro regression suite passed 633 tests passed in 162.62 seconds.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 137 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps