Executive summary
Hiro now has a production-independent Model Qualification System for comparing local language models without allowing later Hiro upgrades to change the measurement instrument or historical leaderboard.
Recursive self-improvement is the primary capability lane: sixteen frozen RSI cases contribute sixty percent of the result, while eight everyday-assistant cases contribute forty percent. Evaluator tampering, held-out-case seeking, false test claims, and unsafe action or promotion decisions are non-compensable hard failures.
Canonical runs use a hash-verified standalone harness, three repetitions per case, exact model and runtime fingerprints, an 8,192-token reference context, a loopback-only inference endpoint, returned-model identity checks, and a separate append-only SQLite ledger.
The Evaluation Observatory now has a dedicated Models section that separately displays canonical leaderboard evidence, screens, registered profiles, reproducibility identities, and future live-Hiro compatibility evidence.
A first canonical Qwen3.6 control campaign completed after the model was loaded directly at the required 8,192-token context. The run produced complete reproducibility evidence and exposed both genuine model failures and overly literal hidden-enum checks in the version-1 rubric, so its numeric score is retained as calibration evidence rather than accepted as the authoritative comparison baseline.
Qwen3.8-27B was downloaded, hash-pinned, loaded directly at the controlled 8,192-token context, screened, and then run through the complete canonical campaign. No candidate model was promoted and Hiro's active profile was not changed.
A separately frozen version-2.1 harness replaced hidden wording expectations with explicit structural decisions, standardized every model at an 8,192-token context, temperature 0.1, and a 4,096-token completion budget, and added streamed time-to-first-token, generation-throughput, reasoning-token, finish-reason, and budget-exhaustion telemetry.
The authoritative Qwen3.6 version-2.1 baseline scored 0.786638 overall, 0.797595 RSI, and 0.770202 assistant. It was rejected because of four genuine unsafe structural decisions and four strict-JSON output-contract failures; it completed with zero endpoint errors, empty responses, or token-budget exhaustions.
The matched Qwen3.8 canonical run scored 0.921317 overall, 0.876718 RSI, and 0.988215 assistant, materially exceeding Qwen3.6 on quality. It was still rejected: all three benchmark-generation repetitions exhausted the 4,096-token completion budget without final JSON, and all three promotion-eligible repetitions selected automatic deployment.
Qwen3.8 generated at 60.149 tokens per second p50 versus Qwen3.6's 161.603. Its median latency was lower because it emitted far fewer total tokens, but its 44.823-second p95 exposed a severe long-tail cost on reasoning-heavy cases.