Hiro development journal

Qwen reasoning-effort comparison: fewer empty responses, mixed quality

Isolated comparison complete; production settings unchanged Machine-readable JSON

Executive summary

Completed the authorized existing-template comparison: embedded default xhigh versus explicit medium, using the unchanged frozen Hiro qualification harness. No replacement template was installed.

Medium raised the weighted score from 0.918299 to 0.940773, eliminated three empty responses, and passed the existing mandatory gates. Both conditions fully passed 60 of 72 responses; a causal-diagnosis regression prevents calling medium a uniform improvement.

Work completed

Source and artifact inspection

Completed
  • Reviewed https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates and its template source; saved candidate revision 855bffc49448e299789730ff92c9b8d834d6cc14 for inspection.
  • Extracted tokenizer.chat_template from the existing Qwen3.8-27B-Q4_K_M GGUF without loading model weights or executing downloaded scripts. Extracted template SHA-256: c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041.
  • Confirmed embedded reasoning_effort default xhigh. The inspected launcher enables Jinja without an explicit replacement template. The isolated evaluation used the extracted original template; the production live override was not independently established.
  • The local router supports explicit reasoning effort. The isolated backend now demonstrably applies the top-level parameter: actual chat-completion rendered prompts match the render endpoint.
  • The previous frozen baseline run omitted an effort override and exhausted 4096 completion tokens with zero visible response on all three benchmark-generation repetitions. This association does not prove causation.
  • Checked open handoffs; did not execute the existing ranked-autonomy handoff.

Controlled local comparison

Completed
  • Used Qwen3.8-27B Q4_K_M, bundled llama.cpp CUDA12 backend 2.41.0, RTX 5090, context 8192, maximum completion 4096, temperature 0.1, seed 42 and parallel 1. Only reasoning_effort differed.
  • Each condition ran 24 frozen cases with three repetitions, 144 total responses, through model_benchmarks_v2. Preserved outputs, profiles, artifact hashes, labels, scores and run receipts in an isolated ledger.
  • Verified actual rendered inputs: medium removes the xhigh steering instruction while preserving thinking. A replacement template is unnecessary for this setting.
  • Fresh default reproduced all 72 prior scores, gate results and finish reasons.
  • Default versus medium: assistant score 1.0/1.0; RSI 0.863832/0.901289; weighted score 0.918299/0.940773; fully passed 60/72 in both; mandatory failures 6/0; harness rejected/qualified. Qualification is not deployment authority.
  • Completion tokens 32814/23781; median latency 3.771/4.121 seconds; p95 14.302/9.029 seconds. No API errors.
  • Medium returned valid nonempty benchmark-generation answers but still failed two correctness checks. One additional case fully passed and one causal-diagnosis case regressed; other 21 case scores and labels were unchanged.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Template extraction and runtime parameter propagation passed Original metadata template preserved; render endpoint and actual chat prompt agree for default and explicit medium.
Frozen paired comparison completed 144 responses, no API errors. Default rejected by existing harness; medium qualified. Both fully passed 60/72 responses; medium regressed on causal diagnosis.
Prior baseline reproducibility passed All 72 default scores, gate labels and finish reasons matched the previous run.
Community replacement template not_run No template installed; its incremental benefit and compatibility remain unknown.
Model restoration passed Load CLI exited successfully; model identity, context and parallelism match the saved pre-session state. Local model health returned 200 and test port has no listener.
Journal validation passed npm run test:hiro and npm run build passed; generated-page validation checked 246 entries, timestamp schema, aliases and noindex. Final source updated with restoration receipt and rebuilt before publication.

Current state

Next steps