Hiro development journal

Comparing Nex and Empero model candidates for Hiro

Comparison complete; no production replacement Machine-readable JSON

Executive summary

Downloaded and hash-verified Nex-N2.5-mini and Empero Qwen3.8-35B-A3B Distill Q4_K_M artifacts, then compared them with the configured Qwen3.8-27B baseline.

Completed 216 responses through the unchanged frozen Hiro model qualification suite. Nex was the strongest speed/assistant challenger, but neither candidate qualified as a general replacement.

Work completed

Reproducible model comparison

Completed
  • Pinned Nex community and Empero publisher GGUF Q4_K_M artifacts by repository revision and SHA-256.
  • Reused the existing 24-case suite, three repetitions per case, 8192 context and 4096 completion budget at temperature 0.1.
  • Used an isolated loopback server and a separate ledger with existing benchmark code.
  • Temporarily unloaded an unrelated LM Studio model with explicit authorization; restore after evaluation.
  • Nex community artifact: ngquocvinh/Nex-N2.5-mini-GGUF at fe3093c81ec10e0388c5375468d2b6a86e0a6026.
  • Empero publisher artifact: empero-ai/Qwen3.8-35B-A3B-Distill-GGUF at b1f9d1dcc3de8aa867669b0ab919384aeeb9b8d5.
  • Both candidate downloads matched their published SHA-256 values; all three models passed runtime artifact checks.
  • Initial baseline mmap startup was stopped before scoring after no visible progress. Successful comparisons used identical non-mapped loading.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Frozen suite and comparability passed 24 cases per model, three repetitions; suite/harness integrity and existing ledger comparability checks passed.
Artifact verification passed Both downloaded GGUF hashes match publisher/repository LFS hashes. Baseline matches the existing pinned hash.
baseline model qualification rejected Completed 72 responses: assistant 100.0%, self-improvement 86.4%, weighted 91.8%, fully passing 60/72. Median latency 3.97s; hard-gate failures 6.
nex model qualification rejected Completed 72 responses: assistant 100.0%, self-improvement 82.2%, weighted 89.3%, fully passing 54/72. Median latency 0.57s; hard-gate failures 6.
empero model qualification rejected Completed 72 responses: assistant 87.4%, self-improvement 82.7%, weighted 84.6%, fully passing 42/72. Median latency 1.19s; hard-gate failures 3.
Session cleanup passed Previous LM Studio model restored with matching identity, context, parallelism and idle state; isolated benchmark server stopped.
Journal validation passed npm run test:hiro and npm run build passed, including 243-entry generated validation.

Current state

Next steps