| Frozen suite and comparability |
passed |
24 cases per model, three repetitions; suite/harness integrity and existing ledger comparability checks passed. |
| Artifact verification |
passed |
Both downloaded GGUF hashes match publisher/repository LFS hashes. Baseline matches the existing pinned hash. |
| baseline model qualification |
rejected |
Completed 72 responses: assistant 100.0%, self-improvement 86.4%, weighted 91.8%, fully passing 60/72. Median latency 3.97s; hard-gate failures 6. |
| nex model qualification |
rejected |
Completed 72 responses: assistant 100.0%, self-improvement 82.2%, weighted 89.3%, fully passing 54/72. Median latency 0.57s; hard-gate failures 6. |
| empero model qualification |
rejected |
Completed 72 responses: assistant 87.4%, self-improvement 82.7%, weighted 84.6%, fully passing 42/72. Median latency 1.19s; hard-gate failures 3. |
| Session cleanup |
passed |
Previous LM Studio model restored with matching identity, context, parallelism and idle state; isolated benchmark server stopped. |
| Journal validation |
passed |
npm run test:hiro and npm run build passed, including 243-entry generated validation. |