Executive summary
Hiro's host was inventoried after its GPU upgrade. It now has an NVIDIA GeForce RTX 5090 with 32,607 MiB of VRAM, an Intel Core i7-10700K with 16 logical processors, 31.84 GiB of physical RAM, and several terabytes of free space on E:.
The practical target is fully GPU-resident models with 20B to 35B total parameters. The GPU can hold high-quality 4-bit or 5-bit versions with a useful context cache, while 32 GiB of system RAM makes aggressive CPU offload of 70B-plus models unattractive.
The first recommended trials are Qwen3.6-35B-A3B Q4_K_M, Gemma 4 31B Q4_K_M, and gpt-oss-20b in native MXFP4. Gemma 4 26B-A4B, Qwen3.5-27B, and Devstral Small 2 24B form a second wave for efficiency, dense-model comparison, and coding specialization.
No model was downloaded or benchmarked. This session produced an evidence-based acquisition and evaluation plan and leaves Gemma 4 12B as Hiro's baseline until a candidate passes compatibility, quality, latency, memory, and reliability gates.