cha0tiktower · 2× RTX 5060 Ti Blackwell · since April 2026
Famous models.
Your hardware.
215+ real inference experiments on a $2,000 consumer GPU rig. Every model tried, every quantization tested, every failure logged with a date and a reason. No marketing. Just what actually runs, and what actually doesn't.
Throughput · tok/s · apr–jun 2026
every point is a dated experiment →
01 — Right Now
full stack →
This is the exact model serving every agent on the tower today. It changes when a better one wins — that's the whole point of the log below.
02 — The Journey, Five Eras
interactive flowchart →
From a laptop that took four minutes to say “hello” to a production rig doing 130 tokens a second. Every era is a different bottleneck discovered and broken.
Era 01
10.4 tok/s
CPU Era
Beelink EQI12 · i5-1235U · Gemma-4 26B, chronic swap thrashing.
Era 02
32 → 107
Tower Arrives
RTX 5060 Ti ×1 then ×2. GPU offload, SWA discovery, PCIe ceiling found.
Era 03
22 → 88
Genesis
GDN architecture wall in llama.cpp. vLLM + MTP n=3 breaks it wide open.
Era 04
117.6
Nemotron
Mamba/SSM hybrid MoE. New hardware peak on llama.cpp cuda128-clean.
Era 05 · Now
~130
Ornith
35B GGUF beats DeepSeek V3.2 on tool-use. NVFP4 production, 131K ctx.
03 — Field Notes
$ proxy-switch ornith
switching backend → Ornith-1.0-35B-AEON-Ultimate NVFP4 · vLLM 0.23 · port 8010
warm: ~101 tok/s · ctx 131K · fp8 KV · SM_120 Marlin patch applied
ornith is live.