cha0tiktower  ·  2× RTX 5060 Ti Blackwell  ·  since April 2026

Famous models.
Your hardware.

215+ real inference experiments on a $2,000 consumer GPU rig. Every model tried, every quantization tested, every failure logged with a date and a reason. No marketing. Just what actually runs, and what actually doesn't.

Throughput · tok/s · apr–jun 2026 every point is a dated experiment →
0 50 100 tok/s APR MAY JUN 10.4 cpu only 107.2 2× gpu 22 kernel wall 83 genesis 117.6 nemotron ~130 ornith peak ~101 warm now · 131K ctx
~130tok/s peak · Ornith-35B
215+experiments logged
15model families tested
speedup vs CPU baseline
full stack →

This is the exact model serving every agent on the tower today. It changes when a better one wins — that's the whole point of the log below.

In Production

Ornith-1.0-35B-AEON-Ultimate nvfp4

NVFP4A16 (FP4 experts, BF16 attention) on vLLM 0.23 with the SM_120 Marlin compatibility patch and fp8 KV cache. 131K context window. Routed through the fleet proxy at :8010 — switch backends with proxy-switch, no client reconfiguration needed.

~101
tok/s warm · 124 peak
interactive flowchart →

From a laptop that took four minutes to say “hello” to a production rig doing 130 tokens a second. Every era is a different bottleneck discovered and broken.

$ proxy-switch ornith
switching backend → Ornith-1.0-35B-AEON-Ultimate NVFP4 · vLLM 0.23 · port 8010
warm: ~101 tok/s · ctx 131K · fp8 KV · SM_120 Marlin patch applied
ornith is live.
Ornith NVFP4 hits production
The Upgrade That Wasn't
The Local Inference Chronicles: Everything We Tried