Blog
Local LLM speed, measured
Articles grounded in signed, reproducible benchmark runs — every number links to the submission that produced it.
- What hardware do you need to run DeepSeek-V4-Flash?2026-07-08
DeepSeek-V4-Flash is a 284B MoE (13B active). Running it locally needs roughly 96GB VRAM or 128GB unified memory; a 512GB Mac Studio has the memory, but no stable backend loads V4 yet.
Read - How fast is GLM-4.7-Flash on an RTX 4090?2026-07-08
Signed benchmark: GLM-4.7-Flash, a 30B-A3B mixture-of-experts coding model, decodes 130 tok/s on a 24GB RTX 4090. A fast small-MoE coder that fits one card.
Read - How fast is Gemma 3 on an RTX 4090? (4B, 12B, 27B)2026-07-03
Signed benchmarks: Gemma 3 decodes 195 tok/s (4B), 93 (12B), and 47 (27B) on a 24GB RTX 4090; every size fits. Which to run for your workload.
Read - The best model for a 512GB Mac Studio (measured)2026-07-03
On a 512GB M3 Ultra, memory is not the limit, speed is. A large MoE like Qwen3-Next-80B runs 80 tok/s, far faster than a dense 70B at 17. Which model to run.
Read - How fast is Qwen3.6-27B on an RTX 4090?2026-07-03
Signed benchmarks: Qwen3.6-27B, a top single-GPU coding model, decodes 44 tok/s on a 24GB RTX 4090 and 74 on a 5090; the 35B-A3B MoE variant is about 3x faster.
Read - Local LLM inference speed in 2026: what we've measured2026-07-02
Signed, reproducible decode tok/s across consumer GPUs and Apple Silicon: the fastest configs, why small-MoE coders win, and what counts as fast enough.
Read - The fastest local coding models in 2026 (measured on an RTX 5090)2026-07-02
Signed, reproducible decode tok/s for the top local coding models on a single RTX 5090, and why small-MoE coders decode ~4x faster than a dense 32B.
Read - The RTX 3090 in 2026: still the value pick for local LLMs?2026-07-02
Signed benchmarks on an RTX 3090 (24 GB): the decode tok/s a used ~$700 card really delivers for local LLMs in 2026. Measured, not folklore.
Read - Decode vs prefill tok/s: what LLM speed numbers actually mean2026-07-02
The two tok/s numbers behind every LLM benchmark: prefill (reading the prompt) vs decode (writing the answer), why decode is ~25x slower, and which matters.
Read - RTX 3090 vs RTX 4090 for local LLMs: is 2x the price worth it?2026-07-02
Signed benchmarks on both cards, same models: the RTX 4090 decodes about 15-30% faster than the 3090 but costs roughly double. Which is the better buy?
Read - What's the fastest GPU for running Llama 3.1 8B locally?2026-07-02
Measured decode tok/s for Llama 3.1 8B on the RTX 5090, 4090, 3090, and Apple Silicon: which runs it fastest, and which is the best value.
Read - The best GPU for a local coding agent in 20262026-07-02
What we measured for local coding models on the RTX 3090, 4090, 5090, and Apple Silicon: the best value, the fastest, and why the model matters too.
Read - RTX 5090 vs M3 Ultra for local LLMs: speed vs memory2026-07-02
First-party signed benchmarks: the RTX 5090 decodes roughly 2x faster than an M3 Ultra, but the Mac's huge unified memory runs models the GPU cannot. Which to pick.
Read - How fast is Qwen3-Coder locally? RTX 5090, 4090, and Apple Silicon2026-07-02
Signed decode benchmarks for Qwen3-Coder-30B-A3B: 260 tok/s on an RTX 5090, 180 on a 4090, and ~112 on Apple Silicon. Why a small-MoE coder is fast everywhere.
Read - Codestral vs Qwen3-Coder: which local coding model is faster?2026-07-02
Signed benchmarks: Qwen3-Coder-30B-A3B decodes about 2.5x faster than Codestral-22B on the same hardware, because it is a small mixture-of-experts. Which to run.
Read - DeepSeek-R1 on an RTX 4090: which size actually fits?2026-07-02
Signed benchmarks: DeepSeek-R1 8B decodes 134 tok/s and 14B 81 tok/s on a 24GB RTX 4090, but the 32B spills VRAM and crawls. Which size to run for reasoning.
Read