M5 Max (40-core GPU) LLM benchmark
The fastest LLM measured on the M5 Max (40-core GPU) is qwen3.6-agent-q6k-ctx128k at 88.3 decode tok/s via ollama (signed run). Across 1 reproducible run on 1 model, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M5 Max (40-core GPU)
88.3 decode tok/s
qwen3.6-agent-q6k-ctx128k via ollama (Q6_K). see full run
qwen3.6-agent-q6k-ctx128k
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.32.1 | Q6_K | 88.29tok/s | 16.20tok/s | 6,915ms | r_udwgba7udqu |
Community folklore on M5 Max (40-core GPU)
94 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 70%
24.79tok/s — Gemma 3 27B on M5 Max via lm-studio
our signed data: M5 Max · Gemma 3 27B
“W|**+45W** (Thermal throttling risk)| |**Time to First Token (Prefill)**|89.83s|24.35s|**\~3.7x Faster**| |**Generation Speed**|23.16 tok/s|24.79 tok/s|**+1.63 tok/s** (Marginal)| |**Total Time**|847.87s|787.85s|**\~1 minute faster** overall| |**Prompt Tokens**|19,761|19,761|Same…”
- communityconfidence 70%
23.16tok/s — Gemma 3 27B on M5 Max via lm-studio
our signed data: M5 Max · Gemma 3 27B
“|< 70W|< 115W|**+45W** (Thermal throttling risk)| |**Time to First Token (Prefill)**|89.83s|24.35s|**\~3.7x Faster**| |**Generation Speed**|23.16 tok/s|24.79 tok/s|**+1.63 tok/s** (Marginal)| |**Total Time**|847.87s|787.85s|**\~1 minute faster** overall| |**Prompt Tokens**|19,761…”
- communityconfidence 70%
24.79tok/s — Gemma 3 27B on M5 Max via lm-studio
our signed data: M5 Max · Gemma 3 27B
“W|**+45W** (Thermal throttling risk)| |**Time to First Token (Prefill)**|89.83s|24.35s|**\~3.7x Faster**| |**Generation Speed**|23.16 tok/s|24.79 tok/s|**+1.63 tok/s** (Marginal)| |**Total Time**|847.87s|787.85s|**\~1 minute faster** overall| |**Prompt Tokens**|19,761|19,761|Same…”
- communityconfidence 70%
23.16tok/s — Gemma 3 27B on M5 Max via lm-studio
our signed data: M5 Max · Gemma 3 27B
“|< 70W|< 115W|**+45W** (Thermal throttling risk)| |**Time to First Token (Prefill)**|89.83s|24.35s|**\~3.7x Faster**| |**Generation Speed**|23.16 tok/s|24.79 tok/s|**+1.63 tok/s** (Marginal)| |**Total Time**|847.87s|787.85s|**\~1 minute faster** overall| |**Prompt Tokens**|19,761…”
- communityconfidence 60%
31.60tok/s — Qwen 2.5 72B on M5 Max via mlx
our signed data: M5 Max · Qwen 2.5 72B
“oretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community reports that llama.cpp has a known pe…”
- communityconfidence 60%
16.50tok/s — Qwen 2.5 72B on M5 Max via llama.cpp
our signed data: M5 Max · Qwen 2.5 72B
“onsistently achieves \~73-75% of theoretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community r…”
- communityconfidence 60%
31.60tok/s — Qwen 2.5 72B on M5 Max via mlx
our signed data: M5 Max · Qwen 2.5 72B
“oretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community reports that llama.cpp has a known pe…”
- communityconfidence 60%
16.50tok/s — Qwen 2.5 72B on M5 Max via llama.cpp
our signed data: M5 Max · Qwen 2.5 72B
“onsistently achieves \~73-75% of theoretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community r…”
- communityconfidence 60%
31.60tok/s — Qwen 2.5 72B on M5 Max via mlx
our signed data: M5 Max · Qwen 2.5 72B
“oretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community reports that llama.cpp has a known pe…”
- communityconfidence 60%
16.50tok/s — Qwen 2.5 72B on M5 Max via llama.cpp
our signed data: M5 Max · Qwen 2.5 72B
“onsistently achieves \~73-75% of theoretical maximum bandwidth utilization. # 2. MLX is Dramatically Faster for Qwen 3.5 * **llama.cpp**: 16.5 tok/s (Q6\_K, 21GB) * **MLX**: 31.6 tok/s (4bit, 16GB) * **Delta**: MLX is **92% faster** (1.9x speedup) This confirms the community r…”
Models measured on M5 Max (40-core GPU)
Common questions about M5 Max (40-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.