RTX 5090 vs RTX 4090
Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs RTX 4090 — every number links to the run it came from.
Verdict
RTX 5090 and RTX 4090 both have submitted runs, but no single model has been measured on both yet — so there is no apples-to-apples row. Each side's measured decode tok/s is listed separately below (peaks of 247 and 195 tok/s). Submit a shared workload to turn this into a direct comparison.
RTX 5090 benchmarks — no shared model with RTX 4090 yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| Qwen2.5-7B-Instruct | 246.7tok/s | chat-short | r_3yn-4321hp- |
| Llama-3.1-8B-Instruct | 232.2tok/s | chat-short | r_kfrkg-vn376 |
| Qwen3.6-35B-A3B-Q4_K_M.gguf | 223.8tok/s | chat-short | r_10yvku19-xt |
| Qwen3.6-27B-Q4_K_M.gguf | 73.31tok/s | chat-short | r_wax_x2ryqhk |
| Qwen2.5-Coder-32B-Instruct | 70.96tok/s | concurrent-decode | r_v983y0y3r2u |
| Qwen3-32B | 69.45tok/s | chat-short | r_phvxm9dcak0 |
| gemma-2-9b-it | 69.45tok/s | chat-short | r_1_xl4zb5-xj |
| gpt-oss-20b | 69.42tok/s | chat-short | r_r9h57uts9lr |
| gemma-4-31B-it-Q4_K_M.gguf | 67.37tok/s | chat-short | r_plujbqnef08 |
RTX 4090 benchmarks — no shared model with RTX 5090 yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| gemma3 | 195.0tok/s | chat-short | r_dlanfbgym0h |
| qwen3-coder | 179.9tok/s | concurrent-decode | r_wjq32z47vlp |
| qwen2.5-coder | 161.1tok/s | concurrent-decode | r_mv8n8k9wu1e |
| llama3.1 | 154.4tok/s | chat-short | r_h1ub_1uxzdh |
| gpt-oss | 141.7tok/s | chat-long | r_iu2sfa9ykvw |
| deepseek-r1 | 133.8tok/s | agent-trace | r_fg77v2hhohb |
| glm-4.7-flash | 129.9tok/s | agent-trace | r_o636l3cc-rr |
| qwen3.6 | 44.18tok/s | agent-trace | r_h_659oy695r |
See also: RTX 5090 benchmarks · RTX 4090 benchmarks · All hardware · All models · Methodology