Up to 183 tok/s · 128K ctx
2× RTX 4090 48 GB total · 64-core CPU · 128 GB RAM · 4 TB NVMe
Single-stream LLM inference, measured on a bare-metal 2× RTX 4090 (48 GB total VRAM) machine, both GPUs working in parallel across the model layers.
Same models, same test method as the 1× RTX 4090 results — so the two pages are directly comparable.
| Model | Quant | Context | Generation |
|---|---|---|---|
| Llama 3.1 8B Instruct | Q4_K_M | 4K | 183 tok/s |
| Llama 3.1 8B Instruct | Q4_K_M | 32K | 160 tok/s |
| Llama 3.1 8B Instruct | Q4_K_M | 128K | 112 tok/s |
| Qwen3 14B | Q4_K_M | 4K | 128 tok/s |
| Qwen3 14B | Q4_K_M | 32K | 111 tok/s |
| Qwen3 14B | Q4_K_M | 128K | 77 tok/s |
| Gemma 3 12B | Q4_K_M | 4K | 139 tok/s |
| Gemma 3 12B | Q4_K_M | 32K | 121 tok/s |
| Gemma 3 12B | Q4_K_M | 128K | 84 tok/s |