Benchmark

2× RTX 4090 — LLM Inference Benchmarks

Up to 183 tok/s · 128K ctx

2× RTX 4090 48 GB total · 64-core CPU · 128 GB RAM · 4 TB NVMe

Single-stream LLM inference, measured on a bare-metal 2× RTX 4090 (48 GB total VRAM) machine, both GPUs working in parallel across the model layers.

Same models, same test method as the 1× RTX 4090 results — so the two pages are directly comparable.

How we test

Results — generation speed

Model Quant Context Generation
Llama 3.1 8B Instruct Q4_K_M 4K 183 tok/s
Llama 3.1 8B Instruct Q4_K_M 32K 160 tok/s
Llama 3.1 8B Instruct Q4_K_M 128K 112 tok/s
Qwen3 14B Q4_K_M 4K 128 tok/s
Qwen3 14B Q4_K_M 32K 111 tok/s
Qwen3 14B Q4_K_M 128K 77 tok/s
Gemma 3 12B Q4_K_M 4K 139 tok/s
Gemma 3 12B Q4_K_M 32K 121 tok/s
Gemma 3 12B Q4_K_M 128K 84 tok/s

Notes