Benchmark

1× RTX 4090 — LLM Inference Benchmarks

Up to 102 tok/s · 128K ctx

1× RTX 4090 24 GB · 64-core CPU · 64 GB RAM · 4 TB NVMe

Single-stream LLM inference, measured on a bare-metal 1× RTX 4090 (24 GB) machine — the “Single GPU” preset as configured on our pricing page.

All layers offloaded to the GPU, no other workload running on the machine. This is the speed one user gets — not batched multi-tenant throughput.

How we test

Results — generation speed

Model Quant Context Generation
Llama 3.1 8B Instruct Q4_K_M 4K 102 tok/s
Llama 3.1 8B Instruct Q4_K_M 32K 89 tok/s
Llama 3.1 8B Instruct Q4_K_M 128K 61 tok/s
Qwen3 14B Q4_K_M 4K 71 tok/s
Qwen3 14B Q4_K_M 32K 62 tok/s
Qwen3 14B Q4_K_M 128K 43 tok/s
Gemma 3 12B Q4_K_M 4K 78 tok/s
Gemma 3 12B Q4_K_M 32K 68 tok/s
Gemma 3 12B Q4_K_M 128K 47 tok/s

Notes