Up to 102 tok/s · 128K ctx
1× RTX 4090 24 GB · 64-core CPU · 64 GB RAM · 4 TB NVMe
Single-stream LLM inference, measured on a bare-metal 1× RTX 4090 (24 GB) machine — the “Single GPU” preset as configured on our pricing page.
All layers offloaded to the GPU, no other workload running on the machine. This is the speed one user gets — not batched multi-tenant throughput.
| Model | Quant | Context | Generation |
|---|---|---|---|
| Llama 3.1 8B Instruct | Q4_K_M | 4K | 102 tok/s |
| Llama 3.1 8B Instruct | Q4_K_M | 32K | 89 tok/s |
| Llama 3.1 8B Instruct | Q4_K_M | 128K | 61 tok/s |
| Qwen3 14B | Q4_K_M | 4K | 71 tok/s |
| Qwen3 14B | Q4_K_M | 32K | 62 tok/s |
| Qwen3 14B | Q4_K_M | 128K | 43 tok/s |
| Gemma 3 12B | Q4_K_M | 4K | 78 tok/s |
| Gemma 3 12B | Q4_K_M | 32K | 68 tok/s |
| Gemma 3 12B | Q4_K_M | 128K | 47 tok/s |