News icon

Kimi K3 is now available on Runpod

GPU Benchmarks Directory

Runpod's GPU comparison tool benchmarks more than 20 GPUs on the metrics that decide an inference deployment: VRAM, tokens per second per user, time to first token, and cost per unit of output. Every figure is a measured run across 33 models spanning LLM, image, video, and speech workloads, not a spec-sheet estimate. Whether you are serving a 7B model on a 24 GB card or a 35B model on Blackwell, start with what fits, then compare what it costs to run.

Purple glow background

GPU Comparison Tool

Real benchmark results, not spec sheets. Pick a model and compare GPUs on speed, latency and cost per output — across LLM, image, video and speech workloads.

Llama 3.1 405B
vLLM
SGLang
Llama 4 Maverick
vLLM
SGLang
Llama 4 Scout
vLLM
TensorRT-LLM
DeepSeek-R1
vLLM
SGLang
DeepSeek-V3.2 / V4
vLLM
SGLang
DeepSeek-R1-Distill
vLLM
TensorRT-LLM
DeepSeek-Coder-V2 / V2.5
vLLM
TensorRT-LLM
DeepSeek-VL2
vLLM
SGLang
DeepSeek-Math-7B
vLLM
Llama.cpp
DeepSeek-OCR
vLLM
TensorRT-LLM
Devstral 2
vLLM
SGLang
Flux 2 Max
vLLM
TensorRT-LLM
Flux.1 Dev
vLLM
TensorRT-LLM
Flux.2 Pro
vLLM
TensorRT-LLM
Gemma 3
vLLM
TensorRT-LLM
GPT-OSS 120B
vLLM
SGLang
GPT-OSS 20B
vLLM
TensorRT-LLM
Kimi K2.5
vLLM
SGLang
Mistral Large 3
vLLM
TensorRT-LLM
Mixtral 8x22B
vLLM
SGLang
Qwen2.5 / Qwen3
vLLM
SGLang
Qwen3-Coder-480B
vLLM
SGLang
Qwen3.5
vLLM
SGLang
QwQ-32B
vLLM
SGLang
Qwen2.5-7B-Instruct
No items found.
Qwen2.5-14B-Instruct (FP8)
No items found.
Qwen3.5-9B (FP8)
No items found.
Qwen3.6-27B (FP8)
No items found.
Qwen3.5-9B
No items found.
Qwen3-VL-8B-Instruct
No items found.
Llama-3.1-8B-Instruct
No items found.
Qwen3.8-27B (FP8)
No items found.
Qwen3.6-35B-A3B (FP8)
No items found.
Qwen3.8-27B
No items found.
Qwen3-8B
No items found.
Qwen2.5-VL-7B-Instruct
No items found.
Gemma 4 31B IT
No items found.
Qwen3.6-27B
No items found.
Gemma 4 26B-A4B IT
No items found.
High-Throughput Inference
Fine-Tuning (LoRA / QLoRA)
Full Parameter Training
Multi-Agent
Batch Image / Video Gen
Single GPU
Dual GPU
8× GPU SXM/HGX Node
Multi-Node (32×)
Multi-Node (64×)
Multi-Node (128×)
vLLM
TensorRT-LLM
SGLang
Llama.cpp
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU
VRAM
Tokens/Sec Per User
Time to First Token
$ / 1M Output Tokens
Price (Secure/Hr)
Availability
B300
SGLang
TensorRT-LLM
288
—
—
—
7.89
H200
SGLang
141
—
—
—
4.59
B200
SGLang
TensorRT-LLM
180
—
—
—
6.79
H100
No items found.
80
—
—
—
2.89
H100 NVL
vLLM
SGLang
94
—
—
—
3.19
H100 PCIe
vLLM
80
—
—
—
2.89
H100 SXM
vLLM
SGLang
80
—
—
—
3.49
A100 PCIe
vLLM
80
—
—
—
1.59
A100 SXM
vLLM
80
—
—
—
1.59
Pro 6000 MIG 48GB
No items found.
48
—
—
—
1.09
L40S
vLLM
48
—
—
—
1.09
RTX 6000 Ada
vLLM
48
—
—
—
0.84
A40
vLLM
48
—
—
—
0.49
L40
vLLM
48
—
—
—
0.82
RTX A6000
vLLM
48
—
—
—
0.53
RTX 5090
vLLM
Llama.cpp
32
—
—
—
0.99
Pro 6000 MIG 24GB
No items found.
24
—
—
—
0.59
L4
Llama.cpp
24
—
—
—
0.49
RTX 3090
Llama.cpp
24
—
—
—
0.5
RTX 4090
Llama.cpp
vLLM
24
—
—
—
0.74
RTX A5000
Llama.cpp
24
—
—
—
0.27
RTX A4000
Llama.cpp
16
—
—
—
0.25
RTX 2000 Ada
Llama.cpp
16
—
—
—
0.24
MI300X
No items found.
192
—
—
—
2.39

Benchmarks by Workload

The tool above compares cards interactively. These four guides publish the full measured results for each workload, including the figures that change the answer.

Questions? Answers.

There is no single answer, because the fastest card and the cheapest card are almost never the same one. Decide which you are optimizing for first.

For lowest cost per token, the older workstation cards win and it is not close. On Qwen2.5-7B-Instruct, an RTX A5000 at $0.27/hr serves tokens for $0.64 per million. An H200 on the same model costs $10.20 per million – roughly 16 times more – while delivering about four and a half times the tokens per second. If your workload tolerates 40-odd tokens per second per user, the cheap card is the rational choice.

For lowest latency per user, Blackwell leads. B300 and B200 top the throughput ranking on 62 of the 68 LLM benchmarks in this tool, reaching 303 tokens per second on Qwen3.5-9B (FP8) and 272 on Qwen2.5-7B-Instruct.

For the middle, the RTX 4090 at $0.74/hr and the RTX 5090 at $0.99/hr are the most balanced cards on the page. On Qwen2.5-7B-Instruct the 5090 returns 101 tokens per second at $2.24 per million – about 60% of an A100 SXM's speed at 62% of its cost per token.

Sort the table by $ / 1M output tokens for your own model rather than taking a general recommendation. The ranking changes with model size and quantization.

Match VRAM to the model first, then optimize within what fits.

Match VRAM to the model first, then optimize within what fits.
Model size Fits on Typical pick
7B–9B, quantized 24 GB – RTX A5000, RTX 3090, RTX 4090, L4 RTX A5000 for cost, RTX 4090 for balance
7B–14B, FP8/BF16 32–48 GB – RTX 5090, A40, L40, RTX 6000 Ada RTX 5090 or RTX 6000 Ada
27B–35B 48–80 GB – RTX 6000 Ada, A100, H100 RTX 6000 Ada at $2.11 per million output tokens on Qwen3.8-27B (FP8)
Above 35B 141 GB+ – H200, B200, B300 Not yet benchmarked here; size on VRAM

Two results worth knowing before you default to an H100. On Qwen2.5-7B-Instruct an H100 SXM delivers 165 tokens per second but costs $7.78 per million output tokens, placing it among the most expensive cards on the page per unit of work. And on Qwen3.6-35B-A3B (FP8), an L40 at $0.82/hr returns 99 tokens per second for $1.86 per million – a mixture-of-experts model rewards cheap VRAM more than it rewards bandwidth.

Sparse and dense models of the same nominal size behave very differently. Qwen3.6-27B (FP8) tops out at 105 tokens per second on a B200; the 35B-A3B sparse model on the same class of hardware reaches 260. Benchmark the model you will actually run.

On fine-tuning and training: the figures in this tool are inference benchmarks and should not be read as training guidance. For sizing, a LoRA or QLoRA job needs roughly the model's weights plus adapter and activation overhead, so a 7B at 4-bit fits comfortably in 24 GB while a 27B at BF16 needs 80 GB or more. Full-parameter fine-tuning needs several times the weight memory once gradients and optimizer state are included, which is multi-GPU territory at any size above about 13B.

For general-purpose work – prototyping, mixed model sizes, a bit of everything – the useful question is what you can leave running.

Start on an RTX A5000 or RTX 4090. At $0.27 and $0.74 per hour they carry most experimentation, and both have the lowest cost per token in their VRAM class. The A5000 is the cheapest card per million output tokens on four of the fifteen chat benchmarks in this tool.

Move up when VRAM forces you to, not when throughput disappoints you. The usual reason to leave a 24 GB card is a model that will not fit, not one that runs slowly. A40 and L40 at 48 GB ($0.49 and $0.82/hr) are the cheapest step up.

Reserve Blackwell and Hopper for work where wall-clock time has a price – a production endpoint under load, an experiment you are waiting on, a deadline. At $4.59 to $7.89 per hour, H200, B200 and B300 pay for themselves through speed or not at all.

One caveat on all of this: these benchmarks run at 1 req/s, which is a light load. Cards with large VRAM and high bandwidth pull further ahead as concurrency rises, because they hold bigger batches. If you are serving many users at once, treat the cost-per-token column here as a floor for the cheap cards rather than a promise.

It depends on the workload, and the ranking genuinely inverts between them.

Image generation. B200 is both the fastest and the cheapest option on several models – 0.57 seconds per image on FLUX.1-schnell at $0.00107, and 1.51 seconds on Z-Image-Turbo. That combination is unusual and worth checking for your model. Where you are optimizing purely for cost, an RTX A5000 produces a FLUX.2 Klein 9B image for $0.00049 and an L40S an SDXL image for $0.00067. Note that the RTX 4090 is not a strong diffusion card in these runs: on Qwen-Image 2512 (ComfyUI fp8) it takes 44.7 seconds per image, and it is the most expensive card per image on that model.

Video generation. B200 leads on the heavy models – 119 seconds for Wan 2.2 T2V A14B and 37.6 seconds for Wan 2.2 TI2V-5B. For cost rather than speed, an RTX Pro 6000 produces Wan 2.2 T2V A14B at $0.215 per video against the B200's $0.237, and on the distilled LTX-2.3 22B an RTX 4090 costs $0.00335 per video. Distilled and quantized video models change the hardware answer completely – check which variant you are running.

Speech to text. This is the one workload where a mid-range card wins outright. An L40S transcribes at 58.3 times real time on Whisper large-v3-turbo (int8_float16), ahead of the H100 SXM's 55.5 times on the same model family. For cost, an RTX A5000 runs an audio hour for $0.0071.

Vision-language. Qwen3-VL-8B-Instruct and Qwen2.5-VL-7B-Instruct both run on consumer cards – an RTX A5000 serves Qwen3-VL-8B at $0.65 per million output tokens. Switch the workload selector to Vision Q&A to compare them properly.

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background

Browse all GPU comparisons

Explore every GPU comparison in the Runpod directory.

Show all comparisonsHide all comparisons