GPU Benchmarks Directory
Runpod's GPU comparison tool benchmarks more than 20 GPUs on the metrics that decide an inference deployment: VRAM, tokens per second per user, time to first token, and cost per unit of output. Every figure is a measured run across 33 models spanning LLM, image, video, and speech workloads, not a spec-sheet estimate. Whether you are serving a 7B model on a 24 GB card or a 35B model on Blackwell, start with what fits, then compare what it costs to run.

GPU Comparison Tool
Real benchmark results, not spec sheets. Pick a model and compare GPUs on speed, latency and cost per output — across LLM, image, video and speech workloads.
Benchmarks by Workload
The tool above compares cards interactively. These four guides publish the full measured results for each workload, including the figures that change the answer.
GPU Comparison Directory
Side-by-side specs and benchmarks for every GPU Runpod offers.
Blackwell Architecture (NVIDIA B-Series)
Hopper Architecture (NVIDIA H-Series)
NVIDIA Ampere Architecture
NVIDIA L-Series
NVIDIA RTX Series, Current Gen
NVIDIA RTX Series, Previous & Legacy Gen
FAQs
Questions? Answers.
There is no single answer, because the fastest card and the cheapest card are almost never the same one. Decide which you are optimizing for first.
For lowest cost per token, the older workstation cards win and it is not close. On Qwen2.5-7B-Instruct, an RTX A5000 at $0.27/hr serves tokens for $0.64 per million. An H200 on the same model costs $10.20 per million – roughly 16 times more – while delivering about four and a half times the tokens per second. If your workload tolerates 40-odd tokens per second per user, the cheap card is the rational choice.
For lowest latency per user, Blackwell leads. B300 and B200 top the throughput ranking on 62 of the 68 LLM benchmarks in this tool, reaching 303 tokens per second on Qwen3.5-9B (FP8) and 272 on Qwen2.5-7B-Instruct.
For the middle, the RTX 4090 at $0.74/hr and the RTX 5090 at $0.99/hr are the most balanced cards on the page. On Qwen2.5-7B-Instruct the 5090 returns 101 tokens per second at $2.24 per million – about 60% of an A100 SXM's speed at 62% of its cost per token.
Sort the table by $ / 1M output tokens for your own model rather than taking a general recommendation. The ranking changes with model size and quantization.
Match VRAM to the model first, then optimize within what fits.
Two results worth knowing before you default to an H100. On Qwen2.5-7B-Instruct an H100 SXM delivers 165 tokens per second but costs $7.78 per million output tokens, placing it among the most expensive cards on the page per unit of work. And on Qwen3.6-35B-A3B (FP8), an L40 at $0.82/hr returns 99 tokens per second for $1.86 per million – a mixture-of-experts model rewards cheap VRAM more than it rewards bandwidth.
Sparse and dense models of the same nominal size behave very differently. Qwen3.6-27B (FP8) tops out at 105 tokens per second on a B200; the 35B-A3B sparse model on the same class of hardware reaches 260. Benchmark the model you will actually run.
On fine-tuning and training: the figures in this tool are inference benchmarks and should not be read as training guidance. For sizing, a LoRA or QLoRA job needs roughly the model's weights plus adapter and activation overhead, so a 7B at 4-bit fits comfortably in 24 GB while a 27B at BF16 needs 80 GB or more. Full-parameter fine-tuning needs several times the weight memory once gradients and optimizer state are included, which is multi-GPU territory at any size above about 13B.
For general-purpose work – prototyping, mixed model sizes, a bit of everything – the useful question is what you can leave running.
Start on an RTX A5000 or RTX 4090. At $0.27 and $0.74 per hour they carry most experimentation, and both have the lowest cost per token in their VRAM class. The A5000 is the cheapest card per million output tokens on four of the fifteen chat benchmarks in this tool.
Move up when VRAM forces you to, not when throughput disappoints you. The usual reason to leave a 24 GB card is a model that will not fit, not one that runs slowly. A40 and L40 at 48 GB ($0.49 and $0.82/hr) are the cheapest step up.
Reserve Blackwell and Hopper for work where wall-clock time has a price – a production endpoint under load, an experiment you are waiting on, a deadline. At $4.59 to $7.89 per hour, H200, B200 and B300 pay for themselves through speed or not at all.
One caveat on all of this: these benchmarks run at 1 req/s, which is a light load. Cards with large VRAM and high bandwidth pull further ahead as concurrency rises, because they hold bigger batches. If you are serving many users at once, treat the cost-per-token column here as a floor for the cheap cards rather than a promise.
It depends on the workload, and the ranking genuinely inverts between them.
Image generation. B200 is both the fastest and the cheapest option on several models – 0.57 seconds per image on FLUX.1-schnell at $0.00107, and 1.51 seconds on Z-Image-Turbo. That combination is unusual and worth checking for your model. Where you are optimizing purely for cost, an RTX A5000 produces a FLUX.2 Klein 9B image for $0.00049 and an L40S an SDXL image for $0.00067. Note that the RTX 4090 is not a strong diffusion card in these runs: on Qwen-Image 2512 (ComfyUI fp8) it takes 44.7 seconds per image, and it is the most expensive card per image on that model.
Video generation. B200 leads on the heavy models – 119 seconds for Wan 2.2 T2V A14B and 37.6 seconds for Wan 2.2 TI2V-5B. For cost rather than speed, an RTX Pro 6000 produces Wan 2.2 T2V A14B at $0.215 per video against the B200's $0.237, and on the distilled LTX-2.3 22B an RTX 4090 costs $0.00335 per video. Distilled and quantized video models change the hardware answer completely – check which variant you are running.
Speech to text. This is the one workload where a mid-range card wins outright. An L40S transcribes at 58.3 times real time on Whisper large-v3-turbo (int8_float16), ahead of the H100 SXM's 55.5 times on the same model family. For cost, an RTX A5000 runs an audio hour for $0.0071.
Vision-language. Qwen3-VL-8B-Instruct and Qwen2.5-VL-7B-Instruct both run on consumer cards – an RTX A5000 serves Qwen3-VL-8B at $0.65 per million output tokens. Switch the workload selector to Vision Q&A to compare them properly.