How many GPUs do you need for LLM inference?

howmanygpus.ai is a free GPU sizing calculator for large language models. Pick a model from the catalog (Llama, Qwen, DeepSeek, Mistral, gpt-oss and 130 more) or paste any HuggingFace id, set the context window, concurrency, quantization and serving framework, and get the VRAM per GPU, tokens per second, time to first token and cost per million tokens for every NVIDIA and AMD GPU that fits, with tensor, pipeline and expert parallelism worked out for you.

The estimates follow the roofline model behind NVIDIA's aiconfigurator and are calibrated against published vLLM, SGLang and TensorRT-LLM benchmarks.