Pick a model, tell us how many people will use it, and see the GPUs, memory, power and rack space it takes. A starting point in a minute; the quote measures your real workload.
Fig. 01 · Lathe area of a machine shop, Paterson, NJ1994
AI hardware estimator
Workload in, configuration out.
Every input changes the result live, and the address bar keeps your settings so you can share an estimate. No prices, by policy: every system is quoted.
The same first-pass sizing our engineers do before they measure anything.
MEM
Memory first
Weights take parameters × bytes per parameter: 2 at 16-bit, 1 at 8-bit, about 0.55 at 4-bit. Each request adds a KV cache that grows with its context; we assume requests fill about half their window. Runtime overhead is added, and only 92% of GPU memory is counted as usable.
BW
Then speed
Generating each token reads the weights in play plus every request's KV cache, so speed per user is mostly memory bandwidth divided by bytes read. We assume 60% of peak bandwidth, less a penalty when one model spans several GPUs, and hold 15% back for reading prompts.
MoE
Mixture-of-experts
Models like gpt-oss, Qwen3 235B-A22B and DeepSeek V3 read only the experts each token uses. With many requests at once most experts get touched anyway; the estimate models that, so large batches don't look unrealistically cheap.
TUNE
Fine-tuning
Training runs off-hours on the same GPUs, so it sets a memory floor rather than adding GPUs. LoRA adapters need little beyond the frozen model; a full fine-tune needs roughly 16 bytes per parameter for weights, gradients and optimizer state.
GPUs in the estimator
Specified from more than one vendor.
Public spec-sheet figures, rounded. The quote chooses from whatever fits and can be delivered that quarter, including parts not listed here.
GPU figures the estimator uses
GPU
Memory
Bandwidth
Power
Goes in
NVIDIA RTX 6000 Ada
48 GB
0.96 TB/s
300 W
Desk, Servers
NVIDIA RTX PRO 6000 Blackwell
96 GB
1.79 TB/s
600 W
Desk, Servers
NVIDIA L40S
48 GB
0.86 TB/s
350 W
Servers
NVIDIA H100 SXM
80 GB
3.35 TB/s
700 W
8-GPU servers
NVIDIA H200 SXM
141 GB
4.8 TB/s
700 W
8-GPU servers
NVIDIA B200
180 GB
7.7 TB/s
1000 W
8-GPU servers
AMD Instinct MI300X
192 GB
5.3 TB/s
750 W
8-GPU servers
AMD Instinct MI325X
256 GB
6 TB/s
1000 W
8-GPU servers
NVIDIA L4
24 GB
0.3 TB/s
72 W
Edge
NVIDIA Jetson AGX Thor (128 GB)
110 GB
0.27 TB/s
130 W
Edge
FAQ
Sizing questions, answered.
How much GPU memory does an LLM need?
Roughly the parameter count times the bytes per parameter, plus room for each conversation's context. A 70B model needs about 140 GB at 16-bit, about 70 GB at 8-bit and under 40 GB at 4-bit, before any context. Twenty users with long documents can add another 20–30 GB, which is why the estimator asks how many people are active at once.
How many GPUs do I need to run Llama 70B on-prem?
For a few users at 4-bit, one 96 GB workstation GPU can do it. For a couple of hundred employees with 10% active at peak, 8-bit weights and reading-speed responses, expect a 4–8 GPU server. The estimator shows the trade-offs for your numbers; discovery measures them on your own task.
Is this a price estimate?
No. We don't publish prices, because GPU supply and pricing move week to week and the right part list depends on your site. The estimator gives a starting configuration; the quote turns it into a dated part list with a price.
How accurate is it?
Good enough to tell a workstation from a server and a server from a cluster. Real throughput depends on the serving software, the prompt mix and the model's architecture details, so we benchmark your workload on candidate hardware before quoting. Expect the real number to land within a size of the estimate.
Why does the estimate change GPU vendor when I change the workload?
Because the cheapest way to meet a workload depends on whether it's limited by memory, bandwidth or compute. Big-memory accelerators win on large models; workstation cards win on small ones. We're vendor-neutral, so the estimator ranks everything the same way.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.