hardware.co · AI hardware for on-prem LLM deploymentsVendor-neutral · Quote onlyBuilt to spec
CSTM-EST · REV AcstmAI™ · Hardware · Estimator

Size your on-prem LLM hardware.

Pick a model, tell us how many people will use it, and see the GPUs, memory, power and rack space it takes. A starting point in a minute; the quote measures your real workload.

Fig. 01 · Lathe area of a machine shop, Paterson, NJ1994
AI hardware estimator

Workload in, configuration out.

Every input changes the result live, and the address bar keeps your settings so you can share an estimate. No prices, by policy: every system is quoted.

WorkloadForm HWCO-E · Rev A
Weight precision
Active at the busiest moment ≈ 20 requests at once
Context window per request
Speed per user Tokens / second
Fine-tuning on the same hardware
Deployment
GPU vendor
Estimates from public spec-sheet figures. No prices: every system is quoted.
Starting configurationSized by speed per user
cstmAI™ Rack →

8 × NVIDIA RTX PRO 6000 Blackwell2 model copies of 4 GPUs

GPU memory
768 GB
Speed per user
≈ 30 tok/s
Total output
≈ 608 tok/s
Requests at once
20
Power, GPUs + chassis
≈ 6.3 kW
Space
4U of rack
GPU memory per model copy · 4 × 96 GB98 of 353 GB usable
  • Weights 71 GB
  • KV cache 13 GB
  • Runtime 10 GB
  • Search models 4 GB
  • Headroom 286 GB
Llama 3.3 70B · 71 GB of weights · 1.34 GB KV cache per request
Other configurations that meet the same workload
Also fitsSizeGPU memoryPer userPower
16 × NVIDIA RTX 6000 AdaCluster768 GB≈ 33 tok/s≈ 7.2 kW
8 × NVIDIA H100 SXMRack640 GB≈ 38 tok/s≈ 9.1 kW
8 × NVIDIA H200 SXMRack1,128 GB≈ 33 tok/s≈ 9.1 kW
8 × AMD Instinct MI300XRack1,536 GB≈ 36 tok/s≈ 9.5 kW
8 × AMD Instinct MI325XRack2,048 GB≈ 41 tok/s≈ 12 kW
How the estimate works

The arithmetic, in the open.

The same first-pass sizing our engineers do before they measure anything.

MEM

Memory first

Weights take parameters × bytes per parameter: 2 at 16-bit, 1 at 8-bit, about 0.55 at 4-bit. Each request adds a KV cache that grows with its context; we assume requests fill about half their window. Runtime overhead is added, and only 92% of GPU memory is counted as usable.

BW

Then speed

Generating each token reads the weights in play plus every request's KV cache, so speed per user is mostly memory bandwidth divided by bytes read. We assume 60% of peak bandwidth, less a penalty when one model spans several GPUs, and hold 15% back for reading prompts.

MoE

Mixture-of-experts

Models like gpt-oss, Qwen3 235B-A22B and DeepSeek V3 read only the experts each token uses. With many requests at once most experts get touched anyway; the estimate models that, so large batches don't look unrealistically cheap.

TUNE

Fine-tuning

Training runs off-hours on the same GPUs, so it sets a memory floor rather than adding GPUs. LoRA adapters need little beyond the frozen model; a full fine-tune needs roughly 16 bytes per parameter for weights, gradients and optimizer state.

GPUs in the estimator

Specified from more than one vendor.

Public spec-sheet figures, rounded. The quote chooses from whatever fits and can be delivered that quarter, including parts not listed here.

GPU figures the estimator uses
GPUMemoryBandwidthPowerGoes in
NVIDIA RTX 6000 Ada48 GB0.96 TB/s300 WDesk, Servers
NVIDIA RTX PRO 6000 Blackwell96 GB1.79 TB/s600 WDesk, Servers
NVIDIA L40S48 GB0.86 TB/s350 WServers
NVIDIA H100 SXM80 GB3.35 TB/s700 W8-GPU servers
NVIDIA H200 SXM141 GB4.8 TB/s700 W8-GPU servers
NVIDIA B200180 GB7.7 TB/s1000 W8-GPU servers
AMD Instinct MI300X192 GB5.3 TB/s750 W8-GPU servers
AMD Instinct MI325X256 GB6 TB/s1000 W8-GPU servers
NVIDIA L424 GB0.3 TB/s72 WEdge
NVIDIA Jetson AGX Thor (128 GB)110 GB0.27 TB/s130 WEdge
FAQ

Sizing questions, answered.

How much GPU memory does an LLM need?

Roughly the parameter count times the bytes per parameter, plus room for each conversation's context. A 70B model needs about 140 GB at 16-bit, about 70 GB at 8-bit and under 40 GB at 4-bit, before any context. Twenty users with long documents can add another 20–30 GB, which is why the estimator asks how many people are active at once.

How many GPUs do I need to run Llama 70B on-prem?

For a few users at 4-bit, one 96 GB workstation GPU can do it. For a couple of hundred employees with 10% active at peak, 8-bit weights and reading-speed responses, expect a 4–8 GPU server. The estimator shows the trade-offs for your numbers; discovery measures them on your own task.

Is this a price estimate?

No. We don't publish prices, because GPU supply and pricing move week to week and the right part list depends on your site. The estimator gives a starting configuration; the quote turns it into a dated part list with a price.

How accurate is it?

Good enough to tell a workstation from a server and a server from a cluster. Real throughput depends on the serving software, the prompt mix and the model's architecture details, so we benchmark your workload on candidate hardware before quoting. Expect the real number to land within a size of the estimate.

Why does the estimate change GPU vendor when I change the workload?

Because the cheapest way to meet a workload depends on whether it's limited by memory, bandwidth or compute. Big-memory accelerators win on large models; workstation cards win on small ones. We're vendor-neutral, so the estimator ranks everything the same way.

Get a quote

Spec your system.

Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.

Form CSTM-Q · Quote only