hardware.co · AI hardware for on-prem LLM deploymentsVendor-neutral · Quote onlyBuilt to spec
hardware.co · October 11, 2026

Sizing GPU Capacity by Team Size and Active Use

How to size on-prem GPU volume by seat count, concurrency and active use patterns, so your AI hardware matches real team load without overbuying.

By Eric Lamanna

A row of desks with workers connected by varying threads of light to a central server rack, suggesting concurrent GPU load.

Most on-prem LLM sizing conversations start at the model and work outward: pick a 70B, do the VRAM math, buy a box that fits it. That math is real, but it answers the wrong question first. The first question is who is using the system, how many of them are typing at the same instant, and what each of them is actually doing when they are.

Seat counts, concurrency ratios and the shape of the work (interactive chat, batch document extraction, long-running agents) drive GPU volume more than parameter count does, and they decide whether you need a cstmAI Desk, a Rack, or a Cluster long before you load a model. This piece lays out the framework, and the numbers behind each lever, so you can pick a system class before you open the AI hardware estimator.

Seats Are Not the Same as Load

A 500-seat deployment almost never means 500 simultaneous requests on the GPUs. In a steady business day, a reasonable active-use ratio for interactive chat sits somewhere between 5% and 15% of licensed seats at peak minute, and the ratio of peak-minute users who are actively holding a decode slot is lower still. The reason is simple: humans read.

A chat interface reads like a page, and a human consumes streamed text at roughly 4 to 5 tokens per second. Once the server clears that, the user is reading, thinking, or typing the next prompt. The GPU is not generating for them during any of that. So a 500-seat site at 10% peak active-use is 50 people with a session open, but perhaps 10 to 15 of them are mid-decode at any given tick. That number, not the seat count, is what the scheduler sees.

The shape of the workload changes the ratio sharply:

  • Interactive chat and RAG: short bursts of decode (hundreds of tokens), long human pauses. Active-concurrency ratio of 2% to 5% of seats at peak is typical.
  • Document AI and batch extraction: no human in the loop, so every queued job holds a slot until it finishes. The ratio is set by your queue depth and SLA, not seat count.
  • Agent workloads: one user triggers many sequential calls. A single operator can hold 10 to 50 decode slots across a running plan, and the context grows across the session.

The third row is where sizing goes wrong most often. Agents multiply the effective concurrency per seat by an order of magnitude, and they do it with longer contexts, which eats VRAM through KV cache rather than weights.

Where GPU Decode Time Actually Goes
Where GPU Decode Time Actually GoesAgent sessions (10 operators): 65; Document AI batch queue: 20; Interactive chat and RAG: 12; Fine-tune or eval jobs: 3Agent sessions (10 operators)65 · 65%Document AI batch queue20 · 20%Interactive chat and…12 · 12%Fine-tune or eval… 3
Illustrative share of concurrent decode slots per 100 seats, by workload shape. Agent sessions dominate even at small headcounts because one user holds many slots. Illustrative: a visual comparison, not measured data.

The Three Numbers That Decide the Box

Before you look at a GPU, write down three numbers for your peak hour. Everything downstream falls out of these.

  1. Peak concurrent active sessions. Not seats. Sessions that have a request in flight or in the queue at the busy minute.
  2. Target tokens per second per user. For readable chat, 15 to 25 tok/s is comfortable, and above 30 tok/s is invisible to a human reader. For agents and batch, you tune for aggregate throughput instead.
  3. Median and P95 context length. This drives KV cache per session, which lives in the same VRAM as the weights.

Context length matters more than people expect. A 70B model that "fits" in 140GB at FP16 does not fit 100 concurrent users at 8K context, because 15 users at 32K context create roughly the same KV pressure as 60 users at 8K. Doubling your context budget does not double your VRAM headroom. It can quarter your safe concurrency.

On real hardware, published figures give you anchor points. Llama-3.3-70B on four H100s with tensor parallelism scales almost linearly up to 500 users on a 200-in, 200-out workload, peaking around 7,000 total tokens per second. That is one 4-GPU node absorbing a mid-sized department under short-context load. Lengthen the context, and the same node tops out far sooner.

A stopwatch beside three stacked blocks of increasing size representing context length filling a fixed memory budget.
The Three Numbers to Write Down First
50
Peak concurrent active sessions
~5–10% of seats · GPU count and replicas
20
Target tokens/sec per user
15–25 tok/s for chat · Per-GPU throughput budget
32,000
P95 context length (tokens)
8K baseline, 32K for agents · KV cache and safe concurrency
Illustrative anchor values for the three inputs that drive system class, each on its own scale. Illustrative: a visual comparison, not measured data.

How Workload Shape Maps to System Class

With the three numbers in hand, the choice between Desk, Rack and Cluster stops being about team size and becomes about active load.

cstmAI Desk: 1 to 2 GPUs

Right for pilots, individual power users, and small teams running an 8B to 14B model for coding, RAG and drafting. In practice this covers up to roughly 10 to 20 concurrent chat users at interactive speed on a quantized mid-sized model, or a single analyst running document extraction overnight. It is also the right place to prove a workflow before committing rack space and power. The first 90 days of an on-prem deployment break more pilots than any sizing math does, and a Desk is a cheap place to find out which assumptions were wrong.

cstmAI Rack: 4 to 8 GPUs

The default for a department or a single mid-market company serving a 32B to 70B model. One 4-GPU node with NVLink will carry a few hundred interactive seats at low double-digit concurrency, or a steady document AI pipeline running at tens of thousands of pages per day. Eight-GPU nodes are where agent workloads start to breathe, because the long contexts and multi-step plans need both VRAM and inter-GPU bandwidth. For enterprise serving, large models and many concurrent users need 80GB-plus VRAM on enterprise-class GPUs, which is the Rack's design center.

cstmAI Cluster: multi-node

Required when you need linear scale-out across replicas rather than one bigger box: hundreds of concurrent agent sessions, document AI at millions of pages per day, or an organization-wide chat service with hard latency SLOs. Clusters also let you dedicate nodes to workload shape (one for interactive chat, one for batch, one for fine-tuning) so a long batch job does not starve the trading desk at 09:30. The GPU cluster options are also where you build in headroom for the growth path, since Deloitte's December 2025 survey found 86% of enterprise respondents expect AI infrastructure budgets to more than triple over the next three years.

System Class by Workload Shape and Load
System Class by Workload Shape and LoadcstmAI Desk (pilots, small teams): 1; cstmAI Rack (department, mid-market): 15; cstmAI Cluster (enterprise, agents at scale): 20015011,0011,5002,000cstmAI Desk (pilots,…1–20cstmAI Rack (departme…15–250cstmAI Cluster (enter…200–2,000
Illustrative overlap of where each cstmAI class is the right starting point. Bars span the active-session range each class comfortably serves; overlaps are where workload shape, not seat count, decides. Illustrative: a visual comparison, not measured data.

Agents and Batch Break the Simple Ratio

The clean "5% of seats at peak" rule holds for interactive chat. It falls apart the moment you add agents or heavy batch, and both are now the main reason mid-market buyers come to us.

An agentic session is not one request. A single user interaction with an AI agent can trigger dozens or hundreds of sequential inference calls, each consuming context that grows over the session. If you build custom agents for a 50-person operations team and each operator runs two concurrent plans averaging 40 calls each, your peak concurrent sessions on the GPU are not 50 but 100, and the context on each is drifting toward 32K as the plan accumulates tool output. That is a Rack-class problem on paper, often a two-Rack or small-Cluster problem in practice.

Batch is the opposite shape and the easier one. Document extraction jobs do not need interactive tok/s; they need queue depth and aggregate throughput. A single 8-GPU node with continuous batching can keep hundreds of concurrent sequences in flight because every weight is read from memory once and reused across the batch. If your workload is 80% nightly document AI and 20% daytime chat, you are better off sizing for the batch window and letting chat ride on the headroom, not the other way around.

A Framework You Can Apply Before the Estimator

Walk through this in order. It will not give you a part number, but it will put you in the right system class and stop you from either underbuying for agents or overbuying for a chat pilot.

Six Steps From Seats to a System Class
Six Steps From Seats to a System ClassCount peak-minute active sessions, not seats: 1; Classify workload: chat, batch, or agent: 2; Set target tokens per second per user: 3; Record median and P95 context length: 4; Multiply by KV cache per session for VRAM: 5; Pick Desk, Rack or Cluster, then size GPUs: 61Count peak-minuteactive sessions,not seats2Classify workload:chat, batch, oragent3Set target tokensper second peruser4Record median andP95 context length5Multiply by KVcache per sessionfor VRAM6Pick Desk, Rack orCluster, then sizeGPUs
Illustrative order of operations before you open the estimator. Illustrative: a visual comparison, not measured data.

Two sanity checks before you commit. First, utilization. An enterprise infrastructure audit found average GPU utilization in the tech industry sits at just 5%, which is the price of sizing to a theoretical peak that never arrives. Build for the P95 hour, not the worst minute of the worst day. Second, growth. The gap between a Desk and a Rack is one procurement cycle; the gap between a Rack and a Cluster is a facilities conversation. If your three-numbers calculation puts you within 30% of the next class up, start there.

Sizing Is a Workload Question, Not a GPU Question

The box you need is a function of how your people actually work, not a function of which model got attention this quarter. Count the sessions that will be in flight at the busy minute, not the seats on the license. Know whether you are serving readers, batch queues, or agents, because each one bends the ratio differently. Then size for the shape of the work, and let the model and the GPU fall out of that.

If you want help turning your three numbers into a configuration, the companion piece on GPU count covers the VRAM side of the math, and the four-size comparison lays out what each system class is built for. From there, a two-week discovery sprint is usually enough to confirm the shape before any hardware ships.

Get a quote

Spec your system.

Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.

Form CSTM-Q · Quote only