Most on-prem LLM sizing conversations start at the model and work outward: pick a 70B, do the VRAM math, buy a box that fits it. That math is real, but it answers the wrong question first. The first question is who is using the system, how many of them are typing at the same instant, and what each of them is actually doing when they are.
Seat counts, concurrency ratios and the shape of the work (interactive chat, batch document extraction, long-running agents) drive GPU volume more than parameter count does, and they decide whether you need a cstmAI Desk, a Rack, or a Cluster long before you load a model. This piece lays out the framework, and the numbers behind each lever, so you can pick a system class before you open the AI hardware estimator.
Seats Are Not the Same as Load
A 500-seat deployment almost never means 500 simultaneous requests on the GPUs. In a steady business day, a reasonable active-use ratio for interactive chat sits somewhere between 5% and 15% of licensed seats at peak minute, and the ratio of peak-minute users who are actively holding a decode slot is lower still. The reason is simple: humans read.
A chat interface reads like a page, and a human consumes streamed text at roughly 4 to 5 tokens per second. Once the server clears that, the user is reading, thinking, or typing the next prompt. The GPU is not generating for them during any of that. So a 500-seat site at 10% peak active-use is 50 people with a session open, but perhaps 10 to 15 of them are mid-decode at any given tick. That number, not the seat count, is what the scheduler sees.
The shape of the workload changes the ratio sharply:
- Interactive chat and RAG: short bursts of decode (hundreds of tokens), long human pauses. Active-concurrency ratio of 2% to 5% of seats at peak is typical.
- Document AI and batch extraction: no human in the loop, so every queued job holds a slot until it finishes. The ratio is set by your queue depth and SLA, not seat count.
- Agent workloads: one user triggers many sequential calls. A single operator can hold 10 to 50 decode slots across a running plan, and the context grows across the session.
The third row is where sizing goes wrong most often. Agents multiply the effective concurrency per seat by an order of magnitude, and they do it with longer contexts, which eats VRAM through KV cache rather than weights.
The Three Numbers That Decide the Box
Before you look at a GPU, write down three numbers for your peak hour. Everything downstream falls out of these.
- Peak concurrent active sessions. Not seats. Sessions that have a request in flight or in the queue at the busy minute.
- Target tokens per second per user. For readable chat, 15 to 25 tok/s is comfortable, and above 30 tok/s is invisible to a human reader. For agents and batch, you tune for aggregate throughput instead.
- Median and P95 context length. This drives KV cache per session, which lives in the same VRAM as the weights.
Context length matters more than people expect. A 70B model that "fits" in 140GB at FP16 does not fit 100 concurrent users at 8K context, because 15 users at 32K context create roughly the same KV pressure as 60 users at 8K. Doubling your context budget does not double your VRAM headroom. It can quarter your safe concurrency.
On real hardware, published figures give you anchor points. Llama-3.3-70B on four H100s with tensor parallelism scales almost linearly up to 500 users on a 200-in, 200-out workload, peaking around 7,000 total tokens per second. That is one 4-GPU node absorbing a mid-sized department under short-context load. Lengthen the context, and the same node tops out far sooner.

How Workload Shape Maps to System Class
With the three numbers in hand, the choice between Desk, Rack and Cluster stops being about team size and becomes about active load.
cstmAI Desk: 1 to 2 GPUs
Right for pilots, individual power users, and small teams running an 8B to 14B model for coding, RAG and drafting. In practice this covers up to roughly 10 to 20 concurrent chat users at interactive speed on a quantized mid-sized model, or a single analyst running document extraction overnight. It is also the right place to prove a workflow before committing rack space and power. The first 90 days of an on-prem deployment break more pilots than any sizing math does, and a Desk is a cheap place to find out which assumptions were wrong.
cstmAI Rack: 4 to 8 GPUs
The default for a department or a single mid-market company serving a 32B to 70B model. One 4-GPU node with NVLink will carry a few hundred interactive seats at low double-digit concurrency, or a steady document AI pipeline running at tens of thousands of pages per day. Eight-GPU nodes are where agent workloads start to breathe, because the long contexts and multi-step plans need both VRAM and inter-GPU bandwidth. For enterprise serving, large models and many concurrent users need 80GB-plus VRAM on enterprise-class GPUs, which is the Rack's design center.
cstmAI Cluster: multi-node
Required when you need linear scale-out across replicas rather than one bigger box: hundreds of concurrent agent sessions, document AI at millions of pages per day, or an organization-wide chat service with hard latency SLOs. Clusters also let you dedicate nodes to workload shape (one for interactive chat, one for batch, one for fine-tuning) so a long batch job does not starve the trading desk at 09:30. The GPU cluster options are also where you build in headroom for the growth path, since Deloitte's December 2025 survey found 86% of enterprise respondents expect AI infrastructure budgets to more than triple over the next three years.
Agents and Batch Break the Simple Ratio
The clean "5% of seats at peak" rule holds for interactive chat. It falls apart the moment you add agents or heavy batch, and both are now the main reason mid-market buyers come to us.
An agentic session is not one request. A single user interaction with an AI agent can trigger dozens or hundreds of sequential inference calls, each consuming context that grows over the session. If you build custom agents for a 50-person operations team and each operator runs two concurrent plans averaging 40 calls each, your peak concurrent sessions on the GPU are not 50 but 100, and the context on each is drifting toward 32K as the plan accumulates tool output. That is a Rack-class problem on paper, often a two-Rack or small-Cluster problem in practice.
Batch is the opposite shape and the easier one. Document extraction jobs do not need interactive tok/s; they need queue depth and aggregate throughput. A single 8-GPU node with continuous batching can keep hundreds of concurrent sequences in flight because every weight is read from memory once and reused across the batch. If your workload is 80% nightly document AI and 20% daytime chat, you are better off sizing for the batch window and letting chat ride on the headroom, not the other way around.
A Framework You Can Apply Before the Estimator
Walk through this in order. It will not give you a part number, but it will put you in the right system class and stop you from either underbuying for agents or overbuying for a chat pilot.
Two sanity checks before you commit. First, utilization. An enterprise infrastructure audit found average GPU utilization in the tech industry sits at just 5%, which is the price of sizing to a theoretical peak that never arrives. Build for the P95 hour, not the worst minute of the worst day. Second, growth. The gap between a Desk and a Rack is one procurement cycle; the gap between a Rack and a Cluster is a facilities conversation. If your three-numbers calculation puts you within 30% of the next class up, start there.
Sizing Is a Workload Question, Not a GPU Question
The box you need is a function of how your people actually work, not a function of which model got attention this quarter. Count the sessions that will be in flight at the busy minute, not the seats on the license. Know whether you are serving readers, batch queues, or agents, because each one bends the ratio differently. Then size for the shape of the work, and let the model and the GPU fall out of that.
If you want help turning your three numbers into a configuration, the companion piece on GPU count covers the VRAM side of the math, and the four-size comparison lays out what each system class is built for. From there, a two-week discovery sprint is usually enough to confirm the shape before any hardware ships.



