Running LLM inference on a dedicated 96 GB GPU: what fits and what it costs
A single 96 GB GPU is now enough to serve serious open-weight models to a product. This guide covers what fits, how fast it goes, and the arithmetic that decides between renting by the hour and a dedicated card.
9 minute read
What fits in 96 GB
Model memory is roughly parameters times bytes per parameter, plus the key-value cache for the context in flight. In 16-bit a parameter is two bytes; in 8-bit one; in 4-bit about half a byte. The cache grows with context length and concurrent requests.
| Model size | 16-bit | 8-bit | 4-bit | Comfortable on 96 GB |
|---|---|---|---|---|
| 7 to 9B | 16 to 18 GB | 8 to 9 GB | 4 to 5 GB | Yes, with very long contexts and high concurrency |
| 27 to 32B | 54 to 64 GB | 27 to 32 GB | 14 to 16 GB | Yes in 16-bit with moderate context; 8-bit for high concurrency |
| 70 to 72B | 140 GB | 70 GB | 35 to 40 GB | 8-bit fits with limited cache; 4-bit is the practical choice |
| Mixture-of-experts, 100B+ total | Too large | Often too large | Depends on active parameters | Case by case |
The practical sweet spot for one card is a 30B-class model in 8- or 16-bit for quality, or a 70B-class model in 4-bit for capability. Both serve production traffic for a product with thousands of daily users.
Throughput to expect
With a modern inference server such as vLLM and continuous batching, a 30B model in 8-bit on this class of GPU typically delivers 1,500 to 3,000 output tokens per second aggregate across concurrent requests, with 30 to 60 tokens per second per stream. A 70B model in 4-bit lands around a third to half of that. Prompt processing is far faster than generation.
Dedicated versus hourly
Hourly GPU rental for a comparable card runs €2 to €4 per hour in 2026, so a GPU that is busy around the clock costs €1,500 to €3,000 per month in rental alone, before the host it runs on, the storage, the egress and the operations work. Serverless inference APIs charge per token and look cheap until volume arrives.
- Rent by the hour when utilisation is below about 30 percent, or when you are still choosing a model.
- Buy dedicated when the model serves traffic all day, when the data cannot leave your control, or when latency variance from shared infrastructure is a problem.
- Infraexa Accel 96 is €6,780 per month with the host, 512 GB of system memory, NVMe, 20 TB of traffic, the inference server configured, monitoring, backups and engineers included.
What the operations layer does for a GPU
- Drivers and CUDA pinned to versions the inference server supports, upgraded in a window with rollback.
- GPU memory, utilisation, temperature and power on the dashboard, with Exa watching for thermal drift and memory leaks in long-running servers.
- Model store on NVMe with nightly backups so a rebuild is a restore, not a re-download.
- A load test against your prompts before handover, so the throughput numbers are yours and not a benchmark's.
Fine-tuning on the same card
Parameter-efficient fine-tuning with LoRA or QLoRA on a 7 to 30B model fits comfortably alongside inference if runs are scheduled in quiet hours. Full fine-tuning of 70B models needs more than one card. Checkpoints go to the NVMe and into the nightly backup.
Data residency
Prompts and completions never leave the server unless your application sends them somewhere. For products handling personal or regulated data, a dedicated GPU in a German or Finnish data centre, under a data processing agreement, is the simplest residency story there is.
Questions
Can I run two models on one card?
Yes, if their combined memory including cache fits. A 7B and a 30B model in 8-bit share a 96 GB card comfortably.
Can I add a second GPU later?
Multi-GPU builds are quoted individually. Many teams instead add a second Accel 96 and load-balance, which also gives redundancy.
Tell us what you run. We will tell you what it costs to run it properly.
A quote within one business day, from an engineer rather than a sales script. No setup fee, three-month minimum, delivery in 48 hours.