Skip to content

Running LLM inference on a dedicated 96 GB GPU: what fits and what it costs

A single 96 GB GPU is now enough to serve serious open-weight models to a product. This guide covers what fits, how fast it goes, and the arithmetic that decides between renting by the hour and a dedicated card.

9 minute read

What fits in 96 GB

Model memory is roughly parameters times bytes per parameter, plus the key-value cache for the context in flight. In 16-bit a parameter is two bytes; in 8-bit one; in 4-bit about half a byte. The cache grows with context length and concurrent requests.

Model size16-bit8-bit4-bitComfortable on 96 GB
7 to 9B16 to 18 GB8 to 9 GB4 to 5 GBYes, with very long contexts and high concurrency
27 to 32B54 to 64 GB27 to 32 GB14 to 16 GBYes in 16-bit with moderate context; 8-bit for high concurrency
70 to 72B140 GB70 GB35 to 40 GB8-bit fits with limited cache; 4-bit is the practical choice
Mixture-of-experts, 100B+ totalToo largeOften too largeDepends on active parametersCase by case
Approximate weights-only memory and headroom on a 96 GB card

The practical sweet spot for one card is a 30B-class model in 8- or 16-bit for quality, or a 70B-class model in 4-bit for capability. Both serve production traffic for a product with thousands of daily users.

Throughput to expect

With a modern inference server such as vLLM and continuous batching, a 30B model in 8-bit on this class of GPU typically delivers 1,500 to 3,000 output tokens per second aggregate across concurrent requests, with 30 to 60 tokens per second per stream. A 70B model in 4-bit lands around a third to half of that. Prompt processing is far faster than generation.

Dedicated versus hourly

Hourly GPU rental for a comparable card runs €2 to €4 per hour in 2026, so a GPU that is busy around the clock costs €1,500 to €3,000 per month in rental alone, before the host it runs on, the storage, the egress and the operations work. Serverless inference APIs charge per token and look cheap until volume arrives.

  • Rent by the hour when utilisation is below about 30 percent, or when you are still choosing a model.
  • Buy dedicated when the model serves traffic all day, when the data cannot leave your control, or when latency variance from shared infrastructure is a problem.
  • Infraexa Accel 96 is €6,780 per month with the host, 512 GB of system memory, NVMe, 20 TB of traffic, the inference server configured, monitoring, backups and engineers included.

What the operations layer does for a GPU

  • Drivers and CUDA pinned to versions the inference server supports, upgraded in a window with rollback.
  • GPU memory, utilisation, temperature and power on the dashboard, with Exa watching for thermal drift and memory leaks in long-running servers.
  • Model store on NVMe with nightly backups so a rebuild is a restore, not a re-download.
  • A load test against your prompts before handover, so the throughput numbers are yours and not a benchmark's.

Fine-tuning on the same card

Parameter-efficient fine-tuning with LoRA or QLoRA on a 7 to 30B model fits comfortably alongside inference if runs are scheduled in quiet hours. Full fine-tuning of 70B models needs more than one card. Checkpoints go to the NVMe and into the nightly backup.

Data residency

Prompts and completions never leave the server unless your application sends them somewhere. For products handling personal or regulated data, a dedicated GPU in a German or Finnish data centre, under a data processing agreement, is the simplest residency story there is.

Questions

Can I run two models on one card?

Yes, if their combined memory including cache fits. A 7B and a 30B model in 8-bit share a 96 GB card comfortably.

Can I add a second GPU later?

Multi-GPU builds are quoted individually. Many teams instead add a second Accel 96 and load-balance, which also gives redundancy.

Tell us what you run. We will tell you what it costs to run it properly.

A quote within one business day, from an engineer rather than a sales script. No setup fee, three-month minimum, delivery in 48 hours.