GPU guide
CUDA out of memory: causes, fixes, and when to move to a GPU with more VRAM
Updated on
Short answer
First reduce the batch size, then try mixed precision, gradient checkpointing and quantization: these fixes often solve the error without changing cards. If the model still does not fit, rent more VRAM at GPU Cloud, in Canada: 48 GB from CA$0.67/h, 80 GB from CA$1.81/h, 96 GB from CA$2.48/h.
What the error means
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate … means PyTorch asks for more memory than is left free on the card. That memory holds the model weights, the gradients, the optimizer state and the activations, which grow with the batch size and the sequence length.
Measure before you fix: the "peak" line shows how much your training really asks for.
import torch
print(torch.cuda.get_device_name(0))
print(f"total : {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")
print(f"peak : {torch.cuda.max_memory_allocated() / 1e9:.1f} GB")
print(torch.cuda.memory_summary(abbreviated=True))The fixes, in order
1.Reduce the batch size
Halve
batch_sizeuntil the error goes away. To keep the same effective batch, accumulate gradients over several steps.accumulation_steps = 4 # effective batch = batch_size x 4 optimizer.zero_grad() for step, (inputs, targets) in enumerate(loader): loss = criterion(model(inputs), targets) / accumulation_steps loss.backward() if (step + 1) % accumulation_steps == 0: optimizer.step() optimizer.zero_grad()2.Use mixed precision
16-bit activations take about half the memory of 32-bit ones.
torch.autocastpicks the precision operation by operation.import torch # bf16 when the card supports it, otherwise fp16 with a GradScaler dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16 scaler = torch.amp.GradScaler("cuda", enabled=(dtype == torch.float16)) for inputs, targets in loader: optimizer.zero_grad() with torch.autocast(device_type="cuda", dtype=dtype): loss = criterion(model(inputs), targets) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update()3.Turn on gradient checkpointing
Intermediate activations are recomputed during backpropagation instead of being kept in memory: less VRAM, a bit more compute.
# Hugging Face Transformers model model.gradient_checkpointing_enable() # Your own PyTorch module from torch.utils.checkpoint import checkpoint def forward(self, x): x = checkpoint(self.block1, x, use_reentrant=False) return self.block2(x)4.Quantize the model
For inference or a QLoRA fine-tune, loading the weights in 4-bit makes them four times smaller than in 16-bit.
# pip install transformers accelerate bitsandbytes import torch from transformers import AutoModelForCausalLM, BitsAndBytesConfig quantization = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model = AutoModelForCausalLM.from_pretrained( "Qwen/Qwen2.5-7B-Instruct", quantization_config=quantization, device_map="auto", )5.Turn off gradients for inference
Without
torch.inference_mode(), PyTorch keeps what a backward pass would need, even though none will happen.model.eval() with torch.inference_mode(): outputs = model(inputs)6.Limit fragmentation
When the message reports a lot of memory "reserved but unallocated", let the allocator grow its segments.
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python train.py
How much VRAM the weights of a model take
| Model size | 16-bit (fp16, bf16) | 8-bit | 4-bit |
|---|---|---|---|
| 7 billion parameters | 14 GB | 7 GB | 3.5 GB |
| 8 billion parameters | 16 GB | 8 GB | 4 GB |
| 13 billion parameters | 26 GB | 13 GB | 6.5 GB |
| 34 billion parameters | 68 GB | 34 GB | 17 GB |
| 70 billion parameters | 140 GB | 70 GB | 35 GB |
Formula: number of parameters × bytes per parameter. This is a floor: activations, the KV cache, gradients and optimizer state come on top, and full training needs several times the size of the weights.
Example: a 7 billion parameter model takes 14 GB in 16-bit, 7 GB in 8-bit and 3.5 GB in 4-bit, before activations and the rest.
When to move to a GPU with more VRAM
If your model still does not fit after these fixes, or if quantization hurts the results too much, you need more memory. At GPU Cloud the smallest card has 48 GB, and a server goes up to 8 GPUs, 768 GB in total (8 × NVIDIA RTX PRO 6000).
| Memory per GPU | GPU | On demand, per hour | Spot, per hour |
|---|---|---|---|
| 48 GB | NVIDIA RTX A6000 | $0.67 | $0.54 |
| 48 GB | NVIDIA L40 | $1.34 | $1.07 |
| 80 GB | NVIDIA A100 PCIe | $1.81 | $1.45 |
| 80 GB | NVIDIA H100 PCIe | $3.35 | $2.68 |
| 96 GB | NVIDIA RTX PRO 6000 | $2.48 | $1.98 |
Prices in Canadian dollars, taxes extra, read from the catalog when the page is built.
With several GPUs, the model must be split across the cards (for example device_map="auto" with Transformers, or FSDP for training). The guide to run your script on a GPU shows how to copy your project and start training on the server.
Good to know before you launch
- Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
- A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
- The minimum top-up is CA$25.00, paid by credit card.
- Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.
Frequently asked questions
Why CUDA out of memory when nvidia-smi shows free memory?
PyTorch keeps the memory it already allocated in reserve, and that reserve can be fragmented: a large request fails even when small free blocks remain. Another process may also be using the card. expandable_segments:True limits fragmentation.
Does torch.cuda.empty_cache() fix the error?
Rarely. It returns reserved memory that no tensor uses any more, but it does not free tensors your code still references.
How much VRAM does a 7 billion parameter model need?
The weights alone take 14 GB in 16-bit, 7 GB in 8-bit and 3.5 GB in 4-bit. Add the activations and the KV cache for inference, and the gradients and optimizer state for training.
What is the most VRAM available at GPU Cloud?
Per GPU, 96 GB. Per server, up to 8 GPUs, 768 GB in total (8 × NVIDIA RTX PRO 6000).
How much does an hour with more VRAM cost?
48 GB: from CA$0.67/h; 80 GB: from CA$1.81/h; 96 GB: from CA$2.48/h, in Canadian dollars, taxes extra, billed by the hour with no commitment.
Need GPUs? We’ve got you.
Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.
Other guides
- Your AI says you need a GPU? Run your Python script on a GPU in minutes
- torch.cuda.is_available() returns False: causes and fixes (Mac, PC without NVIDIA, Docker)
- How much does a GPU cost per hour in Canada, in Canadian dollars?
- OpenAI API too expensive? Host your LLM on a GPU by the hour, in Canada
- Rent a GPU by the hour in Canada: prices, steps and billing
- GPU cloud in Montreal, Canada: NVIDIA GPU servers hosted in the country
- Rent an H100 in Canada: PCIe, NVLink or SXM, by the hour
- Which GPU to fine-tune an LLM? Memory, machine and cost
- Cheap GPU cloud: how to pay as little as possible for a GPU