GPUcloud

GPU guide

CUDA out of memory: causes, fixes, and when to move to a GPU with more VRAM

Updated on

Short answer

First reduce the batch size, then try mixed precision, gradient checkpointing and quantization: these fixes often solve the error without changing cards. If the model still does not fit, rent more VRAM at GPU Cloud, in Canada: 48 GB from CA$0.67/h, 80 GB from CA$1.81/h, 96 GB from CA$2.48/h.

What the error means

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate … means PyTorch asks for more memory than is left free on the card. That memory holds the model weights, the gradients, the optimizer state and the activations, which grow with the batch size and the sequence length.

Measure before you fix: the "peak" line shows how much your training really asks for.

import torch

print(torch.cuda.get_device_name(0))
print(f"total : {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")
print(f"peak  : {torch.cuda.max_memory_allocated() / 1e9:.1f} GB")
print(torch.cuda.memory_summary(abbreviated=True))

The fixes, in order

  1. 1.Reduce the batch size

    Halve batch_size until the error goes away. To keep the same effective batch, accumulate gradients over several steps.

    accumulation_steps = 4  # effective batch = batch_size x 4
    
    optimizer.zero_grad()
    for step, (inputs, targets) in enumerate(loader):
        loss = criterion(model(inputs), targets) / accumulation_steps
        loss.backward()
        if (step + 1) % accumulation_steps == 0:
            optimizer.step()
            optimizer.zero_grad()
  2. 2.Use mixed precision

    16-bit activations take about half the memory of 32-bit ones. torch.autocast picks the precision operation by operation.

    import torch
    
    # bf16 when the card supports it, otherwise fp16 with a GradScaler
    dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
    scaler = torch.amp.GradScaler("cuda", enabled=(dtype == torch.float16))
    
    for inputs, targets in loader:
        optimizer.zero_grad()
        with torch.autocast(device_type="cuda", dtype=dtype):
            loss = criterion(model(inputs), targets)
        scaler.scale(loss).backward()
        scaler.step(optimizer)
        scaler.update()
  3. 3.Turn on gradient checkpointing

    Intermediate activations are recomputed during backpropagation instead of being kept in memory: less VRAM, a bit more compute.

    # Hugging Face Transformers model
    model.gradient_checkpointing_enable()
    
    # Your own PyTorch module
    from torch.utils.checkpoint import checkpoint
    
    def forward(self, x):
        x = checkpoint(self.block1, x, use_reentrant=False)
        return self.block2(x)
  4. 4.Quantize the model

    For inference or a QLoRA fine-tune, loading the weights in 4-bit makes them four times smaller than in 16-bit.

    # pip install transformers accelerate bitsandbytes
    import torch
    from transformers import AutoModelForCausalLM, BitsAndBytesConfig
    
    quantization = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
    )
    model = AutoModelForCausalLM.from_pretrained(
        "Qwen/Qwen2.5-7B-Instruct",
        quantization_config=quantization,
        device_map="auto",
    )
  5. 5.Turn off gradients for inference

    Without torch.inference_mode(), PyTorch keeps what a backward pass would need, even though none will happen.

    model.eval()
    with torch.inference_mode():
        outputs = model(inputs)
  6. 6.Limit fragmentation

    When the message reports a lot of memory "reserved but unallocated", let the allocator grow its segments.

    PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python train.py

How much VRAM the weights of a model take

Memory taken by the weights alone, by precision
Model size16-bit (fp16, bf16)8-bit4-bit
7 billion parameters14 GB7 GB3.5 GB
8 billion parameters16 GB8 GB4 GB
13 billion parameters26 GB13 GB6.5 GB
34 billion parameters68 GB34 GB17 GB
70 billion parameters140 GB70 GB35 GB

Formula: number of parameters × bytes per parameter. This is a floor: activations, the KV cache, gradients and optimizer state come on top, and full training needs several times the size of the weights.

Example: a 7 billion parameter model takes 14 GB in 16-bit, 7 GB in 8-bit and 3.5 GB in 4-bit, before activations and the rest.

When to move to a GPU with more VRAM

If your model still does not fit after these fixes, or if quantization hurts the results too much, you need more memory. At GPU Cloud the smallest card has 48 GB, and a server goes up to 8 GPUs, 768 GB in total (8 × NVIDIA RTX PRO 6000).

Memory per GPU at GPU Cloud, and price of a 1-GPU server
Memory per GPUGPUOn demand, per hourSpot, per hour
48 GBNVIDIA RTX A6000$0.67$0.54
48 GBNVIDIA L40$1.34$1.07
80 GBNVIDIA A100 PCIe$1.81$1.45
80 GBNVIDIA H100 PCIe$3.35$2.68
96 GBNVIDIA RTX PRO 6000$2.48$1.98

Prices in Canadian dollars, taxes extra, read from the catalog when the page is built.

With several GPUs, the model must be split across the cards (for example device_map="auto" with Transformers, or FSDP for training). The guide to run your script on a GPU shows how to copy your project and start training on the server.

Good to know before you launch

  • Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
  • A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
  • The minimum top-up is CA$25.00, paid by credit card.
  • Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.

Frequently asked questions

Why CUDA out of memory when nvidia-smi shows free memory?

PyTorch keeps the memory it already allocated in reserve, and that reserve can be fragmented: a large request fails even when small free blocks remain. Another process may also be using the card. expandable_segments:True limits fragmentation.

Does torch.cuda.empty_cache() fix the error?

Rarely. It returns reserved memory that no tensor uses any more, but it does not free tensors your code still references.

How much VRAM does a 7 billion parameter model need?

The weights alone take 14 GB in 16-bit, 7 GB in 8-bit and 3.5 GB in 4-bit. Add the activations and the KV cache for inference, and the gradients and optimizer state for training.

What is the most VRAM available at GPU Cloud?

Per GPU, 96 GB. Per server, up to 8 GPUs, 768 GB in total (8 × NVIDIA RTX PRO 6000).

How much does an hour with more VRAM cost?

48 GB: from CA$0.67/h; 80 GB: from CA$1.81/h; 96 GB: from CA$2.48/h, in Canadian dollars, taxes extra, billed by the hour with no commitment.

Need GPUs? We’ve got you.

Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.

Other guides