GPUcloud

GPU guide

Which GPU to fine-tune an LLM? Memory, machine and cost

Updated on

Short answer

A LoRA fine-tuning of an 8 billion parameter model fits on an RTX A6000 with 48 GB at CA$0.67/h, and a QLoRA of a 70 billion parameter model, with little headroom, on 1 × RTX A6000 (48 GB) at CA$0.67/h. Full fine-tuning needs about 16 bytes per parameter, that is 128 GB for 8 billion parameters before activations.

How much GPU memory to fine-tune an LLM

GPU memory to plan for to fine-tune an LLM, and the cheapest machine that has it
MethodParametersMemory before activationsCheapest machine that fitsPrice per hour
Full fine-tuning (AdamW)8 billion128 GB4 × RTX A6000 (192 GB)$2.68
Full fine-tuning (AdamW)70 billion1,120 GBno server online: on requeston request
LoRA (16-bit weights)8 billion16 GB1 × RTX A6000 (48 GB)$0.67
LoRA (16-bit weights)70 billion140 GB4 × RTX A6000 (192 GB)$2.68
QLoRA (4-bit weights)8 billion4 GB1 × RTX A6000 (48 GB)$0.67
QLoRA (4-bit weights)70 billion35 GB1 × RTX A6000 (48 GB)$0.67

Usual estimates: 16 bytes per parameter for full fine-tuning (weights, gradients, 32-bit copy and Adam states), 2 for LoRA, 0.5 for QLoRA. The machine is chosen with a 25% margin for activations; sequence length and batch size can need more. Prices in Canadian dollars, taxes extra.

Why such a gap: in full fine-tuning, every parameter keeps its gradients and the optimizer states. With LoRA, the model is frozen and only small adapters train; with QLoRA, the frozen model is also loaded in 4-bit. The CUDA out of memory guide explains how to reduce activations.

The machines to rent, with their availability

The machines of the table, with their availability
GPUGPUs per serverServer GPU memoryOn demand, per hourAvailability
NVIDIA RTX A6000148 GB$0.67
NVIDIA RTX A60004192 GB$2.68

Prices in Canadian dollars, taxes extra, read from the catalog when the page is built. Availability is read live; a machine out of stock shows the closest equivalent machine in stock.

How much a fine-tuning costs

Fine-tuning cost examples (assumed durations)
JobMachineAssumed durationOn-demand cost
Llama 3.1 8B, LoRA (16-bit weights)1 × RTX A6000 (48 GB)3 h$2.01
Llama 3.1 8B, QLoRA (4-bit weights)1 × RTX A6000 (48 GB)4 h$2.68
Llama 3.1 8B, Full fine-tuning (AdamW)4 × RTX A6000 (192 GB)10 h$26.80
Llama 3.1 70B, QLoRA (4-bit weights)1 × RTX A6000 (48 GB)24 h$16.08

The durations are hypotheses to show the calculation, not measurements: they depend on the dataset, the number of epochs, the sequence length and the code. Formula: hourly rate × hours, rounded to the cent, taxes extra.

For example, a LoRA of Llama 3.1 8B that runs 3 hours on an RTX A6000 costs CA$2.01. To know your real duration, run a short pass (a few hundred steps) and multiply.

Start a LoRA fine-tuning

Deploy the PyTorch (CUDA) template: PyTorch and CUDA are ready in /opt/pytorch. Then install the fine-tuning libraries and load the model with its LoRA adapters.

On the server (pip install transformers peft accelerate)
import torch
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM

# Gated model: accept its license on Hugging Face, then `huggingface-cli login`
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B", torch_dtype=torch.bfloat16, device_map="auto"
)
model.gradient_checkpointing_enable()
config = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
model = get_peft_model(model, config)
model.print_trainable_parameters()

For full fine-tuning on several GPUs (for example 4 × RTX A6000, 192 GB), shard the optimizer states across the GPUs with FSDP or DeepSpeed ZeRO. Full fine-tuning of a 70 billion parameter model (about 1,120 GB before activations) goes beyond a server sold online: reserve GPUs to discuss it.

Serve the fine-tuned model

Once the adapters are merged into the model, the vLLM one-click template serves it through an OpenAI-compatible API, from CA$0.67/h. The guide host your LLM on a GPU compares that cost with a per-token API.

Good to know before you launch

  • Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
  • A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
  • The minimum top-up is CA$25.00, paid by credit card.
  • Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.

Frequently asked questions

Which GPU to fine-tune Llama 3.1 8B?

With LoRA, an RTX A6000 with 48 GB (CA$0.67/h) is enough: the 16-bit weights take about 16 GB. Full fine-tuning needs about 128 GB before activations, so 4 × RTX A6000.

Can I fine-tune a 70 billion parameter model on a single server?

Yes with QLoRA: the 4-bit model takes about 35 GB, which fits on 1 × RTX A6000 (48 GB). It is tight: it fits in 4-bit with a small batch, short sequences and gradient checkpointing; for more headroom, take 2 × RTX A6000 or an 80 GB GPU. Full fine-tuning needs about 1,120 GB before activations, more than one server sold online.

LoRA or full fine-tuning?

LoRA first: it needs much less memory and time, and is often enough to adapt a model to a domain or an answer format. Full fine-tuning makes sense when LoRA is not enough.

How long does a fine-tuning take?

It depends on the dataset, the number of epochs and the sequence length. Measure a short pass, then multiply: the cost is the hourly rate times the hours.

Does my training data stay in Canada?

Yes. The servers and their disks stay in Canada, in Montreal. Delete the server at the end to erase its disks.

Need GPUs? We’ve got you.

Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.

Other guides