GPU guide
Which GPU to fine-tune an LLM? Memory, machine and cost
Updated on
Short answer
A LoRA fine-tuning of an 8 billion parameter model fits on an RTX A6000 with 48 GB at CA$0.67/h, and a QLoRA of a 70 billion parameter model, with little headroom, on 1 × RTX A6000 (48 GB) at CA$0.67/h. Full fine-tuning needs about 16 bytes per parameter, that is 128 GB for 8 billion parameters before activations.
How much GPU memory to fine-tune an LLM
| Method | Parameters | Memory before activations | Cheapest machine that fits | Price per hour |
|---|---|---|---|---|
| Full fine-tuning (AdamW) | 8 billion | 128 GB | 4 × RTX A6000 (192 GB) | $2.68 |
| Full fine-tuning (AdamW) | 70 billion | 1,120 GB | no server online: on request | on request |
| LoRA (16-bit weights) | 8 billion | 16 GB | 1 × RTX A6000 (48 GB) | $0.67 |
| LoRA (16-bit weights) | 70 billion | 140 GB | 4 × RTX A6000 (192 GB) | $2.68 |
| QLoRA (4-bit weights) | 8 billion | 4 GB | 1 × RTX A6000 (48 GB) | $0.67 |
| QLoRA (4-bit weights) | 70 billion | 35 GB | 1 × RTX A6000 (48 GB) | $0.67 |
Usual estimates: 16 bytes per parameter for full fine-tuning (weights, gradients, 32-bit copy and Adam states), 2 for LoRA, 0.5 for QLoRA. The machine is chosen with a 25% margin for activations; sequence length and batch size can need more. Prices in Canadian dollars, taxes extra.
Why such a gap: in full fine-tuning, every parameter keeps its gradients and the optimizer states. With LoRA, the model is frozen and only small adapters train; with QLoRA, the frozen model is also loaded in 4-bit. The CUDA out of memory guide explains how to reduce activations.
The machines to rent, with their availability
| GPU | GPUs per server | Server GPU memory | On demand, per hour | Availability |
|---|---|---|---|---|
| NVIDIA RTX A6000 | 1 | 48 GB | $0.67 | |
| NVIDIA RTX A6000 | 4 | 192 GB | $2.68 |
Prices in Canadian dollars, taxes extra, read from the catalog when the page is built. Availability is read live; a machine out of stock shows the closest equivalent machine in stock.
How much a fine-tuning costs
| Job | Machine | Assumed duration | On-demand cost |
|---|---|---|---|
| Llama 3.1 8B, LoRA (16-bit weights) | 1 × RTX A6000 (48 GB) | 3 h | $2.01 |
| Llama 3.1 8B, QLoRA (4-bit weights) | 1 × RTX A6000 (48 GB) | 4 h | $2.68 |
| Llama 3.1 8B, Full fine-tuning (AdamW) | 4 × RTX A6000 (192 GB) | 10 h | $26.80 |
| Llama 3.1 70B, QLoRA (4-bit weights) | 1 × RTX A6000 (48 GB) | 24 h | $16.08 |
The durations are hypotheses to show the calculation, not measurements: they depend on the dataset, the number of epochs, the sequence length and the code. Formula: hourly rate × hours, rounded to the cent, taxes extra.
For example, a LoRA of Llama 3.1 8B that runs 3 hours on an RTX A6000 costs CA$2.01. To know your real duration, run a short pass (a few hundred steps) and multiply.
Start a LoRA fine-tuning
Deploy the PyTorch (CUDA) template: PyTorch and CUDA are ready in /opt/pytorch. Then install the fine-tuning libraries and load the model with its LoRA adapters.
import torch
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM
# Gated model: accept its license on Hugging Face, then `huggingface-cli login`
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B", torch_dtype=torch.bfloat16, device_map="auto"
)
model.gradient_checkpointing_enable()
config = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
model = get_peft_model(model, config)
model.print_trainable_parameters()For full fine-tuning on several GPUs (for example 4 × RTX A6000, 192 GB), shard the optimizer states across the GPUs with FSDP or DeepSpeed ZeRO. Full fine-tuning of a 70 billion parameter model (about 1,120 GB before activations) goes beyond a server sold online: reserve GPUs to discuss it.
Serve the fine-tuned model
Once the adapters are merged into the model, the vLLM one-click template serves it through an OpenAI-compatible API, from CA$0.67/h. The guide host your LLM on a GPU compares that cost with a per-token API.
Good to know before you launch
- Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
- A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
- The minimum top-up is CA$25.00, paid by credit card.
- Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.
Frequently asked questions
Which GPU to fine-tune Llama 3.1 8B?
With LoRA, an RTX A6000 with 48 GB (CA$0.67/h) is enough: the 16-bit weights take about 16 GB. Full fine-tuning needs about 128 GB before activations, so 4 × RTX A6000.
Can I fine-tune a 70 billion parameter model on a single server?
Yes with QLoRA: the 4-bit model takes about 35 GB, which fits on 1 × RTX A6000 (48 GB). It is tight: it fits in 4-bit with a small batch, short sequences and gradient checkpointing; for more headroom, take 2 × RTX A6000 or an 80 GB GPU. Full fine-tuning needs about 1,120 GB before activations, more than one server sold online.
LoRA or full fine-tuning?
LoRA first: it needs much less memory and time, and is often enough to adapt a model to a domain or an answer format. Full fine-tuning makes sense when LoRA is not enough.
How long does a fine-tuning take?
It depends on the dataset, the number of epochs and the sequence length. Measure a short pass, then multiply: the cost is the hourly rate times the hours.
Does my training data stay in Canada?
Yes. The servers and their disks stay in Canada, in Montreal. Delete the server at the end to erase its disks.
Need GPUs? We’ve got you.
Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.
Other guides
- Your AI says you need a GPU? Run your Python script on a GPU in minutes
- CUDA out of memory: causes, fixes, and when to move to a GPU with more VRAM
- torch.cuda.is_available() returns False: causes and fixes (Mac, PC without NVIDIA, Docker)
- How much does a GPU cost per hour in Canada, in Canadian dollars?
- OpenAI API too expensive? Host your LLM on a GPU by the hour, in Canada
- Rent a GPU by the hour in Canada: prices, steps and billing
- GPU cloud in Montreal, Canada: NVIDIA GPU servers hosted in the country
- Rent an H100 in Canada: PCIe, NVLink or SXM, by the hour
- Cheap GPU cloud: how to pay as little as possible for a GPU