GPU guides for developers
Your AI says you need a GPU? These guides answer the questions that follow: what to fix first, which GPU to rent, what it costs in Canadian dollars and how to run your code. Prices are read from the GPU Cloud catalog, from CA$0.67/h in Canada.
Your AI says you need a GPU? Run your Python script on a GPU in minutes
Updated on
Rent an NVIDIA GPU server by the hour: at GPU Cloud, an RTX A6000 with 48 GB costs CA$0.67/h in Canada, with PyTorch and CUDA already installed by the one-click template. Copy your project with rsync, run your script over SSH, then hibernate or delete the server to stop the charges.
Read the guideCUDA out of memory: causes, fixes, and when to move to a GPU with more VRAM
Updated on
First reduce the batch size, then try mixed precision, gradient checkpointing and quantization: these fixes often solve the error without changing cards. If the model still does not fit, rent more VRAM at GPU Cloud, in Canada: 48 GB from CA$0.67/h, 80 GB from CA$1.81/h, 96 GB from CA$2.48/h.
Read the guidetorch.cuda.is_available() returns False: causes and fixes (Mac, PC without NVIDIA, Docker)
Updated on
On a Mac or a PC without an NVIDIA card, CUDA does not exist:
Read the guidetorch.cuda.is_available()always returnsFalseand no setting changes that, so you need an NVIDIA GPU, for example rented by the hour at GPU Cloud from CA$0.67/h in Canada. On an NVIDIA machine, the usual causes are a PyTorch build without CUDA, a missing or outdated driver, or a container started without--gpus all.How much does a GPU cost per hour in Canada, in Canadian dollars?
Updated on
At GPU Cloud, a 1-GPU server costs CA$0.67/h (NVIDIA RTX A6000, 48 GB) to CA$3.35/h (NVIDIA H100 PCIe, 80 GB) on demand, and from CA$0.54/h on Spot, taxes extra. Billing is hourly from prepaid credit with no commitment, and a hibernated server costs $0.02 per hour (about $14.60 per month) and keeps its disk.
Read the guideOpenAI API too expensive? Host your LLM on a GPU by the hour, in Canada
Updated on
Deploy the one-click OpenAI-compatible vLLM template: you get an HTTPS /v1 address and a key, and your existing OpenAI code works by changing only
Read the guidebase_urlandapi_key, on a GPU in Canada from CA$0.67/h. The server is paid by the hour, not by the token: it beats the API once your volume passes the break-even point, computed below with a formula where you plug in your own numbers.Rent a GPU by the hour in Canada: prices, steps and billing
Updated on
At GPU Cloud, you rent an NVIDIA GPU server by the hour, from CA$0.67/h (RTX A6000) to CA$3.35/h (H100 PCIe) for 1 GPU, with no subscription, from prepaid credit paid by credit card. You pick the GPU in the console, the server starts in Montreal within minutes, then you hibernate or delete it when you are done.
Read the guideGPU cloud in Montreal, Canada: NVIDIA GPU servers hosted in the country
Updated on
GPU Cloud rents NVIDIA GPU servers hosted in Montreal (region canada-montreal), and their disks never leave Canada. Prices are shown and billed in Canadian dollars, from CA$0.67/h for an RTX A6000, taxes extra.
Read the guideRent an H100 in Canada: PCIe, NVLink or SXM, by the hour
Updated on
At GPU Cloud, an NVIDIA H100 PCIe (80 GB) rents by the hour from CA$3.35/h, and 8-GPU H100 servers (H100 NVLink and H100 SXM) go up to CA$34.27/h, taxes extra. The servers are in Montreal, billed in Canadian dollars, and an H100 out of stock shows the closest equivalent machine in stock.
Read the guideWhich GPU to fine-tune an LLM? Memory, machine and cost
Updated on
A LoRA fine-tuning of an 8 billion parameter model fits on an RTX A6000 with 48 GB at CA$0.67/h, and a QLoRA of a 70 billion parameter model, with little headroom, on 1 × RTX A6000 (48 GB) at CA$0.67/h. Full fine-tuning needs about 16 bytes per parameter, that is 128 GB for 8 billion parameters before activations.
Read the guideCheap GPU cloud: how to pay as little as possible for a GPU
Updated on
The cheapest GPU at GPU Cloud is an NVIDIA RTX A6000 (48 GB) at CA$0.67/h, or CA$0.54/h on Spot, taxes extra. To pay as little as possible, take the smallest memory that fits your model, and hibernate or delete the server as soon as you are done.
Read the guide