GPU guide
OpenAI API too expensive? Host your LLM on a GPU by the hour, in Canada
Updated on
Short answer
Deploy the one-click OpenAI-compatible vLLM template: you get an HTTPS /v1 address and a key, and your existing OpenAI code works by changing only base_url and api_key, on a GPU in Canada from CA$0.67/h. The server is paid by the hour, not by the token: it beats the API once your volume passes the break-even point, computed below with a formula where you plug in your own numbers.
What the vLLM template installs
- An OpenAI-compatible API (
/v1) served over HTTPS, protected by a key generated for you and shown in the console. - The model of your choice among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3, served with a 8,192-token context. It needs a GPU with at least 24 GB.
- Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3 require accepting their license on the model hub, then adding an access token to /etc/gpucloud/app.env on the server.
- Only the
/v1and/healthroutes are exposed; you keep full SSH access to the server.
The steps
1.Deploy the vLLM template
In the console, pick the OpenAI-compatible vLLM template, the language model and a GPU, then confirm the price. Installation takes a few minutes; the console shows the progress.
2.Get the address and the key
When the server answers, the console shows its HTTPS address and the generated API key.
3.Check that the server answers
/healthanswers once the model is loaded;/v1/modelslists the served model.Check the server export VLLM_API_KEY="<key shown in the console>" curl https://SERVER_ADDRESS/health curl https://SERVER_ADDRESS/v1/models -H "Authorization: Bearer $VLLM_API_KEY"4.Change two lines in your code
Keep the OpenAI SDK: replace
base_urlandapi_key, and use the name of the served model.Your existing OpenAI code: only two lines change # pip install openai from openai import OpenAI client = OpenAI( base_url="https://SERVER_ADDRESS/v1", # shown in the console api_key="VLLM_KEY", # generated key, shown in the console ) response = client.chat.completions.create( model="Qwen/Qwen2.5-7B-Instruct", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content)
What the server costs
| GPU | VRAM | Per hour | Office hours (176 h) | Full month by the hour (730 h) | Monthly term |
|---|---|---|---|---|---|
| NVIDIA RTX A6000 | 48 GB | $0.67 | $117.92 | $489.10 | $474.50 |
| NVIDIA L40 | 48 GB | $1.34 | $235.84 | $978.20 | $941.70 |
| NVIDIA A100 PCIe | 80 GB | $1.81 | $318.56 | $1,321.30 | $1,270.20 |
| NVIDIA RTX PRO 6000 | 96 GB | $2.48 | $436.48 | $1,810.40 | $1,737.40 |
| NVIDIA H100 PCIe | 80 GB | $3.35 | $589.60 | $2,445.50 | $2,343.30 |
Office hours: 8 h a day for 22 days. Prices in Canadian dollars, taxes extra, read from the catalog when the page is built.
The server costs the same whether you send it one request or a million. Outside your usage hours, hibernate it to keep the downloaded model on disk without paying for the GPU.
The break-even: per-token API or GPU by the hour
An API bills every token; your server bills every hour. The math fits in three formulas:
cost per million tokens = hourly price ÷ (tokens per second × 3,600) × 1,000,000
break-even throughput (tokens/s) = hourly price × 1,000,000 ÷ (3,600 × API price per million)
break-even volume (million tokens per month) = monthly GPU cost ÷ API price per millionExample with two hypotheses, which are neither measurements nor a real price list: an API at CA$2.00 per million tokens, and a sustained throughput of 200 tokens per second. On an RTX A6000 at CA$0.67/h, a million tokens then costs CA$0.93. The server comes out ahead as soon as it produces more than 93 tokens per second on average over the hour.
Per month, with the same hypotheses: in office hours (176 h, $117.92), the break-even is 59 million tokens per month; with the monthly term ($474.50), it is 237.3 million tokens per month.
Replace the hypotheses with your numbers: the API price you pay, and the throughput you measure on your server with your model and your real requests.
import time
start = time.perf_counter()
response = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": PROMPT}],
max_tokens=512,
)
elapsed = time.perf_counter() - start
print(response.usage.completion_tokens / elapsed, "tokens/s")When to stay on the API
- Your volume is low or irregular: per token, you do not pay for idle hours.
- You need a model that is not open weights, or a longer context than the one-click template serves.
- You do not want to run a server, even one installed for you.
Your requests stay in Canada
The server runs in Montreal: requests, responses and the model stay on your machine, in Canada. See the Data sovereignty page.
Good to know before you launch
- Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
- A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
- The minimum top-up is CA$25.00, paid by credit card.
- Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.
Frequently asked questions
Do I have to rewrite my OpenAI code?
No. vLLM exposes an OpenAI-compatible API: with the OpenAI SDK, you change base_url, api_key and the model name.
Which models are offered in one click?
Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3, with a 8,192-token context. You can install other models yourself over SSH.
How much does a vLLM server cost?
From CA$0.67/h on an RTX A6000 with 48 GB, or $474.50 per month with the monthly term, in Canadian dollars, taxes extra.
Is it less expensive than the OpenAI API?
It depends on your volume. The page gives the break-even formula: above a certain sustained throughput, or a certain monthly volume, the hourly server costs less than the per-token API.
Does my data stay in Canada?
Requests and responses are processed on your server, in Montreal. Account information may be processed outside Quebec, as the privacy policy explains.
Need GPUs? We’ve got you.
Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.
Other guides
- Your AI says you need a GPU? Run your Python script on a GPU in minutes
- CUDA out of memory: causes, fixes, and when to move to a GPU with more VRAM
- torch.cuda.is_available() returns False: causes and fixes (Mac, PC without NVIDIA, Docker)
- How much does a GPU cost per hour in Canada, in Canadian dollars?
- Rent a GPU by the hour in Canada: prices, steps and billing
- GPU cloud in Montreal, Canada: NVIDIA GPU servers hosted in the country
- Rent an H100 in Canada: PCIe, NVLink or SXM, by the hour
- Which GPU to fine-tune an LLM? Memory, machine and cost
- Cheap GPU cloud: how to pay as little as possible for a GPU