GPUcloud

GPU guide

OpenAI API too expensive? Host your LLM on a GPU by the hour, in Canada

Updated on

Short answer

Deploy the one-click OpenAI-compatible vLLM template: you get an HTTPS /v1 address and a key, and your existing OpenAI code works by changing only base_url and api_key, on a GPU in Canada from CA$0.67/h. The server is paid by the hour, not by the token: it beats the API once your volume passes the break-even point, computed below with a formula where you plug in your own numbers.

What the vLLM template installs

  • An OpenAI-compatible API (/v1) served over HTTPS, protected by a key generated for you and shown in the console.
  • The model of your choice among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3, served with a 8,192-token context. It needs a GPU with at least 24 GB.
  • Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3 require accepting their license on the model hub, then adding an access token to /etc/gpucloud/app.env on the server.
  • Only the /v1 and /health routes are exposed; you keep full SSH access to the server.

The steps

  1. 1.Deploy the vLLM template

    In the console, pick the OpenAI-compatible vLLM template, the language model and a GPU, then confirm the price. Installation takes a few minutes; the console shows the progress.

  2. 2.Get the address and the key

    When the server answers, the console shows its HTTPS address and the generated API key.

  3. 3.Check that the server answers

    /health answers once the model is loaded; /v1/models lists the served model.

    Check the server
    export VLLM_API_KEY="<key shown in the console>"
    curl https://SERVER_ADDRESS/health
    curl https://SERVER_ADDRESS/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
  4. 4.Change two lines in your code

    Keep the OpenAI SDK: replace base_url and api_key, and use the name of the served model.

    Your existing OpenAI code: only two lines change
    # pip install openai
    from openai import OpenAI
    
    client = OpenAI(
        base_url="https://SERVER_ADDRESS/v1",  # shown in the console
        api_key="VLLM_KEY",  # generated key, shown in the console
    )
    
    response = client.chat.completions.create(
        model="Qwen/Qwen2.5-7B-Instruct",
        messages=[{"role": "user", "content": "Hello!"}],
    )
    print(response.choices[0].message.content)

What the server costs

GPUs recommended for the vLLM template, 1-GPU server
GPUVRAMPer hourOffice hours (176 h)Full month by the hour (730 h)Monthly term
NVIDIA RTX A600048 GB$0.67$117.92$489.10$474.50
NVIDIA L4048 GB$1.34$235.84$978.20$941.70
NVIDIA A100 PCIe80 GB$1.81$318.56$1,321.30$1,270.20
NVIDIA RTX PRO 600096 GB$2.48$436.48$1,810.40$1,737.40
NVIDIA H100 PCIe80 GB$3.35$589.60$2,445.50$2,343.30

Office hours: 8 h a day for 22 days. Prices in Canadian dollars, taxes extra, read from the catalog when the page is built.

The server costs the same whether you send it one request or a million. Outside your usage hours, hibernate it to keep the downloaded model on disk without paying for the GPU.

The break-even: per-token API or GPU by the hour

An API bills every token; your server bills every hour. The math fits in three formulas:

cost per million tokens = hourly price ÷ (tokens per second × 3,600) × 1,000,000
break-even throughput (tokens/s) = hourly price × 1,000,000 ÷ (3,600 × API price per million)
break-even volume (million tokens per month) = monthly GPU cost ÷ API price per million

Example with two hypotheses, which are neither measurements nor a real price list: an API at CA$2.00 per million tokens, and a sustained throughput of 200 tokens per second. On an RTX A6000 at CA$0.67/h, a million tokens then costs CA$0.93. The server comes out ahead as soon as it produces more than 93 tokens per second on average over the hour.

Per month, with the same hypotheses: in office hours (176 h, $117.92), the break-even is 59 million tokens per month; with the monthly term ($474.50), it is 237.3 million tokens per month.

Replace the hypotheses with your numbers: the API price you pay, and the throughput you measure on your server with your model and your real requests.

import time

start = time.perf_counter()
response = client.chat.completions.create(
    model=MODEL,
    messages=[{"role": "user", "content": PROMPT}],
    max_tokens=512,
)
elapsed = time.perf_counter() - start
print(response.usage.completion_tokens / elapsed, "tokens/s")

When to stay on the API

  • Your volume is low or irregular: per token, you do not pay for idle hours.
  • You need a model that is not open weights, or a longer context than the one-click template serves.
  • You do not want to run a server, even one installed for you.

Your requests stay in Canada

The server runs in Montreal: requests, responses and the model stay on your machine, in Canada. See the Data sovereignty page.

Good to know before you launch

  • Billing is hourly, at the server’s hourly rate, from prepaid credit. There is no per-minute billing, and the first hour is charged when you deploy.
  • A stopped server is still billed at the full hourly rate. To pay less, hibernate it ($0.02 per hour (about $14.60 per month), disk kept) or delete it. A Spot machine cannot be hibernated: delete it when you are done.
  • The minimum top-up is CA$25.00, paid by credit card.
  • Servers and their disks stay in Canada (region canada-montreal, in Montreal). Account information (identity, billing, e-mails) may be processed outside Quebec, as the privacy policy explains.

Frequently asked questions

Do I have to rewrite my OpenAI code?

No. vLLM exposes an OpenAI-compatible API: with the OpenAI SDK, you change base_url, api_key and the model name.

Which models are offered in one click?

Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3, with a 8,192-token context. You can install other models yourself over SSH.

How much does a vLLM server cost?

From CA$0.67/h on an RTX A6000 with 48 GB, or $474.50 per month with the monthly term, in Canadian dollars, taxes extra.

Is it less expensive than the OpenAI API?

It depends on your volume. The page gives the break-even formula: above a certain sustained throughput, or a certain monthly volume, the hourly server costs less than the per-token API.

Does my data stay in Canada?

Requests and responses are processed on your server, in Montreal. Account information may be processed outside Quebec, as the privacy policy explains.

Need GPUs? We’ve got you.

Launch an NVIDIA GPU server by the hour in Canada, paid in Canadian dollars by credit card, or reserve a GPU for a given date.

Other guides