One-click ready machine
OpenAI compatible vLLM server
A vLLM server with an OpenAI compatible API, protected by a generated key.
Serve an open language model behind an OpenAI compatible API (/v1). Your existing applications connect to it by simply changing the address and the key.
What you get
- An OpenAI compatible API (/v1) served over HTTPS
- The model of your choice among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3
- An API key generated for you, shown in the console
- Full SSH access to the server
What it is for
- Connect an existing application to an open model hosted in Canada
- Offer a private inference endpoint to your team
- Try an open model in real conditions before production
Included with the template
Ready in a few minutes
Installation starts on its own at the server's first boot. The console shows the progress and the access address as soon as the application answers.
HTTPS included
The application is served over HTTPS with a public certificate obtained automatically, behind authentication. Nothing is exposed unprotected.
Data hosted in Canada
The server, your models and your files are hosted in Canada. Billing is in Canadian dollars.
Recommended GPUs
Price per server in Canadian dollars, taxes extra. Hourly billing with no commitment, the template is included.
| GPU | Memory per GPU | Hourly price | Order |
|---|---|---|---|
| 1 × NVIDIA L40 | 48 GB GDDR6 | $1.34/h | Deploy with this GPU |
| 1 × NVIDIA A100 PCIe | 80 GB HBM2e | $1.81/h | Deploy with this GPU |
| 1 × NVIDIA RTX PRO 6000 | 96 GB GDDR7 | $2.48/h | Deploy with this GPU |
| 1 × NVIDIA H100 PCIe | 80 GB HBM2e | $3.35/h | Deploy with this GPU |
Launch your machine now
The template is already selected in the console. Choose your SSH key, confirm the price and the server starts.
Deploy this templateFrequently asked questions
OpenAI compatible vLLM server: how long before I can use it?
Usually a few minutes after the server starts. The time varies with the GPU and the amount to download; the console shows you the progress.
OpenAI compatible vLLM server: how much does it cost?
The template is included. You only pay for the GPU server, from $1.34/h in Canadian dollars, taxes extra, with no commitment.
OpenAI compatible vLLM server: how do I access it securely?
The console gives you an HTTPS address with a public certificate and the credentials generated for your server. The application never listens directly on the Internet: only an authenticated HTTPS proxy gives access to it.
Where is my data hosted?
In Canada. The server and its disks are hosted in Canada and billing is in Canadian dollars.
Which models are offered?
You choose among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3. Some models require accepting their license on the model hub and adding an access token on the server.