GPUcloud

One-click ready machine

OpenAI compatible vLLM server

A vLLM server with an OpenAI compatible API, protected by a generated key.

Serve an open language model behind an OpenAI compatible API (/v1). Your existing applications connect to it by simply changing the address and the key.

What you get

  • An OpenAI compatible API (/v1) served over HTTPS
  • The model of your choice among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3
  • An API key generated for you, shown in the console
  • Full SSH access to the server

What it is for

  • Connect an existing application to an open model hosted in Canada
  • Offer a private inference endpoint to your team
  • Try an open model in real conditions before production

Included with the template

  • Ready in a few minutes

    Installation starts on its own at the server's first boot. The console shows the progress and the access address as soon as the application answers.

  • HTTPS included

    The application is served over HTTPS with a public certificate obtained automatically, behind authentication. Nothing is exposed unprotected.

  • Data hosted in Canada

    The server, your models and your files are hosted in Canada. Billing is in Canadian dollars.

Recommended GPUs

Price per server in Canadian dollars, taxes extra. Hourly billing with no commitment, the template is included.

See all GPU servers

Launch your machine now

The template is already selected in the console. Choose your SSH key, confirm the price and the server starts.

Deploy this template

Frequently asked questions

OpenAI compatible vLLM server: how long before I can use it?

Usually a few minutes after the server starts. The time varies with the GPU and the amount to download; the console shows you the progress.

OpenAI compatible vLLM server: how much does it cost?

The template is included. You only pay for the GPU server, from $1.34/h in Canadian dollars, taxes extra, with no commitment.

OpenAI compatible vLLM server: how do I access it securely?

The console gives you an HTTPS address with a public certificate and the credentials generated for your server. The application never listens directly on the Internet: only an authenticated HTTPS proxy gives access to it.

Where is my data hosted?

In Canada. The server and its disks are hosted in Canada and billing is in Canadian dollars.

Which models are offered?

You choose among Qwen2.5 7B Instruct, Llama 3.1 8B Instruct and Mistral 7B Instruct v0.3. Some models require accepting their license on the model hub and adding an access token on the server.

Other one-click templates