Hourly GPU Servers

Choose a graphics card and a ready-made app (Ollama, ComfyUI, Jupyter), and pay only for the hours it runs. Automatic delivery with a dedicated HTTPS endpoint and API you can call from your own server or app.

From €0.03 per hour
Hourly billing Instant delivery Dedicated URL & API
💡 How is this price possible?

This service runs on a distributed network of thousands of GPUs instead of an expensive datacenter - which is why it costs a fraction of a traditional GPU server. If a node drops out, your app is automatically brought back up on another node, and you only pay for the hours it runs. A great fit for AI inference, image generation, rendering and batch jobs; if your workload cannot tolerate even a short relocation, our regular VPS is the right choice. Machine disk space is ephemeral - store your results externally.

Choose your graphics card

Prices are hourly and come from the same rate that appears on your invoice.

How many machines?

Each unit is a separate machine with its own card - not several cards in one machine.

1
Total · per hour €0.03
roughly per day —
Continue and create

How it works

1 Pick a card and an app

Choose a GPU and one of the ready-made apps - nothing to install or configure.

2 Pay by the hour

Top up credit; only the hours your machine runs are deducted.

3 Get your dedicated endpoint

Your HTTPS address and private token appear in your panel when ready. Allow about one hour and follow live progress in your panel. A confirmed build failure or no first delivery within two hours automatically cancels the order and returns deducted funds to your account credit. Preparation time does not count toward metered usage.

Ready-made apps

Pick one at checkout; it arrives configured and ready.

Ollama LLM

LLM with an OpenAI-compatible API - you pull the model you want (Qwen, Llama, ...) and call it from your own code. No web UI.

ComfyUI Stable Diffusion

Image generation with Stable Diffusion - ready API with web documentation.

Jupyter PyTorch

New Jupyter orders are temporarily paused while code execution connectivity is verified. Your existing services are unchanged.

Which GPU is right for my workload?

The number that matters most is GPU memory (VRAM): the model has to fit in it. The table below is an approximate guide; quantisation (e.g. Q4) lets larger models fit.

GPU memory Good for Example cards
4-8 GB Stable Diffusion 1.5, Whisper speech-to-text, small language models up to 3B parameters GTX 1650 · GTX 1060 · GTX 1070, 1080, 1080Tifrom €0.03 per hour
10-12 GB SDXL, 7–8B language models (Q4) such as Llama 3.1 8B and Qwen 7B, light LoRA training RTX 3060 · RTX 2080 Ti · RTX 5070 Ti Laptopfrom €0.10 per hour
16 GB Flux in fp8, 13–14B language models (Q4), Blender rendering RTX 3080 Ti Laptop · RTX 4060 Ti · RTX 5060 Tifrom €0.13 per hour
20-24 GB Full Flux, 30–34B language models (Q4) such as Qwen 32B, LoRA training for SDXL and 7B models RTX 3090 · RTX 3090 Ti · RTX 5090 Laptopfrom €0.22 per hour
32 GB Video generation models, 32B language models with long context, heavy batch processing RTX 5090from €0.65 per hour
48-96 GB 70B language models (Q4) on a single card, fine-tuning mid-size models, several models at once RTX PRO 6000 Blackwellfrom €2.57 per hour

AMD cards work well with Ollama and language models, but many image-generation and training tools require CUDA and an NVIDIA card.

What common jobs really cost

No minimum runtime: work one hour, pay for one hour. Figures come from this page's live rates.

Card One hour 8 hours (a working day) Full 24 hours
RTX 3090 (24 GB) €0.22 €1.76 €5.29
RTX 4090 (24 GB) €0.43 €3.42 €10.26
RTX 5090 (32 GB) €0.65 €5.18 €15.54
How is it billed?

Pay-as-you-go, by the hour. You need credit for at least 24 hours to start; only the hours the machine is running are charged, and billing stops the moment you delete it.

How do I use it?

Pick a ready-made app at checkout and get a dedicated HTTPS address. Jupyter opens in your browser; Ollama and ComfyUI have no web UI and are called as an API from your own server or code. These are not VPS machines with SSH and root; they are a ready app on a dedicated GPU.

Model inference

Serve language and image models without buying hardware.

Rendering and video

Generate images with ComfyUI; process video in Jupyter using your own tools and models. A ready-made video generation model is not included.

Training and experiments

Train with checkpoints so an interruption does not cost you the whole run.

What exactly happens when credit runs out?

The most common question from hourly customers. The answer is fixed, with no exceptions:

  1. 1

    Warning before it runs out

    When your spendable credit covers about 4 more hours, you get an SMS and email warning.

  2. 2

    Powered off, not deleted

    When credit runs out the machine stops and its settings and service address are kept for 24 hours. The machine's own disk is temporary and wiped on every stop — always save your output elsewhere.

  3. 3

    Top up and it powers on automatically

    Top up within those 24 hours and the server powers back on automatically within an hour — no ticket, no call.

  4. 4

    After 24 hours: permanent deletion

    Without a top-up the server and all its data are deleted permanently; no backup remains and recovery is impossible.

The cost of this 24-hour hold

The 24-hour hold is free on GPU servers: a stopped machine costs us nothing, so we charge you nothing.

Powering off yourself

GPU servers bill running hours only; preparation time and automatic moves between nodes are free.

Your choice

When ordering, and any time afterwards in your panel, you can choose "switch to monthly" or "delete immediately" instead of powering off.

Hourly GPU server FAQ

What is an hourly GPU server?
A dedicated GPU with a ready application, paid hourly from your account credit. Allow about one hour for first delivery. A confirmed build failure or no first delivery within two hours of payment automatically cancels the order and returns deducted funds to your account credit.
How is it billed and what is the minimum?
Each card has a current hourly rate and only the hours the machine is running are charged. You need credit for 24 hours of that card to start, but there is no minimum runtime: work one hour and delete the machine, and you have paid for one hour while the rest stays in your wallet.
What happens when credit runs out?
About 4 hours before it runs out you get an SMS and email warning. When credit runs out the machine stops and its settings and service address are kept for 24 hours free of charge; top up within that window and it starts again automatically within an hour. Without a top-up it is deleted permanently after 24 hours.
Is preparation time or a stopped machine billed?
No. The clock starts when the machine is running; downloading and preparing the app, automatic moves between nodes and a stopped machine cost nothing.
Do I get SSH and root access?
No. The hourly GPU server is a ready-made app on a dedicated card with its own HTTPS address and API - not a virtual machine; there is no SSH, root or OS change. Our hourly and cloud VPS have root but no GPU. If you need a dedicated GPU server with full root, SSH and persistent disk, open a support ticket for a consultation and we will quote the right option.
What exactly does each ready-made app deliver?
Ollama: an OpenAI-compatible API for language models; you download the model yourself with pull, and there is no web UI. ComfyUI: an image-generation API with one preinstalled Stable Diffusion model; there is no web page, and opening the address in a browser returns 404. Jupyter: JupyterLab in the browser with PyTorch, CUDA and a terminal - the only option where you run your own code and libraries.
Does ComfyUI have a graphical interface? Can I add my own nodes or models?
No. Our ComfyUI ships as an API only, with its preinstalled model; there is no node canvas, and custom nodes or other models cannot be installed. If you need your own workflow or nodes, choose Jupyter and install your tools there.
Can I run tools that are not listed (such as Whisper, Rembg, AnimateDiff or LTX)?
On Jupyter, yes - with pip or conda from the Jupyter terminal. There is no root or apt, so anything that needs a system package has to come from conda. We have not tested every tool in advance; installing and configuring them is up to you. Recent video models usually need a card with 24 GB of memory or more; see the card guide table on this page.
Can I run several apps on one GPU at the same time?
Each machine runs one ready-made app; Ollama and ComfyUI side by side need two separate machines (raising the quantity also creates independent machines, not several apps on one card). On Jupyter you can install several libraries together, but they all share that one card memory, so load heavy models one at a time.
Can I build a custom workflow, or only use the predefined apps?
Custom workflows are possible only on Jupyter. Because the machine can be relocated and its disk is temporary, keep your installs in a script or notebook so one run restores them after a move, and store outputs outside the machine.
Can I switch the app after purchase?
Not on the same machine; its app and OS cannot be swapped. For a different app, delete the current machine (billing stops immediately) and create a new one with the app you want.
Which card is enough for my model?
Rule of thumb: a quantised (Q4) language model needs roughly half its parameter count (in billions) in GB of memory plus a little for context — an 8B model needs about 6 GB and a 32B model about 20 GB. SDXL wants at least 10–12 GB and full Flux 24 GB. See the guide table on this page.
Are my files and outputs kept?
The machine's disk is temporary and wiped on every stop or move. Transfer your output (images, model files, results) to your own storage or server as it is produced.
What does "interruptible" mean, and what is it not suited for?
The machine runs on a distributed network of graphics cards and may occasionally move to another node — that is what cuts the price several times over. It is excellent for inference, image generation, rendering and batch work; it is not suited to a service that cannot tolerate even a few minutes of interruption.
Is there an OpenAI-compatible API?
Yes. The Ollama app exposes an OpenAI-compatible API; just point the base URL in your code or tool at your service address and send the gateway token.
How do I pay?
You top up your ServerNet wallet; the same wallet pays for servers, hosting and domains too.
How many machines can I run at once?
Up to 8 machines per order. Each machine is independent with its own card (not several cards in one machine), which suits parallel batch work or serving more users.
Is renting by the hour better than buying a GPU?
If you do not need the card full-time every day, hourly rental is cheaper: no upfront cost, power, cooling or depreciation, and you can switch to a stronger card at any moment. For constant, continuous use, talk to us about a dedicated GPU server.