GPU VPS in 2026: What 24GB of VRAM Really Costs
Hetzner's new GEX45 puts a dedicated Blackwell GPU at EUR 214/month. We compare fixed-rate servers with RunPod, Vast.ai, Vultr and hyperscalers.
A Blackwell GPU for the price of a phone bill
On September 1, 2026, Hetzner quietly did something the GPU rental market rarely does: it published a fixed price. The new GEX45 dedicated server pairs an NVIDIA RTX PRO 4000 Blackwell SFF Edition, 24 GB of GDDR7 ECC memory and all the fifth-generation Tensor Core FP4 goodness with a 14-core Intel i5-13500, 64 GB of DDR4 and two 512 GB NVMe drives. Monthly rent: 214 euros, plus a one-time setup fee of 209 euros. Helsinki only, for now.
That number deserves attention because almost nothing else in this market holds still long enough to be compared. Hourly GPU marketplaces move week to week. Hyperscaler pricing is a spreadsheet of attached storage, egress and region multipliers. Hetzner just says: here is a whole machine with a whole GPU, 214 euros a month, unlimited traffic at 1 Gbit/s, go build something. Spread the setup fee over a year and you are paying roughly 231 euros a month for a dedicated Blackwell card. That is, as of this writing, the cheapest published flat-rate ticket into current-generation 24 GB VRAM.
Why 24 GB is the number that matters
Not everyone needs an H100. Heresy in some circles, but the numbers back it up.
Twenty-four gigabytes of VRAM is the threshold where most deployed open-source models actually run without compromise. A Llama-class 8B model at full 16-bit precision fits comfortably. A 30B-class model fits with 4-bit quantization, with enough headroom left over for a 16k-token context window before you hit out-of-memory errors. Image generation pipelines, Whisper transcription, embedding models, CAD and Blender rendering: all of them land in the same 24 GB bucket. Double the memory and the card price roughly doubles; halve it and you cannot load the models people actually want to run.
The practical sweet spot in 2026: one 24 GB card, owned or flat-rented, covers inference, fine-tuning and rendering for a solo developer or small team. Everything above it is for multi-tenant serving or serious training runs.
The fixed-price option: Hetzner GEX45 and GEX44
Hetzner now runs two entry GPU servers. The GEX44 (older, still sold) carries an RTX 4000 SFF Ada with 20 GB GDDR6 at around 234 euros a month in Falkenstein. The GEX45 replaces the Ada card with the Blackwell generation: 20 GB becomes 24 GB, GDDR6 becomes GDDR7, CUDA cores jump from 6,144 to 8,960, Tensor cores from 192 to 280 with FP4 support, and power draw stays at 70 W. For less money per month you get meaningfully more card.
- GEX45 - RTX PRO 4000 Blackwell SFF, 24 GB GDDR7 ECC, i5-13500, 64 GB DDR4, 2x 512 GB NVMe, 1 Gbit/s unlimited traffic, Helsinki. 214 EUR/month + 209 EUR setup.
- GEX44 - RTX 4000 SFF Ada, 20 GB GDDR6, i5-13500, 64 GB DDR4, 2x 1.92 TB NVMe. Around 234 EUR/month, Falkenstein.
- GEX131 - for when one card is not enough: RTX PRO 6000 Blackwell Max-Q (96 GB GDDR7) with Xeon Gold and up to 768 GB RAM, 10 Gbit/s uplink option.
The catch list is short but real. You get one location, no hourly billing, a setup fee that stings on short experiments (across a two-week proof of concept it is the larger half of the bill), and Hetzner's GPU stock has a history of selling out. The i5-13500 host CPU is fine for inference but will bottleneck aggressive multi-instance serving. And if your workload only runs eight hours a day, a fixed monthly rate is paying for sixteen hours of idle silicon.
The hourly marketplace route: RunPod and Vast.ai
If your GPU work is bursty, hourly billing beats flat rent. This is where the marketplace clouds come in, and where honest comparison gets hard, because the prices genuinely move.
RunPod splits its inventory into Secure Cloud (RunPod-vetted infrastructure) and Community Cloud (third-party hosts). As of August and September 2026 tracking, an RTX 4090 runs about $0.69/hr on Secure and $0.34/hr on Community. An A100 PCIe 80 GB lands around $1.39/hr. The same trackers caught RunPod's L4 price jumping 25.6 percent in eight days in August, from $0.39 to $0.49 an hour, which tells you how much to trust any single snapshot.
Vast.ai is a pure marketplace: hosts list their own hardware, prices float with supply and demand. The listed floor for an RTX 4090 has touched $0.13-0.14/hr; realistic non-interruptible rates cluster at $0.34-0.50/hr. H100 SXM has been seen around $1.49/hr at the floor, against $3.99 at Lambda and $4-5 at the hyperscalers.
Rule of thumb from the marketplace data: budget 20-40 percent above the lowest listed rate for real planning. Interruptible listings are cheap because you accept the risk of being kicked off when a higher bidder shows up.
So how does hourly stack up against Hetzner's flat rate? An RTX 4090 at $0.34/hr works out to roughly $248 a month if it runs 24/7. Hetzner asks about $270 a month (rent plus amortized setup) for a Blackwell card that is one full generation newer, with ECC memory, a whole dedicated machine around it and unmetered traffic. The moment your utilization crosses roughly half of full-time, the dedicated server wins. Below that, rent by the hour.
The big-cloud escape hatch: Vultr, DigitalOcean and friends
The traditional VPS vendors have not ignored GPUs, but their pricing sits in an awkward middle.
- Vultr - A100 (80 GB) instances from roughly $0.55-0.81/hr, A10 and L40S also available, 32 regions, hourly billing. Convenient if you already run Vultr, rarely the cheapest GPU hour on the market.
- DigitalOcean - H100-backed GPU Droplets for those who want managed Kubernetes around their inference; CPU Droplets still start at $4/month. Pricing lands well above the marketplaces.
- Hyperscalers (AWS, Azure, GCP) - still $4-5/hr for an H100 class instance once attached compute and storage are counted, roughly 80-100 percent over the specialist clouds. You are paying for compliance, support contracts and existing vendor lock-in, not for GPU hours.
There is also a quieter tier worth naming: CPU-only inference. If you are running a quantized 7B-13B model on llama.cpp, or just hosting Open WebUI and n8n wired to external APIs, a 60 GB RAM Contabo box at under $30/month will serve tokens all day. The reflex of reaching for a GPU because the project says "AI" wastes more money than any pricing gap between providers.
A decision framework that survives contact with reality
After laying out the numbers, the decision collapses to three questions.
- How many hours a day does the GPU actually compute? Under 12 hours: hourly marketplace (Vast.ai or RunPod Community). Over 12: fixed-rate dedicated (Hetzner GEX45) or your own hardware.
- Can your job tolerate interruption? Interruptible listings save another 20-40 percent but you can lose the machine mid-run. Training jobs with checkpointing survive fine. Production inference does not.
- What happens to your data? Marketplace hosts are strangers with a GPU. Egress fees and persistence costs on hourly platforms quietly close the gap that made them look cheap.
For a solo developer running local LLM inference around the clock, the GEX45 math is hard to beat: current-generation 24 GB Blackwell, ECC memory, dedicated metal, flat bill. For a team training overnight and reading results in the morning, the marketplaces print the same work for a third of the price. For everyone serving a quantized model to a few hundred users, the answer is still a plain CPU VPS with a lot of RAM, and the savings fund something else entirely.
What to watch for the rest of 2026
Three things are worth tracking. First, whether Hetzner's pricing holds: GPU stock at these rates tends to evaporate, and the GEX44's history suggests the GEX45 will spend stretches of the year out of stock rather than repriced upward. Second, whether marketplace floors stay this low: the same trackers that show $0.13/hr 4090 listings also show weekend spikes past $1.20/hr when demand concentrates. Third, Blackwell trickling down: the RTX PRO 4000 appearing in a 214-euro server means the previous generation (and its clones at GPU Mart and similar dedicated-GPU hosts, where RTX Pro 4000 VPS configurations list around $189/month) will be repriced within the year.
The GPU rental market in 2026 is finally splitting into two sane halves: flat-rate dedicated machines for steady workloads, liquid hourly marketplaces for bursty ones. Pick the half that matches your duty cycle and stop paying for the other.