Table of Contents
- How this page selects, and what it does not claim
- Memory first: the sizing rule
- Vultr GPU plans in US regions
- Linode GPU plans
- The large GPUs: listed, but where
- CPU and RAM: the plans that run small models without a GPU
- Hourly first, term never (for this workload)
- What this page could not verify
- FAQ
- Sources
How this page selects, and what it does not claim
A plan is listed if a provider's public API or documentation shows it with a GPU and a US region, or, for the CPU section, if it is a verified high-memory plan from this site's reviews. Prices are the API's monthly and hourly figures; where an API gives only an hourly price (Linode's GPU plans), a 730-hour figure is shown as an estimate and labelled so. The page does not rank by performance and quotes no tokens per second, training times or benchmark scores, because it has none it could source; the earlier version's figures were removed for that reason. GPU prices and availability change more often than ordinary plans, so the date on each source matters more here than elsewhere on this site.
Memory first: the sizing rule
Whatever the hardware, the model's weights have to be in memory to run: in VRAM on a GPU, in RAM on a CPU. The memory a model needs is roughly its parameter count multiplied by the bytes per weight at the precision you load (16-bit weights take two bytes each; the common 4-bit quantizations take about half a byte plus overhead), plus room for the context and the runtime. That arithmetic, not a provider's marketing, decides the plan: a 2 GB GPU share holds a very small model; an 8 GB share or a 16 GB card holds small quantized language models; larger models need larger cards or several. On CPU the same rule applies to RAM, with the difference that RAM is cheap on the fixed-price hosts and the trade is speed. This page states the rule and leaves the arithmetic to the reader's model, because it has measured nothing.
Vultr GPU plans in US regions
Vultr's plans API lists a vcg (Cloud GPU) family of NVIDIA A16 and A40 shares. Prices are stable; the region column is not: it is what the API returned at 14:42 UTC on 2026-09-21, and two reads a few hours apart on the same day differed (Atlanta absent from the 2 GB plan in one, present in the other; the $688 plan absent in one, present in the other). Treat the regions as a snapshot and the plan ids and prices as the durable part:
| Plan | GPU | vCPU / RAM / disk | Monthly | Hourly | US regions listed at 14:42 UTC on 2026-09-21 |
|---|---|---|---|---|---|
| vcg-a16-2c-8g-2vram | NVIDIA A16, 2 GB VRAM | 2 / 8 GB / 50 GB | $43 | $0.059 | ewr, ord, atl, sjc |
| vcg-a40-1c-5g-2vram | NVIDIA A40, 2 GB VRAM | 1 / 5 GB / 90 GB | $55 | $0.075 | ewr |
| vcg-a16-2c-16g-4vram | NVIDIA A16, 4 GB VRAM | 2 / 16 GB / 80 GB | $86 | $0.118 | ewr, ord, sjc |
| vcg-a16-3c-32g-8vram | NVIDIA A16, 8 GB VRAM | 3 / 32 GB / 170 GB | $172 | $0.236 | ewr, ord, sjc |
| vcg-a16-12c-128g-32vram | NVIDIA A16, 32 GB VRAM | 12 / 128 GB / 700 GB | $688 | $0.942 | sjc |
The A16 shares are the entry point: $43 buys a 2 GB slice with a small VM and $172 an 8 GB slice with 32 GB of RAM; the whole 32 GB of an A16 is $688. Larger A16 and A40 allocations (up to 128 GB of VRAM at $2,750) were in the API without a location at the time of reading. Before relying on a region, query https://api.vultr.com/v2/plans?type=vcg or open the order form; GPU stock at Vultr is allocated per region and the list on this page will be stale within hours. Billing is hourly like the rest of Vultr, so a plan can be run for an afternoon. The API also lists preemptible pricing fields for some plans; they are not reproduced here.
Linode GPU plans
Linode's types API lists GPU plans with an hourly price and no monthly figure (the monthly field is null), and its regions API marks six US regions with GPU capability: Newark, Atlanta, Miami, Chicago, Los Angeles and Seattle. The monthly column below is 730 hours at the hourly rate, an estimate this page computed rather than a Linode price:
| Plan | GPU | vCPU / RAM | Hourly (API) | 730 hours (estimate) |
|---|---|---|---|---|
| g2-gpu-rtx4000a1-s | RTX 4000 Ada x1 | 4 / 16 GB | $0.52 | about $380 |
| g2-gpu-rtx4000a1-m | RTX 4000 Ada x1 | 8 / 32 GB | $0.67 | about $489 |
| g2-gpu-rtx4000a1-l | RTX 4000 Ada x1 | 16 / 64 GB | $0.96 | about $701 |
| g2-gpu-rtx4000a2-s | RTX 4000 Ada x2 | 8 / 32 GB | $1.05 | about $767 |
| g1-gpu-rtx6000-1 | RTX 6000 x1 | 8 / 32 GB | $1.50 | about $1,095 |
| g2-gpu-rtx4000a4-m | RTX 4000 Ada x4 | 48 / 196 GB | $3.57 | about $2,606 |
The RTX 4000 Ada plans are whole cards rather than shares, from a single card with 4 vCPUs and 16 GB of RAM at $0.52 an hour to four cards with 48 vCPUs and 196 GB at $3.57. The older RTX 6000 line starts at $1.50 an hour. VRAM per card is not a field in the types API and is not stated here; NVIDIA's specifications for each card are the reference. Which of the six regions has stock for a given plan is shown at deploy time.
The large GPUs: listed, but where
Vultr's plans API also lists L40S shapes (vcg family) and A100, H100, MI325X, MI355X and B200 shapes (vdm family), from an L40S with 48 GB VRAM at $1,122.91 a month ($1.671 an hour) through an A100 at $1,750, an eight-A100 configuration at $14,000, an H100 configuration at $16,074.24 and a B200 configuration at $45,696, all with an empty location list on 2026-09-21, so they are quoted as listed prices without a place to deploy them; Vultr's site did not load for this site's tools, so its GPU product page is not cited. DigitalOcean's regional availability page lists GPU Droplets in its Paperspace datacenters, NY2 (near New York City), CA1 (near Santa Clara) and AMS1, separate from the regions that sell standard Droplets, with H100 and MI300X models named on its pricing page and their prices on a GPU page this version did not read. AWS, Azure and Google Cloud sell the same classes of card by the hour in their calculators. For H100-class work the buying unit is hours, and the Google Cloud and Azure pages describe how those bills are built.
CPU and RAM: the plans that run small models without a GPU
A GPU is not required to run a small model; memory is. The fixed-price hosts sell it cheaply, from the ladders verified on this site:
| Plan | vCPU / RAM / disk | Monthly |
|---|---|---|
| Hostinger KVM 8 | 8 vCPU / 32 GB / 400 GB NVMe | $25.99 (24-month; renews $49.99) |
| Contabo Cloud VPS 12 | 12 vCPU / 48 GB / 400 GB | $30 (introductory) |
| InterServer 16 slices | 8 cores / 32 GB / 640 GB | $48 |
| OVHcloud VPS-4 | 8 vCore / 24 GB / 200 GB | $27.50 |
| Vultr vc2-8c-32gb | 8 vCPU / 32 GB / 640 GB | $160, hourly |
| Linode Dedicated 32GB | 16 vCPU dedicated / 32 GB / 640 GB | $288, hourly |
On any of these a quantized model of a few billion parameters loads and answers one user at a time; how fast depends on the host's CPU generation, memory bandwidth and how much of each vCPU is actually yours, none of which the plan pages state and none of which this page has measured. The dedicated-CPU plans (Linode Dedicated, Vultr's dedicated lines, DigitalOcean's CPU-Optimized) remove the sharing variable at a higher price. The 32 GB page and dedicated CPU page rank those. Embeddings for a modest corpus, classical machine learning, and scheduled batch inference are the workloads that fit here; interactive chat for several users does not.
Hourly first, term never (for this workload)
Every GPU plan on this page bills by the hour, and that is the right unit for machine learning: a model is tried, a fine-tune runs for a weekend, a demo serves traffic for a launch, and the server is destroyed. Create the plan, load the model, measure the thing that matters to you (latency for one request, throughput for many, epochs per hour), and destroy it; the cost of finding out is a few dollars. Do the same on a CPU plan before buying a term at Contabo or Hostinger. Keep the weights and the data on object storage or a cheap storage server (the storage page lists them) so that the expensive machine is only rented while it is computing.
What this page could not verify
- VRAM per Linode GPU plan: not a field in the types API.
- DigitalOcean GPU Droplet prices: on a GPU pricing page not read for this version.
- A region for Vultr's larger A16 and A40 allocations and its L40S, A100, H100, MI-series and B200 shapes: the API listed none at 14:42 UTC on 2026-09-21; and the regions shown for the small plans are a snapshot that changed within the same day.
- Vultr's GPU product page: its site did not load for this site's tools.
- Any performance figure, tokens per second or training time on any plan: none measured, none quoted.