VPS for Machine Learning (2026): GPU Plans With a US Region, Hourly Prices From the APIs, and When CPU and RAM Are Enough

Machine learning on a rented server is a question of memory first and a GPU second. This page lists the GPU plans that the providers on this site actually sell in a US region, with prices from their public APIs as of September 2026, the large GPUs they list without a US region, and the high-memory CPU plans that run small models without a GPU. It contains no inference benchmarks; an earlier version described a "$14 VPS experiment" with token rates and quantization results that could not be sourced, and has been replaced.

✓ GPU plans and prices from Vultr and Linode APIs
✓ No tokens-per-second or benchmark figures
✓ Updated September 2026

Short answer

  • Cheapest GPU in a US region: Vultr A16 shares from $43 a month (2 GB VRAM) to $172 (8 GB VRAM), hourly, in the US regions its API lists at the moment (New Jersey, Chicago, Atlanta, Silicon Valley at 14:42 UTC on 2026-09-21); an A40 share at $55.
  • A whole card by the hour: Linode RTX 4000 Ada from $0.52 an hour (16 GB VRAM), RTX 6000 from $1.50, in six US regions.
  • H100, L40S, B200: listed by Vultr without a US region; DigitalOcean's GPU Droplets sit in its Paperspace datacenters (NY2, CA1); the hyperscalers price them by the hour.
  • No GPU, lots of RAM: Contabo 48 GB for $30, InterServer 32 GB for $48, for small quantized models and batch work.
  • Rule: the model has to fit in memory (VRAM on a GPU, RAM on a CPU); size the plan to the model, then test on an hourly server before any term.

How this page selects, and what it does not claim

A plan is listed if a provider's public API or documentation shows it with a GPU and a US region, or, for the CPU section, if it is a verified high-memory plan from this site's reviews. Prices are the API's monthly and hourly figures; where an API gives only an hourly price (Linode's GPU plans), a 730-hour figure is shown as an estimate and labelled so. The page does not rank by performance and quotes no tokens per second, training times or benchmark scores, because it has none it could source; the earlier version's figures were removed for that reason. GPU prices and availability change more often than ordinary plans, so the date on each source matters more here than elsewhere on this site.

Memory first: the sizing rule

Whatever the hardware, the model's weights have to be in memory to run: in VRAM on a GPU, in RAM on a CPU. The memory a model needs is roughly its parameter count multiplied by the bytes per weight at the precision you load (16-bit weights take two bytes each; the common 4-bit quantizations take about half a byte plus overhead), plus room for the context and the runtime. That arithmetic, not a provider's marketing, decides the plan: a 2 GB GPU share holds a very small model; an 8 GB share or a 16 GB card holds small quantized language models; larger models need larger cards or several. On CPU the same rule applies to RAM, with the difference that RAM is cheap on the fixed-price hosts and the trade is speed. This page states the rule and leaves the arithmetic to the reader's model, because it has measured nothing.

Vultr GPU plans in US regions

Vultr's plans API lists a vcg (Cloud GPU) family of NVIDIA A16 and A40 shares. Prices are stable; the region column is not: it is what the API returned at 14:42 UTC on 2026-09-21, and two reads a few hours apart on the same day differed (Atlanta absent from the 2 GB plan in one, present in the other; the $688 plan absent in one, present in the other). Treat the regions as a snapshot and the plan ids and prices as the durable part:

PlanGPUvCPU / RAM / diskMonthlyHourlyUS regions listed at 14:42 UTC on 2026-09-21
vcg-a16-2c-8g-2vramNVIDIA A16, 2 GB VRAM2 / 8 GB / 50 GB$43$0.059ewr, ord, atl, sjc
vcg-a40-1c-5g-2vramNVIDIA A40, 2 GB VRAM1 / 5 GB / 90 GB$55$0.075ewr
vcg-a16-2c-16g-4vramNVIDIA A16, 4 GB VRAM2 / 16 GB / 80 GB$86$0.118ewr, ord, sjc
vcg-a16-3c-32g-8vramNVIDIA A16, 8 GB VRAM3 / 32 GB / 170 GB$172$0.236ewr, ord, sjc
vcg-a16-12c-128g-32vramNVIDIA A16, 32 GB VRAM12 / 128 GB / 700 GB$688$0.942sjc

The A16 shares are the entry point: $43 buys a 2 GB slice with a small VM and $172 an 8 GB slice with 32 GB of RAM; the whole 32 GB of an A16 is $688. Larger A16 and A40 allocations (up to 128 GB of VRAM at $2,750) were in the API without a location at the time of reading. Before relying on a region, query https://api.vultr.com/v2/plans?type=vcg or open the order form; GPU stock at Vultr is allocated per region and the list on this page will be stale within hours. Billing is hourly like the rest of Vultr, so a plan can be run for an afternoon. The API also lists preemptible pricing fields for some plans; they are not reproduced here.

Linode GPU plans

Linode's types API lists GPU plans with an hourly price and no monthly figure (the monthly field is null), and its regions API marks six US regions with GPU capability: Newark, Atlanta, Miami, Chicago, Los Angeles and Seattle. The monthly column below is 730 hours at the hourly rate, an estimate this page computed rather than a Linode price:

PlanGPUvCPU / RAMHourly (API)730 hours (estimate)
g2-gpu-rtx4000a1-sRTX 4000 Ada x14 / 16 GB$0.52about $380
g2-gpu-rtx4000a1-mRTX 4000 Ada x18 / 32 GB$0.67about $489
g2-gpu-rtx4000a1-lRTX 4000 Ada x116 / 64 GB$0.96about $701
g2-gpu-rtx4000a2-sRTX 4000 Ada x28 / 32 GB$1.05about $767
g1-gpu-rtx6000-1RTX 6000 x18 / 32 GB$1.50about $1,095
g2-gpu-rtx4000a4-mRTX 4000 Ada x448 / 196 GB$3.57about $2,606

The RTX 4000 Ada plans are whole cards rather than shares, from a single card with 4 vCPUs and 16 GB of RAM at $0.52 an hour to four cards with 48 vCPUs and 196 GB at $3.57. The older RTX 6000 line starts at $1.50 an hour. VRAM per card is not a field in the types API and is not stated here; NVIDIA's specifications for each card are the reference. Which of the six regions has stock for a given plan is shown at deploy time.

The large GPUs: listed, but where

Vultr's plans API also lists L40S shapes (vcg family) and A100, H100, MI325X, MI355X and B200 shapes (vdm family), from an L40S with 48 GB VRAM at $1,122.91 a month ($1.671 an hour) through an A100 at $1,750, an eight-A100 configuration at $14,000, an H100 configuration at $16,074.24 and a B200 configuration at $45,696, all with an empty location list on 2026-09-21, so they are quoted as listed prices without a place to deploy them; Vultr's site did not load for this site's tools, so its GPU product page is not cited. DigitalOcean's regional availability page lists GPU Droplets in its Paperspace datacenters, NY2 (near New York City), CA1 (near Santa Clara) and AMS1, separate from the regions that sell standard Droplets, with H100 and MI300X models named on its pricing page and their prices on a GPU page this version did not read. AWS, Azure and Google Cloud sell the same classes of card by the hour in their calculators. For H100-class work the buying unit is hours, and the Google Cloud and Azure pages describe how those bills are built.

CPU and RAM: the plans that run small models without a GPU

A GPU is not required to run a small model; memory is. The fixed-price hosts sell it cheaply, from the ladders verified on this site:

PlanvCPU / RAM / diskMonthly
Hostinger KVM 88 vCPU / 32 GB / 400 GB NVMe$25.99 (24-month; renews $49.99)
Contabo Cloud VPS 1212 vCPU / 48 GB / 400 GB$30 (introductory)
InterServer 16 slices8 cores / 32 GB / 640 GB$48
OVHcloud VPS-48 vCore / 24 GB / 200 GB$27.50
Vultr vc2-8c-32gb8 vCPU / 32 GB / 640 GB$160, hourly
Linode Dedicated 32GB16 vCPU dedicated / 32 GB / 640 GB$288, hourly

On any of these a quantized model of a few billion parameters loads and answers one user at a time; how fast depends on the host's CPU generation, memory bandwidth and how much of each vCPU is actually yours, none of which the plan pages state and none of which this page has measured. The dedicated-CPU plans (Linode Dedicated, Vultr's dedicated lines, DigitalOcean's CPU-Optimized) remove the sharing variable at a higher price. The 32 GB page and dedicated CPU page rank those. Embeddings for a modest corpus, classical machine learning, and scheduled batch inference are the workloads that fit here; interactive chat for several users does not.

Hourly first, term never (for this workload)

Every GPU plan on this page bills by the hour, and that is the right unit for machine learning: a model is tried, a fine-tune runs for a weekend, a demo serves traffic for a launch, and the server is destroyed. Create the plan, load the model, measure the thing that matters to you (latency for one request, throughput for many, epochs per hour), and destroy it; the cost of finding out is a few dollars. Do the same on a CPU plan before buying a term at Contabo or Hostinger. Keep the weights and the data on object storage or a cheap storage server (the storage page lists them) so that the expensive machine is only rented while it is computing.

What this page could not verify

  • VRAM per Linode GPU plan: not a field in the types API.
  • DigitalOcean GPU Droplet prices: on a GPU pricing page not read for this version.
  • A region for Vultr's larger A16 and A40 allocations and its L40S, A100, H100, MI-series and B200 shapes: the API listed none at 14:42 UTC on 2026-09-21; and the regions shown for the small plans are a snapshot that changed within the same day.
  • Vultr's GPU product page: its site did not load for this site's tools.
  • Any performance figure, tokens per second or training time on any plan: none measured, none quoted.

Frequently Asked Questions

Which US VPS providers sell GPU plans?

From their public APIs and documentation in September 2026: Vultr's plans API lists a vcg (Cloud GPU) family of NVIDIA A16 and A40 shares: 2 GB of VRAM with 2 vCPUs and 8 GB for $43 a month, 2 GB of an A40 with 1 vCPU and 5 GB for $55, 4 GB VRAM for $86, 8 GB VRAM for $172, and 32 GB VRAM for $688. Which US regions carry each plan changes from hour to hour; at 14:42 UTC on 2026-09-21 the 2 GB plan listed New Jersey, Chicago, Atlanta and Silicon Valley, the $86 and $172 plans New Jersey, Chicago and Silicon Valley, the A40 New Jersey only and the $688 plan Silicon Valley only, and the API is the only current answer. Larger A40 and L40S shapes were listed without a location. Linode lists RTX 4000 Ada plans from $0.52 an hour (one GPU, 4 vCPU, 16 GB) to $3.57 (four GPUs, 48 vCPU, 196 GB) and RTX 6000 plans from $1.50 an hour, with GPU capability in six US regions (Newark, Atlanta, Miami, Chicago, Los Angeles, Seattle). DigitalOcean's regional availability page lists GPU Droplets in NY2, CA1 and AMS1, its Paperspace datacenters, distinct from its standard Droplet regions. The hyperscalers sell GPUs by the hour in many US regions and are priced in their calculators.

Do I need a GPU for machine learning on a VPS?

Not for everything. Inference with a small quantized language model, classical machine learning (scikit-learn, XGBoost), embeddings for a small corpus, and most data processing run on CPU with enough memory, slowly but usably; a model's weights must fit in RAM, and the memory it takes is roughly its parameter count times the bytes per weight at the quantization you load, plus working room. Training anything beyond toy size, fine-tuning, image generation and serving a model to many users at once want a GPU, where the same rule applies to VRAM. This page does not report tokens per second for any plan; it has not measured any.

What is the cheapest GPU VPS in a US region?

Vultr's vcg-a16-2c-8g-2vram at $43 a month ($0.059 an hour): a 2 GB slice of an NVIDIA A16, 2 vCPUs, 8 GB of RAM and 50 GB of disk, with New Jersey, Chicago, Atlanta and Silicon Valley listed at 14:42 UTC on 2026-09-21; the region list moves, so check the API or the order form for the current one. 2 GB of VRAM holds only the smallest models; the $172 plan (8 GB VRAM, 3 vCPU, 32 GB RAM) is the first that fits a small quantized language model comfortably. Linode's smallest GPU plan is $0.52 an hour, about $380 at 730 hours, for a full RTX 4000 Ada with 16 GB of VRAM, so Vultr is cheaper for a slice and Linode for a whole card.

Where are the big GPUs (H100, L40S, B200)?

Listed, but not in a US VPS region on the pages and APIs this site read. Vultr's plans API carries L40S shapes in its vcg family and A100, H100, MI300-series and B200 shapes in its vdm family (from $1,122.91 a month for an L40S with 48 GB VRAM to $45,696 for a B200 configuration) with no location listed at all on 2026-09-21, so they are quoted here as listed prices without a region. DigitalOcean sells H100 and MI300X GPU Droplets in its Paperspace datacenters (NY2 and CA1 in the US) on a GPU pricing page this version did not read. AWS, Azure and Google Cloud sell them by the hour in their calculators. For a large model, the practical question is usually hours of rental rather than a monthly VPS.

What about running models on CPU with lots of RAM?

The fixed-price hosts sell memory cheaply: Contabo Cloud VPS 12 (12 vCPU, 48 GB) at $30 introductory, InterServer 16 slices (8 cores, 32 GB) at $48, Hostinger KVM 8 (8 vCPU, 32 GB) at $25.99 on a 24-month term, and the 32 GB page ranks the rest. On any of them a quantized 7 or 8 billion parameter model fits in memory and answers, slowly, for one user at a time; larger models need proportionally more RAM and get slower. CPU inference speed depends on the CPU generation, memory bandwidth and how many vCPUs are actually yours, none of which the plan pages quantify, so try it on an hourly provider before committing to a term.

Which should I pick?

A slice of an A16 at Vultr ($43 to $172) to learn, to serve a small model, or to run inference that a CPU makes too slow, in a US region with hourly billing. A whole RTX 4000 Ada at Linode ($0.52 an hour) for fine-tuning or a model that needs 16 GB of VRAM, billed by the hour so a weekend of training costs a weekend. A high-memory CPU plan at Contabo or InterServer for batch work, embeddings and small models where time is not the constraint. A hyperscaler or DigitalOcean's Paperspace GPUs for H100-class work by the hour. Nothing on this page is ranked by speed, because nothing was measured.

Sources

GPU plans and prices are from the Vultr plans API as read on 2026-09-21 and the Linode types and regions APIs as read in September 2026; DigitalOcean GPU locations from its regional availability documentation; CPU plans from the sources cited on each review. This page contains no measurements.

Disclosure: BestUSAVPS has no affiliate relationship with Vultr, Linode, DigitalOcean or the GPU vendors; their links are plain links. The InterServer link leads to its review, which contains affiliate links. This page ranks nothing by performance and contains no affiliate links.