Table of Contents
- 1. The trap: a closed-model test measures itself
- 2. Which WordPress are you testing?
- 3. Running one that produces a usable number
- 4. Reading the result: where the ceiling usually is
- 5. Which tool for which job
- 6. What providers publish that helps you set a target
- 7. What a load test cannot tell you
- FAQ
- Sources
1. The trap: a closed-model test measures itself
The most common way a WordPress load test produces a reassuring and useless number is the scheduling model, and the tool that documents it most plainly is k6. In its words: "In the closed model, VU iterations start only when the last iteration finishes. In the open model, on the other hand, VUs arrive independently of iteration completion."
Why that matters is on the same page under a name worth learning: "When the duration of the VU iteration is tightly coupled to the start of new VU iterations, the target system's response time can influence the throughput of the test. Slower response times means longer iterations and a lower arrival rate of new iterations, and vice versa for faster response times. In some testing literature, this problem is known as coordinated omission."
k6's own worked example makes the size of the gap concrete. A test configured with constant-vus, one virtual user and a duration of one minute, against a request that takes six seconds, produces this: "The following request will take roughly 6s to complete, resulting in an iteration duration of 6s. As a result, New VU iterations will start at a rate of 1 per 6s, and we can expect to get 10 iterations completed in total for the full 1m test duration." The same script switched to constant-arrival-rate at one iteration per second produces 60 in the same minute, whatever the response time turns out to be.
So a fixed-VU test against a WordPress site that has started to struggle will quietly reduce its own request rate, and the latency graph will look acceptable while the site is falling over. The fix is the open model: k6 implements it in the constant-arrival-rate and ramping-arrival-rate executors, which "decouple the start of new VU iterations from the iteration duration", so "the response times of the target system no longer influence the load on the target system". A closed-model test is still useful for a different question, how a known set of users experiences a degrading site, but it cannot answer how much traffic the site serves.
2. Which WordPress are you testing?
A WordPress site is not one thing to a load test. With a page cache in front of it, most requests never reach PHP or the database; without one, every request does. WordPress's own performance documentation describes the first case and puts a size on it: static files are served to users, "reducing the processing load on the server. This can improve performance several hundred times over for fairly static pages."
The same documentation describes the second lever, the object cache, as "the act of moving data from a place of expensive and slow retrieval to a place of cheap and fast retrieval", typically persistent, and notes that "the storage engine for an object cache can be a number of technologies" including Redis, Memcached, APC and the file system, with the choice dictated by the application. It also states the rule that governs all of it: "cached data should always be replaceable and regenerable."
Practically, that means three tests rather than one, and each one labelled:
- Cached pages, which measure the web server, the network and the cache layer. This is what most visitors see most of the time on a content site.
- Uncached pages, which measure PHP, the theme, the plugins and the database. This is what a cache miss, a logged-in user, a comment submission or a WooCommerce cart actually costs.
- A mixed profile, with a realistic share of misses, logged-in sessions and dynamic endpoints, which is the only one that resembles a real afternoon.
Reporting one of those three as "the site's capacity" is the second most common way a load test misleads. The first is section 1.
3. Running one that produces a usable number
- Run the load from somewhere else. A generator on the same machine competes with the server for CPU and network, and on a small VPS it becomes the bottleneck first. That is the usual explanation for a result that plateaus at a round number.
- Use arrival-rate scheduling (section 1) and hold the rate steady rather than ramping while you watch a dashboard. A ramp is for finding a breakpoint; a hold is for measuring a capacity.
- Warm the cache first, then say which state you tested. A cold start and a warm run are different tests, and the first request after a purge is not representative of anything except the first request after a purge.
- Run long enough to get past the first minute. k6's test-type guidance is to start with a smoke test that validates the script under minimal load before moving to "higher loads and longer durations", and to use more than one kind of test: "no single test can uncover all issues."
- Record server-side counters throughout, not just at the end (section 4).
- Repeat it. Two runs that disagree by more than a little mean the environment moved, and a single run is an anecdote with a chart.
- Write down the versions: WordPress, PHP, the cache plugin, MySQL or MariaDB, the theme and the plugin count. A result without those is not reproducible, including by you in six months.
4. Reading the result: where the ceiling usually is
Client-side latency tells you that something is saturated. The server counters tell you what. The pairings below are the ordinary diagnosis order, and each one has a signature in the metrics the test can record alongside the response times.
| Signature | Usually means | Confirm with |
|---|---|---|
| Latency rises while CPU sits low, and the request rate stops growing | The database, or lock contention behind it | The slow query log, the connection count, and whether the same query repeats across requests |
| CPU is pinned and PHP-FPM shows no idle workers, or a request queue is growing | The PHP worker pool is saturated, or a plugin is doing expensive work per request | The PHP-FPM pool status, which shows the busy, idle and queue counts separately |
| Load average climbs while CPU utilisation looks moderate | I/O wait, usually the disk | The disk queue depth and await times, which the storage page covers across providers |
| Response times are fine in the middle and terrible at the top | A queue that only appears under concurrency | Percentiles rather than the average, and the queue counters at the same timestamps |
| Everything sags at a round number of requests per second | The test client, or a provider-side port or transfer limit | Running the generator from a second machine, and the published port figures on the network page |
Two habits make those readings possible. Record percentiles rather than averages, because a site that answers 99 requests in 200 ms and one request in 30 seconds has an average that describes neither. And decide the pass mark before the test rather than after it: k6 calls these thresholds and treats them as pass/fail gates, which is the difference between a measurement and a story about a graph.
5. Which tool for which job
| Tool | What it is | Reach for it when |
|---|---|---|
k6 | Scripted HTTP load in JavaScript, with arrival-rate executors for the open model and thresholds as pass/fail gates | The default choice when you want the test in version control |
wrk | A single binary that hammers one URL with a fixed number of connections | Raw throughput on a page that needs no login and no form |
ab | Apache Bench, on most systems already | A first sanity check in one command; single-threaded, so it peaks early |
hey | A modern replacement for ab with concurrency control | Same job as ab without the client becoming the bottleneck first |
| Locust | Python tests with a web UI and distributed workers | When the test needs logic a small team can read |
| JMeter | A GUI, recordings and a large plugin ecosystem | When the test must cover a flow a browser walks through |
| k6 browser | A real browser engine inside the load test | When the page depends on JavaScript the HTTP clients never run |
sysstat, mysqladmin, slowlog | Server-side counters read while the test runs | Which is what turns a number into a diagnosis |
The tool matters less than the scheduling model and the counters, and the fastest way to waste a day is to spend it choosing between them.
6. What providers publish that helps you set a target
Very little of it predicts page performance, and knowing which numbers exist stops you from shopping for the wrong one. UpCloud is the clearest case on this site: the UpCloud specs page records over 100,000 read IOPS at 4k block size for its MaxIOPS tier and up to 10,000 for standard storage, both qualified by the block size and by "up to". Linode's types API page publishes a network_out figure per instance type, from 1 to 16 Gbit/s, plus a transfer allowance on every plan. Most providers state a storage type and a port speed and stop; the storage page and the network page collect those labels.
What none of them publish is a requests-per-second figure for a WordPress site, because it does not exist independently of the site: the same 4 GB server serves a cached brochure site and a WooCommerce catalogue at rates that differ by orders of magnitude. That is why this page is a method rather than a table of scores. Buying decisions belong on the WordPress page, which ranks plans by what they include, and the sizing arithmetic on the calculator.
7. What a load test cannot tell you
- Whether the site is correct. A load test hammers whatever URLs you give it; it does not click through a checkout.
- How a different traffic mix behaves. Ten URLs tested hard is not the same profile as a thousand URLs tested lightly, and the cache behaves differently in each.
- What happens when a dependency is slow rather than down. That is a fault-injection exercise, and it belongs in a different test.
- Whether the server is the limit. Everything upstream shares the blame: DNS, the CDN, the visitor's network, and the test client itself.
- Anything about the provider you did not test on. A result on one plan in one datacentre is evidence about that plan in that datacentre.
The last one is the point of the whole cluster. Numbers on a provider's page describe hardware; numbers from your own test describe a workload; and the two are related by nothing you can read off a plan card. The benchmarks hub explains why this site publishes none of its own and gives the commands for the ones you can run in an afternoon on an hourly-billed server.