Host Talk / PHP-FPM, and why your pages wait in line

PHP-FPM, and why your pages wait in line

HOST TALK

8 min read · 1,812 words

Most WordPress sites run on PHP, and PHP does not run inside the web server any more. Modern setups use PHP-FPM, a separate service that keeps a small pool of PHP workers ready. The web server passes each dynamic request to the pool, a free worker picks it up, runs the code, and returns the page.

The important number is the size of that pool. Suppose your account has ten workers and each page takes half a second to build. Ten workers can finish twenty pages per second. If thirty visitors arrive at once, the extra requests wait in a queue. The visitors at the back of the queue see a slow site, and if they wait long enough, an error.

This is the single most common explanation for a site that is fine on a quiet Tuesday and falls over when a newsletter goes out. It is also one of the least visible, because nothing in the site itself looks wrong. The code is the same, the database is the same, and yet pages that took half a second now take twenty. The delay is not in the work. It is in the waiting.

How the pool actually works

The web server, whether Apache or Nginx, is very good at serving static things: images, stylesheets, scripts, anything that is just a file on disk. It can handle thousands of those at once with little effort. A .php file is different, because something has to run it. In the older arrangement, the web server loaded PHP into itself. In the current one, it hands the request over a local socket to a separate PHP-FPM service (FPM stands for FastCGI Process Manager).

That service starts a number of worker processes. Each worker handles exactly one request at a time. When a request arrives, a free worker takes it. When none is free, the request waits in a small queue managed by the socket, and when the queue is full or the wait is too long, the web server gives up and returns an error, usually 502 or 504, or a 503 where the host enforces its own limits.

Visitors 30 at once Web server Nginx / Apache Static files Queue waiting PHP Worker pool worker 1: busy worker 2: busy worker 3: busy worker 4: busy ... up to 10 one request each
Static files bypass the pool. Every PHP request needs a free worker, and the rest wait.

This is why caching has such a dramatic effect. A page served from a full-page cache is just a file, and it never enters the pool.

The arithmetic of a queue

The capacity of the pool is the number of workers divided by the time each request takes. Ten workers and 0.5 seconds per page gives twenty pages per second. Halve the page time and you double the capacity without adding a single worker. This is the lever people overlook, because it is cheaper to make the code faster than to buy a bigger plan.

WorkersSeconds per pagePages per secondWhat 30 simultaneous visitors see
100.520Ten served, twenty queued, last one waits about a second
102.05Ten served, twenty queued, last one waits around four seconds
100.1 (cached)100All thirty cleared in a third of a second
202.010Doubling the pool only halves the queue

These are illustrative figures, not measurements, but the shape is real. The last row is the trap. Buying twice the workers when each page takes two seconds improves matters less than fixing the two seconds.

A subtler effect is that slow requests hold a worker for their whole duration. If a few visitors trigger an expensive report that takes thirty seconds, they can occupy several workers for that time while everyone else queues behind them. One slow page can hurt a hundred fast ones.

What makes workers slow

The usual suspects: database queries that take too long, plugins that call out to other servers and wait for replies, heavy image processing, and code that loads far more than the page needs. A worker held up by a slow external call is a worker that cannot serve anyone else.

A few specific patterns I see in support:

Reading the symptoms

A site that is fast most of the time and falls over during short bursts usually has a worker shortage. A site that is slow all the time usually has a slow page, not too few workers. The difference tells you whether to look at traffic or at code.

Spikes during bursts Slow all the time worker shortage (illustrative) slow code or database (illustrative) response time response time
Where the slowness appears in time says more than how bad it is.

If your host's panel or monitoring shows a counter for "entry processes" or "concurrent connections", watch it during one of the slow periods. A line pinned at the limit, with 503 errors in the access log at the same moment, settles the question.

What helps

Caching is the biggest lever, because a cached page never reaches PHP at all. After that comes keeping PHP current, since newer versions are faster, and removing plugins you do not use. If your host shows a counter for "entry processes" or "concurrent connections", that is this pool, and hitting the limit is the reason for many mysterious 503 errors on busy days.

In rough order of effort against reward:

  1. Turn on page caching for logged-out visitors. On WordPress that means a caching plugin or the host's own cache. It typically turns a half-second page into a few milliseconds.
  2. Move to a current PHP 8.x release. Each major version since 7.0 has brought measurable speed gains, and you also get security fixes that older versions no longer receive.
  3. Remove or replace plugins that call external services on every request, and check that the ones you keep are maintained.
  4. Add an object cache if your host offers one, so repeated database lookups are answered from memory.
  5. Put a CDN in front for static assets and, where it suits the site, for whole pages.
  6. Only then consider a bigger plan with more workers.

A worked example: the newsletter that took the shop down

A small online shop sells garden tools on a plan with ten PHP workers. On an ordinary day it sees perhaps three visitors a minute, and every page is served from cache in a few milliseconds. At ten o'clock on a Thursday it sends a newsletter to eight thousand people, with a link to a sale page.

In the first two minutes, about six hundred people click. The sale page itself is cached and copes. But a good third of those visitors add something to the basket, and the basket, checkout and account pages cannot be cached because they differ for each person. Each of those requests needs a worker for around 0.8 seconds, because the shop runs several plugins that each add a database query or two.

Two hundred uncacheable requests over a couple of minutes is not a huge number: under two a second. But the clicks do not arrive evenly. Most land in the first thirty seconds, so the peak is nearer ten per second, and ten workers at 0.8 seconds each can clear only twelve per second at best. The queue starts to build, each page takes longer, which holds workers for longer, which lengthens the queue. By the time the owner opens the site to check, the checkout is returning 503 to roughly one visitor in five.

Nothing was broken. The fix was in three parts. Next time the owner spaced the newsletter into batches over an hour. A plugin that checked a licence server on every request was replaced. And the sale landing page now links straight to cached product pages rather than the basket. Peak demand on the pool fell by roughly two thirds, with no change of plan.

Shared, VPS and dedicated: who sets the pool size

On shared hosting the pool size is part of your plan and you cannot change it. The limit is deliberately low, because the server is divided among many accounts. On a VPS or dedicated server, you or your host configure it, usually in a file such as /etc/php/8.3/fpm/pool.d/www.conf. The key setting is pm.max_children, the maximum number of workers. Set it from memory, not optimism: if each worker uses about 80 MB and you can spare 4 GB, then around 50 is the ceiling. Set it higher and the server begins to swap, which is slower than queueing.

pm = dynamic
pm.max_children = 50
pm.start_servers = 8
pm.min_spare_servers = 4
pm.max_spare_servers = 12
pm.max_requests = 500
Raising pm.max_children without checking memory is a classic way to turn a slow site into a crashed one. Look at the real memory per worker first, using ps --no-headers -o rss -C php-fpm8.3 and averaging the figures.

Commands worth running

Our troubleshooting page covers the wider set of causes for 5xx errors, and the status code reference explains how 502, 503 and 504 differ.

What else people want to know

Will a CDN fix this?

Partly. It removes static files and cached pages from your server's workload. It cannot help with requests that must run PHP, such as logins, carts and searches.

Is more RAM the same as more workers?

Not directly. RAM lets you configure more workers, but the number is still a setting, and your plan may cap it regardless.

Why is the admin area slow when the front end is fine?

Logged-in pages skip the page cache, so each one reaches PHP. The admin is the real test of how quickly your site builds pages without help.

Does PHP-FPM apply to non-WordPress sites?

Yes. Any PHP application on a modern host goes through the same pool, and the same arithmetic applies.

PreviousWhat a web server actually doesNextThe database behind your website

More from Host Talk

Host Talk

When a CDN helps, and when it gets in the way

A content delivery network sits between your visitors and your server. It keeps copies of your files in many...

Host Talk

How automatic certificate renewal works

Automatic renewal sounds like magic, but the mechanism is simple: a program proves to a certificate authority...

Host Talk

Why your site slows down at 3 p.m.

A common support ticket reads something like: the site is fast in the morning and crawls in the afternoon...