DevOps

504 gateway timeout: how to fix it when you run the server

504 gateway timeout: how to fix it when you run the server

A 504 gateway timeout means the server that answered the browser was a middleman, and the server behind it didn’t reply in time. Nginx waited on PHP-FPM, Apache waited on an app server, or Cloudflare waited on your origin, and a timer ran out. So the fix is rarely “raise the timeout.” It’s finding which hop gave up and why the thing behind it was slow.

This guide is for the person with root, not the WordPress user waiting on a host.

What a 504 actually means

RFC 9110 defines it as a server “acting as a gateway or proxy” that “did not receive a timely response from an upstream server it needed to access in order to complete the request.”

Gateway means something sits in front: nginx, Apache with mod_proxy, a load balancer, a CDN. Timely means a timer fired. A 502 is the cousin, where the upstream answered with garbage or hung up. Every hop has its own timer, so the same slow request can surface as a 504, a 524 or a 502, depending on who gives up first.

Read the error log before you touch a config

The proxy that returned the 504 logged why. Start there.

# nginx
grep "upstream timed out" /var/log/nginx/error.log | tail -20

# Apache (apache2/error.log on Debian/Ubuntu, httpd/error_log on RHEL-family)
grep "timeout specified has expired" /var/log/apache2/error.log | tail -20

# PHP-FPM (log path varies by distro)
grep -E "max_children|execution timed out" /var/log/php*-fpm.log

The nginx line reads upstream timed out (110: Connection timed out) while reading response header from upstream, and its upstream: field tells you whether it was PHP-FPM’s socket, an app on a local port or a remote backend.

Behind a CDN, take it out of the path and time the origin directly:

curl -s -o /dev/null -w "%{http_code} %{time_total}s\n" \
  --resolve example.com:443:203.0.113.10 https://example.com/slow-page

A 200 that takes 45 seconds tells you more than any dashboard. The page works. It’s just slow.

Which layer sent it

What the visitor seesWho generated itDefault timerFirst thing to check
504 from nginxnginx, waiting on PHP-FPM or an app server60 s (proxy_read_timeout, fastcgi_read_timeout)nginx error.log: “upstream timed out”
504 from Apachemod_proxy or proxy_fcgiProxyTimeout, inherited from Timeout: 60 s stock, 300 s in Debian and Ubuntu packagesApache error log: “The timeout specified has expired”
524Cloudflare: it reached your origin but got no response125 sOrigin requests that run close to two minutes
504 on a Cloudflare-branded pageYour origin returned the 504; Cloudflare passed it onYour origin’s timerThe origin’s own error log
502 instead of 504The upstream died or was killed mid-requestrequest_terminate_timeout, off by defaultPHP-FPM log: “execution timed out”

nginx: proxy_read_timeout and fastcgi_read_timeout

Both default to 60 seconds (proxy module, FastCGI module). The first covers HTTP backends, the second PHP-FPM.

The detail people miss: the timer runs “only between two successive read operations,” not across the whole response. A backend that sends something every 50 seconds never trips it. One that thinks for 61 seconds before sending headers always does.

If one endpoint legitimately takes minutes, raise the timer for that location only:

location /reports/ {
    proxy_pass http://127.0.0.1:8080;
    proxy_read_timeout 300s;
}

Don’t set 600 seconds server-wide. Slow requests just pile up longer, each one holding a worker.

Apache: ProxyTimeout and Timeout

In Apache 2.4, ProxyTimeout defaults to the value of Timeout, and Timeout defaults to 60. Debian’s packaged apache2.conf, which Ubuntu inherits, sets Timeout 300, so the same slow backend can time out after one minute on Apache’s stock setting and after five on Debian or Ubuntu.

For one slow backend, override it on the ProxyPass line instead of globally:

ProxyPass "/reports/" "http://127.0.0.1:8080/reports/" timeout=300

PHP-FPM: the usual suspect

If a PHP site 504s under load, this is where I’d look first. Every PHP-FPM worker handles one request at a time, and pm.max_children caps how many workers exist. PHP’s sample pool file ships with 5. When every worker is busy, new requests wait in the socket’s backlog while nginx waits too, and 60 seconds later nginx returns a 504 even though no single script ran long.

The log says so plainly: WARNING: [pool www] server reached pm.max_children setting (5), consider raising it. Turn on pm.status_path and watch listen queue and max children reached during a busy hour.

Size it from memory, not hope. Say the box has 16 GB, MySQL takes 6 and everything else takes 2, leaving 8 GB for PHP. If workers average 80 MB resident (ps -o rss= -C php-fpm; the process name varies by distro), you can afford about 100 children. Past that you’re swapping, which is worse than queueing.

Then line up the timers so the innermost layer gives up first. PHP’s max_execution_time defaults to 30 seconds (php.ini docs), but on Linux it isn’t affected by system calls or stream operations, so a script stuck on a slow API or database can blow straight past it. That’s what request_terminate_timeout is for. It’s off by default; set it under nginx’s timer, say 45 seconds against 60, and when it fires the FPM log names the script and nginx returns a 502 instead of an anonymous 504.

Also set request_slowlog_timeout = 10s with slowlog pointing at a file. FPM then dumps a PHP backtrace for any request over 10 seconds, which usually points straight at the plugin, query or outbound HTTP call behind the 504.

LiteSpeed

LiteSpeed uses its own names. Connection Timeout on the server tuning page is the “maximum connection idle time allowed during processing one request.” The external app’s Initial Request Timeout (docs) caps how long it waits for lsphp to answer the first request on a new connection. Read the values in the WebAdmin console rather than assuming defaults. And if every lsphp process is busy, requests wait, same as with PHP-FPM.

Cloudflare 524 vs an origin 504

A 524 means Cloudflare connected to your origin, but the origin “did not provide an HTTP response before the default 125 seconds Proxy Read Timeout” (Cloudflare’s 524 guide). Your nginx was still waiting at the two-minute mark, probably because someone raised its timer. Only Enterprise plans can raise Cloudflare’s limit. Cloudflare suggests a DNS-only subdomain for longer jobs; I’d make the job a background task that emails a link when it’s done.

A 504 on a Cloudflare-branded page is your origin’s own 504, passed through (502/504 guide), so go back to your nginx or Apache log. A 504 on a blank, unbranded page came from Cloudflare itself.

When it’s capacity, not config

More workers only help if the box has headroom. Check that first:

  • uptime against nproc: a load average well above your core count means requests are queueing for CPU.
  • vmstat 1 5: a high r column is CPU queueing, nonzero si and so mean swapping, high wa means waiting on disk.
  • iostat -x 1: await is milliseconds per I/O request including queue time (man page), and %util near 100 means the device is saturated.
  • MySQL: SHOW FULL PROCESSLIST; during an incident, and the slow query log the rest of the time.

If raising pm.max_children pushes you into swap, you need RAM. A pinned CPU means more or faster cores. A climbing await means faster disks, and our NVMe vs SATA comparison shows what that buys. A single 40-second query is different: no hardware fixes that, but an index might.

Once the numbers say capacity, stop tuning and move the site to a bigger dedicated server.

FAQ

What’s the difference between a 502 and a 504?

Both come from a proxy. A 504 means the upstream never answered before the timer ran out. A 502 means it answered with something invalid or dropped the connection, which is what you’ll see when PHP-FPM kills a worker mid-request.

Should I just raise proxy_read_timeout?

Only for endpoints you know are slow, like exports, and only in that location block. Raised server-wide, it hides the real problem. Behind Cloudflare it won’t help past 125 seconds anyway.

Why does Cloudflare show 524 instead of 504?

A 524 means Cloudflare reached your origin and waited 125 seconds for a response that never came. If your own proxy times out first, at nginx’s 60-second default for instance, visitors see your origin’s 504 on a Cloudflare-branded page instead.

Is a 504 my host’s fault?

Usually not. On a server you run, it’s almost always the application, the PHP-FPM pool or the database being too slow for the proxy’s timer. Host-side hardware or network trouble tends to look like refused connections and packet loss, not tidy 504s.

Not sure where the time is going? Our slow website assessment is free: an engineer looks at your traffic, server load and resource use and tells you what the site needs. If you’d rather hand the box over, managed servers are $99 a month per server ($79 with cPanel) and cover OS patching, monitoring and troubleshooting, or send your specs through the custom quote form.

Related articles