Your site goes down every few hours, for two to five minutes. Then it comes back on its own. Hosting monitoring says everything is green, the customer says they saw a white page, and you're stuck between "maybe it was their internet" and "something must be wrong somewhere." This is the most frustrating type of site outage, because it never happens at the same time as you.
The first thing you should do is turn the outage from a "feeling" into a "number." As long as you only have a verbal account, any change to the server is a gamble.
Why Host Panel Monitoring Doesn't See Site Outages
Hosting panels usually ping ICMP or send an HTTP request to the main page once every 5 minutes. There are two structural problems here. First, the 5-minute interval completely misses 90-second outages; if an outage falls between two checks and ends before the next check, it's never recorded. Second, an ICMP ping only tells you the server's network card is responding, not that PHP or MySQL or the web server upstream is healthy.
Set up an external monitoring service with a 30 or 60 second interval and configure it on a real endpoint that depends on the database and application code, not on a static file. For example, a health path like /healthz that runs a lightweight query to the database and returns a 200 code. If you only check the main page, the page cache may respond and you won't see the database outage.
A point that's rarely mentioned: monitoring from one location isn't enough. If the monitor checks from a datacenter inside Iran and your user connects from abroad, the problem may be the international network route, not your server. Set up at least two monitoring locations with different ISPs. To check the route and domain records, you can also use the DNS and network lookup tool to see whether domain resolution is happening correctly at the moment of the outage.
Time Correlation: What Should We Compare a Site Outage To?
When an outage window is recorded, the right question isn't "why did it go down?" but "what exactly was running on the server during that window?" You need to put three sources side by side:
- Web server logs (e.g.,
/var/log/nginx/error.logand access log) with precise timestamps - System and cron logs:
journalctl -u cron --since "2025-01-14 03:00" --until "2025-01-14 04:00" - Server resource graphs (CPU, RAM, disk I/O) with one-minute granularity, not hourly averages
In most cases I've seen myself, the pattern is one of these: a heavy cron job (backup, import, table cleanup) that coincides with real traffic; or a script with a memory leak that triggers the OOM killer every few hours; or an external process like a crawler bot hammering heavy pages at a high rate.
To see which process was under the most pressure at the moment of the outage, if atop is installed, use atop -r /var/log/atop/atop_20250114 -b 03:00. If you don't have it, install it today; without historical data, troubleshooting irregular outages is practically a coin toss.
Don't Blame Cron First, But Check It Before Everything Else
A common mistake: the site admin disables the backup cron, there's no outage for two days, and they conclude the problem is solved. Then three weeks later the outage comes back. The reason is that cron was only the synchronizer, not the cause. If your backup occupies disk I/O for 40 minutes and the site is served from the same disk, the real problem is that the backup and traffic fall on a shared resource. The right solution is to move the backup time or limit the I/O rate with ionice -c2 -n7, not to delete the backup.
This Is Where They Go Wrong
The sentence I hear more than any other: "The site is up right now, so the problem isn't the server." This argument is wrong and is exactly what delays troubleshooting for weeks. Short outages are usually of the momentary saturation kind: a queue fills up, a timeout occurs, a service restarts, and everything returns to normal within a few seconds. By the time you SSH in and see everything healthy, that window has closed.
The sign is also this: the user says "I got a 502," you see the line upstream timed out (110: Connection timed out) while reading response header from upstream in the web server log, but when you manually send the same request, you get a response in 200 milliseconds. This contradiction is itself evidence, not proof of the server's innocence.
What Should We Measure to Make Outages Predictable
You should always have three numbers. First, response time at the 95th and 99th percentiles, not the average. An average of 200 milliseconds with a 99th percentile of 8 seconds means some of your users effectively can't open the site. Second, the number of active connections versus the configured limit; if worker_connections is set to 1024 and you reach 900 during peak hours, you're a short distance from an outage. Third, the memory usage of PHP-FPM or application processes, separately for each pool.
| Monitoring Method | What It Catches | What It Misses |
|---|---|---|
| ICMP ping every 5 minutes | Complete server shutdown | Application outage, database outage, short windows |
| HTTP check every 30 seconds on /healthz | Code errors, database outage, severe slowness | Network route problems for specific users |
| Log-based with alerts on error rate | Gradual patterns, scattered errors | When the server itself isn't writing logs |
If your site is on a shared server or a small VPS and you don't have the control needed to install an agent, combining external HTTP monitoring with the application's own logs is enough. But when traffic rises and outages have a real cost, a dedicated server gives you the ability to control the kernel, network parameters, and cron scheduling yourself; something that isn't available in a shared environment.
Backup, Recovery, and Their Difference at the Moment of an Outage
A point that becomes clear during a crisis: having a backup isn't the same as having fast recovery. If your backup is 20 gigabytes and stored on the same server, in a disk failure you have neither. Put the backup on a separate destination and measure the actual recovery time at least once a month. The number you get (say, 45 minutes for a full restore) is what you should tell stakeholders, not "we have a backup."
For sites running on managed hosting, Linux hosting with standard file and database management tools makes recovery and log review easier; but you still need to know yourself which table is how large and how long its recovery takes.
A Practical Checklist for Tonight
- Create a health endpoint that depends on the database and monitor it at a 30-second interval from two locations.
- Send web server and cron logs with identical timestamps to a central destination so correlation becomes possible.
- Store resource graphs with one-minute granularity, not hourly averages.
- Compare the scheduling of heavy crons with peak traffic hours and, if they overlap, move or limit them.
- Practice a restore from backup once and note its actual time.
If you want a more accurate picture of your site's current state before any infrastructure change, use the free webmaster tools to check speed and responsiveness. And if after these steps you still haven't found a pattern, the problem is probably in a layer you're not measuring; usually DNS or the network route. On the ServerNet blog, real examples of this type of troubleshooting are documented.
Frequently Asked Questions
Why does my site only go down for some users?
It usually means the problem isn't in the server, it's in the route to the server. Differences in DNS resolver, the international route, or even the user's local DNS cache can cause one person to see the site and another not to. To diagnose, test from several different geographic locations and compare the domain resolution results.
Are a few-second outages really important?
Yes, if they fall on endpoints where the user is in the middle of a transaction. A 10-second outage in the middle of a payment leaves the order half-finished and may create inconsistent data. In addition, crawlers also see these errors, which affects page ranking.
What interval is suitable for monitoring?
For commercial sites, 30 to 60 seconds is a reasonable starting point. An interval shorter than 30 seconds usually increases cost and alert noise without showing anything new, unless your SLA is truly strict.
How do I tell whether the outage is from the database or the web server?
Create a separate health endpoint for the database that only runs a simple query. If that endpoint errors during the outage window but a static page is healthy, the problem is the database. If both error, the problem is probably in the web server or network layer.
Comments 0
No comments yet — be the first!