Your web server writes down every request it handles. That file, the access log, is the most reliable account of what happens on your site. Analytics scripts miss bots, visitors who block trackers, and anything that fails before a page finishes loading. The log sees all of it, because it is written by the server itself at the moment the request arrives and is answered.
It is also the first thing a support engineer asks for when a site is slow, odd or under attack. Not because logs are glamorous, but because they settle arguments. Somebody says the site was down at three; the log shows a steady stream of 200 responses at three. Somebody says a crawler is eating the bandwidth; the log shows which address, how often, and what it asked for.
You do not need special software to read one. A text editor will do for a small file, and a few shell commands will do for a large one. This article takes a single line apart, shows where the file lives on different kinds of hosting, runs through the questions worth asking, and ends with the privacy side, which people tend to forget until a regulator reminds them.
Anatomy of a line
A typical entry in the common "combined" format looks like this:
203.0.113.7 - - [03/Oct/2026:09:14:22 +0300] "GET /blog/ HTTP/2.0" 200 18432 "https://example.com/" "Mozilla/5.0 ..."
Reading left to right: the visitor's address, two fields you can usually ignore (they were meant for an old identity service and a login name, and are nearly always a hyphen), the time with its offset from UTC, the request itself (method, path, protocol), the status code, the size of the response body in bytes, the page that referred the visitor, and finally the browser or bot identifying itself.
Two details trip people up. The size is the body only, so a 304 "not modified" reply shows a tiny number or a hyphen even though the visitor saw a full page from their own cache. And the time is the server's time at the moment the request was logged, in the server's configured zone, which is not necessarily yours. Check the offset before you hunt for "what happened at three".
Where the file lives
On shared hosting you rarely see the real file in the real place. The panel usually offers a log viewer, a download link for the current month, or a raw archive per day. If you cannot find any of those, ask support where they are; it is a routine question and the answer is often one click away in a menu you have not opened.
| Hosting type | What you normally get | Typical location |
|---|---|---|
| Shared | Panel viewer or download, often one file per domain, kept for weeks | A "logs" folder above the web root, or a panel page |
| VPS or dedicated, Apache | Full access, you set the format | /var/log/apache2/access.log or /var/log/httpd/ |
| VPS or dedicated, nginx | Full access, format defined in the config | /var/log/nginx/access.log |
| Managed WordPress | Often a dashboard with filtered views, sometimes a raw download on request | Provider's control panel |
If you run your own server and cannot find the file, ask the web server where it writes. For nginx, nginx -T | grep access_log prints every configured path. For Apache, look for CustomLog in the virtual host with grep -r CustomLog /etc/apache2/ (or /etc/httpd/ on Red Hat style systems). One server can have a separate log per site, which is helpful until you search the wrong one for an hour.
One more thing to know: if your site sits behind a CDN or a reverse proxy, the address in the first field may be the proxy's, not the visitor's. Servers can be told to trust a forwarding header and log the real address instead, but only from proxies you actually use. Otherwise anyone can claim to be anyone.
Five questions a log answers fast
The commands below assume the combined format, where the address is field 1, the status is field 9 and the path is field 7. Adjust if your format differs.
Which status codes are you returning?
awk '{print $9}' access.log | sort | uniq -c | sort -rn
A healthy brochure site is mostly 200 and 304, with a sprinkling of 301 and 404. A sudden pile of 500s means the application is failing; a pile of 403s often means a security rule is firing. The status code reference covers what each family means.
Who is hitting you hardest?
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head
Which pages are requested most? Same idea with $7. If a path you have never heard of tops the list, that is worth a look.
Is a bot hammering the login page?
grep 'wp-login.php' access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head
What happened at three in the afternoon? The time stamp is plain text, so filter on it:
grep '03/Oct/2026:15:' access.log | less
A worked example: the afternoon slowdown
A small online shop (a composite, not a real one) reports that the site crawled for about twenty minutes in the afternoon. Nothing was deployed. Analytics shows nothing unusual, which is the first clue, since analytics only counts what runs JavaScript in a browser.
Start with the clock. The server logs in UTC+3, and the owner says "around half past three", so filter on 03/Oct/2026:15: and count requests per minute:
grep '03/Oct/2026:15:' access.log | awk '{print substr($4,15,5)}' | sort | uniq -c
The output shows roughly 40 requests a minute until 15:22, then 900 a minute until 15:41, then back to normal. So it was traffic. Now who?
grep '03/Oct/2026:15:[2-4]' access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head -3
8412 203.0.113.50
311 198.51.100.23
102 192.0.2.88
One address made most of the requests. Look at what it asked for:
grep '^203.0.113.50 ' access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head
Say the answer is the search page with a different query string each time. That explains the slowdown: search results cannot be cached and each one costs a database query. The fix is a rate limit on that path, or a block on the address if it is plainly abusive, plus a look at whether the search feature needs a cache of its own. Without the log, this would have been a shrug and a vague promise to "keep an eye on it".
Telling bots from people
The user agent field names the client, but anyone can write anything there. A request claiming to be a famous search crawler is only that if its address agrees. The check is a reverse DNS lookup followed by a forward one:
host 203.0.113.50
50.113.0.203.in-addr.arpa domain name pointer crawler-203-0-113-50.search.example.
host crawler-203-0-113-50.search.example
crawler-203-0-113-50.search.example has address 203.0.113.50
The names above are placeholders, but the logic is real: the reverse name must belong to the crawler operator's published domain, and that name must resolve back to the same address. If either half fails, it is an impostor. Major search engines publish the domains to expect, and some publish address lists as well.
Unwanted bots fall into two groups. Many announce themselves accurately, with a name in the user agent, and can be refused by that name in a server rule or a robots.txt entry (the polite ones obey it). The sneaky ones pretend to be an ordinary browser, rotate addresses, and need rate limits or a web application firewall rather than a block list. A good rough tell is behaviour: no referrer, no requests for images or stylesheets, and a steady rhythm of one page every second for an hour.
Retention, rotation and privacy
Logs contain visitors' IP addresses, which count as personal data under some laws, the GDPR among them. That does not make logging unlawful. Security and fault-finding are legitimate reasons to keep them. It does mean you should decide how long you keep them and be able to say why.
Rotation is the mechanical side. On a VPS the usual tool is logrotate, which renames the current file on a schedule, compresses the old ones and deletes the oldest. A common pattern is daily rotation, fourteen to thirty days kept. A log that is never rotated eventually fills the disk, and a full disk takes down far more than logging. On shared hosting the provider normally handles this, and your job is to download anything you need before it ages out.
If you pass logs to a third party, for example a consultant or a log analysis service, think about what is in the request line. Query strings sometimes carry email addresses, order numbers or tokens. Anything sensitive should not be in a URL in the first place, but logs are where such mistakes become visible.
Commands worth running
- Find the log for your site: the panel viewer on shared hosting, or
nginx -T | grep access_logon a server you manage. - Run
tail -f access.logand load a page in your browser. You should see your own request appear within a second. If it does not, you are reading the wrong file. - Note your own address (search for "what is my IP" or use
curl ifconfig.me), then confirm it appears in the first field. - Run the status code count above and compare the ratio of 200s to 4xx and 5xx responses.
- Request a page that does not exist and look for the 404 line. This tells you that error pages are logged at all.
The troubleshooting guide has a longer list of first checks for slow and failing sites, and most of them start with a look at the log.