Hosting Autopsy / The firewall rule that blocked the payment provider

The firewall rule that blocked the payment provider

HOSTING AUTOPSY

6 min read · 1,348 words

This is a composite case written by the editors. It is built from patterns that come up often in support work and is not the account of a particular named person or company.

An online shop selling craft supplies had a bad few days. A burst of suspicious traffic hit the site: login attempts against the admin area, scripted probes for old plugin files, and a scatter of checkout requests that looked like card testing. The developer who looked after the site did what many of us would do at that point. They added a rule that blocked all requests from addresses outside the country where most of the shop's customers live.

It worked. The noise dropped away within minutes. The logs went quiet, the load graph flattened, and the developer went to bed pleased with a tidy solution.

It also blocked the payment provider's servers, which send a confirmation message to the shop when a payment completes. Customers paid successfully and then received no order confirmation, because the shop never heard about it. Support emails piled up. This is a composite, but the mechanism is one of the most common ways a well-meant firewall change causes a quiet commercial problem.

How a payment actually completes

Most people picture checkout as one conversation between the customer and the shop. In practice there are three parties, and the shop learns the result by two separate routes.

The customer is sent to the provider's payment page, or a payment form served by the provider, and enters their card details there. When the payment succeeds, the customer's browser is redirected back to the shop's "thank you" page. That is route one, and it relies on the customer's browser behaving: they might close the tab, lose signal, or never click through.

Route two is the one that matters. The provider's own servers make an HTTP request to a URL on the shop, often called a webhook or callback, saying "payment 4411 is complete, amount £38.50". The shop checks the message is genuine, marks the order as paid, reduces the stock and sends the confirmation email. This is a server-to-server call from addresses that belong to the provider, and those addresses are probably not in your country.

Customer Payment provider Shop 1 pays 2 browser sent back: "thank you" page Country block returns 403 3 callback "payment complete" from the provider's servers The customer sees success; the shop never learns the order is paid.
The browser route works, so the customer thinks all is well, but the server-to-server callback is the one the shop depends on.

What the customers and the staff saw

The first symptoms were customer-side and vague. Emails arrived saying "I paid but have no confirmation". Two said their bank had taken the money and the shop's site still showed the order as "pending payment". A few had paid twice, reasonably assuming the first attempt had failed.

Staff looked at the order list and saw a growing column of pending orders. Stock levels were not dropping. The shop's own contact form was fine, the product pages were fine, and the checkout page loaded quickly. Everything the staff could test by clicking around worked, because staff were in the home country and could pass the rule.

That detail is why it went unnoticed for most of two days. The firewall rule treated staff exactly as it treated customers, and neither of them was the party being blocked.

Wrong turns

The first theory was a plugin conflict. The shop's payment plugin had updated the previous week, so they rolled it back. Nothing changed, because nothing about the plugin was faulty, and a rollback briefly broke stock reporting as a bonus.

The second theory was a problem at the provider. They contacted the provider's support, who said, correctly, that payments were succeeding and that callbacks were being sent. The provider's dashboard even showed the delivery attempts, each marked failed with a 403 response, but nobody on the shop's side had looked at that page.

The third theory was email: perhaps confirmations were being sent but landing in spam. The mail logs showed no confirmation emails being generated at all, which was the important clue. A mail server cannot send an email that the shop software never decided to create.

Diagnosis

The access log settled it. The developer searched the web server log for the callback path and saw nothing, not even failed requests, because the firewall in front of the server dropped them before they were logged there. The firewall's own log did show them.

$ grep "/payment/callback" firewall.log | tail -3
2025-03-04 14:02:11 DENY geo=NL src=192.0.2.45 POST /payment/callback 403
2025-03-04 14:07:12 DENY geo=NL src=192.0.2.45 POST /payment/callback 403
2025-03-04 14:19:40 DENY geo=IE src=192.0.2.46 POST /payment/callback 403

Note the countries. The provider ran its servers abroad, and the callbacks arrived from several places, none of them the shop's home country. A rule meant to stop foreign attackers had stopped a foreign partner with exactly the same effect.

Request 1 Allow list(partners) 2 Rate limitsand targeted rules 3 Broadblocks last known partner: straight through to the shop Order matters: a country block placed first catches the payment provider.
Put the partners you depend on at the top of the rule list, and broad blocks at the bottom.

The fix and the clean-up

The immediate fix was to add the provider's published callback addresses to an allow rule placed above the country block. Providers publish those ranges in their documentation and update the list from time to time, so the developer subscribed to the provider's change notices and wrote the date of the last review next to the rule.

A second useful change was to verify the callback with the provider's signature instead of trusting the source address alone. An address allow-list reduces exposure, but a signature check makes the endpoint safe even if someone spoofs a request.

The clean-up was manual. Many providers retry failed callbacks for a while, with growing gaps, but the retries here had mostly run out by the time the rule was fixed. So the shop opened the provider's dashboard, exported the list of completed payments for the affected days, and compared it order by order with their own pending list. About ninety orders needed marking as paid, confirmations resending, and in around eight cases refunding the duplicate payment. It took two people most of a day.

Why the block looked like a success

It is worth understanding why the developer was so pleased, because the same reasoning will be tempting next time. The traffic that had been bothering the shop really was mostly foreign. Cheap hosting abroad is where a lot of scripted probing starts, so a geographic filter removes a large share of it at a stroke. The graphs improved, the log files shrank, and the rule cost nothing to write.

What the graphs could not show was the harm. Blocked requests produce no errors on the shop's side, because the application never sees them. The only visible effect is an absence: confirmations that are never created, stock that is never reduced. Absences do not trigger alerts unless you have defined what normal looks like, such as "at least a few paid orders every hour in the daytime".

The lesson is not that the developer was careless. It is that a security change has two sets of outcomes, the attacks it stops and the legitimate traffic it stops, and only the first is easy to measure.

Checking it yourself

Make a list of everything that legitimately calls your site from outside. For a shop, that normally includes the payment provider, shipping and label services, the email platform's bounce and event hooks, uptime monitors, and sometimes accounting or stock-sync tools. For each, find the path it calls and the source it uses.

curl -I -X POST https://example.com/payment/callback
# a 400 or 405 from your app is healthy; a 403 from the firewall is not

After any firewall change, place a small real test order and confirm it moves from pending to paid. Check the provider's dashboard for delivery failures too, since most have a webhook log with response codes. The redirect generator and the status code reference help if you are unsure what a code means.

What would have caught it

PreviousThe scheduled job that ran twiceNextThe sale that started at the wrong hour

More from Hosting Autopsy

Autopsy

The backup that could not be restored

A hobbyist forum had a nightly backup running faithfully for over a year. Every morning the control panel...

Autopsy

The certificate on one server, but not its twin

A company ran its site behind a load balancer with two identical web servers. When the certificate was...

Autopsy

The CDN that served one customer's basket to another

A boutique put its site behind a CDN and turned on the option to cache everything, including HTML. Speed...