Reading proxy logs to find the exact request that triggered a block
The block is never random
When a scraping job or an automated account gets flagged, the instinct is to blame the IP. Swap the proxy, restart the job, hope it doesn’t happen again. That works often enough that people never learn what actually triggered the block, which means it happens again a week later on a different IP.
Every block is a response to something. A status code, a redirect to a captcha page, a session that silently stops returning real data. Somewhere in your logs there is a specific request, at a specific timestamp, that changed how the target site treated you. Finding it is a log reading exercise, not a guessing exercise.
What you actually need logged
Most people running scraping jobs only log what the scraper needs to keep working: URL, status code, maybe response time. That’s not enough to reconstruct a block. To do proxy log analysis properly you need two logs that you can align on a timestamp:
- The proxy-side log: source IP or device ID, destination host, timestamp, bytes in and out, and the upstream response code the proxy itself saw.
- The application-side log: the full request (method, path, query string, headers actually sent), the response status, and at least the first few hundred bytes of the response body.
The body sample matters more than people think. A 200 status code with a captcha page in the body looks identical to a successful response if you’re only logging status codes. We’ve seen jobs run for hours logging “success” against every request while quietly scraping captcha pages the whole time.
If you’re running this against a mobile proxy setup, also log which physical SIM or device handled the request. On a farm with SingTel, M1, and StarHub lines running side by side, carrier behavior isn’t identical. Some carriers rotate the carrier-grade NAT IP more aggressively than others, so a device ID to IP mapping that was true five minutes ago might not be true now. Without that device ID column, you can’t tell later whether a flagged session got flagged because of what the automation did, or because the carrier moved it onto an IP that already had a bad reputation before you touched it.
Finding the boundary
Once you have both logs, the actual technique is simple: find the last request that got a clean, expected response, and the first one that didn’t. Everything you need is usually in the gap between those two lines.
Sort both logs by timestamp and merge them, or just eyeball them side by side if the volume is small. Walk forward until the response shape changes: a 200 becomes a 403, a JSON payload becomes an HTML error page, a normal page becomes a redirect to a login or verification screen. That line is your boundary.
Then diff the request that came right before it against the one that crossed it. Look at:
- Headers. Did a header get dropped or change value? A missing
Referer, aUser-Agentthat doesn’t match the one used earlier in the session, anAccept-Languagethat suddenly doesn’t match the geography the IP suggests. - Timing. Was there a burst, requests fired closer together than the rest of the session? Rate-based blocking usually triggers on a threshold within a rolling window, not a hard cap, so the trigger request often isn’t unusual on its own, it’s unusual next to the ones just before it.
- Cookies or session tokens. Did the session cookie change, get dropped, or get sent from a different IP than the one that received it? This is common when a proxy rotates mid-session on a job that assumed a sticky connection.
- The endpoint itself. Some paths are watched harder than others. A request to a checkout or account page gets scrutinized differently than a request to a public listing page, even from the same session.
Nine times out of ten, once you line the two requests up next to each other, one of these will be visibly different. The value of doing this over time is that you start recognizing which of these four categories is your recurring problem, instead of treating every block as a fresh mystery.
Timestamps have to agree
This sounds obvious but it’s the thing that actually derails most log analysis: the proxy server, the scraping client, and any browser automation layer in between often don’t agree on clock time, or log in different timezones by default. If your proxy log is in UTC and your scraper log is in local time, you’ll spend an hour convinced a request happened before a block when it actually happened after.
Before you trust a correlation, confirm both logs are in the same timezone, ideally UTC, and that clocks on the machines producing them are actually synced. On a hardware farm this matters more than it seems like it should, because devices sitting on different carriers can drift slightly if NTP isn’t enforced consistently across every unit.
What mobile IPs change about the picture
A lot of block analysis assumes the IP is stable for the length of a session, because that’s how datacenter-style setups usually behave. Real mobile connections don’t work that way. Carrier-grade NAT means the IP your request goes out on can be shared with other devices you don’t control, and it can change if the carrier reassigns the NAT pool, which happens on its own schedule, not yours.
That means two things for log reading. First, don’t assume a block followed a specific request just because the IP looked “used” beforehand. Other traffic, from other customers behind the same carrier NAT, can affect an IP’s reputation before you ever send a request through it. Second, log the IP on every single request rather than once per session. If the IP changed partway through and you didn’t record it per-request, you’ll misattribute a carrier-side rotation as a session anomaly, and you’ll never find it because you’re looking at the wrong variable.
Building a habit out of it
The point of doing proxy log analysis isn’t to solve one block. It’s to build a small library of “this is what a rate block looks like in our logs” versus “this is what a header mismatch looks like” versus “this is what a stale cookie looks like.” After you’ve done this a handful of times, the pattern recognition gets fast. You’ll open a log, scan for the boundary, and know within a minute or two which of the usual suspects it is, because you’ve seen the shape before.
Keep the logs. Don’t just discover the cause and move on, save the request pair (last good, first bad) somewhere searchable, tagged with what actually caused it. Six months from now when a similar block shows up on a different job, you’ll have a reference instead of starting from zero again.
None of this requires special tooling. A merged, timestamp-sorted log and a text diff between two requests is enough. The discipline is in logging the right fields in the first place, and in actually looking at the boundary instead of just restarting the job and hoping.
If you want to see how we run and log real SIM-based mobile proxy traffic on actual SingTel, M1, and StarHub hardware, take a look around singaporemobileproxy.com.
Get new guides and videos first — join the Telegram channel.