Automated health checks that catch a dead proxy endpoint early
A dead proxy endpoint rarely announces itself. The modem is still powered on. The SIM still shows registered on the network. A basic ping might even come back fine. But the moment you try to push actual traffic through it, nothing comes out the other side, or it comes out with the wrong IP, or it takes eleven seconds to answer a request that should take one. If you’re only checking whether a device is “up,” you’ll miss all of this. Proxy health check automation exists because uptime and usefulness are two different things, and on a real mobile network, they drift apart constantly.
Why a proxy can look alive and still be dead
We run physical hardware: racks of modems, each fed by a SingTel, M1, or StarHub SIM, each holding a cell connection the same way a phone in your pocket does. That means every endpoint inherits everything that happens to a normal mobile subscriber. A SIM can run out of its data allocation mid-month and get throttled to near-zero. A tower handoff can leave a session in a broken state where the radio link is technically registered but data doesn’t move. A modem’s software stack can lock up after days of uptime and stop responding to new connections while still answering old ones. Power can blip on one rack unit without touching the ones next to it. None of these show up as “device offline.” They show up as a proxy that accepts a connection and then goes nowhere, or one that silently reroutes to a different IP than the one assigned to it.
This is the core reason a TCP handshake test isn’t a health check. A handshake only proves the port is listening. It says nothing about whether a request routed through that endpoint actually reaches the open internet, gets a real response, and comes back with the IP you expect.
What we actually check on every endpoint
An endpoint check that means anything has to move at the application layer, not just the network layer. For each endpoint in the pool, we send a real HTTP request through the proxy to an external target and look at three things: did we get a valid response, did it come back fast enough to be usable, and does the IP it reports match the IP that’s supposed to be assigned to that SIM right now.
That last check matters more than people expect. Mobile carriers rotate IPs on their own schedule, and a modem that’s lost its data session can sometimes fail over to a stale cached address or answer through the wrong interface entirely. If you don’t compare the returned IP against what the endpoint is supposed to be carrying, you can end up handing a client an endpoint that’s quietly proxying through the wrong SIM, or not proxying at all, just echoing a cached response.
Response time is the second signal, and it needs a baseline, not a fixed number. A cell connection that normally answers in under a second and suddenly takes six seconds isn’t dead, but it’s degraded, usually from a congested tower or a SIM approaching a throttling threshold. Flagging that early lets us pull it before it fails outright, instead of waiting for a client’s scraper to time out first.
Why a single failed check isn’t enough
Mobile networks have jitter built into how they work. A tower handoff, a moment of weak signal indoors, a brief queue on the carrier side, any of these can cause one check to fail on a perfectly healthy endpoint. If the automation marks something dead after one bad response, you’ll spend more time chasing false positives than real failures, and worse, you’ll pull working endpoints out of rotation for no reason.
So the check needs consecutive failures before it acts, not a single miss. We run checks on a short interval and only mark an endpoint down after it fails several times in a row, spaced a few minutes apart. That window is long enough to ride out normal mobile jitter but short enough that a genuinely dead endpoint gets caught before it does much damage. The interval itself is a tradeoff too: check too often and you’re adding load to a SIM that might already be data-constrained, or generating a traffic pattern that looks like abuse to whatever target site you’re testing against. Check too rarely and a dead endpoint sits in the pool for longer than it should.
Pulling a dead endpoint out of rotation before a client hits it
The point of catching a failure isn’t just to know about it, it’s to act on it before it reaches someone’s scraping job or browser automation session. Once an endpoint fails its consecutive-check threshold, it gets marked unavailable and taken out of whatever pool assigns endpoints to sessions. That has to happen automatically and immediately, because the alternative is a client’s request landing on a dead IP, timing out, and either breaking their job or, worse, getting logged by the target site as suspicious traffic from a number that never actually connected.
This is also where sticky sessions get tricky. If a client’s job depends on holding the same IP across multiple requests, and that IP dies mid-session, the automation needs to flag the session as broken rather than silently keep routing to a corpse. Quietly rerouting to a different IP mid-session can be just as damaging as no response at all, since it breaks the continuity the client was relying on in the first place.
What automatic recovery can and can’t fix
Some failures are self-correcting with a nudge. A modem that’s locked up can often be recovered by toggling the SIM’s network registration, essentially forcing it to drop and rejoin the network, the mobile equivalent of turning airplane mode off and on. That fixes a real share of the soft failures we see, the ones caused by a stuck data session or a bad tower handoff that the radio never recovered from on its own.
What it doesn’t fix is anything physical or carrier-side. A SIM that’s actually run out of data for the month isn’t coming back until the allocation resets. A modem with a failed USB controller isn’t coming back until someone swaps it. A carrier-side block on a specific IP range isn’t something a software reset touches. So the automation’s job is narrow: try the cheap recovery step first, recheck, and if it’s still failing after that, stop trying and flag it for a person to look at physically. Retrying a hardware fault forever just delays the moment someone actually walks over to the rack.
What the logs tell us over time
A single dead endpoint is a maintenance task. A pattern across dead endpoints is information. We keep timestamped logs per endpoint, and over weeks you start to see things a one-off check would never surface: a SIM that fails reliably at the same time every day, which usually points to a data cap resetting or a carrier maintenance window; a specific device that degrades after a few days of continuous uptime, pointing to a software or thermal issue on that unit rather than the network; a cluster of endpoints on the same carrier failing around the same time, which is a carrier-side event, not something wrong with our hardware.
None of that comes from watching one check pass or fail. It comes from the automation being consistent enough, over enough time, to make the pattern visible instead of just the individual incident.
What this looks like from the client side
None of this is visible to a client running a scraping job or a browser automation session, and that’s the point. What they should see is an endpoint that either works or gets swapped out before they ever touch it, not a job that hangs for thirty seconds because a dead SIM sat in rotation. A proxy pool without automated health checks eventually hands out failures as if they were working endpoints, and the client finds out the same way we would if we weren’t checking: a broken request, wasted time, and no idea why.
If you’re evaluating mobile proxies for scraping or automation work and want to know how the pool behind them actually gets kept clean, take a look at what we run.
Get new guides and videos first — join the Telegram channel.