MTU, fragmentation and the requests that hang forever
A customer told me his line was dead. It wasn’t. Small requests through it were coming back in under a second, the IP check page loaded fine, and his own health probe had been green all night.
One particular site hung, every time. The connection opened, the handshake finished, and then nothing came back at all until something timed out ninety seconds later.
That pattern is one of the most misdiagnosed things in this business, and it almost always comes down to a number that nobody thinks about until it bites them.
What MTU actually is
Every link between two machines has a maximum size of packet it will carry in one piece. That’s the MTU. On ordinary Ethernet it’s 1500 bytes, and it’s been that number for so long that most software just assumes it.
A mobile carrier isn’t Ethernet. The mobile core wraps your traffic in its own tunnelling before it ever reaches the internet, and that wrapper takes bytes out of the space available for your actual data. So the usable MTU on a mobile line is commonly something below 1500, and the exact figure depends on the carrier and the APN.
Then you add whatever you’re running on top. A VPN takes another chunk. A tunnel between your machine and the proxy takes another. Each layer wraps the last one, and each wrapper eats into the payload.
By the time your request is actually moving over the air, the biggest packet that can survive the whole path might be 1400 and something, and your operating system is still cheerfully building 1500-byte packets because that’s what the network card in front of it said.
The part that’s supposed to handle this
TCP does have a mechanism for this, and it works most of the time. During the handshake, both ends advertise a maximum segment size, the largest chunk of data they want in a single packet. Whoever has the smaller number wins, and both sides stay under it.
The problem is that the handshake only knows about the two endpoints. It doesn’t know about a link somewhere in the middle with a smaller limit than either end.
For that there’s path MTU discovery. Your machine sets a flag on the packet saying don’t fragment this. If it reaches a link that can’t carry something that big, the router at that point is supposed to throw the packet away and send back an ICMP message saying fragmentation needed, here’s the size that would have fit.
Your machine reads that, lowers its idea of the path size, and resends. That’s the design, and when it works you never see any of it.
Why it stops working
That ICMP message is the single point of failure in the whole arrangement.
ICMP gets blocked constantly. Firewall rules written by somebody who decided ICMP was a security problem and dropped all of it. Cloud security groups that only allow TCP. Middleboxes that filter what they don’t recognise.
When that message is dropped, your machine never finds out its packet was too big. It just knows it didn’t get an acknowledgement, so it does what TCP always does, which is send it again, at the same size, into the same link that will discard it exactly the same way.
That’s a path MTU black hole. It isn’t a failure. Nothing reports an error. The connection is technically alive the entire time, and no data is ever going to cross it.
The signature you should learn to recognise
This failure has an extremely distinctive shape, and once you’ve seen it you won’t misread it again.
Small requests work perfectly. Anything that fits in a single packet goes through, comes back, and looks completely healthy.
The connection establishes. The TCP handshake is tiny, so it never trips the limit. The TLS handshake usually completes too, since those packets are mostly small.
Then the first large thing hangs. A POST with a real body in it. A response with a full page of HTML. A file. Anything that needs the sender to fill a packet to the top.
And it hangs rather than failing. No reset, no error code, just silence until a timeout somewhere gives up.
If you have a proxy where the IP check page loads instantly and one specific site never returns anything, you’re almost certainly looking at this rather than a broken line.
How to prove it in two minutes
The test is simple and it works from any machine.
Send a ping with the don’t-fragment flag set and a payload size you choose, then walk the size down until it starts getting through. The largest payload that survives, plus 28 bytes for the headers, is the real path MTU.
On Windows that’s ping -f for don’t fragment and -l for the length. On Linux it’s -M do and -s for the size.
Start at 1472, which corresponds to a 1500-byte path. If that fails, try 1400, then 1350. The point at which it starts working tells you the number, and if 1472 fails while 1350 succeeds, you’ve found your problem and you didn’t need a packet capture to do it.
Do this test from the machine that’s actually having the trouble, going to the host that’s actually hanging. Testing to a different destination proves nothing, because the small link might be on that specific path and nowhere else.
What to actually do about it
There are three fixes and they sit at different points.
The cleanest one is MSS clamping. The device at the edge of the tunnel rewrites the maximum segment size during the handshake so both ends agree on something that fits the real path. No ICMP needed, no discovery, no black hole, because the size was negotiated correctly in the first place.
That’s the right place to solve it, and it’s where we solve it. The tunnel interfaces on our side clamp the segment size to the real usable figure for the carrier path, so a customer connecting through a line never has to know any of this exists. Done properly it’s invisible. Left undone it quietly costs somebody a week.
The second fix is lowering the MTU on your own interface, which is blunt but effective if you control the machine and can’t control anything else. Set it to something safely under the real figure and your stack stops building packets that can’t survive.
The third is to stop dropping ICMP. If you administer the firewall in front of your own scraper, let the fragmentation-needed messages through. That costs you nothing security-wise and it restores the one mechanism designed for exactly this situation.
The DNS version of the same problem
There’s a version of this that hits DNS rather than your requests, and it confuses people badly because it presents as a name resolution problem. A DNS answer over UDP that’s too large for the path gets fragmented, the fragments hit a link that won’t carry them, and the reply never arrives. So lookups for most names work and one particular name times out every time, usually one with a lot of records behind it.
The tell there is that the same lookup succeeds when you force it over TCP. If a query hangs on UDP and answers instantly over TCP, you’ve found the same problem wearing a different hat, and the fix is in the same place.
What to log so you recognise it next time
Two numbers make this obvious in hindsight, and almost nobody records them.
Time to first byte, separately from total time. This failure has a very particular shape, where the connection and the headers arrive quickly and then the clock runs out during the body. If your logs only hold a single duration, every slow request looks the same as every other slow request.
And the size of what you were sending or expecting. A log line that says the requests which hang are the ones over a certain size is the entire diagnosis, handed to you, before you’ve opened a single capture.
If you do end up taking a capture, the giveaway is a segment of the same size being sent over and over with nothing coming back for it. That’s the pattern worth being able to spot on sight.
What makes it worse
IPv6 removes in-path fragmentation entirely. On IPv6 a router is never allowed to break a packet up, so the only thing standing between you and a black hole is that ICMP message getting back to you. That makes this failure more common on IPv6 rather than less.
Tunnels stacked on tunnels are the other multiplier. Every layer subtracts, and people rarely add up the total. A WireGuard tunnel inside a carrier tunnel inside whatever your provider runs is three subtractions from 1500, and the number at the end is smaller than most people would guess.
And QUIC sidesteps some of this by running over UDP with its own path discovery, which is one of the few genuinely good things about HTTP/3 from a debugging point of view. It changes the problem rather than removing it.
Why this gets blamed on the proxy
Here’s why I see this land in my inbox rather than somewhere else.
The customer’s own machine works fine. Their home connection has a normal MTU and no small links in the path, so every test they run locally passes.
The moment they route through a mobile line, the path gets a smaller link in the middle of it, and the same code that’s worked for a year starts hanging on one site.
So the change that broke it was the proxy, from where they’re sitting, and that’s a completely fair inference. It’s just the wrong one. The proxy exposed a size assumption that had never been tested before, and the assumption was in their stack the whole time.
The honest limit
Some of this genuinely is the provider’s job, and you should expect it to be handled. If a line hands you a path where standard-sized packets vanish silently, that’s something the operator should have clamped, and you should say so.
What nobody can fix from my side is a firewall you administer dropping ICMP, or a container image with an MTU set for a datacentre network that doesn’t match the tunnel it’s now running inside. Those are yours, and this is the failure that most rewards knowing which half is which.
So the takeaway is short. Small requests fine, large ones hanging forever, connection alive throughout. That’s MTU. Walk the ping size down, find the real number, and either clamp the segment size or lower your interface, and the site that’s been hanging for a week starts answering immediately.
If you want Singapore lines where the segment size is already clamped to the real carrier path so this never becomes your afternoon, everything is at singaporemobileproxy.com.
Get new guides and videos first — join the Telegram channel.