Reading your own proxy usage data
Reading your own proxy usage data
Nobody messages me about their usage chart. They message me about their invoice.
That is most of the problem in one line. Both are built from the same counting. One of them arrives a month too late to do anything with.
I run the hardware the numbers come off. Real SIMs in real modems, one port per modem, sitting on a shelf in Singapore. Every byte gets counted at my end before it becomes a line on anyone’s bill, so I get to watch what people do with the counting. Mostly they wait for it to turn into money, then react to the money.
A total is a receipt, a shape is a signal
A total answers one question: how much. You can accept it or you can argue about it, and that is the end of what it can do for you.
The chart underneath it answers a different question, which is what has been happening. Volume against time, on a line you control, is a live picture of whether your automation is doing what you told it to. It is the cheapest monitoring most people own and almost none of them open it until an invoice does something strange.
The readings below take about two minutes between them, and need no tool you do not already have.
The flat line where there should not be one
Easiest to spot and the most expensive to miss.
You have jobs on a schedule overnight. Open the chart at the daily or hourly resolution and look at the hours those jobs are supposed to occupy. If nothing moved, the jobs did not run.
Not “the jobs ran slowly” or “the target was down”. Nothing left the line. A scheduler that stopped firing, a token that rotated and never got picked up, a VPS the host rebooted for maintenance, a container that came back without its environment file. All of these produce the same picture: an hourly bar chart with a hole in it.
The monthly total will not show you this. Daytime traffic papers over an overnight gap easily, so the invoice looks unremarkable and you find out at the end of the month, usually when somebody downstream asks where a week of data went.
Weekly resolution makes it even easier, because most jobs have a rhythm that repeats. Once you have seen four normal weeks, a missing Tuesday is obvious without reading a single number. You are comparing pictures, not values.
The reverse reading matters too. Traffic in hours you never scheduled anything for is either a job that did not stop when you thought it had, or credentials being used by somebody who is not you.
Steps, not spikes
A spike goes up and comes back down, and there is usually an obvious story attached to the day it happened.
A step is different, and it is the one that produces the surprise bill. Usage sits at one level, moves to a higher level, and stays there. Nothing came back down. If you changed nothing on your side in the week that step appeared, something in your stack is retrying.
Of every conversation I have had about an unexpected number, this is the most common cause by a distance.
The mechanism is boring. A request fails in a way your code treats as temporary. Your code retries. The retry fails the same way. There is no cap on the attempts, or the backoff resets on a code path nobody considered, and the loop finds a steady rhythm and holds it. Every attempt is a real fetch that really moved bytes. The connection was fine. Nothing in your monitoring is unhappy, because retrying is the correct behaviour and your code is executing it correctly, indefinitely.
A lot of retry logic also refetches from the beginning instead of resuming, so each attempt costs the full price of the original.
The loop is frequently not in code you wrote. It is a default retry policy inside a client library, or a queue that returns a failed message to the front and hands it straight back. Nobody chose that behaviour deliberately, which is why nobody thinks to check it.
Which of the two it is comes down to the next reading.
Bytes per request
Take the data moved over a period, divide by the requests in that period, and you have an average weight for what you are pulling down.
This should be dull. The same job against the same site should produce a figure that barely shifts from one week to the next. What matters is the movement, not the value, so there is no target number I could usefully give you. A job pulling JSON off an API and a job driving a full browser through a checkout flow live in different universes.
When it drifts upward, two things can have happened.
The target started serving heavier pages. New media on the page, a bigger JSON payload with fields you never asked for, an ad slot that was not there in March. Heavy assets pushing a bill up is a subject of its own, so treat this reading as detection rather than diagnosis.
Or you started fetching things you do not need. That one usually arrives with a change nobody connected to bandwidth: a page moved to rendering in JavaScript, so you switched from plain requests to a headless browser, and now you are pulling every image, font and analytics script on the page instead of one document. Invisible in a code review, unmissable in the bytes per request.
A heavier target moves your bytes per request and leaves the request count alone. A retry loop moves the request count and leaves the bytes per request roughly where it was. That distinction costs you two minutes and saves you an afternoon in the wrong half of the system.
Requests against successes
If your provider exposes it, or you log it yourself, the ratio of requests to successful responses is the closest thing to a smoke alarm in this whole picture. It is also the number people are least likely to log.
Flat total, falling success rate, means you are paying the same money for less data every week. That is usually a site tightening up on you gradually rather than refusing you outright, and the gradual version is much harder to notice and much more expensive to sit through than a clean block would be.
The shape a target can see
Two jobs move the same volume in a month. One spreads it through the working day, the other arrives in a dense block between three and five in the morning. Same total, same bill, completely different traffic.
The uncomfortable part is that the target sees roughly the same picture you do, minus the label that says it is you. They cannot see your account or your invoice. They can see volume against time from an address, and volume against time is precisely what your usage chart draws.
So your own chart is a rough preview of how you look from the outside. It is close enough to be worth looking at. If yours is a wall of traffic starting at the same minute every night and stopping dead an hour later, that is what they see as well. Machines start on the minute. People do not.
I am not going to tell you to spread everything out. Sometimes the block is the right call and the target does not care. But you should know which one you are running, and most people have never checked.
What it will not tell you
Usage data cannot tell you whether the data you collected was worth anything.
A job can run for a week with a flat baseline, a stable bytes per request and a perfect success rate while fetching the same stale error page every single time. The response arrived. It was the right size. It carried the right status code. It contained nothing you wanted. I have watched that happen and nothing in the usage picture so much as twitched, because at the network layer everything really was fine.
It cannot attribute anything either. Aggregate usage carries no per request detail, so the moment you need the exact call that got a session flagged, you are into your logs and the chart is finished helping you.
It also lags. Counting happens in one place, aggregation in another, and the page renders after both. On our side that delay is small, but it is never zero, and a chart that looks calm at nine in the morning is not a statement about half past eight.
And it cannot tell you whether your number is high, because you have nothing to compare it against. Somebody running the same job might move half your volume or four times it, and there is no way to learn that from inside your own account.
The number I put in the wrong place
I built the page these numbers land on, and for over a year the figure in the largest text was the month to date total. That is the least useful number on the page. It tells you what you have spent and nothing about what is happening. The chart underneath, the one you had to scroll for, is the part that helps anybody.
Nobody complained, because people expect a usage page to look like a bill and I had given them one.
What changed my mind was the shape of the support messages. Customers who look weekly ask whether I can see anything odd on a particular port on Tuesday night. Customers who look monthly ask me to explain a number. The first conversation almost always ends with something fixed. The second one ends in an argument about a refund that neither of us enjoyed.
If you want lines where the usage data is yours to read per port, with the shape and not just the total, that is what I run at singaporemobileproxy.com.
Get new guides and videos first — join the Telegram channel.