← back to blog

Spreading one job across several proxy lines

mobile-proxy scraping architecture attribution

Spreading one job across several proxy lines

The moment a job outgrows one line, you have written an assignment policy whether you meant to or not. Most people write it by accident, in about four characters, by taking the next item from a list.

That default is round robin, and it costs you two things that only become visible later: session continuity, and any ability to say which line caused a failure.

I run the hardware these jobs exit through. One SIM per modem, one port per modem, Singapore carriers, a rack in my house. When a multi-line run goes badly, the code usually arrives in my inbox with it, so I have read a lot of other people’s assignment logic. It is reliably the thinnest layer in an otherwise careful system.

What an average does to a fleet

Five lines. One of them is serving challenge pages from the first hour. The other four are clean for the entire run.

Your dashboard reads eighty percent success.

Eighty percent is a survivable looking number. It reads like a target that tightened up, or a rate limit you could pace around, or a parser that needs another selector. So you tune your delays. Then you add a retry tier. Then you rotate your headers. None of it moves, because four fifths of the fleet never needed any of that, and the remaining fifth was never coming back on its own.

An aggregate converts a specific, cheap, single line problem into a vague, expensive, whole system problem. That conversion happens silently and it happens by default, which is what makes it worth designing against before you have a reason to.

The remedy is one field, carried further downstream than it feels necessary to carry it.

Pacing and backoff have their own trade-offs and I have written about those separately. This piece is about which line carries what, and about being able to answer that question after the run is over.

Decide what stays put first

Round robin spreads evenly, and evenness feels like the goal until you look at what it breaks.

Continuity goes first, immediately. Anything that carries state between calls comes apart when the second half of the conversation arrives from a different address than the first: a logged-in session, a cart, a paginated result set with a cursor, a target that issues a token on request one and checks it on request two.

Attribution goes second, and that one waits until the run is finished before it hurts. When every line touches every unit of work, no result belongs to a line, so the failure is smeared across the pool by design.

So the first decision is what stays bound to one line for the duration. Everything about how you spread the rest falls out of that answer.

The four things you can bind to

The request. A fresh line per request. Correct for genuinely independent fetches against a target that keeps no state about you between calls. That describes fewer jobs than people assume and almost nothing that involves being signed in.

The worker. One line per process or thread, held for the worker’s lifetime. This is convenient because it matches how the code is already shaped. It is also arbitrary: which unit of work lands on which worker is whatever the queue handed out that second, so you get stability with no meaning attached to it.

The target domain. One line per site. This is the right answer for most scraping. A line’s standing with one target is its own property, and keeping targets separate means a site that starts fighting you leaves your other jobs alone. It also lets you pace each target at its own rate without the pacing bleeding across.

The account. One line per identity, held for the whole job and for considerably longer than the job. This is the only one of the four where the binding has to survive the run ending, which is why it belongs somewhere other than your job config.

Most real work wants the target or the account holding the line. Per-request spread is the exception you reach for deliberately.

A map that outlives the run, and a lease that does not

Two layers, answering two different questions.

The map is long-lived. Account or target on one side, line identifier on the other, plus the date the binding was made. It lives outside your code in something a human can open. A job does not pick a line for an account; it looks up the line that account already has, and if there is no row, that is a decision for a person rather than a scheduler.

The lease is short-lived. A worker about to process a unit of work claims the line that unit maps to, holds it for the duration, releases it on completion. The lease exists to stop two workers driving the same line at once, which is how you accidentally double the rate on one address while the rest of the pool idles.

Give every lease an expiry and something that reclaims stale ones. A worker that dies mid-unit releases nothing, and without a reaper that line is out of circulation until a human notices. A ten-line pool quietly becoming a three-line pool over a week, with the job getting slower and nobody able to say why, is a normal outcome of leaving that out.

Practically: a held_by field, a held_since timestamp, and a sweep that frees anything older than your longest plausible unit of work.

The column that turns a rerun into a query

The line identifier travels with the unit of work through everything downstream. Every log line. Every stored row. Every error record. Every metric.

It is one small column on a table you already have, and it is the piece almost nobody has, because at the time you wrote the job you had one line and the question did not exist.

Record two fields, not one. The line identifier says which physical card did the work. The exit address observed at the time says what the target actually saw, and on mobile that changes underneath you on the carrier’s schedule rather than yours. They answer different questions and you will want both.

With those on every row, the five-line weekend above becomes a group-by. Count the rows that failed your content check, split by line, and you have your answer in under a second. Without them you have a rerun and a guess.

Put the identifier on the failures too. Those are the rows most worth grouping by line, and they are the ones people discard.

Per-line health, and quarantine instead of blind retry

Health is a rolling window. A line that had a bad hour on Tuesday is not a bad line on Thursday, and a lifetime counter will keep insisting otherwise long after the problem has passed.

So track outcomes over the last N attempts or the last N minutes, per line. At this layer you are counting one thing: did the response contain what you wanted. A timeout, a refusal and a page full of challenge HTML all collapse into the same bucket for this purpose.

When a line crosses your threshold, remove it from the pool of leasable lines. That is the whole action. The unit of work that was on it goes back on the queue.

The default behaviour most pools fall into is to retry the failed unit and let the scheduler place it. Frequently that is the same line that just failed, so you pay for the same bad response twice and your error rate on that line looks even worse, which is at least honest.

What quarantine means next depends entirely on the binding, and the two cases are opposites.

Target bound work rebalances. Another line picks up the target, the run continues with less capacity, and the only casualty is throughput.

Account bound work pauses. An account that has only ever appeared from one address should not arrive from a new one because an error counter crossed a threshold at three in the morning. Moving an account between lines is a deliberate operation with its own sequence and its own waiting periods, and automating it inside a scraper is how a quiet Tuesday becomes a verification wall on Wednesday. Park the unit, mark the line, and let a person look in the morning.

Writing that distinction into the code takes five minutes and separates a scraper that heals itself from one that quietly damages the accounts it was built to serve.

Changing the pool without stopping the run

Keeping the map and the lease table outside the process is what makes everything above operable.

A flat file, a small table, whatever store you already run. Workers read the state at the top of each unit of work rather than loading it once at startup.

Then adding a line mid-run is a row insert. Pulling a sick line is a state change. Neither needs a restart, and neither loses the position of a job that has been running for eleven hours.

The alternative, where the pool is a list in code loaded at boot, means every change to your fleet costs you a full stop and a restart. So you stop making changes, and you run the whole weekend down a line you already knew was bad.

Where I keyed it wrong

I built this for my own price monitoring and keyed the per-line counters on the endpoint host and port.

That looked obvious at the time. The port is how you reach the line, so the port is the line.

Ports get recycled here. A card comes out of a modem, the port goes back into stock, and a few days later a different SIM on a different carrier answers on the same number. My counters had a fortnight of one card and a week of another sitting under a single key.

The chart showed a line gradually degrading. What it was actually showing was one healthy line and one bad line averaged together, with the switchover buried somewhere in the middle where no reading of the chart would surface it. I spent an afternoon tuning delays against a line that had nothing wrong with it.

So attribute on an identifier you generate once and never reuse. Not the port, not the address, not the friendly name you may rename next month. A value that means one card for as long as you rent it, written onto every row it ever touched.

That single habit is what separates a fleet you can debug from a fleet you can only average.

If you want lines with stable identifiers and per-port counters you can read for yourself, that is what I run at singaporemobileproxy.com.

Get new guides and videos first — join the Telegram channel.

ready to try Singapore mobile proxies?

24-hour free trial. no credit card required.

start free trial
message me on telegram