Price Monitoring With Static Residential Proxies
· 3 min read
Price monitoring consists of checking the same product pages day after day and trusting what comes back, and it is the trusting that proves difficult. A retailer suspecting a bot will often refrain from blocking it and serve a different page instead: an outdated price, an "out of stock" notice, or the price stripped of its promotion. The scraper carries on undisturbed while the data quietly goes wrong.
Avoiding that comes down to looking like a regular customer, from the same place, every single time, which a small pool of static residential IPs achieves far more reliably than a large rotating one.
Why static residential fits
- It reads as a household, so the page you receive is the one a real customer receives.
- The address does not change, so today's price and yesterday's were observed from the same vantage point and can be compared.
- It is billed per IP, so heavy pages and headless browsers leave the bill untouched.
- It is fast, which counts for flash sales and stock that moves quickly.
Background on the proxy type itself can be found in what is a static residential proxy.
Sizing the pool
Begin by deciding how many requests a single IP may send to a single site per day while still passing for a keen shopper, then divide your daily request count by that budget.
requests per day = products x checks per day
IPs needed = requests per day / daily budget per IP
4,000 products x 6 checks = 24,000 requests per day
24,000 / 1,500 per IP = 16 IPs For large retailers, somewhere between a few hundred and a couple of thousand requests per IP per day makes a reasonable starting point. Start at the low end, watch for blocks and odd data, and raise the figure gradually.
Each IP can be made to go further in a few ways:
- Check volatile products frequently and stable ones once a day, bearing in mind that most catalogues are largely stable.
- Rely on the structured data embedded in the page, or the JSON the page itself loads, in place of rendering everything.
- Spread requests across the day as opposed to sending them in a single burst.
Keep each IP consistent
- Assign each IP to one site, or a small group of sites, and leave it there.
- Keep cookies per IP per site and carry them over between runs.
- Hold the browser fingerprint, language and timezone steady and in line with the IP's location.
- When an IP gets challenged, slow down, as retrying harder only makes matters worse.
Catching bad data
- Cross-check a sample. Fetch a handful of products from two IPs, and if the results disagree, one of them is being served a different page.
- Watch the data, not the status code. A jump in "out of stock", missing promotions or identical prices across a category is a block that happened to return 200.
- Keep control products. Track a few prices you are able to confirm elsewhere and raise an alert whenever the scraper disagrees.
- Save the raw page every time a check fails.
When to use something else
- Many countries. Local prices call for local IPs, and ours are in Ashburn, Virginia, which provides a US East view and no other.
- Millions of pages a day. At that volume a rotating network is the practical choice.
The differences are covered in static vs rotating vs datacenter proxies.
On Leastslow
Static residential IPs on a US consumer ISP in Ashburn, with no fees on bandwidth or connections. The pool sized above costs the same every month however heavy the pages turn out to be, and volume pricing applies to the order as a whole. See the price list.