Collecting public data at scale (prices, catalogs, listings, reviews) quickly runs into two obstacles: websites limit how much a single IP address can request, and what they display depends on the visitor's country. Proxies address both constraints. Used badly, though, they mostly produce blocks, skewed data and legal risk. Here is what makes a scraper reliable and respectful.
Why a scraper needs proxies
A server identifies each visitor first and foremost by their IP address: it is the simplest unit for measuring activity, and therefore for limiting it. Two reasons make proxies useful for data collection:
- Per-IP rate limits. Many websites cap the number of requests an address can make over a given period. Beyond that, they slow down, return a 429 error or block. Spreading requests across several IPs lets each one stay below those thresholds.
- Geolocated content. Prices, currency, availability, language, search results: many pages change with the country of the IP. To collect the version seen in France or Spain, you need an IP from that country.
A proxy does not, however, turn something forbidden into something allowed, and it does not relieve you of keeping the total load on the site under control.
Choosing the right proxy type for your volume and target
The right choice depends on volume, session length and how sensitive the target site is. Our comparison of ISP, datacenter, residential and mobile proxies covers each family in detail. For data collection, keep this in mind:
- Dedicated ISP proxies for steady volumes and stable sessions (a list of sites to monitor, flows that involve cookies or a login): the IP is static and fast, and its history is yours alone.
- Rotating pools for very large, scattered volumes, where each IP only makes a handful of requests. The addresses are shared and their quality varies.
- A mobile IP for highly suspicious targets: since a mobile carrier address is shared by many subscribers, it is hard to block without affecting real users.
| Criterion | Dedicated ISP | Rotating pool | Mobile 4G IP |
|---|---|---|---|
| IP stability | Static | Changes constantly | Fixed until the next rotation |
| Long sessions | Well suited | Limited | Between two rotations |
| Very high volumes | Depends on how many IPs you have | Strong point | A single line |
| Highly suspicious targets | ISP reputation | Variable | Strong point |
| Predictable cost | Per IP, monthly | Often by volume | Monthly |
At Airproxy, the mobile 4G IP is a French line reserved for a single customer. You switch to a new address with a button or through the API, in around twenty seconds, the time it takes the line to reconnect. Trigger rotations between two batches, never in the middle of a session. Our guide to mobile 4G proxies, rotating or sticky IP explains when to change address.
Sizing your setup and spreading the load
Think in requests per IP per minute
The right question is not “how many proxies do I need?” It is “how many requests per minute can one IP send to this site without degrading the service or my results?” There is no universal threshold: it varies with the site, the type of page, the time of day and whether you are logged in. So measure it:
- Start at a low rate, with one IP per site.
- Increase in steps while watching for 429 and 403 errors and response times.
- Settle on a rate well below the point where errors start to appear.
- Divide your volume per minute by that rate, then add a buffer.
With made-up numbers: if you need to fetch 50 pages per minute and one IP can sustain 5 requests per minute, you need 10 IPs, plus a buffer. Keep an eye on the total load as well: what is trivial for a large platform can weigh heavily on a small site.
Rotate IPs cleanly
- One session, one IP. A flow that depends on cookies (login, cart, pagination) keeps the same address from start to finish. Switching midway often breaks the session.
- Rotate between tasks, not between requests, with a rate limiter for each IP and site pair.
- Smooth out your pace with random delays, and schedule big runs during the target site's off-peak hours.
- Do not fetch what you do not need: caching, conditional requests (
If-Modified-Since,ETag) and incremental collection all reduce the load.
The Airproxy customer dashboard exports all your proxies in host:port:username:password format, ready to feed into your scraper, and each proxy can be given an alias (per site or per task) that ties it to your logs.
Headers, fingerprints and blocks
Consistent headers and fingerprint
A website does not look only at your IP: it compares what your HTTP client claims to be with how it actually behaves. Avoid these common inconsistencies:
- a browser
User-Agentsent alongside an HTTP library's default headers; - an
Accept-Languageheader that has nothing to do with the IP's country; - a user agent that changes on every request within the same session;
- a TLS or HTTP/2 fingerprint that gives away a different tool from the one announced.
If the site requires JavaScript, a headless browser renders it more faithfully but puts more strain on the server, so block the resources you do not need (images, fonts, videos).
429, 403, CAPTCHAs: slow down, do not force your way in
- 429 (Too Many Requests): honor the
Retry-Afterheader. If there is none, space out your retries more each time (exponential backoff) and lower the rate for that IP. - 403 (Forbidden): check your headers and the address you used, without retrying in a loop. If the refusal is deliberate, stop.
- CAPTCHA: the site wants to make sure it is dealing with a human. Reduce your pace or look for access designed for automation. Do not bypass it.
- 407, connection refused, timeout: the problem often lies with the proxy (credentials, format, protocol). Check the address with our proxy checker.
Log response codes per IP and per site, and add a circuit breaker that stops the job once the error rate goes above a threshold you set. Switching to another IP right after a block amounts to forcing the door: let that address rest instead.
The legal and ethical framework
Proxies solve a technical problem, not a legal one. Before launching a collection job, check at least the following:
- The
robots.txtfile, which tells you which areas the site owner does not want crawled. Respect it. - The terms of service: some websites prohibit automated collection, especially behind a login, where you have accepted those terms.
- The GDPR and other data protection laws: as soon as you collect personal data (names, profiles, contact details, signed reviews), you need a legal basis, collection limited to what is necessary, a defined retention period, and you must inform the people concerned. Data protection authorities, such as France's CNIL, publish guidance on the subject.
- Rights in content and databases: analyzing is not republishing, and extracting a substantial part of a database can infringe the rights of whoever built it.
- Official APIs: when they exist, they are generally more stable, lighter on the site and safer for you.
Common mistakes
- Sending everything through a single IP, or switching IPs on every request in the middle of a session.
- Ramping up all at once, then retrying in a loop after a 429 or a 403, until a temporary limit turns into a lasting block.
- Launching without testing, then mistaking a proxy error for a block. Our guide on how to test a proxy lists the checks worth running.
- Picking a protocol at random: HTTP(S) and SOCKS5 come with the same Airproxy access. See the differences between HTTP and SOCKS5.
- Letting your proxies expire mid-collection: a rental lasts 30 days. Renew in one click, or turn on automatic renewal, an option that is never ticked by default.
To put this into practice, our guide to using a proxy in Python, Node.js and curl shows how to plug an authenticated proxy into a script. Dedicated ISP proxies in France, Spain and Europe are available on the offers page, with instant delivery.
Frequently asked questions
How many proxies do I need to scrape a website?
It depends on your volume and on the rate a single IP can sustain on that site without errors. Measure that rate in steps, divide your volume per minute by it and add a buffer.
Does a proxy make web scraping legal?
No. It spreads the load and gives you the right location, but legality depends on the data you collect, the site's terms of service and data protection law such as the GDPR.
Dedicated ISP or rotating proxies for scraping?
Dedicated ISP proxies suit regular collection and sessions that rely on cookies. Rotation is mainly useful for very large, scattered volumes, and a mobile IP for the most suspicious targets.
HTTP or SOCKS5 for a scraper?
Scraping libraries generally support HTTP(S) natively, which is all you need for web pages. SOCKS5 is useful if your tool prefers it, and both work on the same Airproxy access.
What should I do when a CAPTCHA appears?
Slow down and check that your headers are consistent. A CAPTCHA means the site wants to limit automation: look for an API or access designed for that purpose, and do not bypass it.
