Anti-Scraping Protection: What Works and What Does Not
Which protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
All articles in the guide Парсинг данных · 11
The view from the other side is useful to both: it shows a site owner what is worth doing, and it explains to anyone collecting data why some platforms are easy and others are not.
Start with a question
What exactly are you protecting, and from whom?
The answers differ, and so do the measures:
- Prices from competitors. They will be checked by hand anyway, so protection only raises the cost of automation.
- Unique content from copying. Legal measures are more effective here than technical ones.
- Users’ personal data. Protection is mandatory, and that is a security question rather than a scraping one.
- The server from load. Then the goal is rate limiting, not bot detection.
The practical point: different goals require different measures, and there is no universal protection.
What does not work
Measures that provide a feeling of protection without providing protection:
Disabling right-click and selection. Bypassed by disabling scripts and irritating for ordinary users.
Obfuscating client code. The data still reaches the browser, so it is still available.
Hiding text with styles. The text remains in the markup.
Checking the user agent. Forged in one line.
Requiring script execution. It stops naive scripts and does not stop browser automation, which is now the norm.
Bot traps - hidden links a human never follows. They catch naive crawlers and risk catching search engines.
What they share: they raise the cost of bypassing by minutes and degrade the site for people and search engines permanently.
What does work
Rate limiting. The basic and most effective measure. It restricts not the fact of collection but its speed, making bulk collection expensive in time.
Authentication for valuable data. Wholesale prices and detailed specifications behind a login. That changes the legal framing: collecting from behind authentication means breaching the terms under which the account was issued.
Behavioural analysis. Navigation speed, missing human-typical actions, paths through the site. It yields probability rather than fact, so it is used alongside other signals.
Watching for anomalies. A sharp rise in requests to particular pages, unusual sequences, a spike from new addresses. Often more useful than any technical measure, because it tells you what is actually happening.
Result caps. A maximum number of records per page and per request. It does not inconvenience users and noticeably raises the cost of bulk collection.
Challenges on suspicion, not for everyone. A captcha for all visitors costs conversion; a captcha on an anomaly is a reasonable measure.
Choosing a level
Proportionality is the main criterion: protection must not cost more than the losses from collection.
A sensible sequence:
- Rate limiting. Costs nothing and solves half the problem.
- Observation. Understand who is collecting what before reacting.
- A soft response. Slow the suspicious visitor rather than block them.
- Authentication for what is genuinely valuable.
- Challenges on suspicion, not for everyone.
What not to do: block on a single signal. A shared address is also an office, a mobile carrier, a provider. Over-strict protection turns away real customers, which costs more than the collected prices.
And the honest conclusion: complete protection does not exist. Data available to a person is available to a program. The realistic goal is making collection cost more than it yields.
How this looks from the collection side: the scraping guide. On captchas and their economics: captcha solving.
FAQ
Can a site be fully protected from scraping?
No. Anything a user can see, a program can see: in the limit it opens the same page in the same browser. The realistic goal is making collection cost more than it yields, not making it impossible.
Which measures are useless?
Disabling right-click, obfuscating client code, hiding text with styles, and checking the user agent. All are bypassed in minutes while degrading the site for real users and search engines.
What actually works against data collection?
Rate limiting, requiring authentication for valuable data, behavioural analysis, and watching for traffic anomalies. What works is the combination plus a sensible response: slow the suspicious visitor rather than block everyone.
- Web Scraping: How It Actually WorksGuide
- Scraping Competitor Sites: What to Collect and WhyWhich competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
- Web Scraping in Python: Playwright, Retries, DeduplicationHow to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
- Scraping Products: Prices, Stock and ListingsHow to collect product listings, why price is the least reliable field, how to match products across sources, and what to validate in the collected data.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."