Web Scraping: How It Actually Works
What scraping is used for, how the browser approach differs from HTTP requests, what breaks scrapers, and where the legal lines are.
All articles in the guide Парсинг данных · 11
Scraping is extracting data from pages built for humans. The task looks simple right up to the first run at volume.
What scraping is used for
The practical jobs behind it:
- Monitoring prices and stock at competitors and suppliers.
- Collecting catalogues to populate your own.
- Market analysis: assortment, positioning, changes.
- Feeding internal systems with external data.
- Checking your own data on platforms: how your products appear.
What they share: the data exists but is not offered in a convenient form. If the source has an API, scraping is unnecessary - working with an API is cheaper and more reliable in every respect.
Browser versus HTTP
The fork that determines both cost and reliability.
HTTP requests. You fetch the page and parse the response.
Advantages: fast, cheap in resources, easy to parallelise. Thousands of pages an hour on an ordinary server.
Disadvantages: you see only what the server returned immediately. If content is drawn by scripts, it will not be there.
A browser. The page genuinely opens, scripts run, and you work with the final result.
Advantages: you see everything a person sees. Clicks, scrolling and authentication all work.
Disadvantages: tens of times slower and heavier in memory. A hundred concurrent browsers is a serious server.
The practical rule: first check whether the data is in the raw response. If it is, no browser is needed and the solution costs a fraction. Often the data arrives as a structured payload in a background request, and calling that directly is enough.
Tooling for both approaches is covered separately: scraping in Python and a survey of tools.
What breaks scrapers
Four causes, each needing a different fix.
The markup changed. The most common. The selector points at an element that no longer exists. Fixed with resilient selectors and an explicit error instead of an empty result: a scraper that silently collected zero records is worse than one that crashed.
Anti-scraping protection. The address fell under restrictions and a challenge appeared. Fixed by changing address type and slowing down - see the proxy guide and captcha solving.
Volume. It worked on a hundred pages and stalled at ten thousand. Fixed with batching, a queue and saved progress.
The data is wrong. A field format changed, a value became a different type, duplicates appeared. Fixed by validating output rather than trusting the source.
Separately: a source may serve different content to different visitors. Prices and availability often depend on region, sign-in state and history. Collected data can be correct and still not match what your customer sees.
The legal lines
Three different things that get conflated.
Site rules. Most platforms forbid automated collection in their terms. That is contractual: the consequence is losing access, and at scale, claims.
Personal data. This is where law begins. Collecting individuals’ contact details is regulated, and “the data was public” is not sufficient grounds to process it - see scraping contacts.
Copyright. Text, photographs and descriptions are protected works. Collecting them for analysis and republishing them are legally different acts.
The practical stance: facts and figures are safer than content; aggregating is safer than copying; internal use is safer than publishing. And if the task is commercial and large, settle the question with a lawyer before development rather than after the first complaint.
Where to go next
- Scraping competitor sites - the most common task.
- Scraping in Python - selectors, retries, deduplication.
- Scraping products and tables.
- Exporting to Excel.
- Scraping tools - if you would rather not write code.
- Anti-scraping protection - the view from the other side.
- What scraping costs - how the price is formed.
- Scraping marketplaces - a case with its own quirks.
- Scraping contacts - with the legal caveat it requires.
If you want the data rather than the skill, see the services page.
In this guide
- Scraping Competitor Sites: What to Collect and WhyWhich competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
- Anti-Scraping Protection: What Works and What Does NotWhich protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
- Web Scraping in Python: Playwright, Retries, DeduplicationHow to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
- Scraping Products: Prices, Stock and ListingsHow to collect product listings, why price is the least reliable field, how to match products across sources, and what to validate in the collected data.
- Scraping a Table From a Website: Three ApproachesHow to get a table off a page: by hand, with ready tools, and in code. What to do about merged cells, pagination, and tables loaded by scripts.
- What Web Scraping Costs: How the Price Is FormedWhat drives the cost of data collection, why superficially identical tasks differ several-fold, what the price consists of, and what to ask before commissioning work.
- Scraping Contacts From Websites: What Is LawfulWhere the line runs between collecting open company data and processing personal data, what the law requires, which tasks are lawful, and why mailing a scraped list does not work.
- Scraping to Excel: Exporting the ResultsWhich format to deliver collected data in, why a spreadsheet is not always the right choice, how to avoid the usual encoding and number problems, and how to arrange regular updates.
- Web Scraping Tools: A SurveyThe categories of data collection tooling, what ready-made programs cover, where they hit a ceiling, and how to choose one for your task.
- Scraping Marketplaces: Wildberries and OzonHow marketplaces differ from ordinary sites for data collection, what official seller interfaces provide, which tasks sellers actually have, and where the limits are.
FAQ
Is scraping legal?
Collecting publicly available data is not forbidden in itself, but there are three lines: site rules, personal data, and copyright in the content. The first is contractual; the second and third are legal, and both belong before the work starts rather than after.
How does browser scraping differ from HTTP requests?
An HTTP request takes the raw server response: fast and cheap, but blind to anything drawn by scripts. A browser opens the whole page and sees the final result, at tens of times the cost in speed and resources. The choice follows where the data actually lives.
Why did my scraper stop working?
Usually the markup changed and the selector points at an element that no longer exists. The next most common case is anti-scraping protection: the address fell under restrictions and a challenge arrives instead of data. Telling them apart is easy - look at what the site actually returned.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."