Skip to content
PD
Парсинг данных Guide

Web Scraping: How It Actually Works

What scraping is used for, how the browser approach differs from HTTP requests, what breaks scrapers, and where the legal lines are.

All articles in the guide Парсинг данных · 11

Scraping is extracting data from pages built for humans. The task looks simple right up to the first run at volume.

What scraping is used for

The practical jobs behind it:

  • Monitoring prices and stock at competitors and suppliers.
  • Collecting catalogues to populate your own.
  • Market analysis: assortment, positioning, changes.
  • Feeding internal systems with external data.
  • Checking your own data on platforms: how your products appear.

What they share: the data exists but is not offered in a convenient form. If the source has an API, scraping is unnecessary - working with an API is cheaper and more reliable in every respect.

Browser versus HTTP

The fork that determines both cost and reliability.

HTTP requests. You fetch the page and parse the response.

Advantages: fast, cheap in resources, easy to parallelise. Thousands of pages an hour on an ordinary server.

Disadvantages: you see only what the server returned immediately. If content is drawn by scripts, it will not be there.

A browser. The page genuinely opens, scripts run, and you work with the final result.

Advantages: you see everything a person sees. Clicks, scrolling and authentication all work.

Disadvantages: tens of times slower and heavier in memory. A hundred concurrent browsers is a serious server.

The practical rule: first check whether the data is in the raw response. If it is, no browser is needed and the solution costs a fraction. Often the data arrives as a structured payload in a background request, and calling that directly is enough.

Tooling for both approaches is covered separately: scraping in Python and a survey of tools.

What breaks scrapers

Four causes, each needing a different fix.

The markup changed. The most common. The selector points at an element that no longer exists. Fixed with resilient selectors and an explicit error instead of an empty result: a scraper that silently collected zero records is worse than one that crashed.

Anti-scraping protection. The address fell under restrictions and a challenge appeared. Fixed by changing address type and slowing down - see the proxy guide and captcha solving.

Volume. It worked on a hundred pages and stalled at ten thousand. Fixed with batching, a queue and saved progress.

The data is wrong. A field format changed, a value became a different type, duplicates appeared. Fixed by validating output rather than trusting the source.

Separately: a source may serve different content to different visitors. Prices and availability often depend on region, sign-in state and history. Collected data can be correct and still not match what your customer sees.

Three different things that get conflated.

Site rules. Most platforms forbid automated collection in their terms. That is contractual: the consequence is losing access, and at scale, claims.

Personal data. This is where law begins. Collecting individuals’ contact details is regulated, and “the data was public” is not sufficient grounds to process it - see scraping contacts.

Copyright. Text, photographs and descriptions are protected works. Collecting them for analysis and republishing them are legally different acts.

The practical stance: facts and figures are safer than content; aggregating is safer than copying; internal use is safer than publishing. And if the task is commercial and large, settle the question with a lawyer before development rather than after the first complaint.

Where to go next

If you want the data rather than the skill, see the services page.

In this guide

FAQ

Is scraping legal?

Collecting publicly available data is not forbidden in itself, but there are three lines: site rules, personal data, and copyright in the content. The first is contractual; the second and third are legal, and both belong before the work starts rather than after.

How does browser scraping differ from HTTP requests?

An HTTP request takes the raw server response: fast and cheap, but blind to anything drawn by scripts. A browser opens the whole page and sees the final result, at tens of times the cost in speed and resources. The choice follows where the data actually lives.

Why did my scraper stop working?

Usually the markup changed and the selector points at an element that no longer exists. The next most common case is anti-scraping protection: the address fell under restrictions and a challenge arrives instead of data. Telling them apart is easy - look at what the site actually returned.

Done for you

I will build a parser for your source

With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.

from $300 · 3 to 7 days

Similar caseEtsy Keyword FinderA queue-driven keyword audit app for Etsy sellers: submit a listing ID plus up to 20 keywords, and background workers walk the real search results step by step with live screenshots.

"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."

sotasoftdv · KworkTranslated from Russian