Web Scraping in Python: Playwright, Retries, Deduplication
How to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
All articles in the guide Парсинг данных · 11
The difference between a script that works once and a scraper that runs for months is not the libraries. It is how it handles everything that does not go to plan.
Choosing the stack
A check order that saves weeks:
1. Is there an API? Sometimes an undocumented one: look at the background requests the page makes. Often the data arrives as a structured payload and there is no markup to parse at all. The best possible outcome.
2. Is the data in the raw response? Fetch the page with an ordinary HTTP request and see whether what you need is there. If it is, use an HTTP client plus a markup parser. Fast, cheap, easy to parallelise.
3. Only if the data is absent should you drive a real browser. Tens of times more expensive in resources and time.
In practice: a browser is reached for more often than it is needed. The check takes five minutes and often saves most of the work.
Selectors that survive markup changes
Markup changes, and that is the main cause of breakage.
Anchor on meaning, not position. Attributes with meaningful names and text labels outlive long tree paths and auto-generated class names.
Do not depend on ordering. The third block in a list stops being third at the first change.
Search from a stable anchor. Find the product card container by a meaningful marker and search inside it.
Verify what you found. An empty result must be an explicit error, not an empty cell in the export. A scraper that silently collected a thousand empty records is worse than one that crashed: you learn about a crash immediately.
A useful habit: a separate structure check. A small set of reference pages and an assertion that every mandatory field is found. Run it before the main collection and it catches markup changes before you collect noise.
Retries
Networks are unreliable, sites return errors, connections drop. Without retries a scraper will not survive an hour.
What matters:
Distinguish error types. Timeouts and transient server errors are retried. A “page not found” is not: a retry gives the same result.
Increasing delay. Retrying immediately means finishing off a site that is already struggling.
An attempt cap. Otherwise one problematic URL consumes all the time.
Retry from a different address if you work through proxies: the same address will most likely be refused again - see rotation.
Separate handling for blocks. If a challenge arrived instead of data, retrying is pointless - the collection conditions have to change.
Deduplication and saved progress
Two things that look optional and become critical at volume.
Deduplication. The same record arrives several times: pagination overlaps, retries repeat pages, a resumed run starts mid-way.
You need a stable key: a SKU, an identifier from the page URL, a barcode. Name and price cannot be keys, because they change. An existence check on that key before writing solves the problem entirely.
Saved progress. Collecting ten thousand pages will be interrupted at some point: network, error, restart. Without saved progress you start over.
In practice: mark processed pages so a resumed run continues where it stopped. That also lets you collect in batches instead of one long run.
What else to do immediately
- Limit the pace. A pause between requests is both courtesy and a way to avoid restrictions.
- Log the URL and result for every page. Diagnosis is impossible without it.
- Validate output data. A price as a string, a negative quantity, a future date - all signs that parsing broke.
- Save the raw response on a parse error. A day later you will not reproduce the page that failed.
Exporting results: to Excel. No-code options: scraping tools. The overview is in the scraping guide.
FAQ
What should I use for scraping in Python?
For pages where the data is in the raw response, ordinary HTTP requests with a markup parser: fast and cheap. For pages where content is drawn by scripts, driving a real browser. Check in that order: a browser is needed less often than people assume.
Why does my scraper run and collect empty data?
Because the selector found nothing and the code treated that as a missing value. A scraper must distinguish "the field is not on the page" from "the field is empty": the first is a failure needing attention, the second is valid data. Without that you get thousands of empty rows and learn about it late.
Why is deduplication necessary?
Because the same record arrives several times: pagination overlaps, retries repeat pages, a resumed run restarts mid-way. Without a stable key and an existence check you accumulate duplicates that later have to be cleaned by hand.
- Web Scraping: How It Actually WorksGuide
- Scraping Competitor Sites: What to Collect and WhyWhich competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
- Anti-Scraping Protection: What Works and What Does NotWhich protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
- Scraping Products: Prices, Stock and ListingsHow to collect product listings, why price is the least reliable field, how to match products across sources, and what to validate in the collected data.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."