Scraping Marketplaces: Wildberries and Ozon
How marketplaces differ from ordinary sites for data collection, what official seller interfaces provide, which tasks sellers actually have, and where the limits are.
All articles in the guide Парсинг данных · 11
Marketplaces are the most common collection target among Russian platforms, and they have quirks that make a naive approach produce wrong data.
How they differ
Personalisation. Price and search results depend on region, browsing history, device and subscription status. A collected value is correct for its collection conditions and need not match what any particular buyer sees.
Dynamic pricing. Prices change often, sometimes within a day. A daily snapshot may not reflect reality.
Search position is not fixed. The same product for the same query sits at different places for different users. Any position measurement here is approximate and should be treated as an estimate.
Serious protection. Platforms of this scale invest in anti-scraping measures, and getting past them is an ongoing race rather than a one-time setup.
The rules explicitly forbid automated collection. That is part of the user agreement, and the consequence is losing access - for a seller, with additional risk to their account.
What is available officially
Before building a scraper, check whether the task is solved through supported channels.
Platforms provide seller interfaces: your own products, orders, stock, financial reports, some analytics. Anything concerning your own data is better taken from there:
- The data is official and complete.
- No rules are broken.
- It does not break when markup changes.
- There is support.
The practical conclusion: scraping marketplaces is justified where the official interface provides nothing - mostly competitor data and the general picture of a niche.
What sellers actually need
Competitor price monitoring. The most common task - see scraping competitors.
Position tracking. Where a product sits for key queries, with the caveat that the figure is approximate.
Niche analysis before launch. How many sellers, what the price spread is, how listings are presented.
Collecting reviews. Your own, to work on quality; competitors’, to understand complaints about the category.
Checking your own listings. What the listing actually looks like: whether an image was replaced, whether the description disappeared, what price the buyer is shown. This often matters more than everything else, because a seller sees their listing in the seller portal while a buyer sees it in search, and the two differ.
Technical notes
Data often arrives as a structured payload. The page calls internal endpoints and receives ready structures. No markup parsing needed - see the scraping guide.
Region is mandatory. Without specifying one you collect data from an unknown city. Region is set both by the exit address and by request parameters - see the proxy guide.
Volume is large. A category with tens of thousands of items requires a queue, resumption and a sensible pace.
Structures change. Platforms revise internal formats regularly. The scraper needs maintenance - see cost.
Pace matters. Aggressive collection triggers restrictions quickly.
For ongoing platform work, some tasks fit well in Browser Automation Studio, where collection sits alongside authentication, proxies and concurrency in one tool.
The limits
Three things to keep in mind:
Platform rules. Automated collection is forbidden. For a seller that is an additional risk: complaints can reach their working account.
Personal data. Reviews contain names. Collecting reviews to analyse complaints is one thing; building a database of their authors is another, and the second is regulated - see scraping contacts.
Listing content. Competitors’ photographs and descriptions are copyrighted works.
The sensible stance: take your own data through official interfaces, take competitors’ data only in the volume analysis requires, and do not move their content onto your own listings.
What this looks like on real projects: the portfolio. The overview is in the scraping guide.
FAQ
Is there an official way to get marketplace data?
Yes, the platforms provide seller interfaces: your own products, orders, stock, reports. Anything concerning your own data is better taken from there - more reliable and within the rules. Scraping remains for what the official interfaces do not expose.
Why do sellers collect marketplace data?
Competitor price monitoring, tracking search positions, analysing a niche before launching a product, collecting reviews, and checking their own listings. The last often matters most: a listing can look different from what the seller assumes.
Why does collected data differ from what a buyer sees?
Because results and prices are personalised by region, history and device. A product search position is not a fixed value at all: it differs between users, so any measurement here is an approximation.
- Web Scraping: How It Actually WorksGuide
- Scraping Competitor Sites: What to Collect and WhyWhich competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
- Anti-Scraping Protection: What Works and What Does NotWhich protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
- Web Scraping in Python: Playwright, Retries, DeduplicationHow to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."