Scraping Competitor Sites: What to Collect and Why
Which competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
All articles in the guide Парсинг данных · 11
Competitor monitoring is the most common scraping task and the most underestimated: collecting the data is easy, making it comparable is not.
What gets collected
Prices. The most sought-after figure, and the most temperamental: price often depends on region, sign-in state, quantity and active promotions.
Stock. Sometimes more important than price: a competitor with nothing in stock is not competing.
Assortment. What appeared, what disappeared. Assortment changes are often more informative than a static snapshot.
Specifications, for matching products across platforms.
Promotions and terms. Delivery, discounts, bundles.
What not to collect: somebody else’s text and photographs for use on your own site. That is not analysis but copyright infringement, and the consequences are not technical.
Building the monitoring
The difference between a one-off collection and ongoing monitoring is fundamental, and the second requires several more things.
Product matching. The main difficulty. You have one SKU, the competitor another, names differ, and the bundle may not be the same. Without reliable matching you compare unlike things and reach wrong conclusions.
Practical anchors: barcode where available, manufacturer part number, exact brand and model match. Matching by name produces errors and needs manual review.
History. Store changes, not the current value. The meaning is in the movement: how the price shifted, when an item went out of stock.
Recording the conditions. A collected price must carry the conditions it was collected under: region, sign-in state, date. Otherwise a month later you cannot explain a discrepancy.
A reaction to empty results. Zero collected items is a failure, not “the competitor closed”. Monitoring must distinguish the two and say so.
A sensible frequency, set by how fast the data changes. Hourly collection where prices move weekly creates load and ban risk with no new information.
Dealing with discrepancies
The practical part that rarely gets written up.
Price depends on region. Collect using addresses in the target regions and label the result - see the proxy guide.
Price depends on sign-in. Wholesale prices are visible only to authenticated users. This is where platform rules are broken explicitly, and the decision is yours, made with the consequences in view.
The product is not the same. Routine when matching by name: different bundles, different volumes, different model years.
The data is stale. A cache on the site side serves yesterday’s price. Detected by comparing against a weekly manual check.
The general rule: any automated monitoring needs periodic manual verification. Ten items checked weekly tell you whether the system works better than any internal check.
The line of acceptable practice
Three things worth respecting beyond mere caution:
Load. Collect at a rate that does not interfere with somebody else’s site. That is both good manners and a way of not attracting attention.
Platform rules. Read them. Most forbid automated collection, and that defines what you are risking.
Content. Collecting facts for analysis is one thing; copying descriptions and photographs is another. The second is not competitor analysis.
How collection works technically: scraping in Python. Dealing with protection: anti-scraping. Where to put the results: Excel. The overview is in the scraping guide.
FAQ
Is scraping competitor sites legal?
Collecting publicly available prices and specifications is not forbidden in itself. Limits appear in three places: site rules forbidding automated collection, copyright in the content, and load on somebody else server that can count as abuse. Collecting facts for analysis is safer than copying content.
What competitor data is collected most often?
Prices, stock, assortment changes and promotions. Price is the most sought-after figure and the least reliable: on many platforms it depends on region and sign-in state, so a collected value may not match what a buyer sees.
How often should price data be collected?
At the rate the data changes, not the rate you would like. If prices change weekly, hourly collection creates load and ban risk without new information. Start daily and increase only if the data genuinely moves faster.
- Web Scraping: How It Actually WorksGuide
- Anti-Scraping Protection: What Works and What Does NotWhich protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
- Web Scraping in Python: Playwright, Retries, DeduplicationHow to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
- Scraping Products: Prices, Stock and ListingsHow to collect product listings, why price is the least reliable field, how to match products across sources, and what to validate in the collected data.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."