Skip to content
PD
Парсинг данных

Scraping Products: Prices, Stock and Listings

How to collect product listings, why price is the least reliable field, how to match products across sources, and what to validate in the collected data.

All articles in the guide Парсинг данных · 11

A product listing is the most common object of collection, and it has quirks that regularly make the collected data wrong while the scraper appears to work.

What a listing contains

Identifiers. SKU, barcode, the identifier in the page URL. The most important field: without it, collected data can neither be matched nor updated.

Name and brand. For humans, not for matching.

Price. The most sought-after and least reliable field - see below.

Stock. Often more important than price.

Specifications, usually as a set of pairs whose structure differs per source.

Images and description. Careful: these are copyrighted works, and collecting them for use on your own site is a different matter from collecting figures for analysis.

Why price is unreliable

Five causes, all of them routine:

It depends on region. One product, different cities, different prices. Collect using an address in the target region and label the result - see the proxy guide.

It depends on sign-in. Wholesale prices are visible only to authenticated users.

It depends on quantity. A per-unit price differs from a per-pack price, and either may be the one on the page.

Conditional promotions. “With a loyalty card”, “on orders over”, “until Friday” - on the page that is several numbers, and the scraper must know which one it is taking.

Personal offers. Some platforms show different prices to different visitors.

The practical conclusion: store the price together with its collection conditions - region, sign-in state, date, price type. Without that, a discrepancy with reality cannot be explained, and you will be asked to explain one.

Matching across sources

The main difficulty when collecting from several sites.

Reliable: barcode, manufacturer part number, exact brand and model match.

Unreliable: the name. “Smartphone X 128 GB black” on two sites can be different bundles, model years and warranties.

What to do:

  • Collect every available identifier, not just one.
  • Match on the stable one and fall back to the name last.
  • Record match confidence: exact, probable, manual.
  • Spot check by hand. Twenty items a week show match quality better than any internal metric.

Stock and change

Stock is expressed differently across platforms: an exact number, “low”, “in stock”, “to order”, a lead time. Normalise to your own scale at collection time rather than at analysis time.

Do not delete disappeared products - mark them unavailable. A product may vanish for a week, and the history of items appearing and disappearing is often more valuable than the current snapshot.

Price history is worth more than the current price: movement shows a competitor’s strategy while a snapshot shows a moment.

What to validate

Automatic checks that catch most problems:

  • Price within a plausible range. Zero, negative, or a hundred times yesterday’s value means parsing broke.
  • Mandatory fields populated. An empty name is a failure, not a nameless product.
  • The item count is comparable with last time. A sharp drop means breakage, not a vanished catalogue.
  • No duplicates by key - see deduplication.
  • Correct types. A price arriving as a string with spaces and a currency symbol needs converting, and that belongs at collection time.

The general rule: a scraper should report the suspicious rather than silently write to the database. A thousand items at zero price is an incident, and you should learn about it immediately rather than from a colleague a week later.

How this works technically: scraping in Python. Marketplace specifics: their own article. The overview is in the scraping guide.

FAQ

Why does a collected price differ from what a buyer sees?

Because price often depends on conditions: region, sign-in state, quantity, active promotions and personal offers. A collected value is correct for the conditions it was collected under, and it must be stored together with them.

How do I match products across sites?

By a stable identifier: a barcode or a manufacturer part number. Matching by name inevitably produces errors, because different bundles, volumes and model years look like the same product. Any automatic matching needs spot checks by hand.

What should I do with products that disappear from a catalogue?

Do not delete them, mark them unavailable. A product can vanish temporarily, and the history of items appearing and disappearing is often more valuable than the current snapshot: it shows what is leaving the range and what is arriving.

More on this topic

Done for you

I will build a parser for your source

With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.

from $300 · 3 to 7 days

Similar caseEtsy Keyword FinderA queue-driven keyword audit app for Etsy sellers: submit a listing ID plus up to 20 keywords, and background workers walk the real search results step by step with live screenshots.

"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."

sotasoftdv · KworkTranslated from Russian