Your Scraper Is Lying: Data Quality Gates for Parsing Pipelines
Scrapers rarely crash. They quietly return plausible garbage. Here are the four validation layers I put in every parsing pipeline so bad data gets quarantined instead of shipped.
Pavel Duglas
AI Automation & MVP Architect
A crashed scraper is a good scraper. It tells you something broke, you get a stack trace, you fix a selector, you move on. The dangerous scraper is the one that returns 200 OK, valid JSON, correct field names, and completely wrong values. Nobody notices for three weeks. By then the client has priced products against fantasy competitor data, or the sales team has been calling numbers that were actually zip codes.
I have shipped a lot of parsers - BAS projects, Python pipelines, LLM extraction layers - and the single biggest quality jump came from a boring idea: treat every source as an untrusted API with a contract, and refuse to load data that violates the contract. Below is the exact structure I use.
How scrapers lie
Before building gates you need to know what you are gating against. These are the failure modes I actually see in production, in rough order of frequency.
- Selector drift. A class name changes from
price__valuetoprice-value. Your extractor returnsNone, your downstream code coerces it to0, and now every product is free. - Soft blocks. The site returns HTTP 200 with a challenge page, a “we noticed unusual traffic” screen, or an empty skeleton waiting for JS. Byte count looks reasonable. Your parser finds zero items and reports a successful run with 0 rows.
- Partial render. Headless browser navigated, but the lazy-loaded grid only painted 8 of 48 cards before the extraction fired.
- A/B and geo variants. Same URL, different DOM depending on cookie bucket or exit IP. Your selectors work 70% of the time and you blame flakiness.
- Locale traps.
1.299,00parsed as1.299. Currency symbol changed from EUR to USD because the proxy moved country. Dates flipped between DD/MM and MM/DD. - Pagination truncation. The “next” button changed markup, so you now scrape page 1 of 40 and the row count drop looks like a slow sales week.
- Silent dedupe collapse. Upstream started reusing IDs, your upsert key is now non-unique, and 12,000 rows quietly became 900.
- LLM hallucination. You handed messy HTML to a model, asked for JSON, and got a beautifully structured object where two fields were invented because they were not on the page at all.
Notice that almost none of these throw an exception. Exception-based monitoring is blind to all of them. You need shape-based monitoring.
Layer 1: fingerprint the page before you parse it
The cheapest gate is the first one. Before extraction runs, assert that you are looking at the page you think you are looking at.
For each source I define a small page contract:
PAGE_CONTRACT = {
"must_contain": ["data-product-grid", "Add to cart"],
"must_not_contain": ["unusual traffic", "cf-challenge", "Enable JavaScript"],
"min_bytes": 40_000,
"max_bytes": 2_500_000,
"min_item_nodes": 10,
}
If must_contain markers are missing, this is not a parse failure, it is a fetch failure. That distinction matters operationally: fetch failures should trigger a retry with a different proxy or a session refresh, parse failures should trigger a human looking at selectors. Mixing them is why teams spend a day “fixing the parser” when the real problem was a banned IP range.
Always store the raw HTML for failed pages. Compressed HTML is cheap. Debugging a drift incident without the original page is not. I keep raw payloads for 14 days keyed by run_id + url_hash, in object storage, and nothing else. That single habit has saved me more hours than any clever selector strategy.
In BAS the same logic is trivial: an If block checking for a marker string in page source, before you enter the item loop, plus a counter that records how many items the loop found. If the marker is absent, throw the thread into a resource-rotate branch instead of continuing.
Layer 2: record schemas that reject instead of coerce
Every extracted record goes through a typed model. Not for elegance, for refusal. The rule is: a field either parses cleanly or the record is quarantined. No silent defaults, no or 0, no try/except: pass.
from decimal import Decimal
from pydantic import BaseModel, field_validator, HttpUrl
class Product(BaseModel):
source_id: str
title: str
price: Decimal
currency: str
in_stock: bool
url: HttpUrl
@field_validator("title")
@classmethod
def sane_title(cls, v: str) -> str:
v = " ".join(v.split())
if len(v) < 3 or len(v) > 300:
raise ValueError("title length out of band")
return v
@field_validator("price")
@classmethod
def sane_price(cls, v: Decimal) -> Decimal:
if v <= 0 or v > Decimal("1000000"):
raise ValueError("price out of band")
return v
@field_validator("currency")
@classmethod
def known_currency(cls, v: str) -> str:
if v not in {"USD", "EUR", "VND", "GBP"}:
raise ValueError(f"unexpected currency {v}")
return v
The currency validator is the kind of thing that feels paranoid until the day your proxy pool shifts region and you catch it in one run instead of one quarter.
Quarantine, do not crash. Failed records go to a rejects table with the field name, the error, and the raw fragment. Then a batch-level rule decides whether the run is usable: I typically allow up to 2% rejects, and treat anything above that as a failed run.
Layer 3: batch statistics, the layer everyone skips
Record validation catches malformed data. It cannot catch plausible data that is wrong. For that you compare this run against the trailing history of the same source.
For every run I record a small stats row: item count, null rate per field, median and p95 of numeric fields, count of distinct values for enum-ish fields, share of records that are new versus seen before.
Then the gates:
def batch_gates(stats, baseline):
problems = []
if stats.rows < baseline.median_rows * 0.7:
problems.append("row volume dropped >30% vs trailing median")
if stats.rows > baseline.median_rows * 2.0:
problems.append("row volume doubled, possible dedupe failure")
for field, rate in stats.null_rate.items():
if rate > baseline.null_rate[field] + 0.15:
problems.append(f"null rate spike in {field}")
if abs(stats.price_median - baseline.price_median) / baseline.price_median > 0.25:
problems.append("price median shifted >25%")
if stats.new_record_share > 0.9 and baseline.new_record_share < 0.3:
problems.append("almost everything is new, ID scheme likely changed")
return problems
Use a 14-run trailing median, not the previous run, so one bad night does not become the new normal. When gates fail, the run does not get promoted. The pipeline keeps serving the last good dataset and posts a message to Telegram with the specific gate that tripped. Not “scraper error”, but “row volume down 61% vs median, price median unchanged”. The second message tells you it is pagination, not selectors, before you open your laptop.
Layer 4: golden-set canaries
Statistics tell you the shape changed. They cannot tell you the values are correct. For that, keep a golden set: 15 to 30 URLs per source with hand-verified expected values, refreshed quarterly.
Run them at the start of every job. Compare against expectations with per-field tolerance: exact match on IDs and titles, tolerance bands on prices that legitimately move, boolean match on stock. If more than one canary fails, do not run the full crawl at all. You just saved yourself 40,000 requests of garbage and a proxy bill.
Canaries are also your regression suite when you edit selectors. Change a rule, run canaries, see exactly which fields you broke. This is the closest thing to unit tests that scraping allows.
Where LLM extraction fits
Models are excellent at messy, low-volume, high-variation pages, and terrible at telling you when they are unsure. Three rules make them safe:
- Require evidence. Ask for the extracted value plus the exact source substring it came from. Then verify programmatically that the substring exists in the HTML text. Hallucinated fields fail this instantly, and it costs zero extra tokens on the way back if you keep the spans short.
- Never let the model produce the final type. It returns strings, your code parses to Decimal, date, enum. The schema in Layer 2 stays the authority.
- Escalate only on disagreement. Run the cheap model first. If the schema rejects or evidence checks fail, re-run with a stronger model. If it fails twice, quarantine and let a human look. I usually see under 5% escalation on stable sources, which keeps the cost sane.
And treat page text as hostile input, not instructions. Extraction prompts should be structured so page content is data, never a place where “ignore previous instructions” can change behavior.
Rollout in one working day
If you have an existing parser and no gates, this is the order that gives the most safety per hour:
- Add a
runstable:run_id, source, started_at, rows, rejects, status, notes. Nothing works without run-level bookkeeping. - Add page markers and raw HTML retention. One hour, catches soft blocks immediately.
- Wrap records in a strict schema, send failures to a
rejectstable. - Log per-run stats, and only after you have ten runs of history, turn on the batch gates.
- Build the golden set last, when you know which fields actually break.
The goal is not zero bad data. It is that bad data never reaches the consumer silently, and that when it does appear, the alert names the failure mode. A pipeline that says “I do not trust this run” is worth far more to a client than one that always says success.
FAQ
How do I set the batch thresholds without a history of runs?
Start in observe-only mode. Log run stats for one to two weeks, alert on nothing, then look at the natural variance in row counts and null rates. Set your gates just outside that observed band, usually a 30% drop in volume and a 15 percentage point rise in null rate for a given field. Tightening thresholds later is easy. Starting tight generates alert fatigue and people stop reading the channel.
Should a failed gate stop the whole pipeline or just skip the bad source?
Skip the source, keep the pipeline. Each source should have its own run status and its own last-good dataset. Downstream consumers read the last promoted snapshot, so a broken source means slightly stale data for that source only, not an empty dashboard. The exception is when a source feeds a decision like pricing or outbound messaging - there I prefer to block the decision and show stale-data age explicitly rather than let anyone act on numbers of unknown freshness.
Does this work in Browser Automation Studio, or is it a Python-only pattern?
It works in BAS with almost no changes. Page markers become an If block on page source before the item loop. Record validation becomes a small function that returns a valid or reject flag and writes rejects to a separate file or table. Run stats are just counters written at the end of a thread. The golden set is a short list of URLs you process in a separate function before the main queue. The concepts are storage-agnostic - what matters is that every run produces a status row and every rejected record is kept for inspection.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks