Skip to content
PD
AI Automation 8 min read

Scrape Like You'll Be Audited: Building a Permission Layer Into Your Parsers

Most scrapers have no concept of permission. I show the policy file, fetch gate, rate budgets and provenance logging I add to every parsing pipeline so it survives a complaint, a block or a client audit.

PD

Pavel Duglas

AI Automation & MVP Architect

Every scraper I have ever inherited had the same missing piece. It knew how to fetch a page and it knew what to extract. It had no idea whether it was allowed to be there. Permission lived in the developer’s head, or in a Slack message from six months ago, or nowhere at all.

That was survivable when scraping was a niche activity. It is not survivable now. Site owners have gotten aggressive about blocking automated collection, and the tooling on their side has improved faster than the tooling on ours. A single sloppy crawler can get a client’s IP range blocked, poison a data product with content the client has no right to redistribute, or trigger an email from someone’s legal department that lands on your desk on a Friday evening.

So I now build a permission layer into every parsing pipeline. It is not a legal opinion generator. It is an engineering artifact that makes the answer to “why did you fetch this page, at this rate, and what did you keep?” a query instead of an archaeology project.

The three failure modes that actually hurt

Before the code, be clear on what you are defending against. In my experience it is never abstract “ethics.” It is three concrete events.

You get blocked mid-project. Your scraper hammers a site, the site adds a WAF rule, and your pipeline silently starts collecting error pages. If you have no rate budget and no per-domain health signal, you find out weeks later when a client asks why the numbers went flat.

You collect something you are not allowed to reuse. Public and reusable are different things. Scraping a page to compute an aggregate is a different act than republishing its text, and different again from feeding it into a training set. Pipelines that do not record what kind of use the data was collected for cannot answer questions later.

Somebody complains and you have no record. The expensive part of a complaint is not the complaint. It is discovering you cannot prove what you did. No fetch logs, no policy history, no way to show you honored the site’s stated preferences.

All three are architecture problems. Let’s fix them.

Rule 1: one policy file per project, not scattered if statements

Permission rules end up spread across cron args, a HEADERS dict, a hardcoded time.sleep(2) and a comment. Pull all of it into one declarative file that a non-developer can read.

# policy/domains.yml
defaults:
  respect_robots: true
  max_rps: 0.5
  max_pages_per_day: 2000
  user_agent: "PavelDuglasBot/1.0 (+https://example.com/bot; ops@example.com)"
  allowed_use: ["aggregate_metrics"]
  retain_raw_html_days: 14

domains:
  shop.example.com:
    source: html
    max_rps: 0.3
    allowed_use: ["aggregate_metrics", "internal_dashboard"]
    notes: "ToS allows personal use only. No republishing of product copy."
    reviewed_at: "2026-01-14"
    reviewed_by: "pavel"

  partner-api.example.net:
    source: api
    api_key_env: PARTNER_API_KEY
    max_rps: 5
    allowed_use: ["aggregate_metrics", "client_delivery"]
    contract: "docs/contracts/partner-2026.pdf"

Three things this buys you immediately. First, adding a new domain becomes a deliberate act with a reviewer’s name on it, not an accident of a wildcard crawl. Second, allowed_use gives your downstream code something to check against. Third, when a client asks “where does this data come from and can we resell it?”, the answer is a file in git with history.

Rule 2: a fetch gate that everything goes through

No module in the pipeline calls requests.get or opens a browser tab directly. Everything goes through one function that can say no.

def fetch(url, *, use_case, ctx):
    domain = urlparse(url).netloc
    policy = ctx.policy.for_domain(domain)

    if policy is None:
        raise PolicyError(f"{domain} is not in policy/domains.yml")

    if use_case not in policy.allowed_use:
        raise PolicyError(f"{use_case} not permitted for {domain}")

    if policy.respect_robots and not ctx.robots.allowed(url, policy.user_agent):
        ctx.log_skip(url, reason="robots_disallow")
        return None

    if not ctx.budget.consume(domain):
        ctx.log_skip(url, reason="budget_exhausted")
        return None

    ctx.limiter.wait(domain, policy.max_rps)
    resp = ctx.session.get(url, headers={"User-Agent": policy.user_agent}, timeout=30)

    if resp.status_code in (429, 503):
        ctx.budget.cool_down(domain, retry_after(resp, default=900))
        ctx.log_skip(url, reason=f"backoff_{resp.status_code}")
        return None

    ctx.log_fetch(url, resp.status_code, len(resp.content), use_case)
    return resp

Unknown domain means hard failure, not a default-allow. That single decision has saved me more grief than everything else on this list. Crawlers drift. A relative link on page 40 sends you to a subdomain nobody reviewed, and suddenly you are collecting from a site with completely different terms.

In Browser Automation Studio the same idea applies, it just lives in a function instead of a Python module. I keep a PolicyCheck function that takes the URL and the use case, reads a JSON policy resource, and either returns the go-ahead or fails the thread with a labeled reason. Every navigation block calls it first. It is boring. That is the point.

Rule 3: identify yourself honestly

I know the temptation. A real Chrome user agent gets through where a bot string gets a 403. And yes, for anti-bot work on targets that expect browser traffic, you are running a real browser with a real fingerprint anyway.

But there is a meaningful line between looking like a normal browser and actively lying about who you are to a site that has asked you not to be there. On the second side of that line, you have no defense at all if it comes up later. My rule: if a site publishes contact-friendly signals (a robots file with crawl directives, an API, a stated data policy), I identify the bot with a URL and an email. If someone wants to throttle me or ask me to stop, I want them to be able to reach me before they reach for legal.

The practical bonus is that identified bots get whitelisted surprisingly often. I have had two site owners raise my rate limit after a single email, which was cheaper than three weeks of proxy rotation.

Rule 4: rate budgets belong to domains, not to jobs

The classic bug: five separate jobs each politely sleep two seconds between requests, and the target sees ten concurrent crawlers. Rate limiting has to be enforced at the domain level, shared across every process.

Redis is enough. One key per domain for the token bucket, one for a daily page counter, one for a cool-down timestamp. When a domain returns 429 or 503, all workers see the cool-down and back off together. When the daily budget is spent, the pipeline stops rather than degrading into an error-page harvester.

I also alert on quiet failures. A domain whose success rate drops below 80% over an hour is either blocking you or has changed layout. Both need a human. Neither should be discovered by a client.

Rule 5: store provenance next to every row

Data without provenance is a liability that looks like an asset. Every extracted record in my pipelines carries:

{
  "source_url": "https://shop.example.com/p/1042",
  "domain": "shop.example.com",
  "fetched_at": "2026-01-20T09:14:02Z",
  "http_status": 200,
  "content_hash": "sha256:9f2c...",
  "policy_version": "domains.yml@a71c3f2",
  "use_case": "aggregate_metrics",
  "parser_version": "shop_v4"
}

policy_version is the one people skip and the one that matters most. When you tighten a rule in March, you need to know which rows were collected under the old rule. Without it, a policy change forces you to either re-scrape everything or hope nobody asks.

A content hash also gives you cheap change detection and lets you drop raw HTML on a retention schedule while keeping the ability to prove what you parsed.

Rule 6: minimize at extraction, not at export

Most pipelines hoover up full HTML, dump it in S3 forever, and filter at the reporting layer. That means your storage bucket contains every email address, phone number and user comment on every page you ever touched.

Extract the fields you actually need at parse time. Set a retention window on raw HTML (14 to 30 days covers almost every debugging need). If a page contains personal data you did not come for, drop it in the parser with a regex pass, and log that you dropped it. The smaller your data footprint, the shorter every uncomfortable conversation.

The escalation ladder

Before writing a parser at all, walk down this list and stop at the first thing that works:

  1. Official API or paid data feed. Slower to set up, immune to layout changes, contractually clear.
  2. Sitemaps, RSS, JSON endpoints the site’s own frontend calls. Structured and cheap.
  3. Plain HTTP HTML fetching with a policy gate.
  4. Headless browser with real rendering.
  5. Full anti-detect browser work with fingerprinting and proxies.

Each step down costs more to build, more to maintain, and carries more risk. I have watched teams start at step 5 because it is the fun part, when step 2 would have delivered the same dataset in an afternoon. Check for a JSON endpoint first. Always.

A 90-minute retrofit for an existing scraper

You do not need a rewrite. In order:

  1. Grep for every direct HTTP call and browser navigation. List the domains. That list is your first domains.yml.
  2. Write the fetch gate. Route the calls through it. Unknown domain raises.
  3. Move rate limits to Redis, keyed by domain.
  4. Add the provenance fields to your output schema. Backfill what you can, mark the rest unknown.
  5. Add a retention job for raw HTML.
  6. Put a real contact address in your user agent for the domains where that makes sense.

Six steps, and your pipeline stops being a black box. The next time a client asks whether they can resell the dataset, or a site owner asks why you are hitting them 40 times a minute, you answer with a query and a git log instead of a shrug.

Data collection is a real engineering discipline with real obligations. Build the permission layer while the project is small. Retrofitting it after a complaint is the same work done at three times the price and under stress.

FAQ

Does honoring robots.txt mean I can never scrape anything useful?

No. Plenty of sites disallow specific paths (search results, checkout, admin) while leaving product or listing pages open, and many disallow only named bots. The point of putting robots handling behind a policy flag is that it becomes a per-domain decision you make consciously and record, rather than a blanket yes or no. For some targets you will decide the directive does not apply to your use case, and that decision now lives in a reviewed file with a name attached to it.

What is the difference between allowed_use values and why does it matter?

Because the same page can be legitimate to fetch for one purpose and problematic for another. Computing an average price across 500 listings is a different act from republishing the seller's product description on your own site, and both differ from putting the text into a training corpus. Tagging each fetch with a use case lets the pipeline block mismatches at the gate, and lets you answer later questions about a specific dataset without guessing.

I use proxies and anti-detect browsers. Does a permission layer even apply to me?

It applies more. Anti-detect work removes the natural feedback loop where a site blocks you and you notice, so it is very easy to keep hammering a target that has clearly signaled it wants you gone. The policy file, shared rate budgets and provenance logging are exactly what give you back the visibility that the proxy layer hides. Fingerprint work is a technical tool for looking like a normal browser, not a substitute for deciding which targets you are willing to collect from.

Related articles