What Web Scraping Costs: How the Price Is Formed
What drives the cost of data collection, why superficially identical tasks differ several-fold, what the price consists of, and what to ask before commissioning work.
All articles in the guide Парсинг данных · 11
Estimates for what looks like the same task differ several-fold, and that is usually not greed but a different understanding of the scope.
What drives the cost
In descending order of impact:
Anti-scraping protection. The main factor. A site with no protection and a site with challenges are different tasks. The second brings proxies, captcha solving and ongoing costs the first does not have at all.
Browser or HTTP. If the data is in the raw response, the solution is several times simpler and cheaper. If a real browser is required, both development and running costs rise - see the scraping guide.
One-off or ongoing. A one-off run has to work once. An ongoing one has to survive markup changes, failures and growing volume. That is different work, not the same work with a markup.
Volume. A thousand pages and a million are different engineering problems: queues, resumption and parallelism appear.
Number of sources. Every site has its own selectors and quirks. And if the data must be reconciled, matching gets added, which is often more expensive than the collection - see scraping products.
Data requirements. “Collect as is” and “collect, normalise, match and validate” are different tasks, and the second is usually the larger.
Authentication. Collecting from behind a login is technically harder and legally riskier.
What the price consists of
Development. One-off, driven by the factors above.
Infrastructure. Server, proxies, captcha solving, storage. Monthly, and on protected sites this can rival development.
Maintenance. Sites change and the scraper needs fixing. That is not a defect in the work but a property of the task: every scraper eventually breaks because its source changed.
Data processing. Normalisation, matching, validation. Routinely underestimated and often the largest part of the work.
Why estimates differ
Five reasons two quotes for one task differ threefold:
- Different assumptions about reliability. A script that runs once and a scraper that runs for six months are different things.
- Whether maintenance is included. Some quotes cover development only.
- Whether infrastructure is included. Proxies and captcha may be in the price or may become your problem.
- Whether data processing is included. “Delivered as is” is cheaper than “delivered ready to use”.
- Whether the site was examined. An estimate without looking at the source is guesswork: protection is established in five minutes and changes the price by a multiple.
The practical conclusion: compare scope, not numbers. A cheap quote often means a one-off script with no maintenance, and a month later you commission the work again.
What to ask before commissioning
- What happens when the markup changes? Is fixing it included, and for how long.
- Who pays for proxies and captcha? And roughly how much per month.
- In what form is the data delivered? A file, a database, an API, a scheduled export.
- What about duplicates and quality checks? Or is that your job.
- What if the site blocks access? “We will figure something out” means the risk is yours.
- Who owns the code? If the contractor leaves, do you still have a working solution.
- Was the source examined? An estimate without that is unreliable.
And a question to ask yourself: does the source have an API or an official export? If it does, the task is solved more cheaply and reliably with no scraping at all.
To discuss a specific task and get an estimate, see the services page. The overview is in the scraping guide.
FAQ
What does scraping a site cost?
The range is wide, and it is set not by page count but by three things: whether the site has anti-scraping protection, whether a browser is required or HTTP requests suffice, and whether this is one-off or ongoing. Superficially identical tasks differ several-fold for exactly those reasons.
Why is ongoing collection more expensive than a one-off?
Because a one-off has to work once while an ongoing process has to survive markup changes, blocks, network failures and growing volume. Maintenance appears too: sites change and the scraper needs fixing. That is different work, not a surcharge.
What should I ask a contractor before commissioning?
What happens when the markup changes, who pays for proxies and captcha solving, whether maintenance is included, in what form the data is delivered, and what happens if the site blocks access. Those answers explain estimate differences better than the numbers do.
- Web Scraping: How It Actually WorksGuide
- Scraping Competitor Sites: What to Collect and WhyWhich competitor data is worth collecting, how to build ongoing monitoring, what to do about discrepancies, and where the line of acceptable practice runs.
- Anti-Scraping Protection: What Works and What Does NotWhich protections against automated collection genuinely work, which only inconvenience users, and how to choose a level of protection for your situation.
- Web Scraping in Python: Playwright, Retries, DeduplicationHow to choose a stack for the task, how to write selectors that survive markup changes, how retries should work, and why deduplication is needed from the start.
Done for you
I will build a parser for your source
With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.
from $300 · 3 to 7 days
"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."