Skip to content
PD
Парсинг данных

Scraping a Table From a Website: Three Approaches

How to get a table off a page: by hand, with ready tools, and in code. What to do about merged cells, pagination, and tables loaded by scripts.

All articles in the guide Парсинг данных · 11

A table is the simplest object to collect and simultaneously a source of non-obvious errors. Here are three approaches and the traps.

Approach 1: by hand

Select, copy, paste into a spreadsheet. The browser usually preserves rows and columns.

When it works: a one-off task, one page, not many rows.

When it stops: the table is paginated, there are thousands of rows, or the task repeats regularly. Copying becomes more expensive than a script quickly - and switching earlier is better, because manual work also makes mistakes.

The trap: not everything copied. If rows load on scroll, only what is visible reaches the clipboard.

Approach 2: ready tools

Spreadsheet applications can import tables from a page by URL, browser extensions exist, and there are visual tools.

When it works: the table is marked up as a table, is available without sign-in, updates regularly and updating is what you need.

Limits:

  • Only works with genuine table markup. A set of blocks will not be recognised.
  • Cannot handle content drawn by scripts.
  • Pagination is not followed.
  • Authentication is usually unavailable.

A useful property: some tools refresh the data automatically. For a regularly changing table that is a complete solution with no code at all.

Approach 3: in code

Full control and the only option for complex cases.

The order is the same as in any collection: first check whether the data is in the raw response. Often the table arrives as a structured payload in a background request - then there is no markup to parse at all, which is the best case.

If the data is in the markup, the task reduces to walking rows and cells. Parsing libraries handle that out of the box, and some tools convert a table into a data structure in one call.

On stack and reliability: scraping in Python.

Common traps

Merged cells. The leading cause of wrong data. A merged cell spans several rows or columns, and without expansion the subsequent values shift. The correct handling is repeating the value in every row it covers.

Multi-level headers. A two-row header requires deciding what to call the columns. Better decided at collection time than untangled in the export.

Pagination. A five-page table is five requests, and the last page is usually partial. Check the total row count if the page states one.

Loading on scroll. Rows appear as you scroll. Not collectable without browser automation.

Not a table. Visually a table, in the markup a set of blocks. Row iteration will not work; you have to parse the block structure.

Numbers as text. Spaces as thousand separators, commas instead of points, a currency symbol in the same cell. Convert at collection time rather than later in the editor.

Hidden rows and columns. Present in the markup, invisible on screen. They land in the export unless filtered out.

Verifying the result

Three quick checks that catch nearly everything:

  1. The row count matches what the page states.
  2. Columns have not shifted: look at several random rows and compare with the page.
  3. Types are right: numbers as numbers, dates as dates.

Exporting the result: scraping to Excel. The overview is in the scraping guide.

FAQ

How do I quickly copy a table from a site into a spreadsheet?

For a one-off, selecting and copying is enough: the browser usually preserves rows and columns. If the table is large or paginated, copying becomes more expensive than writing a script, and switching is worth doing sooner than it feels.

Why does only part of the table copy?

Usually because only some rows are on the page: the rest load on scroll or live on other pages. The other case is a table drawn not as a table but as a set of blocks, so the structure is lost on copy.

What should I do about merged cells?

Expand them at collection time: a merged cell value repeats in every row it covers. Otherwise rows shift and values land in the wrong columns. This is the most common reason an export looks right and contains wrong data.

More on this topic

Done for you

I will build a parser for your source

With protection bypass, proxies and export to a sheet, a database or Telegram. It runs on a schedule without you.

from $300 · 3 to 7 days

Similar caseEtsy Keyword FinderA queue-driven keyword audit app for Etsy sellers: submit a listing ID plus up to 20 keywords, and background workers walk the real search results step by step with live screenshots.

"Very fast parsing, thank you! It even returned a few more numbers than expected, I recommend him to everyone. I have ordered twice now, happy with all of it, and I will be back."

sotasoftdv · KworkTranslated from Russian