Skip to content
PD
AI Automation 7 min read

PDF Is Not Text: Build an Ingestion Layer Before Your LLM Sees It

Most bad RAG answers come from broken PDF extraction, not a weak model. Here is the ingestion layer I build: triage, coordinate-based extraction, structure rebuilding, per-page OCR fallback and quality gates.

PD

Pavel Duglas

AI Automation & MVP Architect

Every second automation project I take on has the same hidden step: “the client sends us PDFs, we extract the text, the LLM does the rest.” And every time, the LLM gets blamed for answers that were doomed before the model saw a single token. Invoices with totals glued to the wrong line item. Contracts where clause 7 is interrupted by a page footer halfway through a sentence. Two-column research papers read straight across both columns, so every line is half of one thought and half of another.

The model is not the problem. The input is. Here is the ingestion layer I now build before any PDF touches a prompt, and the rules I use to keep it cheap and debuggable.

Why PDFs lie to you

A PDF is not a document in the way a Word file or HTML page is. It is a set of drawing instructions: put this glyph at these coordinates in this font. There is no concept of a paragraph, a sentence, a column or a table. Anything that looks like structure is something your eyes reconstruct from spacing.

That gives you a predictable list of failure modes:

  • Reading order. Text is stored in whatever order the generating software wrote it. Sometimes that matches visual order, often it doesn’t.
  • Hard line breaks everywhere. Every visual line ends with a break. Plain extraction gives you a paragraph split into 12 fragments.
  • Hyphenation. “imple-” at the end of one line, “mentation” at the start of the next. Your embedding sees two junk tokens.
  • Repeating noise. Headers, footers, page numbers, “Confidential” watermarks, repeated on every page and injected into the middle of your content.
  • Tables flattened into soup. Cells come out in some order, with no reliable separator, and numbers lose their column.
  • Broken encodings. Some fonts have no proper Unicode mapping. The text layer exists, but it decodes into symbols or shifted letters.
  • No text at all. Scanned pages are just images. Plain extraction returns an empty string and nobody notices.

If you pipe pdftotext output directly into chunking and embeddings, you are shipping all seven of these into your retrieval index.

Step 1: Triage every page before extracting anything

I don’t treat a PDF as one thing. I treat it as a list of pages, and each page gets a quick classification:

  • Born-digital with a clean text layer. The easy case.
  • Born-digital with a broken text layer. Text exists but decodes badly.
  • Scanned. An image with no text, or with a low-quality OCR layer someone else added.
  • Mixed. A digital page with an embedded scanned signature page or stamped appendix.

The checks are cheap and run in milliseconds per page:

  1. Characters per page. Below a threshold (I start at around 50) on a page that has visible content, it is probably scanned.
  2. Ratio of letters and digits to total characters. Broken font mappings produce a lot of symbols and private-use code points.
  3. Dictionary hit rate. Take a sample of words and check what share exists in a word list for the expected language. Clean text sits well above 80%. Garbage decoding falls off a cliff.
  4. Image coverage. If one image covers most of the page area and there is little text, route it to OCR.

The key decision: route per page, not per document. A 40-page contract with one scanned signature page should not be sent to OCR in full. That is slower, more expensive and usually less accurate than the native text layer.

Step 2: Extract with coordinates, not strings

The single biggest improvement is switching from “give me the text” to “give me every word with its bounding box and font size.” PyMuPDF and pdfplumber both do this in Python. Once you have x0, y0, x1, y1, size, text for every word, you can rebuild structure yourself instead of trusting the file’s internal order.

With coordinates you can:

  • Group words into lines by vertical position.
  • Detect columns by looking for a consistent vertical gap where no words appear across many lines.
  • Sort reading order properly: column by column, top to bottom.
  • Spot headings by font size relative to the page median.

This is maybe a day of work the first time, and after that it is a reusable module you drop into every project.

Step 3: Rebuild structure with boring heuristics

You don’t need a model for most of this. Simple, explainable rules handle the bulk of real documents.

Strip repeating headers and footers

A line that appears in roughly the same vertical band on most pages, with the same text (after replacing digits, so page numbers match), is noise. Here is the core of what I use:

import re
from collections import Counter

def find_repeating_lines(pages, band=0.08, min_share=0.6):
    # pages: list of lists of (text, y_ratio) where y_ratio is 0..1 from top
    counter = Counter()
    for lines in pages:
        seen = set()
        for text, y in lines:
            if y < band or y > 1 - band:
                key = re.sub(r'\d+', '#', text.strip().lower())
                seen.add(key)
        counter.update(seen)
    limit = len(pages) * min_share
    return {k for k, n in counter.items() if n >= limit}

Run it once per document, then drop matching lines from the top and bottom bands. On a typical contract that removes the company name, the document title and the “Page 3 of 41” line from every page.

Join lines into paragraphs

For each pair of consecutive lines in the same column, decide whether they belong to the same paragraph. The signals I combine:

  • Vertical gap. If the gap is close to the median line spacing, it is probably a continuation. A gap noticeably larger than median means a new paragraph.
  • Line end. A line ending in a period, colon or question mark is more likely to end a paragraph.
  • Next line start. A lowercase first letter is a strong continuation signal.
  • Line length. A line much shorter than the column width often ends a paragraph.
  • Indentation. A first-line indent signals a new paragraph in many layouts.

For hyphens: if a line ends in a hyphen and the joined word exists in your dictionary, remove the hyphen and join. If it doesn’t (“state-of-the-art”, “COVID-19”), keep the hyphen. That one check fixes a surprising amount of broken retrieval.

Mark headings

Lines with a font size clearly above the page median, or bold and short, become headings. I emit them as markdown headings in the cleaned output. That matters later because chunking on headings gives you chunks that actually mean something.

Step 4: Treat tables as a separate data type

Never let a table flow into paragraph text. Detect it (pdfplumber’s table finder works well for ruled tables, and for unruled ones I look for rows of words aligned on repeated x positions), extract it as rows and cells, and serialize it on its own.

My default is to store tables twice: as JSON rows for anything that needs exact numbers, and as a compact markdown table inside the text chunk so the LLM can read it in context. When a user asks “what was the Q3 total,” you want the model reading a table with a header row, not a string of 30 unlabeled numbers.

Step 5: OCR and vision models only where needed

Pages that fail triage go to OCR. Tesseract is fine for clean scans. For messy ones, a cloud OCR service or a multimodal model is worth the money.

Model prices keep dropping, and it is tempting to just send every page to a vision model as an image and ask for markdown. I have tested that. It works well on hard pages and it is overkill on easy ones. It is also slower, less deterministic, and occasionally invents a number that looks plausible. So I route:

  • Clean digital page: coordinate extraction, zero model calls.
  • Scanned but simple: classic OCR.
  • Scanned with complex tables, handwriting or stamps: vision model, with the output flagged as model-generated.

On a real batch of supplier documents this kept vision calls to roughly one page in ten, and the total ingestion cost per document stayed low enough to not think about.

Step 6: Quality gates before indexing

Every processed document gets a small report card:

  • Pages with no text after all fallbacks.
  • Garbage character ratio.
  • Dictionary hit rate of the final text.
  • Average word length (a sudden jump to 15+ usually means words glued together).
  • Share of pages that needed OCR or a vision model.

If any metric crosses a threshold, the document goes to a quarantine queue instead of the index. Someone looks at it, or it gets reprocessed with a different route. The worst outcome in a RAG system is not a missing document. It is a garbage document that gets retrieved confidently and quoted back to a user.

Step 7: Keep provenance on every chunk

Each chunk I store carries the source file, page number, bounding box of the region it came from, and the extraction route used (native, OCR, vision). This pays off three ways:

  1. Citations. You can show the user the exact page and highlight the region.
  2. Debugging. When an answer is wrong, you open the page and see instantly whether extraction or retrieval failed.
  3. Reprocessing. When you improve the pipeline, you can rebuild only the chunks that came through a weak route.

My default stack

For most client projects this is what I reach for:

  • PyMuPDF for fast word-level extraction with coordinates.
  • pdfplumber for table detection.
  • A small in-house module for triage, header removal, line joining and headings.
  • Tesseract for simple OCR, a multimodal model for the hard pages.
  • Output as markdown with page markers plus a JSON sidecar for tables and provenance.
  • Quality metrics written to the same database as the job, so bad documents are visible in the admin panel.

None of this is glamorous. But when a client says “the AI keeps getting the numbers wrong,” nine times out of ten the fix lives here, not in the prompt and not in a more expensive model. Build the ingestion layer first, measure it, and your LLM suddenly looks a lot smarter.

FAQ

Can I just send every PDF page to a multimodal model and skip all this?

You can, and for a small prototype it is a reasonable shortcut. In production it is slower, costs more at volume, is less deterministic and can occasionally hallucinate values on pages that native extraction would have read perfectly. I route only the hard pages to a vision model and use coordinate-based extraction for clean digital pages.

Which Python library should I start with for PDF extraction?

Start with PyMuPDF for speed and word-level coordinates, and add pdfplumber when you need table detection. Both give you bounding boxes and font sizes, which is what you need to fix reading order, columns, headers and paragraph joining. Avoid building on plain text output, because you lose the layout information you need later.

How do I know my PDF extraction is good enough for RAG?

Measure it per document: pages with no text, garbage character ratio, dictionary hit rate and average word length. Then sample 20 documents by hand and compare the cleaned output to the visual page. If headers leak into paragraphs, columns interleave or tables lose their structure, fix extraction before touching chunking or prompts.

Related articles