RAG

How do you build RAG over PDF documents?

PDF RAG starts with extraction: get clean text out of columns, tables and headers, use OCR for scanned pages, then chunk along document sections and keep page numbers for citations. Plugsky collections accept PDFs and handle chunking, embedding and indexing automatically, returning chunks with page-level source references.

Key facts

FormatsPDF ingestion alongside DOCX, TXT, MD and HTML
Chunking500-token chunks with 50-token overlap by default
CitationsChunks carry source references such as file and page
RetrievalKeyword, vector and hybrid search with optional reranking
Scanned filesOCR the source before upload when a PDF contains images of text
Data handlingPer-collection encryption at rest; API data is not used to train models
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Extraction quality limits retrieval quality more than model choice does.
  • Scanned PDFs need OCR; without it the index contains empty text.
  • Chunk on headings and page boundaries so citations map to a page.
  • Tables need structure preserved or their meaning is lost.
  • Keep the source filename and page on every chunk for verification.

How it works, step by step

  1. Sample your PDFs and check whether text is embedded or only images.
  2. Run OCR for scanned documents and review the output for errors.
  3. Strip repeated headers, footers and page numbers that add noise.
  4. Chunk along headings and sections, keeping tables with their captions.
  5. Upload to a collection and confirm page references appear in results.
  6. Ask known questions to verify the answer cites the correct page.
  7. Re-process documents when extraction settings or sources change.
1Sample your PDFsand check whethertext is embedded or2Run OCR for scanneddocuments andreview the output3Strip repeatedheaders, footersand page numbers4Chunk alongheadings andsections, keeping5Upload to acollection andconfirm page6Ask known questionsto verify theanswer cites the

Try it yourself

Open the Chat with PDF tool →

Extraction is where PDF RAG is won

PDFs are a container, not a text format. A file may hold real text, images of text, or a mix, and multi-column layouts often defeat naive extraction by interleaving lines from different columns. Tables lose their meaning when cells are flattened into a stream of words. If extraction produces garbled text, no embedding model or prompt can rescue retrieval.

Start by sampling the corpus. Documents with selectable text can usually be extracted directly; scanned documents need OCR, and the OCR output should be spot-checked because errors become permanent chunks. Remove repeated headers, footers and page numbers, which otherwise appear in hundreds of chunks and pollute similarity results.

Chunking with citations in mind

Chunk along the document's own structure: headings, numbered sections, clauses and page boundaries. Plugsky collections chunk at 500 tokens with 50-token overlap by default, which suits continuous prose, but the goal is for each chunk to be a coherent unit that a reader would recognise. Tables should stay with their captions and headers so a row can be interpreted.

Preserve the page number on every chunk. Citations that point to a page let a user open the PDF and verify the passage in seconds, which is the practical difference between a demo and a tool people trust.

Managed extraction versus DIY

A DIY pipeline gives you full control over layout analysis, OCR engines and table parsing, at the cost of maintaining all of it and handling the long tail of weird files. A managed path such as Plugsky collections accepts PDFs directly and runs chunking, embedding and indexing, which is usually enough for text-based PDFs and standard layouts.

A common split is to run OCR and cleanup yourself for difficult documents, then upload the cleaned text as TXT or MD. That keeps the hardest step under your control while the retrieval layer stays managed.

Testing PDF retrieval

Build a small set of questions whose answers sit in known pages, then check that retrieval returns those pages and that the generated answer cites them. Test a table question, a multi-column page and a scanned document, because each stresses a different part of the pipeline. Re-run the set after any change to extraction or chunking.

Try document questions with the Chat with PDF tool, then run the pipeline on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ConcernPlugsky collectionsDIY extraction plus APIOutsourced OCR-only
ExtractionUpload PDFs directlyYou control layout and OCRText only, you still index
ChunkingAutomatic with metadata and page referencesYou implement splittingYou implement splitting
TablesHandled as text; complex tables may need preprocessingFull control over structureUsually flattened
Scanned filesPre-process with OCR before uploadYour OCR pipelineBuilt for scans
OperationsManaged retrievalYou run extraction jobsAdditional pipeline to maintain

Frequently asked questions

Can Plugsky read scanned PDFs?

The ingestion path works best with PDFs that contain selectable text. For scans, run OCR first and upload the resulting text or searchable PDF so the index contains real content.

Why do citations matter for PDFs?

A page-level reference lets a user open the document and verify the passage quickly. Without it, an answer is hard to trust or review.

How are tables handled?

Tables are extracted as text, which can lose structure. For table-heavy documents, preprocess them into a structured format before ingestion.

What chunk size works for PDFs?

The 500-token default with 50-token overlap suits most prose. For dense technical documents, test smaller chunks against recall on your own questions.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.