Key facts
| Default chunking | 500-token chunks with 50-token overlap in Plugsky collections |
| Formats | PDF, DOCX, TXT, MD and HTML ingestion |
| Structure signals | Markdown and HTML headings, PDF pages and document sections |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Citations | Chunks return with source references such as file and page |
| Embeddings | plugsky-embed-v1 (1536d) and plugsky-embed-large (3072d) |
| Batch limits | Up to 2,048 embedding inputs per request, max 8,191 tokens each |
| Product status | Live |
TL;DR
- Chunking is a retrieval decision, not a preprocessing detail.
- Structure-aware chunks beat blind splits on documents with headings and sections.
- Overlap protects meaning at boundaries but duplicates storage and cost.
- Store metadata with every chunk so filters and citations work.
- Tune chunk size against recall on your own questions, not a general rule.
How it works, step by step
- Inspect how your documents express structure: headings, sections, tables, pages.
- Pick a strategy that respects those boundaries rather than fixed character counts.
- Set a starting size near 500 tokens with modest overlap.
- Attach metadata such as source, page, section and access level.
- Measure recall@k with a golden question set at two or three chunk sizes.
- Check whether answers cite complete thoughts or truncated fragments.
- Re-run the evaluation whenever documents or chunking change.
Original data
Try it yourself
Open the RAG chunk size calculator →
The main chunking strategies
Fixed-size chunking splits text every N tokens, sometimes with overlap. It is predictable and easy to implement, but it cuts through sentences and sections. Recursive chunking tries larger separators first, such as headings and paragraphs, and falls back to smaller ones only when a piece is still too large. Structure-aware chunking follows the document model itself: Markdown headings, HTML sections, PDF pages, code blocks or table rows.
Semantic chunking groups consecutive sentences whose embeddings are close, aiming for topical units rather than fixed lengths. It can improve coherence but adds an embedding pass during preprocessing and makes chunk sizes less predictable.
Size and overlap trade-offs
Small chunks match focused facts and keep the generator prompt tight, but they fragment explanations that need surrounding context. Large chunks preserve context but dilute relevance, because one vector now represents several topics. Overlap reduces boundary loss by repeating a slice of text at each edge, at the cost of duplicated storage and repeated embedding work.
Plugsky collections default to 500-token chunks with 50-token overlap, which is a reasonable starting point for prose. Adjust from there using measured recall: if answers miss details that exist in the corpus, try smaller chunks; if they miss relationships between facts, try larger ones or parent-child retrieval, where a small chunk is indexed while the larger parent section is passed to the model.
Metadata is part of the chunk
A chunk that loses its source is hard to cite, filter or delete. Store the document identifier, title, page or section, version and access level with every chunk. That metadata powers filtered retrieval for multi-tenant systems, enables citations that point to a page, and makes it possible to remove a document cleanly when it is retired.
Keep identifiers stable across re-ingestion so updates replace existing content instead of creating duplicates. Duplicate chunks are a common cause of repetitive answers and inflated storage.
Testing chunking on Plugsky
Upload a representative sample to a collection and query it with real questions. Keyword, vector and hybrid retrieval can be compared with the same chunks, and optional reranking shows whether ordering or content is the limiting factor. Because every response includes source references, you can see exactly which chunk produced an answer and whether it was truncated.
Estimate token sizes with the RAG chunk size calculator, then run the evaluation loop on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Strategy | Strength | Weakness | Best for |
|---|---|---|---|
| Fixed-size | Simple, predictable cost | Splits sentences and sections | Uniform plain text |
| Recursive | Respects paragraphs and headings | Can still split mid-thought | General prose and docs |
| Structure-aware | Follows headings, pages and code | Needs per-format handling | Markdown, HTML, PDF and code |
| Semantic | Topically coherent units | Extra embedding pass, variable size | Long narrative text |
| Parent-child | Precise matching with richer context | More storage and lookup logic | Technical manuals and policies |
Frequently asked questions
What chunk size should I start with?
Around 500 tokens with a small overlap is a common starting point, and it is the Plugsky collection default. Tune from there against recall on your own question set.
Is overlap always worth it?
Overlap helps when meaning spans a boundary, but it duplicates storage and embedding work. Small overlaps such as 10 percent of chunk size are usually enough.
Should I chunk by tokens or characters?
Tokens align with what the model and embedding endpoint count, so token-based chunking is safer. Character counting can silently overflow model limits.
Can I change chunking after ingestion?
Yes, but existing chunks were created with the old settings. Re-ingest the affected documents so retrieval and citations stay consistent.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.