Key facts
| Parsing | Local tools extract text and structure from PDFs and office files |
| Embeddings | A local embedding model vectorises chunks on-device |
| Storage | Chroma, Qdrant, pgvector or FAISS hold vectors and metadata |
| Answers | A local chat model summarizes and answers with citations |
| OCR | Scanned documents need a local OCR step before embedding |
| Weak spot | Tables, forms and multi-column layouts challenge naive parsing |
| Hosted option | Plugsky offers OpenAI-compatible chat and embeddings |
| Endpoint status | Chat, embeddings and RAG are live; file endpoints coming soon |
TL;DR
- Parse, chunk, embed and answer with every component local.
- Parsing quality sets the ceiling; bad text gives bad answers.
- Always return citations so users can verify the source page.
- Handle scanned pages with a local OCR pass.
- Test with real documents, especially tables and forms.
How it works, step by step
- Inventory document types, languages and whether pages are scanned.
- Choose local parsing tools, adding OCR for image-only pages.
- Chunk text with page and section metadata attached.
- Embed chunks locally and store vectors with metadata.
- Retrieve top chunks per question and answer with citations.
- Review failures: parsing errors, missing chunks, unclear answers.
- Iterate on chunking and metadata, not just prompt wording.
Try it yourself
The document pipeline
A document assistant is a pipeline before it is a model. Files arrive in mixed formats, some with a clean text layer and some as scans. Local parsers extract text and structure; an OCR step handles image-only pages; a chunker splits the text; an embedding model turns chunks into vectors; a local store holds them with metadata. The chat model only writes the final answer.
Every stage can run on one machine. That is what makes the privacy claim credible: no document, embedding or question leaves the device, and the pipeline works with networking disabled.
Why parsing and chunking decide quality
Most disappointing answers trace back to extraction. Multi-column PDFs interleave lines, tables lose their structure, headers repeat on every page, and scans produce nothing at all without OCR. Inspect the extracted text for a sample of documents before you tune anything else.
- Chunk on meaning: paragraph and section boundaries beat fixed character counts.
- Keep identifiers: document id, page and section travel with the chunk.
- Deduplicate: repeated headers and boilerplate create near-identical vectors that crowd out real matches.
- Test retrieval first: if the right chunk is not retrieved, no prompt can recover the answer.
Privacy, evaluation and hosted options
Build a small evaluation set from your own documents: questions with known answers, plus questions whose answer is not in the corpus, to test whether the assistant admits it does not know. Score retrieval separately from answer quality so you can see which stage is failing.
If some documents may leave the device, the same architecture can call a hosted endpoint for generation while keeping parsing and vectors local. Plugsky provides OpenAI-compatible chat and embeddings, both live, so no client rewrite is needed; file and batch endpoints are coming soon, so keep bulk ingestion local for now. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Stage | Offline local stack | Plugsky hosted API | Check before deciding |
|---|---|---|---|
| Parsing | Local tools and OCR | Client-side, unchanged | File complexity |
| Embeddings | Local embedding model | Embeddings endpoint, live | Dimensions and languages |
| Storage | Local vector store | Your store stays as is | Corpus size |
| Answers | Local chat model | 30+ models on one API | Answer quality bar |
| Privacy | Files never leave the device | Requests go to your deployment | Document sensitivity |
Frequently asked questions
Can I chat with PDFs without uploading them?
Yes. Parse the PDFs locally, embed the text with a local embedding model, store vectors locally and generate answers with a local chat model. No step requires a network call.
Why are answers wrong for some documents?
Usually parsing. Multi-column layouts, tables and scanned pages produce jumbled or missing text. Inspect the extracted text before blaming the model.
Do I need OCR?
Only for scanned or image-only pages. Check whether each document contains an embedded text layer; if not, add a local OCR step before chunking.
How do I cite sources?
Store page and section metadata with every chunk and return it with the answer, so users can open the original document and verify the claim.
How large a corpus can it handle?
Local setups handle thousands to hundreds of thousands of chunks depending on memory and vector store. Measure ingestion time and retrieval latency as the corpus grows.
What about tables and forms?
They are the hardest case. Consider extracting them into structured rows before embedding, and test retrieval with questions that require exact values.
Can I keep the local store and use a hosted model?
Yes. Keep parsing and vectors local and route generation, or embeddings, to an OpenAI-compatible endpoint when connectivity and policy allow.