Key facts
| Pipeline | Ingest, extract text, classify, extract fields, route, summarise |
| Extraction | Parser libraries for digital files, OCR for scans, then the model for judgement |
| Structured output | JSON mode constrains field extraction to valid JSON |
| Validation | Check extracted fields against schemas and business rules |
| Confidence | Route low-confidence documents to human review |
| Grounding | Summaries cite page or section references |
| Roadmap | File and vision endpoints are coming soon; text pipelines run today |
| Privacy | Redact and scope per document class; region-locked deployment available |
TL;DR
- Separate deterministic extraction from model judgement; do not use the model for parsing.
- Return typed fields through JSON mode and validate them against schemas.
- Attach a confidence score and route uncertain documents to humans.
- Cite pages in summaries so reviewers can verify quickly.
- Decide residency per document class before ingesting anything.
How it works, step by step
- Inventory document types, volumes and the fields each one must produce.
- Parse digital files with libraries; run OCR only where text is not embedded.
- Classify the document type and route it to the matching extraction schema.
- Extract fields with JSON mode and validate types, ranges and cross-field rules.
- Score confidence from validation failures and ambiguity, and queue low scores.
- Summarise long documents with page references and keep originals linked.
- Log every processing step with the document hash so results are reproducible.
Try it yourself
The document pipeline
Document work decomposes neatly: ingest and store, extract text, classify, extract structured fields, validate, route, and summarise. The agent matters most in the middle — classification, interpretation of ambiguous layouts and normalisation of inconsistent field formats. Everything else should be boring, deterministic code with clear failure modes.
Treat the pipeline as stages with explicit contracts. A stage that fails should stop and quarantine the document rather than passing a half-processed file downstream. That discipline is what keeps error rates measurable instead of mysterious.
Extraction and structured output
- Digital-first: use text extraction libraries before reaching for OCR, because embedded text is exact.
- Schema per document type: invoices, contracts and forms each get their own field set and rules.
- JSON mode: constrain the model to valid JSON so downstream code can parse without defensive hacks.
- Validation: types, ranges, checksums and cross-field consistency catch most extraction errors.
- Normalisation: dates, currencies and names converted to canonical formats at the edge.
- Confidence: combine validation failures and model signals into a routable score.
Review, privacy and throughput
Human review should be targeted. A well-tuned pipeline sends only a small fraction of documents to a queue, and reviewers see the extracted fields beside the source pages that justify them. Track the review rate and the correction patterns, then feed recurring corrections back into prompts, schemas and validation rules.
Documents often contain the most sensitive data in the organisation, so scope access per class, redact before storing derived text where possible, and choose deployment to match policy. Plugsky supports region choice plus VPC, on-prem and air-gapped deployment, and the live API covers function calling, JSON mode, embeddings and RAG. File and vision endpoints are coming soon, so today's pipelines parse text locally and send extracted content. 30+ models on one key let you run classification on a cheap model and complex interpretation on a frontier one. Plans are on the live pricing page.
Honest comparison
| Approach | Accuracy | Cost | Best for |
|---|---|---|---|
| Template or regex parsing | High for fixed layouts | Very low | Stable form formats |
| Text extraction plus model | High | Low to medium | Mixed digital documents |
| OCR plus model | Medium to high | Medium to high | Scanned and image files |
| Vision model end to end | Medium | High | Unknown layouts, low volume |
| Manual review only | Highest | Highest | Rare, high-value documents |
Frequently asked questions
Can the agent process scanned PDFs?
Yes, with OCR before the model stage. Accuracy depends on scan quality, so keep OCR output and confidence and route poor scans to review.
How do I keep extracted data accurate?
Extract with JSON mode, validate against schemas and business rules, normalise formats, and route low-confidence documents to humans rather than guessing.
Should I send whole documents to the model?
Only the relevant sections. Page-level or section-level chunks reduce cost and improve precision, especially for long contracts.
How do I handle multiple document types?
Classify first, then apply a type-specific schema and validation rules. A single generic extraction prompt produces mediocre results across layouts.
What about data residency?
Decide per document class. Plugsky offers region choice plus VPC, on-prem and air-gapped deployment for regulated content.
Can I evaluate extraction quality?
Yes. Build a labelled set per document type and score field-level precision and recall, plus review rate. Re-run it whenever prompts or schemas change.
Are file upload endpoints available?
File endpoints are coming soon. Today you can extract text locally or via your own storage layer and send the content through the live chat and JSON mode API.