Key facts
| Approach | Crawl or export site pages and upload the text to collections |
| Formats | HTML, MD, TXT and DOCX are accepted for ingestion |
| Chunking | 500-token chunks with 50-token overlap by default |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Citations | Chunks return with source references so answers link to pages |
| Refresh | Re-crawl on a schedule and replace changed documents |
| Data handling | Per-collection encryption at rest; API data is not used to train models |
| Product status | Live |
TL;DR
- Strip navigation, cookie banners and footers; they match every query.
- Canonicalise URLs so the same page is not indexed twice.
- Chunk by headings so each chunk is one idea with one source.
- Re-crawl on a schedule; documentation changes faster than you think.
- Only index content you own or are allowed to reproduce.
How it works, step by step
- Prefer a sitemap or content export over blind crawling for coverage and speed.
- Respect robots rules and rate limits, and identify your crawler honestly.
- Extract main content, dropping navigation, sidebars and repeated banners.
- Canonicalise URLs and deduplicate pages that render the same content.
- Chunk by heading hierarchy and keep the page URL and title as metadata.
- Upload to a collection and query with real user questions.
- Schedule re-crawls and replace changed pages rather than appending new copies.
Try it yourself
Crawl or export, then clean
A sitemap is the best starting point: it lists canonical URLs, which avoids the duplicate pages and parameterised links that blind crawling collects. If the site is generated from source files such as Markdown, export those directly; they are cleaner than rendered HTML and cheaper to process.
Extraction should keep the main content and discard the shell. Navigation menus, cookie notices, related-links blocks and footers appear on every page, so they match almost any query and crowd out useful chunks. Convert the remaining content to text or Markdown, preserving headings, lists and code blocks.
Chunking and metadata
Chunk along the heading hierarchy so each chunk covers one topic. Plugsky collections chunk at 500 tokens with 50-token overlap by default, which works well for article-style pages; for long reference pages, smaller chunks keyed to subheadings retrieve more precisely. Keep the page title, URL and section heading as metadata so citations render as links and filters remain possible.
Deduplicate near-identical pages, such as regional variants or print views, before ingestion. They create repeated chunks that return together and make answers look padded.
Freshness and legal care
Web content is a moving target. Schedule re-crawls, compare the new content with what is indexed, and replace changed documents in place using a stable URL-based identifier. Without this, an assistant will confidently quote a version of the page that no longer exists.
Only index content you own or have permission to use. Respect robots rules and terms, identify your crawler, and avoid hammering servers. For third-party documentation, check licensing before reproducing it in an internal assistant.
Retrieval and answers over a site
Once indexed, queries can use keyword, vector or hybrid retrieval with optional reranking, and results carry source references so answers can link back to the exact page. This makes a website corpus useful for support assistants, documentation search and internal knowledge tools without building a search stack by hand.
Prototype with the RAG sandbox, then run ingestion against collections on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Concern | Managed RAG collections | Site search widget | DIY vector pipeline |
|---|---|---|---|
| Ingestion | Upload crawled or exported text | Indexes rendered pages | You build the crawler |
| Chunking | Automatic with overlap and metadata | Page-level only | You implement and tune |
| Retrieval | Keyword, vector and hybrid with reranking | Keyword matching | Whatever you assemble |
| Citations | Chunks with source references | Page links | You implement attribution |
| Ops | Managed | Vendor-hosted | You run every component |
Frequently asked questions
Can Plugsky crawl my website?
No. You crawl or export the content and upload it to collections. A sitemap-based crawler keeps the job small and avoids duplicate pages.
How do I avoid indexing navigation menus?
Extract the main content area and drop repeated shells such as headers, sidebars and footers. Boilerplate matches nearly every query and lowers precision.
How often should I re-index a site?
As often as the content changes. Re-crawl on a schedule, diff against the indexed version, and replace changed pages in place.
Should I index third-party content?
Only with permission. Respect robots rules and licensing, and keep external content in a clearly separated corpus if you do ingest it.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.