RAG

How do you build RAG over website content?

Website RAG starts with collection and cleanup: crawl or export pages, remove navigation and boilerplate, canonicalise URLs, then chunk by heading and index the text. A managed pipeline handles embedding and retrieval, and answers cite the page they came from. Refresh the corpus on a schedule so changed pages are re-indexed.

Key facts

ApproachCrawl or export site pages and upload the text to collections
FormatsHTML, MD, TXT and DOCX are accepted for ingestion
Chunking500-token chunks with 50-token overlap by default
RetrievalKeyword, vector and hybrid search with optional reranking
CitationsChunks return with source references so answers link to pages
RefreshRe-crawl on a schedule and replace changed documents
Data handlingPer-collection encryption at rest; API data is not used to train models
Product statusLive

TL;DR

  • Strip navigation, cookie banners and footers; they match every query.
  • Canonicalise URLs so the same page is not indexed twice.
  • Chunk by headings so each chunk is one idea with one source.
  • Re-crawl on a schedule; documentation changes faster than you think.
  • Only index content you own or are allowed to reproduce.

How it works, step by step

  1. Prefer a sitemap or content export over blind crawling for coverage and speed.
  2. Respect robots rules and rate limits, and identify your crawler honestly.
  3. Extract main content, dropping navigation, sidebars and repeated banners.
  4. Canonicalise URLs and deduplicate pages that render the same content.
  5. Chunk by heading hierarchy and keep the page URL and title as metadata.
  6. Upload to a collection and query with real user questions.
  7. Schedule re-crawls and replace changed pages rather than appending new copies.
1Prefer a sitemap orcontent export overblind crawling for2Respect robotsrules and ratelimits, and3Extract maincontent, droppingnavigation,4Canonicalise URLsand deduplicatepages that render5Chunk by headinghierarchy and keepthe page URL and6Upload to acollection andquery with real

Try it yourself

Open the RAG sandbox →

Crawl or export, then clean

A sitemap is the best starting point: it lists canonical URLs, which avoids the duplicate pages and parameterised links that blind crawling collects. If the site is generated from source files such as Markdown, export those directly; they are cleaner than rendered HTML and cheaper to process.

Extraction should keep the main content and discard the shell. Navigation menus, cookie notices, related-links blocks and footers appear on every page, so they match almost any query and crowd out useful chunks. Convert the remaining content to text or Markdown, preserving headings, lists and code blocks.

Chunking and metadata

Chunk along the heading hierarchy so each chunk covers one topic. Plugsky collections chunk at 500 tokens with 50-token overlap by default, which works well for article-style pages; for long reference pages, smaller chunks keyed to subheadings retrieve more precisely. Keep the page title, URL and section heading as metadata so citations render as links and filters remain possible.

Deduplicate near-identical pages, such as regional variants or print views, before ingestion. They create repeated chunks that return together and make answers look padded.

Web content is a moving target. Schedule re-crawls, compare the new content with what is indexed, and replace changed documents in place using a stable URL-based identifier. Without this, an assistant will confidently quote a version of the page that no longer exists.

Only index content you own or have permission to use. Respect robots rules and terms, identify your crawler, and avoid hammering servers. For third-party documentation, check licensing before reproducing it in an internal assistant.

Retrieval and answers over a site

Once indexed, queries can use keyword, vector or hybrid retrieval with optional reranking, and results carry source references so answers can link back to the exact page. This makes a website corpus useful for support assistants, documentation search and internal knowledge tools without building a search stack by hand.

Prototype with the RAG sandbox, then run ingestion against collections on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ConcernManaged RAG collectionsSite search widgetDIY vector pipeline
IngestionUpload crawled or exported textIndexes rendered pagesYou build the crawler
ChunkingAutomatic with overlap and metadataPage-level onlyYou implement and tune
RetrievalKeyword, vector and hybrid with rerankingKeyword matchingWhatever you assemble
CitationsChunks with source referencesPage linksYou implement attribution
OpsManagedVendor-hostedYou run every component

Frequently asked questions

Can Plugsky crawl my website?

No. You crawl or export the content and upload it to collections. A sitemap-based crawler keeps the job small and avoids duplicate pages.

How do I avoid indexing navigation menus?

Extract the main content area and drop repeated shells such as headers, sidebars and footers. Boilerplate matches nearly every query and lowers precision.

How often should I re-index a site?

As often as the content changes. Re-crawl on a schedule, diff against the indexed version, and replace changed pages in place.

Should I index third-party content?

Only with permission. Respect robots rules and licensing, and keep external content in a clearly separated corpus if you do ingest it.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.