MCP Guides

How to make ChatGPT or Claude read any website (URL ingest MCP)

๐Ÿ”Œ Connect MCP๐Ÿ“ฆ npm package๐Ÿ“š Docs๐Ÿ’ฌ Playground๐Ÿงญ Full guide

How Plugsky MCP works in practice

Everything in this guide runs on Plugskyโ€™s live MCP endpoint (https://plugsky.com/mcp) or the npm bridge (npx -y @plugsky/mcp). A client request flows: JSON-RPC over Streamable-HTTP → auth (API key or OAuth 2.1 + Dynamic Client Registration with fine-grained scopes) → tool router → the service you asked for. Model calls run behind a 7-tier provider failover chain; browsing renders with headless Chromium; transcripts come from real captions and Groq Whisper; images generate on FLUX; code runs in a sandbox with no network; memory persists in your account (Postgres-backed); and every tool call is metered per user on your dashboard.

Relevant live tools for this article: plugsky_web_fetch, plugsky_url_ground and plugsky_site_ingest.

How to make ChatGPT or Claude read any website (URL ingest MCP)

Direct answer

To make ChatGPT or Claude read any website in 2026, install a web-fetch or URL-ingest MCP (Firecrawl, Plugsky plugsky_url_ingest/plugsky_web_fetch, or Exa). Paste the URL in chat; the MCP fetches, cleans, and (optionally) chunks the page into a RAG collection. The LLM then answers grounded in that content. This is what people mean by "train ChatGPT on my website" โ€” technically it's ingestion, not training, but the practical UX is the same.

Key facts

Common misconception"Training" โ€” what actually happens is ingestion or retrieval
Best MCPsFirecrawl (scrape/crawl), Plugsky (plugsky_web_fetch/plugsky_url_ingest/plugsky_web_crawl), Exa (search + contents)
Handles JS-rendered pagesYes (Firecrawl, Plugsky, Playwright-backed)
Handles PDFs at URLYes (Firecrawl, Plugsky)
Crawl multiple pagesYes (Firecrawl crawl, Plugsky plugsky_web_crawl)
Auto-chunk into RAGPlugsky plugsky_url_ingest โ€” one shot
Free tierFirecrawl 500 pages/mo; Plugsky free tier includes URL ingest
Respect robots.txtFirecrawl โœ…, Plugsky โœ…, self-host varies
Auth-required pagesOnly with browser-MCP + session (Playwright, Plugsky browser group)

TL;DR

  • "Training on a website" = ingestion. The LLM doesn't relearn weights; it retrieves the page as context.
  • Single-page: plugsky_web_fetch or Firecrawl scrape.
  • Multi-page: plugsky_web_crawl or Firecrawl crawl.
  • One-shot RAG: plugsky_url_ingest โ€” URL in, queryable collection out.
  • Auth pages: needs browser + session (Playwright MCP or Plugsky browser group when it ships).

How it works

Option 1: Plugsky (one-shot fetch or ingest)

  1. Add https://plugsky.com/mcp (see [Add MCP to Claude](/blog/add-mcp-to-claude)).
  2. In chat: "Use plugsky_web_fetch on https://acme.com/pricing and summarize." โ€” for one-off reads.
  3. Or: "Use plugsky_url_ingest on https://acme.com and add it to my acme-docs collection." โ€” for repeated queries.
  4. Then: "Query my acme-docs collection: what's the enterprise price?" โ€” fast, grounded, cheap.

Option 2: Firecrawl

{ "mcpServers": { "firecrawl": { "url": "https://api.firecrawl.dev/mcp", "headers": {"Authorization":"Bearer fc-..."} } } }

Best for: deep crawls (whole domains), highly JS-heavy sites, PDF-heavy content.

Option 3: Self-host with Playwright

Use Playwright MCP for auth-required pages, then pipe results into Plugsky's RAG collection via plugsky_documents_ingest.

Winning patterns

  • Read a competitor's pricing page โ€” "Fetch competitor.com/pricing and compare to our tiers."
  • Train an agent on your docs โ€” "Crawl docs.mycompany.com to a collection; answer support questions from it."
  • Watch for changes โ€” recurring MCP prompt: "Every Monday, refetch these 5 URLs and alert on diffs." (Plugsky scheduler + plugsky_web_fetch.)
  • Multi-lingual โ€” fetch English page, ask plugsky-arabic to translate it.

"Training on a website" vs ingestion vs fine-tuning โ€” clearing up the terms

Training / fine-tuningUpdate the LLM's weights on your data$$$$DaysRare โ€” specific domain style
RAG (ingestion)Store text as embeddings; retrieve at query time$Minutes95% of "train on my content" use cases
Prompt stuffingPaste content into every prompt$$ per callInstantOne-off, small content
Context cachingReuse a cached long prompt$$ once, $ per reuseInstant after firstLong-form docs used repeatedly

Most users saying "train it" want RAG. That's exactly what plugsky_url_ingest does in one tool call.

Comparison โ€” URL-ingest MCPs

Fetch single URLโœ…โœ…โœ…
Crawl whole siteโœ…โœ… (best in class)โš ๏ธ
Auto-chunk to RAGโœ… (plugsky_url_ingest)โŒ (do it yourself)โŒ
PDF handlingโœ…โœ…โš ๏ธ
JS renderingโœ…โœ…โš ๏ธ
Free tierโœ…โœ… (500 pages/mo)โœ… (small)
Auth pages via sessionโณ (browser group roadmap)โŒโŒ

FAQ

Q: Can Claude/ChatGPT actually "remember" a website?

A: Not natively. They read it in-context (short-term) or via RAG (persistent, recallable). Weights are not updated.

Q: How do I train ChatGPT on my whole documentation site?

A: Use plugsky_web_crawl (or Firecrawl crawl) then push to a Plugsky RAG collection. Query from any client.

Q: Will this work on pages behind a login?

A: Not with plain web-fetch. Use Playwright MCP with a persisted session, or wait for Plugsky's browser group (plugsky_browse_session).

Q: Does the LLM re-fetch every time I ask a question?

A: If you use plugsky_web_fetch, yes. If you use plugsky_url_ingest (RAG), no โ€” it queries the stored embeddings.

Q: What about robots.txt?

A: Firecrawl and Plugsky respect robots.txt by default. You can override on your own sites.

Q: How do I keep a RAG collection fresh?

A: Schedule a recurring plugsky_url_ingest on the URL. Plugsky de-duplicates by URL + content hash so re-ingests are cheap.

Trust & sources

  • Author: Mustafa Hasan.
  • Last updated: {DATE}.
  • References: Firecrawl docs, Exa docs, [Plugsky MCP](/mcp).
  • Related: [YouTube via MCP](/blog/chatgpt-claude-youtube-mcp) ยท [Best free MCP servers](/blog/best-free-mcp-servers-2026).
Cite this page

Plugsky (2026). “How to make ChatGPT or Claude read any website (URL ingest MCP)”. Plugsky. Available at: https://plugsky.com/articles/make-chatgpt-claude-read-website-url-ingest-mcp (last updated 2026-09-30).