How Plugsky MCP works in practice
Everything in this guide runs on Plugskyโs live MCP endpoint (https://plugsky.com/mcp) or the npm bridge (npx -y @plugsky/mcp). A client request flows: JSON-RPC over Streamable-HTTP → auth (API key or OAuth 2.1 + Dynamic Client Registration with fine-grained scopes) → tool router → the service you asked for. Model calls run behind a 7-tier provider failover chain; browsing renders with headless Chromium; transcripts come from real captions and Groq Whisper; images generate on FLUX; code runs in a sandbox with no network; memory persists in your account (Postgres-backed); and every tool call is metered per user on your dashboard.
Relevant live tools for this article: plugsky_web_fetch, plugsky_url_ground and plugsky_site_ingest.
How to make ChatGPT or Claude read any website (URL ingest MCP)
Direct answer
To make ChatGPT or Claude read any website in 2026, install a web-fetch or URL-ingest MCP (Firecrawl, Plugsky
plugsky_url_ingest/plugsky_web_fetch, or Exa). Paste the URL in chat; the MCP fetches, cleans, and (optionally) chunks the page into a RAG collection. The LLM then answers grounded in that content. This is what people mean by "train ChatGPT on my website" โ technically it's ingestion, not training, but the practical UX is the same.
Key facts
| Common misconception | "Training" โ what actually happens is ingestion or retrieval |
| Best MCPs | Firecrawl (scrape/crawl), Plugsky (plugsky_web_fetch/plugsky_url_ingest/plugsky_web_crawl), Exa (search + contents) |
| Handles JS-rendered pages | Yes (Firecrawl, Plugsky, Playwright-backed) |
| Handles PDFs at URL | Yes (Firecrawl, Plugsky) |
| Crawl multiple pages | Yes (Firecrawl crawl, Plugsky plugsky_web_crawl) |
| Auto-chunk into RAG | Plugsky plugsky_url_ingest โ one shot |
| Free tier | Firecrawl 500 pages/mo; Plugsky free tier includes URL ingest |
| Respect robots.txt | Firecrawl โ , Plugsky โ , self-host varies |
| Auth-required pages | Only with browser-MCP + session (Playwright, Plugsky browser group) |
TL;DR
- "Training on a website" = ingestion. The LLM doesn't relearn weights; it retrieves the page as context.
- Single-page:
plugsky_web_fetchor Firecrawlscrape. - Multi-page:
plugsky_web_crawlor Firecrawlcrawl. - One-shot RAG:
plugsky_url_ingestโ URL in, queryable collection out. - Auth pages: needs browser + session (Playwright MCP or Plugsky browser group when it ships).
How it works
Option 1: Plugsky (one-shot fetch or ingest)
- Add
https://plugsky.com/mcp(see [Add MCP to Claude](/blog/add-mcp-to-claude)). - In chat: "Use
plugsky_web_fetchon https://acme.com/pricing and summarize." โ for one-off reads. - Or: "Use
plugsky_url_ingeston https://acme.com and add it to myacme-docscollection." โ for repeated queries. - Then: "Query my
acme-docscollection: what's the enterprise price?" โ fast, grounded, cheap.
Option 2: Firecrawl
{ "mcpServers": { "firecrawl": { "url": "https://api.firecrawl.dev/mcp", "headers": {"Authorization":"Bearer fc-..."} } } }
Best for: deep crawls (whole domains), highly JS-heavy sites, PDF-heavy content.
Option 3: Self-host with Playwright
Use Playwright MCP for auth-required pages, then pipe results into Plugsky's RAG collection via plugsky_documents_ingest.
Winning patterns
- Read a competitor's pricing page โ "Fetch competitor.com/pricing and compare to our tiers."
- Train an agent on your docs โ "Crawl docs.mycompany.com to a collection; answer support questions from it."
- Watch for changes โ recurring MCP prompt: "Every Monday, refetch these 5 URLs and alert on diffs." (Plugsky scheduler +
plugsky_web_fetch.) - Multi-lingual โ fetch English page, ask
plugsky-arabicto translate it.
"Training on a website" vs ingestion vs fine-tuning โ clearing up the terms
| Training / fine-tuning | Update the LLM's weights on your data | $$$$ | Days | Rare โ specific domain style |
| RAG (ingestion) | Store text as embeddings; retrieve at query time | $ | Minutes | 95% of "train on my content" use cases |
| Prompt stuffing | Paste content into every prompt | $$ per call | Instant | One-off, small content |
| Context caching | Reuse a cached long prompt | $$ once, $ per reuse | Instant after first | Long-form docs used repeatedly |
Most users saying "train it" want RAG. That's exactly what plugsky_url_ingest does in one tool call.
Comparison โ URL-ingest MCPs
| Fetch single URL | โ | โ | โ |
| Crawl whole site | โ | โ (best in class) | โ ๏ธ |
| Auto-chunk to RAG | โ
(plugsky_url_ingest) | โ (do it yourself) | โ |
| PDF handling | โ | โ | โ ๏ธ |
| JS rendering | โ | โ | โ ๏ธ |
| Free tier | โ | โ (500 pages/mo) | โ (small) |
| Auth pages via session | โณ (browser group roadmap) | โ | โ |
FAQ
Q: Can Claude/ChatGPT actually "remember" a website?
A: Not natively. They read it in-context (short-term) or via RAG (persistent, recallable). Weights are not updated.
Q: How do I train ChatGPT on my whole documentation site?
A: Use plugsky_web_crawl (or Firecrawl crawl) then push to a Plugsky RAG collection. Query from any client.
Q: Will this work on pages behind a login?
A: Not with plain web-fetch. Use Playwright MCP with a persisted session, or wait for Plugsky's browser group (plugsky_browse_session).
Q: Does the LLM re-fetch every time I ask a question?
A: If you use plugsky_web_fetch, yes. If you use plugsky_url_ingest (RAG), no โ it queries the stored embeddings.
Q: What about robots.txt?
A: Firecrawl and Plugsky respect robots.txt by default. You can override on your own sites.
Q: How do I keep a RAG collection fresh?
A: Schedule a recurring plugsky_url_ingest on the URL. Plugsky de-duplicates by URL + content hash so re-ingests are cheap.
Trust & sources
- Author: Mustafa Hasan.
- Last updated: {DATE}.
- References: Firecrawl docs, Exa docs, [Plugsky MCP](/mcp).
- Related: [YouTube via MCP](/blog/chatgpt-claude-youtube-mcp) ยท [Best free MCP servers](/blog/best-free-mcp-servers-2026).
Plugsky (2026). “How to make ChatGPT or Claude read any website (URL ingest MCP)”. Plugsky. Available at: https://plugsky.com/articles/make-chatgpt-claude-read-website-url-ingest-mcp (last updated 2026-09-30).