Key facts
| What it is | A production Model Context Protocol (MCP) server for Plugsky — remote and local |
| Remote endpoint | https://plugsky.com/mcp (Streamable HTTP, stateless, JSON-RPC 2.0) |
| Local package | npx -y @plugsky/mcp (npm, v1.1.0, zero dependencies, stdio) |
| Tools | Remote: plugsky_chat, plugsky_models, plugsky_fusion · Local: 9 tools in 6 groups |
| Clients | ChatGPT (Connectors), Claude, Cursor, VS Code, Windsurf, Zed — and any MCP client |
| Auth | Authorization: Bearer sk-live-… (revocable per key) |
| Free tier | plugsky-micro and plugsky-lite are free — zero-cost agent inference |
| Stress result | 210/210 requests, 0 failures, 0× HTTP 5xx; chat 30/30 under 10-way concurrency |
| Product status | Live |
TL;DR
- MCP is the 2026 standard for connecting AI apps to tools and data — one connector, no bespoke integrations.
- Plugsky speaks it two ways: a hosted URL (
https://plugsky.com/mcp) and a local npm package (npx -y @plugsky/mcp). - Remote setup in ChatGPT: Settings → Connectors → Add MCP server → paste the URL + API key. Claude, Cursor, VS Code and Windsurf take a two-line JSON.
- The remote endpoint exposes chat (36 models), the model catalogue and multi-model fusion; the local package adds RAG search, document ingestion, platform tools and usage.
- Stressed live on 2026-09-29: 210 requests, zero failures — including 30 concurrent chat calls.
What is MCP (and why should you care)?
The Model Context Protocol is an open standard that lets AI apps call external tools and data through a single connector. Before MCP, every assistant needed a custom integration for every service. With MCP, you point an app at a server URL and it discovers the available tools automatically — the same way a browser discovers what a website offers without a plugin per site.
That matters for teams because the assistant you already use (ChatGPT for writing, Claude for analysis, Cursor for code) can now reach your models, your documents and your internal tools without switching apps or copying data between tabs.
How it works, step by step
- Add the Plugsky server to your MCP client — a URL in ChatGPT, a two-line entry in Claude/Cursor/VS Code.
- Authenticate with your Plugsky API key (Bearer token); keys are revocable from the dashboard.
- The client calls
initialize+tools/listand discovers the available tools along with server instructions. - When the model needs something, it calls a tool — chat with a specific model, search your knowledge base, run a platform tool or merge several models' answers.
- Your agent answers with real data and cites your own documents; every call is metered in your usage view for observability.
- Re-run the same tests whenever you change clients — tool discovery is automatic, but you should verify the connection after updates.
Two ways to connect — remote vs local
| Option | Endpoint / install | Tools | Best for |
|---|---|---|---|
| Remote (recommended) | https://plugsky.com/mcp + Bearer key | plugsky_chat, plugsky_models, plugsky_fusion | ChatGPT connectors, Claude, cloud IDEs — nothing to install |
| Local (stdio) | npx -y @plugsky/mcp + PLUGSKY_API_KEY | 9 tools in 6 groups (chat, models, rag, usage, tools, files) | Cursor/Claude Desktop power users who want RAG + document ingestion + platform tools |
Set up in 2 minutes
ChatGPT — Settings → Connectors → Add MCP server (developer mode): URL https://plugsky.com/mcp, authentication = your API key.
Claude / Cursor / Windsurf — add this to your MCP config (Claude Desktop: claude_desktop_config.json; Cursor: .cursor/mcp.json):
{
"mcpServers": {
"plugsky": {
"url": "https://plugsky.com/mcp",
"headers": { "Authorization": "Bearer sk-live-…" }
}
}
}
Local with all 9 tools (RAG, document ingestion, platform tools, usage):
{
"mcpServers": {
"plugsky": {
"command": "npx",
"args": ["-y", "@plugsky/mcp"],
"env": {
"PLUGSKY_API_KEY": "sk-live-…",
"PLUGSKY_EMAIL": "you@example.com",
"PLUGSKY_PASSWORD": "…"
}
}
}
}
Narrow what a client sees with tool groups — e.g. --groups=rag for a knowledge-only server, or --groups=chat,models to keep just the API-key tools. Create keys in the dashboard (API Keys); the MCP page has ready-made configs and copy buttons: Dashboard → AI → MCP Servers.
Live verification — stress-tested, not just documented
Before publishing, we ran the remote endpoint through a six-phase stress test on 29 September 2026. Results: 210/210 requests succeeded, with zero HTTP 5xx.
| Phase | Load | Result | Latency |
|---|---|---|---|
| tools/list | 60 requests, 20 concurrent | 60/60 OK | p50 0.8s · p95 1.3s |
| ping | 40 requests, 20 concurrent | 40/40 OK | p50 0.9s · p95 1.1s |
| plugsky_models | 30 requests, 15 concurrent | 30/30 OK | p50 1.1s · p95 1.7s |
| plugsky_chat (live inference) | 30 requests, 10 concurrent | 30/30 OK | p50 5.1s · p95 10.3s |
| unknown tool | 20 requests, 10 concurrent | 20/20 clean tool errors | p50 0.7s |
| unauthenticated burst | 30 requests, 15 concurrent | 30/30 → HTTP 401 | p50 0.6s |
The chat phase crosses into live inference on NVIDIA NIM with Plugsky's concurrency guard and automatic model cascade — 30 parallel chat calls, zero failures. On the same day, all 36 model aliases returned valid completions in a separate end-to-end sweep (fastest 75–90ms served responses; slowest worst-case 2.9s).
What teams do with it
- Code assistants with your docs — Cursor queries your RAG collection before suggesting changes, citing company documentation.
- Business copilots — ChatGPT connectors pull answers from uploaded policies and manuals (Arabic and English) instead of hallucinating.
- Agentic workflows — plugsky_fusion fans a high-stakes question out to several models and returns one merged answer.
- Document pipelines — ingest a PDF, image or scan by URL into a knowledge collection, then query it from any MCP client.
- Zero-cost classifications — route high-volume agent chatter to the free plugsky-micro / plugsky-lite models.
Security and limits
- Auth is a revocable API key per client; create separate keys for ChatGPT, Cursor and CI so you can cut one off without touching the others.
- The endpoint is stateless — no session data is persisted server-side beyond normal request logs and usage accounting.
- Plan limits are usage-fair: RPM and concurrency by tier; paid plans are unlimited under fair use (no daily/weekly token caps). 429 responses include
Retry-After. - Free models stay free for agents — the free plan supports plugsky-micro and plugsky-lite with a 30 RPM safety cap.
Frequently asked questions
What is MCP?
The Model Context Protocol is an open standard that lets AI apps call external tools and data through a single connector, instead of a bespoke integration per service.
What is the Plugsky MCP URL?
https://plugsky.com/mcp — remote Streamable HTTP. Authenticate with Authorization: Bearer sk-live-…. Open the URL in a browser to see the live server card.
How do I install the local package?
npx -y @plugsky/mcp — published on npm as @plugsky/mcp v1.1.0 with zero dependencies. Add --groups=chat,models,rag,usage,tools,files to pick tools.
Does it work with ChatGPT?
Yes — Settings → Connectors → Add MCP server (developer mode), enter the URL and your API key as the bearer token. Claude, Cursor, VS Code and Windsurf take the equivalent two-line JSON.
Is there a free tier for agents?
Yes — plugsky-micro and plugsky-lite are free, with tools, vision and JSON mode enabled. Paid plans unlock the full 36-model catalogue.
How is this different from other MCP servers?
Most MCP servers connect you to one service. Plugsky gives agents 36 models behind one key, a free inference tier, multi-model fusion, your own RAG collections and document ingestion — all from the same endpoint.
Where do I manage keys and usage?
In the dashboard: API Keys, Usage, and the MCP setup page at AI → MCP Servers.
Plugsky (2026). “MCP Guide 2026: Connect ChatGPT, Claude and Cursor to 36 AI Models”. Plugsky. Available at: https://plugsky.com/blog/mcp-model-context-protocol-guide-2026 (last updated 2026-09-29).