Key facts
| Endpoint | POST /v1/chat/completions on the live agent API |
| SDK | OpenAI Python SDK with a base_url override, or httpx/requests |
| Auth | Bearer API key in the Authorization header; scoped keys supported |
| Tools | JSON Schema function definitions passed in the tools array |
| Streaming | Token streaming is live for responsive interfaces |
| Structured output | JSON mode is live for typed agent results |
| Free plan | plugsky-micro and plugsky-lite, no credit card |
| Roadmap | Assistants-style managed endpoints are coming soon |
TL;DR
- Install openai, set base_url to Plugsky and keep your existing client code.
- Describe tools as JSON Schema functions; the model decides when to call them.
- Execute tool_calls in your code, append results, loop until the model answers.
- Stream tokens for perceived speed and cap turns to control cost.
- Prototype on plugsky-micro and plugsky-lite free, then test paid models in the trial.
How it works, step by step
- Install the SDK with pip install openai and store your key in an environment variable.
- Create the client with OpenAI(api_key=..., base_url="https://api.plugsky.com/v1").
- Describe each tool with a name, a model-facing description and JSON Schema parameters.
- Send a chat completion with messages and tools, then read message.tool_calls.
- Execute each tool locally, append a tool message with the result and call the endpoint again.
- Stop when the model returns a normal answer; stream tokens as they arrive.
- Log model, tool names, arguments and latency per turn so failures are debuggable.
Try it yourself
Open the OpenAI-compatible API tester →
Point the OpenAI SDK at Plugsky
The migration is a base URL change, so the quickstart is short. Install the official SDK, read your key from the environment and set the base URL to Plugsky. Everything after that — message shapes, tool calling, streaming, JSON mode — follows the OpenAI convention you already know.
from openai import OpenAIclient = OpenAI(api_key=os.environ["PLUGSKY_API_KEY"], base_url="https://api.plugsky.com/v1")
Keep the key out of source control. Agents should use their own scoped key so you can rotate or revoke one workload without touching the rest of your platform.
The tool-calling loop in Python
Define tools as functions with typed parameters. Descriptions are prompts: state what the tool does, when to use it and what it returns. Send them in the tools array of a chat completion request. If the model returns tool_calls, execute each call, append one tool message per call with the result, and call the endpoint again. Repeat until finish_reason is stop or you hit your turn cap.
- Validate arguments against the schema before executing anything.
- Return trimmed results — the five useful fields, not a full database row.
- Cap turns and set per-tool timeouts so one slow dependency cannot stall the run.
- Catch tool errors and return them as text so the model can adapt instead of crashing the loop.
Hardening the quickstart
Three additions turn a demo into something you can ship. First, memory: keep recent turns in the message history, and write durable facts to a vector store for retrieval on later runs. Second, approvals: any tool that writes, sends or pays should pause the loop and surface a confirmation instead of executing. Third, evaluation: keep a frozen set of tasks with expected tool calls and run it whenever you change a prompt or switch models.
Plugsky exposes 30+ models behind the same endpoint, so you can route routine steps to plugsky-micro or plugsky-lite and reserve frontier models for planning. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. Costs and plans are on the live pricing page.
Honest comparison
| Step | Python pattern | Why it matters | Common mistake |
|---|---|---|---|
| Client | OpenAI(base_url=...) with env key | Keeps SDK code and secrets out of git | Hard-coding api.openai.com |
| Tools | JSON Schema function definitions | The model can call tools reliably | Vague or missing descriptions |
| Loop | Append tool results and call again | Completes multi-step tasks | Returning tool output to the user directly |
| Stop | finish_reason stop or turn cap | Controls cost and latency | Unbounded loops |
| Observability | Log tool calls and latency per turn | Debug failures and audit actions | No traces beyond print statements |
Frequently asked questions
Do I need a special agent SDK?
No. Use the OpenAI Python SDK and change the base URL to Plugsky. Request and response shapes stay the same, so existing code keeps working.
Where should the API key live?
In an environment variable backed by a secret manager. Give each agent a scoped key so you can revoke one workload without affecting others.
How do I define a good tool?
Give it a precise name, a description that states when to use it, and JSON Schema parameters. The description is the prompt that drives correct calls.
Is streaming supported?
Yes. Streaming is live, which matters because tool execution adds latency between turns and users should see progress.
What is JSON mode for?
It constrains output to valid JSON, which is useful when the agent must return typed data your Python code parses and stores.
Can I test before writing code?
Yes. Open the OpenAI-compatible API tester to send requests from the browser and confirm keys, models and tool schemas first.
What does it cost to try?
plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers paid models. Plans are listed on the live pricing page.