Key facts
| Install | pip install openai |
| Base URL | https://plugsky.com/v1 |
| Client | OpenAI(api_key=..., base_url=...) |
| Features | Chat, streaming, function calling, embeddings |
| Async | Async client supported through the same OpenAI library |
| Frameworks | LangChain, LlamaIndex, Haystack, AutoGen via base_url |
| Free plan | plugsky-micro and plugsky-lite, no card required |
| Product status | Live |
TL;DR
- pip install openai and change one constant: base_url.
- Streaming iterates delta content exactly as it does on OpenAI.
- Function calling uses the same tools schema; embeddings use the same call.
- Async clients and framework integrations work through the same override.
- Free models are enough to build and test before choosing a plan.
How it works, step by step
- Install the OpenAI SDK with pip install openai.
- Create the client with your Plugsky key and base_url set to api.plugsky.com/v1.
- Send a chat.completions.create call with a Plugsky model.
- Turn on stream=True and iterate chunks for incremental output.
- Define tools with JSON Schema for function calling.
- Call embeddings.create for vectors, and validate quality before production.
Try it yourself
Open the OpenAI-compatible API tester →
Install and first request
Install with pip install openai, then create the client with OpenAI(api_key="sk-live-...", base_url="https://plugsky.com/v1"). A chat call takes model and messages and returns the standard response object, so resp.choices[0].message.content is all you read for a basic application. Keep the key in an environment variable rather than source control.
Streaming responses
Set stream=True and iterate the returned object, printing chunk.choices[0].delta.content when it is present. The delta structure matches OpenAI, so existing streaming UIs and back-pressure handling port unchanged. For long answers, flush output as it arrives rather than buffering the whole completion.
Function calling
Define a tools list with JSON Schema parameters and pass it to chat.completions.create. The model returns structured tool calls instead of free text; your code executes the function, appends the result as a tool-role message and calls the model again. Parallel tool calls are supported, which reduces round trips when several functions are independent.
Embeddings, async and production notes
Embeddings use client.embeddings.create(model="plugsky-embed", input=...) and return the standard data array with vectors and usage. The same OpenAI library supports an async client, which is the better choice for high-concurrency services. When you deploy, pin the model explicitly rather than relying on a default, set sensible timeouts and retries, and watch the 429 Retry-After path under load.
For service accounts, set a per-key rate limit and tag requests with the user field so usage can be attributed per end user. Keep a fallback model configured in your client and log finish reasons to catch truncated answers. For RAG workloads, the RAG API wraps chunking and retrieval so you do not have to manage a vector store yourself.
Honest comparison
| Approach | OpenAI Python SDK + base_url | Plugsky RAG API | Direct HTTP |
|---|---|---|---|
| Setup | pip install, one constant | Same SDK, RAG endpoints | Custom client code |
| Streaming | Built-in | Via chat step | Manual SSE |
| Tool calling | Built-in | Built-in | Manual JSON |
| Embeddings | Built-in | Managed chunking and retrieval | Manual |
| Async support | Async client included | Included | You build it |
| Best for | Most Python services | Document Q&A | Minimal dependencies |
Frequently asked questions
Is there a dedicated Plugsky Python package?
No dedicated package is required. The official OpenAI Python SDK works with a base_url override, which keeps upgrades and type hints familiar.
Does streaming work exactly like OpenAI?
Yes. With stream=True you receive the same delta chunks, so existing streaming code works unchanged.
Can I use async Python?
Yes. The OpenAI library ships sync and async clients; point both at the Plugsky base URL.
How do I do function calling from Python?
Pass a tools array with JSON Schema definitions, execute the tool calls the model returns, then send results back as tool messages.
How do I use embeddings?
Call embeddings.create with model plugsky-embed and your input text; vectors come back in the standard data array.
Does it work with LangChain and LlamaIndex?
Yes. Both support custom OpenAI-compatible base URLs, as do Haystack, AutoGen and Semantic Kernel.
What does the free plan allow?
The free plan includes plugsky-micro and plugsky-lite with no card, and every new account starts with a 14-day full-access trial.
Plugsky (2026). “Plugsky Python SDK — Drop-in OpenAI Client”. Plugsky. Available at: https://plugsky.com/articles/python-sdk (last updated 2026-09-25).