Developer + API

How do you use the Plugsky API from Python?

Install the official OpenAI Python SDK, set base_url to https://plugsky.com/v1, and your existing code works. Chat, streaming, function calling and embeddings all use the OpenAI request shapes, so migration is a base URL and model-name change rather than a rewrite. Start on the free plan with plugsky-micro and plugsky-lite, no card required.

Key facts

Installpip install openai
Base URLhttps://plugsky.com/v1
ClientOpenAI(api_key=..., base_url=...)
FeaturesChat, streaming, function calling, embeddings
AsyncAsync client supported through the same OpenAI library
FrameworksLangChain, LlamaIndex, Haystack, AutoGen via base_url
Free planplugsky-micro and plugsky-lite, no card required
Product statusLive

TL;DR

  • pip install openai and change one constant: base_url.
  • Streaming iterates delta content exactly as it does on OpenAI.
  • Function calling uses the same tools schema; embeddings use the same call.
  • Async clients and framework integrations work through the same override.
  • Free models are enough to build and test before choosing a plan.

How it works, step by step

  1. Install the OpenAI SDK with pip install openai.
  2. Create the client with your Plugsky key and base_url set to api.plugsky.com/v1.
  3. Send a chat.completions.create call with a Plugsky model.
  4. Turn on stream=True and iterate chunks for incremental output.
  5. Define tools with JSON Schema for function calling.
  6. Call embeddings.create for vectors, and validate quality before production.
1Install the OpenAISDK with pipinstall openai.2Create the clientwith your Plugskykey and base_url3Send achat.completions.createcall with a Plugsky4Turn on stream=Trueand iterate chunksfor incremental5Define tools withJSON Schema forfunction calling.6Callembeddings.createfor vectors, and

Try it yourself

Open the OpenAI-compatible API tester →

Install and first request

Install with pip install openai, then create the client with OpenAI(api_key="sk-live-...", base_url="https://plugsky.com/v1"). A chat call takes model and messages and returns the standard response object, so resp.choices[0].message.content is all you read for a basic application. Keep the key in an environment variable rather than source control.

Streaming responses

Set stream=True and iterate the returned object, printing chunk.choices[0].delta.content when it is present. The delta structure matches OpenAI, so existing streaming UIs and back-pressure handling port unchanged. For long answers, flush output as it arrives rather than buffering the whole completion.

Function calling

Define a tools list with JSON Schema parameters and pass it to chat.completions.create. The model returns structured tool calls instead of free text; your code executes the function, appends the result as a tool-role message and calls the model again. Parallel tool calls are supported, which reduces round trips when several functions are independent.

Embeddings, async and production notes

Embeddings use client.embeddings.create(model="plugsky-embed", input=...) and return the standard data array with vectors and usage. The same OpenAI library supports an async client, which is the better choice for high-concurrency services. When you deploy, pin the model explicitly rather than relying on a default, set sensible timeouts and retries, and watch the 429 Retry-After path under load.

For service accounts, set a per-key rate limit and tag requests with the user field so usage can be attributed per end user. Keep a fallback model configured in your client and log finish reasons to catch truncated answers. For RAG workloads, the RAG API wraps chunking and retrieval so you do not have to manage a vector store yourself.

Honest comparison

ApproachOpenAI Python SDK + base_urlPlugsky RAG APIDirect HTTP
Setuppip install, one constantSame SDK, RAG endpointsCustom client code
StreamingBuilt-inVia chat stepManual SSE
Tool callingBuilt-inBuilt-inManual JSON
EmbeddingsBuilt-inManaged chunking and retrievalManual
Async supportAsync client includedIncludedYou build it
Best forMost Python servicesDocument Q&AMinimal dependencies

Frequently asked questions

Is there a dedicated Plugsky Python package?

No dedicated package is required. The official OpenAI Python SDK works with a base_url override, which keeps upgrades and type hints familiar.

Does streaming work exactly like OpenAI?

Yes. With stream=True you receive the same delta chunks, so existing streaming code works unchanged.

Can I use async Python?

Yes. The OpenAI library ships sync and async clients; point both at the Plugsky base URL.

How do I do function calling from Python?

Pass a tools array with JSON Schema definitions, execute the tool calls the model returns, then send results back as tool messages.

How do I use embeddings?

Call embeddings.create with model plugsky-embed and your input text; vectors come back in the standard data array.

Does it work with LangChain and LlamaIndex?

Yes. Both support custom OpenAI-compatible base URLs, as do Haystack, AutoGen and Semantic Kernel.

What does the free plan allow?

The free plan includes plugsky-micro and plugsky-lite with no card, and every new account starts with a 14-day full-access trial.

Cite this page

Plugsky (2026). “Plugsky Python SDK — Drop-in OpenAI Client”. Plugsky. Available at: https://plugsky.com/articles/python-sdk (last updated 2026-09-25).