Agents

What is an AI agent API?

An AI agent API is a model endpoint designed for loops rather than single replies: it supports function calling so the model can request actions, streaming so you can observe progress, JSON mode for machine-readable output, and embeddings for memory and retrieval. Your code runs the loop, executes the tools and enforces limits; the API supplies reasoning, structure and state-free execution.

Key facts

Function callingStructured tool requests with arguments your code validates and executes
StreamingToken-by-token output so long runs show progress and can be cancelled
JSON modeConstrained output for parsers, routers and state updates
Embeddings and RAGVector representation plus retrieval grounding for memory
AgentsAgent patterns run on the chat and tool-calling endpoints
Platform controlsScoped keys, RBAC, SSO/SCIM and audit logs are live
Pricing modelFlat monthly self-serve plans; free tier with two models
Coming soonAssistants, responses, batch, files, fine-tuning, audio, image and moderation endpoints

TL;DR

  • An agent API supports loops: tools, streaming, JSON mode and retrieval.
  • Your code owns the loop, state, tool execution and guardrails.
  • OpenAI compatibility keeps clients, frameworks and skills portable.
  • Check status labels: some assistant-style endpoints are still coming soon.
  • Start free with plugsky-micro and plugsky-lite, then scale.

How it works, step by step

  1. Confirm the endpoint supports function calling, streaming and JSON mode.
  2. Define tools with strict argument schemas and validate every value before execution.
  3. Implement the loop: call, execute tools, append results, repeat until done.
  4. Add limits: maximum turns, token budget, retry ceiling and timeout.
  5. Store state in your own database or a checkpointer you control.
  6. Add embeddings and retrieval for memory and grounding where needed.
  7. Log every step for evaluation, cost tracking and incident review.
1Confirm theendpoint supportsfunction calling,2Define tools withstrict argumentschemas and3Implement the loop:call, executetools, append4Add limits: maximumturns, tokenbudget, retry5Store state in yourown database or acheckpointer you6Add embeddings andretrieval formemory and

Try it yourself

Open the AI agent comparison →

The pieces that make an API agent-ready

A plain completions endpoint returns text. Agent work needs more. Function calling lets the model emit a structured request — a tool name and arguments — that your code executes and feeds back. Streaming exposes tokens as they arrive, which matters for long runs, progress display and cancellation. JSON mode constrains output so downstream parsers and routers can rely on the shape.

Retrieval completes the picture: embeddings turn text into vectors so the agent can search memory, documents and past conversations instead of relying on a growing prompt. Together these capabilities let you build a loop where the model reasons and your code acts.

The loop is still your responsibility

An agent API does not run your agent. It serves model calls. The loop, the tools, the state, the retries and the guardrails live in your code, which is what makes behaviour inspectable and portable. That division matters: the API provides reasoning and structure; you provide action and accountability.

  • Tool execution: validate arguments, enforce permissions, sandbox side effects.
  • State: persist threads and checkpoints where your policies require.
  • Limits: caps on turns, tokens, retries and runtime.
  • Observability: traces of every call, tool result and decision.

How to compare agent APIs

Compare on four axes. Interface: is it OpenAI-compatible, so your SDKs and frameworks keep working? Capability: are function calling, streaming, JSON mode, embeddings and RAG live, and which assistant-style endpoints are still coming soon? Platform: do you get scoped keys, RBAC, audit logs and deployment choices? Economics: is pricing predictable for long-running, multi-step work?

Plugsky answers these with 30+ models on one OpenAI-compatible key, live function calling, streaming, JSON mode, embeddings, RAG and agents, plus scoped keys, RBAC, SSO/SCIM and audit logs, and deployment from shared cloud to VPC, on-prem and air-gapped. Assistants, responses, batch, files, fine-tuning, audio, image and moderation endpoints are coming soon. Self-serve plans are flat monthly and the free tier covers plugsky-micro and plugsky-lite; current plans are on the live pricing page.

Honest comparison

CapabilityPlain chat endpointAgent-ready APIHosted agent product
Function callingNoneLive, structured tool requestsHandled internally
StreamingSometimesLive token streamingYes
RetrievalYou add separatelyEmbeddings and RAG liveBuilt in
Loop controlYou build itYou build itVendor controls it
PortabilityHighHigh with OpenAI compatibilityLow

Frequently asked questions

Is an agent API the same as an assistant API?

No. Assistant-style endpoints typically store threads and files server-side. An agent API exposes model capabilities such as tool calling, and you keep state and orchestration in your code.

Do I need a separate API for agents?

No. Agents run on standard chat completion calls with function calling, streaming and JSON mode. The difference is how you use the endpoint, not a special model class.

What does OpenAI compatibility mean here?

Your existing OpenAI-format SDKs, frameworks and tooling keep working after a base URL change, which lowers migration effort and exit cost.

Where does memory live?

In your systems. Use embeddings and a vector store for semantic recall, and your own database for structured threads and checkpoints.

Which endpoints are not agent-ready yet?

On Plugsky, assistants, responses, batch, files, fine-tuning, audio, image and moderation endpoints are coming soon. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live.

How do I keep an agent safe?

Validate tool arguments, run tools with least privilege, sandbox side effects, cap turns and budget, and require approval for irreversible actions.

How do I start?

Create an API key on the free plan with plugsky-micro and plugsky-lite, run a two-tool loop, then use the 14-day full-access trial to evaluate the wider catalogue.