Agents

How do you control which tools an AI agent can use?

Tool permissions decide what an agent may actually do. Grant each agent only the tools it needs, scope keys to specific capabilities, validate arguments against JSON Schema before execution, enforce tenant and user scope server-side, rate-limit expensive or destructive tools, and require two-person approval for money-moving actions. Prompts describe intent; permissions are enforced in code the model cannot reach.

Key facts

Enforcement pointTool router or gateway, never the system prompt
Grant modelAllowlist per agent, deny by default, reviewed on a schedule
Key scopingOne scoped key per agent, rotatable and revocable
Argument policyJSON Schema validation plus value range checks before execution
Tenant scopeServer-side tenant and user scoping on every tool call
LimitsPer-tool rate limits, spend caps and blast-radius limits
Dual controlTwo-person approval for payments, deletes and permission changes
AuditEvery grant and every denied call is logged with a trace ID

TL;DR

  • Enforce permissions outside the model; prompts are not a security boundary.
  • Give each agent an allowlist and a scoped key, and deny everything else.
  • Validate arguments and check value limits before any tool runs.
  • Scope every call to the tenant and user server-side, not in the prompt.
  • Use two-person approval for irreversible or money-moving tools.

How it works, step by step

  1. List every tool the agent could use and group them by risk: read, write, spend, admin.
  2. Grant the read tools first and add write tools only when evaluation proves they are needed.
  3. Issue a scoped key per agent and enforce the allowlist in your tool router.
  4. Validate arguments with JSON Schema and range checks; reject anything unexpected.
  5. Apply tenant and user scoping inside each tool, based on the authenticated request context.
  6. Rate-limit expensive tools and cap total spend or volume per run and per day.
  7. Add two-person approval for the highest-risk tools, and log grants and denials.
1List every tool theagent could use andgroup them by risk:2Grant the readtools first and addwrite tools only3Issue a scoped keyper agent andenforce the4Validate argumentswith JSON Schemaand range checks;5Apply tenant anduser scoping insideeach tool, based on6Rate-limitexpensive tools andcap total spend or

Try it yourself

Open the tool registry builder →

Permissions live at the tool boundary

Anything written in a system prompt can be argued away by injected content or a clever prompt. Permissions must live where the model cannot reach them: in the router that receives a tool call and decides whether to execute it. That router knows the agent's identity, the granted tool list, the requesting user and the resource being touched, which is exactly the context needed to make an authorization decision.

A useful mental model is a reverse proxy for actions. The model proposes, the router authorizes, the tool executes, the audit log records. Separating those four responsibilities makes each one testable in isolation.

A permission model you can implement

  • Agent grants: a table mapping each agent to allowed tools and scopes.
  • Key scopes: credentials that can only call permitted capabilities, not every endpoint.
  • Argument policy: schema validation, allowed value ranges and required fields.
  • Resource scope: tenant, user and record-level checks inside the tool.
  • Budgets: call limits, spend limits and daily caps per agent.
  • Deny logging: every rejected call is an audit event, not a silent no-op.

High-risk tools need more than scopes

Some actions are not just permission problems. Payments, contract acceptance, production deletions and permission changes carry consequences that scopes cannot mitigate. For these, require a separated approval step: the agent prepares a proposal with the exact arguments, a human reviews it, and a second actor executes. Idempotency keys ensure the approval cannot be replayed twice.

Combine that with spending and volume limits so a compromised or confused agent cannot drain an account before someone notices. On Plugsky, scoped keys, RBAC and audit logging are live on the OpenAI-compatible API, and you can keep a cheap model on routine tool routing while a frontier model handles planning — 30+ models on one key. Plans, including the free plugsky-micro and plugsky-lite tier, are on the live pricing page.

Honest comparison

Tool riskExampleMinimum controlExtra control
Read public dataWeb search, docs lookupRate limitResult size caps
Read private dataCRM lookup, ticket readUser-scoped accessField-level redaction
Write dataUpdate record, send emailAllowlist plus approvalTwo-person rule
Spend moneyRefund, purchase, payoutApproval plus spend capDual control and reconciliation
Admin and deletePermission change, deletionDisabled by defaultBreak-glass with review

Frequently asked questions

Can I control tools with a system prompt?

You can describe intent in the prompt, but you cannot enforce it there. Enforce allowlists, scopes and validation in the tool router where injected content cannot change the rules.

What is a scoped API key?

A credential limited to specific capabilities or endpoints. Give each agent its own scoped key so you can revoke one workload without affecting others.

Should an agent ever have admin rights?

Almost never. Keep high-risk tools disabled by default and grant narrowly for specific, monitored workflows with approval and logging.

How do I prevent cross-tenant data leaks?

Enforce tenant scoping inside each tool using the authenticated request context, never a value the model supplies.

What rate limits make sense for tools?

Set per-tool limits based on the cost and impact of the action, then cap total calls and spend per run and per day.

Do I need approval for every write?

No. Classify writes by reversibility and impact. Reversible internal writes can run automatically; external, financial or destructive actions need a human.

Does Plugsky support scoped keys and RBAC?

Yes. Scoped keys, RBAC, SSO/SCIM and audit logging are live, with VPC, on-prem and air-gapped deployment options.