Key facts
| Keys | Scoped API keys per agent, with rotation and revocation |
| Tool access | Allowlist tools per agent; deny everything else by default |
| Input trust | Retrieved documents and web pages are data, never instructions |
| Execution | Sandbox code and browser tools with no ambient credentials |
| Writes | Human approval before irreversible or money-moving actions |
| Egress | Restrict outbound network access from tool processes |
| Audit | Per-action logs with trace IDs, exported to a SIEM |
| Platform | RBAC, SSO/SCIM and region-locked deployment are live |
TL;DR
- Treat the agent as a privileged user: least privilege, scoped keys, narrow tools.
- Assume retrieved content is hostile; it is data, not instructions.
- Sandbox code execution and browser tools, and restrict their network egress.
- Gate irreversible actions behind human approval and idempotency keys.
- Log every action and review traces after incidents, not only after audits.
How it works, step by step
- Inventory every tool and system the agent can touch, and cut the list to the minimum it needs.
- Issue a scoped key per agent and enforce allowlists at the tool boundary, not in the prompt.
- Validate all model-generated arguments against schemas before executing anything.
- Run code, shell and browser tools in a sandbox without ambient credentials or open egress.
- Add prompt-injection defences: separate trusted instructions from untrusted content and re-check policy after retrieval.
- Require human approval for writes, payments and deletions, with idempotency keys on execution.
- Log actions with trace IDs, export to your SIEM, and rehearse an incident response for agents.
Try it yourself
Open the API key security checklist →
The threat model: an agent is a privileged user
Most agent incidents are not exotic. They are ordinary access-control failures with a generative twist. An agent holds credentials, reads untrusted content and writes to real systems, so an attacker who can influence its input may influence its actions. Treat it exactly like a service account: enumerate what it can reach, grant the minimum, and assume its inputs can be adversarial.
The two failure classes that matter most are prompt injection, where instructions hide in retrieved documents or web pages, and excessive agency, where the agent simply has more permissions and tools than the task requires. Both are design problems, and both are cheaper to prevent than to detect after the fact.
Controls that matter most
- Scoped keys: one key per agent, rotated on schedule, revoked when the workload retires.
- Allowlists: the tool router rejects any tool not explicitly granted to that agent.
- Argument validation: schema checks before execution; never trust model output.
- Sandboxing: code, shell and browser tools run isolated with no ambient secrets and limited egress.
- Approval gates: writes, payments, deletes and permission changes pause for a human.
- Rate and blast-radius limits: caps per tool, per run and per day.
- Audit logging: one trace ID with tool arguments, decisions and outcomes.
Operating security after launch
Security is a loop, not a checklist you complete. Review tool grants quarterly, remove tools nobody calls, rotate keys, and re-run injection tests against the retrieval sources the agent can reach. Watch traces for anomalies: a sudden spike in a rarely used tool, repeated permission errors, or tool calls with unusual argument patterns are all signals worth investigating.
On Plugsky, scoped keys, RBAC, SSO/SCIM and audit logging are live, and deployment options span the shared cloud, your VPC, on-prem and air-gapped environments with region choice, so the security boundary matches your data policy. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. See the live pricing page for plans.
Honest comparison
| Control | Prototype | Production | Regulated |
|---|---|---|---|
| API keys | One shared key | Scoped key per agent | Scoped keys plus rotation policy |
| Tools | All tools exposed | Allowlisted per agent | Allowlist plus approval per tool |
| Execution | Direct host access | Sandboxed tool processes | Sandbox plus egress controls |
| Writes | Automatic | Human approval for risky actions | Two-person approval and limits |
| Logging | Console prints | Trace IDs exported to SIEM | Immutable retention with review |
Frequently asked questions
What is prompt injection?
It is an attack where instructions are hidden in content the agent reads, such as a document or web page, trying to redirect its behaviour. Defend by treating retrieved content as data, isolating instructions, and validating every action at the tool boundary.
Should the agent use my admin credentials?
No. Give it a scoped identity with the minimum permissions it needs, and let it act as the requesting user where possible so existing access rules apply.
How do I sandbox code execution?
Run generated code in a short-lived isolated container with no host mounts, no ambient credentials, restricted network egress and hard resource limits.
Do allowlists belong in the prompt?
No. The prompt can be overridden by injected content. Enforce permissions in the tool router or gateway where the model cannot reach them.
What should trigger a human approval step?
Payments, contract changes, deletions, permission changes, outbound messages to customers and anything legally or financially material.
How do I test agent security?
Run red-team prompts against your real retrieval sources and tool set, then verify that policy enforcement happens outside the model. Re-run after every change to prompts, tools or data sources.
Does Plugsky provide security controls?
Yes. Scoped keys, RBAC, SSO/SCIM and audit logging are live, with VPC, on-prem and air-gapped deployment for stricter boundaries.