Agents

What should be on an AI agent security checklist?

Secure an agent like a privileged service account with a mind. Start with scoped keys and least-privilege tools, allowlist what each agent may call, validate every argument, sandbox code and browser execution, treat retrieved content as untrusted data rather than instructions, gate writes behind human approval, restrict egress, and log every action with a trace ID.

Key facts

KeysScoped API keys per agent, with rotation and revocation
Tool accessAllowlist tools per agent; deny everything else by default
Input trustRetrieved documents and web pages are data, never instructions
ExecutionSandbox code and browser tools with no ambient credentials
WritesHuman approval before irreversible or money-moving actions
EgressRestrict outbound network access from tool processes
AuditPer-action logs with trace IDs, exported to a SIEM
PlatformRBAC, SSO/SCIM and region-locked deployment are live

TL;DR

  • Treat the agent as a privileged user: least privilege, scoped keys, narrow tools.
  • Assume retrieved content is hostile; it is data, not instructions.
  • Sandbox code execution and browser tools, and restrict their network egress.
  • Gate irreversible actions behind human approval and idempotency keys.
  • Log every action and review traces after incidents, not only after audits.

How it works, step by step

  1. Inventory every tool and system the agent can touch, and cut the list to the minimum it needs.
  2. Issue a scoped key per agent and enforce allowlists at the tool boundary, not in the prompt.
  3. Validate all model-generated arguments against schemas before executing anything.
  4. Run code, shell and browser tools in a sandbox without ambient credentials or open egress.
  5. Add prompt-injection defences: separate trusted instructions from untrusted content and re-check policy after retrieval.
  6. Require human approval for writes, payments and deletions, with idempotency keys on execution.
  7. Log actions with trace IDs, export to your SIEM, and rehearse an incident response for agents.
1Inventory everytool and system theagent can touch,2Issue a scoped keyper agent andenforce allowlists3Validate allmodel-generatedarguments against4Run code, shell andbrowser tools in asandbox without5Addprompt-injectiondefences: separate6Require humanapproval forwrites, payments

Try it yourself

Open the API key security checklist →

The threat model: an agent is a privileged user

Most agent incidents are not exotic. They are ordinary access-control failures with a generative twist. An agent holds credentials, reads untrusted content and writes to real systems, so an attacker who can influence its input may influence its actions. Treat it exactly like a service account: enumerate what it can reach, grant the minimum, and assume its inputs can be adversarial.

The two failure classes that matter most are prompt injection, where instructions hide in retrieved documents or web pages, and excessive agency, where the agent simply has more permissions and tools than the task requires. Both are design problems, and both are cheaper to prevent than to detect after the fact.

Controls that matter most

  • Scoped keys: one key per agent, rotated on schedule, revoked when the workload retires.
  • Allowlists: the tool router rejects any tool not explicitly granted to that agent.
  • Argument validation: schema checks before execution; never trust model output.
  • Sandboxing: code, shell and browser tools run isolated with no ambient secrets and limited egress.
  • Approval gates: writes, payments, deletes and permission changes pause for a human.
  • Rate and blast-radius limits: caps per tool, per run and per day.
  • Audit logging: one trace ID with tool arguments, decisions and outcomes.

Operating security after launch

Security is a loop, not a checklist you complete. Review tool grants quarterly, remove tools nobody calls, rotate keys, and re-run injection tests against the retrieval sources the agent can reach. Watch traces for anomalies: a sudden spike in a rarely used tool, repeated permission errors, or tool calls with unusual argument patterns are all signals worth investigating.

On Plugsky, scoped keys, RBAC, SSO/SCIM and audit logging are live, and deployment options span the shared cloud, your VPC, on-prem and air-gapped environments with region choice, so the security boundary matches your data policy. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. See the live pricing page for plans.

Honest comparison

ControlPrototypeProductionRegulated
API keysOne shared keyScoped key per agentScoped keys plus rotation policy
ToolsAll tools exposedAllowlisted per agentAllowlist plus approval per tool
ExecutionDirect host accessSandboxed tool processesSandbox plus egress controls
WritesAutomaticHuman approval for risky actionsTwo-person approval and limits
LoggingConsole printsTrace IDs exported to SIEMImmutable retention with review

Frequently asked questions

What is prompt injection?

It is an attack where instructions are hidden in content the agent reads, such as a document or web page, trying to redirect its behaviour. Defend by treating retrieved content as data, isolating instructions, and validating every action at the tool boundary.

Should the agent use my admin credentials?

No. Give it a scoped identity with the minimum permissions it needs, and let it act as the requesting user where possible so existing access rules apply.

How do I sandbox code execution?

Run generated code in a short-lived isolated container with no host mounts, no ambient credentials, restricted network egress and hard resource limits.

Do allowlists belong in the prompt?

No. The prompt can be overridden by injected content. Enforce permissions in the tool router or gateway where the model cannot reach them.

What should trigger a human approval step?

Payments, contract changes, deletions, permission changes, outbound messages to customers and anything legally or financially material.

How do I test agent security?

Run red-team prompts against your real retrieval sources and tool set, then verify that policy enforcement happens outside the model. Re-run after every change to prompts, tools or data sources.

Does Plugsky provide security controls?

Yes. Scoped keys, RBAC, SSO/SCIM and audit logging are live, with VPC, on-prem and air-gapped deployment for stricter boundaries.