Key facts
| Core record | Run ID, identity, model, prompt hash, tool calls, arguments, results, outcome |
| Approval trail | Who approved which action, when and from where |
| Redaction | Strip credentials, secrets and personal data before logs leave the boundary |
| Correlation | One trace ID across model calls, tools and downstream services |
| Retention | Retention and access rules defined per data class |
| Export | Stream logs to your SIEM or warehouse for correlation |
| Deployment | Region-locked logging with VPC, on-prem and air-gapped options |
| Status | Audit logging for API calls is live |
TL;DR
- Log the run, not just the answer: every model call, tool call and approval decision.
- Use one trace ID so agent, tools and downstream systems correlate.
- Redact secrets and personal data before logs are stored or exported.
- Define retention and read access per data class, and ship logs to your SIEM.
- Approval decisions are audit events: capture who approved what, when and why.
How it works, step by step
- Define what one agent run must prove: inputs, tools called, arguments, outputs and approvals.
- Emit a trace ID at run start and propagate it to every model call, tool and service.
- Record each tool call with name, validated arguments, result status and latency.
- Redact credentials, tokens and personal data before writing the log entry.
- Capture approval events with approver identity, timestamp and the decision.
- Export logs to your SIEM or warehouse and set retention per data class.
- Review traces after incidents and turn findings into regression tests.
Try it yourself
Open the API key security checklist →
What belongs in an agent audit log
Start from the questions you will ask after something goes wrong: which run, triggered by whom, using which model and prompt, calling which tools with which arguments, producing what result, and which human approved the irreversible step. If a field answers one of those questions, log it. If it contains a credential or unnecessary personal data, keep it out or hash it.
- Identity: end user, service account and the scoped key used.
- Model: model name and version, plus a prompt template hash so you know which revision ran.
- Tools: name, arguments, result status, duration and retry count.
- Decisions: routing choice, guardrail result and approval outcome.
- Outcome: final answer reference, cost estimate and total turns.
Making logs tamper-resistant and useful
Audit logs are only evidence if they cannot be quietly edited. Write append-only storage, restrict access with RBAC, and export copies to a system the agent team does not administer. Correlate everything with a single trace ID so a model call, its tool calls and the downstream database write appear as one story rather than four unrelated entries.
Volume is the practical enemy. Logging full prompts and results for every turn is expensive and privacy-heavy. A workable pattern is to log hashes and references by default, then capture full payloads only for sampled runs, errors and workflows that touch money, health or regulated data.
Compliance mapping without theater
Regulators and auditors increasingly ask how autonomous systems are controlled, not whether logs exist. Map each control to concrete evidence: least privilege to scoped keys and RBAC, human oversight to approval events, data governance to redaction and residency, reliability to retry and failure records. When a review asks what the agent did on a specific date, one trace should answer it.
Plugsky provides scoped keys, RBAC, SSO/SCIM and per-action audit logging on the live API, and region-locked deployment options — cloud, VPC, on-prem and air-gapped — so log data stays in the jurisdiction you choose. See the live pricing page for plans and the docs for logging details.
Honest comparison
| Concern | Minimal logging | Production agent logging | Regulated audit trail |
|---|---|---|---|
| Identity | Shared API key | Scoped key per agent | User and agent identity both |
| Tool calls | Errors only | Name, args, status, latency | Full payload with redaction |
| Approvals | Not captured | Approver and timestamp | Signed decision record |
| Retention | Whatever the log tool keeps | Policy per data class | Defined retention with legal hold |
| Access | Team dashboard | RBAC-scoped viewers | External append-only archive |
Frequently asked questions
Do I need to log every model call?
You need enough to reconstruct what happened: model and version, the decisions made, tools called and results. Hash full prompts by default and store payloads for sampled or high-risk runs.
What personal data can appear in agent logs?
User identifiers, message content and tool arguments can all contain personal data. Minimise, redact or hash before storage, and apply the same retention rules as your other systems.
How do I prove which prompt version ran?
Store a hash of the prompt template and system instructions with each run, and version prompts in your repository.
Where should agent logs live?
In append-only storage your team does not solely administer, exported to a SIEM or warehouse. On Plugsky, region-locked deployment keeps logs in your chosen jurisdiction.
Are approval decisions loggable?
Yes, and they should be. Capture the approver identity, timestamp, action details and outcome as first-class events, not chat history.
Does logging add latency?
Write asynchronously or batch log shipping so the agent loop is not blocked by the logging path.
Is audit logging available on Plugsky today?
Yes, API call logging is live with scoped keys, RBAC and SSO/SCIM; deployment options cover cloud, VPC, on-prem and air-gapped environments.