Key facts
| Core tools | Alert read, log and metric query, runbook execution, draft comms |
| Default mode | Read-only triage and recommendation first |
| Reversible actions | Restarting stateless services behind rate limits |
| Approval | Production changes, DNS and data operations require a human |
| Tracing | One trace per incident linking signals, actions and outcome |
| Runbooks | Executable runbooks with idempotent steps and timeouts |
| Model routing | Small models for classification, frontier for root-cause reasoning |
| Status | Function calling, streaming and audit logs are live |
TL;DR
- Earn trust in stages: triage read-only, then reversible actions, then guarded changes.
- Wrap runbooks as idempotent tools with timeouts and clear preconditions.
- One trace per incident should link alert, signals, actions and outcome.
- Never let the agent touch DNS, production data or IAM without approval.
- Measure time to context and time to mitigation, not just actions taken.
How it works, step by step
- Start with alert triage: deduplicate, correlate and summarise with a recommended owner.
- Expose read-only tools for logs, metrics, deploys and status pages.
- Convert stable runbooks into idempotent tools with preconditions and timeouts.
- Allow only reversible actions with rate limits, starting with stateless restarts.
- Require human approval for production changes, DNS, data and permission operations.
- Trace every incident with one ID and capture the outcome for review.
- Run blameless post-incident reviews and fold learnings into tools and prompts.
Try it yourself
Open the agent workflow designer →
Start with triage, not remediation
Incident responders spend the first minutes gathering context: what changed, what is affected, is this a duplicate, who owns it. An agent that reads alerts, deploys, logs and metrics and produces a correlated summary with a probable owner is immediately valuable and carries almost no risk. That alone can cut time to context substantially.
Remediation comes later and in stages. Reversible actions on stateless components are the natural second step. Changes with lasting effect — schema migrations, DNS, access control, data deletion — stay human-only, with the agent preparing the change and the evidence.
Runbooks as tools
- Idempotent: running a step twice must not cause a second outage.
- Preconditions: each tool checks the state it requires before acting.
- Timeouts: bounded execution with a clear failure result the agent can reason about.
- Rate limits: caps per service and per window to prevent action storms.
- Dry run: where possible, report what would change before changing it.
- Traceability: every execution logged with the incident ID and actor.
Tools written this way are reusable by humans too, which improves the runbooks for everyone and keeps automation honest.
Incident trust and post-incident review
Trust is built from evidence. During incidents, the agent should state what it observed, what it did, what changed and what remains uncertain, with links to the underlying signals. On-call engineers trust an agent that shows its work far more than one that declares success.
Afterwards, measure time to context, time to mitigation, actions taken versus actions needed, and false escalations. Feed recurring patterns into new runbook tools and evaluation cases. On Plugsky, function calling, streaming and audit logging are live, and 30+ models on one key let classification run on a small model while root-cause reasoning uses a frontier one. Chat, embeddings and RAG are live; batch endpoints are coming soon. Plans are on the live pricing page.
Honest comparison
| Stage | Agent capability | Risk | Requires approval |
|---|---|---|---|
| Triage and summary | Correlate alerts and logs | Low | No |
| Diagnosis | Query metrics and propose causes | Low | No |
| Reversible action | Restart stateless service | Medium | Rate-limited automatic |
| Configuration change | Deploy, scale, rotate credentials | High | Yes |
| Data or DNS change | Migrations, records, access | Critical | Two-person approval |
Frequently asked questions
Can the agent restart production services?
It can restart stateless services with rate limits once it has earned trust. Stateful services, databases and anything with data-loss risk should remain behind human approval.
How do I stop action storms?
Rate-limit each tool per service and per time window, require idempotency, and cap the number of actions per incident. Alert on repeated identical actions.
What should the agent never touch?
DNS, identity and access rules, production data deletion and schema migrations. These have blast radius beyond the incident and belong to humans.
How does it help if it cannot fix anything?
Most incident cost is context gathering and coordination. A fast, accurate summary with a likely owner and related changes saves responder time with negligible risk.
How do we evaluate an ops agent?
Replay historical incidents, measure how quickly it produces an accurate timeline and cause, and track false escalations. Review results with the on-call team.
What about noisy alerts?
Use the agent to deduplicate and correlate first. Reducing noise improves the agent's signal and the team's response at the same time.
Can it run in a restricted network?
Yes. Plugsky supports VPC, on-prem and air-gapped deployment, so ops tooling and telemetry stay inside your network boundary.