Key facts
| Perception | Screenshots plus accessibility tree or DOM extraction |
| Actions | Click, type, scroll, drag, keyboard shortcuts and clipboard |
| Isolation | Dedicated VM or container with snapshots and no host mounts |
| Credentials | Least-privilege accounts; never ambient production access |
| Speed | Seconds per step, far slower than a direct API call |
| Cost profile | Screenshots and page state consume many tokens per step |
| Model layer | 30+ models on one OpenAI-compatible key with live function calling |
| Roadmap | Image and audio endpoints are coming soon, not live yet |
TL;DR
- Use computer use only when no API or export path exists.
- Run the agent in a disposable VM, never on a real desktop.
- Gate irreversible actions behind human approval.
- Screenshots are token-expensive; cap steps and resolution.
- Log every action with screenshots so failures are reproducible.
How it works, step by step
- Confirm an API, database export or bulk import path does not exist for the task.
- Provision a disposable VM with its own browser profile, disk snapshot and network rules.
- Create a dedicated least-privilege account for the target system.
- Expose a small action set — click, type, scroll, wait — with argument validation.
- Add a step cap, a wall-clock timeout and a human approval gate for irreversible actions.
- Capture screenshots and action logs on every step for replay and debugging.
- Rebuild the task as an API integration as soon as the vendor offers one.
Try it yourself
Open the browser automation generator →
How a computer use agent works
The loop is perceive, decide, act, verify. Perception combines a screenshot with structured context from the accessibility tree or DOM. The model decides on an action, the harness executes it, and the next screenshot confirms the result. Because the agent must look again after every action, tasks take many steps and each step carries an image into context, which is why computer use is the most token-hungry agent pattern.
Reliability comes from grounding actions in stable identifiers where possible. Clicking by element ID or accessible label beats clicking by coordinates, and explicit waits beat fixed sleeps. Where the interface offers keyboard shortcuts or forms, prefer them over pixel hunting.
Sandbox design is the whole safety story
Treat the agent as an untrusted user with a keyboard. Give it a dedicated VM or container with a fresh profile, no shared clipboard, no host filesystem mounts, and outbound network rules limited to the target system. Use a least-privilege account with no admin rights and no access to unrelated data.
- Snapshots: restore a clean state after every run.
- Approval gates: payments, deletions, sends and permission changes pause for a person.
- Rate limits: cap actions per minute so a loop cannot hammer a system.
- Kill switch: one control stops the run and revokes the session.
- Audit: store screenshots, actions and model reasoning for every step.
Cost, reliability and when to walk away
Computer use is the fallback, not the default. A direct API call is faster, cheaper and testable; GUI automation breaks whenever a layout changes. Budget for step caps and image downscaling, and measure success by completed tasks per attempt rather than steps taken.
Plugsky supplies the model layer behind such agents: 30+ models on one OpenAI-compatible key with live function calling, streaming and JSON mode, plus scoped keys and audit logs. It does not provide a desktop sandbox or a browser runtime, so those are yours to operate; image and audio endpoints are coming soon rather than live. Match the model to the job — vision-capable models for screen reading, cheaper models for repetitive confirmation steps — and check current plans and the free plugsky-micro and plugsky-lite models on the live pricing page.
Honest comparison
| Approach | Direct API integration | Structured export or import | Computer use agent |
|---|---|---|---|
| Reliability | Highest, contract-based | High if schema is stable | Breaks on UI changes |
| Speed | Milliseconds to seconds | Minutes per batch | Seconds per screen step |
| Cost | Lowest | Low | Highest, image tokens per step |
| Maintenance | Versioned API | Format drift | Continuous, UI-dependent |
| Fit | Preferred first choice | Good for bulk data | Last resort when no API exists |
Frequently asked questions
When is computer use the right choice?
When the target system has no API, no export and no integration path, and the task value justifies the fragility. If any programmatic route exists, use it instead.
Is it safe to let an agent control a browser?
Only inside a disposable sandbox with least-privilege credentials, network restrictions, action caps and approval gates for irreversible steps. Never point it at your daily desktop.
Why are computer use agents slow?
Every action requires a fresh screenshot and a model decision, so a task with twenty steps means twenty perception and reasoning cycles rather than one API call.
How do I reduce token cost?
Downscale screenshots, crop to the relevant region, use accessibility text where available, cache unchanged screens, and route simple confirmation steps to cheaper models.
Which models support screen understanding?
Vision-capable models can read screenshots; check the live model catalogue for current vision support, since coverage changes as models are added. Image generation endpoints are a separate, coming-soon capability.
How should I log a run?
Store each step's screenshot, chosen action, model output and result together, so a failed run can be replayed and the exact divergence point identified.
Does Plugsky provide the sandbox?
No. Plugsky provides the model API and function calling. The VM, browser profile, action executor and approval flow are infrastructure you build or buy.