Agents

What are computer use AI agents and when should you use them?

Computer use agents operate software the way a person does: they read the screen through screenshots and accessibility trees, then click, type and scroll. Use them when a system has no usable API. Always run them in an isolated VM with no unnecessary credentials, approval gates on irreversible actions, and a kill switch. APIs remain faster, cheaper and more reliable when available.

Key facts

PerceptionScreenshots plus accessibility tree or DOM extraction
ActionsClick, type, scroll, drag, keyboard shortcuts and clipboard
IsolationDedicated VM or container with snapshots and no host mounts
CredentialsLeast-privilege accounts; never ambient production access
SpeedSeconds per step, far slower than a direct API call
Cost profileScreenshots and page state consume many tokens per step
Model layer30+ models on one OpenAI-compatible key with live function calling
RoadmapImage and audio endpoints are coming soon, not live yet

TL;DR

  • Use computer use only when no API or export path exists.
  • Run the agent in a disposable VM, never on a real desktop.
  • Gate irreversible actions behind human approval.
  • Screenshots are token-expensive; cap steps and resolution.
  • Log every action with screenshots so failures are reproducible.

How it works, step by step

  1. Confirm an API, database export or bulk import path does not exist for the task.
  2. Provision a disposable VM with its own browser profile, disk snapshot and network rules.
  3. Create a dedicated least-privilege account for the target system.
  4. Expose a small action set — click, type, scroll, wait — with argument validation.
  5. Add a step cap, a wall-clock timeout and a human approval gate for irreversible actions.
  6. Capture screenshots and action logs on every step for replay and debugging.
  7. Rebuild the task as an API integration as soon as the vendor offers one.
1Confirm an API,database export orbulk import path2Provision adisposable VM withits own browser3Create a dedicatedleast-privilegeaccount for the4Expose a smallaction set — click,type, scroll, wait5Add a step cap, awall-clock timeoutand a human6Capture screenshotsand action logs onevery step for

Try it yourself

Open the browser automation generator →

How a computer use agent works

The loop is perceive, decide, act, verify. Perception combines a screenshot with structured context from the accessibility tree or DOM. The model decides on an action, the harness executes it, and the next screenshot confirms the result. Because the agent must look again after every action, tasks take many steps and each step carries an image into context, which is why computer use is the most token-hungry agent pattern.

Reliability comes from grounding actions in stable identifiers where possible. Clicking by element ID or accessible label beats clicking by coordinates, and explicit waits beat fixed sleeps. Where the interface offers keyboard shortcuts or forms, prefer them over pixel hunting.

Sandbox design is the whole safety story

Treat the agent as an untrusted user with a keyboard. Give it a dedicated VM or container with a fresh profile, no shared clipboard, no host filesystem mounts, and outbound network rules limited to the target system. Use a least-privilege account with no admin rights and no access to unrelated data.

  • Snapshots: restore a clean state after every run.
  • Approval gates: payments, deletions, sends and permission changes pause for a person.
  • Rate limits: cap actions per minute so a loop cannot hammer a system.
  • Kill switch: one control stops the run and revokes the session.
  • Audit: store screenshots, actions and model reasoning for every step.

Cost, reliability and when to walk away

Computer use is the fallback, not the default. A direct API call is faster, cheaper and testable; GUI automation breaks whenever a layout changes. Budget for step caps and image downscaling, and measure success by completed tasks per attempt rather than steps taken.

Plugsky supplies the model layer behind such agents: 30+ models on one OpenAI-compatible key with live function calling, streaming and JSON mode, plus scoped keys and audit logs. It does not provide a desktop sandbox or a browser runtime, so those are yours to operate; image and audio endpoints are coming soon rather than live. Match the model to the job — vision-capable models for screen reading, cheaper models for repetitive confirmation steps — and check current plans and the free plugsky-micro and plugsky-lite models on the live pricing page.

Honest comparison

ApproachDirect API integrationStructured export or importComputer use agent
ReliabilityHighest, contract-basedHigh if schema is stableBreaks on UI changes
SpeedMilliseconds to secondsMinutes per batchSeconds per screen step
CostLowestLowHighest, image tokens per step
MaintenanceVersioned APIFormat driftContinuous, UI-dependent
FitPreferred first choiceGood for bulk dataLast resort when no API exists

Frequently asked questions

When is computer use the right choice?

When the target system has no API, no export and no integration path, and the task value justifies the fragility. If any programmatic route exists, use it instead.

Is it safe to let an agent control a browser?

Only inside a disposable sandbox with least-privilege credentials, network restrictions, action caps and approval gates for irreversible steps. Never point it at your daily desktop.

Why are computer use agents slow?

Every action requires a fresh screenshot and a model decision, so a task with twenty steps means twenty perception and reasoning cycles rather than one API call.

How do I reduce token cost?

Downscale screenshots, crop to the relevant region, use accessibility text where available, cache unchanged screens, and route simple confirmation steps to cheaper models.

Which models support screen understanding?

Vision-capable models can read screenshots; check the live model catalogue for current vision support, since coverage changes as models are added. Image generation endpoints are a separate, coming-soon capability.

How should I log a run?

Store each step's screenshot, chosen action, model output and result together, so a failed run can be replayed and the exact divergence point identified.

Does Plugsky provide the sandbox?

No. Plugsky provides the model API and function calling. The VM, browser profile, action executor and approval flow are infrastructure you build or buy.