Key facts
| Core tools | Repo search, file read, patch write, command and test execution |
| Isolation | Ephemeral sandbox with the target branch and no production secrets |
| Feedback loop | Run linters and tests and feed failures back to the model |
| Output | Propose a diff for review rather than pushing to a protected branch |
| Model routing | Fast models for small edits, frontier models for refactors |
| Context | Retrieve relevant files instead of pasting the whole repository |
| Status | Function calling is live; sandboxing remains your responsibility |
| Roadmap | Fine-tuning endpoints are coming soon |
TL;DR
- Treat the coding agent as a junior engineer: sandbox, branch, tests and review.
- Four tool families cover most work: search, read, patch, execute.
- Run tests after every change and feed real failures back into the loop.
- Never give it production credentials or push access to protected branches.
- Retrieve relevant files; whole-repo context wastes tokens and attention.
How it works, step by step
- Create an ephemeral sandbox with a fresh checkout of the target branch.
- Expose search, read, patch and command tools with tight argument schemas.
- Give the agent the task plus repository conventions and the definition of done.
- Run the test suite after each patch and return failures with the relevant output.
- Cap turns and wall-clock time, and stop when tests pass or progress stalls.
- Produce a diff summary for human review rather than committing directly.
- Collect accepted and rejected diffs as evaluation data for the next iteration.
Try it yourself
Open the best model for coding selector →
The four tool families a coding agent needs
Keep the tool surface small and predictable. Search finds candidate files and symbols. Read returns file contents or ranges. Patch applies a structured edit rather than free-form text replacement, so a malformed edit fails loudly instead of corrupting code. Execute runs commands and tests inside the sandbox with a timeout and captured output.
Resist adding more. A handful of well-described tools beats twenty overlapping ones because tool selection accuracy drops as the set grows. If the agent needs git operations, expose a narrow tool for the specific operations you allow rather than raw shell access.
The loop: propose, patch, test, report
The productive pattern is a short loop with real feedback: the agent reads the relevant code, proposes a patch, applies it in the sandbox, runs the tests, and reads the failures. Real test output is the highest-signal context you can provide, which is why sandboxed execution matters even for simple tasks.
- Plan first for multi-file changes, and keep the plan visible in output.
- Patch narrowly — small diffs are easier to review and less likely to regress.
- Test after each change rather than batching at the end.
- Stop conditions — tests green, turn cap reached or two consecutive failures on the same error.
Keeping it safe and reviewable
The agent runs untrusted, generated commands. Sandbox it in a container or VM with no production credentials, no host mounts beyond the checkout, restricted egress and hard resource limits. Treat repository content, issues and dependency files as untrusted input that may attempt to steer the agent.
On review: the deliverable is a diff with a summary of what changed and why, plus test results. A human approves before merge. Route routine changes to a fast model and complex refactors to a frontier model — Plugsky gives you 30+ models on one OpenAI-compatible key with live function calling and streaming, so routing is configuration. Fine-tuning and batch endpoints are coming soon; plans are on the live pricing page.
Honest comparison
| Capability | Simple autocomplete | Chat-based assistant | Coding agent |
|---|---|---|---|
| Scope | One line or block | Suggestions in conversation | Multi-file task completion |
| Tools | Editor context | Paste and copy | Search, patch, execute, test |
| Feedback | None | Human applies and reports | Tests run automatically |
| Isolation | Editor process | None | Sandbox required |
| Review | Inline accept | Manual | Diff plus summary before merge |
Frequently asked questions
Can the agent push code itself?
It can be allowed to push to an ephemeral branch, but merging to a protected branch should stay behind human review. The safest default is a diff and summary that a person approves.
How do I stop it breaking the build?
Run the test suite after every patch inside an isolated sandbox, and stop on repeated failures. Nothing reaches the main branch without a green run and a review.
What context does a coding agent need?
Repository structure, the files relevant to the task, coding conventions, and the definition of done. Retrieve files rather than dumping the entire repository into the prompt.
Which model should I use?
Fast, low-cost models handle small edits, renames and formatting. Frontier models earn their cost on multi-file refactors and debugging. Route per task on one API key.
Are generated commands dangerous?
Directly executing them is. Sandbox execution with no production credentials, restricted egress and resource limits makes the blast radius local and recoverable.
How do I improve it over time?
Track accepted and rejected diffs, common failure classes and test flakiness. Add those cases to the evaluation set and refine tool descriptions and prompts accordingly.
Does Plugsky support code models?
Yes. The catalogue includes coding-focused models among 30+ options, all behind the same OpenAI-compatible API with live function calling.