Key facts
| Core tools | Web search, page fetch, extract, notes and citation tracking |
| Memory | Findings scratchpad plus a source registry keyed by URL |
| Stopping rule | Coverage threshold, turn cap or token budget |
| Citations | Every claim maps back to a retrieved URL and quoted span |
| Model access | 30+ models behind one OpenAI-compatible API key |
| Retrieval | Embeddings and RAG are live for dedupe and grounding |
| Traffic hygiene | Cache fetched pages, cap concurrency and respect robots rules |
| Roadmap | Batch and files endpoints are coming soon |
TL;DR
- Four tools cover most research work: search, fetch, note and cite.
- Bound the loop with a coverage rule, a turn cap and a token budget.
- Store each finding with its source URL and quoted span so synthesis can cite.
- Cache and rate-limit fetches, and treat page content as untrusted input.
- Use fast models for extraction and a frontier model for the final brief.
How it works, step by step
- Define the research question, the deliverable format and what counts as sufficient coverage.
- Implement search, fetch, extract, note and citation tools with tight argument schemas.
- Give the agent a scratchpad where every finding stores text, URL and retrieval time.
- Add a stopping rule: source coverage reached, turn cap hit or budget spent.
- Extract and summarise sources with fast models, and surface conflicting claims explicitly.
- Synthesise the final brief with inline citations and a list of open questions.
- Evaluate against a frozen question set and keep traces of every run for review.
Try it yourself
The research loop that works
The reliable pattern is plan, search, fetch, extract, note, reflect. The agent decomposes the question into sub-questions, runs several searches per sub-question, fetches the most promising results, extracts answers with quotes, stores them as notes, then reflects on what is still missing. Reflection is what separates a research agent from a search wrapper: after each round it should name the gaps and generate the next queries from them.
Keep the loop bounded. Without explicit limits, agents keep searching forever, re-reading the same pages and inflating cost. A coverage check, a maximum turn count and a token budget all matter, and the agent should stop when any one of them is reached.
Memory and citations are the hard part
Research output is only useful if a reader can verify it. Store findings in a structure that keeps the claim, the supporting quote, the URL and the retrieval timestamp together. Deduplicate near-identical findings with embeddings rather than string matching, and keep a separate source registry so the final reference list is stable even if several sub-questions used the same page.
- Scratchpad: working notes the agent can read and rewrite each round.
- Source registry: canonical list of retrieved URLs with titles and dates.
- Quote spans: the exact passage supporting each claim, not a paraphrase.
- Conflict log: where sources disagree, record both positions instead of silently picking one.
Guardrails, cost and model routing
Fetched pages are untrusted input and may contain instructions aimed at your agent. Treat page text as data, never as instructions, and never let the agent call tools based on content it read. Cache pages so repeated work is free, cap concurrent fetches, and identify your crawler honestly.
Route by task: small, fast models handle extraction, classification and summarisation, while a frontier model writes the final synthesis where reasoning quality matters most. Plugsky exposes 30+ models on one OpenAI-compatible key with live function calling, streaming, JSON mode, embeddings and RAG, so routing is configuration rather than a second integration. Plans, including the free plugsky-micro and plugsky-lite models, are on the live pricing page; batch and files endpoints are coming soon.
Honest comparison
| Capability | Manual research | Single-prompt chatbot | Research agent |
|---|---|---|---|
| Coverage | Whatever the person finds | One answer per question | Multi-query plan with follow-ups |
| Citations | Manual | Sometimes, unverified | Claim linked to URL and quote |
| Memory | Notes app | Conversation only | Scratchpad plus source registry |
| Repeatability | Varies by person | Varies by prompt | Frozen eval set per question type |
| Cost control | Time spent | Per prompt | Budget and turn caps per run |
Frequently asked questions
How is a research agent different from deep research in a chat app?
A chat app's deep research feature is a hosted product. Building your own agent keeps the plan, sources and output format under your control, and lets you run it inside your product or pipeline.
How many tools does a research agent need?
Start with five: search, fetch, extract, note and cite. More tools reduce tool-selection accuracy, so add only what a real task demands.
How do I stop it searching forever?
Set three limits: a coverage rule for the sub-questions, a maximum number of turns, and a token or spend budget. Stop at whichever triggers first.
How do I keep citations accurate?
Store the quoted span with the URL at extraction time and build the reference list from the source registry, not from memory at synthesis time.
Can it read paywalled or private content?
Only with credentials you control and authorise. Keep private sources separate from public ones and log which credentials were used for each retrieval.
Which model should write the final brief?
A frontier-tier model, because synthesis is where reasoning quality shows. Extraction and triage can run on smaller, faster models to control cost.
How do I evaluate a research agent?
Freeze a set of questions, define required sources and required facts, and score citation accuracy, coverage and contradiction handling on every change.
Does Plugsky provide search or browsing?
No. Plugsky provides the model and API layer, including live function calling, streaming and embeddings. Search, fetch and browser tools are yours to choose and operate.