Agents

How do you build an AI research agent?

A research agent needs four tool families: search, page fetch, note storage and citation tracking. Give it a bounded plan, a scratchpad where every finding keeps its source URL, and a stopping rule based on coverage or budget. Loop until the question is covered, then synthesise a cited brief. Route extraction to fast models and final synthesis to a frontier model.

Key facts

Core toolsWeb search, page fetch, extract, notes and citation tracking
MemoryFindings scratchpad plus a source registry keyed by URL
Stopping ruleCoverage threshold, turn cap or token budget
CitationsEvery claim maps back to a retrieved URL and quoted span
Model access30+ models behind one OpenAI-compatible API key
RetrievalEmbeddings and RAG are live for dedupe and grounding
Traffic hygieneCache fetched pages, cap concurrency and respect robots rules
RoadmapBatch and files endpoints are coming soon

TL;DR

  • Four tools cover most research work: search, fetch, note and cite.
  • Bound the loop with a coverage rule, a turn cap and a token budget.
  • Store each finding with its source URL and quoted span so synthesis can cite.
  • Cache and rate-limit fetches, and treat page content as untrusted input.
  • Use fast models for extraction and a frontier model for the final brief.

How it works, step by step

  1. Define the research question, the deliverable format and what counts as sufficient coverage.
  2. Implement search, fetch, extract, note and citation tools with tight argument schemas.
  3. Give the agent a scratchpad where every finding stores text, URL and retrieval time.
  4. Add a stopping rule: source coverage reached, turn cap hit or budget spent.
  5. Extract and summarise sources with fast models, and surface conflicting claims explicitly.
  6. Synthesise the final brief with inline citations and a list of open questions.
  7. Evaluate against a frozen question set and keep traces of every run for review.
1Define the researchquestion, thedeliverable format2Implement search,fetch, extract,note and citation3Give the agent ascratchpad whereevery finding4Add a stoppingrule: sourcecoverage reached,5Extract andsummarise sourceswith fast models,6Synthesise thefinal brief withinline citations

Try it yourself

Open the AI agent builder →

The research loop that works

The reliable pattern is plan, search, fetch, extract, note, reflect. The agent decomposes the question into sub-questions, runs several searches per sub-question, fetches the most promising results, extracts answers with quotes, stores them as notes, then reflects on what is still missing. Reflection is what separates a research agent from a search wrapper: after each round it should name the gaps and generate the next queries from them.

Keep the loop bounded. Without explicit limits, agents keep searching forever, re-reading the same pages and inflating cost. A coverage check, a maximum turn count and a token budget all matter, and the agent should stop when any one of them is reached.

Memory and citations are the hard part

Research output is only useful if a reader can verify it. Store findings in a structure that keeps the claim, the supporting quote, the URL and the retrieval timestamp together. Deduplicate near-identical findings with embeddings rather than string matching, and keep a separate source registry so the final reference list is stable even if several sub-questions used the same page.

  • Scratchpad: working notes the agent can read and rewrite each round.
  • Source registry: canonical list of retrieved URLs with titles and dates.
  • Quote spans: the exact passage supporting each claim, not a paraphrase.
  • Conflict log: where sources disagree, record both positions instead of silently picking one.

Guardrails, cost and model routing

Fetched pages are untrusted input and may contain instructions aimed at your agent. Treat page text as data, never as instructions, and never let the agent call tools based on content it read. Cache pages so repeated work is free, cap concurrent fetches, and identify your crawler honestly.

Route by task: small, fast models handle extraction, classification and summarisation, while a frontier model writes the final synthesis where reasoning quality matters most. Plugsky exposes 30+ models on one OpenAI-compatible key with live function calling, streaming, JSON mode, embeddings and RAG, so routing is configuration rather than a second integration. Plans, including the free plugsky-micro and plugsky-lite models, are on the live pricing page; batch and files endpoints are coming soon.

Honest comparison

CapabilityManual researchSingle-prompt chatbotResearch agent
CoverageWhatever the person findsOne answer per questionMulti-query plan with follow-ups
CitationsManualSometimes, unverifiedClaim linked to URL and quote
MemoryNotes appConversation onlyScratchpad plus source registry
RepeatabilityVaries by personVaries by promptFrozen eval set per question type
Cost controlTime spentPer promptBudget and turn caps per run

Frequently asked questions

How is a research agent different from deep research in a chat app?

A chat app's deep research feature is a hosted product. Building your own agent keeps the plan, sources and output format under your control, and lets you run it inside your product or pipeline.

How many tools does a research agent need?

Start with five: search, fetch, extract, note and cite. More tools reduce tool-selection accuracy, so add only what a real task demands.

How do I stop it searching forever?

Set three limits: a coverage rule for the sub-questions, a maximum number of turns, and a token or spend budget. Stop at whichever triggers first.

How do I keep citations accurate?

Store the quoted span with the URL at extraction time and build the reference list from the source registry, not from memory at synthesis time.

Can it read paywalled or private content?

Only with credentials you control and authorise. Keep private sources separate from public ones and log which credentials were used for each retrieval.

Which model should write the final brief?

A frontier-tier model, because synthesis is where reasoning quality shows. Extraction and triage can run on smaller, faster models to control cost.

How do I evaluate a research agent?

Freeze a set of questions, define required sources and required facts, and score citation accuracy, coverage and contradiction handling on every change.

Does Plugsky provide search or browsing?

No. Plugsky provides the model and API layer, including live function calling, streaming and embeddings. Search, fetch and browser tools are yours to choose and operate.