Agents

How do you build an AI data analysis agent?

A data analysis agent inspects schemas, writes queries, runs them against a read-only replica, validates results and explains findings. Give it schema introspection, a query tool with row and time limits, and a way to output tables or charts. Validate generated SQL before execution, restrict access per user, and require stated assumptions and data freshness.

Key facts

Core toolsSchema introspection, read-only query, result sampling, chart output
AccessRead-only credentials on a replica or warehouse role
SafetyParse and validate SQL, cap rows and runtime, block DDL and DML
GroundingRequire stated assumptions and data freshness in answers
ValidationCross-check totals and row counts before presenting findings
Structured outputJSON mode is live for typed intermediate results
Model routingSmall models for exploration, frontier models for analysis
RoadmapFiles and batch endpoints are coming soon

TL;DR

  • Connect a read-only replica; the agent should never hold write credentials.
  • Validate SQL before execution and cap rows, runtime and scans.
  • Require the agent to state assumptions, filters and data freshness.
  • Sample and cross-check results rather than trusting the first query.
  • Route exploratory steps to small models and interpretation to frontier models.

How it works, step by step

  1. Create a read-only database role with access limited to the tables the agent needs.
  2. Expose schema introspection so the agent can discover tables, columns and types.
  3. Add a query tool that validates SQL, blocks writes and enforces row and time limits.
  4. Give the agent a sampling tool so it can inspect data before writing complex queries.
  5. Require answers to state the query logic, filters, assumptions and data as-of date.
  6. Cross-check aggregate results with an independent query before presentation.
  7. Log every query with the user identity, and review expensive or failed queries.
1Create a read-onlydatabase role withaccess limited to2Expose schemaintrospection sothe agent can3Add a query toolthat validates SQL,blocks writes and4Give the agent asampling tool so itcan inspect data5Require answers tostate the querylogic, filters,6Cross-checkaggregate resultswith an independent

Try it yourself

Open the JSON mode tester →

Why analysis is a strong agent use case

Analysis is iterative by nature: look at the schema, sample the data, form a hypothesis, query, check the result, refine. That loop maps directly onto an agent with tools, and the intermediate results provide natural verification points. Unlike open-ended writing, analysis has ground truth — the numbers either reconcile or they do not.

The risk is confident wrong answers. A plausible query against the wrong table produces a clean-looking result with the wrong story. That is why validation, freshness metadata and stated assumptions are part of the design, not optional polish.

Tools and guardrails for SQL

  • Introspection: list tables, columns, types and relationships before querying.
  • Read-only execution: a role that cannot write, with statement timeout and row caps.
  • SQL validation: parse the statement and reject anything outside a select allowlist.
  • Cost limits: cap scanned bytes or estimated cost on warehouse engines.
  • Result limits: truncate with an explicit notice so the model knows it saw a sample.
  • Lineage: log the query, filters and result hash for every answer.

From numbers to defensible findings

The final step is judgement: is this difference meaningful, is the sample representative, is the data fresh enough to act on. Require the agent to state data vintage, filter assumptions and confidence, and to flag when a result is surprising enough to verify manually. A useful pattern is a self-check pass where the agent writes a second query that should corroborate the first.

For access control, use the same querying identity rules as your BI tools — the agent acts for the requesting user, not as a superuser. On Plugsky, JSON mode and function calling are live for structured intermediate results, and 30+ models on one key let you keep exploration cheap. Regulated teams can run in-region, in a VPC, on-prem or air-gapped. Plans are on the live pricing page; files and batch endpoints are coming soon.

Honest comparison

ConcernNaive text-to-SQLGuarded analysis agentBI dashboard
CredentialsOften read-writeRead-only scoped roleVaries
Query safetyNoneValidation, limits, timeoutsPredefined queries
FreshnessUnstatedExplicit as-of dateScheduled refresh
VerificationNoneCross-checks and samplesManual
FlexibilityHighHigh within guardrailsLow

Frequently asked questions

Can the agent write to the database?

It should not. Use a read-only role on a replica, and handle any write workflow as a separate, approved process outside the analysis loop.

How do I prevent expensive queries?

Set row caps, statement timeouts and, on warehouse engines, estimated cost or scanned-byte limits. Log and review the most expensive queries.

How do I handle wrong but plausible answers?

Require stated assumptions and data freshness, cross-check aggregates with a second query, and sample raw rows for surprising results.

Should the agent show its SQL?

Yes, in an expandable view or appendix. Analysts trust results they can audit, and query visibility is also the fastest debugging tool.

Which model should generate SQL?

Modern mid-tier models handle typical SQL well. Route schema exploration to small models and schema grounding, then use a frontier model for complex analysis and interpretation.

Can it produce charts?

Yes. Have the agent return a typed result set plus a chart specification, then render charts in your own front end rather than in the model.

Does this work with a data warehouse?

Yes. The pattern is engine-agnostic: introspection, validated read-only queries and result limits work on any SQL warehouse with appropriate roles.