Key facts
| Core tools | Schema introspection, read-only query, result sampling, chart output |
| Access | Read-only credentials on a replica or warehouse role |
| Safety | Parse and validate SQL, cap rows and runtime, block DDL and DML |
| Grounding | Require stated assumptions and data freshness in answers |
| Validation | Cross-check totals and row counts before presenting findings |
| Structured output | JSON mode is live for typed intermediate results |
| Model routing | Small models for exploration, frontier models for analysis |
| Roadmap | Files and batch endpoints are coming soon |
TL;DR
- Connect a read-only replica; the agent should never hold write credentials.
- Validate SQL before execution and cap rows, runtime and scans.
- Require the agent to state assumptions, filters and data freshness.
- Sample and cross-check results rather than trusting the first query.
- Route exploratory steps to small models and interpretation to frontier models.
How it works, step by step
- Create a read-only database role with access limited to the tables the agent needs.
- Expose schema introspection so the agent can discover tables, columns and types.
- Add a query tool that validates SQL, blocks writes and enforces row and time limits.
- Give the agent a sampling tool so it can inspect data before writing complex queries.
- Require answers to state the query logic, filters, assumptions and data as-of date.
- Cross-check aggregate results with an independent query before presentation.
- Log every query with the user identity, and review expensive or failed queries.
Try it yourself
Why analysis is a strong agent use case
Analysis is iterative by nature: look at the schema, sample the data, form a hypothesis, query, check the result, refine. That loop maps directly onto an agent with tools, and the intermediate results provide natural verification points. Unlike open-ended writing, analysis has ground truth — the numbers either reconcile or they do not.
The risk is confident wrong answers. A plausible query against the wrong table produces a clean-looking result with the wrong story. That is why validation, freshness metadata and stated assumptions are part of the design, not optional polish.
Tools and guardrails for SQL
- Introspection: list tables, columns, types and relationships before querying.
- Read-only execution: a role that cannot write, with statement timeout and row caps.
- SQL validation: parse the statement and reject anything outside a select allowlist.
- Cost limits: cap scanned bytes or estimated cost on warehouse engines.
- Result limits: truncate with an explicit notice so the model knows it saw a sample.
- Lineage: log the query, filters and result hash for every answer.
From numbers to defensible findings
The final step is judgement: is this difference meaningful, is the sample representative, is the data fresh enough to act on. Require the agent to state data vintage, filter assumptions and confidence, and to flag when a result is surprising enough to verify manually. A useful pattern is a self-check pass where the agent writes a second query that should corroborate the first.
For access control, use the same querying identity rules as your BI tools — the agent acts for the requesting user, not as a superuser. On Plugsky, JSON mode and function calling are live for structured intermediate results, and 30+ models on one key let you keep exploration cheap. Regulated teams can run in-region, in a VPC, on-prem or air-gapped. Plans are on the live pricing page; files and batch endpoints are coming soon.
Honest comparison
| Concern | Naive text-to-SQL | Guarded analysis agent | BI dashboard |
|---|---|---|---|
| Credentials | Often read-write | Read-only scoped role | Varies |
| Query safety | None | Validation, limits, timeouts | Predefined queries |
| Freshness | Unstated | Explicit as-of date | Scheduled refresh |
| Verification | None | Cross-checks and samples | Manual |
| Flexibility | High | High within guardrails | Low |
Frequently asked questions
Can the agent write to the database?
It should not. Use a read-only role on a replica, and handle any write workflow as a separate, approved process outside the analysis loop.
How do I prevent expensive queries?
Set row caps, statement timeouts and, on warehouse engines, estimated cost or scanned-byte limits. Log and review the most expensive queries.
How do I handle wrong but plausible answers?
Require stated assumptions and data freshness, cross-check aggregates with a second query, and sample raw rows for surprising results.
Should the agent show its SQL?
Yes, in an expandable view or appendix. Analysts trust results they can audit, and query visibility is also the fastest debugging tool.
Which model should generate SQL?
Modern mid-tier models handle typical SQL well. Route schema exploration to small models and schema grounding, then use a frontier model for complex analysis and interpretation.
Can it produce charts?
Yes. Have the agent return a typed result set plus a chart specification, then render charts in your own front end rather than in the model.
Does this work with a data warehouse?
Yes. The pattern is engine-agnostic: introspection, validated read-only queries and result limits work on any SQL warehouse with appropriate roles.