Key facts
| API surface | OpenAI-compatible /v1/chat/completions; keep your existing SDK |
| Grounding | Embeddings and RAG are live for campaign, brand and client knowledge search |
| Models | 30+ models behind one API, from free tiers to frontier |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped |
| Access control | Scoped API keys, rotation and usage analytics; enterprise SSO and RBAC options |
| Endpoint roadmap | Audio, images, moderation, batch and fine-tuning are coming soon |
| Pricing model | Flat monthly self-serve plans; no per-token billing on self-serve |
| Free tier | Free plan with plugsky-micro and plugsky-lite; 14-day full-access trial |
TL;DR
- Ground advertising answers in approved internal content with RAG instead of model memory.
- Keep client or unreleased campaign content inside VPC, on-prem or air-gapped deployments where policy requires it.
- Scope retrieval per team, client or site so permissions and confidentiality hold at query time.
- Keep a named human owner for every client-facing decision.
- Embargo unreleased campaign material in a separate scoped workspace.
How it works, step by step
- Inventory the content to make searchable: creative briefs, campaign performance reports, brand guidelines and media plans.
- Define a data boundary for the pilot that excludes restricted material until controls are proven.
- Choose a deployment target: cloud for public content, VPC, on-prem or air-gapped for restricted data.
- Build an evaluation set with account planners and strategists so answer quality is judged by domain experts.
- Ingest, chunk and embed approved documents, and require citations on every answer.
- Add refusal behavior for questions outside the indexed, approved content.
- Review logged interactions on a schedule and expand only after accuracy and access checks pass.
Original data
Try it yourself
Where advertising teams start
Start with internal questions that already have a written answer. High-value first workloads include:
- Creative and brief search: retrieve approved briefs, brand rules and past concepts with citations.
- Campaign performance Q&A: answer questions from post-campaign reports and media plans.
- Client and account knowledge: surface contracts, meeting notes and deliverable history per client.
- Media planning support: pull audience, channel and budget notes from approved plans.
Each use case augments staff with cited answers; none replaces account lead judgment.
A private RAG architecture for advertising knowledge
The stack is consistent across industries: ingest approved creative briefs, campaign performance reports, brand guidelines and media plans, chunk and embed with a multilingual embedding model, store vectors inside your environment, and call chat completions that answer only from retrieved context. Plugsky embeddings and chat completions are OpenAI-compatible, so teams already using OpenAI SDKs change the base URL and keep their code.
Use JSON mode when answers feed reporting or brief-generation tools, and streaming when a strategist is reading. Model routing can send simple asset lookups to a small model and multi-document synthesis to a frontier model.
Access control, confidentiality and audit
Client confidentiality is the central constraint: retrieval must respect which agency team, client or pitch owns each document. Tag content with client and project metadata at ingestion and filter at query time rather than trusting the model to remember who may see what. Flag embargoed or unreleased campaign material for stricter scoping.
The technical controls are consistent: enforce permission-aware retrieval in your own service layer, scope API keys per application, team or tenant, rotate keys, and retain request and response logs on a defined schedule. Regulatory obligations vary by jurisdiction and sector, so map them with counsel rather than assuming one framework covers every deployment; Plugsky supplies the deployment and logging primitives you document.
See what private AI means for related deployment and control detail.
Rollout and human oversight
Pilot on brand and process documentation before touching client or unreleased campaign content. Build a labelled question set with account planners and strategists, then measure retrieval hit rate, citation correctness and answer accuracy before and after every index or model change. Require citations on every answer, refuse out-of-scope questions, and name a human owner for every client-facing decision.
Review logged interactions weekly at first, correct the index rather than the prompt when retrieval misses, and expand the corpus only when accuracy and access checks pass.
Honest comparison
| Capability | Plugsky | Public AI assistants | Building in-house |
|---|---|---|---|
| Data boundary | Cloud, VPC, on-prem or air-gapped | Vendor cloud only | You control fully |
| Grounding | Embeddings and RAG are live for campaign, brand and client knowledge search | Uncontrolled retrieval | You assemble and operate |
| Access control | Scoped keys, usage analytics, enterprise SSO and RBAC options | Account-level only | Custom identity work |
| Auditability | Request and response logging | Limited | You build logging |
| Pricing | Flat monthly self-serve plans; see live pricing | Per-seat or per-token | GPU plus operations cost |
| Time to pilot | Days | Hours, without residency control | Quarters |
Frequently asked questions
Can advertising companies keep data private with Plugsky?
Yes. Choose the deployment boundary that matches the data: Plugsky cloud for public content, or your VPC, on-prem and air-gapped options for restricted material. Access is controlled with scoped API keys and usage analytics.
Do we need to fine-tune on our internal documents?
Not for a first release. Fine-tuning is coming soon and is better for style than facts. RAG keeps answers current, permission-aware and traceable to a source, which matters more for internal knowledge.
Which model should we use?
Start free with plugsky-micro and plugsky-lite to validate retrieval, then evaluate mid-tier and frontier models from the 30+ model catalogue on your own question set.
How do we stop wrong or unsupported answers?
Restrict the assistant to approved indexed content, require citations, refuse out-of-scope questions and keep a human decision-maker for every regulated or client-facing outcome.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage, and there are no per-token charges on self-serve plans. See the live pricing page for current plans and the free tier.
Can the assistant use client campaign data?
Only inside a workspace and deployment boundary the client has approved. Tag documents by client and filter retrieval at query time; never index another client's material into a shared space.
How do we handle embargoed creative?
Keep embargoed material out of the default index and expose it through a separate, tightly scoped workspace with audit logging until the campaign is public.