Key facts
| Language landscape | 22 scheduled languages; production apps usually start with a subset plus English |
| Indic scripts | Devanagari, Bengali, Tamil and others are left-to-right abugidas; NFC-normalise matras and conjuncts |
| Code-mixing | Hinglish and other Roman-script mixes need their own prompts and evaluation slices |
| Embeddings | plugsky-embed-multilingual for Indic languages and English in one collection |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Start with a launch set of languages, not all 22 at once.
- NFC-normalise Indic scripts before counting tokens or embedding.
- Treat Hinglish and Roman-script queries as first-class inputs.
- One multilingual collection can serve Indic languages and English.
- Residency options matter for regulated Indian deployments.
How it works, step by step
- Rank languages by audience — often English and Hindi first, then the regional scripts your users actually write in.
- Create a free Plugsky key and send one prompt per script to /v1/chat/completions to compare models.
- Normalise text (NFC, digit style, punctuation) before measuring tokens or embedding.
- Build a gold set per language plus a separate Hinglish set with real informal phrasing.
- Embed with plugsky-embed-multilingual and test Roman-script queries against Devanagari documents.
- Check residency and deployment requirements, then roll out language by language.
Original data
Try it yourself
Open the embedding model comparison →
India's language reality: one product, many scripts
India has 22 scheduled languages and a digital audience that switches between them mid-sentence. Most products start with English and Hindi, then add regional languages such as Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam and Punjabi. Each addition brings a different script, its own fonts and its own evaluation work.
The API stays constant across all of them: UTF-8 text in, text out, on the same OpenAI-compatible endpoint. Treat language expansion as a rollout plan rather than a technology change — one language at a time, with its own prompts, glossary and review.
Tokenization and normalization across Indic scripts
Indic scripts are abugidas: consonants carry inherent vowels, and marks modify them. Conjuncts and matras can be encoded in more than one code-point order, so the same visible text may tokenise differently unless you normalise to NFC early. Devanagari also uses its own digit set, which behaves differently from ASCII digits in tokenisers and in search matching.
- Normalise to NFC before counting tokens, chunking or embedding.
- Standardise digits and punctuation across the pipeline.
- Measure tokens per script — they do not track characters.
- Keep original text for display; normalise only derived fields.
Hinglish and code-mixing
Hinglish — Hindi syntax with heavy English vocabulary, written in Roman script — is normal in chat, reviews and support threads. Tanglish and similar mixes follow the same pattern in the south. A model that handles clean Hindi can still stumble on Roman-script code-mixing, so define how much mixing each surface accepts.
- Collect real code-mixed examples; do not translate them into standard Hindi for evaluation.
- Keep Hinglish prompts short and explicit about the expected reply script.
- Index Roman-script queries with multilingual embeddings so they can match Devanagari documents.
- Score code-mixed outputs with reviewers who read that register daily.
Evaluation, formats and compliance
Local formats matter: digit grouping follows the lakh and crore system (for example 1,00,000), rupee symbols appear in prices, and date formats vary. Test these explicitly because models often default to Western conventions. Build per-language gold sets and review them with native speakers rather than relying on fluency impressions.
On compliance, India's Digital Personal Data Protection Act shapes how teams handle personal data, and many enterprises pair it with internal residency rules. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options, so confirm the arrangement that fits your obligations with your legal team and the enterprise team. See the docs for deployment details.
Honest comparison
| Capability | Plugsky | India AI stack today | Building in-house |
|---|---|---|---|
| API compatibility | One OpenAI-compatible endpoint for every script | Several providers and SDKs to maintain | Full rewrite |
| Indic text handling | UTF-8 text with NFC guidance, per-language evaluation | Depends on provider tokeniser and preprocessing | You build normalisation and evals |
| Code-mixing | Hinglish treated as its own prompt and eval slice | Often untested until users complain | You collect and curate datasets |
| Multilingual retrieval | plugsky-embed-multilingual for Indic languages and English | Often a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped with residency options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Does Plugsky support Indic languages?
The API accepts UTF-8 text in any Indic script on the OpenAI-compatible chat endpoint. Model quality varies by language, so evaluate candidates on your own gold sets.
Should I launch all 22 scheduled languages at once?
No. Rank by audience and ship one language at a time with its own prompts, glossary and review. Each script has distinct tokenization and formatting issues.
How do I handle Hinglish inputs?
Keep them in their own prompt and evaluation slice. Roman-script code-mixing behaves differently from standard Hindi, and translating it before testing hides real failure modes.
How do I count tokens for Devanagari or Tamil?
Use the token calculator on real text after NFC normalisation. Indic scripts vary in tokens per word, and matras or conjuncts can split into multiple subwords.
Is there a multilingual embedding model?
Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports cross-language retrieval, including Roman-script queries against Indic-script documents.
Can I keep data inside India or my own environment?
Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options. Confirm the specifics with the docs and the enterprise team for your compliance needs.
Is there a free plan for prototyping?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers stronger models for evaluation.
How do I migrate an existing app?
Change base_url to https://api.plugsky.com/v1 and map model names. Chat, streaming, JSON mode, function calling and embeddings keep their request shapes.