Languages

How do you build multilingual AI for India with Plugsky?

India ships many languages at once, and users code-mix freely. Plugsky handles them as UTF-8 text on one OpenAI-compatible endpoint — set base_url to https://api.plugsky.com/v1 and send Devanagari, Tamil, Bengali or Roman-script Hinglish to /v1/chat/completions. Pick a launch set of languages, normalise Indic scripts before indexing, and evaluate code-mixed text separately.

Key facts

Language landscape22 scheduled languages; production apps usually start with a subset plus English
Indic scriptsDevanagari, Bengali, Tamil and others are left-to-right abugidas; NFC-normalise matras and conjuncts
Code-mixingHinglish and other Roman-script mixes need their own prompts and evaluation slices
Embeddingsplugsky-embed-multilingual for Indic languages and English in one collection
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon

TL;DR

  • Start with a launch set of languages, not all 22 at once.
  • NFC-normalise Indic scripts before counting tokens or embedding.
  • Treat Hinglish and Roman-script queries as first-class inputs.
  • One multilingual collection can serve Indic languages and English.
  • Residency options matter for regulated Indian deployments.

How it works, step by step

  1. Rank languages by audience — often English and Hindi first, then the regional scripts your users actually write in.
  2. Create a free Plugsky key and send one prompt per script to /v1/chat/completions to compare models.
  3. Normalise text (NFC, digit style, punctuation) before measuring tokens or embedding.
  4. Build a gold set per language plus a separate Hinglish set with real informal phrasing.
  5. Embed with plugsky-embed-multilingual and test Roman-script queries against Devanagari documents.
  6. Check residency and deployment requirements, then roll out language by language.
1Rank languages byaudience — oftenEnglish and Hindi2Create a freePlugsky key andsend one prompt per3Normalise text(NFC, digit style,punctuation) before4Build a gold setper language plus aseparate Hinglish5Embed withplugsky-embed-multilingualand test6Check residency anddeploymentrequirements, then

Original data

22 scheduled lLanguage landscapeOpenAI-compatiAPI compatibility30+ models behModelsSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the embedding model comparison →

India's language reality: one product, many scripts

India has 22 scheduled languages and a digital audience that switches between them mid-sentence. Most products start with English and Hindi, then add regional languages such as Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam and Punjabi. Each addition brings a different script, its own fonts and its own evaluation work.

The API stays constant across all of them: UTF-8 text in, text out, on the same OpenAI-compatible endpoint. Treat language expansion as a rollout plan rather than a technology change — one language at a time, with its own prompts, glossary and review.

Tokenization and normalization across Indic scripts

Indic scripts are abugidas: consonants carry inherent vowels, and marks modify them. Conjuncts and matras can be encoded in more than one code-point order, so the same visible text may tokenise differently unless you normalise to NFC early. Devanagari also uses its own digit set, which behaves differently from ASCII digits in tokenisers and in search matching.

  • Normalise to NFC before counting tokens, chunking or embedding.
  • Standardise digits and punctuation across the pipeline.
  • Measure tokens per script — they do not track characters.
  • Keep original text for display; normalise only derived fields.

Hinglish and code-mixing

Hinglish — Hindi syntax with heavy English vocabulary, written in Roman script — is normal in chat, reviews and support threads. Tanglish and similar mixes follow the same pattern in the south. A model that handles clean Hindi can still stumble on Roman-script code-mixing, so define how much mixing each surface accepts.

  • Collect real code-mixed examples; do not translate them into standard Hindi for evaluation.
  • Keep Hinglish prompts short and explicit about the expected reply script.
  • Index Roman-script queries with multilingual embeddings so they can match Devanagari documents.
  • Score code-mixed outputs with reviewers who read that register daily.

Evaluation, formats and compliance

Local formats matter: digit grouping follows the lakh and crore system (for example 1,00,000), rupee symbols appear in prices, and date formats vary. Test these explicitly because models often default to Western conventions. Build per-language gold sets and review them with native speakers rather than relying on fluency impressions.

On compliance, India's Digital Personal Data Protection Act shapes how teams handle personal data, and many enterprises pair it with internal residency rules. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options, so confirm the arrangement that fits your obligations with your legal team and the enterprise team. See the docs for deployment details.

Honest comparison

CapabilityPlugskyIndia AI stack todayBuilding in-house
API compatibilityOne OpenAI-compatible endpoint for every scriptSeveral providers and SDKs to maintainFull rewrite
Indic text handlingUTF-8 text with NFC guidance, per-language evaluationDepends on provider tokeniser and preprocessingYou build normalisation and evals
Code-mixingHinglish treated as its own prompt and eval sliceOften untested until users complainYou collect and curate datasets
Multilingual retrievalplugsky-embed-multilingual for Indic languages and EnglishOften a separate embedding vendorYou serve and maintain embeddings
SovereigntyCloud, VPC, on-prem and air-gapped with residency optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Does Plugsky support Indic languages?

The API accepts UTF-8 text in any Indic script on the OpenAI-compatible chat endpoint. Model quality varies by language, so evaluate candidates on your own gold sets.

Should I launch all 22 scheduled languages at once?

No. Rank by audience and ship one language at a time with its own prompts, glossary and review. Each script has distinct tokenization and formatting issues.

How do I handle Hinglish inputs?

Keep them in their own prompt and evaluation slice. Roman-script code-mixing behaves differently from standard Hindi, and translating it before testing hides real failure modes.

How do I count tokens for Devanagari or Tamil?

Use the token calculator on real text after NFC normalisation. Indic scripts vary in tokens per word, and matras or conjuncts can split into multiple subwords.

Is there a multilingual embedding model?

Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports cross-language retrieval, including Roman-script queries against Indic-script documents.

Can I keep data inside India or my own environment?

Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options. Confirm the specifics with the docs and the enterprise team for your compliance needs.

Is there a free plan for prototyping?

Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers stronger models for evaluation.

How do I migrate an existing app?

Change base_url to https://api.plugsky.com/v1 and map model names. Chat, streaming, JSON mode, function calling and embeddings keep their request shapes.