Key facts
| Hindi script | Left-to-right Devanagari abugida; matras and conjuncts need NFC normalisation |
| Tokenisation | Matras and conjuncts can fragment into subwords — measure with the token calculator |
| Code-mixing | Hinglish in Roman script needs its own prompts and evaluation slices |
| Transliteration | Users type Hindi words in Latin letters; multilingual embeddings match them to Devanagari documents |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Hindi runs on the standard endpoint with one base_url change.
- NFC-normalise Devanagari before measuring tokens or embedding.
- Keep a separate prompt and eval slice for Hinglish.
- Roman-script queries can match Devanagari documents via multilingual embeddings.
- plugsky-embed-multilingual serves Hindi and English from one collection.
How it works, step by step
- Create a free Plugsky key and set base_url to https://api.plugsky.com/v1.
- NFC-normalise all Hindi input before counting tokens, chunking or embedding.
- Send real Devanagari prompts to /v1/chat/completions and compare two or three models.
- Measure tokens with the token calculator and size chunks from Devanagari samples, not transliterations.
- Collect Hinglish examples and create a separate prompt template and evaluation set for them.
- For RAG, embed with plugsky-embed-multilingual and test Roman-script queries against Devanagari documents, then review with native speakers.
Try it yourself
How Plugsky handles Hindi text
Hindi needs no special endpoint: Devanagari UTF-8 text goes in and comes back. The language-specific work is normalisation and code-mixing. Devanagari consonants carry an inherent vowel, and matras (vowel signs), nukta marks and conjuncts modify them. Some sequences have more than one valid code-point order, so the same visible text can tokenise differently unless you normalise to NFC at the edge.
The second reality is Hinglish: Hindi grammar in Roman script with heavy English vocabulary, extremely common in chat, reviews and support. A model that writes clean Devanagari can still mishandle Roman-script input, so treat the two as separate input classes with their own prompts.
Tokenisation and cost in Hindi
Devanagari words often split into more subwords than their English translations, because matras and conjuncts fragment under character or byte-based encoding. Devanagari digits behave differently from ASCII digits in tokenisers and in search matching, so standardise them early.
- Normalise to NFC before measuring or embedding.
- Count tokens on real Devanagari text, not transliterations.
- Standardise digits and punctuation across the pipeline.
- Keep the original for display; normalise derived fields.
Hindi retrieval and RAG
One multilingual collection built with plugsky-embed-multilingual serves Devanagari documents with Devanagari, Hinglish or English queries. The common failure is a user typing Hindi phonetically in Latin letters and getting nothing, because keyword matching never bridges the scripts; embeddings handle that case well.
- Embed with plugsky-embed-multilingual and keep one model per collection.
- Add transliteration keys if you also run keyword search.
- Test Latin-script queries against Devanagari documents explicitly.
- Keep Hinglish queries in their own evaluation slice.
Code example: a Hindi request
Point your OpenAI client at https://api.plugsky.com/v1; normalise the input text before sending if your sources may use a different Unicode form.
client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])
client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "हिंदी में विनम्र और स्पष्ट उत्तर दें"}, {"role": "user", "content": "इस अनुबंध को तीन बिंदुओं में सारांशित करें"}])
Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models on a Hindi gold set with native review. See the docs for the API reference.
Honest comparison
| Capability | Plugsky | Hindi apps today | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change base_url and model name | Varies by provider and SDK | Full rewrite |
| Devanagari handling | UTF-8 with NFC guidance and code-mixing support | Depends on provider tokeniser and preprocessing | You build normalisation and evals |
| Token budget | Fixed tokeniser per model; measure Devanagari and chunk to fit | Varies by provider and model | You host and tune each tokeniser |
| Multilingual retrieval | plugsky-embed-multilingual for Hindi, Hinglish and English RAG | Often needs a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped with residency options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Can Plugsky handle Hindi text?
Yes. The API accepts UTF-8 Devanagari input on the OpenAI-compatible chat endpoint. Compare two or three models on your own prompts, since quality varies by task.
Why does Devanagari need NFC normalisation?
Matras, nukta marks and conjuncts can be encoded in more than one code-point order for the same visible text. Normalising to NFC prevents duplicate strings, mismatched search and unstable token counts.
How do I handle Hinglish?
Treat it as a separate input class with its own prompt template and evaluation set. Do not translate it into standard Hindi for testing, or you will hide the failures users see.
Can English-script queries search Hindi documents?
Yes, with multilingual embeddings. Add transliteration keys as well if you run keyword search, since phonetic Roman spellings vary.
Is there a multilingual embedding model?
Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports Hindi and English retrieval from one collection.
Can I keep data in my own environment?
Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options for regulated teams.
Is there a free plan?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the paid catalogue.