Languages

How does Plugsky support different languages?

Plugsky treats language as data, not configuration: every language goes through the same OpenAI-compatible endpoint as UTF-8 text. The differences live in tokenization, script direction, registers and retrieval. These guides cover those details per language — set base_url to https://api.plugsky.com/v1, pick a model, and use plugsky-embed-multilingual when one collection must serve several languages.

Key facts

Language coveragePer-language guides for Arabic, Hindi, Japanese, European and Southeast Asian languages
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Embeddingsplugsky-embed-multilingual for cross-language retrieval
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required
Trial14-day full-access trial for the paid catalogue
DeploymentPlugsky cloud, your VPC, on-prem and air-gapped options
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon

TL;DR

  • One endpoint handles every language — there is no locale parameter to set.
  • Tokenization, script direction and registers are the real per-language work.
  • Use the language guides to plan prompts, chunking and evaluation.
  • plugsky-embed-multilingual keeps multilingual RAG in one collection.
  • Prototype on free models, then use the 14-day trial for stronger ones.

How it works, step by step

  1. Pick the languages you must serve and rank them by user impact.
  2. Open the matching language guide and note its script, direction and tokenization quirks.
  3. Create a free API key, set base_url to https://api.plugsky.com/v1 and test prompts per language against two or three models.
  4. Measure tokens per language and set context and chunk budgets from the worst case.
  5. Embed multilingual content with plugsky-embed-multilingual and test cross-language queries.
  6. Build a gold set per language, review with native speakers and roll out language by language.
1Pick the languagesyou must serve andrank them by user2Open the matchinglanguage guide andnote its script,3Create a free APIkey, set base_urlto4Measure tokens perlanguage and setcontext and chunk5Embed multilingualcontent withplugsky-embed-multilingual6Build a gold setper language,review with native

Original data

OpenAI-compatiAPI compatibility30+ models behModels14-day full-acTrialSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the model picker →

What these language guides cover

Each guide in this hub is written for engineers, not linguists. You get the script and direction facts that affect rendering, the tokenization behaviour that affects cost and context, the register and locale choices that affect output quality, and an evaluation approach that catches regressions before users do.

The API surface never changes. Arabic, Hindi, Japanese or Portuguese all travel as UTF-8 text through the same OpenAI-compatible endpoint, so a single integration serves every market. What changes is the engineering around the call: normalisation, chunking, prompt policy and review.

Script, direction and tokenization: what changes per language

Three axes explain most language-specific work. Direction and script: Arabic is right-to-left with cursive joining, while Indic and Southeast Asian scripts combine base characters with marks that benefit from Unicode normalisation. Segmentation: Japanese and Thai are written without spaces, so indexing needs a segmenter rather than whitespace splitting. Morphology: Finnish, Turkish and Hungarian stack suffixes, while German and the Nordic languages build long compounds — both patterns fragment into more subword tokens than plain English.

  • Latin-script European languages: watch accents, elisions and compounds.
  • Right-to-left languages: plan bidi isolation and mirrored UI.
  • No-space scripts: segment before indexing or chunking.
  • Agglutinative languages: budget extra tokens per word.

One API and one embedding collection for every language

Multilingual products usually fail at retrieval before they fail at generation. A single collection built with plugsky-embed-multilingual lets Devanagari documents answer Roman-script queries and lets an English question retrieve Arabic passages, without maintaining one index per language.

Pair that with deployment choice: Plugsky runs in its own cloud, in your VPC, on-prem or air-gapped, which matters when a market has residency or sovereignty requirements. Because the API is OpenAI-compatible, none of this changes your client code — model names and the base URL stay the only moving parts.

How to use this hub

Start with the guide for your highest-impact language, run a short evaluation on free models, and only then widen scope. Keep per-language prompt templates in version control, re-run gold sets whenever a model or preprocessing step changes, and move traffic market by market so a regression in one language never affects the others.

If you are unsure where to begin, open a language guide, copy its code sample, and change the model name to one of the free models — the first request takes minutes. See the docs for the API reference and the pricing page for plan details.

Honest comparison

CapabilityPlugskyTypical multi-provider setupBuilding in-house
API compatibilityOne OpenAI-compatible endpoint for every languageSeveral provider SDKs to maintainFull rewrite
Language guidancePer-language engineering guides in one hubVendor docs scattered across languagesYou write and maintain them
Multilingual retrievalplugsky-embed-multilingual serves cross-language RAGOften one embedding stack per vendorYou serve and tune embeddings
Token planningMeasure per language and chunk to fit with built-in toolsEstimate per provider and modelYou host and tune each tokeniser
SovereigntyCloud, VPC, on-prem and air-gapped deployment optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Which languages does Plugsky support?

The API accepts UTF-8 text in any language. Models differ in quality per language, so evaluate candidates on your own prompts; the hub guides cover the languages with the highest demand.

Do I need a separate endpoint per language?

No. There is one OpenAI-compatible chat endpoint; language is just text in the request. Any per-language behaviour belongs in your prompts and preprocessing.

How do I choose a model for a specific language?

Build a small gold set in that language, run two or three models from the catalogue, and score correctness, fluency and register separately with native reviewers.

What is plugsky-embed-multilingual?

It is a multilingual embedding model in the 30+ model catalogue, intended for cross-language retrieval so one vector collection can serve several languages.

How are non-Latin scripts measured?

By tokens, not characters. Counts per sentence vary widely between scripts, so use the token calculator on real text before sizing context windows.

Is there a free plan?

Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the paid catalogue for evaluation.

Can I deploy in my own region?

Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options for regulated markets.