Languages

How do you build Arabic AI apps with Plugsky?

Arabic is right-to-left, morphologically rich and usually written without short vowels. Plugsky serves it on the same OpenAI-compatible endpoint as English: set base_url to https://api.plugsky.com/v1 and send UTF-8 text to /v1/chat/completions. Measure tokens rather than characters, pin Modern Standard Arabic or one dialect per surface, and use plugsky-embed-multilingual for cross-language retrieval.

Key facts

Arabic scriptRight-to-left UTF-8 text handled as-is; wrap mixed Latin spans in bidi isolation marks
TokenisationAttached clitics and omitted short vowels add subwords — measure with the Plugsky token calculator
NormalisationUnify alef, ya and ta marbuta forms and Arabic-Indic digits for search only
Embeddingsplugsky-embed-multilingual for Arabic and cross-language RAG
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon

TL;DR

  • Arabic runs on the standard endpoint with one base_url change.
  • Budget by measured tokens — clitics and optional diacritics change the count.
  • Pin Modern Standard Arabic or one dialect per product surface.
  • Normalise letter and digit variants for search, not for display.
  • plugsky-embed-multilingual serves Arabic and English queries from one collection.

How it works, step by step

  1. Create a Plugsky API key on the free plan (no card) and set base_url to https://api.plugsky.com/v1.
  2. Send a set of real Arabic prompts to /v1/chat/completions — including one long document — and compare two or three models.
  3. Measure tokens for those prompts with the token calculator and size chunk lengths to fit your model context.
  4. Decide per surface whether you ship Modern Standard Arabic or a dialect, and state it in the system prompt.
  5. For RAG, embed with plugsky-embed-multilingual and test Arabic-only, English-only and mixed queries against the same collection.
  6. Score candidate models on an Arabic gold set with native-speaker review, then move production traffic.
1Create a PlugskyAPI key on the freeplan (no card) and2Send a set of realArabic prompts to/v1/chat/completions3Measure tokens forthose prompts withthe token4Decide per surfacewhether you shipModern Standard5For RAG, embed withplugsky-embed-multilingualand test6Score candidatemodels on an Arabicgold set with

Original data

Right-to-left Arabic scriptOpenAI-compatiAPI compatibility30+ models behModelsSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the RAG sandbox →

How Plugsky handles Arabic text

Arabic is written right to left in a cursive script, and everyday text usually omits short vowels (tashkeel). Plugsky passes Arabic UTF-8 through unchanged on the OpenAI-compatible endpoint, so there is no separate language route to maintain. What you do need is deliberate handling in three places: your UI, your prompts and your search layer.

In the UI, mixed Arabic and Latin spans — product names, code, URLs — should be wrapped in Unicode isolation marks so the visual order stays correct. In prompts, state whether you expect Modern Standard Arabic or a regional dialect; otherwise outputs drift between registers. In search, normalise letter variants while keeping the original text for display. Arabic diglossia is real: the written standard and spoken dialects differ enough that they deserve separate evaluation slices.

Tokenisation and cost in Arabic

Arabic is morphologically rich, and several high-frequency morphemes attach directly to words: the conjunction و, prepositions such as ب and ل, and the definite article ال. Tokenisers split these clitics into their own subwords, and optional diacritics add more when present. Arabic-Indic digits (٠١٢٣) may tokenise differently from Western digits, so standardise on one set early.

  • Count tokens on real product text, not on dictionary words.
  • Normalise tatweel and repeated characters before measuring.
  • Keep diacritics out of indexed text unless your content requires them.
  • Re-measure whenever you change models, because tokenisers differ.

Arabic retrieval and RAG

One multilingual collection can serve Arabic documents with Arabic or English queries. The main retrieval errors come from spelling variation, not from the model: alef forms (أ إ آ ا), final ya versus alef maqsura, and ta marbuta versus ha all produce different tokens for the same word. Normalise these for indexing and matching, and keep the original for display.

  • Embed with plugsky-embed-multilingual and keep one embedding model per collection.
  • Index dialect queries separately from MSA queries if your users mix both.
  • Test numeric queries with both digit sets.
  • Build a gold set of question-document pairs before launch.

Code example: an Arabic request

Point an existing OpenAI client at https://api.plugsky.com/v1 and send Arabic text in the content field. No language flag and no separate endpoint are required, and streaming, JSON mode and function calling keep the same shapes.

client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])

client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "أجب بالعربية الفصحى"}, {"role": "user", "content": "لخّص هذا العقد في ثلاث نقاط"}])

Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models on your Arabic gold set before cutover. See the docs for request details.

Honest comparison

CapabilityPlugskyArabic apps todayBuilding in-house
API compatibilityOpenAI-compatible — change base_url and model nameVaries by provider and SDKFull rewrite
Arabic text handlingRTL UTF-8, bidi guidance and register control via promptsDepends on provider tokeniser and prompt hygieneYou build normalisation, bidi handling and evals
Token budgetFixed tokeniser per model; measure and chunk to fitVaries by provider and modelYou host and tune each tokeniser
Multilingual retrievalplugsky-embed-multilingual for Arabic and cross-language RAGOften needs a separate embedding vendorYou serve and maintain embeddings
SovereigntyCloud, VPC, on-prem and air-gapped with residency optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Can Plugsky handle Arabic text?

Yes. The API accepts UTF-8 Arabic input on the OpenAI-compatible chat endpoint; output quality depends on the model, so compare two or three on your own prompts before choosing.

How do I estimate token usage for Arabic?

Arabic attaches clitics and often omits short vowels, so tokens do not track characters. Send real prompts through the token calculator and size context windows from those counts.

Should I use Modern Standard Arabic or a dialect?

Choose per surface. MSA suits formal, compliance and government content; a dialect suits chat and consumer apps. State the choice in the system prompt and evaluate it separately.

How do I handle mixed Arabic and Latin text?

Wrap Latin spans in Unicode isolation marks (or CSS bidi isolation) so ordering renders correctly, and test with real product data such as model names and URLs.

Is there a multilingual embedding model?

Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and is built for cross-language retrieval. Keep one embedding model per vector collection.

Can I keep data in my region?

Plugsky supports cloud, VPC, on-prem and air-gapped deployment with data-residency options; confirm requirements against the docs and with the enterprise team.

How do I migrate an existing app?

Change base_url to https://api.plugsky.com/v1 and map the model name. Streaming, JSON mode, function calling and embeddings keep the same request shapes.

Is there a free plan?

Yes — the free plan includes plugsky-micro and plugsky-lite with no card. A 14-day full-access trial unlocks the paid catalogue for evaluation.