Languages

How do you build Arabic AI products with Plugsky?

Arabic AI is less about translation and more about register: Modern Standard Arabic, regional dialects and Latin-script Arabizi behave like different inputs. Plugsky serves all of them through one OpenAI-compatible endpoint — set base_url to https://api.plugsky.com/v1, send UTF-8 text to /v1/chat/completions, pin the register in the prompt and evaluate each variant on its own gold set.

Key facts

Arabic variantsMSA, regional dialects and Arabizi need separate prompts and evaluation slices
RTL textArabic is handled as UTF-8; isolate Latin spans so bidi order stays correct
TokenisationClitics, letter variants and Latin-script Arabizi change token counts — measure with the token calculator
Embeddingsplugsky-embed-multilingual for Arabic, Arabizi and English retrieval
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required; 14-day full-access trial available
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon

TL;DR

  • Decide MSA, dialect or Arabizi per surface before writing prompts.
  • Treat each Arabic variant as its own evaluation slice.
  • RTL rendering needs bidi isolation, not just dir="rtl".
  • Measure tokens on real text; Arabizi and letter variants change counts.
  • One multilingual embedding collection can cover Arabic, Arabizi and English.

How it works, step by step

  1. Pick the variant per surface: Modern Standard Arabic for formal flows, one dialect for chat, Arabizi only where users actually type it.
  2. Create a free API key, set base_url to https://api.plugsky.com/v1 and send one prompt per variant to /v1/chat/completions.
  3. Measure tokens for each variant with the token calculator and set context budgets per surface.
  4. Write register instructions into the system prompt and keep them in version control.
  5. For RAG, embed Arabic, Arabizi and English content with plugsky-embed-multilingual and test mixed queries.
  6. Score models per variant with native speakers, then roll out surface by surface.
1Pick the variantper surface: ModernStandard Arabic for2Create a free APIkey, set base_urlto3Measure tokens foreach variant withthe token4Write registerinstructions intothe system prompt5For RAG, embedArabic, Arabizi andEnglish content6Score models pervariant with nativespeakers, then roll

Original data

Arabic is handRTL textOpenAI-compatiAPI compatibility30+ models behModelsFree plan withFree tierSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the prompt diff and evaluator →

Why Arabic AI needs language-specific engineering

Arabic is a diglossic language: the formal written standard and everyday spoken dialects differ in grammar and vocabulary, and most text is written without short vowels. A product that ignores this ships prompts that drift between registers and evaluations that flatter models which only handle one variety well.

On the API side nothing special is required. Plugsky treats Arabic as UTF-8 text on the same OpenAI-compatible endpoint used for every other language. The engineering work is upstream: deciding what correct Arabic looks like for each surface, then measuring it.

Dialects, Arabizi and code-switching

Egyptian, Gulf, Levantine and Maghrebi Arabic differ enough that a prompt tuned for one can misread another. Arabizi adds a second problem: users type Arabic in Latin letters, often using digits for sounds that have no Latin equivalent — 3 for ع, 7 for ح, 9 for ق. Arabizi is informal, inconsistent and impossible to normalise with a single rule.

  • Name the target variety explicitly in the system prompt.
  • Keep Arabizi inputs in a separate evaluation set from scripted Arabic.
  • Do not transliterate Arabizi before sending it — evaluate the model on the raw form users type.
  • Where a flow needs one canonical output, convert at the edge, not mid-pipeline.

Evaluating Arabic output

Quality checks that work for English transfer poorly. Build a gold set per variant, include real user phrasing, and have native speakers score responses on correctness, register and fluency separately. Diacritic accuracy is a weak proxy for quality — most Arabic readers never see full tashkeel — while numerals, dates and personal names are common failure points worth testing directly.

Run the same gold set against two or three models from the catalogue and pick per surface. A model that wins on MSA may lose on Gulf chat text, and that is a product decision, not a ranking.

Code example: a Gulf Arabic support reply

One client handles every variant. The system prompt carries the register decision, so the request shape stays identical to any English call.

client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])

client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "اكتب باللهجة الخليجية وبأسلوب ودّي"}, {"role": "user", "content": "اكتب رد دعم للعميل يعتذر عن تأخر الطلب"}])

Start free with plugsky-micro and plugsky-lite, test stronger models during the 14-day full-access trial, and keep prompts in version control so register changes are reviewable. See the docs for the full request reference.

Honest comparison

CapabilityPlugskyTypical Arabic AI stackBuilding in-house
API compatibilityOpenAI-compatible — one endpoint for every variantOften a mix of providers and SDKsFull rewrite
Variant handlingMSA, dialects and Arabizi via prompt policy and evalsDepends on each model's training mixYou collect and curate dialect data
Token budgetFixed tokeniser per model; measure and chunk to fitVaries by providerYou host and tune each tokeniser
Multilingual retrievalplugsky-embed-multilingual for Arabic and English RAGOften a separate embedding vendorYou serve and maintain embeddings
SovereigntyCloud, VPC, on-prem and air-gapped deployment optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Does Plugsky support Arabic dialects?

The API accepts any Arabic text, including dialect and Arabizi input. Quality per dialect varies by model, so test each candidate on your own dialect gold set before choosing.

What is Arabizi and can the API handle it?

Arabizi is Arabic written in Latin letters, often with digits for sounds such as 3, 7 and 9. The API accepts it as plain UTF-8 text; treat it as its own input class with dedicated prompts and evals.

How should I mix Arabic and English in one prompt?

It works, but state the expected output language in the system prompt and test mixed prompts explicitly — models often mirror the dominant language of the input.

Is RTL support automatic?

The API returns text; rendering is your responsibility. Set direction on the container and use bidi isolation around Latin spans so embedded names and URLs display in the right order.

How do I estimate Arabic token usage?

Use the token calculator on real prompts. Clitics, optional diacritics and Arabizi all change counts, so characters are a poor proxy.

Is there a free plan to prototype with?

Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the stronger models.

Can I deploy in my own environment?

Yes. Plugsky offers cloud, VPC, on-prem and air-gapped deployment with residency options for regulated teams.