Key facts
| Arabic variants | MSA, regional dialects and Arabizi need separate prompts and evaluation slices |
| RTL text | Arabic is handled as UTF-8; isolate Latin spans so bidi order stays correct |
| Tokenisation | Clitics, letter variants and Latin-script Arabizi change token counts — measure with the token calculator |
| Embeddings | plugsky-embed-multilingual for Arabic, Arabizi and English retrieval |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required; 14-day full-access trial available |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Decide MSA, dialect or Arabizi per surface before writing prompts.
- Treat each Arabic variant as its own evaluation slice.
- RTL rendering needs bidi isolation, not just dir="rtl".
- Measure tokens on real text; Arabizi and letter variants change counts.
- One multilingual embedding collection can cover Arabic, Arabizi and English.
How it works, step by step
- Pick the variant per surface: Modern Standard Arabic for formal flows, one dialect for chat, Arabizi only where users actually type it.
- Create a free API key, set base_url to https://api.plugsky.com/v1 and send one prompt per variant to /v1/chat/completions.
- Measure tokens for each variant with the token calculator and set context budgets per surface.
- Write register instructions into the system prompt and keep them in version control.
- For RAG, embed Arabic, Arabizi and English content with plugsky-embed-multilingual and test mixed queries.
- Score models per variant with native speakers, then roll out surface by surface.
Original data
Try it yourself
Open the prompt diff and evaluator →
Why Arabic AI needs language-specific engineering
Arabic is a diglossic language: the formal written standard and everyday spoken dialects differ in grammar and vocabulary, and most text is written without short vowels. A product that ignores this ships prompts that drift between registers and evaluations that flatter models which only handle one variety well.
On the API side nothing special is required. Plugsky treats Arabic as UTF-8 text on the same OpenAI-compatible endpoint used for every other language. The engineering work is upstream: deciding what correct Arabic looks like for each surface, then measuring it.
Dialects, Arabizi and code-switching
Egyptian, Gulf, Levantine and Maghrebi Arabic differ enough that a prompt tuned for one can misread another. Arabizi adds a second problem: users type Arabic in Latin letters, often using digits for sounds that have no Latin equivalent — 3 for ع, 7 for ح, 9 for ق. Arabizi is informal, inconsistent and impossible to normalise with a single rule.
- Name the target variety explicitly in the system prompt.
- Keep Arabizi inputs in a separate evaluation set from scripted Arabic.
- Do not transliterate Arabizi before sending it — evaluate the model on the raw form users type.
- Where a flow needs one canonical output, convert at the edge, not mid-pipeline.
Evaluating Arabic output
Quality checks that work for English transfer poorly. Build a gold set per variant, include real user phrasing, and have native speakers score responses on correctness, register and fluency separately. Diacritic accuracy is a weak proxy for quality — most Arabic readers never see full tashkeel — while numerals, dates and personal names are common failure points worth testing directly.
Run the same gold set against two or three models from the catalogue and pick per surface. A model that wins on MSA may lose on Gulf chat text, and that is a product decision, not a ranking.
Code example: a Gulf Arabic support reply
One client handles every variant. The system prompt carries the register decision, so the request shape stays identical to any English call.
client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])
client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "اكتب باللهجة الخليجية وبأسلوب ودّي"}, {"role": "user", "content": "اكتب رد دعم للعميل يعتذر عن تأخر الطلب"}])
Start free with plugsky-micro and plugsky-lite, test stronger models during the 14-day full-access trial, and keep prompts in version control so register changes are reviewable. See the docs for the full request reference.
Honest comparison
| Capability | Plugsky | Typical Arabic AI stack | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — one endpoint for every variant | Often a mix of providers and SDKs | Full rewrite |
| Variant handling | MSA, dialects and Arabizi via prompt policy and evals | Depends on each model's training mix | You collect and curate dialect data |
| Token budget | Fixed tokeniser per model; measure and chunk to fit | Varies by provider | You host and tune each tokeniser |
| Multilingual retrieval | plugsky-embed-multilingual for Arabic and English RAG | Often a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped deployment options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Does Plugsky support Arabic dialects?
The API accepts any Arabic text, including dialect and Arabizi input. Quality per dialect varies by model, so test each candidate on your own dialect gold set before choosing.
What is Arabizi and can the API handle it?
Arabizi is Arabic written in Latin letters, often with digits for sounds such as 3, 7 and 9. The API accepts it as plain UTF-8 text; treat it as its own input class with dedicated prompts and evals.
How should I mix Arabic and English in one prompt?
It works, but state the expected output language in the system prompt and test mixed prompts explicitly — models often mirror the dominant language of the input.
Is RTL support automatic?
The API returns text; rendering is your responsibility. Set direction on the container and use bidi isolation around Latin spans so embedded names and URLs display in the right order.
How do I estimate Arabic token usage?
Use the token calculator on real prompts. Clitics, optional diacritics and Arabizi all change counts, so characters are a poor proxy.
Is there a free plan to prototype with?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the stronger models.
Can I deploy in my own environment?
Yes. Plugsky offers cloud, VPC, on-prem and air-gapped deployment with residency options for regulated teams.