RAG

How do you build RAG for Arabic documents?

Arabic RAG needs three adjustments: normalize text so diacritics and letter variants match, use a multilingual embedding model rather than an English-only one, and evaluate retrieval with real Arabic queries. Plugsky provides a multilingual embedding model and RAG collections with hybrid retrieval and citations, so the pipeline stays the same as for English.

Key facts

Embeddingsplugsky-embed-multilingual is listed in the model catalogue
RAG endpointsPOST /v1/rag/collections and POST /v1/rag/query
RetrievalKeyword, vector and hybrid search with optional reranking
FormatsPDF, DOCX, TXT, MD and HTML ingestion
CitationsRanked chunks return with source references such as file and page
Data handlingPer-collection encryption at rest; API data is not used to train models
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Normalize Arabic text before embedding so spelling variants collapse.
  • Use a multilingual embedding model for Arabic or mixed-language corpora.
  • Keep mixed Arabic-English passages together when chunking.
  • Evaluate with native Arabic questions, not translated English ones.
  • Hybrid retrieval helps with names, acronyms and exact phrases.

How it works, step by step

  1. Audit the corpus for language mix, file formats and text quality issues.
  2. Apply consistent normalisation: diacritics, alef variants, tatweel and spacing.
  3. Choose a multilingual embedding model and test it on Arabic queries first.
  4. Chunk along Arabic paragraph and section boundaries, keeping mixed text intact.
  5. Ingest into a collection with metadata for source, language and access level.
  6. Evaluate retrieval with real Arabic questions, including colloquial phrasing.
  7. Tune hybrid weighting and reranking where exact terms or names matter.
1Audit the corpusfor language mix,file formats and2Apply consistentnormalisation:diacritics, alef3Choose amultilingualembedding model and4Chunk along Arabicparagraph andsection boundaries,5Ingest into acollection withmetadata for6Evaluate retrievalwith real Arabicquestions,

Try it yourself

Open the embedding model comparison →

Why Arabic needs specific handling

Arabic text carries optional diacritics, multiple forms of the alef character, tatweel elongation and inconsistent use of spaces around punctuation. The same word can appear several ways in one corpus, and an embedding model trained mostly on English may map those variants to distant vectors. Normalisation before embedding is the single most effective fix: remove or standardise diacritics, unify alef forms, strip tatweel and collapse repeated whitespace.

Consider whether normalisation should also happen at query time. If documents are normalised but queries are not, matching quality suffers, so apply the same function in both paths and test it as part of ingestion.

Choosing embeddings and chunking

A multilingual embedding model is the safe default for Arabic corpora; an English-centric model will work on some queries and fail unpredictably on others. Plugsky lists plugsky-embed-multilingual in the catalogue, alongside plugsky-embed-v1 and plugsky-embed-large for other language profiles. Test candidates on Arabic questions with known answers before committing, because published rankings rarely reflect your domain.

Chunk along Arabic paragraph and section boundaries so each chunk remains a coherent unit. In mixed Arabic-English documents, keep both languages in the same chunk rather than splitting by script; the model sees the relationship, and citations stay meaningful.

Retrieval and evaluation for Arabic

Hybrid retrieval is usually worth enabling for Arabic: vector search handles paraphrase and morphological variation, while keyword search catches names, acronyms and exact phrases that embeddings may blur. Optional reranking helps when several chunks are plausible, and every result returns source references so answers can be traced back to the correct page.

Evaluate with questions written natively in Arabic, including how users actually phrase requests in your organisation, not translations of your English test set. Score recall on those queries and check whether the generated answers cite the right passages. Colloquial phrasing and regional vocabulary are common blind spots that translation-based tests never reveal.

Building it on Plugsky

The pipeline shape does not change for Arabic: upload documents to a collection, let chunking, embedding and indexing run automatically, then query with keyword, vector or hybrid retrieval and pass cited chunks to any of 30+ models. Data is encrypted at rest per collection and API data is not used to train models.

Compare embedding options with the embedding model comparison, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ConcernEnglish-only modelMultilingual modelTranslation-first pipeline
Arabic queriesUnpredictable qualityTrained across languagesTranslated before search
Mixed documentsWeak on non-English passagesHandles both in one spaceMeaning can shift in translation
NormalisationStill requiredStill requiredApplied to source text
Latency and costOne callOne callExtra translation step
CitationsSource references preservedSource references preservedMay lose alignment with original wording

Frequently asked questions

Do I need a special Arabic embedding model?

A multilingual embedding model is the practical choice. It handles Arabic and mixed-language corpora without a separate pipeline.

Should I remove Arabic diacritics?

For retrieval, normalising diacritics usually improves matching because the same word appears in multiple forms across a corpus. Apply the same normalisation to queries.

Can I translate documents to English instead?

Translation adds cost, latency and meaning drift, and citations then point at translated text. Embedding the original Arabic with a multilingual model is simpler and keeps sources intact.

Does hybrid search help Arabic?

Yes. Combining keyword and vector retrieval catches exact names and phrases that embeddings alone may miss, and Plugsky supports keyword, vector and hybrid modes.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.