Key facts
| Embeddings | plugsky-embed-multilingual is listed in the model catalogue |
| RAG endpoints | POST /v1/rag/collections and POST /v1/rag/query |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Formats | PDF, DOCX, TXT, MD and HTML ingestion |
| Citations | Ranked chunks return with source references such as file and page |
| Data handling | Per-collection encryption at rest; API data is not used to train models |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Product status | Live |
TL;DR
- Normalize Arabic text before embedding so spelling variants collapse.
- Use a multilingual embedding model for Arabic or mixed-language corpora.
- Keep mixed Arabic-English passages together when chunking.
- Evaluate with native Arabic questions, not translated English ones.
- Hybrid retrieval helps with names, acronyms and exact phrases.
How it works, step by step
- Audit the corpus for language mix, file formats and text quality issues.
- Apply consistent normalisation: diacritics, alef variants, tatweel and spacing.
- Choose a multilingual embedding model and test it on Arabic queries first.
- Chunk along Arabic paragraph and section boundaries, keeping mixed text intact.
- Ingest into a collection with metadata for source, language and access level.
- Evaluate retrieval with real Arabic questions, including colloquial phrasing.
- Tune hybrid weighting and reranking where exact terms or names matter.
Try it yourself
Open the embedding model comparison →
Why Arabic needs specific handling
Arabic text carries optional diacritics, multiple forms of the alef character, tatweel elongation and inconsistent use of spaces around punctuation. The same word can appear several ways in one corpus, and an embedding model trained mostly on English may map those variants to distant vectors. Normalisation before embedding is the single most effective fix: remove or standardise diacritics, unify alef forms, strip tatweel and collapse repeated whitespace.
Consider whether normalisation should also happen at query time. If documents are normalised but queries are not, matching quality suffers, so apply the same function in both paths and test it as part of ingestion.
Choosing embeddings and chunking
A multilingual embedding model is the safe default for Arabic corpora; an English-centric model will work on some queries and fail unpredictably on others. Plugsky lists plugsky-embed-multilingual in the catalogue, alongside plugsky-embed-v1 and plugsky-embed-large for other language profiles. Test candidates on Arabic questions with known answers before committing, because published rankings rarely reflect your domain.
Chunk along Arabic paragraph and section boundaries so each chunk remains a coherent unit. In mixed Arabic-English documents, keep both languages in the same chunk rather than splitting by script; the model sees the relationship, and citations stay meaningful.
Retrieval and evaluation for Arabic
Hybrid retrieval is usually worth enabling for Arabic: vector search handles paraphrase and morphological variation, while keyword search catches names, acronyms and exact phrases that embeddings may blur. Optional reranking helps when several chunks are plausible, and every result returns source references so answers can be traced back to the correct page.
Evaluate with questions written natively in Arabic, including how users actually phrase requests in your organisation, not translations of your English test set. Score recall on those queries and check whether the generated answers cite the right passages. Colloquial phrasing and regional vocabulary are common blind spots that translation-based tests never reveal.
Building it on Plugsky
The pipeline shape does not change for Arabic: upload documents to a collection, let chunking, embedding and indexing run automatically, then query with keyword, vector or hybrid retrieval and pass cited chunks to any of 30+ models. Data is encrypted at rest per collection and API data is not used to train models.
Compare embedding options with the embedding model comparison, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Concern | English-only model | Multilingual model | Translation-first pipeline |
|---|---|---|---|
| Arabic queries | Unpredictable quality | Trained across languages | Translated before search |
| Mixed documents | Weak on non-English passages | Handles both in one space | Meaning can shift in translation |
| Normalisation | Still required | Still required | Applied to source text |
| Latency and cost | One call | One call | Extra translation step |
| Citations | Source references preserved | Source references preserved | May lose alignment with original wording |
Frequently asked questions
Do I need a special Arabic embedding model?
A multilingual embedding model is the practical choice. It handles Arabic and mixed-language corpora without a separate pipeline.
Should I remove Arabic diacritics?
For retrieval, normalising diacritics usually improves matching because the same word appears in multiple forms across a corpus. Apply the same normalisation to queries.
Can I translate documents to English instead?
Translation adds cost, latency and meaning drift, and citations then point at translated text. Embedding the original Arabic with a multilingual model is simpler and keeps sources intact.
Does hybrid search help Arabic?
Yes. Combining keyword and vector retrieval catches exact names and phrases that embeddings alone may miss, and Plugsky supports keyword, vector and hybrid modes.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.