Key facts
| Polish script | Latin script with diacritics ą, ć, ę, ł, ń, ó, ś, ź and ż |
| Inflection | Seven grammatical cases and gender produce many surface forms per word |
| Register | Formal Pan and Pani versus informal ty — keep one per surface |
| Tokenisation | Inflected endings and consonant clusters split into subwords — measure with the token calculator |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Polish runs on the standard endpoint with one base_url change.
- Inflection multiplies word forms — pair embeddings with stemming for search.
- Add diacritic-insensitive keys for mobile users.
- Pin Pan or Pani formality per surface and keep it consistent.
- plugsky-embed-multilingual covers Polish and English retrieval.
How it works, step by step
- Create a free Plugsky key and set base_url to https://api.plugsky.com/v1.
- Send real Polish text to /v1/chat/completions and compare two or three models.
- Measure tokens with the token calculator on inflected, diacritic-heavy text.
- Add stemming for keyword search so inflected forms match, alongside embeddings.
- Choose Pan or Pani versus ty per surface, and state it in the system prompt.
- For RAG, embed with plugsky-embed-multilingual and review a Polish gold set with native speakers.
Try it yourself
How Plugsky handles Polish text
Polish needs no special endpoint: UTF-8 text with diacritics goes in and comes back. The language-specific work is inflection and formality. Polish marks seven grammatical cases across nouns, adjectives and pronouns, and gender affects endings too, so a single word appears in many forms depending on its role in the sentence.
Formality is explicit rather than implied: Pan and Pani with third-person verb forms for formal address, ty for informal. Mixing them inside one conversation is the failure Polish readers notice first, so fix the choice per surface.
Tokenisation and cost in Polish
Polish text carries diacritics and long consonant clusters, and inflected endings can split into subwords, so counts differ from English in both directions. Formal and legal writing compounds these effects with longer sentence structures.
- Measure on inflected text, not on dictionary forms.
- Keep diacritics for display and embedding; fold them for search keys only.
- Watch hyphenated and abbreviated forms in official writing.
- Re-measure whenever models change.
Polish retrieval and RAG
Inflection is the classic Polish retrieval problem: the same concept appears in several case forms, and keyword search treats them as different strings. One multilingual collection built with plugsky-embed-multilingual handles that better, and stemming covers the keyword path.
- Add stemming or lemmatisation for keyword search on top of embeddings.
- Add diacritic-insensitive keys for mobile users who skip accents.
- Test inflected query forms against base-form documents.
- Keep one embedding model per collection.
Code example: a Polish request
Point your OpenAI client at https://api.plugsky.com/v1; only the content and system prompt change.
client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])
client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "Odpowiadaj po polsku, w formalnym stylu z formą Pan lub Pani"}, {"role": "user", "content": "Streść tę umowę w trzech punktach"}])
Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models on a Polish gold set. See the docs for the API reference.
Honest comparison
| Capability | Plugsky | Polish apps today | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change base_url and model name | Varies by provider and SDK | Full rewrite |
| Polish text handling | UTF-8 with diacritics, inflection-aware search and Pan/Pani prompts | Depends on provider tokeniser and preprocessing | You build normalisation, stemming and evals |
| Token budget | Fixed tokeniser per model; measure inflected text and chunk to fit | Varies by provider and model | You host and tune each tokeniser |
| Multilingual retrieval | plugsky-embed-multilingual for Polish and English RAG | Often needs a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped with residency options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Can Plugsky handle Polish text?
Yes. The API accepts UTF-8 Polish input with diacritics on the OpenAI-compatible chat endpoint. Compare two or three models on your own prompts before choosing.
How does Polish inflection affect search?
The same word appears in many case and gender forms, so keyword search misses inflected queries. Use embeddings and add stemming or lemmatisation for keyword keys.
Should prompts use Pan or Pani?
Formal service, B2B and public-sector flows use Pan or Pani with third-person verb forms; consumer and community products often use ty. State the choice in the system prompt.
Do I need diacritic-insensitive search?
Many users type without Polish diacritics on mobile. Add an accent-insensitive path while keeping diacritics for display and embeddings.
Is there a multilingual embedding model?
Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports Polish and English retrieval from one collection.
Can I keep data in the EU?
Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options; confirm specifics with the docs and the enterprise team.
Is there a free plan?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the paid catalogue.