Key facts
| Portuguese text | Latin script with accents, cedilla and nasal vowels (ã, õ) handled as UTF-8 |
| Locale variants | pt-BR and pt-PT differ in vocabulary, pronouns and clitic placement |
| Clitics | Hyphenated clitic pronouns such as dar-lhe can split into subwords — measure with the token calculator |
| Register | você is standard in Brazil; o senhor or a senhora carry formality in both variants |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Portuguese runs on the standard endpoint with one base_url change.
- Choose pt-BR or pt-PT per surface; do not mix them silently.
- Hyphenated clitics and accents can split into extra subwords.
- Keep accents and cedilla for display and embedding.
- plugsky-embed-multilingual serves both variants from one collection.
How it works, step by step
- Create a free Plugsky key and set base_url to https://api.plugsky.com/v1.
- Send real pt-BR and pt-PT prompts to /v1/chat/completions and compare two or three models on each.
- Measure tokens with the token calculator, including hyphenated clitic forms.
- Choose the variant and formality per surface, and state the choice in the system prompt.
- For RAG, embed with plugsky-embed-multilingual and test both variants plus English queries.
- Score outputs per variant with native speakers before cutover.
Original data
Try it yourself
Open the embedding model comparison →
How Plugsky handles Portuguese text
Portuguese needs no special route: send UTF-8 text and the API returns it with accents, cedilla and nasal vowels intact. The language-level decisions are locale and formality. Brazilian Portuguese uses você widely and places clitic pronouns before the verb; European Portuguese uses clitic forms differently and keeps o senhor and a senhora for formal address.
Vocabulary differs too — celular versus telemóvel, trem versus comboio. If your audience spans both markets, either localise per market or choose wording that reads naturally in both.
Tokenisation and cost in Portuguese
Portuguese sits close to English in token cost. The splitting points are accents and hyphenated clitics: dar-lhe, enviá-lo, fá-lo-á in formal European Portuguese. Brazilian writing uses fewer of these forms, so the same content can measure differently depending on the variant.
- Measure each variant separately, not as one blended sample.
- Normalise hyphen and apostrophe characters early.
- Keep accents for display and embedding; fold for search keys only.
- Re-measure after model changes.
Portuguese retrieval and RAG
One multilingual collection built with plugsky-embed-multilingual serves Portuguese documents from either variant with Portuguese or English queries. Retrieval errors usually come from locale vocabulary and accent folding rather than from the embedding model.
- Add accent-insensitive keys for keyword search alongside embeddings.
- Keep pt-BR and pt-PT query sets separate during evaluation.
- Test English queries against Portuguese documents explicitly.
- Keep one embedding model per collection.
Code example: a pt-BR request
Point your OpenAI client at https://api.plugsky.com/v1; only the content and system prompt change.
client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])
client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "Responda em português do Brasil, com tratamento por você"}, {"role": "user", "content": "Resuma este contrato em três pontos"}])
Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models per variant. See the docs for the API reference.
Honest comparison
| Capability | Plugsky | Portuguese apps today | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change base_url and model name | Varies by provider and SDK | Full rewrite |
| Locale coverage | pt-BR or pt-PT via prompt policy and separate evaluation | Depends on provider training mix | You curate data per variant |
| Token budget | Fixed tokeniser per model; measure each variant and chunk to fit | Varies by provider and model | You host and tune each tokeniser |
| Multilingual retrieval | plugsky-embed-multilingual for Portuguese and English RAG | Often needs a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped with residency options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Can Plugsky handle Portuguese text?
Yes. The API accepts UTF-8 Portuguese input, including accents, cedilla and nasal vowels, on the OpenAI-compatible chat endpoint. Compare models on your own prompts.
pt-BR or pt-PT?
Pick per surface and audience. Brazilian Portuguese uses você and pre-verbal clitics; European Portuguese differs in placement and vocabulary. State the choice in the system prompt.
How do clitics affect token counts?
Hyphenated clitic forms such as dar-lhe can split into several subwords, especially in formal European Portuguese. Measure each variant with the token calculator.
Can one search collection serve both variants?
Yes when you embed with plugsky-embed-multilingual, but keep evaluation slices separate and add accent-insensitive keys for keyword search.
Is there a multilingual embedding model?
Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports Portuguese and English retrieval from one collection.
Can I keep data in my region?
Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options; confirm specifics with the docs and the enterprise team.
Is there a free plan?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the paid catalogue.