Languages

How do you build Portuguese AI apps with Plugsky?

Portuguese needs no special endpoint: point base_url at https://api.plugsky.com/v1 and send UTF-8 text to /v1/chat/completions. Choose pt-BR or pt-PT vocabulary, NFC-normalise accents, and watch hyphenated clitics in token counts. plugsky-embed-multilingual lets one collection serve Portuguese and English queries across both variants.

Key facts

Portuguese textLatin script with accents, cedilla and nasal vowels — handle as UTF-8
Locale variantspt-BR and pt-PT differ in vocabulary and clitic placement
TokenisationAccented and hyphenated forms can split into subwords — measure with the Plugsky token calculator
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required
DeploymentPlugsky cloud, your VPC, on-prem and air-gapped options
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, batch and fine-tuning coming soon

TL;DR

  • Portuguese uses the same API — one base_url change.
  • NFC-normalise accents and cedillas before indexing.
  • Pin pt-BR or pt-PT vocabulary per product.
  • Hyphenated clitics can inflate token counts — measure.
  • plugsky-embed-multilingual covers Portuguese and English retrieval.

How it works, step by step

  1. Create a Plugsky API key on the free plan (no card) and set base_url to https://api.plugsky.com/v1.
  2. Send a small set of real Portuguese prompts to /v1/chat/completions and compare output across two or three models.
  3. Count tokens for those prompts with the Plugsky token calculator and set chunk sizes that fit your model context.
  4. Normalise text before indexing or prompting: NFC-normalise accents and preserve cedillas.
  5. For RAG, embed with plugsky-embed-multilingual and test cross-language queries alongside Portuguese-only queries.
  6. Score candidate models on a Portuguese gold set with native-speaker review, then cut production traffic over.
1Create a PlugskyAPI key on the freeplan (no card) and2Send a small set ofreal Portugueseprompts to3Count tokens forthose prompts withthe Plugsky token4Normalise textbefore indexing orprompting:5For RAG, embed withplugsky-embed-multilingualand test6Score candidatemodels on aPortuguese gold set

Original data

Latin script wPortuguese textOpenAI-compatiAPI compatibility30+ models behModelsSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the OpenAI cost calculator →

How Plugsky handles Portuguese text

Portuguese uses the Latin alphabet with accents, cedilla (ç) and nasal vowels (ã, õ), and reads left to right.

Brazilian and European Portuguese differ in vocabulary, pronoun placement and gerund usage, and formal o senhor or você versus informal você or tu varies by region.

Evaluate on pt-BR and pt-PT samples; check accents, clitic placement and vocabulary consistency, which differ enough to matter for local readers.

Tokenisation and cost in Portuguese

Portuguese is close to English in token efficiency; accented characters and clitic pronouns such as dar-lhe can split into subwords, especially in hyphenated forms.

  • Normalise to NFC so accents stay composed.
  • Pick pt-BR or pt-PT vocabulary and keep it consistent.
  • Keep hyphenated clitics intact in display text.
  • Measure tokens on legal or formal copy where clitics concentrate.

Portuguese retrieval and RAG

Serve pt-BR and pt-PT from one collection but track separate eval slices, since vocabulary differences affect retrieval precision.

  • Use plugsky-embed-multilingual for Portuguese and English retrieval.
  • NFC-normalise accents before embedding.
  • Track pt-BR and pt-PT queries separately.
  • Add synonyms for cross-variant vocabulary.

Code example: a Portuguese request

Point your existing OpenAI client at https://api.plugsky.com/v1 and pass Portuguese text in the content field — no language flag and no separate endpoint. Streaming, JSON mode and function calling keep the same request shapes.

client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])

client.chat.completions.create(model="plugsky-pro", messages=[{"role": "user", "content": "Resuma este contrato em três pontos, em português."}])

Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models on a Portuguese gold set before cutover. See the docs for request details.

For production, log the model name and your normalisation settings with each request, and re-run the Portuguese gold set whenever either changes — language quality regressions usually come from prompt or preprocessing drift, not from the model alone.

Honest comparison

CapabilityPlugskyPortuguese workflow todayBuilding in-house
API compatibilityOpenAI-compatible — change base_url and model nameVaries by provider and SDKFull rewrite
Portuguese text handlingOne variant's vocabulary with NFC-normalised accentsDepends on provider tokeniser and prompt hygieneYou build normalisation, segmentation and evals
Token budgetFixed tokeniser per model; measure with the Plugsky token calculator and chunk to fitVaries by provider and modelYou host and tune each tokeniser
Multilingual retrievalplugsky-embed-multilingual available for cross-language RAGOften needs a separate embedding vendorYou serve and maintain embeddings
SovereigntyCloud, VPC, on-prem and air-gapped with residency optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Can Plugsky handle Portuguese text?

Yes. The API accepts UTF-8 Portuguese input on the OpenAI-compatible chat endpoint; output quality depends on the model, so compare two or three on your own prompts before choosing.

How do I estimate token usage for Portuguese?

Portuguese is close to English in token efficiency; accented characters and clitic pronouns such as dar-lhe can split into subwords, especially in hyphenated forms. Use the token calculator at /tools/llm-token-calculator before sizing context windows or chunk lengths.

Brazilian or European Portuguese?

Pick one variant per product and keep vocabulary consistent. pt-BR has the larger developer base, but pt-PT matters for Portugal-facing services.

Do accents affect retrieval?

NFC-normalise accents before embedding. Keep accents in display text — removing them changes meaning and reads as low quality.

Is there a multilingual embedding model?

Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and is built for cross-language retrieval. Keep one embedding model per vector collection.

Can I keep data in my region?

Plugsky supports cloud, VPC, on-prem and air-gapped deployment with data-residency options; confirm your requirements with the docs and the enterprise team.

How do I migrate an existing app?

Change base_url to https://api.plugsky.com/v1 and map the model name. Streaming, JSON mode, function calling and embeddings keep the same request shapes.

Is there a free plan?

Yes — the free plan includes two free models, plugsky-micro and plugsky-lite, with no card. A 14-day full-access trial unlocks the paid catalogue.