Languages

How do you build Finnish AI apps with Plugsky?

Finnish is agglutinative: case endings and clitics stack onto stems, and long compounds appear in ordinary text, so token counts run higher than English. Plugsky serves Finnish on the standard OpenAI-compatible endpoint — set base_url to https://api.plugsky.com/v1 and post to /v1/chat/completions. Budget tokens per word, choose standard kirjakieli or colloquial puhekieli per surface, and use plugsky-embed-multilingual for cross-language retrieval.

Key facts

Finnish grammarAgglutinative: roughly fifteen grammatical cases and clitics attach to stems, lengthening words
CompoundsLong compound nouns split into many subwords — measure with the tokenizer
RegisterStandard kirjakieli and spoken puhekieli differ; pick one per surface
Function wordsNo grammatical gender or articles, but case endings still add tokens
API compatibilityOpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling
Models30+ models behind one endpoint, from free tiers to frontier reasoning
Free tierFree plan with plugsky-micro and plugsky-lite, no card required
Product statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon

TL;DR

  • Finnish works on the standard endpoint with one base_url change.
  • Expect more tokens per word: case endings and compounds fragment.
  • Choose kirjakieli or puhekieli per surface and state it in the prompt.
  • Measure on official documents, where compounds are longest.
  • plugsky-embed-multilingual handles Finnish and English retrieval.

How it works, step by step

  1. Create a free Plugsky key and set base_url to https://api.plugsky.com/v1.
  2. Send real Finnish text to /v1/chat/completions — include one official document — and compare two or three models.
  3. Measure tokens with the tokenizer; set chunk sizes from compound-heavy samples, not averages.
  4. Decide per surface between standard kirjakieli and conversational puhekieli.
  5. For RAG, embed with plugsky-embed-multilingual and test Finnish queries against English documents and vice versa.
  6. Review outputs with native speakers on a fixed gold set, then move production traffic.
1Create a freePlugsky key and setbase_url to2Send real Finnishtext to/v1/chat/completions3Measure tokens withthe tokenizer; setchunk sizes from4Decide per surfacebetween standardkirjakieli and5For RAG, embed withplugsky-embed-multilingualand test Finnish6Review outputs withnative speakers ona fixed gold set,

Try it yourself

Open the tokenizer →

How Plugsky handles Finnish text

Finnish needs no special route: the API accepts UTF-8 text with ä and ö and returns text in the same encoding. The engineering work is about length and register. Agglutination stacks case endings, possessives and clitics onto stems, so a single word can carry what English expresses with several words — and the tokeniser still splits it into pieces.

Register is the second decision. Standard written Finnish (kirjakieli) is expected in official, legal and news contexts; spoken Finnish (puhekieli) looks very different, with shortened pronouns and dropped endings. Products that mix them read as careless, so pick one per surface and keep it in the system prompt.

Tokenisation and cost in Finnish

Budget more tokens per concept than English. Case endings such as -ssa, -sta, -lle and clitics like -kin and -ko attach to nouns and verbs, and compounds such as tietosuojavaltuutettu split into multiple subwords. Long number words add to the load in invoices and contracts.

  • Measure on official and technical text, where compounds cluster.
  • Leave more context headroom than you would for English.
  • Keep endings attached for display; stem only for search keys.
  • Re-measure when models change.

Finnish retrieval and RAG

Inflected forms are the classic retrieval problem in Finnish: the same noun appears in many case forms, so keyword search misses matches that embeddings catch. One multilingual collection built with plugsky-embed-multilingual serves Finnish documents with Finnish or English queries.

  • Add lemmatisation or stemming for keyword search on top of embeddings.
  • Test queries written in puhekieli against kirjakieli documents.
  • Include compound variants (joined, hyphenated, split) in test sets.
  • Keep one embedding model per collection.

Code example: a Finnish request

Point your OpenAI client at https://api.plugsky.com/v1. The request shape stays the same as any English call; only the content and system prompt change.

client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])

client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "Vastaa suomeksi selkeällä yleiskielellä"}, {"role": "user", "content": "Tiivistä tämä sopimus kolmeen kohtaan"}])

Prototype on the free plan with plugsky-micro and plugsky-lite, then compare paid models on a Finnish gold set. See the docs for the API reference.

Honest comparison

CapabilityPlugskyFinnish apps todayBuilding in-house
API compatibilityOpenAI-compatible — change base_url and model nameVaries by provider and SDKFull rewrite
Finnish text handlingUTF-8 with kirjakieli or puhekieli prompt controlDepends on provider tokeniser and prompt hygieneYou build normalisation and evals
Token budgetFixed tokeniser per model; expect more tokens per word and chunk accordinglyVaries by provider and modelYou host and tune each tokeniser
Multilingual retrievalplugsky-embed-multilingual for Finnish and English RAGOften needs a separate embedding vendorYou serve and maintain embeddings
SovereigntyCloud, VPC, on-prem and air-gapped with residency optionsUsually US/EU public endpointsYou own the full stack

Frequently asked questions

Can Plugsky handle Finnish text?

Yes. The API accepts UTF-8 Finnish input with ä and ö on the OpenAI-compatible chat endpoint. Quality varies by model, so compare candidates on your own prompts.

Why does Finnish use more tokens than English?

Finnish attaches case endings and clitics to stems and forms long compounds, so one word can split into several subwords. Use the tokenizer on real text before sizing context windows.

Should I write prompts in kirjakieli or puhekieli?

Match the surface. Official and legal content expects kirjakieli; chat and consumer flows often sound better in puhekieli. State the choice in the system prompt.

How does Finnish search handle inflected forms?

Inflected forms of the same word differ, so keyword search misses matches. Pair embeddings with lemmatisation or stemming for keyword keys.

Is there a multilingual embedding model?

Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports Finnish and English retrieval from one collection.

Can I deploy in my own environment?

Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options for regulated teams.

Is there a free plan?

Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers stronger models.