Key facts
| Japanese script | Kanji, hiragana and katakana with no word spaces; width variants matter |
| Tokenisation | Character-based encoding and katakana loanwords can raise token counts — measure with the token calculator |
| Register | Keigo, plain and casual forms are distinct — fix one per surface |
| Segmentation | Spaces are not delimiters, so indexing and chunking need a segmenter |
| API compatibility | OpenAI-compatible POST https://api.plugsky.com/v1/chat/completions — chat, streaming, JSON mode and function calling |
| Models | 30+ models behind one endpoint, from free tiers to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Product status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, images, moderation, files, batch, fine-tuning, assistants and responses coming soon |
TL;DR
- Japanese runs on the standard endpoint with one base_url change.
- Normalise full-width and half-width characters before counting or indexing.
- Segment before indexing — spaces do not delimit words.
- Fix one keigo level per surface and state it in the prompt.
- plugsky-embed-multilingual covers Japanese and English retrieval.
How it works, step by step
- Create a free Plugsky key and set base_url to https://api.plugsky.com/v1.
- Send real Japanese text to /v1/chat/completions and compare two or three models.
- Normalise width variants and normalise Unicode before measuring tokens or embedding.
- Add a glossary for product terms so katakana renderings stay consistent.
- Segment text with a Japanese tokenizer before keyword indexing or chunking.
- For RAG, embed with plugsky-embed-multilingual and test Japanese and English queries, then review outputs with native speakers.
Try it yourself
How Plugsky handles Japanese text
Japanese needs no special endpoint: UTF-8 text goes in and comes back. The language-specific work starts with writing system. Text mixes kanji, hiragana, katakana and Latin words in one sentence, and the same character can appear in full-width or half-width form — ABC versus ABC, 123 versus 123. Normalising width early stops near-duplicate strings from fragmenting your token counts and search indexes.
Register is the second decision. Keigo, plain and casual forms are not interchangeable, and business content expects a consistent politeness level across greetings, apologies and requests. Put the level in the system prompt and test it on realistic templates.
Tokenisation and cost in Japanese
Tokenisers split Japanese by characters, subwords or a mix, so counts per sentence depend heavily on the model and script mix. Katakana loanwords — product names, technical terms — often fragment into many pieces, and full-width Latin or digits add further variation.
- Measure on real product text, not on dictionary sentences.
- Normalise full-width and half-width characters first.
- Keep a glossary so loanword renderings stay stable.
- Re-measure whenever you change models.
Japanese retrieval and RAG
Keyword search needs a Japanese segmenter because spaces are not delimiters; embeddings help, and one multilingual collection built with plugsky-embed-multilingual serves Japanese documents with Japanese or English queries. Common failures come from width variants, loanword spelling and mixed scripts.
- Segment before indexing and chunking.
- Normalise width and Unicode form before embedding.
- Test English queries against Japanese documents explicitly.
- Keep one embedding model per collection.
Code example: a Japanese request
Point your OpenAI client at https://api.plugsky.com/v1; only the content and system prompt change.
client = OpenAI(base_url="https://api.plugsky.com/v1", api_key=os.environ["PLUGSKY_API_KEY"])
client.chat.completions.create(model="plugsky-pro", messages=[{"role": "system", "content": "ビジネス向けの丁寧な日本語で回答してください"}, {"role": "user", "content": "この契約書を三つのポイントに要約してください"}])
Start on the free plan with plugsky-micro and plugsky-lite, then compare paid models on a Japanese gold set. See the docs for the API reference.
Honest comparison
| Capability | Plugsky | Japanese apps today | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change base_url and model name | Varies by provider and SDK | Full rewrite |
| Japanese text handling | UTF-8 with width guidance, glossaries and keigo prompt control | Depends on provider tokeniser and prompt hygiene | You build normalisation, segmentation and evals |
| Token budget | Fixed tokeniser per model; measure script mix and chunk to fit | Varies by provider and model | You host and tune each tokeniser |
| Multilingual retrieval | plugsky-embed-multilingual for Japanese and English RAG | Often needs a separate embedding vendor | You serve and maintain embeddings |
| Sovereignty | Cloud, VPC, on-prem and air-gapped with residency options | Usually US/EU public endpoints | You own the full stack |
Frequently asked questions
Can Plugsky handle Japanese text?
Yes. The API accepts UTF-8 Japanese input across kanji, hiragana and katakana on the OpenAI-compatible chat endpoint. Compare models on your own prompts.
How do full-width and half-width characters matter?
The same visible text can be encoded differently, which changes token counts and search matching. Normalise width early and keep the display form separate.
Why does Japanese need segmentation?
Japanese is written without spaces, so keyword indexing and chunking need a segmenter rather than whitespace splitting. Embeddings complement, but do not replace, segmentation.
How do I control politeness level?
State the expected register — keigo, plain or casual — in the system prompt and test it on greetings, apologies and requests, where inconsistency shows first.
Is there a multilingual embedding model?
Yes — plugsky-embed-multilingual is part of the 30+ model catalogue and supports Japanese and English retrieval from one collection.
Can I deploy in my own environment?
Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment with residency options for regulated teams.
Is there a free plan?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial covers the paid catalogue.