Guides · Text models
Choosing a model for Southeast Asian languages.
This GotoRamp guide explains how to evaluate Chinese AI models for Malay, Indonesian, Thai, Vietnamese, Filipino, Chinese and Arabic with your own prompts and native-speaker reviewers, then route each language to the model that passed.
Last updated
Start from your own prompts
A general-purpose score can't tell you how a model handles your customers, your products and your languages. The only reliable way to choose is to run the prompts you actually receive through each candidate model and have people who speak the language judge the answers. Plan it once, keep the test set, and re-run it whenever a model version changes.
What goes in each language's test set
Collect a few dozen real prompts per language from your logs, tickets or chat history. Remove names, phone numbers, addresses and account numbers before the prompts leave your systems. Cover your common tasks (answering, summarizing, translating, extracting) and your hard cases.
| Language | Include | Watch for |
|---|---|---|
| Malay | Standard Malay and casual chat; Malay mixed with English | Formal anda versus casual awak; replies drifting into Indonesian vocabulary |
| Indonesian | Formal written Indonesian and everyday chat with slang and abbreviations | Anda versus kamu; titles such as Bapak and Ibu; replies drifting into Malay wording |
| Thai | Polite service phrasing and casual chat | A consistent polite particle (khrap or kha) for your assistant; the title Khun before names |
| Vietnamese | Messages typed with and without diacritics | Pronouns that depend on age and relationship (anh, chị, em, quý khách); missing or wrong tone marks |
| Filipino (Tagalog) | Tagalog, English and mixed messages | Po and opo for politeness; whether replies follow the customer's language mix |
| Chinese | Simplified and Traditional characters; Mandarin with local words used in Singapore and Malaysia | Replies in the wrong script; formal 您 versus 你 |
| Arabic | Modern Standard Arabic and Gulf dialect; right-to-left text mixed with English product names and numbers | Register (standard versus dialect); punctuation and number direction |
For each prompt, write down what a good answer must contain or must avoid. Reviewers score faster and agree more often when they check against written rules.
Code-switching
Real customers rarely stay in one language. Singlish (Singapore English with Malay, Hokkien and Tamil words and particles such as lah), Manglish (Malaysian English mixed with Malay, Chinese and Tamil) and Taglish (Tagalog and English in the same sentence) are normal in support chats. Include them in your test set in the proportion you actually see.
Test two separate things: whether the model understands the mixed message, and whether its reply follows your policy. Decide the policy first (reply in the customer's main language, reply in standard English, or mirror the mix), write it in the system prompt, and score against it.
Formality and politeness
In several of these languages, politeness is grammar, not just tone: pronouns, particles and titles change with the listener's age, status and relationship to the speaker. Decide the register for each channel, such as support chat, marketing copy and account notices, and state it in the system prompt with one example sentence. Then check that the model keeps the register over a long conversation and when the customer switches language.
Scripts, tokens and cost
Every model family splits text into tokens with its own tokenizer. The same sentence can produce a different number of tokens in each model, and text in some scripts, such as Thai, Arabic, Chinese, or Vietnamese with its diacritics, can take more tokens than the same meaning in English, depending on the tokenizer. GotoRamp bills text models per token, with input and output tokens at separate rates per model, so this affects cost as well as how much fits in the context window.
Measure it instead of estimating: send the same content in each language to each model and compare the token counts in each response's usage. Set max_tokens per language after you've measured. Thai also doesn't put spaces between words, so your own chunking, search and length limits need a Thai-aware word segmenter rather than a split on spaces.
Scoring with native speakers
- Run every prompt through every candidate model with the same system prompt and settings.
- Shuffle the outputs and hide the model names, so reviewers judge the text, not the brand.
- Have native speakers who know your market score each output against the criteria below, ideally two per language.
- Decide per language, and record the chosen model, a fallback and the date of the test.
| Criterion | Reviewer's question |
|---|---|
| Meaning | Is the answer correct and complete? |
| Fluency | Would a native speaker write it this way? |
| Register | Is the politeness level right for this channel? |
| Terminology | Are product names, local terms and units right? |
| Script and format | Right script, diacritics, dates, currency format and text direction? |
| Language choice | Did it reply in the language your policy asks for? |
| Safety | Did it refuse or hand over to a person when it should, and only then? |
Automated checks, including one model grading another, are useful for screening large test sets. Let a native speaker make the final call.
import csv, os, random
from openai import OpenAI
client = OpenAI(base_url="https://api.gotoramp.ai/v1",
api_key=os.environ["GOTORAMP_API_KEY"])
MODELS = ["deepseek-chat", "qwen-plus"] # use the IDs your account lists
SYSTEM = "You are the support assistant for a travel agency. Reply politely in the customer's language."
rows = []
with open("test_set.csv", encoding="utf-8") as f: # columns: id, language, prompt
for case in csv.DictReader(f):
for model in MODELS:
r = client.chat.completions.create(model=model, messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": case["prompt"]}])
rows.append({"id": case["id"], "language": case["language"], "model": model,
"output": r.choices[0].message.content,
"input_tokens": r.usage.prompt_tokens,
"output_tokens": r.usage.completion_tokens})
random.shuffle(rows) # give reviewers a copy without the "model" column
with open("review.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)Which model families to test
GotoRamp's text models open with early access. These are descriptions of each family, not rankings for any language; your test set decides.
| Family | Developer | Described as |
|---|---|---|
| DeepSeek | DeepSeek | Chat and reasoning models, widely used for coding, analysis and agent workflows |
| Qwen | Alibaba Cloud | A broad family (chat, reasoning, vision-language) with wide multilingual coverage |
| Kimi | Moonshot AI | Long-context and agent-oriented models |
| GLM | Zhipu AI (Z.ai) | General-purpose models used for coding and agents |
| Doubao Seed | ByteDance Seed | General chat and multimodal models |
Tool calling, JSON output, vision input and context length vary by model version. Check the console or GET /v1/models for what each one supports.
Route by language
Once each language has a winner, route in code: detect the language of the incoming message, look up the model for that language, and fall back to the second choice if the first fails or is rate-limited. Language detection often confuses Malay and Indonesian, so use the customer's country or app locale as a second signal.
Keep the routing table in configuration, not code. Log language, model, token usage and review results, so you can see cost per language and re-run the comparison when a model version changes.
Related
Start with a small balance. Scale when it works.
We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.