GotoRamp

Guides · Text models

Choosing a model for Southeast Asian languages.

This GotoRamp guide explains how to evaluate Chinese AI models for Malay, Indonesian, Thai, Vietnamese, Filipino, Chinese and Arabic with your own prompts and native-speaker reviewers, then route each language to the model that passed.

Last updated

Start from your own prompts

A general-purpose score can't tell you how a model handles your customers, your products and your languages. The only reliable way to choose is to run the prompts you actually receive through each candidate model and have people who speak the language judge the answers. Plan it once, keep the test set, and re-run it whenever a model version changes.

What goes in each language's test set

Collect a few dozen real prompts per language from your logs, tickets or chat history. Remove names, phone numbers, addresses and account numbers before the prompts leave your systems. Cover your common tasks (answering, summarizing, translating, extracting) and your hard cases.

LanguageIncludeWatch for
MalayStandard Malay and casual chat; Malay mixed with EnglishFormal anda versus casual awak; replies drifting into Indonesian vocabulary
IndonesianFormal written Indonesian and everyday chat with slang and abbreviationsAnda versus kamu; titles such as Bapak and Ibu; replies drifting into Malay wording
ThaiPolite service phrasing and casual chatA consistent polite particle (khrap or kha) for your assistant; the title Khun before names
VietnameseMessages typed with and without diacriticsPronouns that depend on age and relationship (anh, chị, em, quý khách); missing or wrong tone marks
Filipino (Tagalog)Tagalog, English and mixed messagesPo and opo for politeness; whether replies follow the customer's language mix
ChineseSimplified and Traditional characters; Mandarin with local words used in Singapore and MalaysiaReplies in the wrong script; formal 您 versus 你
ArabicModern Standard Arabic and Gulf dialect; right-to-left text mixed with English product names and numbersRegister (standard versus dialect); punctuation and number direction

For each prompt, write down what a good answer must contain or must avoid. Reviewers score faster and agree more often when they check against written rules.

Code-switching

Real customers rarely stay in one language. Singlish (Singapore English with Malay, Hokkien and Tamil words and particles such as lah), Manglish (Malaysian English mixed with Malay, Chinese and Tamil) and Taglish (Tagalog and English in the same sentence) are normal in support chats. Include them in your test set in the proportion you actually see.

Test two separate things: whether the model understands the mixed message, and whether its reply follows your policy. Decide the policy first (reply in the customer's main language, reply in standard English, or mirror the mix), write it in the system prompt, and score against it.

Formality and politeness

In several of these languages, politeness is grammar, not just tone: pronouns, particles and titles change with the listener's age, status and relationship to the speaker. Decide the register for each channel, such as support chat, marketing copy and account notices, and state it in the system prompt with one example sentence. Then check that the model keeps the register over a long conversation and when the customer switches language.

Scripts, tokens and cost

Every model family splits text into tokens with its own tokenizer. The same sentence can produce a different number of tokens in each model, and text in some scripts, such as Thai, Arabic, Chinese, or Vietnamese with its diacritics, can take more tokens than the same meaning in English, depending on the tokenizer. GotoRamp bills text models per token, with input and output tokens at separate rates per model, so this affects cost as well as how much fits in the context window.

Measure it instead of estimating: send the same content in each language to each model and compare the token counts in each response's usage. Set max_tokens per language after you've measured. Thai also doesn't put spaces between words, so your own chunking, search and length limits need a Thai-aware word segmenter rather than a split on spaces.

Scoring with native speakers

  1. Run every prompt through every candidate model with the same system prompt and settings.
  2. Shuffle the outputs and hide the model names, so reviewers judge the text, not the brand.
  3. Have native speakers who know your market score each output against the criteria below, ideally two per language.
  4. Decide per language, and record the chosen model, a fallback and the date of the test.
CriterionReviewer's question
MeaningIs the answer correct and complete?
FluencyWould a native speaker write it this way?
RegisterIs the politeness level right for this channel?
TerminologyAre product names, local terms and units right?
Script and formatRight script, diacritics, dates, currency format and text direction?
Language choiceDid it reply in the language your policy asks for?
SafetyDid it refuse or hand over to a person when it should, and only then?

Automated checks, including one model grading another, are useful for screening large test sets. Let a native speaker make the final call.

Python
import csv, os, random
from openai import OpenAI

client = OpenAI(base_url="https://api.gotoramp.ai/v1",
                api_key=os.environ["GOTORAMP_API_KEY"])

MODELS = ["deepseek-chat", "qwen-plus"]   # use the IDs your account lists
SYSTEM = "You are the support assistant for a travel agency. Reply politely in the customer's language."

rows = []
with open("test_set.csv", encoding="utf-8") as f:   # columns: id, language, prompt
    for case in csv.DictReader(f):
        for model in MODELS:
            r = client.chat.completions.create(model=model, messages=[
                {"role": "system", "content": SYSTEM},
                {"role": "user", "content": case["prompt"]}])
            rows.append({"id": case["id"], "language": case["language"], "model": model,
                         "output": r.choices[0].message.content,
                         "input_tokens": r.usage.prompt_tokens,
                         "output_tokens": r.usage.completion_tokens})

random.shuffle(rows)   # give reviewers a copy without the "model" column
with open("review.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.DictWriter(f, fieldnames=list(rows[0]))
    w.writeheader()
    w.writerows(rows)

Which model families to test

GotoRamp's text models open with early access. These are descriptions of each family, not rankings for any language; your test set decides.

FamilyDeveloperDescribed as
DeepSeekDeepSeekChat and reasoning models, widely used for coding, analysis and agent workflows
QwenAlibaba CloudA broad family (chat, reasoning, vision-language) with wide multilingual coverage
KimiMoonshot AILong-context and agent-oriented models
GLMZhipu AI (Z.ai)General-purpose models used for coding and agents
Doubao SeedByteDance SeedGeneral chat and multimodal models

Tool calling, JSON output, vision input and context length vary by model version. Check the console or GET /v1/models for what each one supports.

Route by language

Once each language has a winner, route in code: detect the language of the incoming message, look up the model for that language, and fall back to the second choice if the first fails or is rate-limited. Language detection often confuses Malay and Indonesian, so use the customer's country or app locale as a second signal.

Keep the routing table in configuration, not code. Log language, model, token usage and review results, so you can see cost per language and re-run the comparison when a model version changes.

Start with a small balance. Scale when it works.

We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.