GotoRamp

Models · Text

Kimi API for businesses in Southeast Asia

Kimi, Moonshot AI's family of long-context and agent-oriented models, is opening with early access on GotoRamp, an OpenAI-compatible API for businesses in Singapore, Malaysia, the Philippines and Thailand.

At a glance

For long inputs and multi-step work.

Kimi is worth testing when a task needs a lot of text in a single request, or an agent that keeps working across many turns.

GotoRamp LLC isn't affiliated with Moonshot AI. We offer a model only once we hold written terms for it through a channel that allows our use in each market.

  • MakerMoonshot AI
  • TypeText: long-context and agent-oriented models
  • How you call itOpenAI-compatible POST /v1/chat/completions through GotoRamp
  • Model IDsReturned by GET /v1/models once your account opens
  • BillingPer input and output token, at separate rates, from a prepaid USD balance
  • StatusOpening with early access
  • CountriesSingapore, Malaysia, the Philippines and Thailand first

Use cases

What teams use it for.

Three patterns that lean on long context or agent loops. Try each on real documents from your business before you commit.

Whole documents in one request

Tender packs, supplier contracts, policy manuals or a month of meeting transcripts, sent without splitting into chunks when the version's context length allows. Ask for answers that quote the section they rely on.

Agents that run for many turns

Research or operations agents that plan, call your tools, read the results and carry on. Your code runs every tool, keeps the history and decides when the agent stops.

Questions over a set of files

Load a staff handbook, a code module or a product catalog as context and let your team ask questions against it, with a reviewer spot-checking answers before anyone acts on them.

Call it with the OpenAI SDK

List the IDs, then send the document.

Set base_url to https://api.gotoramp.ai/v1 and read your key from GOTORAMP_API_KEY. Kimi model IDs aren't printed here: ask GET /v1/models which ones your key can call, and paste one in.

Streaming, stop sequences and error codes are covered in the chat completions reference. For long outputs, streaming shows text as it arrives.

Python
# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gotoramp.ai/v1",
    api_key=os.environ["GOTORAMP_API_KEY"],
)

# 1. See which Kimi model IDs your key can call
for m in client.models.list():
    print(m.id)

# 2. Use one of them here
MODEL = "<model-id from GET /v1/models>"

with open("tender-pack.txt", encoding="utf-8") as f:
    document = f.read()

answer = client.chat.completions.create(
    model=MODEL,
    messages=[
        {"role": "system", "content": "Answer only from the document. Quote the section you used."},
        {"role": "user", "content": document + "\n\nList every submission deadline and its section."},
    ],
)

print(answer.choices[0].message.content)
print(answer.usage.prompt_tokens, "input tokens")

Things to know

Long context, kept under control.

"Long context" describes the family, not a number. Check the version you call, and watch input tokens.

  • Features vary by versionContext length, tool calling and JSON output differ between Kimi versions. Check GET /v1/models and the console before you depend on one.
  • Input adds upEvery input token is billed, so a document you resend on each turn is billed on each turn. Send only the sections a question needs when you can.
  • Spend limitsGive each API key its own spend limit. A key that reaches it gets 402 insufficient_balance instead of running on.
  • Content rulesMoonshot AI's own content policies apply alongside our Acceptable Use Policy, so some prompts may be refused. Refused requests aren't billed.
  • Where data is processedThe processing location for each model is listed in our Privacy Notice before accounts open.
  • TrainingWe don't use your prompts, documents or outputs to train models.

FAQ

Kimi questions.

More on accounts and billing in the general FAQ.

How long a document can Kimi read through GotoRamp?
It depends on the Kimi version you call, and the console lists the context length for each one. Measure your input in tokens rather than pages, because limits and billing both count tokens.
Why isn't a Kimi model ID shown on this page?
Kimi model IDs are confirmed when accounts open. GET /v1/models returns the ones your key can call, so your code uses exactly what is available to you.
How do I keep costs predictable with long documents?
Set a spend limit on each API key and watch usage per key in the console. Trim the context to what each question needs, since every input token is billed.
Can I build an agent on Kimi through GotoRamp?
Yes, on versions that support tool calling. Pass tools and tool_choice in the OpenAI-compatible format and run the loop in your own code.
Who makes Kimi?
Moonshot AI makes Kimi. GotoRamp LLC is a separate US company that offers Kimi models once it holds written terms through a channel that allows its use in each market.

Other models

Also on GotoRamp.

Same account, same balance, same request format. Compare all models.

  • DeepSeekDeepSeek · chat and reasoning models
  • QwenAlibaba Cloud · chat, reasoning and vision-language models
  • GLMZhipu AI (Z.ai) · general-purpose models used for coding and agents
  • Doubao SeedByteDance Seed · general chat and multimodal models
  • SeedanceByteDance · video generation, billed per second

Start with a small balance. Scale when it works.

We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.