GotoRamp

Models · Text

GLM API for businesses in Southeast Asia

GLM, the general-purpose model family from Zhipu AI (known internationally as Z.ai), is opening with early access on GotoRamp, an OpenAI-compatible API for businesses in Singapore, Malaysia, the Philippines and Thailand.

At a glance

GLM on GotoRamp.

A general-purpose family that teams use for coding and agents, called through the same endpoint as every other text model on GotoRamp.

ItemGLM through GotoRamp
MakerZhipu AI, known internationally as Z.ai
TypeText: general-purpose models used for coding and agents
How you call itOpenAI-compatible POST /v1/chat/completions
Model IDsReturned by GET /v1/models once your account opens
BillingPer input and output token, at separate rates, from a prepaid USD balance
StatusEarly access Opening with early access
CountriesSingapore, Malaysia, the Philippines and Thailand first

GotoRamp LLC isn't affiliated with or endorsed by Zhipu AI. We offer a model only once we hold written terms for it through a channel that allows our use in each market.

Use cases

What teams use it for.

Coding and agent work share one pattern: the model proposes, your code acts, a person approves anything that matters.

Coding agents

Draft patches, explain failing tests and write migration scripts from your editor, chat tool or CI pipeline. A developer reviews and merges every change.

Agents over your own APIs

Check stock, draft a quote, update a delivery record. The model picks a function and fills in its arguments; your code validates them and runs it.

Everyday assistant features

Summaries, rewrites, classification and question answering inside your product, on one general-purpose family instead of a separate model for each feature.

Call it with the OpenAI SDK

A tool call, end to end.

Set base_url to https://api.gotoramp.ai/v1, read your key from GOTORAMP_API_KEY, pass your functions in tools and read tool_calls from the reply.

Use a GLM model ID from GET /v1/models whose console entry shows tool calling. Field-by-field details: chat completions parameters.

Python
# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gotoramp.ai/v1",
    api_key=os.environ["GOTORAMP_API_KEY"],
)

tools = [{
    "type": "function",
    "function": {
        "name": "get_stock",
        "description": "Units in stock for one SKU",
        "parameters": {
            "type": "object",
            "properties": {"sku": {"type": "string"}},
            "required": ["sku"],
        },
    },
}]

resp = client.chat.completions.create(
    model="<model-id from GET /v1/models>",  # a GLM version with tool calling
    messages=[{"role": "user", "content": "Do we have SKU TH-2291 in stock?"}],
    tools=tools,
    tool_choice="auto",
)

msg = resp.choices[0].message
if msg.tool_calls:
    call = msg.tool_calls[0]
    print(call.function.name, call.function.arguments)  # run it, then send a "tool" message
else:
    print(msg.content)

How a tool call runs

Four steps, and your code runs the tool.

The model never touches your systems. It returns a request, and you decide whether to act on it.

  1. 1 · Send

    Messages and tools

    The conversation plus tools: the functions your code will run, each with a JSON Schema.

  2. 2 · Choose

    The model asks

    The reply carries tool_calls, a function name and JSON arguments, instead of text.

  3. 3 · Run

    Your code acts

    Validate the arguments, run the function, and add a tool message with the result and its tool_call_id.

  4. 4 · Answer

    The model replies

    Call again. The model answers from the result or asks for another tool.

Things to know

Before you build on GLM.

What to confirm before an agent goes near production data.

  • Features vary by versionTool calling, JSON output and context length differ between GLM versions. Check GET /v1/models and the console before you depend on one.
  • Validate argumentsThe model proposes arguments; it doesn't check them. Validate every call before your code writes data, sends messages or spends money.
  • Content rulesZhipu AI's own content policies apply alongside our Acceptable Use Policy, so some prompts may be refused. Refused requests aren't billed.
  • Where data is processedThe processing location for each model is listed in our Privacy Notice before accounts open.
  • TrainingWe don't use your prompts, code or outputs to train models.

FAQ

GLM questions.

For anything else, write to us.

Is GLM the same thing as Z.ai?
GLM is the model family, and Z.ai is the international name of Zhipu AI, the company that makes it. On GotoRamp you call GLM with a GotoRamp key, not a Z.ai account.
Can I use GLM in a coding agent?
Yes. GLM models are used for coding and agents, and developer tools that let you set a custom OpenAI-compatible base URL and key can usually call them through GotoRamp. Check that the version you pick supports tool calling.
What happens if a GLM request is refused?
You get a 403 with the code content_policy, and the request isn't billed. Rephrase the prompt or check it against our Acceptable Use Policy.
Can I mix GLM with other models in one app?
Yes. Every text model on GotoRamp uses the same endpoint, key format and balance, so routing a task to another model is a change to the model field.
When will GLM open in my country?
Early access opens first in Singapore, Malaysia, the Philippines and Thailand. Indonesia, Vietnam, the United Arab Emirates and Saudi Arabia are planned after local registration; see availability for the full list.

Other models

Route by task, not by vendor.

Switching families is one field in the request. Compare every model.

  • DeepSeekDeepSeek · chat and reasoning models
  • QwenAlibaba Cloud · chat, reasoning and vision-language models
  • KimiMoonshot AI · long-context and agent-oriented models
  • Doubao SeedByteDance Seed · general chat and multimodal models
  • SeedanceByteDance · video generation, billed per second

Start with a small balance. Scale when it works.

We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.