GotoRamp

Guides · Text models

Use your OpenAI SDK code with GotoRamp.

This GotoRamp guide shows how to run existing OpenAI SDK code against GotoRamp's OpenAI-compatible text API: change the base URL and API key, choose a model ID, and check the few things that differ between models.

Last updated

Early access. API keys and https://api.gotoramp.ai/v1 open with your early-access invitation. Request early access to get yours.

What changes

GotoRamp's text API accepts the same chat-completion requests the OpenAI SDKs send, so most code needs two new settings and a new model name. Nothing else in your request code has to change to make a first call.

SettingValue on GotoRamp
Base URLhttps://api.gotoramp.ai/v1
API keyYour GotoRamp key, read from the GOTORAMP_API_KEY environment variable
Auth headerAuthorization: Bearer plus your key (the SDK sets it)
ModelA GotoRamp model ID, for example deepseek-chat
EndpointsPOST /v1/chat/completions, GET /v1/models

Store the key in an environment variable or your secret manager, never in source code. Give each app or environment its own key so you can set a spend limit per key.

Python, Node.js and curl

Python

Python
# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gotoramp.ai/v1",
    api_key=os.environ["GOTORAMP_API_KEY"],
)

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": "You answer questions for a Malaysian online store. Reply in the customer's language."},
        {"role": "user", "content": "Boleh tukar alamat penghantaran selepas buat pesanan?"},
    ],
)

print(resp.choices[0].message.content)
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)

Node.js

Node.js
// npm install openai   (Node.js 18+, ES module)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gotoramp.ai/v1",
  apiKey: process.env.GOTORAMP_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "deepseek-chat",
  messages: [
    { role: "system", content: "You answer questions for a hotel in Phuket. Reply in the guest's language." },
    { role: "user", content: "Can I check in early tomorrow?" },
  ],
});

console.log(resp.choices[0].message.content);
console.log(resp.usage);

curl

curl
curl https://api.gotoramp.ai/v1/chat/completions \
  -H "Authorization: Bearer $GOTORAMP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-chat",
    "messages": [
      {"role": "user", "content": "Translate to Thai: Your order has shipped."}
    ]
  }'

Each response includes a usage object with input and output token counts. Text models are billed per token, with input and output tokens at separate rates for each model, so log these numbers from the start.

List the models you can call

GET /v1/models returns the model IDs your account can use. Copy IDs exactly as listed and keep them in configuration, not scattered through code.

curl
curl https://api.gotoramp.ai/v1/models \
  -H "Authorization: Bearer $GOTORAMP_API_KEY"
Python
for model in client.models.list():
    print(model.id)

Tool calling, JSON output, vision input and context length vary by model version. Check the console before you depend on one of them.

Stream responses

Set "stream": true to receive the reply in pieces as it's generated, which makes chat interfaces feel faster. The API sends server-sent events and the SDK turns them into an iterator of chunks.

Python
stream = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Write a two-line welcome message in Tagalog."}],
    stream=True,
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
Node.js
const stream = await client.chat.completions.create({
  model: "qwen-plus",
  messages: [{ role: "user", content: "Write a two-line welcome message in Vietnamese." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

With curl, add -N so output isn't buffered. If you relay a stream to a browser, stop reading when the user closes the page so you don't keep generating output nobody sees.

Handle errors and retries

Retry only the errors that can succeed on a second try. The others need a fix first.

HTTPMeaningWhat to do
400The request is malformed or outside what the model acceptsFix the request; don't retry it unchanged
401Key missing, revoked or mistypedCheck GOTORAMP_API_KEY; don't retry
402Balance used up, or the key reached its spend limitTop up or raise the key's limit, then retry
403The request came from a country we don't serve, or broke the content rulesDon't retry; check where the call runs from, or change the input
429Too many requests, or too many running at onceRetry with exponential backoff and jitter; if the response has a Retry-After header, wait at least that long
5xxTemporary server or model-side errorRetry with backoff; failed requests aren't billed

The OpenAI SDKs already retry some errors, including 429 and 5xx, a small number of times by default. Either set that count yourself (max_retries in Python, maxRetries in Node.js) or turn it off and use your own loop, but not both, or retries multiply.

Python
import os, random, time
from openai import OpenAI, APIStatusError, APIConnectionError

client = OpenAI(
    base_url="https://api.gotoramp.ai/v1",
    api_key=os.environ["GOTORAMP_API_KEY"],
    max_retries=0,   # this loop handles retries
    timeout=60,
)

RETRY = {429, 500, 502, 503, 504}

def chat(messages, model="deepseek-chat", attempts=5):
    for attempt in range(attempts):
        try:
            return client.chat.completions.create(model=model, messages=messages)
        except APIStatusError as err:
            # 400, 401, 402 and 403 won't succeed on retry: fix the cause
            if err.status_code not in RETRY or attempt == attempts - 1:
                raise
        except APIConnectionError:
            if attempt == attempts - 1:
                raise
        time.sleep(min(30, 2 ** attempt) + random.random())
Node.js
const client = new OpenAI({
  baseURL: "https://api.gotoramp.ai/v1",
  apiKey: process.env.GOTORAMP_API_KEY,
  maxRetries: 4,     // the SDK retries 429, 5xx and connection errors with backoff
  timeout: 60_000,
});

try {
  const resp = await client.chat.completions.create({ model: "deepseek-chat", messages });
} catch (err) {
  if (err.status === 402) {
    // top up, or raise this key's spend limit
  } else if (err.status === 403) {
    // country not served or content rules: don't retry
  } else {
    throw err;
  }
}

Migration checklist

  • Base URLhttps://api.gotoramp.ai/v1 everywhere a client is created: services, workers, scripts and tests.
  • API keyRead GOTORAMP_API_KEY from your secret store. One key per app or environment, each with its own spend limit.
  • Model IDsModel IDs differ. Replace hard-coded model names with IDs from GET /v1/models, such as deepseek-chat or qwen-plus.
  • Tool callingVaries by model. Re-test your tool definitions and the arguments each model returns.
  • JSON outputVaries by model. Validate every response against your schema and handle failures.
  • Vision inputVaries by model. Confirm a model accepts images before you send them.
  • Context lengthVaries by model version. Check the console and trim conversation history to fit.
  • Token countsEach model's tokenizer counts the same text differently. Don't reuse token estimates or max_tokens values from another model; read usage from each response.
  • PromptsRe-test system prompts. A prompt tuned for one model can read differently to another.
  • Where calls runRequests from countries we don't serve are refused. Check where your servers, workers and CI jobs run before you switch.

Start with a small balance. Scale when it works.

We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.