Guides · Text models
Use your OpenAI SDK code with GotoRamp.
This GotoRamp guide shows how to run existing OpenAI SDK code against GotoRamp's OpenAI-compatible text API: change the base URL and API key, choose a model ID, and check the few things that differ between models.
Last updated
Early access. API keys and https://api.gotoramp.ai/v1 open with your early-access invitation. Request early access to get yours.
What changes
GotoRamp's text API accepts the same chat-completion requests the OpenAI SDKs send, so most code needs two new settings and a new model name. Nothing else in your request code has to change to make a first call.
| Setting | Value on GotoRamp |
|---|---|
| Base URL | https://api.gotoramp.ai/v1 |
| API key | Your GotoRamp key, read from the GOTORAMP_API_KEY environment variable |
| Auth header | Authorization: Bearer plus your key (the SDK sets it) |
| Model | A GotoRamp model ID, for example deepseek-chat |
| Endpoints | POST /v1/chat/completions, GET /v1/models |
Store the key in an environment variable or your secret manager, never in source code. Give each app or environment its own key so you can set a spend limit per key.
Python, Node.js and curl
Python
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.gotoramp.ai/v1",
api_key=os.environ["GOTORAMP_API_KEY"],
)
resp = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": "You answer questions for a Malaysian online store. Reply in the customer's language."},
{"role": "user", "content": "Boleh tukar alamat penghantaran selepas buat pesanan?"},
],
)
print(resp.choices[0].message.content)
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)Node.js
// npm install openai (Node.js 18+, ES module)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.gotoramp.ai/v1",
apiKey: process.env.GOTORAMP_API_KEY,
});
const resp = await client.chat.completions.create({
model: "deepseek-chat",
messages: [
{ role: "system", content: "You answer questions for a hotel in Phuket. Reply in the guest's language." },
{ role: "user", content: "Can I check in early tomorrow?" },
],
});
console.log(resp.choices[0].message.content);
console.log(resp.usage);curl
curl https://api.gotoramp.ai/v1/chat/completions \
-H "Authorization: Bearer $GOTORAMP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-chat",
"messages": [
{"role": "user", "content": "Translate to Thai: Your order has shipped."}
]
}'Each response includes a usage object with input and output token counts. Text models are billed per token, with input and output tokens at separate rates for each model, so log these numbers from the start.
List the models you can call
GET /v1/models returns the model IDs your account can use. Copy IDs exactly as listed and keep them in configuration, not scattered through code.
curl https://api.gotoramp.ai/v1/models \
-H "Authorization: Bearer $GOTORAMP_API_KEY"for model in client.models.list():
print(model.id)Tool calling, JSON output, vision input and context length vary by model version. Check the console before you depend on one of them.
Stream responses
Set "stream": true to receive the reply in pieces as it's generated, which makes chat interfaces feel faster. The API sends server-sent events and the SDK turns them into an iterator of chunks.
stream = client.chat.completions.create(
model="deepseek-chat",
messages=[{"role": "user", "content": "Write a two-line welcome message in Tagalog."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)const stream = await client.chat.completions.create({
model: "qwen-plus",
messages: [{ role: "user", content: "Write a two-line welcome message in Vietnamese." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}With curl, add -N so output isn't buffered. If you relay a stream to a browser, stop reading when the user closes the page so you don't keep generating output nobody sees.
Handle errors and retries
Retry only the errors that can succeed on a second try. The others need a fix first.
| HTTP | Meaning | What to do |
|---|---|---|
| 400 | The request is malformed or outside what the model accepts | Fix the request; don't retry it unchanged |
| 401 | Key missing, revoked or mistyped | Check GOTORAMP_API_KEY; don't retry |
| 402 | Balance used up, or the key reached its spend limit | Top up or raise the key's limit, then retry |
| 403 | The request came from a country we don't serve, or broke the content rules | Don't retry; check where the call runs from, or change the input |
| 429 | Too many requests, or too many running at once | Retry with exponential backoff and jitter; if the response has a Retry-After header, wait at least that long |
| 5xx | Temporary server or model-side error | Retry with backoff; failed requests aren't billed |
The OpenAI SDKs already retry some errors, including 429 and 5xx, a small number of times by default. Either set that count yourself (max_retries in Python, maxRetries in Node.js) or turn it off and use your own loop, but not both, or retries multiply.
import os, random, time
from openai import OpenAI, APIStatusError, APIConnectionError
client = OpenAI(
base_url="https://api.gotoramp.ai/v1",
api_key=os.environ["GOTORAMP_API_KEY"],
max_retries=0, # this loop handles retries
timeout=60,
)
RETRY = {429, 500, 502, 503, 504}
def chat(messages, model="deepseek-chat", attempts=5):
for attempt in range(attempts):
try:
return client.chat.completions.create(model=model, messages=messages)
except APIStatusError as err:
# 400, 401, 402 and 403 won't succeed on retry: fix the cause
if err.status_code not in RETRY or attempt == attempts - 1:
raise
except APIConnectionError:
if attempt == attempts - 1:
raise
time.sleep(min(30, 2 ** attempt) + random.random())const client = new OpenAI({
baseURL: "https://api.gotoramp.ai/v1",
apiKey: process.env.GOTORAMP_API_KEY,
maxRetries: 4, // the SDK retries 429, 5xx and connection errors with backoff
timeout: 60_000,
});
try {
const resp = await client.chat.completions.create({ model: "deepseek-chat", messages });
} catch (err) {
if (err.status === 402) {
// top up, or raise this key's spend limit
} else if (err.status === 403) {
// country not served or content rules: don't retry
} else {
throw err;
}
}Migration checklist
- Base URL
https://api.gotoramp.ai/v1everywhere a client is created: services, workers, scripts and tests. - API keyRead
GOTORAMP_API_KEYfrom your secret store. One key per app or environment, each with its own spend limit. - Model IDsModel IDs differ. Replace hard-coded model names with IDs from
GET /v1/models, such asdeepseek-chatorqwen-plus. - Tool callingVaries by model. Re-test your tool definitions and the arguments each model returns.
- JSON outputVaries by model. Validate every response against your schema and handle failures.
- Vision inputVaries by model. Confirm a model accepts images before you send them.
- Context lengthVaries by model version. Check the console and trim conversation history to fit.
- Token countsEach model's tokenizer counts the same text differently. Don't reuse token estimates or
max_tokensvalues from another model; readusagefrom each response. - PromptsRe-test system prompts. A prompt tuned for one model can read differently to another.
- Where calls runRequests from countries we don't serve are refused. Check where your servers, workers and CI jobs run before you switch.
Related
Start with a small balance. Scale when it works.
We're onboarding early-access accounts in small batches. Tell us what you're building; we reply within one business day.