Docs · Chat API reference
Chat completions API reference
GotoRamp's chat completions API is an OpenAI-compatible REST endpoint for the DeepSeek, Qwen, Kimi, GLM and Doubao Seed text models opening with early access, billed per token from a prepaid USD balance.
Early-access reference: the endpoint opens with your account invitation; field names and error codes are frozen at launch.
Base URL and auth
All requests go to https://api.gotoramp.ai/v1 and carry your key as a bearer token:
Authorization: Bearer $GOTORAMP_API_KEY
Keys are scoped to one account. Each key has its own usage, spend limit and optional model allowlist. Requests from countries we don't serve are refused regardless of the key.
The request and response formats are OpenAI-compatible, so the OpenAI SDK works once you change the base URL and the key. The OpenAI SDK guide walks through an existing project.
Create a chat completion
POST /v1/chat/completions returns the model's reply to a list of messages: as one JSON object by default, or as a stream of server-sent events when "stream": true.
Parameters
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A model ID returned by GET /v1/models, for example deepseek-chat. |
messages | array | Yes | The conversation so far, oldest first. Each message has a role (system, user, assistant or tool) and content. See Messages. |
stream | boolean | No | Send the reply as server-sent events while it's generated. Default false. |
temperature | number | No | Sampling randomness; lower values give more focused, repeatable output. The accepted range and default vary by model. |
top_p | number | No | Nucleus sampling: consider only the most likely tokens up to this probability mass. Adjust this or temperature, not both. |
max_tokens | integer | No | Upper limit on output tokens for this reply. When it's reached, finish_reason is length. |
stop | string or array | No | One or more sequences where the model stops generating. The stop sequence isn't included in the reply. |
tools | array | No | Functions the model may ask you to call, each with a name, description and JSON Schema for its parameters. Only on models that support tool calling. |
tool_choice | string or object | No | auto, none, or a named function the model must call. Only on models that support tool calling. |
response_format | object | No | For example {"type": "json_object"} to request valid JSON output. Only on models that support it. |
user | string | No | Your end-user ID. Platforms should always set it so abuse reports can be traced. |
Support for tool calling, JSON output, image input and context length varies by model version. Check the console for the model you plan to use before you depend on a feature.
Messages
| Role | What it carries |
|---|---|
system | Instructions for the whole conversation: tone, scope, format, language. |
user | What your end user or your application asks. Models with image input also accept image parts in content. |
assistant | Earlier replies from the model, including any tool_calls it made. Send them back to continue a conversation. |
tool | The result of a tool call your code ran, with the matching tool_call_id. |
The API is stateless: send the full history you want the model to see on every request, and trim it yourself when it grows.
Examples
The same request in curl, Python and Node.js. The Python and Node.js versions use the OpenAI SDK (pip install openai or npm install openai).
curl https://api.gotoramp.ai/v1/chat/completions \
-H "Authorization: Bearer $GOTORAMP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-chat",
"messages": [
{ "role": "system", "content": "You answer in two sentences or fewer." },
{ "role": "user", "content": "What is a prepaid API balance?" }
],
"temperature": 0.3,
"max_tokens": 200
}'# pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.gotoramp.ai/v1",
api_key=os.environ["GOTORAMP_API_KEY"],
)
completion = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": "You answer in two sentences or fewer."},
{"role": "user", "content": "What is a prepaid API balance?"},
],
temperature=0.3,
max_tokens=200,
)
print(completion.choices[0].message.content)
print(completion.usage) # prompt_tokens, completion_tokens, total_tokens// npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.gotoramp.ai/v1",
apiKey: process.env.GOTORAMP_API_KEY,
});
const completion = await client.chat.completions.create({
model: "deepseek-chat",
messages: [
{ role: "system", content: "You answer in two sentences or fewer." },
{ role: "user", content: "What is a prepaid API balance?" },
],
temperature: 0.3,
max_tokens: 200,
});
console.log(completion.choices[0].message.content);
console.log(completion.usage); // prompt_tokens, completion_tokens, total_tokensResponse
A non-streamed request returns one chat.completion object in the OpenAI format.
{
"id": "chatcmpl-7f3a9c21",
"object": "chat.completion",
"created": 1790000000,
"model": "deepseek-chat",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A prepaid API balance is money you add to your account before you make calls. Each request draws it down by the tokens it uses."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 27,
"completion_tokens": 29,
"total_tokens": 56
}
}| Field | Description |
|---|---|
id | Unique ID for this completion. Quote it when you contact support. |
object | Always chat.completion. |
created | Unix timestamp, in seconds. |
model | The model ID that served the request. |
choices[].index | Position of the choice, starting at 0. |
choices[].message | The reply: role is assistant, content is the text, and tool_calls appears when the model asks for a tool. |
choices[].finish_reason | stop (natural end or a stop sequence), length (hit max_tokens or the context limit) or tool_calls (the model wants you to run a tool). |
usage | prompt_tokens, completion_tokens and total_tokens for this request. These are the numbers you're billed on. |
Streaming
With "stream": true, the response is a stream of server-sent events (text/event-stream). Each event is a line starting with data: followed by a JSON chunk whose object is chat.completion.chunk. New text arrives in choices[].delta.content, the last chunk carries the finish_reason, and the stream ends with data: [DONE].
data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"content":"A prepaid"},"finish_reason":null}]}
data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"content":" API balance"},"finish_reason":null}]}
data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]stream = client.chat.completions.create(
model="deepseek-chat",
messages=[{"role": "user", "content": "What is a prepaid API balance?"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)const stream = await client.chat.completions.create({
model: "deepseek-chat",
messages: [{ role: "user", content: "What is a prepaid API balance?" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}The OpenAI SDK parses the events for you. If you read the stream yourself, split on blank lines, skip anything that doesn't start with data:, and stop at [DONE].
List models
GET /v1/models returns the model IDs your key can call, in the OpenAI list format. Availability can differ by country and by key allowlist, so read the list from your application instead of hard-coding it.
curl https://api.gotoramp.ai/v1/models \
-H "Authorization: Bearer $GOTORAMP_API_KEY"{
"object": "list",
"data": [
{ "id": "deepseek-chat", "object": "model", "created": 1790000000 },
{ "id": "qwen-plus", "object": "model", "created": 1790000000 }
]
}Model families and what each is used for: Models. Each model's context length and supported features are shown in the console.
Billing
- You're billed on the
usagetokens of each request:prompt_tokensas input andcompletion_tokensas output. - Input and output tokens are billed at separate rates, set per model. Rates are shared with early-access accounts and shown in the console; see plans and billing.
- Usage comes out of a prepaid balance in US dollars. Each key can have its own spend limit.
- Failed requests are not billed, including requests refused with
content_policy.
Errors
Errors return a JSON body with a machine-readable code and a human-readable message.
{
"error": {
"code": "model_not_found",
"message": "No model with the ID deepseek-chta is open to this key. Call GET /v1/models for the list."
}
}| HTTP | code | Cause and fix |
|---|---|---|
| 400 | invalid_request | A field is missing, malformed or not supported by this model. The message names the field. |
| 401 | invalid_api_key | Key missing, revoked or mistyped. Check the Authorization header. |
| 402 | insufficient_balance | The balance is used up or the key reached its spend limit. Top up, or raise the key's limit. |
| 403 | region_not_available | The request came from a country we don't serve. |
| 403 | content_policy | The prompt breaks our Acceptable Use Policy or the model provider's content policy. Rephrase or remove the content. Not billed. |
| 404 | model_not_found | Unknown model ID, or one that isn't open to your key or country. Use an ID from GET /v1/models. |
| 429 | rate_limited | Too many requests in a short window. Retry with backoff. |
| 5xx | upstream_error | Temporary model-side failure. Retry with backoff; not billed. |
Rate limits
Each account has rate limits, shown in the console. Pay as you go accounts get standard limits; Volume and Enterprise accounts get higher concurrency, agreed in writing. When you exceed a limit you get 429: retry with exponential backoff and jitter, and cap concurrent requests on your side. The OpenAI SDK already retries some 429 and 5xx responses, and its max_retries option sets how many times.
Versioning and changes
The /v1 path is stable. We add fields without notice, but we announce removals, renames and model-version changes at least 30 days ahead by email to account owners.
Related: Quickstart · Video API reference · OpenAI SDK guide · DeepSeek · Qwen