GotoRamp

Docs · Chat API reference

Chat completions API reference

GotoRamp's chat completions API is an OpenAI-compatible REST endpoint for the DeepSeek, Qwen, Kimi, GLM and Doubao Seed text models opening with early access, billed per token from a prepaid USD balance.

Early-access reference: the endpoint opens with your account invitation; field names and error codes are frozen at launch.

Base URL and auth

All requests go to https://api.gotoramp.ai/v1 and carry your key as a bearer token:

Authorization: Bearer $GOTORAMP_API_KEY

Keys are scoped to one account. Each key has its own usage, spend limit and optional model allowlist. Requests from countries we don't serve are refused regardless of the key.

The request and response formats are OpenAI-compatible, so the OpenAI SDK works once you change the base URL and the key. The OpenAI SDK guide walks through an existing project.

Create a chat completion

POST /v1/chat/completions returns the model's reply to a list of messages: as one JSON object by default, or as a stream of server-sent events when "stream": true.

Parameters

FieldTypeRequiredDescription
modelstringYesA model ID returned by GET /v1/models, for example deepseek-chat.
messagesarrayYesThe conversation so far, oldest first. Each message has a role (system, user, assistant or tool) and content. See Messages.
streambooleanNoSend the reply as server-sent events while it's generated. Default false.
temperaturenumberNoSampling randomness; lower values give more focused, repeatable output. The accepted range and default vary by model.
top_pnumberNoNucleus sampling: consider only the most likely tokens up to this probability mass. Adjust this or temperature, not both.
max_tokensintegerNoUpper limit on output tokens for this reply. When it's reached, finish_reason is length.
stopstring or arrayNoOne or more sequences where the model stops generating. The stop sequence isn't included in the reply.
toolsarrayNoFunctions the model may ask you to call, each with a name, description and JSON Schema for its parameters. Only on models that support tool calling.
tool_choicestring or objectNoauto, none, or a named function the model must call. Only on models that support tool calling.
response_formatobjectNoFor example {"type": "json_object"} to request valid JSON output. Only on models that support it.
userstringNoYour end-user ID. Platforms should always set it so abuse reports can be traced.

Support for tool calling, JSON output, image input and context length varies by model version. Check the console for the model you plan to use before you depend on a feature.

Messages

RoleWhat it carries
systemInstructions for the whole conversation: tone, scope, format, language.
userWhat your end user or your application asks. Models with image input also accept image parts in content.
assistantEarlier replies from the model, including any tool_calls it made. Send them back to continue a conversation.
toolThe result of a tool call your code ran, with the matching tool_call_id.

The API is stateless: send the full history you want the model to see on every request, and trim it yourself when it grows.

Examples

The same request in curl, Python and Node.js. The Python and Node.js versions use the OpenAI SDK (pip install openai or npm install openai).

curl https://api.gotoramp.ai/v1/chat/completions \
  -H "Authorization: Bearer $GOTORAMP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-chat",
    "messages": [
      { "role": "system", "content": "You answer in two sentences or fewer." },
      { "role": "user", "content": "What is a prepaid API balance?" }
    ],
    "temperature": 0.3,
    "max_tokens": 200
  }'

Response

A non-streamed request returns one chat.completion object in the OpenAI format.

Response · 200
{
  "id": "chatcmpl-7f3a9c21",
  "object": "chat.completion",
  "created": 1790000000,
  "model": "deepseek-chat",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "A prepaid API balance is money you add to your account before you make calls. Each request draws it down by the tokens it uses."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 27,
    "completion_tokens": 29,
    "total_tokens": 56
  }
}
FieldDescription
idUnique ID for this completion. Quote it when you contact support.
objectAlways chat.completion.
createdUnix timestamp, in seconds.
modelThe model ID that served the request.
choices[].indexPosition of the choice, starting at 0.
choices[].messageThe reply: role is assistant, content is the text, and tool_calls appears when the model asks for a tool.
choices[].finish_reasonstop (natural end or a stop sequence), length (hit max_tokens or the context limit) or tool_calls (the model wants you to run a tool).
usageprompt_tokens, completion_tokens and total_tokens for this request. These are the numbers you're billed on.

Streaming

With "stream": true, the response is a stream of server-sent events (text/event-stream). Each event is a line starting with data: followed by a JSON chunk whose object is chat.completion.chunk. New text arrives in choices[].delta.content, the last chunk carries the finish_reason, and the stream ends with data: [DONE].

data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"content":"A prepaid"},"finish_reason":null}]}

data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{"content":" API balance"},"finish_reason":null}]}

data: {"id":"chatcmpl-7f3a9c22","object":"chat.completion.chunk","created":1790000000,"model":"deepseek-chat","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

The OpenAI SDK parses the events for you. If you read the stream yourself, split on blank lines, skip anything that doesn't start with data:, and stop at [DONE].

List models

GET /v1/models returns the model IDs your key can call, in the OpenAI list format. Availability can differ by country and by key allowlist, so read the list from your application instead of hard-coding it.

curl https://api.gotoramp.ai/v1/models \
  -H "Authorization: Bearer $GOTORAMP_API_KEY"

Model families and what each is used for: Models. Each model's context length and supported features are shown in the console.

Billing

  • You're billed on the usage tokens of each request: prompt_tokens as input and completion_tokens as output.
  • Input and output tokens are billed at separate rates, set per model. Rates are shared with early-access accounts and shown in the console; see plans and billing.
  • Usage comes out of a prepaid balance in US dollars. Each key can have its own spend limit.
  • Failed requests are not billed, including requests refused with content_policy.

Errors

Errors return a JSON body with a machine-readable code and a human-readable message.

Error · 404
{
  "error": {
    "code": "model_not_found",
    "message": "No model with the ID deepseek-chta is open to this key. Call GET /v1/models for the list."
  }
}
HTTPcodeCause and fix
400invalid_requestA field is missing, malformed or not supported by this model. The message names the field.
401invalid_api_keyKey missing, revoked or mistyped. Check the Authorization header.
402insufficient_balanceThe balance is used up or the key reached its spend limit. Top up, or raise the key's limit.
403region_not_availableThe request came from a country we don't serve.
403content_policyThe prompt breaks our Acceptable Use Policy or the model provider's content policy. Rephrase or remove the content. Not billed.
404model_not_foundUnknown model ID, or one that isn't open to your key or country. Use an ID from GET /v1/models.
429rate_limitedToo many requests in a short window. Retry with backoff.
5xxupstream_errorTemporary model-side failure. Retry with backoff; not billed.

Rate limits

Each account has rate limits, shown in the console. Pay as you go accounts get standard limits; Volume and Enterprise accounts get higher concurrency, agreed in writing. When you exceed a limit you get 429: retry with exponential backoff and jitter, and cap concurrent requests on your side. The OpenAI SDK already retries some 429 and 5xx responses, and its max_retries option sets how many times.

Versioning and changes

The /v1 path is stable. We add fields without notice, but we announce removals, renames and model-version changes at least 30 days ahead by email to account owners.

Related: Quickstart · Video API reference · OpenAI SDK guide · DeepSeek · Qwen