API reference

The Inco API is compatible with the OpenAI and Anthropic SDKs — point an existing client at our base URL and change nothing else. Grab a key on the API Keys page.

Base URL

All endpoints are served from a single host. For the OpenAI SDK, set base_url to the /v1 path; the Anthropic SDK uses the host root.

OpenAI SDK · base_url
https://api.inco.ai/v1
Anthropic SDK · base_url
https://api.inco.ai

Authentication

Authenticate with a bearer token — your sk-inco-… key — in the Authorization header. Keys are hashed at rest and shown only once at creation.

Authorization header
Authorization: Bearer sk-inco-...

Models & pricing

The catalog is machine-readable. GET /v1/models lists every model with its context length and prices; without a key it is the public catalog, with your key any private model your workspace is entitled to appears as well. Prices are USD per 1M tokens: input, cached_input (prompt tokens served from the cache) and output. An entry also carries modalities, license, model_card_url and parameters when the catalog records them. GET /v1/models/{id} returns one entry.

List models
curl https://api.inco.ai/v1/models
One entry · kimi-k3
{
  "id": "kimi-k3",
  "object": "model",
  "created": 1790294854,
  "owned_by": "inco",
  "name": "Kimi K3",
  "description": "",
  "tier": "standard",
  "context_length": 1048576,
  "pricing": {
    "input": 3,
    "cached_input": 0.3,
    "output": 15
  },
  "capabilities": {
    "reasoning_effort": true
  },
  "license": "Kimi K3 License",
  "model_card_url": "https://huggingface.co/moonshotai/Kimi-K3",
  "parameters": 2800000000000
}

Quickstart

One key and one model list, two request styles — pick the tab for your SDK. Use the OpenAI Chat Completions API (POST /v1/chat/completions) or the Anthropic Messages API (POST /v1/messages).

POST /v1/chat/completions — the OpenAI Chat Completions API.

curl
curl https://api.inco.ai/v1/chat/completions \
  -H "Authorization: Bearer $INCO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [{ "role": "user", "content": "Explain RoPE in one sentence." }]
  }'
python · openai
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://api.inco.ai/v1",
    api_key=os.environ["INCO_API_KEY"],
)

chat = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Explain RoPE in one sentence."}],
)
print(chat.choices[0].message.content)
node · openai
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.inco.ai/v1",
  apiKey: process.env.INCO_API_KEY,
});

const chat = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "Explain RoPE in one sentence." }],
});
console.log(chat.choices[0].message.content);

Streaming

Both endpoints support streaming via server-sent events — pass stream: true. Usage is metered on the full response either way.

curl · stream
# Server-sent events: add "stream": true
curl https://api.inco.ai/v1/chat/completions \
  -H "Authorization: Bearer $INCO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "kimi-k3", "stream": true,
        "messages": [{ "role": "user", "content": "Hi" }] }'

Errors

StatusMeaning
401Missing or invalid API key (or the key was revoked).
402Prepaid balance is empty — add credit to resume.
404Unknown or disabled model id.
429Rate limit — x-inco-ratelimit-scope says which: upstream means back off and retry; quota means wait for the window the rate-limit headers report.
5xxUpstream/model backend error — safe to retry with backoff.

Rate limits & billing

Access is prepaid: each request is metered to the token and debited from your balance at the rates above (a cheaper rate applies to cached prompt tokens). When your balance reaches zero, requests return 402 until you add credit. Track spend on the Usage page, and every request, served or refused, on the Logs page.

Rate limits apply to your workspace, per model, on three dimensions: requests, input tokens (prompt-cache reads are free) and output tokensper minute. Every key you create draws from the same pools for a given model, and a second model is a second set of pools. Your workspace's current numbers, per model, are on the Limits page; contact us to raise them.

You can also cap what your workspace spends in a calendar month (UTC) on the same page. The cap is checked before every request, counting requests already in flight, so it cannot be exceeded; once reached, requests get a 429 that says so until the first of the next month, or until you raise it. Alerts at a share of the cap show on your dashboard.

Every response reports the limits in effect for the model it called, in the headers your SDK already understands. OpenAI-style endpoints return x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests (a duration such as 60s), the same three for tokens (input) and for output-tokens; /v1/messages returns anthropic-ratelimit-requests-limit, -remaining, -reset (RFC 3339), the same for input-tokens and output-tokens, and tokens for whichever of the two is tighter. Input tokens are charged against the window as an estimate at admission (prompt plus max_tokens) and corrected to the real uncached count afterwards, so a request larger than the remaining input budget is refused up front; output tokens are counted as generated, so a response is never cut short by a limit.

A 429 carries an x-inco-ratelimit-scope header naming which of the two causes applies. quota means you exceeded one of your own limits above — the message says which one and when the window resets, and the headers above carry the numbers; retrying sooner will not help, so wait, slow down, or send smaller requests. upstream means the model backend is shedding load; retrying after the Retry-After delay is the right response.

Streaming is supported on both dialects — set stream: true (OpenAI) or use the Anthropic SDK's messages.stream helper; responses arrive as server-sent events with the same metering.

Errors follow each dialect's envelope. OpenAI-style endpoints return { error: { message, type, param, code } }; /v1/messages returns Anthropic's { type: "error", error: { type, message }, request_id }. The message is one plain sentence that names the field at fault (messages[0].role: …) and never the backend; codenames the cause, the way both vendors' SDK guidance expects: 400 invalid request (param names the field), 401 invalid_api_key or key_expired, 402 credit_balance_exhausted, 403 model_not_entitled, 404 model_not_found, 429 rate_limit_exceeded (per-minute) or spend_limit_exceeded (your monthly cap), and 5xx a failure on our side: retry with backoff. Every response carries an x-request-id header (the Anthropic dialect also answers request-id, which its SDKs read), the id the request is listed under on the Logs page, whether it was served or refused; a chat completion or message body carries the same id in its id field. Include it in support requests.

FAQ