API reference

Read the docs.

One endpoint, any OpenAI SDK. Change the base URL, the key, and the model name. Existing code works as-is.

01Endpoint

OpenAI-compatible chat completions
https://inference.coret.ai/v1

Authenticate with Authorization: Bearer sk-coret-.... Keys are issued in the console: sign in, name a key, copy it once. POST /chat/completions and GET /models are live. Every response carries an x-request-id header; quote it when reporting an issue.

02Quickstart

curl · Python · JavaScript
curl https://inference.coret.ai/v1/chat/completions \
  -H "Authorization: Bearer $CORET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "coret/glm-5.3-flash-uncensored",
    "messages": [{"role": "user", "content": "Say hello."}]
  }'

Responses are standard OpenAI chat-completion objects. usage reports prompt, completion, and cached token counts. Cached input tokens (prompt_tokens_details.cached_tokens) bill at the lower cached rate automatically.

03Streaming

server-sent events
Python
stream = client.chat.completions.create(
    model="coret/glm-5.3-flash-uncensored",
    messages=[{"role": "user", "content": "Write a haiku."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Set stream: true for token-by-token delivery. The final chunk before [DONE] has an empty choices array and carries usage; guard for it as above.

04Tool calling

functions · parallel calls
Python
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

reply = client.chat.completions.create(
    model="coret/glm-5.3-flash-uncensored",
    messages=[{"role": "user", "content": "Weather in Lisbon?"}],
    tools=tools,
)
call = reply.choices[0].message.tool_calls[0]
# call.function.name == "get_weather"
# call.function.arguments == '{"city": "Lisbon"}'

Pass OpenAI-shape tools; steer with tool_choice and parallel_tool_calls. Return results as role: "tool" messages to continue the exchange.

05JSON mode

structured output
Python
reply = client.chat.completions.create(
    model="coret/glm-5.3-flash-uncensored",
    messages=[
        {"role": "system", "content": "Reply with a JSON object."},
        {"role": "user", "content": "Name three primes."},
    ],
    response_format={"type": "json_object"},
)

response_format: {"type": "json_object"} constrains output to valid JSON. Mention JSON in a message as well; models comply more reliably when the prompt asks for it.

06Models & pricing

USD per 1M tokens
ModelIDAliasesContextInputCachedOutput
GLM 5.3 Flash Uncensoredcoret/glm-5.3-flash-uncensoredglm-5.3-flash-uncensored · glm-5.3-flash · glm5.3-flash256K$0.075$0.015$0.25
Qwen3.8 27B Uncensoredcoret/qwen3.8-27b-uncensoredqwen3.8-27b-uncensored262K$0.20$0.05$0.80
DeepSeek V4 Flashcoret/deepseek-v4-flashdeepseek-v4-flash · deepseek-chat1M$0.054$0.011$0.108
GLM 5.2coret/glm-5.2glm-5.2 · glm5.2256K$0.80$0.22$1.80

Aliases resolve to the same model; use whichever your tooling prefers. GET /v1/models returns the live catalog machine-readably. Prepaid credits, no usage tiers.

07Parameters

supported request fields
messagestemperaturetop_pmax_tokensmax_completion_tokensstopntoolstool_choiceparallel_tool_callsresponse_formatpresence_penaltyfrequency_penaltyseeduserstream

Unlisted parameters are ignored, not rejected. Request bodies are capped at 10 MB.

08Errors & limits

OpenAI-shape error bodies
JSON
{"error": {"message": "...", "type": "rate_limit_exceeded", "param": null, "code": "rate_limit_exceeded"}}
StatusCodeMeaning
400invalid_request_errorMalformed body or missing field.
401authentication_errorMissing or invalid API key.
402insufficient_creditsBalance exhausted. Add credits in the console.
429rate_limit_exceededCap exceeded or capacity squeeze. Honor retry-after.
500internal_errorOur fault. Safe to retry.
503service_unavailableTemporary outage. Retry with backoff.

Accounts run up to 40 concurrent requests by default, with no requests-per-minute cap unless one is set on your account. Over the cap, or during a capacity squeeze, you get 429 with a retry-after header; back off and retry. 5xx responses are retry-safe. Need higher limits? Ask from the console.

Next: get API access or see models & pricing.