API reference
Read the docs.
One endpoint, any OpenAI SDK. Change the base URL, the key, and the model name. Existing code works as-is.
01Endpoint
OpenAI-compatible chat completionshttps://inference.coret.ai/v1Authenticate with Authorization: Bearer sk-coret-.... Keys are issued in the console: sign in, name a key, copy it once. POST /chat/completions and GET /models are live. Every response carries an x-request-id header; quote it when reporting an issue.
02Quickstart
curl · Python · JavaScriptcurl https://inference.coret.ai/v1/chat/completions \
-H "Authorization: Bearer $CORET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "coret/glm-5.3-flash-uncensored",
"messages": [{"role": "user", "content": "Say hello."}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://inference.coret.ai/v1",
api_key="sk-coret-...",
)
reply = client.chat.completions.create(
model="coret/glm-5.3-flash-uncensored",
messages=[{"role": "user", "content": "Say hello."}],
)
print(reply.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://inference.coret.ai/v1",
apiKey: process.env.CORET_API_KEY,
});
const reply = await client.chat.completions.create({
model: "coret/glm-5.3-flash-uncensored",
messages: [{ role: "user", content: "Say hello." }],
});
console.log(reply.choices[0].message.content);Responses are standard OpenAI chat-completion objects. usage reports prompt, completion, and cached token counts. Cached input tokens (prompt_tokens_details.cached_tokens) bill at the lower cached rate automatically.
03Streaming
server-sent eventsstream = client.chat.completions.create(
model="coret/glm-5.3-flash-uncensored",
messages=[{"role": "user", "content": "Write a haiku."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Set stream: true for token-by-token delivery. The final chunk before [DONE] has an empty choices array and carries usage; guard for it as above.
04Tool calling
functions · parallel callstools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
reply = client.chat.completions.create(
model="coret/glm-5.3-flash-uncensored",
messages=[{"role": "user", "content": "Weather in Lisbon?"}],
tools=tools,
)
call = reply.choices[0].message.tool_calls[0]
# call.function.name == "get_weather"
# call.function.arguments == '{"city": "Lisbon"}'Pass OpenAI-shape tools; steer with tool_choice and parallel_tool_calls. Return results as role: "tool" messages to continue the exchange.
05JSON mode
structured outputreply = client.chat.completions.create(
model="coret/glm-5.3-flash-uncensored",
messages=[
{"role": "system", "content": "Reply with a JSON object."},
{"role": "user", "content": "Name three primes."},
],
response_format={"type": "json_object"},
)response_format: {"type": "json_object"} constrains output to valid JSON. Mention JSON in a message as well; models comply more reliably when the prompt asks for it.
06Models & pricing
USD per 1M tokens| Model | ID | Aliases | Context | Input | Cached | Output |
|---|---|---|---|---|---|---|
| GLM 5.3 Flash Uncensored | coret/glm-5.3-flash-uncensored | glm-5.3-flash-uncensored · glm-5.3-flash · glm5.3-flash | 256K | $0.075 | $0.015 | $0.25 |
| Qwen3.8 27B Uncensored | coret/qwen3.8-27b-uncensored | qwen3.8-27b-uncensored | 262K | $0.20 | $0.05 | $0.80 |
| DeepSeek V4 Flash | coret/deepseek-v4-flash | deepseek-v4-flash · deepseek-chat | 1M | $0.054 | $0.011 | $0.108 |
| GLM 5.2 | coret/glm-5.2 | glm-5.2 · glm5.2 | 256K | $0.80 | $0.22 | $1.80 |
Aliases resolve to the same model; use whichever your tooling prefers. GET /v1/models returns the live catalog machine-readably. Prepaid credits, no usage tiers.
07Parameters
supported request fieldsmessagestemperaturetop_pmax_tokensmax_completion_tokensstopntoolstool_choiceparallel_tool_callsresponse_formatpresence_penaltyfrequency_penaltyseeduserstreamUnlisted parameters are ignored, not rejected. Request bodies are capped at 10 MB.
08Errors & limits
OpenAI-shape error bodies{"error": {"message": "...", "type": "rate_limit_exceeded", "param": null, "code": "rate_limit_exceeded"}}| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request_error | Malformed body or missing field. |
| 401 | authentication_error | Missing or invalid API key. |
| 402 | insufficient_credits | Balance exhausted. Add credits in the console. |
| 429 | rate_limit_exceeded | Cap exceeded or capacity squeeze. Honor retry-after. |
| 500 | internal_error | Our fault. Safe to retry. |
| 503 | service_unavailable | Temporary outage. Retry with backoff. |
Accounts run up to 40 concurrent requests by default, with no requests-per-minute cap unless one is set on your account. Over the cap, or during a capacity squeeze, you get 429 with a retry-after header; back off and retry. 5xx responses are retry-safe. Need higher limits? Ask from the console.
Next: get API access or see models & pricing.