GLM 5.3 Flash Uncensored
coret/glm-5.3-flash-uncensoredFast 320B-MoE flash build, no refusal training. Reasoning and tool calls at flash pricing.
- Input
- $0.075
- Cached input
- $0.015
- Output
- $0.25
OpenAI-compatible inference
Run leading open models through the OpenAI SDK with long context, predictable pricing, and no provider lock-in.
Four models. One endpoint. No usage tiers.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.openai.com/v1""https://inference.coret.ai/v1",
)
response = client.chat.completions.create(
model="gpt-4o""coret/deepseek-v4-flash",
messages=[{"role": "user", "content": "Hello"}],
)
Choose for speed, depth, or scale. Switch with a single model name.
coret/glm-5.3-flash-uncensoredFast 320B-MoE flash build, no refusal training. Reasoning and tool calls at flash pricing.
coret/qwen3.8-27b-uncensoredNo refusal training. Built for security research, fiction, and blunt subject matter.
coret/deepseek-v4-flashMillion-token context reads whole repos, logs, and archives in one pass.
coret/glm-5.2Multi-step reasoning and dependable tool calls for work that gets reviewed.
All prices USD per 1M tokens.
Keep your OpenAI SDK. Change the base URL and model name.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://inference.coret.ai/v1",
)
response = client.chat.completions.create(
model="coret/qwen3.8-27b-uncensored",
messages=["role": "user", "content": "Hello"}],
)Everything else stays the same.
Read the API docsGet a key through AntSeed, point your OpenAI client to Coret, and start building.
Provider “Coret” · agent #60566 · antseed.com