A familiar API.
A fresh start.
Connect your app with an OpenAI-compatible SDK. Security models require early access.
Join the waitlist ↗Models and pricing
Prices are per million tokens and charged against your prepaid balance at the price shown when the request runs. Security models marked Soon require early access and cannot be called with a regular API key.
Join the early access waitlist →| Model | Slug | Input | Output | Context | Availability |
|---|---|---|---|---|---|
| GLM-5.3 Raw | glm-5.3-raw | $2.50 | $7.00 | 1,048,576 | Soon → |
Authentication
Create a key under settings → API keys after signing in. Keys start with skrm-, are shown once, and use the same prepaid balance as the web chat. Send them as a bearer token.
Authorization: Bearer skrm-...
Chat completions
POST /v1/chat/completions — request and response bodies follow the OpenAI schema, including stream: true chunks terminated by data: [DONE].
Models run on GPUs that are started on demand, so the first request after a quiet spell waits for a worker to boot — usually one to five minutes. Everything after that answers in seconds. Thanks for your patience. Please allow for it in your client: use stream: true and a timeout of at least 300 seconds. The non-streaming endpoint gives up at 280 seconds with model_cold_start_timeout, and a retry normally lands on a warm worker.
The examples use fast. Your account must have early access before calling it.
curl https://www.rawmodels.ai/v1/chat/completions \
-H "Authorization: Bearer skrm-..." \
-H "Content-Type: application/json" \
-d '{
"model": "fast",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'from openai import OpenAI
client = OpenAI(api_key="skrm-...", base_url="https://www.rawmodels.ai/v1")
response = client.chat.completions.create(
model="fast",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)Claude Code and Codex
Use Raw Models from Claude Code and Codex like any other model. Besides /v1/chat/completions, the API accepts the Anthropic Messages API at /v1/messages and the OpenAI Responses API at /v1/responses, so any client built for those works with a Raw Models key.
curl -fsSL https://www.rawmodels.ai/install | shThen just run claude or codex — every model tier maps to glm-5.3-raw, and /model lists the Raw Models models that run tool calls. Your previous settings are kept and come back with curl -fsSL https://www.rawmodels.ai/install | sh -s -- uninstall. Only Claude Code: sh -s -- claude; only Codex: sh -s -- codex.
Early access is required. Usage is billed to your prepaid balance. macOS and Linux. The first request after an idle period waits for the model to start.
Full Claude Code setup with per-tier model mapping: Claude Code on Raw Models. Typed decisions over text and images: Decisions API.
Early access is required. These agents keep several providers at once: glm-5.3-raw appears next to the models you already use, and those keep working on your own accounts. Set export RAWMODELS_API_KEY=skrm-... in your shell.
Merge into ~/.config/opencode/opencode.json, then pick rawmodels/glm-5.3-raw with /models.
{
"provider": {
"rawmodels": {
"npm": "@ai-sdk/openai-compatible",
"name": "Raw Models",
"options": {
"baseURL": "https://www.rawmodels.ai/v1",
"apiKey": "{env:RAWMODELS_API_KEY}"
},
"models": {
"glm-5.3-raw": {
"name": "GLM-5.3 Raw",
"tool_call": true,
"limit": {
"context": 1048576,
"output": 8192
}
}
}
}
}
}Parameters
modelModel slug from the table above.
messagesArray of {role, content}. Roles: system, user, assistant, tool. Content may be a string or an array of text parts.
streamServer-sent events when true. Defaults to false.
max_tokensOptional. Left out, the model writes until it stops or the context is full. Reasoning tokens count against it.
tools / tool_choiceOpenAI function calling on models that support it. Calls come back as message.tool_calls (or tool_calls deltas when streaming) with finish_reason "tool_calls"; send each result as a {role: "tool", tool_call_id, content} message.
parallel_tool_calls / stopForwarded to the model as in the OpenAI API.
temperature0 – 2. Left out, the upstream provider's default applies.
top_p / top_k / min_pNucleus, top-k and min-p sampling. Only sent upstream when you set them.
repetition_penalty0.5 – 2. Only sent upstream when you set it.
frequency_penalty / presence_penalty-2 – 2. Only sent upstream when you set them.
seed / response_format / reasoningForwarded to the provider unchanged. Reasoning models return their thinking as message.reasoning (or reasoning deltas when streaming).
Models endpoint
GET /v1/models lists the slugs your key can call.
curl https://www.rawmodels.ai/v1/models -H "Authorization: Bearer skrm-..."
Errors
401invalid_api_keyMissing, malformed, or revoked key.
403early_access_requiredThis model is coming soon. Join the early access waitlist at /waitlist.
402insufficient_quotaPrepaid balance cannot cover the request. Top up on the billing page.
429rate_limit_exceeded20 requests/minute, 60 after the first top-up.
400max_tokens_exceededmax_tokens is above the selected model's output limit.
400tools_not_supportedThe request carries tools or tool messages and the selected model does not run tool calling.
504model_cold_start_timeoutThe worker did not produce a first token within 280s. Retry, or use stream: true.
400invalid_request_errorUnknown model, oversized request, or parameters outside the allowed range.
Questions: hello@rawmodels.ai