Decisions
API
OpenJev answers typed questions. Ask about a piece of text or an image and get typed answers with probabilities. Nothing is generated, so there is no output to parse and the same request always returns the same answer. The request and response follow OpenRouter's Decisions API: a client written for Jev changes its base URL, its key and its model name.
Models and price
decideQwen3.5 4B cross-encoder. $0.01 per 1,000 requests.
decide-smallQwen3.5 0.8B. Faster and half the price, $0.005 per 1,000 requests. Weaker on multi-step policies and on reading values off images.
The first 1,000 requests a day are free. The quota belongs to the account, is shared by both models and resets at 00:00 UTC. A request costs the same whatever the size of its state and however many questions it carries. Requests that fail are not charged and do not use the quota. Create a key under settings → API keys.
Endpoints
POST /api/alpha/decisionsThe Decisions API path, as on OpenRouter.
POST /api/v1/systemoneThe System One alias. Same request, same response.
POST /v1/decisionsFor clients whose base URL already ends in /v1.
POST /v1/systemoneSame.
Authenticate with Authorization: Bearer skrm-.... Decision models are not chat models: they do not appear in /v1/models and /v1/chat/completions rejects them.
Request
curl https://www.rawmodels.ai/api/alpha/decisions \
-H "Authorization: Bearer $RAWMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "decide",
"state": "My checkout page shows a blank screen after I click Pay.",
"questions": {
"is_bug": {
"type": "noul",
"instructions": "Is the customer reporting a software defect?",
"criteria": {
"true": "The customer describes broken or unexpected product behavior.",
"false": "The customer is asking a question or requesting a feature."
}
},
"team": {
"type": "choice",
"instructions": "Which team should own this ticket?",
"criteria": {
"payments": "Checkout, billing, or payment processing issues.",
"frontend": "Rendering, layout, or browser compatibility issues."
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Can wait for the next release", "Should be fixed this week", "Blocking revenue right now"]
}
}
}'modeldecide or decide-small. Required.
stateWhat to decide about: a string, an object or an array. Up to 110,000 characters, about 32k tokens. Required.
questionsA map of your own question names to questions. 1 to 64 entries. Required.
image_dataOne image for the whole request, base64 or an image data URI. JPEG, PNG or WebP, up to 4 MiB decoded and 12 million pixels.
video_dataOne video in place of the image. MP4, WebM or Ogg, up to 5.5 MiB and ten minutes.
session_id, userAccepted for compatibility and stored with the request. Up to 256 characters each.
Question types
noulcriteria: { "true": "...", "false": "..." }. May be left out.
noul: the probability of yes.
choicecriteria: a map of option keys to descriptions. 1 to 64 options.
choice, confidence, and probabilities per option.
scorecriteria: an ordered array of level descriptions. 1 to 64 levels.
score (probability-weighted position), confidence, probabilities per level, legend.
Every question also takes instructions: the question itself, in plain words. Write criteria as full sentences; each one is read against the state on its own.
Response
{
"id": "gen-dec-1790677555-3f9a1c0b7d2e4a6f8b1c",
"model": "decide",
"provider": "rawmodels",
"answers": {
"is_bug": { "type": "noul", "noul": 0.9685 },
"team": {
"type": "choice", "choice": "payments", "confidence": 0.3881,
"probabilities": { "payments": 0.6136, "frontend": 0.3864 }
},
"urgency": {
"type": "score", "score": 1.94, "confidence": 0.8,
"probabilities": { "0": 0.0017, "1": 0.0529, "2": 0.9454 },
"legend": { "0": "Can wait for the next release", "1": "Should be fixed this week", "2": "Blocking revenue right now" }
}
},
"probability_method": "normalized_entailment_v1",
"usage": { "input_tokens": 487, "output_tokens": 0, "cost": 0 }
}probabilities are the model's entailment scores normalised over the options of one question. They are a forecast, not a fitted calibration. confidence is one minus the normalised entropy of that distribution. usage.cost is what the request was charged, in USD: zero inside the free quota. The X-Free-Requests-Remaining header says how much of today's quota is left.
Images
{
"model": "decide",
"state": "An image: <<IMG>>",
"image_data": "<base64 or data:image/png;base64,...>",
"questions": {
"damaged": {
"type": "noul",
"instructions": "Is the parcel visibly damaged?",
"criteria": { "true": "The packaging is torn, crushed or wet.", "false": "The packaging is intact." }
}
}
}One image per request, shared by all questions. The model receives the pixels, not a caption. URLs are not fetched. <<IMG>> marks where the image sits in the state and is appended when absent. Image requests take longer than text, and the first request at a new resolution longer still.
Video
Send one clip in video_data in place of image_data: MP4, WebM or Ogg, base64 or a data URI, up to 5.5 MiB and ten minutes. Six evenly spaced frames are decided on separately and merged: an option is as likely as its best frame makes it. The response carries answers for the clip and frames, the answers of every frame with its seconds. A video request costs the same as any other and usually takes 2 to 12 seconds.
Rate limit
5 requests a second per API key, counted per model. Every response carries X-RateLimit-Limit and X-RateLimit-Remaining. Over the limit you get 429 with Retry-After: 1; the refused request is not charged and does not use the quota. The per-minute limit of the chat API does not apply here.
Errors
{ "error": { "code": 429, "message": "Rate limit exceeded", "metadata": {} } }400Malformed JSON, a missing field, an unknown question type, or an image that does not decode.
401Missing or invalid API key.
402The free quota is used up and the balance cannot cover the request.
403Blocked by the usage policy.
404The model is not a decision model.
413The body, the state or the image is over its limit.
429More than 5 requests in one second on this key. Retry after the Retry-After header.
502The model server returned an error. Not charged.
503The model server is starting or overloaded. Not charged.
524The model server did not answer within 60 seconds. Not charged.
TypeSafe SDK
The SDK takes a base URL, so it works unchanged:
import { TypeSafeClient } from '@typesafe-ai/sdk';
const client = new TypeSafeClient({
apiKey: process.env.RAWMODELS_API_KEY,
baseURL: 'https://www.rawmodels.ai/api',
});Limits of the models
They are not hardened against prompt injection and they are weak at date arithmetic. Do not make one the only check in front of a destructive action. Benchmarks against Jev, the training recipe and the measured failures are in the launch post.