October 1, 2026 · 5 min read
Decision models, street fighting
Loading replay…
Every match is a list of decisions. The page runs the game again from that list, frame by frame, so the replay is the match itself and not a video of it. Under each fighter: the probability the model gave every move at its latest decision, and the text it was sent.
A decision model takes a state and a list of options and returns the one to take, with a probability for each. That is also a description of a fighting game: four times a second, here is the screen, here are ten moves, pick one.
So we built a small arcade fighter and put every decision model we could reach in it: OpenJev and OpenJev Small from us, Jev 1.13 from TypeSafe, Kev 4B, Solar Decide from Upstage and Span-01 from Respan through OpenRouter, and d1 from Liquid AI. A button masher that presses at random sets the scale. Every model plays every other, 2 matches per pair with sides swapped, 56 matches in all. OpenJev is on top.
The ladder.
| # | Fighter | Elo | 95% interval | W–L–D | Win rate | Damage per round | Median latency |
|---|---|---|---|---|---|---|---|
| 1 | OpenJev RawModels | 1291 | 1134–1428 | 13–1–0 | 93% | 83 | 692 ms |
| 2 | Kev 4B Jared Palmer | 1210 | 1059–1385 | 11–3–0 | 79% | 77.3 | 1221 ms |
| 3 | d1 Liquid AI | 1210 | 1062–1352 | 11–3–0 | 79% | 71.5 | 301 ms |
| 4 | Solar Decide Upstage | 1069 | 903–1217 | 7–7–0 | 50% | 84.4 | 8825 ms |
| 5 | Jev 1.13 TypeSafe | 1035 | 865–1189 | 6–8–0 | 43% | 69.3 | 272 ms |
| 6 | Button masher uniform random | 1000 | 1000–1000 | 5–9–0 | 36% | 75.8 | — |
| 7 | OpenJev Small RawModels | 927 | 788–1091 | 3–11–0 | 21% | 66.9 | 471 ms |
| 8 | Span-01 Respan | 799 | 682–958 | 0–14–0 | 0% | 12.6 | 497 ms |
Bradley-Terry ratings on the Elo scale, with Button masher pinned at 1000. A 200-point gap means the higher fighter is expected to win three matches out of four. Intervals come from resampling the matches 300 times. A draw counts as half a win. Latency is the median round trip of a decision request; the game does not wait on it, see below.
Head to head
| Row beats column | OpenJev | Kev 4B | d1 | Solar Decide | Jev 1.13 | Button masher | OpenJev Small | Span-01 |
|---|---|---|---|---|---|---|---|---|
| OpenJev | 100% | 100% | 100% | 50% | 100% | 100% | 100% | |
| Kev 4B | 0% | 50% | 100% | 100% | 100% | 100% | 100% | |
| d1 | 0% | 50% | 100% | 100% | 100% | 100% | 100% | |
| Solar Decide | 0% | 0% | 0% | 100% | 50% | 100% | 100% | |
| Jev 1.13 | 50% | 0% | 0% | 0% | 50% | 100% | 100% | |
| Button masher | 0% | 0% | 0% | 50% | 50% | 50% | 100% | |
| OpenJev Small | 0% | 0% | 0% | 0% | 0% | 50% | 100% | |
| Span-01 | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
Rating over the tournament
The rules.
Two fighters, 100 health each, 60 seconds a round, best of three. Ten moves: walk forward, walk back, jump kick, crouch, block, jab, heavy punch, kick, low sweep and fireball. They form a loop of counters. A standing block stops everything but the sweep. A crouch ducks punches and fireballs and blocks the sweep, but kicks and jump kicks land on it. A jump clears sweeps and fireballs. The heavy punch hits hardest and leaves you open longest.
Every 15 frames, a quarter of a second, the game stops and asks both fighters at once. It only continues when both have answered, so a slow model is not punished for being slow: this ranks decisions, not latency. A request that still fails after retries leaves that fighter standing still for a quarter of a second, and it is counted.
What a model sees
Text, from its own side of the screen, with no left and no right: only toward and away. The jab reaches 78 px and the sweep 112 px, and the state says which attacks are in range right now, so no model has to do geometry.
Round 1. Rounds won: you 0, opponent 0. Time left: 39 s. Your health: 64/100. Opponent's health: 81/100. Distance to the opponent: 90 px (close). In range of your kick, sweep. The wall behind you is 340 px away. You are standing still. The opponent is doing a heavy punch, winding up. Your fireball recharges in 1.2 s. The opponent's fireball is ready.
What it is asked
The same choice question for every model and every decision, sent unchanged through each provider's own decisions API. The answer is the option with the highest probability.
{
"move": {
"type": "choice",
"instructions": "You are a fighter in a 2D fighting game. Pick the move to make right now to win the round: deal damage and avoid getting hit.",
"criteria": {
"walk_forward": "Walk toward the opponent to get into attack range.",
"walk_back": "Walk away from the opponent to get out of range.",
"jump_kick": "Jump toward the opponent and kick on the way down. Flies over sweeps and fireballs, and a standing block stops it.",
"crouch": "Crouch: punches and fireballs pass over you and you block sweeps, but kicks and jump kicks still hit you.",
"block": "Stand and block. Stops punches, kicks, jump kicks and fireballs, but a sweep still hits you.",
"light_punch": "Quick jab: 4 damage, reaches 78 px, comes out in 3 frames. Misses a crouching opponent.",
"heavy_punch": "Strong punch: 10 damage, reaches 88 px, slow (8 frames) and open to a counter if it misses. Misses a crouching opponent.",
"kick": "Kick: 7 damage, reaches 106 px, hits a crouching opponent.",
"sweep": "Low sweep: 8 damage, reaches 112 px and knocks down. Beats a standing block, a crouch blocks it, a jump avoids it.",
"fireball": "Throw a fireball: 9 damage at any distance. Crouching or jumping avoids it. Needs 2.5 s to recharge."
}
}
}Span-01 only answers yes-or-no questions, so it gets the same ten moves as ten questions in one request, “is this the best move right now?”, and plays the move with the highest yes. OpenJev goes through our endpoint, Jev, Kev, Solar and Span through OpenRouter's decisions API, d1 through Liquid's. The game runs at 60 frames a second and has no randomness: a match depends only on the decisions in it.
Fine print.
This is a toy, and it measures one thing: how well a model turns a described situation into a sensible action, thousands of times in a row, against an opponent that reacts. Each model gets one prompt we wrote, the same for all. A different description of the state could reorder the ladder. The models do not remember earlier decisions; each request stands alone.
Questions or a model to add: the decisions API is in the guide, and the OpenJev post is here.