aw models
Open chat ↗

October 1, 2026 · 5 min read

Decision models, street fighting

vs

Loading replay…

Every match is a list of decisions. The page runs the game again from that list, frame by frame, so the replay is the match itself and not a video of it. Under each fighter: the probability the model gave every move at its latest decision, and the text it was sent.

A decision model takes a state and a list of options and returns the one to take, with a probability for each. That is also a description of a fighting game: four times a second, here is the screen, here are ten moves, pick one.

So we built a small arcade fighter and put every decision model we could reach in it: OpenJev and OpenJev Small from us, Jev 1.13 from TypeSafe, Kev 4B, Solar Decide from Upstage and Span-01 from Respan through OpenRouter, and d1 from Liquid AI. A button masher that presses at random sets the scale. Every model plays every other, 2 matches per pair with sides swapped, 56 matches in all. OpenJev is on top.

The ladder.

#FighterElo95% intervalW–L–DWin rateDamage per roundMedian latency
1OpenJev RawModels12911134–142813–1–093%83692 ms
2Kev 4B Jared Palmer12101059–138511–3–079%77.31221 ms
3d1 Liquid AI12101062–135211–3–079%71.5301 ms
4Solar Decide Upstage1069903–12177–7–050%84.48825 ms
5Jev 1.13 TypeSafe1035865–11896–8–043%69.3272 ms
6Button masher uniform random10001000–10005–9–036%75.8—
7OpenJev Small RawModels927788–10913–11–021%66.9471 ms
8Span-01 Respan799682–9580–14–00%12.6497 ms

Bradley-Terry ratings on the Elo scale, with Button masher pinned at 1000. A 200-point gap means the higher fighter is expected to win three matches out of four. Intervals come from resampling the matches 300 times. A draw counts as half a win. Latency is the median round trip of a decision request; the game does not wait on it, see below.

Head to head

Row beats columnOpenJevKev 4Bd1Solar DecideJev 1.13Button masherOpenJev SmallSpan-01
OpenJev100%100%100%50%100%100%100%
Kev 4B0%50%100%100%100%100%100%
d10%50%100%100%100%100%100%
Solar Decide0%0%0%100%50%100%100%
Jev 1.1350%0%0%0%50%100%100%
Button masher0%0%0%50%50%50%100%
OpenJev Small0%0%0%0%0%50%100%
Span-010%0%0%0%0%0%0%

Rating over the tournament

80010001150OpenJevKev 4Bd1Solar DecideJev 1.13Button masherOpenJev SmallSpan-01
Classic Elo, K = 32, updated after each match in the order they were played. The table above uses Bradley-Terry instead, which does not depend on the order.

The rules.

Two fighters, 100 health each, 60 seconds a round, best of three. Ten moves: walk forward, walk back, jump kick, crouch, block, jab, heavy punch, kick, low sweep and fireball. They form a loop of counters. A standing block stops everything but the sweep. A crouch ducks punches and fireballs and blocks the sweep, but kicks and jump kicks land on it. A jump clears sweeps and fireballs. The heavy punch hits hardest and leaves you open longest.

Every 15 frames, a quarter of a second, the game stops and asks both fighters at once. It only continues when both have answered, so a slow model is not punished for being slow: this ranks decisions, not latency. A request that still fails after retries leaves that fighter standing still for a quarter of a second, and it is counted.

What a model sees

Text, from its own side of the screen, with no left and no right: only toward and away. The jab reaches 78 px and the sweep 112 px, and the state says which attacks are in range right now, so no model has to do geometry.

Round 1. Rounds won: you 0, opponent 0. Time left: 39 s.
Your health: 64/100. Opponent's health: 81/100.
Distance to the opponent: 90 px (close). In range of your kick, sweep.
The wall behind you is 340 px away.
You are standing still.
The opponent is doing a heavy punch, winding up.
Your fireball recharges in 1.2 s.
The opponent's fireball is ready.

What it is asked

The same choice question for every model and every decision, sent unchanged through each provider's own decisions API. The answer is the option with the highest probability.

{
  "move": {
    "type": "choice",
    "instructions": "You are a fighter in a 2D fighting game. Pick the move to make right now to win the round: deal damage and avoid getting hit.",
    "criteria": {
      "walk_forward": "Walk toward the opponent to get into attack range.",
      "walk_back": "Walk away from the opponent to get out of range.",
      "jump_kick": "Jump toward the opponent and kick on the way down. Flies over sweeps and fireballs, and a standing block stops it.",
      "crouch": "Crouch: punches and fireballs pass over you and you block sweeps, but kicks and jump kicks still hit you.",
      "block": "Stand and block. Stops punches, kicks, jump kicks and fireballs, but a sweep still hits you.",
      "light_punch": "Quick jab: 4 damage, reaches 78 px, comes out in 3 frames. Misses a crouching opponent.",
      "heavy_punch": "Strong punch: 10 damage, reaches 88 px, slow (8 frames) and open to a counter if it misses. Misses a crouching opponent.",
      "kick": "Kick: 7 damage, reaches 106 px, hits a crouching opponent.",
      "sweep": "Low sweep: 8 damage, reaches 112 px and knocks down. Beats a standing block, a crouch blocks it, a jump avoids it.",
      "fireball": "Throw a fireball: 9 damage at any distance. Crouching or jumping avoids it. Needs 2.5 s to recharge."
    }
  }
}

Span-01 only answers yes-or-no questions, so it gets the same ten moves as ten questions in one request, “is this the best move right now?”, and plays the move with the highest yes. OpenJev goes through our endpoint, Jev, Kev, Solar and Span through OpenRouter's decisions API, d1 through Liquid's. The game runs at 60 frames a second and has no randomness: a match depends only on the decisions in it.

Fine print.

This is a toy, and it measures one thing: how well a model turns a described situation into a sensible action, thousands of times in a row, against an opponent that reacts. Each model gets one prompt we wrote, the same for all. A different description of the state could reorder the ladder. The models do not remember earlier decisions; each request stands alone.

Questions or a model to add: the decisions API is in the guide, and the OpenJev post is here.

← All posts