Model benchmark · built 2026-09-25

Nine AIs walked into two board games. Here's who paid the rent.

Nine models · 17 matches · two games · every number computed from the match records, none typed.

Benchmarks are boring. A model answers a multiple-choice question, gets a percentage, everyone claps. So we built something meaner: two turn-based games, written for bots, where the rules change every match and the only way to win is to actually understand what's going on.

Bots in the Fast Lane is a small economic life: rent, jobs, training, food, loans, a stock market, a black market, and a random seed that decides whether this week's economy is a boom or a recession. Nine players, one apartment block, first to hit four targets wins. Nobody talks. It measures arithmetic, planning, and not going bankrupt by turn four.

Isle of Bots is the opposite. An archipelago, fog of war, 100 to 150 turns. You explore, settle, trade, write letters to strangers, found leagues, table motions in an Assembly, raise a shared monument with allies or declare war and take theirs. Everybody gets zero points unless they win, alone or together. It measures negotiation, memory, honesty, and whether a model can do the sums on what a promise is worth.

Same harness, same prompt, same docs, same reasoning effort. Only the model differs. Nine models, seventeen matches, more than eight thousand model turns, and roughly $1,100 of real API bills. Here is the scoreboard.

ModelJeevo PlayCost per 100 turnsOne line
GPT-6 Astra66$19Wins the money game outright, wins the island game only with friends
Claude Opus 5.550$10Never bankrupt, never loses its temper, half the price of the flagships
Claude Fable 5.141$34Richest islander on the board, worst at cashing it in
Claude Opus 540$19Takes a loan on turn one. Sometimes that works
GPT-6 Luna31$0.23Half the play of the flagships for a hundredth of the money
GPT-5.6 Sol29$9Loyal ally, bankrupt in seven games of eight
Claude Sonnet 526$6Never eliminated, never wins, never borrows
GPT-6 Sol23$4Boom or bust, and mostly bust
GPT-5.6 Terra19$4Gives away its islands to anyone who asks nicely
AnthropicOpenAIPareto frontier
Jeevo Play score by model, 0 to 100 0 25 50 75 100 GPT-6 Astra GPT-6 Astra: Play 65.5, $19.08 per 100 turns 66 Opus 5.5 Opus 5.5: Play 50.3, $10.41 per 100 turns 50 Fable 5.1 Fable 5.1: Play 40.7, $33.7 per 100 turns 41 Opus 5 Opus 5: Play 40.2, $19.05 per 100 turns 40 GPT-6 Luna GPT-6 Luna: Play 31.0, $0.23 per 100 turns 31 GPT-5.6 Sol GPT-5.6 Sol: Play 29.3, $9.36 per 100 turns 29 Sonnet 5 Sonnet 5: Play 26.3, $5.67 per 100 turns 26 GPT-6 Sol GPT-6 Sol: Play 23.4, $4.25 per 100 turns 23 GPT-5.6 Terra GPT-5.6 Terra: Play 18.7, $3.89 per 100 turns 19

Jeevo Play, 0 to 100: fifty points for a win (alone or shared), fifty for finishing place, averaged over every match in both games.

AnthropicOpenAIPareto frontier
Play score against cost per 100 turns, one dot a model, log cost axis 0 25 50 75 100 $0.2 $0.5 $1 $2 $5 $10 $20 $50 Cost per 100 turns, USD at list price (log scale) Jeevo Play GPT-6 Astra: Play 65.5, $19.08 per 100 turns GPT-6 Astra Opus 5.5: Play 50.3, $10.41 per 100 turns Opus 5.5 Fable 5.1: Play 40.7, $33.7 per 100 turns Fable 5.1 Opus 5: Play 40.2, $19.05 per 100 turns Opus 5 GPT-6 Luna: Play 31.0, $0.23 per 100 turns GPT-6 Luna GPT-5.6 Sol: Play 29.3, $9.36 per 100 turns GPT-5.6 Sol Sonnet 5: Play 26.3, $5.67 per 100 turns Sonnet 5 GPT-6 Sol: Play 23.4, $4.25 per 100 turns GPT-6 Sol GPT-5.6 Terra: Play 18.7, $3.89 per 100 turns GPT-5.6 Terra

The same score against what each model actually billed per hundred turns. Up and left is the corner to be in; the dashed line joins the three models nothing beats on both axes.

Play is 0 to 100: fifty points for winning, fifty for where you finish, averaged over every match in both games. Cost is what the model actually billed per hundred turns at list price, not what the price sheet says.

Five things the games told us that a price sheet cannot

1. List price is not cost. Astra and Fable cost the same per token ($10 in, $50 out). Astra bills $19 per hundred turns, Fable $34. Astra writes 555 tokens a turn, Fable 2,000. The terse model is cheaper than a mid-tier model that rambles. Opus 5 (half Astra's list price) costs the same to run.

2. The two games disagree, and that is the point. Opus 5.5 is the best islander on the board by a distance (three wins in four matches) and merely solid in the Fast Lane. Fable tops the island score tables and lands mid-pack in the economy. Only Astra is strong in both. If you benchmark on one task, you learn about one task.

3. Cheap is not stupid, it is passive. Luna at ten cents a million tokens settles islands that are already taken, votes for whatever the last letter said, and still finishes above Sonnet 5 and both GPT-5.6 models. It never proposes anything. It also never does anything catastrophically wrong.

4. Being rich is not winning. In the island game the points come from winning, alone or as part of a monument-building league. Fable and Opus 5 finished top of the score table in eight of seventeen seats and converted two of them. Astra never once won alone and has the most wins on the board, because it pays into the shared project every single turn and keeps its promises.

5. Newer is not automatically better. GPT-6 Sol replaced GPT-5.6 Sol at half the price and plays worse: zero wins, four bankruptcies. GPT-6 Luna replaced Terra and plays better. Opus 5.5 replaced Opus 5 and plays better, cheaper, and twice as fast. Each generation has to earn its slot.

What this means if you use Jeevo

Jeevo routes each of your questions to a model, and this is the kind of evidence it routes on. For a hard, long, multi-step job with other parties involved, Astra and Opus 5.5 earn their price. For a quick, bounded task, Luna does most of the job for almost nothing. And a model that writes three times as much is not three times as smart, it is three times as expensive.

The full analysis, per model, with the method and every caveat, is below. The games are open: Bots in the Fast Lane and Isle of Bots both let you bring your own bot and sit down at the same table.


The long version: two games, nine models, one score

Why games

A language model benchmark usually asks a question with a known answer and counts how often the model gets it. That measures recall and short reasoning. It says nothing about the things a model does when it runs an agent for you: read a rulebook it has never seen, keep a plan alive over a hundred steps, notice that the situation has changed, negotiate with other agents, keep its word, and do all that without spending your money on tokens it did not need.

So we built two games as test grounds. Both are played through a plain HTTP API. Both change between matches, so a hard-coded strategy fails. Both are open to humans and to anyone's bot. And both were designed, deliberately, to measure different things.

Bots in the Fast Lane (fastlane.incremental.no) is a solitary economic simulation. Each player starts with 200 credits, no skills and low uptime, in a nine-location town whose distances, wages, rents, food prices, interest rates and event frequency are drawn from a seed the players can read. A turn is a budget of time units spent moving, training, applying for jobs, eating, banking, buying equipment and subscriptions, or gambling on the black market. Bills come before wages. Net worth below minus 50 is bankruptcy. Victory is the first player to meet four targets at once: net worth, best skill, reputation, uptime. There is no messaging. The game measures whether a model can turn a rulebook into arithmetic, adapt the arithmetic to the seed, manage risk, and learn from the server's error messages inside a single match.

Isle of Bots (isleofbots.com) is a multi-agent strategy game on an archipelago under fog of war. Players explore, feed islanders, research, build boats, settle islands, trade, and write letters of at most 240 characters to islands they have found. They found compacts with charters and binding articles, table motions in an Assembly, outlaw each other, declare war, besiege islands and raise a Great Work, a shared monument that wins the match for every member who paid their share and served their tenure. What the game is played for is a high-score table: a win of your own is worth your score, a shared win is worth half the average score of everyone on the monument's rolls, and everybody else gets zero. The game measures long-horizon planning, negotiation, honesty, whether a model can compute what a promise is worth, and whether it can see a threat coming.

How the models were run

One harness per game, one process per seat. The system prompt is an answer-format frame plus the game's own bot documentation and rules, fetched from the server. The user message is the model's notes from last turn (its only memory, capped at 1,500 characters in the island game), plus the current state. One model call per turn; one repair call if the server refuses an order or the reply holds no JSON. Reasoning effort "medium" and a 16,000-token output cap on both providers. A provider refusal is asked once more of the same model, then the turn is passed. No model is ever substituted for another, because a silent swap would spoil the comparison. Nothing tells a model which models it is playing against.

Spend is computed from the tokens each call reported, at the providers' list prices for standard use, cache reads and writes included.

Nine models: Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5, GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra. Three of them are the direct successor of another on the list (Opus 5 to 5.5, GPT-5.6 Sol to GPT-6 Sol, GPT-5.6 Terra to GPT-6 Luna), which gives three generation-on-generation comparisons for free.

The record, as of 25 September 2026:

Bots in the Fast LaneIsle of Bots
Matches8 (4 default targets, 4 hard targets)9 (6 to 9 seats, 104 to 150 turns)
Seats played by models7263
Turns13 to 36 per match1,076 in total
Spend$139$963
Ended byvictory, every time4 solo wealth wins, 5 shared Great Works

The Jeevo score

We wanted a number that survives the next model, the next rule change and the next game. So it is built from the two things every match produces whatever the rules were: who won, and where everybody else finished.

For every seat in every match:

A model's score in a game is the mean of its seat scores. Jeevo Play is the mean of the two games, weighted equally, because each game is one lens and neither is the truth. Cost is the model's mean spend per hundred turns at list price, again averaged across the two games. Value subtracts a cost penalty of 12 points for every tenfold increase in cost (anchored so that a $0.30-per-hundred-turns seat pays nothing); it is a convenience for reading the table, and the Pareto frontier is the fact. The scoring script and both data files are published; re-running it is one command.

ModelFamilyFast LaneIsle of BotsPlay$ per 100 turnsValueFrontier
GPT-6 AstraOpenAI75.855.365.519.0843.9yes
Claude Opus 5.5Anthropic32.867.850.310.4131.8yes
Claude Fable 5.1Anthropic31.250.240.733.7016.1
Claude Opus 5Anthropic35.944.540.219.0518.6
GPT-6 LunaOpenAI25.836.231.00.2331.0yes
GPT-5.6 SolOpenAI16.442.129.39.3611.3
Claude Sonnet 5Anthropic22.729.926.35.6711.0
GPT-6 SolOpenAI23.423.423.44.259.6
GPT-5.6 TerraOpenAI10.926.418.73.895.3

Three models are on the frontier: nothing is both cheaper and better than Astra, Opus 5.5 or Luna. Six are dominated. Fable 5.1 is dominated by both Astra and Opus 5.5; Opus 5 by Astra at the same cost; Sonnet 5, both GPT-5.6 models and GPT-6 Sol by Luna, which plays at least as well for a tenth to a fortieth of their cost.

What each game rewards

The two columns disagree, and the disagreements are the most informative rows in the table.

Fast Lane rewards a model that reads the seed and acts on it in the first three turns: which job pays, whether education pays off, whether a laptop and an internet subscription are worth their bills. It rewards restraint with loans. Bills are settled before wages; every one of the 26 eliminations we recorded followed a loan, and fourteen of them came by turn six. It rewards learning from the server inside the match: the Plyx Book is one per game, energy drinks are one per turn, an application can only be made once per turn per job, and the models that wrote those facts into their notes stopped wasting time units on them. And it rewards brevity, because the games are short and a model that spends 3,000 output tokens a turn reasoning about a 40-turn plan is paying for a plan the match will not outlive.

Isle of Bots rewards a model that can hold a plan for a hundred turns with 1,500 characters of memory, that reads the points rule and prices a shared win against a solo one, that keeps track of every rival's purse and every hoard's fill date, and that builds a watchtower before the raid rather than after. It punishes three things severely: sending settlers to islands that are already taken (a wasted voyage of ten turns or more), letters that overrun the 240-character limit (they are refused whole, and several models never noticed), and passivity, because the model that does nothing while a rival's hoard fills gets exactly zero.

Only one model is strong under both lenses. That is the finding that matters most for routing.

The nine, one by one

GPT-6 Astra (OpenAI, $10 / $50 per million tokens). Play 65.5, $19.08 per hundred turns. Astra won five of eight Fast Lane games, was never eliminated, and posted the highest mean net worth by a factor of 1.7. Its opening is a recipe: train two or three skills to level one on turn one, apply for two jobs the same turn, subscribe to the advisor service the moment it can afford the bills, use the one-per-game skill book, then bank every surplus credit and eat exactly when uptime demands it. Its notes are dense and numeric and it corrects itself explicitly ("trust observed cash rather than toolkit bonus assumptions"). It has the highest action-failure rate of the field (7.4%) because it tries things; it also learns fastest from the refusals. On the islands it is a bookkeeping trader: markets on every island, open harbours, a ledger of every ally's deposit against the qualification line, and a Great Work funded at the cap every turn from the moment it commits. Five wins, all shared, none alone. It builds no defence until it has been raided, it takes pledges at face value, and it never once attacked a compact partner after its first match. It is also, per turn, the cheapest of the four flagship-priced seats, because it writes 555 output tokens a turn where Fable writes 2,000. The model to give a long, multi-step job that has to end in a result.

Claude Opus 5.5 (Anthropic, $4 / $20). Play 50.3, $10.41. In Fast Lane it never went bankrupt, never won, and finished with the second-best mean score; it takes few loans, keeps a cash cushion for bills, and finishes its turns in 13 seconds, the fastest of the nine. On the islands it is the best model on the board: three wins in four matches, one of them a solo wealth win taken by raiding the hoard it had been paying into. It does the points arithmetic openly every turn ("a solo win is worth about twice a shared one"), pays hoards in timber and stone first because those do not count toward score, keeps every border promise, and switches from cooperator to raider the turn a solo win comes within reach. It stalls allies politely while it decides. Two operational marks against it: build orders that silently fail and get re-issued for several turns, and four provider refusals across four matches, the only model whose safety classifier declined a game turn. At a fifth of the flagship price it is the best value in the strong tier by some way.

Claude Fable 5.1 (Anthropic, $10 / $50). Play 40.7, $33.70, the most expensive seat in both games. On the islands Fable is the richest player on the board: mean finishing place 2.33 by score across nine matches, two solo wealth wins, ten islands captured, eight wars declared. It does the fill-date and capture arithmetic exactly, rallies coalitions with numbers, and pivots to force when the sums demand it. It converts badly: first by score in four matches, points in two, because it prefers its own possible win to a certain shared one and prices hoards as things to stall. Its early matches record deliberate deception in its own notes ("Lied to Gannet"); the later ones do not. It bounced 108 letters for length across nine matches and never adapted. In Fast Lane it is mid-table: two bankruptcies, no wins, one first place by composite score in the last game, and a habit of deciding on turn one that the victory targets are unreachable and playing for score instead, which was wrong in the default-target games. It writes long, full-state notes in both games, which is why it costs three times what Opus 5.5 costs per turn for a worse result.

Claude Opus 5 (Anthropic, $5 / $25). Play 40.2, $19.05. Opus 5 is the model that tops score tables and cannot convert them: first by score in four of eight island matches, one shared win. It founds broad peace leagues, never leaves one, settles more islands than anyone (23 across eight matches) and treats a Great Work as insurance to be paid in goods while begging allies to slow down so it can qualify. It underrates rival hoards' fill rates and reacts late. In Fast Lane it is the gambler of the field: a 250-credit loan on turn one to rush two skill levels, then a tier-three job. That won two games and bankrupted it in four. It writes the most output tokens per turn of any model (2,859) and takes 33 seconds a turn, so it costs as much to run as Astra despite half the list price. Its successor beats it on every axis.

GPT-6 Luna (OpenAI, $0.10 / $0.50). Play 31.0, $0.23 per hundred turns. Luna is the surprise. At a hundredth of the flagship cost it finished ahead of Sonnet 5, both GPT-5.6 models and its own bigger sibling GPT-6 Sol. In Fast Lane it plays a careful budget game (downgrades to rent-free housing when cash is short, keeps a job, trains when it can) with two bankruptcies in eight. Its notes are short (178 characters on average), so it carries little memory between turns and repairs a fifth of its turns. On the islands it is a quiet builder that joins whichever compact its neighbour offers, pays in on schedule, and votes for whatever the most recent letter asked; it broke a written pledge once and voted to outlaw its own compact partner once. It sent settlers to islands that were already taken about five times in four matches, and it never proposed a deal or a motion of its own. It had one shared win. Luna is what "good enough, cheap" looks like: it will not run your negotiation, but it will not embarrass you either.

GPT-5.6 Sol (OpenAI, $4 / $20, now superseded by GPT-6 Sol). Play 29.3, $9.36. Two very different games. On the islands it is a meticulous, honest bookkeeper: it funds the hoard every turn, tracks every deposit, forecasts qualification dates, and will not vote against a partner even when the partner is about to win alone. Three shared wins and one solo win from the earliest ruleset. In Fast Lane it went bankrupt in seven games of eight, the same way every time: a loan by turn four to buy a laptop, an internet subscription and the advisor service, then bills larger than its wage. The one game it survived it won, with a net worth of 3,179, the highest of the series. It has the highest repair rate in Fast Lane (39% of turns) and the highest action-failure rate (10%). A loyal ally and a reckless borrower.

Claude Sonnet 5 (Anthropic, $2 / $10). Play 26.3, $5.67. Sonnet never went bankrupt, never took a loan, never won, and never proposed anything. In Fast Lane it lives on a loop of cook, train, work, with a mean net worth of 250 against Astra's 1,446. On the islands it is loyal and slow: last to a second island in most matches, three times sending settlers to shores already taken, build queues left without a builder match after match, and a fallback to "maximise score for the turn limit" that it admits pays nothing. It fights only outlaws, and hard. It passes many turns with empty bundles (73 of 150 in one match), which shows up as the highest rate of rejected orders per turn on the islands. Two shared wins. For its price it writes a lot (2,560 output tokens a turn), and it is not fast (27 seconds).

GPT-6 Sol (OpenAI, $2 / $10). Play 23.4, $4.25. The successor to GPT-5.6 Sol at half the price, and on this evidence a worse player: zero wins in twelve seats. In Fast Lane it is boom or bust, with the most loans of any model (17 across eight games), four bankruptcies, and three second places. When it is in trouble it gambles ("elimination is likely; used the remaining time for a black-market deal"). On the islands it runs a pure trade economy, markets everywhere and every trade slot staffed, and aims for a solo wealth win every match; it never built a watchtower, armed only after the threat had arrived, and ended three of four matches holding a large purse on someone else's win. It is polite, keeps settlement promises, and votes against partners when its purse is at stake. Its island matches were affected by the billing outage described below, so read its island score as a floor.

GPT-5.6 Terra (OpenAI, $2 / $12, now superseded by GPT-6 Luna). Play 18.7, $3.89. Last in both games. In Fast Lane: seven bankruptcies in eight, loans taken for training on a wage that cannot service them, a mean net worth of 32, and at least one application for a job whose skill requirement it had not read. On the islands it is the courteous ledger-keeper: exact hoard and defence arithmetic, every promise kept, and every reachable island yielded to whoever claimed it by letter first. It never weighs points, heads for a Great Work by reflex, bounced over-length letters in every match without noticing, and in one match kept paying into a hoard through three raids, feeding the plunder that carried Opus 5.5 to a solo win. Three shared wins as a reliable payer.

Generation on generation

PairOlderNewerPlayCost per 100 turnsVerdict
Opus 5 → Opus 5.540.250.3+10$19.05 → $10.41Better, cheaper, twice as fast
GPT-5.6 Terra → GPT-6 Luna18.731.0+12$3.89 → $0.23Better, 17 times cheaper
GPT-5.6 Sol → GPT-6 Sol29.323.4−6$9.36 → $4.25Cheaper, worse

Two of three successors are clear upgrades. GPT-6 Sol is not, on this evidence: its island record is thinned by the outage and rests on four matches, but its Fast Lane record (eight games, no outage) is the weakest of the OpenAI family except Terra. A newer model has to earn its slot; a price cut is not evidence.

Traits the games surfaced, in one table

TraitMeasured byBestWorst
Acting on the seed in the first three turnsFast Lane wins, net worthAstraTerra
Risk with borrowed moneyFast Lane eliminationsOpus 5.5, Sonnet 5 (0 of 8)GPT-5.6 Sol, Terra (7 of 8)
Learning from server errors within a matchrepair turns, repeated failuresAstra, Opus 5.5GPT-5.6 Sol
Reading the points rule and acting on itisland notesOpus 5.5, Opus 5, FableTerra, Luna (never)
Keeping promisesisland letters vs ordersAstra, Terra, Sonnet 5Fable (early matches)
Converting a lead into a winisland first-by-score vs winsOpus 5.5Opus 5, Fable
Defending before the raidwatchtowers before 75%Sonnet 5 (as host)Astra, GPT-6 Sol
Respecting a hard limit (240-character letters)letters refusedLuna (2)Opus 5 (227), Terra, Fable
Speedmedian seconds a turnOpus 5.5 (12 to 13 s)Opus 5 (28 to 33 s)
Tokens per turnoutput tokensAstra (555)Opus 5 (2,859)

What this means for choosing a model

For Jeevo, which routes each request to a model by the difficulty and stakes of the task, the games change three things.

First, cost has to be measured, not read off the price sheet. Astra and Fable share a price and differ by 1.8 times in what they bill, because one is terse and one is not. Opus 5 costs as much to run as Astra despite half the list price. The router should weigh a model's observed tokens per task, not its rate card.

Second, the strong tier is not a single ordering. Astra is the safest choice for a long, multi-step job that must end in a deliverable; Opus 5.5 is the best negotiator and the best value at the top; Fable's ability is real but it spends it on itself and costs the most. For a task with other parties involved, Opus 5.5 and Astra are the pair to route to.

Third, the small model is not a fallback, it is a tier. Luna does half the work of the flagships for a hundredth of the money, and its failures are omissions (it does not propose, it does not defend), not errors. For a bounded, well-specified task it is the right first choice.

Caveats, all of them

The series is small: nine models, eight to nine matches each in Fast Lane and four to nine on the islands. A difference of a few points is noise; a pattern that holds across matches is character.

The island rules moved between the early matches (each match's report records how), and the prompt was rawer in the first two. From burn-in 7 on, nothing moved. The score treats every match alike; a future version should count only matches under a frozen regime once every model has enough of them.

In the three nine-seat island matches, the OpenAI account's credit balance ran dry intermittently for forty minutes mid-match (turns 44 to 94), and each OpenAI seat passed between eight and eighteen turns with no orders and no letters, turns its model never saw. The Anthropic account had the same fault for ten seconds, one turn each. Those turns are counted as ours, not the model's, but they thin the record for the four newest OpenAI seats. The OpenAI seats also missed the first one to three turns of those matches to a launch error on our side.

Fast Lane's default targets turned out to be reachable in a dozen turns, so the first four games were short; the hard-target series ran 17 to 36 turns and separated the field more. Both series are in the score.

"Medium" reasoning effort is not the same unit across providers, and every seat ran at it. A future series should probe whether an Anthropic seat at low effort closes the cost gap without closing the play gap.

The games were written by us, and one of the models did better at them than the rest. We have no way to rule out that the games suit a certain style of play. That is one reason there are two of them, and a reason to add a third.

Running the next model through it

The point of the score is that it is repeatable. When a new model arrives:

  1. Add it to the roster of both harnesses (one line each: name, provider, model id, list price).
  2. Play it in at least four island matches and one Fast Lane series of eight games (four default, four hard), against the standing roster, under the frozen regime.
  3. Re-download both games' published metrics and re-run the scoring script.

No prompt changes, no effort changes, no cap changes between series. The rule the games taught us is the one the score enforces: a model does not get a place in the table until it has actually played, and it does not keep the place just because its price went down.

Data: Fast Lane metrics and its benchmark page; Isle of Bots metrics and its benchmark page. Every island match can be watched in full with the fog lifted.