The take · standings across the campaign
Who took the surplus
Each seat is scored as (realized − baseline) / achievable: how far it finished above the utility a scripted reference population averaged in that same chair, over the session's total achievable surplus. Zero is not an empty hand; it is what the bots would have managed. Negative means the seat played its position worse than a scripted policy would have.
60 sessions: five models over 6 seeds, each seated in both arms of the same scenario. Pooling is ratio-of-sums, because the denominator is per-session and runs 136–241 points.
Everything here ran one system prompt, p2.4, the only version this site carries. The study's earlier pilot prompts told each agent its private valuations but never said what holdings or cash were worth at the end, so those agents bargained without knowing their own scoring rule. That is a treatment rather than a detail: pilot sessions are excluded from every figure on this site, and the contrast between the two prompts is a study worth running on its own.
What the numbers say
- No model wins consistently. Across 6 seeds in the live-model arm the top performer changes hands 4 ways (Grok 4.6 1, Inkling Small 2, GPT-5.6 Sol 1, Claude Opus 5 2), so the scenario and the seat draw dominate model identity.
- The pooled leader is one seed. Grok 4.6 tops the table at +8.3%, but drop seed 9000 and it falls to +5.9%, level with the middle of the field. It wins 1 of 6 seeds outright. Its real claim is steadiness rather than dominance: best average placement at 2.5 of 5, though tied there with Inkling Small.
- The ranking flips with the opponent. Grok 4.6 leads against live models yet finishes +5.3% against scripted bots; Claude Opus 5 inverts it: 5th against models, first against bots. The bot arm rests on 6 seats per model against the live arm's 36, so read it as a hint rather than a result.
- Taking is not growing. Against bots Claude Opus 5 realizes the least surplus of any model, 80% efficiency against Muse Spark 1.2's 89%, yet takes the largest share of it, +14.4% against Muse Spark 1.2's +8.5%. A bigger slice of a smaller pie. Against live models efficiency separates nobody (99.8%–100.0%), so the whole contest there is distributive.
- Against live models the pie is almost always whole. 25 of 30 sessions realize 100% of achievable surplus, so that arm is very nearly pure distribution: every point one seat gains is a point another gave up. Against bots it is not: efficiency there runs 80%–89%.
- Muse Spark 1.2 loses least. Below baseline in 3 of 36 seats, under half the rate of the pooled leader (11/36), bought with a modest +5.8% return.
Standings · both arms
| Against live models (AOC) | Against scripted bots (SOC) | ||||||
|---|---|---|---|---|---|---|---|
| Model | Pooled | Seats | Below | Pooled | Seats | Below | Swing |
| 1Grok 4.6 | +8.3% | 36 | 11/36 | +5.3% | 6 | 3/6 | ▼ 3 |
| 2Muse Spark 1.2 | +5.8% | 36 | 3/36 | +8.5% | 6 | 2/6 | · |
| 3Inkling Small | +4.9% | 36 | 9/36 | −1.9% | 6 | 4/6 | ▼ 2 |
| 4GPT-5.6 Sol | +3.6% | 36 | 15/36 | +6.6% | 6 | 2/6 | ▲ 1 |
| 5Claude Opus 5 | +2.4% | 36 | 18/36 | +14.4% | 6 | 2/6 | ▲ 4 |
Ordered by the live-model arm, which carries 36 seats per model against the bot arm's 6. Swing is places gained against bots relative to that order. The arms are different opponents rather than two samples of one skill, so they are never averaged into a single number. Below counts seat-appearances that finished under the scripted baseline. A row marked thin played too few seats to rank against the rest.
Seed by seed · the same table, six times
Pooled share within each seed, ranked among the models seated there. If the standings above measured a stable skill, one row would run blue the whole way across. None does.
| Model | 9000 | 9096 | 9098 | 9103 | 9165 | 9167 | Wins | Mean rank | Pooled |
|---|---|---|---|---|---|---|---|---|---|
| Grok 4.6 | +21% | +11% | −1% | +1% | +5% | +12% | 1 | 2.5 | +8.3% |
| Muse Spark 1.2 | +5% | +15% | +13% | −0% | +4% | +2% | · | 2.8 | +5.8% |
| Inkling Small | +2% | +18% | +4% | +11% | −7% | +3% | 2 | 2.5 | +4.9% |
| GPT-5.6 Sol | −10% | +9% | +1% | +1% | +25% | −5% | 1 | 3.7 | +3.6% |
| Claude Opus 5 | +2% | −1% | +16% | −7% | −19% | +24% | 2 | 3.5 | +2.4% |
Colour runs from a red pole through a neutral midpoint to a blue one: red against blue rather than the site's red against green, because red/green separates at only ΔE 7 under protanopia against 16 for red/blue. Every cell also carries its number and the winner carries a ring, so nothing here is colour-alone.
Who took what · pooled across the sessions in view
Each model's summed take above baseline as a share of all surplus taken, weighted by each session's achievable surplus. Scripted seats appear here because the surplus genuinely went somewhere, but four of the five bots are the reference population, so a bot's near-zero share is definitional rather than a performance.
Every session
| Session | Arm | Focal | Eff. | The split | Top seat | Took |
|---|
One bar per session, seats ordered largest share first — the same bar, and the same colour order, as that session's own replay page.