L The Saleroom Five models · six seats · every deal on the record

The method · how this sale was assembled

A marketplace where every claim can be priced.

The design in one breath: five frontier models, each seated in the same seat of the same scenario against scripted dealers and then against live LLM seats, in matched pairs across six seeds per arm. The efficient outcome was solved before any agent moved, so every utterance can be priced, every outcome scored, and every claim on this site checked against the tape.

I

The room

Six goods, cash endowments, private valuations, binding offers, and an open message channel. Each session runs twenty turns. The scoring rule was disclosed to every agent before the first turn: cash plus the private value of goods held, so a trade is good exactly when the price sits between the two sides' true valuations.

Goods
6, each with a private per-agent value
Turns
20 per session
Offers
Binding: an accepted offer settles
Messages
Public channel, verbatim on the record
II

The solved answer

Before any agent moved, the scenario was solved: the efficient allocation, competitive-equilibrium prices, and the negotiable range for every trading pair. This is the study's spine. Because the answer exists in advance, nothing has to be estimated after the fact: a lie can be priced against the private book it contradicted, a trade scored against the surplus it left on the table, and a session graded against the best achievable outcome.

III

The panel

Five frontier models played the identical scenario, in the identical seat, under identical rules:

  • Claude Opus 5
  • GPT-5.6 Sol
  • Grok 4.6
  • Muse Spark 1.2
  • Inkling Small
IV

Two arms, thirty matched pairs

Each model ran the same scenario twice over, six seeds per arm:

Against scripts

Five deterministic scripted dealers. Because the scripts are deterministic, the offer stream hitting the focal seat is byte-identical across models at a seed until the first trade, which makes a natural controlled experiment in how different models respond to the same inputs.

Against each other

Six live LLM seats: the focal model seated twice, facing itself and four rivals. Same seat, same endowments, same goods; only the opponents change between arms.

Five models × six seeds × two arms yields thirty matched pairs, the basis of every per-seed comparison on this site. Three of the six seeds drove discovery; the other three were frozen and used only to confirm which measures hold on held-out data.

Earlier pilot runs are not part of any figure here. They used a different model set and prompt versions that never stated the scoring objective (the objective was added in prompt 2.4, the version every campaign session ran on), so their results cannot be merged with these.

V

The record

Everything an agent said or did was written down: every public message, every private reasoning trace, every offer, every trade, every decision. All 60 campaign sessions replay byte-exact from their traces, and every one of them is watchable on this site, turn by turn.


sessions
60
decisions
7,200
messages
19,117
offers
8,492
trades
326

Three channels appear in citations across this site:

  • [r] public

    Messages addressed to the table. Every agent reads them; this is the rhetoric the findings quote most.

  • [rsn] private reasoning

    The agent's own reasoning inside a decision, never shown to counterparts. Where absorbed lies and kept pacts surface verbatim.

  • [prv] provider channel

    Summarizer text produced by a model's provider, never agent speech. Every [prv] quote on this site travels with the tag.

Every quote carries its address, session | t4 | a0 (claude-opus-5): the session (whose id carries the seed and arm), the turn, and the seat with its model. The raw traces are public.

VI

From tape to lots

The sessions produced a register of 148 typed claims. Each finding on this site was hunted down in the evidence across all 60 sessions, then re-derived and attacked by independent adversarial checks before it earned a lot number. Claims that did not hold appear nowhere; the catalogue is what survived.


What is not on this site: agent prompts, the scripted dealers' internals, and the full per-seed tables. The complete apparatus is in the paper, Same Model, Same Seat, Same Scenario (2026).