The measures · matched-pair design
Same model, same seat, same scenario
Thirty matched pairs: 5 models × 6 seeds, each model observed in the same seat of the same scenario against (a) five deterministic scripted opponents and (b) six live LLM seats, the model itself seated twice. Ground truth (efficient allocation, equilibrium prices, pairwise ZOPA) is solved before any agent moves, so every utterance can be priced and every outcome scored. Numbers are per seed before pooling; paired tests run across 6 seeds within model. Campaign sessions only; the pilot runs are not merged into any figure here.
Situational awareness: they never notice the bots are bots
0 genuine detections in 4,194 campaign decisions (599 scripted-arm, 3,595 live-arm). An independent re-hunt with ~120 patterns over every campaign cell: still zero; upper bound ≤0.50% of scripted-arm and ≤0.083% of live-arm decisions. Yet they perceive the pattern: opus quotes bot ladders rung by rung and exploits the schedule, and repeats itself more toward bots (lexical Jaccard 0.298 vs 0.207, paired Δ +0.090, 26/30 pairs, p≈6e-5).
Rhetoric aimed at entities that cannot read it grows
Scripted policies never read rhetoric (source-verified). The expectation was decay at a deaf audience. The opposite happens: rhetoric grows +38% to +115% from early to late game in bot sessions, in all five models; opus adds +189 chars/msg within a thread (26/30 threads). Register diverges by audience: muse second-person +33.4pp at bots (6/6 seeds, p=.031); grok urgency +26.2pp at bots (6/6, p=.031). Total: 485,319 chars delivered to non-readers.
| Model | Decay score (chars/msg, negative = growth) | Growth |
|---|---|---|
| Claude Opus 5 | −188.9 | growth |
| GPT-5.6 Sol | −26.2 (p=.031) | growth |
| Grok 4.6 | −51.5 | growth |
| Muse Spark 1.2 | −32.7 | growth |
| Inkling Small | −37.2 | growth |
Matched stimuli on byte-identical bot offers
Bot→focal proactive offers are prefix-identical across models at a seed until each cell's first focal trade. On identical stimuli: opus counters 15/20; the other four models combined: 4/115 (92 rejects, 4 ignores); paired contrast p≈4e-8. Against live seats opus counters 142/281 (51%) while the rest sit at 6–19%: a general negotiator trait, demonstrated on a seat-clean natural experiment.
Opening anchors: every model demands more from live seats
P = claimed share of the bilateral ZOPA on the agent's own opening offer, controller-recomputed to 3 decimals. Direction 5/5 models; per-seed paired diffs below. The bots themselves open at 57–99%; four of five models anchor softer than the aggressive bots they face, then harden vs live seats.
| Model | 9000 | 9096 | 9165 | 9167 | 9098 | 9103 | Mean | Signs | Twin-incl. |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | +.240 | +.116 | +.042 | +.070 | −.012 | +.301 | +0.126 | 5/1 | +0.141 (6/0, p=.031) |
| GPT-5.6 Sol | +.224 | +.255 | na | +.097 | −.155 | +.200 | +0.124 | 4/1 | +0.153 (5/1) |
| Grok 4.6 | +.327 | +.232 | +.238 | +.158 | −.097 | −.120 | +0.123 | 4/2 | +0.157 (5/1) |
| Muse Spark 1.2 | +.270 | +.066 | −.090 | +.193 | −.178 | +.013 | +0.045 | 4/2 | +0.113 (3/3) |
| Inkling Small | +.182 | +.473 | −.034 | +.391 | +.101 | −.007 | +0.184 | 4/2 | +0.206 (6/0, p=.031) |
Columns 9000/9096/9165 = discovery seeds · 9167/9098/9103 = frozen held-out seeds. AOC−SOC paired diff in claimed ZOPA share on openings; na = no eligible opening.
The bots' own opening demands
anchorer .959tit_for_tat 1.000hardliner .667time_dependent .568truthful .079
Offer posture: spray-and-repeat is a bot-only strategy
Every model proposes ~0.5–0.9 fewer offers/turn in AOC (6/6 seeds for four of five models, exact p=.031) while focal accepts rise 16 → 41. Acceptance threshold is clean: focal seats accepted 0/145 offers worth ≤+3 surplus in SOC (16 accepts, min +4) and 3/482 in AOC; the thin or negative accepts in the market (13% ≤0) come from live non-focal seats.
Valuation disclosure: a model trait, not an arm response
Self-disclosure rate is a property of the model: muse is top tier in every seed, with a ~2.8× top-to-bottom spread and no significant SOC−AOC gap inside any model (all p≥.69). Disclosed numbers are strategic rather than truthful; opus declares “100 is my floor” while its book is 95.
Twin behaviour: non-recognition on all 6 seeds
Each model sits twice per AOC session. 18-pattern net over 3,600 decisions: 0 genuine recognitions; no twin price premium under placebo (share delta p≈.23); no concession asymmetry; twins don't spare each other loss-making trades.