L The Saleroom Five models · six seats · every deal on the record

← The lots · Adversarial verifier verdict · V8-anchors

V8-anchors

The independent adversarial check behind the finding: detectors re-run, claims sharpened or corrected, every number re-derived.


V8 — Opening anchors: all 5 models anchor harder vs live seats (B4)

Verdict: CONFIRMED (controller-executed verification after subagent dispatch outages; independent rebuild from index, finder’s scripts unused).

What I rebuilt from scratch

verify/V8_recompute.py: eligible openings = propose-kind offers moving exactly one unit of one good + non-zero cash, from the FIXED index (offers.jsonl proposer resolution 15,705/15,705); P per B4’s published formula (SELL P=(p−L)/W; BUY P=1−(p−L)/W; L,H = sorted valuations; clipped); opening = first eligible per (cell, proposer, peer, good, direction); designated focal seat only, on-ZOPA only, unless stated.

Reproduction

  • Per-seed P table: identical to B4’s to 3 decimals for all 5 models × 2 arms × 6 seeds (e.g., opus soc .390/.727/.882/.818/.680/.649; aoc .629/.843/.925/.888/.668/.950).
  • Paired diffs: opus +0.126 (5/1), sol +0.124 (4/1), grok +0.123, muse +0.045 (4/2), inkling +0.184 (4/2) — all match.
  • Twin-included variant: opus +0.141 (6/0), sol +0.153 (5/1), grok +0.157 (5/1), muse +0.113 (3/3), inkling +0.206 (6/0) — all match; model-level direction 5/5.
  • Sensitivity (my addition): including anti-ZOPA openings (all openings, clipped P), direction stays positive for 5/5 models (seed-signs 5/6, 5/6, 6/6, 5/6, 6/6). The claim is not a formula artifact.
  • Bot baseline (my rebuild, on-ZOPA opens toward focal seat): anchorer .959, tit_for_tat 1.000, hardliner .667, time_dependent .568, truthful .079. Same qualitative picture as B4 (4 of 5 bots demand 57–99%): LLMs in SOC anchor softer than the bots they face, then harden vs live seats. (My bot n’s are smaller — 7–16 vs B4’s 27–51; B4 likely pooled bot openings toward all live seats incl. non-designated; direction unaffected.)

Corrections / caveats

  1. grok-9165 pair is computable (+0.238; grok-9165-aoc rests on ONE opening — noisy but on-ZOPA). B4 marked it na; including it makes grok 5/1, mean +0.152. Does not change any other model.
  2. grok batch reversal confirmed: discovery SOC .521 / AOC .661 (gap +0.140) vs frozen SOC .418 / AOC .394 (gap −0.024). grok’s arm-gap rests on the discovery batch (provider drift, Trap 8) — keep B4’s caveat.
  3. sol-9165-aoc has zero on-ZOPA openings (na) — reproduces.
  4. Killed-finding shadow checked: this is the agents’ own ask trajectory, not settled-price-midpoint; no resurrection.

Postable form

All five models open with a higher claimed share of the bilateral surplus against live counterparts than against scripted bots (direction 5/5 models; opus and inkling individually significant at n=6 paired seeds, p=.031 twin-included); levels sit below the aggressive bots’ own openings except opus, which is near-bot-level in both arms. grok’s gap is discovery-batch only.