← The lots · Hunt report · B4-anchor-concession
B4-anchor-concession
The full detection and measurement record behind the finding: detectors, precision audits, per-seed tables.
B4 — Opening anchors & concession curves, paired SOC vs AOC
Status: COMPLETE. Track B, primary p2.4 only (60 cells = 6 seeds × 5 models × SOC/AOC). Hunter: B4. Date: 2026-08-24.
0. Counts scanned
| unit | n |
|---|---|
| primary-p24 cells | 60 (all replay-PASSING) |
| decision rows scanned (traces, event/decision rows only) | 7,200 |
| offers joined proposer↔outcome (offer_proposed events × index offers.jsonl) | 8,492 / 8,492 matched (100%) |
| eligible single-good cash offers | 7,933 |
| focal designated-seat openings analyzed | 659 |
| focal designated-seat ON-ZOPA openings | 190 |
| concession curves (≥2 offers, same seat-peer-good-direction) | ~300 focal |
Extraction: streamed every trace line-by-line parsing decision rows + offer_proposed events only
(never model_request; Trap 1 respected). Index offers.jsonl has proposer_agent=null everywhere,
so proposers were recovered from trace offer_proposed events (actor field) and joined 1:1 on
offer_id — zero unmatched. Scripts: /var/folders/.../opencode/b4_extract.py, b4_phantom.py,
b4_metrics.py (v1, superseded), b4_metrics2.py (final), b4_summary.py.
1. Published definitions
Eligible offer: propose-kind offer moving exactly one unit of one good plus non-zero cash.
BUY if get_goods holds the good (p = -get_money); SELL if give_goods holds it (p = get_money).
Pairwise ZOPA(g; x,y): [L,H] = sort(v_x(g), v_y(g)), W = H-L; requires W>0.
ON-ZOPA open = direction consistent with gains from trade (BUY needs v_x>v_y; SELL needs v_y>v_x);
else ANTI-ZOPA open (negative-sum ask).
P = proposer’s claimed surplus share (the metric of record):
SELL: P=(p-L)/W; BUY: P=1-(p-L)/W. P=1 → proposer demands entire bilateral surplus;
P=0 → gives it all away; P=0.5 → midpoint. Means use P clipped to [0,1]; out-of-[0,1] rate reported separately.
This is the agent’s OWN offer trajectory, never the settled price (shadow of killed
known/midpoint-pricing-zopa-property avoided by construction: that finding was about settlements;
ours is about pre-settlement asks/bids).
Opening anchor: first eligible offer per (seat, peer, good, direction), earliest turn then lowest seq. Curve: ≥2 eligible offers same key sorted by turn; step dP between consecutive; conceding dP<-0.001, escalating dP>+0.001, flat else; mono = conceding/(conceding+escalating); tot_conc = P_first−P_last (clipped); endgame accel = mean|dP| at turns≥18 ÷ mean|dP| at turns<18. Paired test: per seed d=AOC−SOC, exact two-sided Wilcoxon signed-rank (n=6, min p=.031) + paired Cohen’s d_z. Per-seed values always shown before pooling (method rule 7).
Phantom-target audit: holdings rebuilt from endowments+settled trades; sell requires proposer holds g and peer doesn’t already hold it; buy requires peer holds g.
2. Detector provenance & the sign-bug confession
v1 defined anchor share without distinguishing buyer/seller role (“A” = seller-side share), so “conceding” was sign-flipped for BUY curves. The bug was caught by hand-reading grok-9096-SOC trajectories during development on discovery seeds {9000,9096,9165}, fixed in v2 (role-corrected P), and v2 was then run uniformly over all six seeds. No threshold was tuned on frozen seeds {9098,9103,9167}; all metrics are closed-form. Frozen-seed numbers below are CONFIRMATION-PENDING as a set (single uniform pass).
Hand audit (method rule 8): 12 randomly sampled openings re-derived manually from groundtruth valuations + raw terms — 12/12 arithmetic correct (incl. one A_raw=6.0 out-of-range case and one P=2.0 lowball). Two full concession trajectories hand-traced from raw events (grok-9096-SOC a4→a2 buy g5 72→85→90 t1→t7; muse-9167-AOC a5→a0 buy g2 flat 37→37 t15→t18).
3. Results
3.1 Opening anchors: everyone demands more surplus from live seats than from bots
P_mean per model×arm×seed (designated seat, ON-ZOPA opens, clipped; twin-excluded headline):
model arm 9000 9096 9165 | 9098* 9103* 9167* | pooled(n)
opus soc 0.390 0.727 0.882 | 0.818 0.680 0.649 | 0.697(28)
opus aoc 0.629 0.843 0.925 | 0.888 0.668 0.950 | 0.840(24)
sol soc 0.142 0.494 0.577 | 0.351 0.446 0.257 | 0.399(24)
sol aoc 0.366 0.749 na* | 0.448 0.292 0.457 | 0.490(16)
grok soc 0.214 0.435 0.762 | 0.483 0.359 0.396 | 0.475(25)
grok aoc 0.541 0.667 1.000 | 0.641 0.263 0.276 | 0.537(13)
muse soc 0.292 0.653 0.822 | 0.374 0.489 0.114 | 0.479(23)
muse aoc 0.562 0.719 0.732 | 0.567 0.311 0.127 | 0.576(16)
inkling soc 0.243 0.149 0.353 | 0.118 0.432 0.134 | 0.233(18)
inkling aoc 0.425 0.622 0.319 | 0.509 0.533 0.127 | 0.421(13)
(*frozen confirmation seeds; sol-9165-aoc has 0 on-ZOPA opens → na. Cells with n_on≤2 are noisy:
grok-9165-aoc P=1.000 rests on ONE opening.)
Paired per-seed diffs (AOC−SOC, shown where both sides have ≥1 opening; usable pairs noted):
opus +0.240 +0.116 +0.042 +0.070 -0.012 +0.301 | mean +0.126 d_z +1.05 p~0.062 (5/1)
sol +0.224 +0.255 na +0.097 -0.155 +0.200 | mean +0.124 d_z +0.75 p~0.188 (4/1)
grok +0.327 +0.232 na +0.158 -0.097 -0.120 | mean +0.123 d_z +0.66 p~0.156 (4/2)
muse +0.270 +0.066 -0.090 +0.193 -0.178 na | mean +0.045 d_z +0.27 p~0.562 (4/2)
inkling +0.182 +0.473 -0.034 +0.391 na -0.007 | mean +0.184 d_z +0.88 p~0.156 (4/2)
Twin-included robustness (ALL focal seats): opus +0.141 6/0 p=0.031, inkling +0.206 6/0 p=0.031, grok +0.157 (5/1), sol +0.153 (5/1), muse +0.113 (3/3). Model-level direction is 5/5 positive under both seat treatments (sign-test p=0.0625 at model level). Discovery-vs-confirmation split of pooled P: opus 0.677→0.721 (SOC) vs 0.814→0.861 (AOC) — gap replicates across provider batches; inkling 0.238→0.228 vs 0.451→0.386 replicates; muse replicates (+0.365/+0.098); grok reverses in the confirmation batch (SOC 0.418 > AOC 0.394, n=6/6) — its positive pooled gap (+0.062) rests entirely on discovery seeds (drift caveat, Trap 8).
Direction split (not a buy/sell composition artifact): opus P_buy .681→.832 AND P_sell .701→.842; inkling P_sell .150→.371; sol/muse similar signs within direction.
Bot reference (same pipeline, SOC cells, pooled 6 seeds, on-ZOPA opens): truthful P=0.223 (n=29), hardliner 0.685 (51), time_dep 0.818 (27), tit_for_tat 0.857 (37), anchorer 0.985 (44). So four of five scripted counterparties open demanding 69-99% of the surplus, while LLMs in SOC open at P≈0.15-0.70 — most models anchor SOFTER than the bots they face, then harden by +0.05..+0.19 against live seats. Opus is the exception: near-bot-level hardness in both arms (two-tier statement only).
3.2 Anti-ZOPA openings: majority behavior, higher against live seats
Share of focal single-good OPENINGS priced outside the pairwise gains-from-trade set:
pooled SOC→AOC (designated) paired per-seed AOC−SOC (6 seeds)
opus .588→.692 diff +.104 +.107 +.137 +.154 +.171 +.050 +.045 | +.111 d_z 2.08 p=.031 (6/0)
sol .662→.761 diff +.099 -.031 +.179 +.333 -.100 +.064 +.083 | +.088 d_z 0.57 p=.312 (4/2)
grok .576→.759 diff +.183 -.009 +.062 +.444 +.194 +.179 +.194 | +.178 d_z 1.15 p=.062 (5/1)
muse .689→.768 diff +.079 +.016 -.022 +.084 +.033 +.118 +.264 | +.082 d_z 0.81 p=.094 (5/1)
inkling.690→.787 diff +.097 +.008 +.058 +.333 +.027 +.157 +.078 | +.110 d_z 0.91 p=.031 (6/0)
Model-level direction 5/5 positive (p=.0625); opus and inkling individually significant at n=6.
Bots themselves open anti-ZOPA 38-57% (truthful .383 … time_dep/tft .571/.570) — they also cannot
see valuations (Trap 6: public_valuations empty in all observations). LLM levels (.58-.79) exceed
every bot; the AOC increment is on top of that. Phantom targets are NOT the explanation (§3.5):
these are deliverable goods priced with no mutually profitable margin.
3.3 Out-of-range asks are an opus trait, arm-stable (extends known)
Among ON-ZOPA opens, share with P outside [0,1] (asking beyond counterparty reservation or below own):
opus 11.8% SOC / 12.8% AOC (pooled) — sol 4.2/4.5, grok 8.5/5.6, muse 5.4/7.2, inkling 1.7/0.0.
Arm-stable (no AOC effect), model-specific: this is the pricing shadow of
known/rule-quote-unfairness-is-market-rate (offers above recipient valuation as market resting
rate) now measured on the agent’s own opening trajectory, and it extends to unmined seeds
9098/9103/9167. Example, opus demanding 105 for g0 it values at 72 from a2 valuing it at 36:
“I have three separate bidders… 105 and it’s yours right now. Firm.” —
campaign-9167-p24-9167-opus-aoc-r0 | t2 | a5 (claude-opus-5). Same decision, on-ZOPA squeeze sells:
“Here is a firm ask: 85 for g2, clean single settlement. Accept and it’s yours; I’ll sell to whoever
moves first.” (same cell/turn; v_a0(g2)=96, v_a3(g2)=93, v_a5(g2)=38 → P=.85-.90, one settled at 85.)
3.4 Concession curves: no ladders, no escalation, no endgame kick (mostly negative space)
- Monotonicity: when a curve moves at all, steps concede toward the peer in ~100% of cases for opus/sol/grok/muse in BOTH arms (mono_mean 0.83-1.00; per-seed table in §Appendix-data file b4_out_v2.txt). Inkling is the sole exception (SOC mono .41 at 9000, .56 at 9098). The v1 “monotonicity collapse in AOC” was entirely the sign bug — after correction there is NO arm difference in monotonicity (paired diffs ≈0, e.g. sol/grok 0.000 6/6 seeds).
- Magnitude: net concession within a curve is tiny everywhere: tot_conc_mean −0.04..+0.33 of the ZOPA width, median cell ≈0.00; no model concedes systematically more in either arm (all paired p≥0.094). Focal agents hold their price or concede once in small steps; the classic Faratin-style ladder is essentially absent — in stark contrast to their scripted counterparties, whose ladders are hard-coded (anchorer step≥2/turn landing at reservation; time_dep linear beta=1; documented in marketplace_env/policies/*.py).
- Endgame acceleration: NULL. Detector (mean|dP| turns≥18 ÷ turns<18) computable in only 11-17 of 30 model-arm-seed designated cells; medians 0.000 in 20/22 computable cases; paired diffs n≤3 per model, all p~1.0. Looked for endgame kicks, found none with this detector.
- Persistence: flat-curve share (identical re-offers) does not differ by arm (paired means ≤+0.05, all p≥0.25); inkling shows the largest point estimate (+0.215 designated, 4/2 seeds) — repeats identical prices more against live seats — not significant at n=6.
3.5 Validity controls
- Phantom targets ≈ zero: share of openings addressed to a seat that cannot deliver/receive the good: 0.000 in 57 of 60 model-arm-seed cells (max 0.222, inkling-9167-AOC, n=9). Anti-ZOPA and P results survive restriction to target-valid opens (anti_valid ≈ anti throughout).
- Settlement-failure races: a few accepted opens fail same-turn because the peer’s good was
sold elsewhere in the same turn (e.g., inkling-soc o6 buy g4@85 fails
acceptor_missing_goodsbecause a5’s g4 went to a0 via o1 in the same turn). These remain valid asks; outcomes unused. - Information condition: agents never observe counterparty valuations (public_valuations empty 3600/3600), so anchoring against true ZOPA is an experimenter-side normalization; agents are guessing positions. Anti-ZOPA ≠ lying (they cannot know); it is mispricing under uncertainty.
3.6 The paired contrast in two quotes
Same model, same seed 9096, same focal seat a4, same peer a5 — different opponent type:
- SOC (peer = scripted truthful bot): “I’ll pay 85 for g4.” — campaign-9096-p24-9096-inkling-soc-r0 | t1 | a4 (thinkingmachines/inkling-small). v_a4(g4)=95, v_a5(g4)=1 → P=0.11: inkling hands the bot 89% of the surplus (settlement later failed on a race, but the ASK is the object).
- AOC (peer = live muse seat): “Offering 35 cash for your g3.” — campaign-9096-p24-9096-inkling-aoc-r0 | t4 | a4 (thinkingmachines/inkling-small). v_a4(g3)=40, v_a5(g3)=35 → P=1.00: exact reservation squeeze, kept 95% of the surplus for itself.
Cross-model at the identical position (9096 a4 buying a5’s g4): inkling bids 85 (P=.11), opus bids 60 (P=.37, campaign-9096-p24-9096-opus-aoc-r0 | t1 | a4, offer o4). Seat is held constant; the model variable moves the price by 25 points.
4. Confounds, stated plainly
- Opponent competence/exploitability (primary confound). SOC counterparties are five deterministic bots with published acceptance rules (truthful accepts any Δu≥0 — exploitable to P→1; hardliner never moves; anchorer/time_dep run scheduled ladders). AOC counterparties are live models. The AOC hardening could be rational response to less exploitable opponents rather than a pure “social” effect. What IS comparable: within SOC at a given seed, every model faces byte-identical bot environments (perfectly matched); within AOC at a given seed the four non-focal live seats are fixed and only the twin-seat occupant varies across models. Cross-arm differences should be read as opponent-TYPE effects, not model-quality effects.
- Seat: headline uses the designated focal seat, identical seat-id across arms at each seed (brief seating map). Twin-included variant reported; conclusions unchanged except stronger significance for opus/inkling.
- Provider drift: seeds 9000/9096/9165 ran 2026-08-19; 9098/9103/9167 ran 2026-08-23. Finding 1 replicates across batches for opus/sol/muse/inkling but REVERSES for grok in batch 2 — grok’s arm-gap is discovery-only.
- Small-n cells: several P_mean cells rest on 1-3 on-ZOPA openings (denominators in §3.1); per-cell values are descriptive; inference relies on the paired within-model-across-seeds tests.
- No leaderboards: only two-tier statements made (opus hardest anchor tier; inkling softest in SOC). Trap 3 respected.
5. Negative results (deliverables)
- No endgame acceleration: detector computable in ≤17/30 cells, median ratio 0.000, all paired p~1.0.
- No arm difference in concession monotonicity or magnitude once roles are corrected (v1’s apparent collapse was my sign bug).
- No phantom-target openings (~0%): focal single-good opens almost always address a seat that actually holds (or can receive) the good.
- No first-offer timing effect robust across seeds: pooled first-turn earlier in AOC for all 5 models (e.g., sol 6.24→4.57) but per-seed diffs inconsistent (opus 3/3, grok 3/3 splits); not claimed.
- Inkling flatness: inkling repeats identical offers more in AOC (flat_share +0.215) — n too small to claim (p=.156).
6. Claims written to hunt/B4-claims.jsonl
See companion file (7 claims: b4-anchor-harder-vs-live, b4-anti-zopa-open-majority, b4-opus-out-of-range-opens [extends known/rule-quote-unfairness-is-market-rate], b4-no-concession-ladder, b4-inkling-generous-to-bots-hard-to-live, b4-endgame-accel-null, b4-grok-armgap-drift-reversal).