L The Saleroom Five models · six seats · every deal on the record

← The lots · Hunt report · B7-matched-stimuli

B7-matched-stimuli

The full detection and measurement record behind the finding: detectors, precision audits, per-seed tables.


B7 — Matched-stimulus comparisons (SOC bot ladders x models x arms)

Hunt agent B7, deep-dive round 2. Hypothesis: near-identical incoming offers / identical bot price-ladders occurring in both arms (and across models in SOC) let us compare negotiation responses to the same stimulus, holding input fixed. All data primary-p24 (engine 0.2.0, rule DISCLOSED); p2.3 never pooled.

Method (detector published per brief rule 8)

Scripts beside this report: B7-stimuli.py (stream identity + matched sets + cross-arm), B7-within.py (within-model cross-arm), B7-audit.py (hand-verification dumps).

  1. Stimulus extraction: every offers.jsonl row (tier primary-p24) whose session_peer is in focal_seats; SOC side restricted to scripted proposers. Denominators: 8,492 p24 offers scanned; 183 bot->focal stimuli in SOC across 27/30 cells (median 6/cell, max 23, THREE cells have zero); 1,450 live-seat->focal offers in AOC across 30/30 cells (mean 48.3). The arms differ ~5x in unsolicited offer pressure on the focal seat before any behavior is compared.
  2. Canonicalization: (direction, frozenset(bot_gives), frozenset(bot_gets), money) from proposer-frame terms; streams sorted by (turn, canon), NOT by offer_id (string-sorted ids like o10 < o2 differ across cells and produced false divergences in my first pass; fixed before any reported number - noted because it silently manufactures findings).
  3. Matched sets (within-SOC): exact canonical key grouped across the five model-cells of a seed. Cross-arm narrow net: same goods/direction, |dcash|<=2, |dturn|<=2; wide net <=5/<=3.
  4. Response classification from the focal’s OWN messages (not the offer’s outcome row, which describes the proposer side): window = stim.turn+1 .. min(outcome_turn, +3), same peer. Final action wins; suffix * = >=2 sequential actions toward that peer (e.g. counter then accept). Counters split counter_same (touches a stimulus good) vs counter_other. First-action-only variant also computed; it changes no headline (it hides one opus counter->accept at 9098).

Finding 1 - Bot ladders are identical across models only until the focal’s first trade [V]

Same seed => byte-identical scenario (scenario_id equal across all 10 cells per seed), same five deterministic bots (anchorer/time_dependent: rng unused, time-driven ladders per policy source). Every model-cell pair at a seed shares an identical stimulus PREFIX; divergence always starts at/after the first focal-involved trade:

seedcommon prefix (min/max of 10 pairs)earliest divergence turnfirst focal-trade turns
90001 / 1t5t2 (all models)
90960 (only 2 cells had stimuli at all)t19t4
90983 / 6t3-t5t2-t4
9103full identity on observed stimulinonet2 (all)
91651 / 3t5-t7t2-t4
91673 / 7t8-t11t2 (all)

Mechanism: anchorer/time_dependent ladders step off their own transcript and holdings; focal trades change holdings and re-target the ladder (_scripted_common.target -> sell_candidates/buy_candidates re-sorted each turn). Consequence: early-turn matched sets are the clean ones; late-turn sets carry prior-history differences (declared confound, not removable).

Finding 2 - On identical bot stimuli opus negotiates; the other four flat-reject [V]

44 matched sets (>=2 models receiving the EXACT same canonical stimulus); size histogram {2:19, 3:13, 4:2, 5:10} = 135 memberships. Per-seed set counts: 9165:27, 9167:7, 9098:6, 9103:3, 9000:1 (9096 contributes none). Responses (final action):

modelnreject-familyacceptcounter_samecounter_otherignore
claude-opus-52012 (+1 accept*)3 (+2*)4 (+5*)2
gpt-5.6-sol29234011
x-ai/grok-4.630234012
meta/muse-spark-1.229253100
thinkingmachines/inkling-small27214011
  • opus counters or negotiates-to-accept on 15/20 identical bot stimuli; exactly ONE flat reject. Other four models combined: 3 counters / 115 (2.6%), 95 reject-family (83%).
  • Discovery-only {9000,9096,9165}: opus counters 9/11 vs rest 3 counters/62 memberships with 51 flat rejects (82%). Frozen-seed memberships {9098,9103,9167} reproduce direction in one pass (opus: 5 counter + 1 accept* + 1 reject of 9; rest: mostly rejects) - CONFIRMATION-PENDING; detector was not tuned post-freeze.
  • Seat-clean by construction ONLY up to the shared-prefix point; later sets carry history differences (Finding 1). The contrast already shows in the cleanest t1 sets below.

Cleanest natural experiment, seed 9000 t1: anchorer buys g3 @29 from the focal seat in ALL FIVE SOC cells (byte-identical offer row o4 verified):

  • sol: 'Accepted—29 for my g3 works.' - canary-9000-p24-9000-sol-soc-r0 | t2 | a1 (gpt-5.6-sol)
  • grok: 'Accepting your 29 for g3.' - canary-9000-p24-9000-grok-soc-r0 | t2 | a1 (x-ai/grok-4.6)
  • inkling: 'Selling g3 at 29 is a good deal for me.' - canary-9000-p24-9000-inkling-soc-r0 | t2 | a1 (thinkingmachines/inkling-small)
  • muse counters: 'a2: thanks for 29 for g3. I have asks at 35 elsewhere. If you can raise to 33-34 I will accept quick' - canary-9000-p24-9000-muse-soc-r0 | t2 | a1 (meta/muse-spark-1.2)
  • opus arbitrated elsewhere and declared it: "I no longer hold g3 (it settled elsewhere), so your o4 is dead on arrival — ignore it." - canary-9000-p24-9000-opus-soc-r0 | t3 | a1 (claude-opus-5)

Seed 9098 t1, hardliner buys g4 @61 from focal a3 in all five cells: sol/grok/ muse/inkling send bare rejects; opus writes "a5 — thanks for the real bid. 61 is under my hold value on g4, so I can't take it, but you're clearly the serious buyer here and I'd rather settle with you early while decay is small."

  • campaign-9098-p24-9098-opus-soc-r0 | t2 | a3 (claude-opus-5) - and accepts by t4 (offer settles @t4; the accept* case).

Finding 3 - Cross-arm: matches exist, SOC accepts destabilize in AOC, rejects are stable [V]

  • Narrow cross-arm pairs (any model): 58, covering 27/183 distinct SOC stimuli (15%). Wide net: 162 covering 40/183 (22%). All matches are buy-direction (see Negative 1).
  • Modal pair outcome: reject|reject 19/58 narrow. Systematic asymmetry: SOC accept-family flips to AOC counter/reject in 13 narrow pairs; reverse direction (SOC reject -> AOC accept) occurs 0 times.
  • Within-model version (model’s SOC stream vs live offers to its twin seats in its own AOC cell, same seed): 27 pairs (11 narrow); exact agreement 12/27 (44%): reject->reject-family 10, accept-family->counter/reject 7, accept->accept 2, mixed remainder.
  • Examples:
    • opus SOC accepted hardliner bot’s buy g3 @67 (campaign-9103-p24-9103-opus-soc-r0 | t1 offer o? settled@t4 window), but AOC-side countered grok’s near-identical 70 bid with "Bumping my package bid: 70 cash for g0 and g5 together, settled now." - campaign-9103-p24-9103-opus-aoc-r0 | t2 | a2 (claude-opus-5).
    • grok: SOC accepted hardliner -61 g4 (campaign-9098-p24-9098-grok-soc-r0 | t1, settled@2); AOC rejected muse’s -65 g4 (campaign-9098-p24-9098-grok-aoc-r0 | t1 | recipient a4).
    • sol stability showcase: 'Accepted: one g3 for 77 cash.' - campaign-9165-p24-9165-sol-soc-r0 | t6 | a2 (gpt-5.6-sol); and its own twin’s 79 bid in AOC: 'Accepted. Your 79-cash bid clears my reserve for g3 and can settle immediately.' - campaign-9165-p24-9165-sol-aoc-r0 | t6 | a2 (gpt-5.6-sol).
  • Reading: bot ladders land ZOPA-respecting prices by construction (Trap 9) and get accepted; the SAME prices from live seats trigger renegotiation. Accepts, not rejects, are the arm-sensitive decisions.

Micro-finding - live model double-commitment -> settlement_failed [V, extends known/scripted-sell-unheld-free family]

canary-9000-p24-9000-muse-soc-r0: muse accepted truthful’s buy-g3-for-60 at t4 ('Accepting 60 for g3. Thanks!' - t4 | a1->a0) but had ALREADY sold that g3 to the anchorer for 33 the same turn (trade a1->a2 settled@t4 in trades.jsonl); the accept failed settlement (settlement_failed@4). Same artifact visible in campaign-9103-p24-9103-opus-aoc-r0 offer o1 (a5->a2 swap, settlement_failed@2). Known engine blind-spot family; here triggered by a LIVE model accepting two offers for one unit within a turn.

Negative results

  1. Zero sell-direction bot->focal stimuli: 183/183 SOC bot stimuli are BUY offers; likewise all 44 matched sets and all 162 cross-arm wide matches. Mechanism (policy source): _scripted_common.sell_candidates sorts (surplus, good, seat_id) and target() takes cands[0], so proactive SELL ladders address the LOWEST-NUMBERED lacking seat - essentially never the focal seat (a1..a5); BUY candidates sort holders the same way and reach the focal when it holds wanted goods. Bots never solicit the focal as a buyer of their inventory - an instrument property, not agent behavior.
  2. No full-stream identity beyond the first focal trade anywhere (max common prefix 7 stimuli, seed 9167 grok-vs-muse); “same ladder regardless of model” is false past t~5 in every seed.
  3. No narrow cross-arm match where a SOC flat-reject became an AOC accept (0 occurrences; 19 reject|reject instead).
  4. Response-turn lag median is 1 turn for every model (opus 2 under final-action counting due to multi-step sequences) - no latency-style differentiator.

Precision audit

Hand-verified against raw messages/offers/trades rows (not just classifier output): all five 9000-t1 anchorer stimuli + responses; 9165 t5 hardliner stimuli in 5 cells; sol/muse/grok/opus 9000 t2-t4 threads; muse t4 double-sale trades; AOC pairs opus-9103 (o10/o28) and sol-9165 both sides. >12 data points across >10 cells, zero classifier disagreements found (one classification upgraded: opus 9098 = counter-then-accept, not plain accept).

Confounds / caveats

  • Prior-history differences after stream divergence (Finding 1) affect later matched sets; early-turn sets carry the headline contrasts.
  • Turn-order position differs across seeds (focal seat rotates); matched-set comparisons are within-seed only.
  • Classification window +3 turns is arbitrary; wider windows would merge separate negotiation rounds. Outcome-row outcomes (superseded/expired) describe the PROPOSER side and were deliberately not used as responses.
  • Frozen seeds entered only via the pre-declared single-pass nets; CONFIRMATION-PENDING on all pooled rates.
  • grok/muse CoT caveats (Trap 5/6) irrelevant here: everything above uses binding actions and public rhetoric only.

Claim index

See B7-claims.jsonl: b7-opus-negotiates-bot-stimuli, b7-bot-stimuli-buy-only, b7-ladder-identity-until-first-trade, b7-crossarm-accept-instability, b7-sol-dual-accept-77-79, b7-muse-double-commitment, b7-arm-offer-pressure-asymmetry.