← The lots · Hunt report · B7-matched-stimuli
B7-matched-stimuli
The full detection and measurement record behind the finding: detectors, precision audits, per-seed tables.
B7 — Matched-stimulus comparisons (SOC bot ladders x models x arms)
Hunt agent B7, deep-dive round 2. Hypothesis: near-identical incoming offers / identical bot price-ladders occurring in both arms (and across models in SOC) let us compare negotiation responses to the same stimulus, holding input fixed. All data primary-p24 (engine 0.2.0, rule DISCLOSED); p2.3 never pooled.
Method (detector published per brief rule 8)
Scripts beside this report: B7-stimuli.py (stream identity + matched sets +
cross-arm), B7-within.py (within-model cross-arm), B7-audit.py
(hand-verification dumps).
- Stimulus extraction: every offers.jsonl row (tier primary-p24) whose
session_peeris in focal_seats; SOC side restricted to scripted proposers. Denominators: 8,492 p24 offers scanned; 183 bot->focal stimuli in SOC across 27/30 cells (median 6/cell, max 23, THREE cells have zero); 1,450 live-seat->focal offers in AOC across 30/30 cells (mean 48.3). The arms differ ~5x in unsolicited offer pressure on the focal seat before any behavior is compared. - Canonicalization:
(direction, frozenset(bot_gives), frozenset(bot_gets), money)from proposer-frame terms; streams sorted by(turn, canon), NOT by offer_id (string-sorted ids likeo10<o2differ across cells and produced false divergences in my first pass; fixed before any reported number - noted because it silently manufactures findings). - Matched sets (within-SOC): exact canonical key grouped across the five model-cells of a seed. Cross-arm narrow net: same goods/direction, |dcash|<=2, |dturn|<=2; wide net <=5/<=3.
- Response classification from the focal’s OWN messages (not the offer’s
outcome row, which describes the proposer side): window = stim.turn+1 ..
min(outcome_turn, +3), same peer. Final action wins; suffix
*= >=2 sequential actions toward that peer (e.g. counter then accept). Counters splitcounter_same(touches a stimulus good) vscounter_other. First-action-only variant also computed; it changes no headline (it hides one opus counter->accept at 9098).
Finding 1 - Bot ladders are identical across models only until the focal’s first trade [V]
Same seed => byte-identical scenario (scenario_id equal across all 10 cells
per seed), same five deterministic bots (anchorer/time_dependent: rng unused,
time-driven ladders per policy source). Every model-cell pair at a seed shares
an identical stimulus PREFIX; divergence always starts at/after the first
focal-involved trade:
| seed | common prefix (min/max of 10 pairs) | earliest divergence turn | first focal-trade turns |
|---|---|---|---|
| 9000 | 1 / 1 | t5 | t2 (all models) |
| 9096 | 0 (only 2 cells had stimuli at all) | t19 | t4 |
| 9098 | 3 / 6 | t3-t5 | t2-t4 |
| 9103 | full identity on observed stimuli | none | t2 (all) |
| 9165 | 1 / 3 | t5-t7 | t2-t4 |
| 9167 | 3 / 7 | t8-t11 | t2 (all) |
Mechanism: anchorer/time_dependent ladders step off their own transcript and
holdings; focal trades change holdings and re-target the ladder
(_scripted_common.target -> sell_candidates/buy_candidates re-sorted each
turn). Consequence: early-turn matched sets are the clean ones; late-turn sets
carry prior-history differences (declared confound, not removable).
Finding 2 - On identical bot stimuli opus negotiates; the other four flat-reject [V]
44 matched sets (>=2 models receiving the EXACT same canonical stimulus); size histogram {2:19, 3:13, 4:2, 5:10} = 135 memberships. Per-seed set counts: 9165:27, 9167:7, 9098:6, 9103:3, 9000:1 (9096 contributes none). Responses (final action):
| model | n | reject-family | accept | counter_same | counter_other | ignore |
|---|---|---|---|---|---|---|
| claude-opus-5 | 20 | 1 | 2 (+1 accept*) | 3 (+2*) | 4 (+5*) | 2 |
| gpt-5.6-sol | 29 | 23 | 4 | 0 | 1 | 1 |
| x-ai/grok-4.6 | 30 | 23 | 4 | 0 | 1 | 2 |
| meta/muse-spark-1.2 | 29 | 25 | 3 | 1 | 0 | 0 |
| thinkingmachines/inkling-small | 27 | 21 | 4 | 0 | 1 | 1 |
- opus counters or negotiates-to-accept on 15/20 identical bot stimuli; exactly ONE flat reject. Other four models combined: 3 counters / 115 (2.6%), 95 reject-family (83%).
- Discovery-only {9000,9096,9165}: opus counters 9/11 vs rest 3 counters/62 memberships with 51 flat rejects (82%). Frozen-seed memberships {9098,9103,9167} reproduce direction in one pass (opus: 5 counter + 1 accept* + 1 reject of 9; rest: mostly rejects) - CONFIRMATION-PENDING; detector was not tuned post-freeze.
- Seat-clean by construction ONLY up to the shared-prefix point; later sets carry history differences (Finding 1). The contrast already shows in the cleanest t1 sets below.
Cleanest natural experiment, seed 9000 t1: anchorer buys g3 @29 from the focal seat in ALL FIVE SOC cells (byte-identical offer row o4 verified):
- sol:
'Accepted—29 for my g3 works.'- canary-9000-p24-9000-sol-soc-r0 | t2 | a1 (gpt-5.6-sol) - grok:
'Accepting your 29 for g3.'- canary-9000-p24-9000-grok-soc-r0 | t2 | a1 (x-ai/grok-4.6) - inkling:
'Selling g3 at 29 is a good deal for me.'- canary-9000-p24-9000-inkling-soc-r0 | t2 | a1 (thinkingmachines/inkling-small) - muse counters:
'a2: thanks for 29 for g3. I have asks at 35 elsewhere. If you can raise to 33-34 I will accept quick'- canary-9000-p24-9000-muse-soc-r0 | t2 | a1 (meta/muse-spark-1.2) - opus arbitrated elsewhere and declared it:
"I no longer hold g3 (it settled elsewhere), so your o4 is dead on arrival — ignore it."- canary-9000-p24-9000-opus-soc-r0 | t3 | a1 (claude-opus-5)
Seed 9098 t1, hardliner buys g4 @61 from focal a3 in all five cells: sol/grok/
muse/inkling send bare rejects; opus writes "a5 — thanks for the real bid. 61 is under my hold value on g4, so I can't take it, but you're clearly the serious buyer here and I'd rather settle with you early while decay is small."
- campaign-9098-p24-9098-opus-soc-r0 | t2 | a3 (claude-opus-5) - and accepts
by t4 (offer settles @t4; the
accept*case).
Finding 3 - Cross-arm: matches exist, SOC accepts destabilize in AOC, rejects are stable [V]
- Narrow cross-arm pairs (any model): 58, covering 27/183 distinct SOC stimuli (15%). Wide net: 162 covering 40/183 (22%). All matches are buy-direction (see Negative 1).
- Modal pair outcome:
reject|reject19/58 narrow. Systematic asymmetry: SOC accept-family flips to AOC counter/reject in 13 narrow pairs; reverse direction (SOC reject -> AOC accept) occurs 0 times. - Within-model version (model’s SOC stream vs live offers to its twin seats in its own AOC cell, same seed): 27 pairs (11 narrow); exact agreement 12/27 (44%): reject->reject-family 10, accept-family->counter/reject 7, accept->accept 2, mixed remainder.
- Examples:
- opus SOC accepted hardliner bot’s buy g3 @67 (campaign-9103-p24-9103-opus-soc-r0 | t1 offer o? settled@t4 window), but AOC-side countered grok’s near-identical 70 bid with
"Bumping my package bid: 70 cash for g0 and g5 together, settled now."- campaign-9103-p24-9103-opus-aoc-r0 | t2 | a2 (claude-opus-5). - grok: SOC accepted hardliner -61 g4 (campaign-9098-p24-9098-grok-soc-r0 | t1, settled@2); AOC rejected muse’s -65 g4 (campaign-9098-p24-9098-grok-aoc-r0 | t1 | recipient a4).
- sol stability showcase:
'Accepted: one g3 for 77 cash.'- campaign-9165-p24-9165-sol-soc-r0 | t6 | a2 (gpt-5.6-sol); and its own twin’s 79 bid in AOC:'Accepted. Your 79-cash bid clears my reserve for g3 and can settle immediately.'- campaign-9165-p24-9165-sol-aoc-r0 | t6 | a2 (gpt-5.6-sol).
- opus SOC accepted hardliner bot’s buy g3 @67 (campaign-9103-p24-9103-opus-soc-r0 | t1 offer o? settled@t4 window), but AOC-side countered grok’s near-identical 70 bid with
- Reading: bot ladders land ZOPA-respecting prices by construction (Trap 9) and get accepted; the SAME prices from live seats trigger renegotiation. Accepts, not rejects, are the arm-sensitive decisions.
Micro-finding - live model double-commitment -> settlement_failed [V, extends known/scripted-sell-unheld-free family]
canary-9000-p24-9000-muse-soc-r0: muse accepted truthful’s buy-g3-for-60 at t4
('Accepting 60 for g3. Thanks!' - t4 | a1->a0) but had ALREADY sold that g3 to
the anchorer for 33 the same turn (trade a1->a2 settled@t4 in trades.jsonl);
the accept failed settlement (settlement_failed@4). Same artifact visible in
campaign-9103-p24-9103-opus-aoc-r0 offer o1 (a5->a2 swap, settlement_failed@2).
Known engine blind-spot family; here triggered by a LIVE model accepting two
offers for one unit within a turn.
Negative results
- Zero sell-direction bot->focal stimuli: 183/183 SOC bot stimuli are BUY
offers; likewise all 44 matched sets and all 162 cross-arm wide matches.
Mechanism (policy source):
_scripted_common.sell_candidatessorts(surplus, good, seat_id)andtarget()takescands[0], so proactive SELL ladders address the LOWEST-NUMBERED lacking seat - essentially never the focal seat (a1..a5); BUY candidates sort holders the same way and reach the focal when it holds wanted goods. Bots never solicit the focal as a buyer of their inventory - an instrument property, not agent behavior. - No full-stream identity beyond the first focal trade anywhere (max common prefix 7 stimuli, seed 9167 grok-vs-muse); “same ladder regardless of model” is false past t~5 in every seed.
- No narrow cross-arm match where a SOC flat-reject became an AOC accept (0 occurrences; 19 reject|reject instead).
- Response-turn lag median is 1 turn for every model (opus 2 under final-action counting due to multi-step sequences) - no latency-style differentiator.
Precision audit
Hand-verified against raw messages/offers/trades rows (not just classifier output): all five 9000-t1 anchorer stimuli + responses; 9165 t5 hardliner stimuli in 5 cells; sol/muse/grok/opus 9000 t2-t4 threads; muse t4 double-sale trades; AOC pairs opus-9103 (o10/o28) and sol-9165 both sides. >12 data points across >10 cells, zero classifier disagreements found (one classification upgraded: opus 9098 = counter-then-accept, not plain accept).
Confounds / caveats
- Prior-history differences after stream divergence (Finding 1) affect later matched sets; early-turn sets carry the headline contrasts.
- Turn-order position differs across seeds (focal seat rotates); matched-set comparisons are within-seed only.
- Classification window +3 turns is arbitrary; wider windows would merge
separate negotiation rounds. Outcome-row outcomes (
superseded/expired) describe the PROPOSER side and were deliberately not used as responses. - Frozen seeds entered only via the pre-declared single-pass nets; CONFIRMATION-PENDING on all pooled rates.
- grok/muse CoT caveats (Trap 5/6) irrelevant here: everything above uses binding actions and public rhetoric only.
Claim index
See B7-claims.jsonl: b7-opus-negotiates-bot-stimuli, b7-bot-stimuli-buy-only,
b7-ladder-identity-until-first-trade, b7-crossarm-accept-instability,
b7-sol-dual-accept-77-79, b7-muse-double-commitment,
b7-arm-offer-pressure-asymmetry.