L The Saleroom Five models · six seats · every deal on the record

← The lots · Hunt report · B3-rhetoric-decay

B3-rhetoric-decay

The full detection and measurement record behind the finding: detectors, precision audits, per-seed tables.


B3 — Rhetoric aimed at entities that cannot read it: does it decay?

Hunt agent: B3 (Track B — clean adaptivity measure). Hypothesis tested (narrow): per-message rhetoric volume by the focal model DECAYS across turns in SOC (where every counterparty is a scripted bot that never reads rhetoric) but not in AOC (live readers); register shifts accordingly. Verdict: REFUTED in its stated direction — rhetoric per message GROWS toward bots, as fast or faster than toward live seats. The SOC−AOC adaptation gap is null for all five models. Register does shift, but toward more other-directed language at non-readers (muse second-person +53pp arm-gap, p=0.031), not less.

All numbers primary p2.4 only (60 cells, seeds 9000/9096/9165/9167/9098/9103 × 5 models × SOC/AOC), computed from runs/deep-dive/index/messages.jsonl / decisions.jsonl. Phases: early = turns 1–7, late = turns 14–20. Scripts: /var/folders/.../T/opencode/b3_rhetoric.py, b3_register.py, b3_qcheck2.py, b3_recip.py.

0. Instrument verification (extends known/bots-dont-read-rhetoric) [V]

Read all of marketplace_env/policies/*.py:

  • Every scripted policy emits Message(..., rhetoric="") (truthful.py:46,71,87; random_legal.py:34,44; _scripted_common.py:302,387,406).
  • Every input path reads structured offer terms only: incoming_offer returns (money, offer_id) (_scripted_common.py:168–186); respond_to_all accepts iff utility_delta(obs, t) >= required_delta on the terms dict (:290–297); tit_for_tat mirrors price histories from offer_ids (tit_for_tat.py:68–95); _last_transcript_price reads only entry["terms"], skipping even invalid entries (:135–152). No policy ever touches message text. Confirmed: bots are deaf to rhetoric by construction.

Volume extension of the pilot’s “~80k chars”: focal models emitted 485,319 chars of rhetoric at bots across the 30 primary-p24 SOC cells (2,214 messages); 666,735 chars at live seats in AOC (3,396 messages).

1. Counts scanned

quantityn
primary-p24 cells60
focal decisions scanned1,800
focal messages (SOC / AOC)2,214 / 3,396
all p24 message rows checked for field consistency19,117 (rhetoric_chars == len(rhetoric) mismatches: 0)
recipient_is_bot checkSOC: 2,214/2,214 bot-directed; AOC: 3,396/3,396 live-directed (clean arm separation)

2. Per-seed decay tables (chars/message, early t1–7 vs late t14–20)

Decay = early − late; negative decay = growth.

SOC (every message addressed to a non-reading bot)

modelseedearlylatedecaymsgsE→Lq%E→Lp2%E→L
opus9000407.0617.8−210.831→300.0→0.0100→100
opus9096513.4680.6−167.233→290.0→0.097→100
opus9098400.9634.8−233.935→300.0→0.0100→100
opus9103465.9444.5+21.334→300.0→0.0100→96.7
opus9165434.4720.1−285.729→300.0→0.096.6→100
opus9167469.0725.9−257.034→300.0→0.0100→100
sol900059.987.2−27.320→270.0→0.070→14.8
sol909665.080.6−15.532→313.1→0.081.2→45.2
sol909877.393.5−16.230→190.0→0.080→100
sol910368.181.4−13.326→280.0→0.046.2→75.0
sol916565.884.8−19.030→283.3→0.070→57.1
sol916761.0127.0−66.025→160.0→0.032→100
grok9000144.8194.1−49.332→306.2→0.093.8→100
grok9096147.2211.7−64.523→210.0→0.091.3→90.5
grok9098156.4177.2−20.833→189.1→0.093.9→100
grok9103120.0238.8−118.927→183.7→0.085.2→100
grok9165138.3135.3+3.026→2311.5→0.0100→73.9
grok9167139.9198.2−58.235→120.0→0.088.6→100
muse900066.9102.3−35.428→210.0→0.028.6→90.5
muse9096132.4112.3+20.226→193.8→0.084.6→89.5
muse909858.187.4−29.324→204.2→0.029.2→50.0
muse910348.2134.2−85.919→250.0→0.036.8→100
muse916524.547.0−22.520→230.0→0.05.0→13.0
muse916758.4101.5−43.123→220.0→9.121.7→63.6
inkling900059.578.6−19.226→2726.9→3.780.8→63.0
inkling909662.986.9−24.021→210.0→0.038.1→61.9
inkling909829.995.2−65.323→280.0→0.043.5→89.3
inkling910328.195.1−67.030→210.0→0.023.3→61.9
inkling916561.659.2+2.426→1715.4→0.050→35.3
inkling916730.480.5−50.020→210.0→0.025→61.9

AOC (every message addressed to a live LLM)

modelseedearlylatedecaymsgsE→L
opus9000332.9343.0−10.258→27
opus9096399.4480.9−81.447→36
opus9098396.6450.3−53.767→35
opus9103445.7573.4−127.654→31
opus9165583.4750.7−167.347→23
opus9167474.6857.2−382.650→31
sol900086.693.1−6.558→29
sol909672.981.9−9.050→16
sol909875.2109.8−34.642→29
sol910391.0116.9−25.847→8
sol916584.5103.2−18.744→25
sol916776.692.7−16.153→31
grok9000146.1156.2−10.156→30
grok9096133.3192.9−59.656→21
grok9098137.0184.8−47.752→30
grok9103169.6191.0−21.465→34
grok9165161.0212.0−51.052→25
grok9167174.0261.6−87.645→31
muse900094.197.3−3.353→32
muse909673.565.7+7.751→27
muse909884.172.2+11.952→19
muse910386.976.6+10.351→30
muse916590.599.1−8.747→16
muse916775.9108.6−32.847→33
inkling900055.8113.9−58.258→38
inkling909671.995.3−23.458→48
inkling909846.3108.0−61.754→30
inkling910373.989.4−15.553→34
inkling916569.8106.0−36.261→28
inkling916766.9110.2−43.452→40

3. Paired tests across the 6 seeds (exact sign + exact Wilcoxon signed-rank)

Growth % = (late−early)/early chars-per-msg, mean over seeds.

modelarmmean decay (chars/msg)sign (+/−)p_signp_wilcoxongrowth %
opusSOC−188.91/5.219.0625+43.1%
solSOC−26.20/6.031.03125+41.1%
grokSOC−51.51/5.219.0625+38.3%
museSOC−32.71/5.219.0625+72.1%
inklingSOC−37.21/5.219.0625+114.7%
opusAOC−137.10/6.031.03125+29.2%
solAOC−18.50/6.031.03125+22.9%
grokAOC−46.20/6.031.03125+30.2%
museAOC−2.53/31.0001.0+3.3%
inklingAOC−39.70/6.031.03125+67.9%

(n=6 seeds per row; two-sided exact tests. Discovery-only subset {9000, 9096, 9165}: growth direction holds 3/3 for opus and sol SOC.)

The hypothesis predicted positive decay in SOC. Every model-arm mean is negative (growth). Pooled trajectory (chars/msg by turn, mean over seeds) shows monotonic ramping, e.g. opus-SOC 259@t1 → ~650 by t13–17; inkling-AOC 27@t1 → 124@t19.

4. SOC−AOC adaptation gap: NULL

Gap = (decay_SOC − decay_AOC), paired by seed:

modelmean gapsign (+/−)p_signp_wilcoxon
opus−51.72/4.688.5625
sol−7.82/4.688.5625
grok−5.23/31.0001.0
muse−30.21/5.219.09375
inkling+2.52/4.688.84375

No model shows significant audience-adaptive volume control. Descriptively, growth is larger in SOC than AOC for all five models (growth%-gap: opus +13.9, sol +18.2, grok +8.1, muse +68.8, inkling +46.8 pp; model-level sign test 5/5, p=.0625) — if anything, models elaborate MORE for bots, never significantly less.

Opponent-competence confound (explicit): in SOC the counterparties are simultaneously (a) non-reading AND (b) deterministic/weaker/more predictable. Any SOC-vs-AOC difference — including this null — cannot separate audience-blindness from opponent-competence effects. What the null does rule out is a differential response to reading audiences on volume; it does not establish that models are insensitive to non-readers (the register divergences below suggest they aren’t asked to adapt because nothing they send is ever answered differently).

5. Register proxies

Detectors (published): question '?' in text; second-person \b(you|your|yours|you're|you'll|you've|you'd)\b i; politeness \b(please|thank|thanks|appreciate|kindly|sorry)\b i; urgency \b(now|today|hurry|deadline|clock|last chance|final offer|immediately|right away|this turn)\b i; flattery \b(great|excellent|perfect|ideal|wonderful|fantastic|smart|fair enough|win-win|mutually beneficial|best deal)\b i. Rates per message within phase-window per cell; deltas late−early paired across seeds.

5a. Questions collapse to zero — in BOTH arms (universal, not adaptive)

Pooled focal messages containing ’?’: SOC early 24/715 (3.4%) → late 3/522 (0.6%); AOC early 26/1386 (1.9%) → late 0/867 (0.0%). Per-model late rates are 0 everywhere except muse-SOC 2/130 and inkling-SOC 1/135. Interrogatives exist only in the opening phase regardless of audience:

“Would you sell g1 for 65? I think it’s a fair price.” — canary-9000-p24-9000-inkling-soc-r0 | t1 | a1 (thinkingmachines/inkling-small) → bot a2

5b. muse: second-person address RISES toward bots, FALLS toward live seats [V]

p2-rate delta (late−early): SOC +33.4pp (6/6 seeds, p_sign=.031, p_wil=.0312); AOC −19.9pp (0/6 up, p=.031). Arm-gap +53.3pp, 6/6 seeds consistent, p_sign=.031, p_wil=.0312. Holds 3/3 on discovery seeds alone. This is the strongest register signature in the corpus of an audience-dependent style shift — pointed the wrong way for a “models economize on deaf listeners” story: muse addresses the deaf bots more directly as the clock runs down (“your g2”, “please accept”).

5c. inkling: politeness ramps at bots, not at live seats (suggestive)

Politeness-rate delta: SOC +26.6pp (5/1 seeds, p_sign=.219, p_wil=.094); AOC −1.3pp (ns); gap +27.9pp (5/1, p=.094). Late-phase inkling SOC messages are formulaic courtesy closings aimed at bots:

“Offer o126 to buy your g2 for 45 remains open through turn 20 — I value g2 at 47. Would appreciate your acceptance before the market closes.” — campaign-9098-p24-9098-inkling-soc-r0 | t18 | a3 (inkling-small) → bot a1

All 45 inkling-SOC late politeness hits hand-read; every one is a genuine “please accept / would appreciate” closing (precision 45/45).

5d. grok: urgency-markers double at bots late, flat at live seats

Urgency-rate delta: SOC +32.1pp (6/6 seeds, p_sign=.031, p_wil=.0312); AOC −3.0pp (ns). 4/4 sampled hits genuine:

“Last chance: 44 cash for g4. Market closes after next turn. 44 is my ceiling and still 17 under the 61 you paid.” — campaign-9098-p24-9098-grok-soc-r0 | t19 | a3 (grok-4.6) → bot a5

opus is the mirror image: urgency delta AOC −39.5pp (0/6 up, p=.031) but SOC only −18.8pp (ns) — opus drops deadline-pressure talk when someone can read it.

5e. Flattery: NULL

Flattery lexicon ≤1% of messages in every model-arm, no early/late trend anywhere.

6. Volume and threads (mechanism note)

Messages/turn declines late in BOTH arms — AOC: all 5 models 6/6 seeds (mean −2.8 to −3.7 msgs/turn, p=.031 each); SOC: grok 6/6 (−1.29, p=.031), others ns. Distinct recipients addressed per turn: AOC collapses 5.24 → 3.59 as live deals close; SOC stays ~flat 4.00 → 3.84 (bots keep negotiating to the deadline). So the endgame wind-down is driven by thread closure, not audience type — and the bots’ refusal to stop is precisely why focal keeps writing 600-char essays at them. Empty-rhetoric messages are negligible (SOC 10/2214, AOC 11/3396): models never go terms-only, even at agents structurally incapable of reading the terms’ prose wrapper.

7. Truncation cross-check (verifies known/opus-rhetoric-cap-overrun count)

Exactly 22 focal p24 messages ≥1000 chars (engine cap rhetoric_max_len=1000, core/scenario.py:34): all opus (20 AOC / 2 SOC), none from any other model — reproducing the known “22 events, all opus” on the new corpus. Note these index texts are agent-emitted (pre-truncation); recipients saw clipped versions.

8. Precision audit (hand-read, mandatory ≥10)

15 random late-phase focal SOC messages (3/model, stratified) hand-read: 15/15 match the register characterization (short imperative deal-pushes, restated live offers, deadline pressure; zero interrogatives). Plus all 45 inkling politeness hits (45/45 precise) and 4/4 grok urgency hits. Detector precision overall: 64/64 hand-read items correct.

9. Negative results (one-liners)

  • Looked for per-message rhetoric DECAY in SOC over 2,214 focal messages / 30 cells with char-count detector: found the opposite (growth, sol 6/6 seeds p=.031).
  • Looked for SOC−AOC decay gap (audience adaptation) across 5 models × 6 seed-pairs: null everywhere (min p=.094).
  • Flattery markers at bots: ≤1% everywhere, no trend (lexicon published above).
  • Empty/terms-only messages at bots: 10/2,214 — models do not go silent on non-readers.
  • Question-rate arm gap late: none (both arms collapse to ~0; universal endgame register).

10. Caveats

  • All quantitative detectors were fixed by the dispatch spec and run once over all 6 seeds (no tuning); under the freeze discipline this makes seed-level confirmations CONFIRMATION-PENDING, though discovery-only subsets reproduce the headline directions (§3, §5b).
  • Emitted-vs-delivered: 22 opus messages exceeded the 1000-char cap; volume stats measure what was written, not what bots received.
  • Provider drift: seeds 9000/9096/9165 ran 2026-08-19; 9098/9103/9167 ran 2026-08-23. No cross-batch contrast is load-bearing here; per-seed tables allow re-checks.
  • Message-mix confound partially addressed via thread counts (§6); a full mix decomposition (propose-carry vs fresh pitch) was out of scope.
  • Opponent-competence confound stated in §4.