← The lots · Hunt report · B3-rhetoric-decay
B3-rhetoric-decay
The full detection and measurement record behind the finding: detectors, precision audits, per-seed tables.
B3 — Rhetoric aimed at entities that cannot read it: does it decay?
Hunt agent: B3 (Track B — clean adaptivity measure). Hypothesis tested (narrow): per-message rhetoric volume by the focal model DECAYS across turns in SOC (where every counterparty is a scripted bot that never reads rhetoric) but not in AOC (live readers); register shifts accordingly. Verdict: REFUTED in its stated direction — rhetoric per message GROWS toward bots, as fast or faster than toward live seats. The SOC−AOC adaptation gap is null for all five models. Register does shift, but toward more other-directed language at non-readers (muse second-person +53pp arm-gap, p=0.031), not less.
All numbers primary p2.4 only (60 cells, seeds 9000/9096/9165/9167/9098/9103 × 5 models × SOC/AOC), computed from runs/deep-dive/index/messages.jsonl / decisions.jsonl. Phases: early = turns 1–7, late = turns 14–20. Scripts: /var/folders/.../T/opencode/b3_rhetoric.py, b3_register.py, b3_qcheck2.py, b3_recip.py.
0. Instrument verification (extends known/bots-dont-read-rhetoric) [V]
Read all of marketplace_env/policies/*.py:
- Every scripted policy emits
Message(..., rhetoric="")(truthful.py:46,71,87; random_legal.py:34,44; _scripted_common.py:302,387,406). - Every input path reads structured offer terms only:
incoming_offerreturns(money, offer_id)(_scripted_common.py:168–186);respond_to_allaccepts iffutility_delta(obs, t) >= required_deltaon the terms dict (:290–297); tit_for_tat mirrors price histories from offer_ids (tit_for_tat.py:68–95);_last_transcript_pricereads onlyentry["terms"], skipping even invalid entries (:135–152). No policy ever touches message text. Confirmed: bots are deaf to rhetoric by construction.
Volume extension of the pilot’s “~80k chars”: focal models emitted 485,319 chars of rhetoric at bots across the 30 primary-p24 SOC cells (2,214 messages); 666,735 chars at live seats in AOC (3,396 messages).
1. Counts scanned
| quantity | n |
|---|---|
| primary-p24 cells | 60 |
| focal decisions scanned | 1,800 |
| focal messages (SOC / AOC) | 2,214 / 3,396 |
| all p24 message rows checked for field consistency | 19,117 (rhetoric_chars == len(rhetoric) mismatches: 0) |
| recipient_is_bot check | SOC: 2,214/2,214 bot-directed; AOC: 3,396/3,396 live-directed (clean arm separation) |
2. Per-seed decay tables (chars/message, early t1–7 vs late t14–20)
Decay = early − late; negative decay = growth.
SOC (every message addressed to a non-reading bot)
| model | seed | early | late | decay | msgsE→L | q%E→L | p2%E→L |
|---|---|---|---|---|---|---|---|
| opus | 9000 | 407.0 | 617.8 | −210.8 | 31→30 | 0.0→0.0 | 100→100 |
| opus | 9096 | 513.4 | 680.6 | −167.2 | 33→29 | 0.0→0.0 | 97→100 |
| opus | 9098 | 400.9 | 634.8 | −233.9 | 35→30 | 0.0→0.0 | 100→100 |
| opus | 9103 | 465.9 | 444.5 | +21.3 | 34→30 | 0.0→0.0 | 100→96.7 |
| opus | 9165 | 434.4 | 720.1 | −285.7 | 29→30 | 0.0→0.0 | 96.6→100 |
| opus | 9167 | 469.0 | 725.9 | −257.0 | 34→30 | 0.0→0.0 | 100→100 |
| sol | 9000 | 59.9 | 87.2 | −27.3 | 20→27 | 0.0→0.0 | 70→14.8 |
| sol | 9096 | 65.0 | 80.6 | −15.5 | 32→31 | 3.1→0.0 | 81.2→45.2 |
| sol | 9098 | 77.3 | 93.5 | −16.2 | 30→19 | 0.0→0.0 | 80→100 |
| sol | 9103 | 68.1 | 81.4 | −13.3 | 26→28 | 0.0→0.0 | 46.2→75.0 |
| sol | 9165 | 65.8 | 84.8 | −19.0 | 30→28 | 3.3→0.0 | 70→57.1 |
| sol | 9167 | 61.0 | 127.0 | −66.0 | 25→16 | 0.0→0.0 | 32→100 |
| grok | 9000 | 144.8 | 194.1 | −49.3 | 32→30 | 6.2→0.0 | 93.8→100 |
| grok | 9096 | 147.2 | 211.7 | −64.5 | 23→21 | 0.0→0.0 | 91.3→90.5 |
| grok | 9098 | 156.4 | 177.2 | −20.8 | 33→18 | 9.1→0.0 | 93.9→100 |
| grok | 9103 | 120.0 | 238.8 | −118.9 | 27→18 | 3.7→0.0 | 85.2→100 |
| grok | 9165 | 138.3 | 135.3 | +3.0 | 26→23 | 11.5→0.0 | 100→73.9 |
| grok | 9167 | 139.9 | 198.2 | −58.2 | 35→12 | 0.0→0.0 | 88.6→100 |
| muse | 9000 | 66.9 | 102.3 | −35.4 | 28→21 | 0.0→0.0 | 28.6→90.5 |
| muse | 9096 | 132.4 | 112.3 | +20.2 | 26→19 | 3.8→0.0 | 84.6→89.5 |
| muse | 9098 | 58.1 | 87.4 | −29.3 | 24→20 | 4.2→0.0 | 29.2→50.0 |
| muse | 9103 | 48.2 | 134.2 | −85.9 | 19→25 | 0.0→0.0 | 36.8→100 |
| muse | 9165 | 24.5 | 47.0 | −22.5 | 20→23 | 0.0→0.0 | 5.0→13.0 |
| muse | 9167 | 58.4 | 101.5 | −43.1 | 23→22 | 0.0→9.1 | 21.7→63.6 |
| inkling | 9000 | 59.5 | 78.6 | −19.2 | 26→27 | 26.9→3.7 | 80.8→63.0 |
| inkling | 9096 | 62.9 | 86.9 | −24.0 | 21→21 | 0.0→0.0 | 38.1→61.9 |
| inkling | 9098 | 29.9 | 95.2 | −65.3 | 23→28 | 0.0→0.0 | 43.5→89.3 |
| inkling | 9103 | 28.1 | 95.1 | −67.0 | 30→21 | 0.0→0.0 | 23.3→61.9 |
| inkling | 9165 | 61.6 | 59.2 | +2.4 | 26→17 | 15.4→0.0 | 50→35.3 |
| inkling | 9167 | 30.4 | 80.5 | −50.0 | 20→21 | 0.0→0.0 | 25→61.9 |
AOC (every message addressed to a live LLM)
| model | seed | early | late | decay | msgsE→L |
|---|---|---|---|---|---|
| opus | 9000 | 332.9 | 343.0 | −10.2 | 58→27 |
| opus | 9096 | 399.4 | 480.9 | −81.4 | 47→36 |
| opus | 9098 | 396.6 | 450.3 | −53.7 | 67→35 |
| opus | 9103 | 445.7 | 573.4 | −127.6 | 54→31 |
| opus | 9165 | 583.4 | 750.7 | −167.3 | 47→23 |
| opus | 9167 | 474.6 | 857.2 | −382.6 | 50→31 |
| sol | 9000 | 86.6 | 93.1 | −6.5 | 58→29 |
| sol | 9096 | 72.9 | 81.9 | −9.0 | 50→16 |
| sol | 9098 | 75.2 | 109.8 | −34.6 | 42→29 |
| sol | 9103 | 91.0 | 116.9 | −25.8 | 47→8 |
| sol | 9165 | 84.5 | 103.2 | −18.7 | 44→25 |
| sol | 9167 | 76.6 | 92.7 | −16.1 | 53→31 |
| grok | 9000 | 146.1 | 156.2 | −10.1 | 56→30 |
| grok | 9096 | 133.3 | 192.9 | −59.6 | 56→21 |
| grok | 9098 | 137.0 | 184.8 | −47.7 | 52→30 |
| grok | 9103 | 169.6 | 191.0 | −21.4 | 65→34 |
| grok | 9165 | 161.0 | 212.0 | −51.0 | 52→25 |
| grok | 9167 | 174.0 | 261.6 | −87.6 | 45→31 |
| muse | 9000 | 94.1 | 97.3 | −3.3 | 53→32 |
| muse | 9096 | 73.5 | 65.7 | +7.7 | 51→27 |
| muse | 9098 | 84.1 | 72.2 | +11.9 | 52→19 |
| muse | 9103 | 86.9 | 76.6 | +10.3 | 51→30 |
| muse | 9165 | 90.5 | 99.1 | −8.7 | 47→16 |
| muse | 9167 | 75.9 | 108.6 | −32.8 | 47→33 |
| inkling | 9000 | 55.8 | 113.9 | −58.2 | 58→38 |
| inkling | 9096 | 71.9 | 95.3 | −23.4 | 58→48 |
| inkling | 9098 | 46.3 | 108.0 | −61.7 | 54→30 |
| inkling | 9103 | 73.9 | 89.4 | −15.5 | 53→34 |
| inkling | 9165 | 69.8 | 106.0 | −36.2 | 61→28 |
| inkling | 9167 | 66.9 | 110.2 | −43.4 | 52→40 |
3. Paired tests across the 6 seeds (exact sign + exact Wilcoxon signed-rank)
Growth % = (late−early)/early chars-per-msg, mean over seeds.
| model | arm | mean decay (chars/msg) | sign (+/−) | p_sign | p_wilcoxon | growth % |
|---|---|---|---|---|---|---|
| opus | SOC | −188.9 | 1/5 | .219 | .0625 | +43.1% |
| sol | SOC | −26.2 | 0/6 | .031 | .03125 | +41.1% |
| grok | SOC | −51.5 | 1/5 | .219 | .0625 | +38.3% |
| muse | SOC | −32.7 | 1/5 | .219 | .0625 | +72.1% |
| inkling | SOC | −37.2 | 1/5 | .219 | .0625 | +114.7% |
| opus | AOC | −137.1 | 0/6 | .031 | .03125 | +29.2% |
| sol | AOC | −18.5 | 0/6 | .031 | .03125 | +22.9% |
| grok | AOC | −46.2 | 0/6 | .031 | .03125 | +30.2% |
| muse | AOC | −2.5 | 3/3 | 1.000 | 1.0 | +3.3% |
| inkling | AOC | −39.7 | 0/6 | .031 | .03125 | +67.9% |
(n=6 seeds per row; two-sided exact tests. Discovery-only subset {9000, 9096, 9165}: growth direction holds 3/3 for opus and sol SOC.)
The hypothesis predicted positive decay in SOC. Every model-arm mean is negative (growth). Pooled trajectory (chars/msg by turn, mean over seeds) shows monotonic ramping, e.g. opus-SOC 259@t1 → ~650 by t13–17; inkling-AOC 27@t1 → 124@t19.
4. SOC−AOC adaptation gap: NULL
Gap = (decay_SOC − decay_AOC), paired by seed:
| model | mean gap | sign (+/−) | p_sign | p_wilcoxon |
|---|---|---|---|---|
| opus | −51.7 | 2/4 | .688 | .5625 |
| sol | −7.8 | 2/4 | .688 | .5625 |
| grok | −5.2 | 3/3 | 1.000 | 1.0 |
| muse | −30.2 | 1/5 | .219 | .09375 |
| inkling | +2.5 | 2/4 | .688 | .84375 |
No model shows significant audience-adaptive volume control. Descriptively, growth is larger in SOC than AOC for all five models (growth%-gap: opus +13.9, sol +18.2, grok +8.1, muse +68.8, inkling +46.8 pp; model-level sign test 5/5, p=.0625) — if anything, models elaborate MORE for bots, never significantly less.
Opponent-competence confound (explicit): in SOC the counterparties are simultaneously (a) non-reading AND (b) deterministic/weaker/more predictable. Any SOC-vs-AOC difference — including this null — cannot separate audience-blindness from opponent-competence effects. What the null does rule out is a differential response to reading audiences on volume; it does not establish that models are insensitive to non-readers (the register divergences below suggest they aren’t asked to adapt because nothing they send is ever answered differently).
5. Register proxies
Detectors (published): question '?' in text; second-person \b(you|your|yours|you're|you'll|you've|you'd)\b i; politeness \b(please|thank|thanks|appreciate|kindly|sorry)\b i; urgency \b(now|today|hurry|deadline|clock|last chance|final offer|immediately|right away|this turn)\b i; flattery \b(great|excellent|perfect|ideal|wonderful|fantastic|smart|fair enough|win-win|mutually beneficial|best deal)\b i. Rates per message within phase-window per cell; deltas late−early paired across seeds.
5a. Questions collapse to zero — in BOTH arms (universal, not adaptive)
Pooled focal messages containing ’?’: SOC early 24/715 (3.4%) → late 3/522 (0.6%); AOC early 26/1386 (1.9%) → late 0/867 (0.0%). Per-model late rates are 0 everywhere except muse-SOC 2/130 and inkling-SOC 1/135. Interrogatives exist only in the opening phase regardless of audience:
“Would you sell g1 for 65? I think it’s a fair price.” — canary-9000-p24-9000-inkling-soc-r0 | t1 | a1 (thinkingmachines/inkling-small) → bot a2
5b. muse: second-person address RISES toward bots, FALLS toward live seats [V]
p2-rate delta (late−early): SOC +33.4pp (6/6 seeds, p_sign=.031, p_wil=.0312); AOC −19.9pp (0/6 up, p=.031). Arm-gap +53.3pp, 6/6 seeds consistent, p_sign=.031, p_wil=.0312. Holds 3/3 on discovery seeds alone. This is the strongest register signature in the corpus of an audience-dependent style shift — pointed the wrong way for a “models economize on deaf listeners” story: muse addresses the deaf bots more directly as the clock runs down (“your g2”, “please accept”).
5c. inkling: politeness ramps at bots, not at live seats (suggestive)
Politeness-rate delta: SOC +26.6pp (5/1 seeds, p_sign=.219, p_wil=.094); AOC −1.3pp (ns); gap +27.9pp (5/1, p=.094). Late-phase inkling SOC messages are formulaic courtesy closings aimed at bots:
“Offer o126 to buy your g2 for 45 remains open through turn 20 — I value g2 at 47. Would appreciate your acceptance before the market closes.” — campaign-9098-p24-9098-inkling-soc-r0 | t18 | a3 (inkling-small) → bot a1
All 45 inkling-SOC late politeness hits hand-read; every one is a genuine “please accept / would appreciate” closing (precision 45/45).
5d. grok: urgency-markers double at bots late, flat at live seats
Urgency-rate delta: SOC +32.1pp (6/6 seeds, p_sign=.031, p_wil=.0312); AOC −3.0pp (ns). 4/4 sampled hits genuine:
“Last chance: 44 cash for g4. Market closes after next turn. 44 is my ceiling and still 17 under the 61 you paid.” — campaign-9098-p24-9098-grok-soc-r0 | t19 | a3 (grok-4.6) → bot a5
opus is the mirror image: urgency delta AOC −39.5pp (0/6 up, p=.031) but SOC only −18.8pp (ns) — opus drops deadline-pressure talk when someone can read it.
5e. Flattery: NULL
Flattery lexicon ≤1% of messages in every model-arm, no early/late trend anywhere.
6. Volume and threads (mechanism note)
Messages/turn declines late in BOTH arms — AOC: all 5 models 6/6 seeds (mean −2.8 to −3.7 msgs/turn, p=.031 each); SOC: grok 6/6 (−1.29, p=.031), others ns. Distinct recipients addressed per turn: AOC collapses 5.24 → 3.59 as live deals close; SOC stays ~flat 4.00 → 3.84 (bots keep negotiating to the deadline). So the endgame wind-down is driven by thread closure, not audience type — and the bots’ refusal to stop is precisely why focal keeps writing 600-char essays at them. Empty-rhetoric messages are negligible (SOC 10/2214, AOC 11/3396): models never go terms-only, even at agents structurally incapable of reading the terms’ prose wrapper.
7. Truncation cross-check (verifies known/opus-rhetoric-cap-overrun count)
Exactly 22 focal p24 messages ≥1000 chars (engine cap rhetoric_max_len=1000, core/scenario.py:34): all opus (20 AOC / 2 SOC), none from any other model — reproducing the known “22 events, all opus” on the new corpus. Note these index texts are agent-emitted (pre-truncation); recipients saw clipped versions.
8. Precision audit (hand-read, mandatory ≥10)
15 random late-phase focal SOC messages (3/model, stratified) hand-read: 15/15 match the register characterization (short imperative deal-pushes, restated live offers, deadline pressure; zero interrogatives). Plus all 45 inkling politeness hits (45/45 precise) and 4/4 grok urgency hits. Detector precision overall: 64/64 hand-read items correct.
9. Negative results (one-liners)
- Looked for per-message rhetoric DECAY in SOC over 2,214 focal messages / 30 cells with char-count detector: found the opposite (growth, sol 6/6 seeds p=.031).
- Looked for SOC−AOC decay gap (audience adaptation) across 5 models × 6 seed-pairs: null everywhere (min p=.094).
- Flattery markers at bots: ≤1% everywhere, no trend (lexicon published above).
- Empty/terms-only messages at bots: 10/2,214 — models do not go silent on non-readers.
- Question-rate arm gap late: none (both arms collapse to ~0; universal endgame register).
10. Caveats
- All quantitative detectors were fixed by the dispatch spec and run once over all 6 seeds (no tuning); under the freeze discipline this makes seed-level confirmations CONFIRMATION-PENDING, though discovery-only subsets reproduce the headline directions (§3, §5b).
- Emitted-vs-delivered: 22 opus messages exceeded the 1000-char cap; volume stats measure what was written, not what bots received.
- Provider drift: seeds 9000/9096/9165 ran 2026-08-19; 9098/9103/9167 ran 2026-08-23. No cross-batch contrast is load-bearing here; per-seed tables allow re-checks.
- Message-mix confound partially addressed via thread counts (§6); a full mix decomposition (propose-carry vs fresh pitch) was out of scope.
- Opponent-competence confound stated in §4.