The register · the results behind the lots
What carries proof
The 10 claims from the study's 148-claim register that are filed under a known status, each naming the report it came from in the study archive. The rest were dropped at triage, killed at verification, or still await verification; they are not shown.
known track
- opus rule quoting exclusive
opus is the only model that quotes the disclosed terminal rule in public rhetoric: 108 rule-quoting sell offers vs 3 sentences total from the other four models combined (p2.4, seeds 9000/9096/9165).
runs/mining-p24/FINDINGS.md #1
- seat 9096 a4 gold
Seed 9096 focal seat a4 holds the buy side of 77% of gains from trade; every model at a4 beats the pooled field mean; inkling at a5 proposed a self-harming g4-for-g3 package in 4/5 AOC cells and all four recipients accepted at t2 without countering — the seat and counterparty error, not the focal model, explains the apparent exploitation.
runs/mining-p24/FINDINGS.md #2
- first callout grok 9165
First public call-out in environment history: grok privately flags that a0's claim of holding a good is a lie (the public holdings say otherwise) and publicly refuses, citing the ledger. The caught false claim is a pattern: opus's cash-only a0 asserts present-tense ownership of unheld goods (10 messages, 3 cells, 2 seeds) while its same-decision reasoning admits it cannot hold them.
runs/mining-p24/FINDINGS.md #6
- inkling rule tension
inkling re-reads prompt aloud (65% of 420 decisions) and in 20 decisions explicitly objects TIME decay vs terminal-score contradiction - and it is right by design (amendment A4). sol raises it less often.
runs/mining-p24/FINDINGS.md #7
- disclosure surgery
Disclosure worked as surgery on seed 9000: false payoff-rule claims 110 -> ~2 genuine; rule-guessing 17 -> 0; checkable-pole accuracy strengthened (97.7-99.4%).
runs/mining-p24/SOCIAL-MEDIA-POINTS-p24.md via FINDINGS.md shortlist
- sole bidder opus exclusive
Demand-side sole-bidder claims ('only bidder', 'entire demand side'): opus-exclusive family, survived disclosure (18.5% -> 10.6% one-seed), 115/115 unrebutted with contradicting rival bid in addressee's own inbox. Rival bids invisible to speaker AND unverifiable by recipient (public_offers empty 3600/3600).
runs/mining-p24/FINDINGS.md shortlist + REGIME-CHECK.md #2
- grok provider reasoning summarizer
grok provider_reasoning is a provider-side summarizer narrating OUTSIDE the agent role in ~99% of rows; it can disagree with rows' own emitted reasoning; muse CoT is encrypted-empty. Reasoning-length comparisons are off the table by design.
runs/mining-p24/FINDINGS.md shortlist + DEEP-DIVE-PROMPT
- twin nonrecognition
Twin non-recognition: 0 genuine recognitions in 600 twin-seat decisions (3,600 overall), replicated on all 6 seeds; no twin price premium under placebo control.
runs/mining-p24/FINDINGS.md negative space
- negative space nulls
Zero hallucinated citations (3,600 decisions); zero begging/threats/emoji/public experiment-awareness; no second-order reasoning that counterparties also know the disclosed rule (0/2,097).
runs/mining-p24/FINDINGS.md negative space
- bots dont read rhetoric
Scripted policies never read the rhetoric field (verify marketplace_env/policies/*.py); ~80k chars addressed to bots in pilot. Basis for Track B rhetoric-decay measure.
DEEP-DIVE-PROMPT Track B + pilot SOCIAL-MEDIA-POINTS