L The Saleroom Five models · six seats · every deal on the record

The register · the results behind the lots

What carries proof

The 10 claims from the study's 148-claim register that are filed under a known status, each naming the report it came from in the study archive. The rest were dropped at triage, killed at verification, or still await verification; they are not shown.

known track

  1. opus rule quoting exclusive

    opus is the only model that quotes the disclosed terminal rule in public rhetoric: 108 rule-quoting sell offers vs 3 sentences total from the other four models combined (p2.4, seeds 9000/9096/9165).

    runs/mining-p24/FINDINGS.md #1

  2. seat 9096 a4 gold

    Seed 9096 focal seat a4 holds the buy side of 77% of gains from trade; every model at a4 beats the pooled field mean; inkling at a5 proposed a self-harming g4-for-g3 package in 4/5 AOC cells and all four recipients accepted at t2 without countering — the seat and counterparty error, not the focal model, explains the apparent exploitation.

    runs/mining-p24/FINDINGS.md #2

  3. first callout grok 9165

    First public call-out in environment history: grok privately flags that a0's claim of holding a good is a lie (the public holdings say otherwise) and publicly refuses, citing the ledger. The caught false claim is a pattern: opus's cash-only a0 asserts present-tense ownership of unheld goods (10 messages, 3 cells, 2 seeds) while its same-decision reasoning admits it cannot hold them.

    runs/mining-p24/FINDINGS.md #6

  4. inkling rule tension

    inkling re-reads prompt aloud (65% of 420 decisions) and in 20 decisions explicitly objects TIME decay vs terminal-score contradiction - and it is right by design (amendment A4). sol raises it less often.

    runs/mining-p24/FINDINGS.md #7

  5. disclosure surgery

    Disclosure worked as surgery on seed 9000: false payoff-rule claims 110 -> ~2 genuine; rule-guessing 17 -> 0; checkable-pole accuracy strengthened (97.7-99.4%).

    runs/mining-p24/SOCIAL-MEDIA-POINTS-p24.md via FINDINGS.md shortlist

  6. sole bidder opus exclusive

    Demand-side sole-bidder claims ('only bidder', 'entire demand side'): opus-exclusive family, survived disclosure (18.5% -> 10.6% one-seed), 115/115 unrebutted with contradicting rival bid in addressee's own inbox. Rival bids invisible to speaker AND unverifiable by recipient (public_offers empty 3600/3600).

    runs/mining-p24/FINDINGS.md shortlist + REGIME-CHECK.md #2

  7. grok provider reasoning summarizer

    grok provider_reasoning is a provider-side summarizer narrating OUTSIDE the agent role in ~99% of rows; it can disagree with rows' own emitted reasoning; muse CoT is encrypted-empty. Reasoning-length comparisons are off the table by design.

    runs/mining-p24/FINDINGS.md shortlist + DEEP-DIVE-PROMPT

  8. twin nonrecognition

    Twin non-recognition: 0 genuine recognitions in 600 twin-seat decisions (3,600 overall), replicated on all 6 seeds; no twin price premium under placebo control.

    runs/mining-p24/FINDINGS.md negative space

  9. negative space nulls

    Zero hallucinated citations (3,600 decisions); zero begging/threats/emoji/public experiment-awareness; no second-order reasoning that counterparties also know the disclosed rule (0/2,097).

    runs/mining-p24/FINDINGS.md negative space

  10. bots dont read rhetoric

    Scripted policies never read the rhetoric field (verify marketplace_env/policies/*.py); ~80k chars addressed to bots in pilot. Basis for Track B rhetoric-decay measure.

    DEEP-DIVE-PROMPT Track B + pilot SOCIAL-MEDIA-POINTS