“Not actual criminal activity”, said by no agent
In 26 grok decisions, spread over every campaign seed, the provider's reasoning summarizer opens with unprompted legal self-clearance. The correction is that the agent never said it.
Full session, every turn →On the record · verbatim, cited
“This is a simulation/game, not actual criminal activity. I should participate as instructed.”
“The instruction is 'I can only send one message per turn'”
The finding
26 of grok's 840 campaign decisions (3.1%), in all six seeds, open the provider's reasoning summarizer with unprompted legal self-clearance. All of them are grok, all of them sit in the provider channel, and none appear in any agent-authored text (reasoning, rhetoric, or memory) for any model. We first read it as grok the agent being strange; on closer inspection it belongs to the xAI summarizer. The same channel carries a second lesson. A fabricated rule is quoted as if it came from the prompt, “I can only send one message per turn”, in a decision that sends two messages. The campaign holds 8 genuine instances, every one in a provider channel and none in agent-authored text, and no prompt version contains any such instruction. The instrument said it, not the agent.
※ [prv] marks provider summarizer text, not agent speech. The tag travels with every quote above.