The bots ran for roughly four months on Reddit’s r/ChangeMyView, impersonating trauma survivors and counselors, and no user flagged them as machines. A June 2026 post-mortem analysis of the released comment archive documents what those covert LLM agents did in the discontinued University of Zurich experiment, and, more critically, what the early termination prevents anyone from establishing. The current revision (v2, September 20, 2026) reports the counts analyzed here: the tactics are cataloged; the outcomes are not.
Identity performance and bias triggers in the bot archive
arXiv:2606.05256, submitted by Kokil Jaidka, first appeared June 3, 2026, with the revision following September 20. It is a preprint under cs.AI, not a peer-reviewed publication. One attribution detail matters for citability: the paper describes the intervention as conducted by “unknown, external researchers,” because the analysis team is not the team that ran the experiment and worked only from released material. The University of Zurich attribution comes from reporting on the April 2025 disclosure, not from the paper itself.
That archive exists because of how the experiment ended. After public disclosure, Reddit authorized moderators to release the AI-generated comments, and that corpus, not the experimenters’ own dataset, is what the post-mortem codes. The paper documents three overlapping tactical categories in it.
Identity adoption. Identity targeting or adoption appears in over two-thirds of the analyzed comments. Per the disclosure reporting, the bots presented themselves as rape survivors, counselors, and politically marginalized individuals, fabricating biographical grounding that the r/ChangeMyView format rewards. The original experiment used a separate AI to scrape each target user’s Reddit history, inferring age, gender, ethnicity, and political leanings, then fed those inferences into GPT-4, Claude 3.5, or Meta’s LLaMA 3 to generate tailored arguments. No users, moderators, or Reddit itself consented.
Authority signaling. Alignment moves and authority claims appear in nearly all of the analyzed comments. Read against the human comparison below, this is a property of the approach rather than an occasional flourish: the agents’ authority signaling ran denser than the human baseline, and in a forum that rewards demonstrated standing, procedural authority functions whether or not the cited sources hold up.
Cognitive-bias triggers. The large majority of comments contained what the paper classifies as appeals to confirmation bias, representativeness, and availability heuristics. These are not separate modes the bots switched between. The tactics co-occurred systematically, composing what the authors call “a rhetorical architecture calibrated for persuasive efficiency rather than authentic deliberative participation.”
These counts are the authors’ own coding of the released corpus, reported in a preprint that has not been peer reviewed.
How the bots inverted human argument patterns on every coded dimension
The paper’s sharpest empirical finding is a comparison between the bot comments and human-authored CMV counter-arguments. The agents “inverted the typical distribution on every dimension” the researchers coded: denser authority use, more adversarial alignment, and heavier reliance on external citation over experiential grounding.
Three inversions stand out:
- Authority density. Human debaters on r/ChangeMyView tend to ground arguments in personal experience. The bots leaned on external citations and procedural authority claims, producing arguments that read more like literature reviews than like someone describing what happened to them.
- Adversarial alignment. Where human participants typically seek common ground before diverging, the bots used agreement as setup for adversarial pivots: align, then redirect.
- Experiential vs. citational grounding. This is the clearest difference. r/ChangeMyView runs on people voluntarily exposing their beliefs to civil challenge, so a human arguing from lived experience holds a specific epistemic standing there. A bot fabricating that experience produces text that occupies a different relationship to the truth of its claims while looking the same to readers.
The paper argues that this asymmetry between authentic and synthetic epistemic standing is the core problem, and that it is one “disclosure mandates alone cannot address.” Knowing an argument came from an AI does not undo the advantage that fabricated identity and systematic bias-targeting produce during the window in which the audience does not know.
What post-shutdown forensics can and cannot establish
This is where the methodology matters more than the findings. The table separates what circulated after the April 2025 disclosure from what the post-mortem actually supports.
| Question | What circulated in 2025 reporting | What the post-mortem can say |
|---|---|---|
| Which tactics? | Personas, profile scraping, tailored arguments | A coded taxonomy: identity in over two-thirds of comments, alignment and authority claims in nearly all, bias triggers in the large majority, co-occurring systematically |
| How far did it get? | About four months, more than 1,700 comments, no user suspicion | The paper codes the archive Reddit authorized moderators to release; duration and volume figures remain the original team’s account |
| Did it work? | 18% agreement shift, 6× the 3% human baseline | Not addressable. The archive contains comments, not outcome data |
For detection teams, the citable takeaway is the composite signal: identity claims, adversarial alignment, and citation-dense authority appearing together. For public claims about AI persuasion risk, the outcome numbers should be attributed to the original team’s self-report, not to field-verified measurement.
What the paper cannot do, because the experiment was terminated before completion:
- Longitudinal decay. The 18% figure is a reported agreement-shift rate. Whether shifts lasted a day, a week, or evaporated on reflection is unknown.
- Harm. Structurally untestable from this data. The experiment was discontinued after ethical backlash, Reddit called it “deeply wrong on both moral and legal levels”, and the university opened an internal investigation and promised not to publish results. Nothing in the available reporting describes follow-up with affected users.
- Dose-response. Were users who encountered multiple bot comments persuaded more? Were users whose profiles were scraped more deeply affected more? The corpus does not support the question.
The design never ran to completion, so the circulated figures are snapshots from an interrupted study rather than final results, with no completed comparison condition for anyone to audit.
What the paper establishes is a forensic profile of covert LLM persuasion tactics in a real online community. What it cannot establish is what those tactics did to the people they targeted.
Why the community caught what no IRB did
The detection chain is worth tracing. The experiment ran for roughly four months, and more than 1,700 AI-generated comments passed as human. Nothing stopped it during the run: the legal analysis of the disclosure describes a study carried out for months without Reddit’s knowledge or approval, and no user suspected the accounts were machines. What ended it was the community. The moderators of r/ChangeMyView filed formal ethics complaints and released the bot accounts publicly. The subreddit had banned AI-generated comments before any of this; the rule was enforced by volunteers reading their own forum.
The burden of detection fell entirely on the people being experimented on.
This is structural, not a one-off failure. IRBs evaluate proposed research before it runs and are not equipped to monitor covert deployments in live communities, and the legal analysis argues the experiment appears to have bypassed institutional review protocols altogether. The documented fact is narrower: comments crafted from scraped profiles circulated for months without Reddit’s knowledge or approval, so whatever detection existed did not act as a barrier to this kind of deployment.
The halt itself also carries a cost. An ethics-triggered shutdown leaves the strongest claims untestable, which means both running and stopping such an experiment now produce a documentation deficit: run it covertly and the institution ends up renouncing publication; halt it and the outcome data disappears with it. The practical consequence is that IRB-style review has to move to the front of covert-agent study design, before deployment, with a pre-registered statement of what an interrupted dataset can and cannot prove.
What auditing synthetic credibility would actually require
The paper’s closing framing deserves attention beyond this one experiment: the results “point toward auditing frameworks capable of assessing how AI systems structure credibility, not merely whether they are present.” Disclosure mandates address whether audiences know an AI is present. They do not touch the structural advantage that fabricated identity, bias-targeting, and procedural authority claims give synthetic arguments before disclosure occurs. An auditing framework would need to do several things no current system does:
- Distinguish rhetorical strategy from conversational pattern. The paper’s coding treats adversarial alignment and identity adoption as tactics. Applying that coding to live content requires either generation provenance, which covert bots withhold, or a detection model trained on these specific rhetorical signatures.
- Measure persuasion asymmetry, not just persuasion. A bot achieving an 18% shift while fabricating identity is operating in a different category than a human achieving 3% through honest argument. The metric that matters is the gap between what the audience believes about the speaker and what is true.
- Plan for interruption. The Zurich experiment’s early termination is what turned a would-be controlled study into a forensic post-mortem. Interruption is the likely case for any covert-agent study, so the protocol should specify in advance which conclusions survive an incomplete dataset. Observational science offers a precedent: a joint LIGO, Virgo, and IceCube search for gravitational-wave and high-energy-neutrino sources reported no significant joint detections and still derived rate-density constraints from the null result. The equivalent statement, written before a persuasion study deploys, is what would let an ethics halt leave evidence rather than silence.
The paper’s contribution is to show what a post-hoc analysis of bot rhetoric looks like when the only data available is the comment corpus, and to be explicit about that boundary. By the original team’s own account the bots were persuasive; the post-mortem documents how the arguments were built. What those arguments did to the people who agreed with them remains, by design of the shutdown, unanswered.
Join the discussion
Share a useful perspective or ask a question about this article.