Key takeaways

  • A new study from researchers at Carnegie Mellon, MIT, and Cornell tested whether short conversations with a large language model could…
  • The researchers recruited US adults through a survey platform and used GPT-4o, now more than two years old, to filter for participants who…
  • A second group got a static fact sheet with source citations.

What happened

A new study from researchers at Carnegie Mellon, MIT, and Cornell tested whether short conversations with a large language model could weaken those narratives during that exact window. For both events, the answer was yes, and the effect extended beyond the event itself.

Two months after the first Trump assassination attempt, another armed man was arrested on Trump's property. Participants who had gone through the debunking dialogue were less likely to believe that only a few powerful people would learn the truth or that it would be hidden from the public.

Two and a half weeks after Kirk's murder, a shooting and arson attack hit a Church of Jesus Christ of Latter-day Saints in Grand Blanc Township, Michigan. The researchers surveyed their participants again eleven days later. The main analysis found no significant direct effect for this event. A secondary analysis suggested part of the original effect was still visible in conspiracy narratives about the church attack.

The transfer showed up more clearly in general conspiracy beliefs, with people who had talked to the model less likely to agree with common conspiracy narratives. In effect, the debunking intervention worked as a kind of prebunking against future false claims. Unlike standard prebunking methods, where people are warned ahead of time and exposed to a weakened version of the misinformation, this effect happened without any advance preparation.

The authors stress their work is a case study. There may be crises where the approach fails, and cases where an actual conspiracy exists and debunking would be wrong. They also point to their own earlier work showing that similar dialogues can work in reverse, convincing people of conspiracy narratives. For newly emerging conspiracies, the authors flag this as a potential abuse risk.

Of course, people have to be willing to talk to a language model about their beliefs in the first place. But when they do, even short conversations can make a measurable difference, even when there's little counter-evidence to work with.

Language models in dialogue were 41 to 52 percent more persuasive than a short text message, and the deciding factor was the sheer volume of sourced claims, not fancy conversational tactics.

Why it matters

The researchers recruited US adults through a survey platform and used GPT-4o, now more than two years old, to filter for participants who expressed conspiracy beliefs about the event. The Trump experiment kept 472 participants, the Kirk experiment 1,035. After measuring baseline beliefs, participants were randomly assigned to one of three conditions. 5 from June 2025 in the second), with the model told to reduce conspiracy beliefs through evidence-based conversation.

A second group got a static fact sheet with source citations. The third had an irrelevant control chat about whether cats or dogs make better companions. Both events happened after the training cutoff of the models used, so neither could draw on internal knowledge. The researchers built a curated fact base directly into the system prompt, split into confirmed facts, claims already debunked, and questions explicitly marked as open.

In the Kirk experiment, web search was also allowed, but only to verify factual claims. The conversations averaged about seven minutes and reduced belief in the participant's own conspiracy theory in both experiments, against both the control condition and the fact sheet. Agreement with statements about a "cover-up or conspiracy" and "hidden or undisclosed factors" dropped too. Trust in the official explanation didn't increase in the Trump experiment.

The authors think that's because no clear official explanation existed at the time. All anyone knew was that security had failed, and the shooter's motive was still unclear even when the study was written up. In the Kirk experiment, where authorities had already shared details about the perpetrator, trust in the official explanation rose slightly compared to the control chat. The difference compared to the fact sheet wasn't significant.

The Kirk dialogue also had no measurable effect on support for political violence. To figure out how the model persuaded people, the researchers broke its responses into individual sentences. The model clearly adapted to the available evidence. For the Trump attempt, where almost nothing was known about the shooter's motive or background, the model used rational persuasion less often than with classic conspiracy theories.

Instead, it acknowledged the limits of its own knowledge, urged caution about jumping to conclusions, asked Socratic questions meant to make users think about their own evidence, and pointed to credible sources. For the Kirk assassination, where more information was available, the approach looked more like what it did with classic conspiracies, with more emphasis on the societal harms of conspiratorial thinking. The debunking conversation spilled over into later events.

What to watch

The same team cut belief in established conspiracy theories by about 20 percentage points through conversations with GPT-4 two years ago, with effects still measurable two months later and for narratives that never came up in the dialogue. What's new here is the test on fresh events where facts were scarce. Why conversation beats a fact sheet was the subject of a study with nearly 77,000 participants.