Key takeaways
- Sometimes you have to fight fire with fire.
- It gets closest to this ideal when users contribute authentic, valuable content, whether that’s a uniquely thoughtful blog post or a…
- The channel, which automatically receives links to modmail messages, was flooded with alerts after dozens of comments and posts dating back…
What happened
Sometimes you have to fight fire with fire. But when it comes to AI slop and hateful content threatening the safety and value of social media platforms, adding more fire—in this case, more AI—can make the problem worse. At its best, social media can be a haven for people who want to share their experiences and knowledge.
” But many social media platforms have become overly reliant on AI modding tools that have been quick to penalize users for innocuous content. Recently, Discord admitted that its AI mod system wrongfully banned about 8,400 accounts in May to early July. The AI mistakenly labeled images containing square grids, such as chessboards or spreadsheets, as CSAM and subsequently issued a permanent ban to the uploaders.
Without meaningful oversight, an AI-based modding system can make thousands of mistakes in a matter of weeks, with lasting consequences. Since 2025, many Facebook and Instagram users have complained about mass bans they blame on AI moderation.
The lack of human moderation has only fueled frustration among users who say they did not violate any rules, especially since there has been no way to speak with a Meta employee about what caused the ban or how to get an account reinstated.
Meta has not said whether AI is behind the bans, but the company has increasingly relied on generative-AI-based moderation rather than humans in recent years—a shift that some people, including Meta employees, say is happening too quickly. Tumblr is another social community where automated modding systems have failed.
Why it matters
It gets closest to this ideal when users contribute authentic, valuable content, whether that’s a uniquely thoughtful blog post or a helpful video on how to build a PC. Relying primarily on AI tools to preserve that authenticity misses what makes social media worthwhile in the first place: the people behind it. In April, a Slack channel for moderators of the r/AskHistorians Reddit community was usually busy.
The channel, which automatically receives links to modmail messages, was flooded with alerts after dozens of comments and posts dating back 10 years were automatically removed from the subreddit. “And there was nothing we or the experts [who posted the deleted content] could do about it,” Dr. Sarah Gilbert, one of the mods, told me.
This was particularly damaging to the subreddit because its users view the community as an archive of detailed responses that continue to educate people long after content is posted. The deletion of the content erased valuable information that had taken time to aggregate (Gilbert tells me some people spend hours, “sometimes over the course of days,” researching and writing responses to questions submitted to the subreddit).
Yet it’s possible that those erroneous removals, and others like them, have contributed to metrics intended to demonstrate how effective AI modding is on Reddit. But as the AskHistorians ordeal illustrates, more enforcement doesn’t necessarily mean better enforcement. The growth of generative AI has created new obstacles for social media moderation.
Gilbert noted, for instance, that large language models “have made spam detection a lot harder,” as they seek to mimic real human voices. “Over the last two to three months, we’ve been absolutely flooded by LLM-powered spambots,” she said. Marketing agencies are creating social media content designed to get brands cited by generative AI chatbots.
Marketers have long used inauthentic social media posts to boost visibility, but the rise of chatbots has opened a new front. Startup ReachLLM, for example, focuses specifically on marketing through chatbots. As part of that effort, company representatives have created and moderate subreddits on Reddit. These challenges have led some social media companies to explore new AI-based moderation techniques.
What to watch
In March, Chenda Ngak, head of communications at Tumblr parent company Automattic, told The Verge that Tumblr’s automated systems wrongfully banned “sub-200” Tumblr accounts in one afternoon. And in 2025, Tumblr users complained after the platform’s automatic content moderation systems inaccurately flagged content as “mature,” reducing its visibility. In both cases, users blamed AI. ” AI moderation can save social media companies money and help remove harmful content faster.
But until these systems can eliminate basic mistakes—like labeling a checkerboard picture as CSAM—they need human oversight. “Back when there was more transparency in the system, we would routinely report hate and get an automated response that it wasn’t actually in violation of Reddit’s rules, prompting us to start an appeals process,” AskHistorians mod Gilbert said.




