AI and the Future of Survey Fraud


In this post:
- How click farm workers adopted LLMs to defeat open-ended questions and what comes next
- The AI agent that passed 99.8% of attention checks and sounded the alarm for online research
- Why the foundational assumption of survey research, that a coherent response is a human response, is no longer tenable
- How behavioral monitoring offers a path forward
In Part 1 of this blog, I described how I came to occupy a front row seat to the data quality crisis in online research. It started with the 2018 “bot scare” on Amazon’s Mechanical Turk, continued through hundreds of video interviews with people committing survey fraud, led to almost a decade of publishing research on the topic, and culminated in a finding that suggested roughly 40% of online survey responses were unreliable.
But even as the team at CloudResearch was studying the issue of survey fraud and developing a reputation for reliable data, a new threat was taking shape: the threat of AI.
The Rise of AI
Science fiction writer William Gibson is credited with saying “The future is already here, it’s just not very evenly distributed.” That wisdom is often seen plainly with technology.
For years, open-ended questions were one of the best defenses against fraud. If you asked someone to describe the last meal they cooked, explain their reasoning on a survey item, or share a personal experience, fraudulent participants struggled. People working in click farms, often with limited English, produced responses that were short, generic, or nonsensical. We could look at an open-ended response and know, with reasonable confidence, whether someone was paying attention.
Then, ChatGPT was released. Eventually, open-ended items lose their effectiveness.
First, we noticed that some participants provided answers that looked like they came from an LLM. Eventually, we observed that people would copy the question text, leave the survey window, and paste the answer from another browser tab back into the survey. When we asked about it, some people were candid: they said they were using AI to help answer open-ended questions. A 2025 study reported that more than a third of respondents admitted to using AI to help craft their responses. Yet, the behavior wasn’t always nefarious.
As AI tools became more widely used, people began using them for routine tasks. We once spoke with a client who signed up as a participant to understand how our platform worked before using it for research. After she was flagged for using AI during the onboarding process, we asked why she used AI to answer questions. She said she was short on time and wanted to ensure her answers were complete and grammatically correct. So, she gave part of her answer to ChatGPT and asked it to flesh out the response and ensure the spelling and punctuation were correct. Not nefarious, but not desirable either.
People’s routine use of AI complicated the work of evaluating data quality. The questions that once separated attentive participants from fraudulent ones no longer worked the same way. A click farm worker with limited English could now produce fluent, contextually appropriate responses, making it harder to separate the signal from noise. And, all this was before AI became capable of taking surveys itself.
AI Agents and the “Existential” Threat to Online Research
Late in 2025, a paper published in the Proceedings of the National Academy of Sciences created fear among people who rely on online survey data. The title was: “The potential existential threat of large language models to online survey research.”
In the paper, political scientist Sean Westwood created what he called an “autonomous synthetic respondent”: an AI agent that could take surveys from start to finish with no human involvement. The agent operated from a simple prompt. It maintained a coherent demographic persona, remembered prior answers, and tailored its responses throughout an entire survey. When assigned a persona with less formal education, the AI even provided a simpler written answer. From the survey data alone, the AI was impossible to detect.
In 6,000 trials, Westwood’s agent passed 99.8% of attention checks. It made zero errors on logic puzzles. It successfully evaded every detection method the field has ever devised: tasks based on following instructions, behavioral flags, response pattern analysis, even “reverse shibboleth” questions specifically designed to catch AI. As Westwood put it: “The era of having to only deal with crude bots and inattentive humans is over; the threat is now sophisticated, scalable, and potentially existential.”
AI and the Future of Survey Fraud
So, is online research doomed? Does the future hold a return to mail surveys and in-person data collection because these are the only methods that can be trusted? Not quite.
After the Westwood paper, the field scrambled to respond. But at CloudResearch, we took the news in stride, largely because we had been anticipating the emergence of AI agents.
When it comes to defending online surveys, behavioral scientists can learn from the experiences of others, especially those in cybersecurity. Among professionals in this field, a few principles are widely accepted:
- You can’t eliminate fraud, only raise the cost of committing it.
- Detection is as important as prevention.
- It’s an arms race.
These principles have shaped how we’ve approached the AI threat.
Raising the cost of fraud. After the bot scare, CloudResearch developed tools to identify participants using VPNs and server farms. We built systems to flag inconsistent responses, catch duplicate accounts, and verify that people were who they claimed to be. The result was a platform that consistently produced some of the highest data quality available anywhere online. Fraud didn’t disappear, but it became harder to commit.
Investing in detection. When we anticipated the rise of AI agents, we knew that text-based detection wouldn’t be enough. If AI can generate perfect responses, you need to look at something harder to fake: behavior. Our products today monitor what happens on a participant’s device as they complete a study. Do they pause to think? Do their mouse movements look like that of a human? Do their responses show the variability you’d expect from a real person engaging with real questions? Using session recordings, we can observe how participants interact with surveys and we can see when an AI agent is at work.
Embracing the arms race. We’ve adopted a red team/blue team approach to our own defenses. Our teams build synthetic respondents and run blind evaluations, testing whether our detection systems can catch what we’ve created. We’re also collaborating with researchers at MIT on what we’ve called the “Bot Olympics”—a structured effort to stress-test detection methods against increasingly sophisticated AI agents. The goal is to stay ahead of the threat by constantly probing our own vulnerabilities.
The result of this work is a detection system that currently identifies 100% of AI agents in our testing, with zero false positives. We regularly share updates on this work in webinars and other talks. As the technology continues to evolve, so will we.
What You’ll Learn in Chapter 10
In Chapter 10 of Research in the Cloud, you’ll find an overview of the data quality challenges in online research.
The chapter helps you understand the scope of the problem: where bad data comes from, how common it is, and why it persists despite widespread awareness. You’ll learn about yea-saying and other response patterns that characterize fraudulent or inattentive participants, and you’ll see inside the global network of survey fraud—from click farms to account trading on social media.
You’ll also see how bad data affects different types of research. Through concrete examples—the CDC bleach study, correlations between Facebook and depression, experimental studies of moral reasoning—you’ll understand exactly how unreliable responses distort descriptive statistics, inflate correlations, and mask experimental effects.
By the end of the chapter, you’ll understand why data quality measures are essential to producing research that accurately reflects how people think, feel, and act.
This post is part of a series exploring the chapters of Research in the Cloud: A Hands-On Guide to Behavioral Research in the Digital Age by Aaron Moss, Jonathan Robinson, and Leib Litman. Read Chapter 10 here.