Inside the Mind of a Fraudster: What a Fraud Farm's Own Playbook Reveals About Survey Data Quality


Key Takeaways:
- Fraud escalates in layers, not all at once. Fraudsters start with the cheapest tactic available and only add complexity when simpler approaches fail.
- Proxies now hide inside real residential networks. Apps that pay people to share their home or mobile connection let fraudsters route traffic through ordinary IP addresses anywhere in the world.
- The most advanced setups are also the rarest. Renting a same-country virtual machine defeats location checks, but the cost and complexity mean most fraud farms haven’t adopted it yet.
- Device detection and behavior detection catch different things. Knowing who someone really is and watching what they do inside a survey closes gaps that neither approach catches alone.
Watch the Webinar
Henry LeGardFounder & CEO, VerisoulHow fraudsters escalate, one layer at a time4:20
How we doing? Doing great. Really excited for this. Likewise, you’ve got a killer cast. It’s a really all-star studded group.
Thanks for having me. Super excited about the conversation today. We will be going through actual tactics that fraudsters use to scale survey fraud. I think one of the misconceptions when it comes to fraud is that all fraud is more or less the same. What we’ve found through conversations with survey fraudsters, as well as our own internal broad team, which we have fraudsters on payroll, is it’s very layered.
And that fraudsters have a series of steps that they can and will take over time, but they only need to take those steps if you actually detect the prior tactic that they’ve tried. So with each additional layer that they add, it adds cost, complexity, sophistication, and time. And so fraudsters are reticent to add things that they don’t need. And so what we find is that fraudsters take a kind of iterative approach where they start with one clean session, see what happens. Then they just try a simple cookie clear going into duplicate devices.
Then if you’re catching that, they actually start going to the more sophisticated tactics. How do we actually know this? Two main reasons that I mentioned. One is we literally have fraudsters on payroll in Eastern Europe. These are kind of double agent slash reformed good guys now that are on our team.
They’re informing a lot of our strategy, red teaming our product, red teaming markets around the world. Two is we work with industries all over the internet. So we work with industry… customers across twelve different industries from fintech to payments to gambling to gaming. And third is we actually talk with some of the survey fraudsters and Leib’s going to show some killer videos of what that actually looks like.
But let’s dive into it so step one is actually just trying kind of a clean session just a normal device normal network uh that generally will get detected and so what they or they’ll get through what will happen next is they’ll try super simple cookie clear um using the same device and seeing if you can actually catch that device coming back so the second thing they do is same device same kind of device fingerprint with a cookie clear in between Now, if you have what most people call device fingerprinting, you’re often going to catch that duplicate device. And so the fraudster needs to turn to something more sophisticated. So that’s where they bring out the anti-detect browser. So if we swing over to that page, Sebastian, one more pass this. Yep.
So there’s an entire, you know, you guys may have noticed fraud. feels more advanced than ever before. It also may feel more scaled than ever before. And it is because it’s true. There are now products in market, literally operated like businesses, SaaS businesses.
That are building software for fraudsters to take advantage of. So this one I’m showing here is called Ads Power. It is in what they call an anti-detector fingerprint browser. And it allows you to create a fresh fingerprint out of thin air on your device. So my Mac machine can be running a hundred different you know, windows, each of which has a unique fingerprint that are very hard to distinguish.
So standard device fingerprinting will be completely dismantled by this technique. And these are proliferating. Basically, every fraud farm we talked to has access to more than four of these different platforms. And they test all of them to see which one is going to be most effective and go, I would say, stealth for a given task. So this is what they step up from.
They have clean browser, then they just try the same browser again. If you catch that, then they go to this anti-dedect browser. It adds a little bit of cost, so that’s why they reserve it. What’s super interesting is from the data we’ve seen all over market research, Over fifteen percent of all claimed Mac browsers are actually on anti detect browsers coming from Windows machines. So one of the things you kind of the only way to detect these is to try to escape the context of the browser.
So basically it creates a perfect browser fingerprint. One of the ways you can detect it is actually looking beyond the browser to things that exist on the operating system. like the chips, the fonts, the emojis, things like that, which are operating system based. And so through that, we’ve been able to identify that. Fifteen percent of all claimed Mac OS sessions are actually Windows machines in disguise using anti-theft browsers.
Pretty cool. Our team is constantly working on new techniques to detect these. Some of the ones that we think are most interesting are literally invisibly rendering emojis that only the Mac or the Windows machine would have to distinguish which is which. And then, so if you start detecting this, or for example, if the fraudster needs to change their location, you might have IP deduplication checks on, or you might have location validation. the fraudster will use a proxy.
And these are not your grandfather’s data center proxies. These are modern residential and mobile proxies. What’s super powerful about these is that there are networks like Honeygain and other gaming applications which install on mobile devices and they pay users to use their network connection. So for example, I can actually just download Honeygain on my phone. I could then open up my network connection.
I’m in Utah right now. And my AT&T IP connection could actually be used by fraudsters all over the world. Now, what this actually means is that fraudsters can effectively hide in real IP networks. So for example, A fraudster in Nigeria can now be on the exact same AT&T mobile IP address that you or I are on. And so it’s much harder to detect those.
However, we’ve built systems that actually can, and that’s one of the ways we detect bad behavior. And so basically, if you can catch that, they go to something even more sophisticated, which is on the next slide. which is a premium VPN. So for example, NordVPN, et cetera, they offer dedicated private residential IPs where they basically buy a residential IP connection in kind of like a almost like a computer farm in random states around the US, like Kentucky or Ohio or whatever. And they actually have dedicated private VPNs for their users to use.
So you can get a VPN that no one else is using. So there’s no risk of a shared connection throwing risk against your IP. The way to detect these is actually with location spoofing detection. So the VPN itself is very challenging to detect the technology. And so you have to do it with location-based signals.
So going to the next slide here, what we found is almost all of the survey fraud that we see comes from other geographies outside of tier one markets where you’re likely surveying. Fraud comes from places like Syria, China, Philippines, Venezuela, and most critically, Nigeria, which is the largest survey fraud hotspot by a factor of, you know, ten. And what we find is that almost all of these locations, spoofers are going to the US and UK. The majority of them, sixty percent, are traveling or using proxies and VPNs to pretend to be in the US and the UK. Now, one of the ways to detect this, as I mentioned, is with kind of latency.
So we use the speed of light to our advantage. If the device is sitting in Egypt, pretending to be in the US, data will take a long time to move from that device to the US. And we can exploit that speed of light time and figure out they’re not where they say they are. So the last step they use, the most sophisticated thing they do, if we go to the next page, Sebastian, is they actually rent a virtual machine, or you may have heard VPS or data emulator in the US. So they get a data center machine.
They then run an anti-deduct browser and they then run a proxy. This is the most sophisticated because the machine’s actually sitting in the US. So you can no longer use that kind of latency signal we just mentioned. But it is the most costly, most complex, and most challenging for them to set up, which is why they’re reluctant to do. And most of these device farms actually don’t have this figured out yet.
And so most of them are operating their own devices in their fraud farm, not operating remote data center machines, though this is something that we’re tracking as one of the kind of emerging trends. So just to recap, each additional fraud layer that they add adds cost and complexity, which is why they don’t do it on every session, just on the ones where you’ve already caught the stuff ahead of time.
Leib Litman, PhDChief Research Officer, CloudResearchBehavioral fraud detection with Sentry12:10
Awesome. Thanks, Henry. A little bit terrifying, I will say. I am on a Mac here. I swear it’s not an anti-detect browser.
But thanks for running over those technical checks. Now Leib can talk more about the fraudulent survey behaviors that fraudsters use to fill out surveys. Take it away, Leib. Yeah, thank you, Sebastian. I’m going to focus more on the behavioral side.
And as Henry just showed us, there’s this increasingly sophisticated series of ways that fraudsters can disguise who they are and where they’re coming from. And as he was just showing us, there’s a global network of fraud from a variety of studies that have been conducted in the industry over the last several years. These fraud farms can account for anywhere between twenty and of any particular survey online. And what I’m showing you right now is what those actual fraud farms actually look like. And also what kind of tools they use, because we’re able to actually observe these tools as they’re being used by these click farms.
So for the last ten years or so, when we’ve encountered fraudulent participants, we followed up with them and then we asked them to take a study on video. And what you’re seeing right now is a video where there was a study that is, it’s supposed to be a US-based study that only US participants are eligible to take. However, this is a click farm that is in India and people who are in this click farm speak any English, but they use all kinds of sophisticated software that allows them to translate the survey, as Henry was saying, to evade detection so that it’s not clear that they’re not coming from the United States. And they engage in a variety of behaviors that are particularly problematic for research. So on this slide, this is an example of us showing a question where these people in these fraud farms, where they systematically provide data that’s not random, but are biased in a specific way.
So here we’re asking people, are you a petroleum engineer? Right. And the answer is yes. And then like an overwhelming number of times they’re going to say yes or they’re going to be positive and they’re going to be agreeable to. Oh, eighty percent of questions that are posed to them.
And this systematically biases the data. And again, the biggest issue is that going back to the way that Henry sort of introduced this cottage industry, there is just a tremendous amount of these click farms all over the world, and they can make up for anywhere between twenty and forty percent of a survey. And so you could imagine if twenty percent of or more of participants are systematically saying yes to something that that creates really a fundamental bias. So on the next slide, I just want to show you another tool that demonstrates just how much more sophisticated the tools that are available to these fraudsters has become. So this is Claude taking a survey with just a simple prompt of a single line.
All you have to do is just take a link to a survey, paste it into Claude, and say, here, go take a survey. And it’s going to open up a browser, just like you’re seeing right now. It’s going to spin up its own browser. And it’s going to take the survey in a way that is really, really sound. So it’s not going to fall for all kinds of red herrings.
It’s going to have really nice open ends, right? And we don’t think that these kinds of tools are being used en masse right now by these click farms, but they’re available and over time going to become increasingly more widely used. And so on the next slide, I want to talk about some ways that Sentry addresses these problems both from these human click farms and from increasingly sophisticated tools that they’re using. What we’re seeing is that we can look at the device level. This is what Henry talked about and Verisoul is amazing at being able to identify by looking at the device and looking at various signals, whether there’s fraud coming from a different country or there’s fraud coming from a click farm.
What Sentry does is that it provides an additional layer where it looks at the behavior of the respondents and what the respondent is doing as they’re moving through the survey. So how are participants answering structured questions, whether their respondents reveal inattention, whether they’re systematically saying yes in a way that’s indicative of someone coming from a click farm. And then we also look at evidence that something outside the responded is helping them to generate the answers. For example, translation software that we see they’re using when we engage with them on video, right? When they’re using various other software that helps them to generate open ends, right?
The ability to detect that they’re, for example, using an LLM and then copy and pasting that answer into an open-ended field. So the software that we use for this is called an event streamer. And the events framework allows us to not only look at the response that’s made, but actually how the response is made. So we’re able to see every single event within that engagement and to be able to see whether people are attentive, whether they provide good open ends, whether they’re using software like translation software, whether they’re likely to be using an AI agent. Right.
And in general, whether they’re generally coming from click farm.
Ellen LeeVP of Strategy and Insights, KADROExpanding the view of good data19:30
And you’re muted. Hello. All right. Thanks, Sebastian. So we just saw Henry and Leib kind of walk us through the different ways that modern fraudsters are manipulating their ways into surveys.
And I’m actually going to show us all how as an insight consultancy, we’re going to apply both of these techniques together to make sure that we’re ensuring that we’re weeding out bad respondents, but also bad responses still and ensuring that we can deliver like high quality results. And so on the next slide, we’re going to show kind of how all of those components are being brought together. So just quick two by two here, but essentially historically in the insight space, I know we’ve done a lot to focus on that quality of response, making sure that we have internal quality control flags, making sure we have red herrings, et cetera, et cetera, to make sure that people are signaling that they’re paying attention while they’re taking the survey. That being said, as you know, technology has gotten a lot more sophisticated. It’s really important that we kind of also layer in that quality of respondent as well.
So beyond someone just maybe not paying attention while they’re taking a survey, we wanna make sure that they’re also a real respondent and giving us good responses. So taking it all together, we’re gonna look at how Verisoul’s showing us real time detection terms of like, are they a real quality respondent, right? Like nothing from bot farms, nothing from using tons of VPNs, but also are they still providing the good quality response? And we’ll get that through Sentry as well. And so the big things that we tend to look at are, is this a real person?
Are they where they say they are? And lastly, the thing that we’ve always still needed to benchmark everything on, are they a good respondent? All right. And so with that, we’re going to walk through what the actual like century dashboard looks like to get a sense of how we’re going to piece folks and three and various all results together to kind of identify whether someone’s, which of these different types of fraudsters. So this is the first example I’m going to walk through.
This is a real respondent with a good response, really the kind of ideal person that we want to base our strategy and results around. So this person, as you can see on the first set of buttons under status, they’re not showing any inattentive behavior. They’re passing through the event streamer, which means they’re probably an engaged person and they’re taking the survey seriously. So they’re going to pass through the sensory prescreener. And then also through Verisoul, we are going to see that this person isn’t using any proxy, their device isn’t associated with fifteen other accounts on a panel, and they’re actually where they say they are.
So as an example, if we have a survey based in the US, they are truly in the United States and they’re going to be paying attention. So based on all the stories we’ve talked about, it’s probably the most vanilla of respondent type. So with that, we’re going to get to the next type of respondent. And here we have a real respondent, but still a bad response. So this is the type of person that we might have internal survey checks, trying to keep them blocked from the survey.
But what we’re going to find is with Sentry that they were blocked from getting into the survey in the first place because they’re going to show inattentive behavior. They’re just showing way too much yay saying, and their open end is a little bit problematic. I’m going to show the demo of that, of what that actually looks like through the event streamer. This person is still where they say they are. They’re a real panelist.
And so what this is showing is that this is testing a different layer of security in terms of we have someone that’s not paying attention, but they’re where they say they are. But we just don’t want them in our survey to begin with. And what we found is when we applied both of these softwares together, the in-house checks that we do internally, aside from all of these technologies, we’re going to be able to essentially pare down the amount of time we’re spending having to weed out people that aren’t paying attention. They’re real, but we still don’t want to hear their responses. And on the next slide, we will go through a bad respondent with a good response.
So this is a little bit scary in that Someone is looking normal. They have a good response at first glance, but we’re going to find that the fraud is really hidden in the signals that Verisoul is detecting that’s a little bit outside of kind of standard checks. So in this example, we’re seeing that the person is paying attention in this entry pre-screener. There’s nothing finicky happening there. But what we’re finding through Verisoul is the device is being associated with multiple accounts.
The respondent is using a proxy network. And as Henry mentioned earlier, these are the techniques that fraudsters are using kind of layer by layer in order to kind of bypass the different checks that different companies are using. So as an example, we had a project that was being run out of the US. It was a B to B survey, which means it’s typically a higher incentive. And what we found is that we were getting a lot lower response rate than we were typically used to.
And once we dug under the hood in the platform, We found that actually, as the team was saying before, a lot of the responses were coming from Nigeria, India. And so they were doing a really good job in terms of like weeding out people from getting in just based on their locations moving and their multi-accounting. And so that’s sort of one way that the two work together. And then on the next slide, this is the example of a bad respondent and a bad response together. And so this is a good way of showing how the two different softwares basically are going to be complementing each other by detecting different types of fraud that might have been missed.
So for example, on the Verisoul side, we’re finding that this person is associated with multiple different accounts. They’re not using the right browser. They’re spoofing their location. But also on the Sentry side, we’re seeing that there’s going to be a lot of yay saying, and they’re also using translation software. So what that might mean is we are deploying a survey in Japan, it’s happening in Japanese, and we’re finding that they’re translating it into a non-Japanese language in order to be able to bypass the survey.
And one just piece that I want to pause here is that because we have the event streamer, we’ll be able to record and be able to communicate with our panel partners and saying like, Hey, there may have been a miss here. We’re getting some issues and we’ll have recorded evidence to be able to show this is kind of what the supply chain of data quality is breaking down. And we can kind of figure out in an astute way where the issue is happening rather than having kind of a blind box understanding of. hey, we think this is a bad response. So being able to kind of have that detailed information of exactly where the problems are happening are good for us to be able to clear out bad respondents, but also to be able to ensure that the panels can weed out potentially bad sources of data as well.
And so here I’m gonna show, oh, there we go. So what’s happening right now is an actual event streamer recording And what we’re seeing is that the person is really, really rapidly just saying yes. They’re just trying to get into the screener. And so what we’re seeing is that that’s being recorded at an inattentive speed. We’re just going to pause maybe right at the open end, but What we’re seeing, I’m just going to recap really quickly.
The video of the respondents going through is super quick. They’re clearly clicking yes without even reading anything. And then one thing that we’re seeing at this very last question is it’s a pretty straightforward question. What’s your preferred flavor of ice cream? But what we saw is someone paused the survey.
They took the survey. They copy pasted the question. And then you see at the very end that just a giant block of text is being entered in. So in the past, we might have seen that someone had an overly detailed response to what might be your favorite ice cream flavor. And you might say to yourself, this seems like fraudulent behavior, but you might not necessarily have the evidence to detect, to be able to explain why that seems fishy.
Now with Sentry, we have the actual recording and we can say, this is clearly not something that someone can produce in an organic way. So it’s able to show both the, and just give us the evidence that we need to be able to kind of indicate like where exactly the issues are happening.
Sarah ZurillaVP, Client Services and Data Operations, ProdegeQualityShield and layered checks29:05
Cool. Thanks, Sebastian. I have to agree. Mint chocolate chip is also my favorite type of ice cream, but thank you to you and the Verisoul and Sentry teams. So it’s always a pleasure to work with you guys.
So the team and I at Prodege, we are dedicated to delivering exceptional service, ensuring data quality, and maintaining a client-focused approach. So I assume most of you know us, but we operate consumer reward platforms where people come to our sites, to take surveys, play games, shop and receive cash back and discover new products. So we’ve paid over two point five billion, not million, billion in rewards to our members. So it’s safe to say to operate at this scale, everything has to be built on trust and creating a safe environment for our clients, our advertisers and our members. So this means that data quality for us is our number one priority.
It’s something that we always are It’s the standard of how we work. So Quality Shield is our system to protect the research from the moment a respondent enters our ecosystem. And we choose to partner with industry leading tools as well as having our own proprietary software to ensure that anyone enters the survey is a real person. on a trusted device. And because people can do more outside of just surveys, we’re able to monitor every interaction that the members or different respondents have on our sites.
So this is a constant way of validating identity, creating different sessions. So from anywhere from registration all the way through the various sessions on all the different platforms. So the Prodege Quality Shield is a three-layered protection Secure Forward is our proprietary tool that instantly blocks low quality traffic, duplicate devices, and then Verisoul detects bots, emulators, and the professional survey farms and Sentry does the pre-survey checks that ensure respondents are attentive and authentic humans. So as I think a lot of you might know that fraud in this space is moving incredibly fast. So gone are the days where you can like focus on a single factor.
Or just kind of rely on platforms that aren’t evolving as quickly. So this is why we’ve chose to work with these partners and pairing with our internal compliance team. We have a process for ensuring that the response is valid and it’s a unique model. So that’s like our quality shield promise to you. And to show you how this actually works in real life, I want to go through a quick example of a fraud ring that tried to attack us, but the quality shield protected our client’s data.
So Sebastian, if we can go to the next slide. So this is an attempted attack. And so as Henry was mentioning, and this is a specific incident actually highlights why we chose to partner with Verisoul, because most other fraud detection platforms will rely on a pretty basic checklist. They might analyze the respondent’s device and the IP address and location. But as Henry mentioned earlier, these IP addresses can be spoofed.
And this is exactly what was happening in this case. So you can see that we were able to uncover that they were in Nigeria, but spoofing to say that they were in New York. Pennsylvania, Texas, and Oregon. So that was exactly what was happening and using our combined defenses. So it was both the compliance team at Prodege leveraging data from Secure Forward, Verisoul, and Century.
We caught something really sophisticated and this was the bot ring coming out of Nigeria. So as Henry mentioned, they were proxying through data center IPs right here in the U.S. And then actively spoofing these locations, making it look like they were authentic domestic respondents. But once we started peeling back the layers, we started seeing consistencies. And so this was a playbook that Henry mentioned, that they started small and then thinking with a few minor attacks.
And then seeing that they thought they found a loophole, they tried to automate it. Um, but weren’t successful because we don’t rely just on simple IP address checks. Uh, we have this triple layer shield in place. So we were able to see through the spoofing, identify the automation and stop the attack before it touched our client’s data. So then how did we catch them?
So you might say, okay, well, there’s a lot of different signals, but we looked under the hood and we found five things that were pretty similar across. So the first one was an email risk that Verisoul offers. So almost every account in this ring triggered Verisoul’s email risk flag. If any of you get to meet with Henry one-on-one or in a small group setting, he can look up, we did an SMR last year, my email address to see what my risk score was. I think it was like Ninety eight percent validated humans.
I don’t know what I did on that other two percent, but a loss that, you know, these ones had not did not have any history. So then we also looked at those geographic inconsistencies. So on the surface, we were they were routing through these U.S. data centers. But underneath, we saw that it was revealing that they were actually coming from Nigeria.
Nigeria. So the third thing, which is another thing that I do love about the Verisoul team is that we were able to catch them in a lie. It was a mismatch between the device they claim to be using. So claiming that they were on an iPhone, but when you looked at it, the hardware data tells us that they were actually on a desktop computer. So that mismatch is an incredibly strong red flag and just something that does not happen with a legitimate respondent.
The fourth thing was the AI generated responses. So these teams here both can identify if someone is using AI and give you that indication score. So that was something that clearly indicated that it was from AI. And then there was some ISP patterns that the traffic was flowing through uncommon and high risk internet service providers. So that was where we were able to work with Verisoul really quickly.
Having that data in hand, we ran some different queries in Snowflake. And that ultimately is how our system works really well because we combine our own proprietary data as well as the signals from our partners to not just play whack-a-mole with bot removal, but to eliminate that total network. So that also that was without ever disrupting any of our real members. And so then let’s go to the next slide, Sebastian. This is a visual of how it looks in practice.
So if you look, this is a real account cluster that Verisoul created. But if you look down at this, this graph, every circle on this graph represents an account. But the one that really matters is this one in the center. So that’s what we call the main account. And you could think of it, if we go back to the COVID days of everyone trying to figure out who patient zero was.
So we were able to identify patient zero and we talked about how they started small to test out the waters. But then from that single center, they were trying to send out different other attacks. But because we looked at those deeper signals, we were able to identify all of the consistency and go back to the source and not just eliminate these peripheral accounts, but really get to the source of the data to eliminate the whole fraud network. So that was taking down an entire cluster all at once rather than just removing people one by one. So thanks again to the Verisoul and Sentry team.
We very much appreciate the partnership and ability to be able to eliminate this fraud attack.
Jonathan BacchieriVP Supply, PotlocPremium proprietary respondents37:20
, who has some really interesting data to share, like actual hard data on how using these tools to stop modern fraudsters has impacted potlucks. So Jonathan, take it away. Thank you, Sebastian. So you’ve now seen this from multiple angles. Henry showed you how fraudsters climb the sophistication ladder.
Lee broke down the behavioral signals. Helen reframed what good data even means. And Sarah will walk you through Prodigy’s approach to data quality and the live attack getting caught. So what I’m going to show you today is how we wired Verisoul and Sentry into Potloc’s production stack. And just a quick context before I get the ball rolling.
So at Potloc, we run a multi-source supply ecosystem, two proprietary panels, one for B to C, one for B to B. Also a pioneering social media sampling technology and a vetted partner network for the world’s top consulting and PE firms. So the data has to be clean and sharp. And here’s how we keep it that way. So before any tool, quality starts with the supply chain itself.
So we think of it as four ingredients when sampling with our proprietary panels. The first one, obviously, is the source. So we acquire proprietary respondents through premium channels, Meta, Google, Snapchat, LinkedIn, and also a network of vetted affiliates. We don’t buy from cheap lead lists or incentivized offer roles. The second ingredient is members engagement.
So we pay real cash, pretty generous incentives, and we use AI to match the right respondent to the right survey. All of that to create a virtuous lifecycle. So we have better conversion rates, frequent payouts, and respondents who don’t feel they have to cheat the system to get something out of it. The third ingredient, survey design. So that might sound obvious, but it’s not always the case.
Short, logical surveys of the right length, so good respondents stay engaged instead of spinning through. And last but not least, quality controls. so this is where we are also generally different because we program and host about ninety percent of the surveys we run meaning that we own the entire chain scoping scripting survey design programming sampling data quality all the way to the insights a pure sample provider only gets to check the respondents at the door we get to check them at the door and all the way through the survey so beyond pre-service screening we also run in survey checks so the usual attention traps coherence checks to that catch people contradicting themselves, duplicate email detection mid-survey, and then we also check the data post-survey complete before anything reaches the client. So the speeder detection, open-end quality checks, including now LLM-generated answer checks, plus similar answer and copy-paste detection across respondents with a human review layer on top. So various one century, they plug into that complex ecosystem.
They don’t replace the chain. They’re just part of it. So now let me show you where each one of the tools sits. So first, Verisoul sits at the very front door the moment someone tries to register on one of our panel before they’re ever respondent. So every registration gets screened in real time against a wide signal set.
Bot detection, account sharing, account takeovers, synthetic identities, VPN and proxy use, duplicate accounts. And for BtoB panel, we also have a selfie to ID face match that actually kills impersonation and deepfakes. Then there is a verdict. So real joiners go into the panel, suspicious ones get flagged for an extra check on the first survey attempt, and fake ones get automatically blocked. And if they came from a paid affiliate, the payment is automatically reversed through an API.
That last part actually matters a lot because it’s a closed loop with zero analyst time, and it’s a big part of why the stack pays for itself. More on that in a second. If you go to the next slide. So if very small secures who joins century actually secures how they behave so century sits on the survey itself at the pre survey checkpoint it screens. Before the respondents even see the question one across our whole online ecosystem it runs to net on the left side, the technical nets that catches what shouldn’t be there.
So duplicates, high thread bots and fraudsters, people trying to tamper with the data. So using translation tools, service holders. And on the right, you have the behavioral net that catches how people answer. So inattentiveness, session recording, the sort of stuff that a technical check alone will never see. And that behavioral layer is the difference maker.
And I’ll show you the numbers on it right now. So here’s the whole thing in one line. VerySource secures the respondent. Sentry secures the response. And the results speak for themselves.
I’m a data guy, so I couldn’t end the presentation without talking about data. So on our proprietary B-to-B panel, we have up to thirty three quality controls per respondent. Our client reconciliation rate is sub one percent and our own post cleaning rate is sitting at eleven percent. On Verisoul specifically, we have ninety one percent of our new joiners that are real and we register one hundred and thirty three percent year one ROI thanks to the automated affiliates reconciliation that we have. So that’s what I was saying before.
It literally pays for itself. And on Sentry, it catches forty seven percent of all our pre-serve determinations, which is nearly half of them over the last twelve months coming straight from our analytics. So that’s how we turn quality from a checkbox into an entire architecture. And I want to thank also the Sentry and the Verisoul team for their great work on all that. And thank you for your time.
Leib Litman, PhDChief Research Officer, CloudResearchCase study: Verisoul and Sentry combined43:55
To close off this webinar before the Q&A. Thank you. Awesome. Thank you so much, Jonathan. I think I speak for anyone here who ever has bought research when I say that a sub one percent reconciliation rate is truly, truly impressive.
So thank you for walking through that. And I’ll pass it over to Leib, who can walk through why this matters for your data quality. Yeah, and Thank you. Thank you, Jonathan. I think seeing that less than one percent reconciliation rate is amazing.
And I will just take a minute or two to expand on that, because really everything that we’ve discussed today ultimately comes down to the main question, which is, do these protections produce more accurate data? Ultimately, that’s really what we care about. In the video that I showed earlier, there was a slide where I asked one of these respondents from a click farm whether he was a petroleum engineer, and he said yes. If we were to do that in a survey, if we did a representative survey within the US and we were to ask what percentage of people should say that they’re a petroleum engineer, the answer should be somewhere around one percent or less. Yet, when we do these studies online, what we see typically is that the answer is not less than one percent.
The rate is actually closer to over ten percent, which is really astonishing, which means that if something is not actually happening, is happening at a rate of zero, we will think that it’s actually happening to twelve percent of people. If no one likes a product with these click farms, we will think that twelve percent of people like that product. We ask questions like this all the time just to see what is the rate of bias in online surveys. So, for example, we will get. Fifteen percent Tesla ownership where the expected rate in the US is closer to two percent.
Right. We will get almost eighteen percent of people saying that they went scuba diving in last year, whereas in the actual population, it’s one percent. So all those numbers are numbers that we would see without any kind of data quality protection. However. When we have Verisoul and Sentry protecting the survey, what we see, and this is what you’re seeing here in the second column, that’s more in green, is that the reported rates of all these behaviors now starts to match what it’s supposed to be.
So petroleum engineer reporting is close to zero percent. What does that mean? That means that we’ve gotten rid of all the fraudsters who are misreporting who they actually are and what they actually do. And that’s ultimately the goal of what we’re trying to achieve here is to get a clean and a reliable signal for our market research. So overall, the central takeaway is that data quality is not a single check.
It requires protections at different levels that validate both the respondent at the device level and the response at the behavioral level. And when these layers work together, we’re ultimately doing the most important thing for the survey, which is to protect the signal that researchers and decision makers actually need in their data.
A Rare Look Behind the Curtain
Most conversations about survey fraud stay abstract, focusing on bad actors, bot traffic, low-quality data. Inside the Mind of a Fraudster, a recent webinar hosted by Verisoul, took a more direct approach. Leib Litman, Chief Research Officer at CloudResearch, joined panelists from Potloc, Prodege, and Kadro to break down the actual tools fraud farms use to slip past detection, and how researchers are catching up.
Fraud Escalates One Layer at a Time
The panel opened by pushing back on a common assumption: that all fraud looks the same. In reality, fraudsters work through something closer to a decision tree. They start with the cheapest, simplest tactic available and only add complexity once that tactic gets flagged.
The escalation typically looks like this:
- A clean session. Normal device, normal network. It’s usually the first thing tried, and usually the first thing caught.
- A simple cookie clear. Same device, same fingerprint, just wiped between attempts: an easy catch for standard device fingerprinting.
- An anti-detect browser. Once device fingerprinting starts working, fraudsters switch to software built to generate a fresh, convincing fingerprint out of thin air. A single machine can run dozens of these fingerprints at once, each one looking like a separate device.
- Residential and mobile proxies. Apps that pay ordinary users to share their internet connection let fraudsters route traffic through real home and mobile IP addresses, so a fraudulent session can look like it’s coming from the same network as a genuine participant.
- Premium VPNs with dedicated residential IPs. A step up from shared proxy pools, these give a fraudster a private connection with no risk of overlapping with other fraudulent sessions.
- A rented virtual machine in the target country. The most expensive and complex setup: a data center machine in the target country, running an anti-detect browser through a proxy. It defeats even latency-based location checks, but it’s also the rarest tactic, since most fraud farms haven’t built the infrastructure to pull it off yet.
Each added layer costs more time and money, which is exactly why fraudsters save it for cases where a simpler approach has already failed.

The Tells That Survive the Disguise
Anti-detect browsers are built to make device fingerprinting useless. But some signals live above the browser, at the operating system level: fonts, emoji rendering, and hardware behavior that a spoofed browser can’t fully fake. Verisoul has used exactly this kind of signal to confirm that a meaningful share of sessions claiming to be on Mac are in fact Windows machines running anti-detect software.
Location is trickier to verify directly, so detection leans on physics instead: data traveling from a real device to a distant server takes measurable time. A session claiming to be in the US while the device sits overseas will show latency that gives it away, at least until a fraudster upgrades to a same-country virtual machine, which is why that tactic is the trend worth watching most closely.
Fraud Farms Are Global, But Concentrated
Survey fraud isn’t evenly spread across the globe. It overwhelmingly originates from outside tier-one research markets, with Nigeria standing out as the largest single source by a wide margin, alongside Syria, China, the Philippines, and Venezuela. Most of these sessions are built to appear as though they’re coming from the US or UK.

Behavioral Fraud Needs a Different Kind of Detection
Where device-level tools focus on who a respondent really is, CloudResearch’s Sentry looks at how they behave inside the survey. Fraud farms tend to produce systematically biased data: a fraudulent respondent asked “are you a petroleum engineer?” will say “yes” far more often than a real population should, simply because agreeing is the fastest path through a screener.
In addition, AI agents can now complete a survey with nothing more than a single-line prompt and a link, navigating cleanly and producing coherent open-ended responses. These tools aren’t yet widely used by fraud farms, but the barrier to entry is low enough that this is unlikely to stay true for long.
Sentry’s approach is to track not just what answer a respondent gives, but how they arrive at it: attention patterns, response speed, signs of translation software, and indicators that an AI tool generated an open-ended answer before it was pasted into the field.

Four Respondent Types, and Why One Layer Isn’t Enough
One panelist mapped respondents along two dimensions: real vs. fraudulent, and good vs. bad response.
- Real respondent, good response. No inattentive behavior, a clean pass through device-level checks, genuinely located where they claim. The ideal case, and the most common one.
- Real respondent, bad response. A genuine panelist who isn’t paying attention (e.g. heavy yea-saying, thin open ends) but who is exactly where they say they are. This is the profile traditional in-survey checks are built to catch.
- Bad respondent, good response. The unsettling middle case: the response itself looks fine, but device-level signals show a proxy connection or an account linked to multiple other panel profiles.
- Bad respondent, bad response. Both layers catch something. Device signals reveal a spoofed location and linked accounts, while behavioral signals show excessive agreement and translation software in use.
Running device-level and behavior-level detection together meaningfully cuts down the time spent manually reviewing questionable responses and reduces the odds of discovering after fieldwork closes that a chunk of the sample wasn’t who it claimed to be.
Practical Takeaways From the Q&A
- Screener questions need validation, not just cleverness. Ambiguous or tricky screeners can catch fraud, but they can just as easily flag a genuine respondent with an unusual but real answer. Pulling from a large, validated question library and rotating a random subset per survey helps avoid overfitting to any one pattern.
- Friction should scale with what’s at stake. For high-value audiences, such as C-level executives and highly strategic B2B studies, added friction like biometric or ID verification is worth the tradeoff. For broader consumer studies, standard technical checks are typically enough.
- Low-friction signals are an emerging layer. One example raised was checking how many established online accounts are tied to a respondent’s email address. Genuine respondents tend to have a real digital footprint; fraudsters are more likely to be operating from freshly created addresses with none.
- IP-based detection has moved past static blocklists. Proxy pools rotate too quickly for a “known bad IP” list to keep up. The more durable approach is real-time session verification, checking whether the network signals a device produces match the device it claims to be.
Why Layered Detection Matters More Than Ever
No single check catches everything, because fraud itself isn’t a single thing. It’s a moving target that adapts to whatever defense it meets first. Device-level, network-level, and behavioral detection each close a different gap, protecting against fraud farms who count on researchers to rely on just one.
Want to join us live next time? See upcoming CloudResearch events.