Skip to main content
Research in the Cloud Textbook, forthcoming with Cambridge University Press

Data Cleaning in Online Research: A Practical Guide

Aaron Moss, PhD5 min read

Data Cleaning in Online Research: A Practical Guide

In this post:

  • Why data cleaning is essential before running any statistical analysis
  • How to spot problematic responses, from failed attention checks to AI-generated answers
  • How to set exclusion thresholds based on your research goals
  • Step-by-step guidance for cleaning data in SPSS

Within the United States and other Western nations, participation in organized religion has been on a decades long decline. Year after year, organizations like Gallup and Pew have reported that the number of people who identify with a religion or regularly attend religious services has steadily decreased. That means it came as a surprise when the Bible Society, a non-profit dedicated to translating and sharing the Bible, published a report in 2024 claiming that church attendance in the UK had increased, and increased the most among adults aged 18-24. They called it the “Quiet Revival.

The revival garnered a lot of attention. In the US and UK, news outlets reported that the trend toward secularization was over. Some less restrained pundits and outlets spun up theories about why Gen Z suddenly found going to church cool: it was an anecdote to loneliness, polarization, uncertainty, a salve for anxiety and depression, a search for meaning. While the theories made sense, experts urged caution and, ultimately, the uptick in church attendance turned out to be nothing more than survey error.

Early in 2026, the Bible Society and the polling firm YouGov retracted the surveys they conducted on church attendance. As the statement released by YouGov said, it wasn’t that the data were not checked for quality in 2024, but that “some of the anti-fraud measures available in 2024 were not administered in the optimal way.” In other words, someone made a mistake.

Whenever researchers gather data online, it is important to include validated measures of data quality. Then, after the data are gathered, these measures must be assessed and the researchers has to decide which participants to include in the analysis and which to exclude. Getting these decisions wrong, can lead to cases like the Bible Society. Chapter 12 of Research in the Cloud walks you through the data cleaning process step by step. This blog gives an overview of the process.

What is Data Cleaning?

Data cleaning is the process of reviewing raw data, identifying the problematic responses, and deciding which participants to remove from your analysis. It happens after data collection is complete and before statistical analysis begins.

Beyond ensuring the data were properly collected and each person has a complete response, the process typically involves three things.

  1. Reviewing attention checks. Data cleaning begins with examining how participants performed on attention check questions and other indicators of data quality. It is best to use validated measures of attention so you know you are removing the right people. This chapter walks you through how to score these items and flag participants who fail to meet your threshold.
  2. Examining open-ended responses. Open-ended questions show whether participants engaged meaningfully with your survey and are often useful for spotting participants who used AI. Responses that are nonsensical, copied from the question, or clearly generated with AI are signs of fraud. These participants should be dropped.
  3. Checking for impossible or inconsistent responses. Some responses are simply not credible. No one, for instance, can complete a 15-minute survey in three minutes. People with ages that don’t match their reported life experiences or contradictory answers to related questions also raise suspicion. When the contradictions are clear enough, the data should be discarded.

Quick Data Cleaning Checklist:

  • Score attention check items and flag failures
  • Review open-ended responses for nonsense, copy-paste, or AI patterns
  • Check completion times against expected survey length
  • Look for contradictory or impossible response patterns
  • Document your exclusion criteria before applying them

How to Decide Which Participants to Exclude

Once you’ve flagged problematic responses, you need to decide what to do with them. This is where the fit-for-purpose framework from Chapter 11 comes into play.

For some studies, a moderate threshold makes sense. You may exclude participants who failed two or more attention checks while retaining the rest. For studies that examine a rare behavior or a small sub-group of participants, you may need stricter criteria. Chapter 12 provides practical guidance on setting exclusion thresholds and how to handle edge cases. It also recommends pre-registering your data cleaning plan so that your exclusion decisions are transparent and defensible.

Tools for Data Cleaning

The chapter includes hands-on examples using common tools. You’ll learn how to implement data cleaning procedures in SPSS with step-by-step instructions for scoring attention checks, flagging problematic responses, and filtering your dataset.

Why Data Cleaning Matters

Data cleaning is easy to neglect. When you first download a data file, few things are as dispiriting as realizing you must spend an hour or more cleaning the data before you get to answer your research questions. But, without taking the time to clean the data, the answers you pull out of your analysis may be unreliable.

A few minutes spent cleaning your data can prevent problems down the road. Until you’ve developed your own workflow, Chapter 12 of Research in the Cloud provides a hands-on guide to walk you through the process step by step.

This post is part of a series exploring the chapters of Research in the Cloud: An Introduction to Modern Methods in Behavioral Science by Aaron Moss, Jonathan Robinson, and Leib Litman. Explore Chapter 12 here.

Share this blog

Subscribe to Receive Updates

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.