Calendar Icon - Dark X Webflow Template
August 25, 2026
Clock Icon - Dark X Webflow Template
14
 min read

Data quality in surveys: what researchers must audit and fix

Learn how to improve survey data quality using Total Survey Error, pre-field QC, paradata, fraud checks and transparent reporting.

Data quality in surveys: what researchers must audit and fix

Data quality in surveys means the extent to which a dataset is fit for its intended use once every error source in the Total Survey Error (TSE) framework, coverage, sampling, nonresponse, measurement and processing, has been accounted for. Two standards worth knowing before anything else: AAPOR’s disclosure guidance and GESIS’s Survey Quality Predictor (SQP), which scores measurement risk before a single response comes in.

Here’s the uncomfortable context: close to 50% of published articles using online surveys report no procedure at all for evaluating the quality of the data behind their conclusions.

Three things to do before reading further:

  • Pre-specify your quality rules and exclusion thresholds in writing, before fieldwork opens.
  • Switch on paradata collection and basic bot/duplicate detection at the platform level.
  • Save the raw, untouched data file separately from your working file, and never edit it directly.

Key Takeaways

Reliable survey data comes from pre-registering quality rules against the Total Survey Error framework, not from cleaning up problems after fieldwork ends.

Point Details
Use TSE, not one metric Judge fitness for use across coverage, sampling, nonresponse, measurement and processing together.
Pre-specify thresholds Lock exclusion rules and cut-offs before fieldwork to avoid post-hoc bias in cleaning decisions.
Expect 5 to 15% exclusions Rates above 20% to 32% usually signal a questionnaire design problem, not just bad respondents.
Match QC to platform MTurk needs heavier fraud screening than proprietary panels or Qualtrics-managed samples.
Document for audit Publish frame details, rates, thresholds and exclusion counts, and keep flagged records rather than deleting them.
Bring in specialist support Skopos applies TSE-consistent design, piloting and QC reporting across its market research projects.

    Why the Total Survey Error framework should organise every quality check

    Most researchers still ask “is this survey any good?” as if quality were one number. It isn’t. The Total Survey Error paradigm treats quality as a judgement about fitness for use, built from several distinct error sources that behave independently. A survey can have a superb response rate and still be measuring the wrong thing entirely.

    Map your checks against the five components and the gaps become obvious fast:

    • Coverage — does your sampling frame actually represent the population you're claiming to describe? Check the frame against a known population benchmark before fielding.
    • Sampling — how were respondents recruited, and does the recruitment path introduce systematic skew? Compare achieved sample against target quotas.
    • Nonresponse — who declined or dropped out, and do they differ from completers on anything you can observe? Run cross-tabs against known population parameters.
    • Measurement — does the question wording produce the same meaning for every respondent? This is where cognitive testing and SQP earn their keep.
    • Processing — are your cleaning, coding and weighting steps documented and reproducible, or did someone edit values by hand in a spreadsheet?

    A single design flaw, say, recruiting through a Facebook ad, shows up in every one of these boxes at once. It skews coverage (only Facebook users are reachable), skews sampling (algorithmic ad delivery favours certain demographics), inflates nonresponse bias risk, and can even distort measurement if the recruited group interprets questions differently from your target population. Judging fitness for use means asking whether that combined distortion still lets you answer your actual research question, not whether any single metric looks acceptable in isolation.

    How do you protect measurement quality and comparability?

    Measurement error is the one TSE component researchers most often skip, because it requires slowing down before fieldwork rather than cleaning up after it. Get the checklist right and you avoid rebuilding trust in results later:

    • Define what each question is meant to measure before you draft the wording, not after.
    • Run cognitive interviews or think-aloud pretests with a handful of target respondents.
    • Translate and back-translate any multi-market instrument, then check equivalence with anchoring vignettes.
    • Pilot with enough completes to see genuine variance, not just "it ran without crashing."

    SQP scores individual survey items on expected reliability, validity and method effects before you field a single response, using characteristics of the question’s wording, response scale and context that its underlying meta-analysis links to measurement error. Feed it your draft items during questionnaire development and it flags which questions are likely to underperform, letting you fix wording before the cost of a bad question compounds across a full sample. GESIS maintains it as a free tool, and it’s particularly valuable when you’re running the same instrument across multiple countries, where subtle translation drift can quietly break comparability across markets.

    Watch for the usual culprits: double-barrelled items asking two things at once, jargon that means one thing to researchers and nothing to respondents, and order effects where earlier questions prime later answers. Cheap diagnostics catch most of this after the fact: item non-response rates that spike on a specific question, unusually low variance suggesting respondents didn’t engage with the scale, and item-total correlations that reveal a question isn’t measuring what the rest of the battery measures.

    Pro Tip: If you’re choosing between a bigger pilot sample and more cognitive interviews, pick the interviews first. A pilot with 200 completes tells you a question performs badly; five cognitive interviews tell you why, and only the second gets you to a fix.

    What sample composition checks catch problems early?

    A response rate on its own tells you almost nothing about whether your sample represents the population you care about. Run these checks before you trust the topline numbers:

    • Document the sampling frame: where did the list or panel come from, and who is structurally excluded from it?
    • Trace the recruitment path end to end, including any subcontracted panel providers.
    • Record inclusion and exclusion criteria exactly as applied, not as intended.
    • Track quota performance by cell, not just in aggregate.
    • Check geographic coverage against known population distribution.

    Four metrics are worth computing and reporting every time: frame coverage (the proportion of your target population that could theoretically be reached through your frame), response or participation rate (completes divided by eligible contacts), cooperation rate (completes divided by all contacts who engaged at all), and completion rate (finished surveys divided by started ones). None of these numbers is inherently good or bad. A 15% response rate from a well-matched frame can outperform a 40% response rate from a skewed one.

    Spotting nonresponse bias doesn’t require sophisticated modelling. Cross-tab your achieved sample against two or three variables you can benchmark externally, age, region, or a behaviour tracked in official statistics, and see where the gaps sit. If the gaps are structural rather than random, weighting and post-stratification correct for known skew, but they cannot fix a frame that never reached the right people in the first place.

    Pro Tip: Longer field periods and larger incentives both tend to raise response rates, but incentives that are too generous relative to the task can attract satisficers chasing the reward rather than engaged respondents. Match incentive size to task length, not to what maximises completes.

    How do you detect careless, inattentive or fraudulent responses?

    Not every problem response is the same problem, and treating them identically wastes effort. A framework for assessing response quality sorts problematic answers into three distinct categories: response styles (consistent tendencies like always picking extreme options), insufficient-effort responding (straight-lining, speeding through without reading), and response manipulation (bots, duplicate submissions, coordinated deception for incentive fraud). Each needs a different diagnostic.

    Your practical toolkit:

    • Attention checks / instructed-response items (IMCs) — direct instructions embedded in the questionnaire that a careful respondent will follow.
    • Speeding thresholds — flagging completion times below a defensible minimum reading speed.
    • Straight-lining detection — identical or near-identical responses across a grid battery.
    • Individual response variance (IRV) — low variance across items that should logically differ signals disengagement.
    • Odd-even consistency checks — comparing responses to conceptually similar items split across the questionnaire.
    • Mahalanobis distance — a multivariate outlier statistic that flags response patterns far from the norm.
    • Open-text screening and duplicate/IP checks — catching bots and repeat submissions from the same source.

    How many responses should you expect to lose? Industry baselines sit around 5% to 15%, though some published cleaning studies report removal rates as high as 20% to 32% depending on platform and topic sensitivity.

    The safest decision rule combines indices with an AND condition rather than flagging on any single failed check, which cuts false positives sharply. Never delete flagged records outright: keep them with flags attached, then run your headline analysis both with and without them as a sensitivity check. If the conclusions hold either way, you’ve earned confidence in the finding rather than just in the cleaning.

    Pro Tip: Lock every threshold, speeding cut-off, IRV minimum, Mahalanobis critical value, before you look at the outcome data. Setting thresholds after seeing which respondents “look wrong” is how post-hoc bias creeps into supposedly objective cleaning.

    What paradata should you monitor while fieldwork is running?

    Paradata are the digital exhaust of a survey: timestamps, page-by-page timing, break-off points, device and browser type, and geolocation or IP fingerprints, plus whatever metadata your panel vendor supplies about respondent history. Waiting until fieldwork closes to look at any of this wastes the one advantage online surveys have over paper ones: you can watch quality in real time and intervene while there’s still time to fix it.

    Three alerts are worth building even in a modest dashboard:

    1. A burst of completions clustered in an implausibly short window, often the signature of a bot script or a panel source gaming completes.
    2. An abnormal shift in the device mix partway through fieldwork, which can signal a new, lower-quality traffic source has been switched on by a vendor.
    3. Quotas filling unusually fast from one geolocation cluster while others lag, hinting at a single source dominating the sample.

    A short implementation checklist keeps this manageable: enable paradata capture at setup, define alert thresholds before launch, review a daily QC dashboard rather than waiting for a weekly report, and agree in advance who has authority to pause a source. In practice, that might mean pausing recruitment from a specific panel supplier the moment speeding rates from that source spike, rather than discovering the damage during post-field cleaning.

    Do platforms like Qualtrics, SurveyMonkey and MTurk differ on quality?

    Yes, and the differences are large enough to change your entire QC strategy depending on where your sample comes from. The European Social Survey (ESS) remains the closest thing to a methodological exemplar for probability-based, cross-national fieldwork, built around exhaustive documentation and TSE-consistent reporting. That’s the standard to aspire to; it’s rarely the standard commercial fieldwork can fully replicate on commercial timelines and budgets.

    Platform-level evidence backs this up directly. A study applying 20 standardised quality checks across three crowdsourcing platforms found good-quality response rates ranging widely across platforms, from low on Amazon Mechanical Turk (MTurk) to relatively high on SurveyMonkey, a sevenfold gap driven largely by vendor-level fraud controls and audience composition rather than anything the researcher did.

    • Qualtrics panels: strong built-in fraud detection and quota tools; still needs attention checks layered on top.
    • SurveyMonkey audiences: comparatively higher baseline quality in the cited study, but verify recruitment source per project.
    • MTurk / crowdsourcing: highest fraud and inattention risk of the three; needs the heaviest QC layer, including CAPTCHA-style screening and duplicate-worker checks.

    Pro Tip: Don’t run the same QC package across every platform. A rule set calibrated for a proprietary panel will under-catch fraud on a crowdsourced sample, and an MTurk-grade screening battery is often overkill, and annoying, for a well-managed panel.

    What belongs in a survey data quality report?

    Transparency isn’t a courtesy to readers, it’s the mechanism by which anyone can judge whether your fitness-for-use call was reasonable. AAPOR’s guidance treats disclosure and routine publication of quality metrics as essential, not optional, precisely because errors are only assessable when the procedures behind them are visible.

    A publishable quality report should disclose:

    • Frame and recruitment details, including any subcontracted panel sources.
    • Field dates and total field period length.
    • Response, cooperation and completion rates, defined explicitly.
    • Every quality check applied and the exact threshold used for each.
    • Numbers excluded at each stage of cleaning, not just a final total.
    • Weighting methodology and the variables used for post-stratification.
    • A summary of paradata patterns observed during fieldwork.

    A reusable outline follows a simple order: methods, QC steps applied, exclusions by stage, sensitivity analyses run, and limitations. Write cleaning steps as syntax or scripts rather than manual spreadsheet edits, which preserves an auditable trail that anyone can rerun and verify.

    Pro Tip: Attach a short machine-readable QC appendix, a table of flags and counts, alongside the narrative report, and keep excluded records rather than deleting them. Auditability is worth far more after publication than it feels like at the time.

    What’s the step-by-step checklist from pre-field to post-field?

    Treat quality control as three phases, each with its own actions rather than one generic “clean the data” step tacked on at the end.

    1. Pre-field: define research objectives precisely, pre-register your quality rules and exclusion thresholds in writing, run a pilot with cognitive interviews, switch on paradata and vendor-level bot protections, and choose platform-appropriate checks based on where your sample will come from.
    2. In-field: monitor a daily paradata dashboard, run speed and attention summaries every day fieldwork is open, apply pausing rules and reallocate quotas away from underperforming sources, and log every intervention with a timestamp and reason.
    3. Post-field: apply your pre-registered rule-based filters, compute IRV and Mahalanobis distance for multivariate outliers, consolidate every exclusion flag into one master file, run sensitivity analyses comparing results with and without flagged cases, and produce the QC appendix for publication.

    Threshold examples worth adopting as defaults: a completion-time floor set relative to your median reading speed for the instrument, and an IRV cut-off calibrated against your pilot data rather than borrowed from an unrelated study. On exclusion rates, hold the same benchmark from earlier in mind: 5% to 15% is typical, and rates above 20% to 32% usually mean the questionnaire, not the respondent pool, needs a second look.

    How does Skopos apply these standards in practice?

    A typical Skopos survey moves through the same TSE-consistent sequence this guide describes: objective setting and questionnaire design, a pilot with cognitive checks, live fieldwork with paradata monitoring switched on, structured post-field QC, and a transparent debrief that shows the client exactly what was excluded and why.

    Each stage maps to something the client actually receives:

    • Design and pilot stage feeds the Advertising and creative testing and concept work where measurement validity determines whether a result is actionable.
    • Live paradata monitoring uses the same real-time logic behind PulseCheck™, Skopos's continuous sentiment tracking tool.
    • Post-field cleaning and sensitivity checks are documented into an executive debrief, not left as an internal-only appendix.
    • Multi-touchpoint programmes route feedback through Loopback™ to keep QC consistent across channels.

    Pro Tip: Agree the QC plan, thresholds and all, with the client before fieldwork opens, not after seeing the results. It turns “why did you exclude these responses?” from an awkward post-hoc justification into a decision everyone already signed off on.

    What trade-offs actually matter when you’re the one making the call?

    Every quality rule in this guide sounds clean on paper. In practice, you are usually choosing between two imperfect options under time pressure, and the honest answer is that neither rigour nor speed wins every time.

    Favour conservative exclusion rules when the decision riding on the data is high-stakes and hard to reverse, a pricing decision, a market-entry call, anything a client will act on immediately. Favour a lighter touch with heavier sensitivity reporting when the sample is already small or hard to replace, and losing another 15% of it would leave you with too little to say anything useful at all.

    What I’d push back on is the instinct to treat exclusion rate as a badge of rigour. The better discipline is documenting the trade-off explicitly and discussing it with stakeholders before fieldwork, not defending it after the fact. Clients respect a pre-agreed threshold far more than a post-hoc explanation, however statistically sound that explanation is.

    How Skopos can help you get this right

    If you’re weighing up whether to build this QC infrastructure in house or bring in a partner who already runs it on every project, Skopos supports exactly the stages this guide covers: questionnaire design with pretesting, pilot fieldwork, real-time QC monitoring, and the exclusion documentation you need for a defensible quality report.

    Skopos works to a data quality excellence pledge that puts TSE-consistent checks and transparent reporting into every project rather than treating them as an optional add-on. Relevant services include:

    • Questionnaire design and cognitive pretesting for measurement validity.
    • Pilot fieldwork with paradata monitoring built in from day one.
    • Post-field QC appendices and sensitivity reporting suitable for publication or board presentation.
    • Ongoing tracking programmes where consistent QC matters as much as the topline numbers, supported by tools like Brand Tracking.

    For readers wanting third-party guidance on converting cleaned data into decisions, Kontrol Media’s guide to informed decision-making is a useful companion read.

    If your next project needs this level of rigour built in from the start, get in touch with the Skopos team to scope a questionnaire design and QC plan before you go to field.

    Frequently asked questions


    Run attention checks, speeding thresholds and straight-lining detection together with an AND rule, then compare your headline result with and without the flagged cases before deciding anything is final.

    How much survey data should I expect to exclude?

    Does SQP replace the need for pretesting? No. SQP predicts measurement error from question characteristics before fieldwork, but cognitive interviews catch comprehension problems SQP's scoring model can't detect on its own.

    Why do MTurk and SurveyMonkey show such different quality rates? Vendor-level fraud controls and audience composition differ sharply between platforms, which is why a study applying identical checks found good-quality rates ranging from roughly one in ten completes to over seven in ten depending on the platform.

    What should always go in a survey quality report? Frame and recruitment details, field dates, response and completion rates, every threshold applied, exclusion counts by stage, and the weighting method used, published alongside the results rather than held back.


    Sources

    Data quality in surveys: what researchers must audit and fix

    Author: Michael King

    Latest articles

    Browse all
    Get In Touch