A few years ago, "low-quality survey response" meant someone typing "n/a" in every open-ended box or straight-lining a rating grid to get to the gift card faster. That problem hasn't gone away, but it has a new, better-dressed cousin. Open-ended answers are showing up in research datasets that are a little too well-written: no typos, no rambling, no half-finished thought, just clean, complete paragraphs that read like they were drafted by someone who had all the time in the world.
Increasingly, they were, just not a person. This is the AI-generated survey response problem, and it is no longer a hypothetical.
Researchers Are Already Catching This in Their Data
This isn't a fringe concern being raised by one lab. A Stanford Graduate School of Business researcher first noticed the pattern after a colleague flagged that some open-ended answers "seemed nonhuman," and a follow-up study formalized the finding.
34% of survey participants in that study admitted to using an LLM to help answer open-ended questions.
Separately, NORC at the University of Chicago, one of the oldest and most respected survey research organizations in the country, has built its own internal tool to flag AI-written responses in survey data and reports catching them with over 99% precision on its own training set.
Neither of those numbers is a fluke or a scare tactic. They're a signal that the tools researchers built to catch bots, duplicate IP addresses, and speeding through a survey were never designed for a respondent who can generate a thoughtful, on-topic paragraph in two seconds. The homogenization is the part that should worry you most: when a meaningful share of "different" opinions in your dataset were actually written by the same handful of language models, your results start to look more consistent than your customers or employees actually are, which is the opposite of what a feedback survey is supposed to tell you.
How to Spot an AI-Written Survey Response
You don't need a research lab to start noticing the pattern. A few signals tend to repeat across the responses researchers have flagged as likely AI-assisted:
- Suspiciously clean grammar and structure, especially from a channel where typos and fragments are normal
- Generic phrasing that could apply to almost any company or product, with no specific names, dates, or details from the actual experience
- Uniform length and tone across many responses, even when the ratings attached to them are wildly different
- Answers that restate the question back to you before answering it, a habit LLMs default to
- A response speed that's too fast for the apparent length and thoughtfulness of the answer
- Near-identical phrasing appearing across multiple respondents who have no reason to know each other
Why This Hits Online Panels Harder Than In-Person Feedback
None of those signals is proof on its own, and plenty of genuinely thoughtful people write clean, well-organized feedback. The pattern only becomes meaningful when several of these signals cluster together across a batch of responses, which is tedious to catch by eye in a spreadsheet with a thousand rows and much easier to catch systematically. Almost everything written about this problem so far comes out of academic survey panels, market research firms, and crowdsourced platforms like Mechanical Turk, where a respondent is alone at a keyboard, unmonitored, and often being paid per completed survey.
That's the exact environment where opening a second tab and asking an AI assistant to "write a thoughtful answer to this" is frictionless. It's a different story at the point of a real-world interaction. A customer answering a kiosk survey on their way out of a clinic, or an employee tapping through a one-click pulse check between meetings, is responding in the moment, in public, without a laptop and a chatbot open.
That doesn't make in-person and one-tap feedback immune. Someone can still type a short answer with their phone's AI keyboard turned all the way up, and short-answer fields on any channel are worth watching. But the friction is real, and it's a big part of why this particular threat to data quality has grown up around long-form, unattended online surveys rather than the kiosk and QR code feedback most SurveyStance customers rely on.
What We're Building Into the SurveyStance Analyzer
We built the free survey response analyzer to help teams turn a spreadsheet of open-ended comments into sentiment and themes without paying for enterprise research software. AI-written response detection is the next piece we're adding to that same tool, so a flag for likely AI-generated or templated answers shows up alongside sentiment and theme, not as a separate upload or a separate bill. We're being careful about how we build it.
Generic AI detectors are notoriously unreliable on short text, and a tool that falsely accuses real customers or employees of not being human is worse than not having the feature at all. NORC's approach, training on real panelists and real LLMs answering the same survey questions side by side, is the right model, and it's the one we're following rather than trying to bolt a one-size-fits-all detector onto survey data it was never built for.
The Bottom Line
AI-generated survey responses are not a distant research problem. Roughly a third of participants in a recent academic study admitted to getting help from an LLM, and organizations like NORC are investing in dedicated detection because the stakes for data integrity are real. If you're running open-ended surveys today, whether through an online panel, an email link, or a kiosk in your lobby, it's worth looking at your open-text responses with this pattern in mind.
Try the free survey response analyzer on a batch of your own comments, and check back soon for the AI-response flag we're building into it.
FEEDBACK PROGRAMS
Turn this research into a live feedback program.
Compare kiosk, QR code, email signature, and OneClick feedback options built for faster response and real-time visibility.