What it catches
Six questionnaire defects that logic and language alone can expose
None of these depend on what real respondents believe. They surface as soon as someone reads the question and tries to answer it, which is exactly where simulation does well.
Ambiguous wording
A question supports two readings, and different simulated respondents interpret it differently. In "Do you exercise often these days?", neither "often" nor "these days" is defined.
Double-barreled questions
One item asks about two things, such as "Are you satisfied with the price and quality?" Someone happy with the price but not the quality has no honest answer.
Incomplete or overlapping options
The choices leave out cases or overlap, as with "1-3 years" and "3-5 years", so respondents cannot find the one option that fits.
Inconsistent scale direction
Positively and negatively worded items are mixed without warning, which invites response errors and makes scoring and cleaning harder.
Broken skip logic
A respondent type is routed to a question it should never see, or lands in a branch with no exit, and part of the sample becomes unusable.
Jargon barriers
Terms the researcher takes for granted may confuse the target population. Human pretests usually involve only a handful of people, so this is easy to miss.
Limits
What simulated responses can and cannot tell you
Ask whether a finding depends only on the questionnaire itself or on the attitudes and behavior of a real population. Simulation can find the first kind; only real respondents can answer the second.
| Question | AI simulation | Real respondents |
|---|---|---|
| Is the wording clear and are the options complete? | Good fit, cheap to screen | Good fit, but small pretests miss things |
| Does every skip path work? | Good fit, can walk each path by type | Only covers paths that actually occur |
| Completion time and fatigue | Rough estimate at best | Rely on a human pretest |
| Distributions, means and proportions | Cannot estimate | The only credible source |
| Correlation and causation between variables | Not evidence | Requires a proper sample and analysis |
Method
Run a simulated survey pretest in five steps
- 01
Document the survey and its context
Prepare the full questionnaire, including skip rules, and state the research goal, target population and distribution channel. Online self-completion and interviewer-administered surveys surface different problems, so say which one applies.
- 02
Define five to eight respondent types
Cover the differences the study needs to compare, such as age group, occupation, familiarity with the topic and willingness to answer carefully. Deliberately include an unfamiliar respondent and a low-effort respondent; they expose the most problems.
- 03
Simulate answers type by type, item by item
For every answer, record three things: how the respondent read the question, what they would pick if no option fits, and whether they would want to skip it. Answers without reasoning will not reveal ambiguity.
- 04
Compile an issue list
Collect the items that several types read differently, the options that keep coming up as missing, and the skip paths that dead-end. Rank fixes by how much of the sample each one affects.
- 05
Revise and rerun
Revise the survey, rerun it with the same respondent types, confirm the old issues are gone and no new ones appeared, then move on to a human pretest.
Going further
Simulate in code so hunches become records you can check
Asking a chatbot to play a few people by hand is hard to review and makes before-and-after comparisons unreliable. Writing the simulation as a program fixes the respondent profiles, lets you rerun it, and gives you numbers to decide which items need work.
Generate respondents from structured profiles
Describe each type as parameters such as demographics, topic familiarity and response style, and let a script combine them into a batch of virtual respondents instead of improvising descriptions each time. The profile file becomes part of the pretest record.
Repeat runs on the same profile
Have each type answer several times and check whether answers hold steady. Answers that swing for the same type usually point to ambiguous wording; identical answers across every type may mean the item does not discriminate.
Measure differences across types
For each item, tally interpretation differences, how often respondents chose "Other" or asked for a missing option, and how often they flagged wanting to skip. Sort the results to surface the most suspect questions.
Walk every skip path
Encode the skip rules and trace the route each respondent type actually takes, flagging unreachable items, branches with no exit and contradictory conditions.
Scientify
Turn a survey pretest into reproducible analysis with Scientify
Scientify is a science agent that runs on an isolated cloud computer. Give it the questionnaire, the research goal and the respondent types, and the agent can write code to generate virtual respondents, run simulated responses in batches, measure how answers differ across types, suggest revisions and rerun the new version.
Code, dependencies, respondent profiles, simulated data, logs and charts all stay in one workspace, so you can compare results before and after a revision and reproduce them exactly. If a large batch takes a while, the task keeps running after you close your laptop, and you can check progress and give feedback from your phone.
- The agent, files and runtime live on an isolated cloud computer, so tasks keep running after you shut down
- Chat history and workspace files are stored only on the isolated cloud computer; Scientify servers do not store this research data
- The core agent is open source, with 2k+ stars on GitHub
- New users get a free cloud computer and $5 in model credit
Pitfalls
Common mistakes when simulating surveys with AI
- Treating simulated answers as a sample: you cannot compute proportions, means or significance tests from them, and they cannot top up a real sample.
- Collecting answers without reasoning: if virtual respondents do not explain how they read each item, ambiguity and missing options stay hidden.
- Respondent types that barely differ: renaming the same profile five times produces near-identical results and amounts to a single test.
- Skipping the human pretest: simulation screens out obvious defects, but completion time and reactions to sensitive items still need real people.
- Keeping no records: without saved profiles, versions and results, you cannot tell whether a fix worked or describe the process in your methods section.