7 Oct 2026 · From the team
How Many Survey Responses Do You Actually Need?
A practical guide to choosing a sample size for the decision you actually need to make.
A survey with 100 responses can give you a useful read on your customers. Split those same responses across four customer groups, and you have about 25 people in each. The headline number hasn't changed, but the questions you can answer have.
I think you should choose your survey sample size around the decision you need to make, with enough responses in each group you plan to act on.
Your sample size is the number of usable responses behind a result. If 200 people start your survey but only 120 answer the question you're analysing, that result has a sample size of 120. Invitations, starts and completed responses are different counts.
For a first directional read, 50 to 100 relevant responses can be useful. Around 400 gives you a narrower estimate of a yes/no share, roughly ±5 percentage points under standard sampling assumptions. Comparing groups can require hundreds in each group. For open-ended exploration, a few dozen thoughtful answers may be enough to reveal recurring themes, though you can't assume you've heard everything.
Those are starting points. Here's how I'd choose between them.
What decision will the survey change?
Before choosing a number, finish this sentence: “Once we have these answers, we will decide whether to…”
Maybe you'll rewrite a confusing delivery page. Maybe you'll choose which product concept deserves a larger test. Or you're deciding whether customers who buy once have different needs from customers who subscribe.
These decisions demand different levels of certainty. I would accept more uncertainty when choosing which message to test next than when deciding to stop selling a product.
A useful distinction is between a directional read and a comparison:
- A directional read helps you see broad patterns and decide what to investigate or test. You can tolerate an estimate being several points off.
- A comparison asks whether one group, concept or time period differs from another. You need enough responses on both sides to separate a real difference from sampling noise.
“Most respondents seem confused by this wording” might be enough to justify a cheap rewrite. “Subscribers are six percentage points more likely to want this feature” is a much harder claim to support.
I'd write down the action and the smallest difference that would change it before opening a sample-size calculator. Otherwise, it's easy to spend money measuring something more precisely than the decision requires.
What is the margin of error for a survey?
Imagine asking several different groups of customers the same yes/no question. Even if nothing about your customer base changes, the percentage saying yes will move around. One group happens to contain more enthusiastic customers; another contains fewer.
Increasing your sample size is like taking more measurements with a slightly shaky scale. The random wobble gets smaller. But if the scale is badly calibrated, taking more measurements won't fix it. In a survey, biased recruitment or a leading question can create that calibration problem.
For a yes/no share near 50%, the rough margins of error are:
| Usable responses | Approximate margin of error at 95% confidence |
|---|---|
| 50 | ±14 percentage points |
| 100 | ±10 percentage points |
| 200 | ±7 percentage points |
| 400 | ±5 percentage points |
So if 50 out of 100 respondents say yes, the rough interval runs from 40% to 60%. That can support a different decision from an estimate between 45% and 55%.
The calculation behind the table is:
Margin of error ≈ 1.96 × √[p × (1 − p) / n]
Here, p is the observed share and n is the number of responses. Using 50% gives the widest margin for a fixed sample size. These figures are percentage points, not a percentage of your result.
The uncertainty shrinks with the square root of the sample size. To halve your margin of error, you need roughly four times as many responses. Going from 100 to 200 helps, but it doesn't make the estimate twice as precise.
What 95% confidence assumes
These calculations assume independent responses from a simple random sample, with a population much larger than the sample. In repeated sampling under those assumptions, about 95% of intervals calculated this way would contain the true population share.
Most customer surveys don't perfectly meet those conditions. People choose whether to answer your email. A panel may use targeting and quotas rather than random selection from everyone in your market. Weighting and more complex sampling designs can also change the uncertainty.
For those studies, treat the table as a planning benchmark for sampling noise, rather than a guarantee of total accuracy. It doesn't account for people misunderstanding the question or the wrong people entering the study.
If you're surveying a large fraction of a small, known population, the calculation can change too. But having a small customer list doesn't remove nonresponse bias. The customers who ignore your survey may differ from those who answer.
How many responses do you need for each segment?
Count the people in the smallest group you'll use to make a decision.
Suppose you collect 200 responses and split them evenly between first-time and repeat buyers. You have 100 per group, so each group's yes/no percentage has a rough margin of ±10 points near 50%.
The difference between those percentages has uncertainty from both groups. Under the same assumptions, its rough 95% margin is about ±14 points. A result of 54% versus 46% doesn't establish a clear difference with that sample.
This catches people because the overall count still looks comfortable. I've got 200 answers, right? Yes, but no individual segment has 200 answers behind it.
If you want roughly ±5-point precision for each of four groups, you're looking at about 400 responses per group, or 1,600 total. That assumes you actually recruit 400 into each. A total of 1,600 won't help a segment represented by only 40 people.
Comparing two groups is also a different calculation from estimating each group's share. If a five-point difference would change your decision, use a sample-size calculation for two proportions. Specify the expected baseline, the smallest difference you care about, the confidence level and the statistical power you want. Power describes how likely your study is to detect a difference of that size if it exists.
The margin-of-error table alone can't answer that question. Nor can it tell you how many responses you need for a rating average or a complex pricing study.
My default would be to reduce the number of planned comparisons before increasing the sample. Choose the one split that could change what you do. Exploring ten more splits afterwards can produce apparent differences just by chance.
How many open-ended responses are enough?
Open-ended answers have a different job. You're often trying to discover the language people use, the problems they describe, or an explanation you hadn't considered.
For a narrow question asked of a fairly similar audience, I'd expect common themes to start repeating after a few dozen thoughtful answers. That's a planning expectation, not a universal threshold established for your study.
Repetition is sometimes called thematic saturation. But hearing “shipping cost” repeatedly doesn't prove you've found every reason someone abandons a checkout. A less common problem may take longer to appear, especially when your audience contains people with very different circumstances.
I would read answers in batches and keep a simple record:
- What genuinely new theme appeared?
- Did an existing theme gain a different explanation?
- Which relevant customer groups have we barely heard from?
If each new batch mostly repeats what you already understand, you may have enough material to write better questions or choose an issue to investigate. If new explanations keep appearing, keep going or narrow the scope.
Discovering a theme and measuring how common it is require different evidence. Five detailed answers about confusing returns can give you something concrete to inspect. They don't establish what percentage of all customers find returns confusing.
And if the written answers leave you guessing about what happened, I'd move to interviews. More short answers won't necessarily resolve an unclear explanation.
A sample-size example for a DTC survey
Imagine you sell skincare and want to understand why first-time buyers haven't placed another order. This is a hypothetical study, not a Probe customer result.
Your decision is whether to test a replenishment reminder or improve product-use guidance. You could screen for people whose first order arrived long enough ago for them to have used the product, excluding anyone who has already reordered.
I'd ask about what happened before asking what they want:
“What happened with the product after it arrived?”
Then a closed question about their current situation, with options such as still using it, finished it, stopped using it and haven't started. You'd want wording and options that fit the product, plus room for another answer.
For an initial directional read, 100 usable responses could help you choose the next test. If a hypothetical 65% report that they're still using the product, the rough sampling margin would be around nine points under the standard assumptions. That gives you a broad read, while the open-ended answers help explain it. It still doesn't prove a reminder will increase repeat purchases.
Now suppose you also want separate conclusions for four acquisition channels. An evenly split sample would leave only 25 respondents per channel. I'd either drop that comparison or plan a larger study around it.
I'd also resist expanding this survey into a general brand audit. You started with one decision. Every extra objective adds questions and can create new demands on the sample.
Which survey responses should count?
I'd rather make a modest claim from a relevant sample than a confident claim from people who couldn't answer the question properly.
Start with eligibility. “Bought skincare recently” is broader than “received a first order of this product and hasn't reordered.” Screen for the behaviour your decision concerns. Avoid screener wording that makes the qualifying answer obvious.
Then make the survey answerable. Ask one thing at a time, avoid praising the product inside the question, and let people say they don't know when that's a real possibility. Keep the survey short enough that you're asking only for information you'll use.
Before launch, decide how you'll review suspicious responses. Contradictory eligibility answers, irrelevant text and implausibly fast completion can be signals. None is conclusive on its own. A brief answer may be perfectly valid, and a fast reader may still be paying attention. Attention checks should be clear rather than tricky.
Apply those rules consistently, without looking for reasons to exclude people whose opinions you dislike. Report the usable count after exclusions, and use the count for each question when some respondents skipped it.
More responses reduce random noise. They can also make a biased result look reassuringly precise.
How Probe handles survey sample size
With Probe, you can design a survey or interview study and send it to your own audience. Own-audience studies are free up to 10 questions and 25 responses.
I'd use that allowance for a small exploratory study, checking whether questions make sense, or collecting early open-ended answers. Twenty-five responses can expose a confusing question or give you something specific to investigate. I wouldn't use them for precise market-wide percentages or fine comparisons between groups.
If you need recruited participants, participants come from our recruitment partner's vetted panel. Recruited panel studies have one all-in price per study, and you see the exact price before launch. That lets you weigh the cost of more responses against the decision you're trying to make.
You still need eligibility criteria that match your question. Our guide to Screen participants for your research covers that part of setup.
Results come back as charts, themes and quotes. Those help you inspect the answers; the size and composition of your sample still limit the conclusions you can draw.
Before you launch, write down the decision, the smallest group you'll report separately, and how much uncertainty you can tolerate. If you can't explain what another 100 responses would let you do differently, pause before paying for them.