to tell the truth
Traps · 10 of 15

The Illusion Every Tool Creates

Surveys collect politeness, focus groups collect conformity, and A/B tests collect tunnel vision — every research tool bends the truth in its own way.
Questions this piece answers
  • Why do survey answers differ from actual customer behavior?
  • How does groupthink distort focus group results?
  • What can't A/B testing tell you about a product decision?
  • How is ethnography different from surveys and focus groups?
Written forMarketers, researchers, and PMs who set product and marketing direction based on surveys, interviews, or A/B test results

How you ask the question already decides half the answer. No instrument is neutral.

This series has covered two traps so far. The gap between what customers say and what they actually want. And the gap between the loudest voices and the silent majority. Both were, at bottom, questions of whose words you listen to, and how you read them. But there is a third trap, one layer deeper. Even when you ask the right person, in the right spirit, it doesn't disappear. The instrument itself bends the answer.

Surveys distort in one direction. Focus groups distort in another. A/B tests distort in a third. None of this is researcher error. It's built into how each tool works. No matter how carefully you design around it, the distortion doesn't vanish — it can only be managed. Call the trace each instrument leaves on an answer its fingerprint. Read the data without knowing the fingerprint, and you'll mistake the pattern the tool made for the customer's true feeling.

The trouble is that every instrument's fingerprint looks different. A survey's fingerprint and a focus group's fingerprint bend answers into different shapes. Overlay the distortion from one tool onto the results of another, and the outline of the fingerprint finally becomes visible. But most research stops at a single tool. Budgets and timelines are finite, so one survey, a handful of interviews, and a conclusion follows. The fingerprint gets mistaken for the customer's face.

Why Do Surveys Lie Politely?

Surveys carry what's known as social desirability bias. Respondents tend to choose the answer that sounds right, not the one that's true. Ask "would you like more sustainable packaging?" and most people say yes. Yet far fewer of those same people actually pay more for it at checkout. Respondents aren't lying. The person filling out the survey and the person standing at the register are judging by different standards. In front of the survey, they become the person they want to be. In front of the register, they become the person who opens their wallet.

This bias grows stronger the more morally loaded the question. On topics like health, the environment, or fairness — where society has already settled on a "correct" answer — respondents report who they wish to be rather than what they actually do. The trouble is that this answer never sounds like a lie. In the moment, respondents believe it themselves. Which means the researcher has no reason to doubt it either.

The design of the answer choices quietly does the same work. "Which would you like us to improve: design, features, price, or shipping?" — there's no option for "none of the above." Respondents have no choice but to pick one. The result gets written up as "customers most want better design." In reality, a customer who wanted nothing changed may simply have picked the least objectionable box on the list. Whatever the designer failed to imagine never had a place in the data to begin with.

These two biases pull in different directions but land in the same place. The answer sheet ends up narrower than what customers actually think, and that narrow sheet becomes the basis for strategy. The results report keeps response rates and rankings. It keeps no record of whether the question led the answer, or whether some truer feeling existed just outside the given choices. The numbers look precise. But that precision only holds within the boundaries the question allowed.

A survey doesn't collect what customers actually think. It collects what's easiest for them to say.

Groupthink in the Room

Focus groups carry a different kind of distortion: groupthink. Put several people in a room, and their statements start shaping each other. Once one person voices a strong opinion, the rest feel pressure to agree, or at least to lean the same way. Opinions that would have surfaced in a one-on-one interview get amplified inside the group — or suppressed and erased. A focus group's output shouldn't be read as "customers think this." It should be read as "this particular room, on this particular day, arrived at this conclusion." In practice, though, it's treated as if it spoke for customers at large.

Add to this the expert-consumer effect. Participants know they're being studied, and from that moment on, they think harder and speak more actively than they normally would. Ask for creative ideas, and they'll pour out suggestions with real enthusiasm. But those ideas were conceived inside that room, under those conditions — there's no guarantee they match what the same person wants in front of a real shelf or a real app. This is why a feature that tested well in a focus group so often goes unused after launch.

These two biases reinforce each other. Everyone in the room is thinking harder than usual, while simultaneously adjusting what they say to read the room. What emerges sounds smooth and persuasive — but it may be a different story from what actually happens at the point of purchase, outside that room. The moment this result gets presented back in a conference room, its smoothness is easily misread as truth. The more polished the statement, the more it deserves suspicion.

Then there's the problem of speaking order. Who answers the moderator's first question is mostly a matter of chance. But that chance first answer becomes the reference point for everything that follows. Later speakers tend not to contradict it outright — they nudge and adjust around it instead. By the end of the session, the room has produced one tidy narrative. That narrative isn't the average of what everyone privately thought going in. It's closer to a polished version of whoever happened to speak first.

Selection Bias in What We Read

Analyzing reviews and social mentions runs into selection bias. The people who leave reviews, who mention a brand online, are only a slice of all customers. Nothing guarantees their traits or motivations resemble those of the buyers who stay silent. The ratio of the delighted-enough-to-review to the furious-enough-to-review shifts from category to category. Read the content of reviews without knowing that ratio, and you'll mistake a far more extreme picture than reality for genuine customer sentiment.

What makes this bias tricky is that every individual review is real. None of it is fabricated, none of it is a lie. The data simply doesn't contain the answer to whether that voice represents the whole. Stare at the review panel for as long as you like — the opinions of people who never left a review will never appear on that screen. The moment you infer what's missing from the screen using only what's on it, your judgment already stands on a distorted sample.

Review data has another trap: it's unusually easy to work with. Star ratings can be averaged. Mention frequency can be ranked. Keywords can be charted. Easy-to-handle data is easy to present in a meeting. So review analysis keeps showing up in strategy decks, regardless of how representative it actually is. The paradox is that the more polished the data looks, the more that polish conceals the limits of the sample beneath it.

The Limits of A/B Testing

A/B testing is a form of listening too, but it's never enough on its own. It tells you which of two options performs better. But when both options point in the wrong direction, A beating B still leaves you heading the wrong way overall. A test can tell you which of two button colors gets clicked more. It cannot tell you whether that button should exist on that page at all, or whether the page's structure is right to begin with.

A/B testing becomes dangerous exactly when teams forget this boundary. Because the data comes back as a clean number, it's tempting to believe that number answers the bigger question too. But A/B testing only optimizes within the choices it's given. Outside the boundary of the question, it never generates data in the first place. This is how a team ends up polishing something small to a fine edge while the larger question behind it never once makes it onto the test bench.

What makes this trap especially quiet is that A/B testing feels more "scientific" than any other tool. Statistical significance, sample size, confidence intervals — this language reads as certainty. But a statistically significant result only means the difference between two options isn't due to chance. Whether those two options came from the right question in the first place is not something statistics can answer. The language of rigor doesn't vouch for the validity of the question.

Confirmation Bias in the Numbers

Quantitative data carries its own trap. Click-through rates, conversion rates, repeat-purchase rates — these metrics tell you what happened, not why it happened. If conversion rose after a campaign, there could be several reasons. Maybe the message worked. Maybe a competitor pulled back its ads that same week. Maybe it was seasonal. What the data shows is correlation. Turning that correlation into causation is a job left entirely to the person reading it.

The tendency to lean, in that interpretive step, toward confirming a hypothesis you already held is confirmation bias. Data that supports the existing hypothesis catches the eye easily; data that contradicts it gets filed away as an exception and quietly passed over. The numbers look objective, but which of them get treated as evidence and which get dismissed as noise is decided by belief that was already there. The same bias operates just as strongly at the step of interpreting what you've heard.

Confirmation bias is dangerous precisely because, unlike the others, it doesn't live in the instrument — it lives in the person. No matter how carefully a survey is designed, how attentively a focus group is run, how rigorously an A/B test is executed, one more layer of distortion happens the moment a human reads the result. Erase every fingerprint the instruments left behind, and the interpreter's own fingerprint is still the last one on the page.

Laid side by side, each tool's limits come into focus. Surveys are wide but shallow, and they collect politeness. Focus groups run deep but get stained by the dynamics of the room. Review analysis is only open to those who chose to speak. A/B tests can't see past the choices they were given. Behavioral data shows the outcome and stays silent about the reason. Understanding these limits and cross-checking multiple tools against each other is always more accurate than trusting any single one too much.

What Ethnography Sees That Surveys Don't

All of these instruments share one root. They summon the customer into a research situation and ask a question there. That situation is artificial, and whatever answer comes out of it will inevitably differ from what happens in the moment of an actual purchase. Ethnography is the approach that touches this root directly — observing customers where they actually use the product, without asking anything at all.

In real settings, customers reveal behaviors they never mentioned in an interview, or never even noticed in themselves. What they touch first. Where they hesitate. What they skip entirely. No question was asked, so there's no reason to dress up an answer. These behaviors may be the most honest customer data available. The only price is time and money. Given what it recovers — everything the traditional instruments miss — that price is often worth paying.

Ironically, it's exactly because ethnography is slow and expensive that organizations push it to the back of the queue. For a team that has to report results within a quarter, following customers around for weeks just to watch feels like a luxury. They reach instead for the survey that delivers results in a day, the focus group that wraps up by lunch. The faster the tool, the deeper the fingerprint it leaves — and the deeper that fingerprint, the faster and more confidently the decision gets made. This is exactly where speed and accuracy start moving in opposite directions.

Ethnography doesn't replace the other tools. What sets it apart is that it's the one method that steps outside the premise every other tool shares — that an answer can only come from asking. While surveys, focus groups, and A/B tests each distort the answer in their own way, observation never demands an answer to begin with. Which means there's no answer left to distort.

Adding more tools isn't, by itself, the fix. Knowing that no tool is complete is the fix. Not locking in a direction on the strength of one survey. Not mistaking a focus group's consensus for the market's consensus. Asking, even after A beats B, whether the question behind the test was the right one at all. That second-guessing won't erase an instrument's fingerprint. But it will stop you from mistaking that fingerprint for the customer's true feeling.

In the end, the first question that gets you out of this trap isn't "what did we hear," but "how did we hear it." The moment you accept that the same customer, asked the same question, can give you a different answer on a survey than in observation, whatever certainty came stapled to that one page of data starts to thin out. That thinning can feel uncomfortable. But the discomfort itself is the sign that you've started telling the instrument's fingerprint apart from the customer's true feeling.

Every instrument leaves its own fingerprint on the answer. Read the fingerprint before you read the answer.