From Customer-Led Development: How to Build What Your Customers Want When AI Can Build Anything — the operating manual for building what customers want when AI can build anything. See every chapter.
Every month, thousands of founders send their users the same question: how would you feel if you could no longer use this product? They tally the “very disappointed” answers, compare them to a benchmark, and announce they’ve found product-market fit.
Then the service goes down at 2am and nobody calls.
If you have to send a PMF survey, you don't have PMF.
That sounds glib. It isn't. It's what happens when you take one of the oldest findings in behavioral economics seriously and apply it to how product teams actually make decisions.
What people say vs. what people do
Economists call it stated preference versus revealed preference. Stated preference is what people tell you they'll do. Revealed preference is what they actually do when something is on the line. These aren't slightly different. They're not even close.
Stated intentions substantially overstate real behavior (Raassens & Haans, “NPS and Online WOM,” Journal of Service Research, 2017). In one study, donors' intended donation frequency ran 30–40% above how often they actually gave (de Corte et al., “Stated versus Revealed Preferences,” Health Economics, 2021). Doctors said they'd prescribe one thing, then prescribed another (Mark & Swait, “Using Stated Preference and Revealed Preference Modeling to Evaluate Prescribing Decisions,” Health Economics, 2004).
Psychologists have a name for the root failure: affective forecasting. Humans are reliably bad at predicting how they'll feel (Wilson & Gilbert, “Affective Forecasting,” Advances in Experimental Social Psychology, 2003). Now add the fact that a survey answer is non-binding. There's no consequence to saying “I would definitely pay for that” or “I would be very disappointed.” So people answer the way they think they should feel, not the way they do.
Rule of thumb: If you're making product decisions based on what customers say they'll do, you're building for a fictional customer base.
Why the popular surveys mislead
The NPS myth
Net Promoter Score is everywhere. Two-thirds of Fortune 1000 companies use it, according to Bloomberg (2016), all tracing back to Fred Reichheld's “The One Number You Need to Grow” in Harvard Business Review (2003). It's simple. It's memorable. It's also completely misleading.
When researchers compared NPS ratings against what people actually did, behavior often didn't line up with the categories (Harvard Business Review, “Where Net Promoter Score Goes Wrong,” 2019). Worse: 52% of people who had actively discouraged others from using a brand had also actively recommended it. The same person was both promoter and detractor.
Marketing professor Roland Rust has said NPS is “now widely discredited in the academic literature” (quoted in The Wall Street Journal, via Itamar Gilad's “Net Promoter Score: Helpful or Harmful?”). The research that launched it appeared in HBR, not a peer-reviewed journal, and the underlying data was never made public.
NPS survives not because it works, but because it's easy. It produces a number, and executives love numbers. “Our NPS went from 42 to 47” sounds like progress. Research on the measure keeps concluding it might be noise (Dawes et al., “The Predictive Validity of the Net Promoter Score,” Journal of Advertising Research, 2024).
The “very disappointed” test
Sean Ellis popularized the PMF survey: “How would you feel if you could no longer use [product]?” If 40% or more answer “very disappointed,” you have product-market fit. Three problems:
- Nothing peer-reviewed validates the 40% threshold. It was pattern-matched, not statistically tested (MeasuringU, “What Is the Product-Market Fit (PMF) Item?”).
- The threshold is arbitrary. Depending on industry and audience, a different percentage might be appropriate.
- Responders aren't representative. The people who answer are your most engaged users. Of course they'd be disappointed.
Superhuman is the most famous user of this survey, and its story shows how much the number depends on who you count. In the summer of 2017, two years into building an email client, the company still wasn't ready to launch. CEO Rahul Vohra surveyed users who had used the product at least twice in the prior two weeks: 22% said they'd be very disappointed without it. When he counted only the kinds of people who loved it most, the number rose to 33%. The product hadn't changed. Who got counted had (How Superhuman Built an Engine to Find Product/Market Fit, First Round Review, 2018).
The useful part wasn't the score. It was the written answers. Users who valued speed but were only somewhat disappointed kept asking for a mobile app and calendar features. So the team split its roadmap between making the product faster and filling those gaps. Within three quarters, the score reached 58%.
The number told Superhuman who to count. The words told it what to build.
The silent majority
Survey response rates have collapsed. Telephone polling fell from 36% response in 1997 to 6% by 2018 (Pew Research Center, Response Rates in Telephone Surveys Have Resumed Their Decline). In-product surveys aren't phone polls, but they obey the same law: the more surveys people get, the fewer they answer. In one study, college seniors who had recently received a survey answered the next one at 57%, versus 67% for seniors who hadn't (Porter, Whitcomb & Weitzer, Multiple Surveys of Students and Survey Fatigue, 2004).
So you're deciding based on a small slice of your customers. And it's not a random slice. The people who complete surveys are disproportionately the highly motivated, the ones holding extremely positive or extremely negative views (Clootrack, “How to Decode the Silent Majority in CX”). You hear from the ecstatic and the furious. The satisfied middle stays silent.
What works instead
Forget hypothetical feelings. Ask about money, and watch behavior.
The cash test
There's no better truth than cold hard cash. Ask your customers directly:
“How much would I have to pay you to stop using our product?”
“How disappointed would you be” measures feelings. A dollar amount measures economics. A customer who says they'd be “very disappointed” might accept $500 to switch. A customer who says “you couldn't pay me enough” is sticky.
Then follow up: What would you lose by losing our service? Push for specifics: revenue impact, development time, operational efficiency, competitive advantage. When customers quantify the cost of losing you, you learn what you're worth to them in dollars, not how they feel about you.
The switching cost test
Don't ask whether customers would be disappointed. Ask what it would cost them to leave.
| Real stickiness (high switching cost) | Precarious position (low switching cost) |
|---|---|
| Data lives in your system and is hard to export | Easy export |
| Workflows are built around your product | No workflow integration |
| The team is trained on your interface | Commodity features available elsewhere |
| Integrations with other tools | Interchangeable with competitors |
| Relationships with your support team | Nothing that would be missed |
iPhone users don't stay because they'd be “very disappointed” to leave. They stay because their photos, messages, apps, and contacts make leaving painful. That's revealed preference. Apple earns loyalty and engineers switching costs, and a survey can't tell you which one you have.
Read the data you already have
You already have every click, session, churn event, and expansion. Swap each hypothetical question for an observable behavior:
- Instead of “How likely are you to recommend us?” Track organic signups. Start a referral program. See who actually recommends you.
- Instead of “How disappointed would you be?” Watch what happens during downtime. Count the support tickets. Follow engagement trends.
- Instead of “Would you pay for this feature?” Build a prototype. Charge for early access. See who pays.
Then there's the qualitative side: what customers say unprompted. Use AI tools to analyze sales calls, onboarding calls, customer success check-ins, and support tickets, and hear it in their own words:
- Which objections come up again and again?
- Where do customers express frustration?
- What signals churn risk before it happens?
- What do happy customers say differently than struggling ones?
No survey required. These are conversations already happening: customers telling you the truth in real contexts, with no hypothetical questions and no predictions about future behavior.
The 2am test
One of my investors gave me the best definition of product-market fit I've heard:
“You have PMF when your service goes down at 2am and someone calls to complain.”
That's behavior, not prediction. Frustrated customers pick up the phone. They don't wait for a survey.
I know founders who are convinced they have PMF because ten people said they'd be “very disappointed.” But when their service goes down, nobody calls. Nobody complains. The disappointment evaporates the moment it requires the inconvenience of doing something.
The 2am test isn't scientifically validated either. It's a heuristic. But it's a heuristic built on what customers do, not on what they predict they'd feel.
Better questions, if you must survey
Some survey questions are less terrible than others. The difference is simple: ask about concrete past behavior, not hypothetical future feelings. Two favorites from Blanks & Jesson's Making Websites Win (2017):
- “Before you decided to go with us, what almost convinced you not to?”
- “Before picking us, which other products did you seriously evaluate, and why did you decide not to go with them?”
And replacements for “How disappointed would you be?”:
- “If you couldn't use [product] today, what would you do instead?”
- “What were you using before [product]? How does it compare?”
- “When was the last time you almost left? What happened?”
Rule of thumb: Concrete, specific, past tense. No prediction required.
AI removed the reason surveys existed
Surveys were the best available tool when we couldn't process qualitative data at scale. That constraint is gone. AI can read every support ticket, sales call, feature request, complaint, and compliment.
The old model: ask 1,000 customers a hypothetical question, get 60 responses, make decisions based on 6%.
The new model: analyze 10,000 interactions, find patterns across all of them, make decisions based on what customers actually say and do.
The silent majority isn't silent. They're talking to support, on calls, and in the product itself. You just weren't listening.
When surveys actually help
I'm not saying never send a survey. I'm saying stop using surveys as validation crutches. They earn their place in three situations:
- You need a specific answer fast. “Which of these three names do you prefer?” Quick, low-stakes, concrete.
- You're exploring a problem space, not validating a solution. Early discovery questions like “What's frustrating about X?” can surface themes. Just don't treat the responses as proof of anything.
- You're measuring change over time. Ask the same question the same way and watch the trend. The absolute number means little. The direction tells you something.
Surveys don't help when you're trying to validate PMF, deciding whether to build a feature, predicting customer behavior, or using them to feel good about your product.
The signal was always in the behavior. Now you can finally hear it.
If you want the full system for replacing stated-preference guesswork with what customers actually do, it's laid out in Customer-Led Development.
And if this changed how you'll read your next NPS report, subscribe to the Context Limit newsletter for more essays on watching what people do, not what they say.
