What Is Product-Market Fit? Definition, the 40% Test, and How to Measure It in 2026
Product-market fit is when a defined market pulls your product out of the team. The 40% disappointment test, 2026 failure data, and the signals that tell you not to scale yet.
Product-market fit (PMF) is the state in which a specific group of customers needs your product badly enough to buy it, keep using it, and tell other people about it — without you having to push. Marc Andreessen's 2007 definition still holds: "being in a good market with a product that can satisfy that market." In a strong market the customers pull the product out of the team. In a weak one, a polished product and a star team still stall.
Growth Centr publishes evergreen, research-backed analysis on go-to-market strategy for founders, marketers, and operators who need a measurement they can defend in a board meeting, not a vibe.
That distinction is expensive to get wrong. CB Insights' analysis of 431 VC-backed shutdowns since 2023 found that running out of capital is how the story ends (70%), but poor product-market fit is one of the reasons the capital ran out (43%), ahead of bad timing (29%) and unsustainable unit economics (19%). Two-thirds of those PMF failures were early-stage companies that never found a market. Twenty Series B-and-later companies cited poor fit as a primary cause — they had raised on traction that never widened. The 431 firms had raised a combined $17.5 billion. Median equity: $11 million. Median time from last raise to death: 22 months.
Key Takeaways
- Product-market fit is a state, not a GTM motion. It means a defined segment pulls the product from you: they buy, retain, and refer. Andreessen's test is still the feel; Sean Ellis's 40% "very disappointed" survey is the leading indicator.
- Ellis benchmarked nearly 100 startups: companies that later grew almost always cleared 40% "very disappointed"; those that stalled almost always sat below it. Superhuman used that survey as an engine — 22% to 58% in three quarters — as documented by First Round Review.
- CB Insights' 2026 cut of 431 post-2023 shutdowns puts poor PMF at 43%. The older 101-post-mortem series put "no market need" at 42%. Cash is the last line on the death certificate; fit is often the cause.
- Confirm fit with more than a survey: flattening retention cohorts, organic and referral share that holds when paid is paused, compressing sales cycles, and net revenue retention that does not need a new-logo hose. Benchmarkit's 2026 private median NRR is 102%; top quartile is 110%.
- Do not scale customer acquisition cost or a demand-gen machine until the 40% test, retention, and organic pull agree. A 30-day test can tell you without a reorg.
How Product-Market Fit Actually Works
Andy Rachleff, who named the idea after Don Valentine's market-first investing, put the hierarchy bluntly: when a great team meets a lousy market, market wins; when a lousy team meets a great market, market wins. Andreessen's corollary is the operator version: the only thing that matters, before you scale, is getting to product/market fit.
That is a different job from building a feature list. The market is a set of people with a recurring, expensive problem and a budget to solve it. The product is acceptable if it solves that problem well enough that they will not go back. Fit is the overlap. You can miss it in two directions: a sharp product for a market that is too small or not desperate, or a real market with a product that does not deliver the job.
Andreessen's qualitative tell is still useful as a lagging check. When fit is missing, word of mouth is quiet, usage grows slowly, reviews are "fine," the sales cycle drags, and deals die in committee. When fit is present, customers arrive faster than you can onboard them, money shows up without a discount war, and hiring lags demand. Those are outcomes. You need a leading indicator before you have investment bankers "staking out your house."
That leading indicator is a survey, plus behavior. Paul Graham's shorter phrasing — make something a small number of people want a large amount — is the segmentation rule. A 25% "very disappointed" score across a mixed audience can hide a 50% score inside one persona. The work is to find that pocket, serve it, and refuse to average it away.
Product-led growth is one way to discover fit, not a synonym for it. A PLG motion puts the product in users' hands so retention and referral can show up in the data. You can have PMF in a sales-led company. You can have a free trial and no fit. Do not confuse the distribution model with the state.
The 40% Test
Sean Ellis, after running early growth at Dropbox, LogMeIn, and Eventbrite, asked active users one question: "How would you feel if you could no longer use this product?" The answers are "very disappointed," "somewhat disappointed," and "not disappointed." After nearly 100 startups, the pattern held: teams above 40% "very disappointed" tended to grow; teams below it struggled no matter how hard they pushed marketing.
The score is (very disappointed ÷ valid responses) × 100. Ellis's sampling rule, as Superhuman applied it: poll people who used the core product at least twice in the last two weeks. About 40 responses are directional; 100-plus is a number you can take to a board. Do not survey people who signed up and bounced — they never saw the product.
First Round Review documents the Superhuman engine Rahul Vohra built on that question. In summer 2017 the score was 22%. After segmenting to the personas inside the "very disappointed" group (founders, managers, executives, business development), it read 33%. The team then split the roadmap: half doubling down on what fans already loved (speed, shortcuts, design), half removing what held "somewhat disappointed" fans of that same benefit back (mobile, integrations, calendar). They ignored users who would not be disappointed. Within three quarters the score nearly doubled, to 58%. NPS moved with it. Fundraising got easier because the metric was visible.
The same survey, run by Hiten Shah on 731 Slack users in 2015, landed at 51% "very disappointed" — already above the line when Slack had on the order of half a million paying users. Clearing 40% is not a participation trophy.
Four questions are enough:
- How would you feel if you could no longer use [product]? (very / somewhat / not disappointed)
- What type of people would most benefit from it? (fans describe themselves)
- What is the main benefit you receive?
- How can we improve it for you?
Score question 1. Build the high-expectation customer from question 2. Double down on question 3 from the "very disappointed" group. Act on question 4 only from "somewhat disappointed" users who already name the same main benefit. Everyone else is a lost cause for this version of the product.
You Have It vs You Don't
A single survey can lie if the sample is fans, or if you have not shipped the job yet. Pair it with behavior.
| Signal | Likely no PMF | Likely PMF |
|---|---|---|
| Sean Ellis score (engaged users) | Under 40%; "somewhat" dominates | 40%+ "very disappointed," and a named persona inside that group is higher still |
| Retention cohorts | Curve decays toward zero after week 4–8 | Curve flattens at a non-trivial level; new cohorts look like old ones |
| Acquisition mix | Growth dies when paid pauses; almost all new logos are outbound | Organic, referral, and inbound hold when you cut spend for two weeks |
| Sales motion | Long cycles, discounting, "come back next quarter," low win rate | Cycles compress, inbound asks for the close, win rate rises in the target ICP |
| Net revenue retention | NRR well below 100%; GRR leaking faster than expansion can cover | NRR at or above the 2026 private median (102%); expansion is not a one-off rescue |
| Word of mouth | You have to ask for reviews and intros | Users bring users; "how did you hear about us?" names a peer |
| Roadmap pressure | Every persona wants a different product | One job shows up in every fan interview; other requests are friction around that job |
| Team feel | Marketing is "the problem"; sales wants more leads | The constraint is onboarding, support, or supply — not demand |
Net revenue retention is the paid, B2B version of a flattening cohort. Benchmarkit's 2026 private-SaaS median is 102%; top quartile is 110%. Usage-based models print 108%; seat-based 98%. If NRR needs a constant flood of new logos to look fine, you are measuring a sales funnel, not fit. SaaS churn benchmarks for 2026 are the other side of the same coin: logo churn that does not flatten with tenure is a PMF problem dressed as a CS problem.
PMF vs Problem-Solution Fit vs Product-Led Growth
Three labels get stacked because they happen in sequence.
Problem-solution fit is evidence that a defined person has a painful problem and will consider a solution. Interviews, waitlists, and paid pilots can get you here. It is not PMF. People will nod at a problem and still not change how they work.
Product-market fit is evidence that your product is the thing they keep. Retention, "very disappointed," and unpaid pull. The market is specified: ICP, use case, and willingness to pay.
Product-led growth is a go-to-market motion: the product does acquisition, activation, and expansion. It is a way to generate the behavioral evidence of PMF (activation, PQLs, free-to-paid). It is not the evidence. A sales-led shop can have PMF; a PLG shop can have a leaky trial.
The failure mode is treating problem-solution interviews as permission to scale spend. CB Insights' later-stage PMF deaths are that movie: early design-partner love, a raise, then a market that never widened. Rachleff's common mistakes match the table: chasing well-known customers instead of desperate ones; iterating the product when the who is wrong; buying growth before value; and treating fit as a one-time certificate.
Metrics That Matter (and the Ones That Fake It)
The 40% score, segmented. Blended 38% with one persona at 55% is a wedge, not a miss. Blended 45% with no persona above 40% is a mushy maybe. Re-run on new users; do not re-survey the same people or you break the benchmark.
Retention that flattens. Plot weekly or monthly retention by signup cohort. If every cohort slides to zero, the product is a tour. If it levels off, you have a core who got the job done. That core is your market.
Organic and referral share. Pause paid for 14 days on one channel. If pipeline and signups collapse, you were renting attention. Fit shows up as branded search, direct traffic, and "a colleague sent this." That is the same memory demand generation is supposed to build — but demand gen cannot create a must-have that is not there.
Cycle time and win rate in the target ICP. Lengthening cycles and falling win rates while MQL volume rises is the opposite of pull. For handoff quality, see MQL vs SQL. More SQLs from a segment that does not retain is not PMF.
NRR and GRR, not vanity MRR. GrowthCentr's B2B SaaS CAC compilation puts median payback at 16 months and blended S&M at $1.30 per $1 of new ARR. If you have not cleared the 40% test, that payback clock is a bet that fit will appear after you have spent it. It usually does not.
What not to use as proof. Press, waitlist size, a single whale, conference applause, or "we closed the last five outbound deals with a discount." Those are activity. Customer lifetime value only becomes a useful ratio once people stay; a 5:1 LTV:CAC on a 60-day sample is fiction.
A 30-Day Operator Test
Do not hire a growth lead to paper over an unloved product. Run the engine Superhuman ran, at your current volume.
Days 1–7 — Sample and ask. List users who hit the core action at least twice in the last 14 days. Send the four-question survey. Target 40 responses; 100 if you have them. Write one sentence for the job the product does and one sentence for who it is not for. Freeze new paid experiments this week so the baseline is clean.
Days 8–14 — Segment, don't average. Split "very / somewhat / not." Build the high-expectation customer from the "very disappointed" answers to question 2. Recalculate the score inside that persona. Plot 4-, 8-, and 12-week retention for that persona versus everyone else. If the persona is above 40% and the rest are not, you have a wedge. If nobody is, you do not have a messaging problem.
Days 15–21 — Roadmap split. Half the next sprint: make the main benefit (question 3) sharper. Other half: remove the top friction named by "somewhat disappointed" users who already cite that benefit. Do not ship the feature list from people who would not miss you. Pause one paid channel for 14 days and watch organic and referral.
Days 22–30 — Re-score and decide. Survey a new engaged cohort. Compare the score, the persona mix, retention, and the paid-pause result. If the score crossed 40% in the wedge and retention flattened, you may scale acquisition into that wedge — not into "the market." If the score is stuck and retention is still sliding, cut spend, change the who or the job, and do not staff a demand-gen pod to fill a hole the product created.
Methodology
This is a definitional brief, not a survey we ran. The definition follows Marc Andreessen, "The only thing that matters" (2007), archived at Pmarchive, which credits Andy Rachleff. The 40% test, Slack 51% (731 users), and Superhuman 22% → 33% → 58% figures are from Rahul Vohra's First Round Review account of the Superhuman product-market-fit engine, which cites Sean Ellis's survey of nearly 100 startups. 2026 failure figures: CB Insights' analysis of 431 VC-backed shutdowns since 2023 (385 with identifiable reasons) — poor PMF 43%, capital 70%, timing 29%, unit economics 19%, $17.5B raised, median $11M, 22 months from last raise. The classic "no market need" 42% figure is CB Insights' earlier 101-post-mortem series, kept as historical context. NRR/GRR and CAC: GrowthCentr's Benchmarkit 2026 compilation. No statistic appears here unless it was on a page we fetched.
Read Next
- What Is Product-Led Growth (PLG)? Definition, Motion, and When It Beats Sales-Led
- What Is Demand Generation vs Lead Generation? Definition, Metrics, and How They Work Together
- B2B SaaS CAC Benchmarks 2026: What It Costs to Acquire a Customer
- What Is Net Revenue Retention (NRR)? Formula, 2026 Benchmarks, and What Good Looks Like
FAQs
1. What is product-market fit?
Product-market fit is the degree to which a product satisfies strong demand in a defined market. Andreessen: you are in a good market with a product that can satisfy it, and the market pulls the product out of the company. Practically: a specific segment buys, retains, and refers without heavy persuasion.
2. How do you measure product-market fit?
Use a leading indicator and lagging ones together. Leading: Sean Ellis's survey — 40%+ of engaged users would be "very disappointed" to lose the product. Lagging: flattening retention, organic/referral growth that survives a paid pause, compressing sales cycles, and NRR that does not depend on new logos. One number is not enough; Superhuman treated the score as an engine, not a trophy.
3. What is the 40% product-market-fit rule?
Ellis found that startups where at least 40% of active users would be "very disappointed" without the product tended to grow, and those below 40% tended not to, across nearly 100 companies. It is a heuristic. A 35% score rising inside a sharp persona beats a 45% score from a mixed, one-time sample. Get to ~40 responses before you react; 100 before you treat it as a board metric.
4. Can you have product-market fit without product-led growth?
Yes. PMF is a state; PLG is a distribution motion. Enterprise sales-led companies have PMF when a named ICP pulls deals through and retains. PLG is useful when users can reach value without a demo, because the retention curve becomes visible faster. A free trial is not proof of fit.
5. When should you start scaling marketing spend?
After the 40% test, retention, and organic pull agree in a defined wedge — not after a good month of outbound. CB Insights' 2026 cohort shows later-stage companies that scaled on early traction and still died of poor fit. Median CAC payback in 2026 is already 16 months; spending that clock before fit is how "ran out of capital" becomes the listed cause.
Disclaimer: This content is provided for informational purposes only and does not constitute financial, investment, or operating advice. Figures reflect publicly reported research as of August 2026, from studies with different sample frames, years, and formulas. The 40% rule is a heuristic from Ellis's startup sample, not a census of your category. Treat every benchmark as a directional peer check, not a board target without your own cohort data.