Startup case studies.
With the receipts attached.

Fourteen teardowns of how real companies found their first customers, tested demand before building, and in four cases ran out of road. Every number carries a link to where it came from, every source is marked primary or reported, and every case says plainly which of its moves you can still use in 2026 and which one belongs to a year that has closed.

Most startup case studies are written backwards. Someone succeeded, so the things they did become the reasons they succeeded, and the hundred companies that did the same things and disappeared are not available to argue. This page is an attempt at the honest version: fourteen companies, every factual claim linked to its source, each source labelled primary or reported, and a standing note on each card about what was true for those founders that is probably not true for you. Four of the fourteen are failures, because a library of winners teaches you almost nothing about what works.

1. How to read a startup case study

The format has three structural problems. None of them are fixable, and all of them are manageable once named.

Selection. Companies get written about because they survived long enough to be interesting. The same decisions, taken by companies that then folded, are not documented anywhere. When you read that Notion scrapped its codebase, laid off its team and moved to a cheaper city, you are seeing one draw from a distribution whose other outcomes were never written down. The decision looks brilliant because it worked, and we have no way to count the times it did not.

Retrospective coherence. A founder recounting 2012 in 2019 produces a cleaner story than the one they lived through. Dead ends get compressed, luck gets reinterpreted as judgement, and the order of events gets tidied into cause and effect. This is not dishonesty, it is how memory works on a narrative. It does mean the causal claims in any retrospective are the weakest part of it, and the specific facts are the strongest.

Era effects. Several of the most famous moves in startup history are not merely harder now, they are closed. Airbnb’s Craigslist cross-posting would today run into terms of service, rate limiting and legal exposure. The 2007 Digg front page and the 2012 Hacker News front page were different economic objects from anything available today. A tactic has a shelf life, and nobody puts an expiry date on the retelling.

What survives all three. The specific facts: who the first customers were, what they used before, how long it took, what it cost in hours. Those are checkable and they transfer. The lessons are the part to distrust. So this page leads with facts and sources, and marks every lesson as a judgement.

2. What the base rates actually say

Before any individual case, the denominator. Most of the numbers circulating about startup failure are either wrong or a decade stale.

The figure everyone repeats is that ninety percent of startups fail. It is not supported by the data most people think it comes from. The US Bureau of Labor Statistics tracks every private sector establishment from the year it opens, and in the most recent cohort 77.9 percent survived their first year. Of the establishments that opened in March 2015, 34.7 percent were still operating ten years later, and roughly half are gone by year five. The ninety percent number is generally traceable to venture-backed companies failing to produce a venture-scale return, which is a different question from whether a business closes its doors.

The second stale number is the reason. For years the standard citation has been a CB Insights analysis of 101 post-mortems in which "no market need" led at 42 percent, and that figure is still what most startup content quotes. CB Insights published a far larger update in March 2026, covering 431 venture-backed companies that shut down since 2023, which between them had raised 17.5 billion dollars, at a median raise of 11 million. The ranking changed.

What is still being repeated
  • 90 percent of startups fail, often stated as failing in the first year.
  • No market need, 42 percent, from a 101-company sample now several years old.
  • Running out of money is a cash problem, to be solved by raising more.
What the current sources say
  • 77.9 percent survive year one; about 34.7 percent reach year ten. (BLS)
  • Ran out of capital, 70 percent. Poor product-market fit, 43 percent. Bad timing, 29 percent. Unit economics, 19 percent. From 431 post-mortems, March 2026. (CB Insights)
  • CB Insights says capital is the final cause, not the root one. The other three are why the capital stopped.

Two things follow. First, the base rate for a business surviving is much better than founder folklore suggests, and the base rate for a venture-backed company returning a fund is much worse. Those are separate games with separate odds, and conflating them is how founders end up optimising for the wrong one, which is exactly the trap Gumroad’s founder describes below. Second, the percentages add to well over a hundred because failures are multi-causal. No company in that sample died of one thing.

3. How this library was built

Stated so you can judge it, and so you can tell which parts to distrust.

  • Inclusion. A case had to have a publicly documented account of a specific early decision, not just an outcome. Companies with famous trajectories and no documented early mechanics were left out, however well known.
  • Source tiering. Each case is marked primary where the strongest evidence is the founder’s own account, and reported where it is journalism or third-party analysis. 7 of the 14 rest on primary sources.
  • Numbers. Figures are quoted only where a credible source states them. Where sources disagree, the disagreement is stated rather than averaged, which is why Olive AI’s funding total appears as a range.
  • No invented detail. Where a widely circulated number has no credible origin, it is omitted and flagged. Stripe’s "first twenty customers" is one of these.
  • Failures included by design. Four of the fourteen did not work, and one is a near-death recovery. A library of survivors cannot distinguish a method from a coincidence.
  • What this is not. None of these companies is a Cafiyn customer and none of these numbers is ours. This is a public-record teardown. Cafiyn appears in one section, clearly marked, about running the same method on your own market.

4. Part one: how six companies got their first customers

The common thread is not a channel. It is that somebody did something by hand that could not have continued at scale, and did it for longer than felt reasonable.

Read these for the mechanics rather than the inspiration. In every case the interesting detail is the one that sounds like a chore: the laptop handed across a table, the eight months of forum replies before a single inbound, the flight to New York with a camera. The parts that sound like strategy are mostly reconstruction.

Stripe

Payments API, 2010 to 2011
By handevery early install
What they sold
A payments API, to developers who already had a product and could not get a merchant account without weeks of paperwork.
The move
The founders did not send documentation. When someone at Y Combinator said they might try it, Patrick or John Collison would ask for their laptop and integrate Stripe there and then. Paul Graham named the practice the "Collison installation" and used it as the central example in his essay on doing things that do not scale.
The evidence
Graham, who ran the accelerator both founders went through, describes the behaviour first hand. The number of installs done this way is not published anywhere credible, so treat any specific figure you see as invented.
Replicable
Closing the gap between stated interest and working integration yourself. For any developer tool, the distance between "that sounds useful" and "it is running in my codebase" is where almost all early interest dies, and in the first few months you can simply carry people across it.
Not replicable
The population. Stripe had a dense, pre-qualified, physically co-located set of technical founders who all had the exact problem. Most founders do not, and manufacturing one is a years-long project.

Zapier

Integration platform, 2011 to 2012
8 monthsto first inbound request
What they sold
Connections between SaaS tools, to people who wanted two products to talk to each other and had no engineer to ask.
The move
Wade Foster went into the support forums of Evernote, Salesforce, Dropbox and others, found the threads where users were asking for an integration that did not exist, and answered each one with an explanation and a link to an early Zapier demo. Each link brought roughly ten to twenty visitors. Around half of them became beta users. He then set the integrations up manually on Skype calls, one customer at a time, including for Mixergy founder Andrew Warner.
The evidence
Widely reported and consistent across several accounts drawing on Foster interviews. The beta carried a one hundred dollar one-time fee that later dropped to ten dollars. It took roughly eight months of outbound effort before a single request arrived unprompted.
Replicable
Going to where the problem is already described in public, in the buyer’s own words, and answering the specific person who described it. Support forums, review sites, community threads and issue trackers are all still full of these.
Not replicable
Nothing much. This is the most portable move in the whole library, and the least glamorous. The constraint is patience: eight months of manual work before the first inbound is longer than most founders will tolerate.

Segment

Customer data platform, December 2012
25 to 1,000+github stars in a day
What they sold
One JavaScript library that sent analytics events to many destinations, to engineers tired of adding a new snippet for every tool.
The move
Segment began as ClassMetric, a tool telling lecturers when a class was confused. It did not work, in part because the students opened Facebook instead. The team spent about a year trying to build a differentiated analytics product and failed to find an angle. With the money nearly gone, co-founder Ian Storm Taylor suggested open-sourcing a small internal library. He posted it to Hacker News on 12 December 2012 as "Show HN: Analytics.js". It reached the front page, collected several hundred points, and GitHub stars went from roughly twenty-five to over a thousand within the day.
The evidence
Peter Reinhardt has told this story in detail for Y Combinator, and the Hacker News thread and GitHub history are both still public and checkable.
Replicable
Shipping the by-product. The library was infrastructure the team had built for itself and did not consider the product. If you have built an internal tool to survive your own problem, other people with that problem may want it more than they want your actual product.
Not replicable
The 2012 Hacker News front page. The audience was smaller, less saturated, and far more likely to convert into paying infrastructure customers than it is now. A Show HN post is still worth the fifteen minutes, but the expected value has fallen a great deal.

Loom

Video messaging, 2015 to 2016
$600revenue in 7 months, pre-pivot
What they sold
Screen-and-face recording in a browser extension, to people who needed to explain something that did not justify a meeting.
The move
The company started as Opentest, a usability-testing marketplace. It made around six hundred dollars in seven months and the founders were about two weeks from shutting down. Then they noticed a Harvard research team using the Chrome extension in a way they had not designed for: the researcher recorded himself summarising the feedback he had gathered, to send to his team. They rebuilt around that single behaviour, launched it as OpenVid on Product Hunt, finished first for the day with roughly 2,500 downloads, and reached more than 12,800 users inside 100 days. OpenVid became Loom.
The evidence
Reported consistently across several accounts based on founder interviews. The revenue and user figures come from those retellings rather than from filings, so treat them as approximate.
Replicable
Watching for the unintended use. The signal was not in a survey or a roadmap vote: it was one user doing something off-label because it solved a problem better than the thing they came for. That is visible in session recordings, support tickets and onboarding calls, and it is almost always under-investigated.
Not replicable
Product Hunt finishing first as a growth event. It produced real distribution in 2016. Today it mostly produces a one-day spike among other builders, which is a different and much weaker thing.

Airbnb

Marketplace, 2008 to 2010
Door to doorhost onboarding
What they sold
Short stays in other people’s homes, to travellers who did not yet believe that was a normal thing to do.
The move
Two moves, and only one of them is the famous one. The famous one is technical: Airbnb built a way to cross-post listings to Craigslist, which had no public API, and harvested Craigslist listers directly. The more instructive one is manual: the founders flew to New York, knocked on hosts’ doors, and photographed the listings themselves, because the listings were failing on photography rather than on demand.
The evidence
Graham documents the door-knocking and photography first hand in his essay. The Craigslist integration is extensively reported and the mechanics are well established.
Replicable
Going to the supply side in person when a marketplace is not converting, and fixing the asset quality yourself before concluding the market is absent. "The listings look bad" is a diagnosis you can only reach by looking.
Not replicable
The Craigslist arbitrage. That specific gap closed, and the general pattern of scraping and cross-posting onto someone else’s platform now runs into terms of service, rate limiting and legal exposure that did not meaningfully exist in 2009. Do not read this as a tactic.

Linear

Issue tracking, 2019
Invite onlythrough its first year
What they sold
An issue tracker, to software teams who found the incumbent slow and were willing to switch tools mid-project.
The move
Linear launched as a private beta in 2019 and stayed invite-gated, admitting companies from a waitlist rather than opening up. Over that period it took on hundreds of companies, including Render, Curology, Compound and a number of Y Combinator startups. Sequoia’s Stephanie Zhan has said the company first came onto her radar when that private beta launched, and a Sequoia scout had already invested earlier in the year.
The evidence
Reported by TechCrunch at the time of the 4.2 million dollar seed round in November 2019, including the Sequoia account of how the company was found.
Replicable
Using a gate as a quality filter rather than as marketing. Admitting teams in small batches let a three-person company give real attention to each one and keep the product coherent while the core abstractions were still moving.
Not replicable
The scarcity effect on its own. A waitlist with nothing behind it produces a list, not demand. This worked because the founders had built the tools they were replacing, at Airbnb, Coinbase and Uber, and the product was visibly better on the first screen.

What to take from part one. Four of these six first channels were places where the problem was already being described by the people who had it, in their own words, in public. Not one of the six started from a channel plan. They started from a complaint, and went to where it was being made.

5. Part two: how four companies tested demand

Three of these four spent days, not months, on the test. The fourth spent four years, and is the case most often misused to justify doing so.

Validation is the stage founders most want to skip and most often fake. The pattern worth noticing across these four is the difference between testing curiosity and testing intent. An email capture measures the first. A price, a survey with a forced choice, or an observed behaviour measures the second, and only the second predicts anything.

Dropbox

File sync, 2007
5k to 75kwaitlist, overnight
What they sold
File sync across machines, to technical people who had already tried and abandoned every existing option.
The move
Drew Houston had a prototype that was not ready to launch and wanted to know whether anyone else felt the problem. He made a short screencast of the product working, and seeded roughly a dozen in-jokes aimed specifically at the Digg audience, including references to Chocolate Rain, Office Space and XKCD. The video was the minimum viable product. It passed ten thousand Diggs within a day and the beta waiting list went from about five thousand people to seventy-five thousand overnight.
The evidence
Houston has described this directly, including the deliberate Easter eggs and the waitlist numbers, in interviews from 2011 onward.
Replicable
Testing the demand for the outcome before finishing the thing that delivers it, and tailoring the artefact to one specific community rather than to a general audience. The Easter eggs are the real lesson: the video worked because it was made for Digg and nobody else.
Not replicable
The channel and the format. A product demo video is now the default, not a novelty, and no equivalent of the 2007 Digg front page exists. The transferable version is a landing page, a short video, or a working prototype aimed at one named community you already understand.

Buffer

Social scheduling, 2010
120signups in 7 weeks
What they sold
Scheduled social posting, to people already posting manually and losing the habit.
The move
Joel Gascoigne built a two-page site before writing the product. Page one explained the benefit. Page two showed pricing. He tweeted the link and asked people what they thought. Enough people clicked through to the five and ten dollar plans to suggest a willingness to pay, which was the actual question. Over seven weeks he collected about 120 signups, spoke to a large share of them, and when he launched, roughly fifty started using it. The first payment arrived three days later.
The evidence
Gascoigne wrote the method up himself at the time, and Buffer published its own account of going from idea to paying customers in seven weeks. Both are still available.
Replicable
All of it, and this is the most directly copyable case in the library. The important detail is the second page. A page collecting emails tests curiosity; a page showing a price tests intent, and the two are not correlated.
Not replicable
The headline numbers as a benchmark. 120 signups in seven weeks is a small number, and it was enough because he talked to them. The conversations, not the count, were the validation.

Superhuman

Email client, 2017 to 2019
22% to 58%very disappointed
What they sold
A faster email client, at a price point nobody believed an email client could carry.
The move
Rahul Vohra took Sean Ellis’s product-market fit survey, which asks users how they would feel if they could no longer use the product, and turned the share answering "very disappointed" into the company’s primary metric. Ellis’s benchmark is that products above roughly forty percent tend to grow and products below it tend to struggle. Superhuman started at twenty-two percent. Rather than treat that as a verdict, Vohra segmented the respondents, ignored the people who would not be disappointed, and built only for the ones who already loved it while fixing the specific blockers named by the people who were merely somewhat disappointed. The score reached fifty-eight percent.
The evidence
Vohra published the full method, including the survey wording and the segmentation, in First Round Review. It is the most replicable artefact on this page.
Replicable
The method, exactly as written. It costs one survey and an afternoon of reading, and it converts an unanswerable question into a number you can move. It works at any size above roughly forty respondents.
Not replicable
The starting conditions. Superhuman had a waitlist and a concierge onboarding call for every user, which is what made the qualitative follow-up possible. Run the survey anyway: the number is useful even without the onboarding apparatus.

Figma

Design tool, 2012 to 2016
~4 yearsbefore general release
What they sold
Collaborative design in a browser, to design teams whose incumbent tool was a desktop application with files.
The move
Figma was founded in 2012 and did not release publicly until September 2016, with a preview in December 2015. It raised both a seed and a Series A with no public product and no user metrics to show. The time went into the technical bet underneath, which was rendering a professional design tool in a browser well enough that professionals would accept it.
The evidence
The timeline is well established. Dylan Field has discussed building and marketing during the stealth period in interviews.
Replicable
Matching the validation window to the technical risk. If the hard part is whether the thing can exist at all, shipping early does not answer the question, and a long build is the correct answer rather than a failure of discipline.
Not replicable
Almost everything else, and this is the case most often misapplied. Figma is the exception that gets quoted to justify years of building without customer contact. The distinguishing feature is that the risk was genuinely technical and genuinely resolvable. If your risk is whether anyone wants it, four years in stealth is not courage.

What to take from part two. The cheapest reliable test in the set is Buffer’s second page: show a price and see who clicks. The most reusable artefact is Superhuman’s survey, which is one question and a segmentation rule and costs an afternoon. Figma is the case to handle with care, because its four years in stealth were buying down a technical risk, not avoiding a demand question.

6. Part three: four failures, up close

Including one company that nearly died and did not, because the decisions in that case are the same shape as the ones in the three that ended.

Failure accounts are rarer and more useful than success accounts, for the obvious reason: nobody is incentivised to write them, so the ones that exist tend to be unusually candid. Two of these four have reasonably detailed public records. One has a founder’s own essay, which is close to unique. One, Notion, survived, and is included precisely because the survival makes the selection problem visible.

Fast

One-click checkout, 2019 to April 2022
$600krevenue, $120M+ raised
What they sold
A one-click checkout button, to online merchants, in a market where Shopify and Amazon already owned the checkout.
The move
Fast raised more than 120 million dollars, including a 102 million dollar round led by Stripe, and hired hundreds of people. In its final full year it generated roughly 600,000 dollars of revenue. It shut down in April 2022. Reporting at the time described heavy spending on brand-building, including sports team partnerships, while the core distribution problem went unsolved.
The evidence
NPR reported the shutdown, the funding total and the revenue figure in April 2022, drawing on employee accounts.
Replicable
Nothing to copy. The lesson is diagnostic: a revenue-to-funding ratio this extreme is visible from inside long before the end, and it is the single most legible warning sign in the library.
Not replicable
The framing that this was a market timing problem. One-click checkout was a real category and competitors in it survived. The gap was distribution into merchants who already had a checkout they were not unhappy with.

Olive AI

Healthcare automation, to 2023
>$800Mraised before winding down
What they sold
Administrative automation, to US hospitals, against a backdrop of legacy clinical and billing systems.
The move
Olive AI reached a reported four billion dollar valuation and raised a sum reported variously between roughly 830 and 856 million dollars, depending on the source. It laid off 450 people in July 2022 and 200 more the following February, and wound down in 2023, selling remaining assets. Reporting points to three causes: capability claims that outran what was delivered, including a widely repeated claim about reducing administrative costs several times over; integration with hospital legacy systems proving far harder than modelled; and expansion across too many products at once.
The evidence
Reported across trade and general press. Note that the funding total differs between accounts, which is itself worth knowing: when a number varies by tens of millions across sources, do not quote it to the dollar.
Replicable
Nothing to copy. The diagnostic value is in the integration point. If your product has to work inside systems you do not control and cannot test against before a contract is signed, your real delivery risk sits outside your own roadmap.
Not replicable
Reading this as an AI story. The failure mode is enterprise integration and overclaiming, and it predates the current generation of AI products by decades.

Gumroad

Creator payments, 2011 to 2019
75%of staff laid off
What they sold
A way for creators to sell a file directly, to people who did not want to build a store.
The move
Sahil Lavingia left Pinterest as its second employee, raised around eight million dollars in the first year, and built toward a billion dollar outcome that did not arrive. He laid off about seventy-five percent of the team, including friends, considered selling, and eventually ran the company largely alone. He then rebuilt it as a small, profitable business and began buying investors out. He wrote the whole thing up publicly in 2019 under the title "Reflecting on My Failure to Build a Billion-Dollar Company".
The evidence
The founder’s own published essay, which is unusually specific about the financing structure and the decisions. One of the few genuinely primary failure accounts available.
Replicable
The re-scoping. The company was failing against a venture outcome and succeeding as a business, and those are separate questions. Knowing which one you are being measured against, and by whom, changes what counts as a good decision.
Not replicable
The sequence, for most people. Raising eight million dollars first and discovering the true size of the market afterwards is the expensive order to do it in, and the essay says so.

Notion

Workspace software, 2013 to 2016
2015the year it nearly ended
What they sold
A single tool combining documents, wikis and databases, to teams running on four separate products.
The move
Three years in, Notion was out of money and crashing constantly on a technology stack that could not carry the product. Ivan Zhao and Simon Last laid off their only colleagues, scrapped the codebase, and moved from San Francisco to Kyoto, where living costs were less than half. They rewrote the core from scratch around a single block-based model that treated text, databases and wikis as one underlying structure, and shipped version 1.0 in 2016.
The evidence
Reported with direct Zhao quotes, including his account of letting the early team go being the hardest part and permanently changing how he hires.
Replicable
Rewriting rather than patching when the architecture is the constraint, and cutting burn hard enough to buy the time to do it. Notion did not get a better growth plan in 2015, it got a different data model.
Not replicable
Survival as the expected outcome. This is the clearest survivorship problem in the library: the same three decisions, taken by companies that then failed, are not written up anywhere. We can see the version that worked and not the distribution it came from.

What to take from part three. None of these four died of a bad idea. Fast had a real category with surviving competitors. Olive AI had a real administrative burden to automate. Gumroad had, and still has, a working business. The failures were in distribution, in delivery against claims, and in which game the company thought it was playing. That matches the CB Insights ranking: capital runs out because something upstream of capital was wrong.

7. The patterns across all fourteen

Counted rather than asserted. Fourteen is far too small a sample to prove anything, which is the point of showing the counts.

PatternCountWhat it looks like
The first customers were acquired manually6 of 14Stripe, Zapier, Airbnb, Linear, Buffer and Superhuman all did work in the first year that could not have continued at ten times the size: hand installs, Skype setup calls, door knocking, onboarding calls for every single user.
The channel was a place the problem was already described in public4 of 14Zapier in support forums, Segment and Dropbox in communities they understood intimately, Loom in an observed off-label use. None of them started from a channel strategy. They started from where the complaint was.
Demand was tested before the product was finished3 of 14Dropbox with a video, Buffer with a pricing page, Segment with a library it already had. In each case the artefact cost days, not months.
The company pivoted, keeping the same customer3 of 14Segment, Loom and Notion all changed what they shipped without changing who they were shipping to. The three cases where the pivot held are all of this kind.
The founders had the problem themselves8 of 14Stripe, Segment, Dropbox, Buffer, Superhuman, Figma, Linear and Notion. Useful, and frequently overstated: it tells you the problem is real for one person, not that a market will pay.
The early numbers would look like failure on a dashboard5 of 14120 signups in seven weeks. Six hundred dollars in seven months. A twenty-two percent product-market fit score. Eight months of outbound before one inbound reply. Every one of these preceded a working business.

The last row is the one worth sitting with. Every single one of those numbers, seen on a dashboard in isolation, would read as a failing company. A hundred and twenty signups in seven weeks. Six hundred dollars in seven months. Eight months of outbound with no inbound. A product-market fit score of twenty-two percent. Four of the five preceded a company worth more than a billion dollars, and the fifth preceded a profitable business its founder chose on purpose.

That is not an argument for patience in general. Plenty of companies posted numbers like these and then died, and they are not in this library because nobody wrote them up. It is an argument for knowing which number is supposed to be moving at your stage, and measuring that one honestly, rather than reading a small absolute figure as a verdict.

8. First-customer channels, scored for 2026

Every channel these fourteen companies used, with an honest note on whether it still works.

ChannelSeen inEffortCeilingWhere it stands now
Direct outreach to people who described the problemZapierHighHighStill the most reliable first channel in B2B, and the one most improved by better targeting. The work is research, not volume.
Founder-installed onboardingStripe, SuperhumanVery highMediumStill works and still underused. Caps out around the first hundred customers, which is exactly the range where it matters.
Open-sourcing an internal toolSegmentMediumHighStill viable, now crowded. Works when the tool solves a problem that is obviously annoying and obviously narrow.
Community launch (Hacker News, Reddit, niche forums)SegmentLowVery high, very rareCheap to attempt, far lower expected value than in 2012. Worth the fifteen minutes; not worth a plan.
Product HuntLoomHighLow to mediumA one-day spike among other builders, a backlink and a badge. Rarely a customer channel for vertical B2B.
Demo artefact seeded to one communityDropboxMediumMedium to highThe format is commoditised, the principle is not. Tailoring to one named community still outperforms general polish.
Priced landing page before buildBufferLowDiagnostic onlyStill the cheapest honest demand test available. Show a price or you are testing curiosity.
Invite-gated betaLinearMediumMediumWorks as a quality filter when there is something worth gating. Produces nothing on its own.
Marketplace supply-side fieldworkAirbnbVery highHighStill correct for marketplaces, still mostly skipped in favour of demand-side spend.
Press and PRNone of the fourteenVery highLowNot one of these fourteen companies traces its first customers to press coverage. Worth sitting with.

The final row is the finding most likely to be useful and least likely to be welcome: not one of these fourteen companies traces its first customers to press coverage, and press is the single most effort-intensive item on the list. Founders routinely spend weeks on it. The same weeks spent on the first row of this table have a documented track record in this sample and press does not.

If you want the longer version of this argument with the channel mechanics, the launch guide ranks the same channels by what they return for a B2B product, and the founder rulebook covers the first-customer chapter in more depth.

9. What is replicable, and what belongs to its year

The sorting almost no case study does, because doing it reduces the apparent value of the advice.

Still works, essentially unchanged

  1. 1
    Answer the public description of the problem, one person at a time.

    Zapier’s forum method. Support forums, review-site complaints, community threads and issue trackers are all still full of people describing your problem in their own words. The work is research and reply, and it does not expire.

  2. 2
    Show a price before you build.

    Buffer’s second page. Email capture measures curiosity; a price measures intent. This test costs a day and most founders still skip it.

  3. 3
    Run the forty percent survey and segment the answers.

    Superhuman’s method, published in full. One question, roughly forty responses, and a rule for whose feedback to ignore. The highest-value-per-hour artefact in the library.

  4. 4
    Install the product for people yourself.

    The Collison installation. Caps out in the low hundreds of customers, which is precisely the range where founders need it most.

  5. 5
    Go and fix the supply side in person.

    Airbnb’s photography, not its Craigslist integration. If a marketplace is not converting, look at the asset quality before concluding the demand is absent.

  6. 6
    Investigate the off-label use.

    Loom’s researcher recording a summary for his team. Someone using your product wrong, because it solves a different problem better, is the highest-signal event in your analytics.

Degraded, closed, or dependent on circumstances you do not have

  • Cross-posting arbitrage onto another platform. Airbnb and Craigslist. Terms of service, rate limiting and legal exposure now make this a liability rather than a tactic.
  • The community front page as a business event. Digg in 2007, Hacker News in 2012. Still worth fifteen minutes, no longer worth a plan. The expected value has fallen by an order of magnitude.
  • Product Hunt as an acquisition channel. It delivers a one-day spike among other builders, a backlink and a badge. Those are real and they are not customers for most vertical B2B products.
  • A dense pre-qualified network on tap. Stripe’s early users came through Y Combinator. If you do not have an equivalent, the install tactic still works but the pipeline into it does not exist and has to be built separately.
  • Years in stealth. Figma earned it by carrying a real technical risk. Used to avoid a demand question, it is the most expensive mistake available.

10. What this page cannot tell you

Four limits, stated plainly, because a case study library that claims more than this is selling something.

The sample is fourteen, and it is not random. These companies were selected because their early mechanics are documented, which correlates strongly with having survived long enough for someone to ask. Every count in section seven is a count within that biased sample. None of them is a rate.

The failures that match the successes are missing. Somewhere there is a founder who ran the forum-reply method for eight months, got nothing, and stopped. Somewhere there is a team that rewrote its codebase in a cheap city and still ran out of money. Those cases are the control group, they are the thing that would tell you whether any of this is a method, and they are almost entirely unwritten.

Era effects are not fully priced. Section nine sorts the moves as honestly as we can, but the judgement calls in it are ours, not data. Some of the things in the "still works" column will have degraded by the time you read this.

None of this is about your market. Fourteen companies in payments, analytics, design, email, file sync, social scheduling, issue tracking, travel, healthcare administration and creator payments do not add up to information about whether your specific buyer, in your specific segment, has your specific problem badly enough to pay. That question is only answerable against your own market, which is the next section.

11. Running the same teardown on your own market

Where Cafiyn fits, marked clearly so you can skip it.

Strip the narrative out of the fourteen cases above and the same four questions sit underneath all of them. Who precisely has this problem. What are they using instead. Where do they describe it in public. What would have to be true for them to switch. Every company in this library answered those by hand, slowly, and several of them nearly died doing it.

Two of those four questions are research. Cafiyn Lens is the structured version: an opportunity assessment that maps the buying committee, sizes the segment, tears down the alternatives people are actually using, and returns a scored verdict rather than a pile of tabs. It is the same work Airbnb did by flying to New York and Zapier did by reading forums, with the synthesis done for you. Lens starts at 14.99 dollars a month for five assessments.

The other two are outreach. Cafiyn FlyWheel is the part that goes and talks to the people the assessment named: researched, verified, personalised per account, from your own deliverability-managed infrastructure, priced in Wheels rather than seats. One Wheel is one target account through the full workflow, from 29 dollars a month. It is the Zapier forum method at a volume a founder cannot sustain by hand, which is the honest description of what it does and does not do.

The connection between them is the point: the Blueprint is the shared model both products read from and write to, so what the outreach learns revises the next assessment. That loop is the thing none of the fourteen companies had. They ran it manually, in their heads, over years. If you want the longer argument for why the loop matters more than either half, the go-to-market engine page makes it.

Nothing here requires a tool. The template in the next section is the manual version, and it is free. Several companies in this library succeeded with nothing more than a spreadsheet and an uncomfortable amount of persistence.

See pricing

12. The teardown template

Nine questions to run against any case study you read, including this one. If an account cannot answer five of them, it is a story.

  1. 1
    Who exactly were the first ten customers, by name or role?

    If the account cannot answer this, it is a story about a company, not a case study. Vague first customers usually mean the real answer was a network the writer does not want to foreground.

  2. 2
    What did those customers use the day before?

    Every first customer switched from something, including from doing nothing. The thing they left is the real competitor, and it is usually a spreadsheet or an intern.

  3. 3
    Which single channel produced them, and what did it cost in hours?

    Accounts list channels. They rarely price them. A channel that worked at forty hours a week is a different recommendation from one that worked at four.

  4. 4
    How long between starting and the first unprompted inbound?

    This number is the honest measure of whether anything compounds yet. Zapier took roughly eight months. Most case studies omit it entirely.

  5. 5
    What was true about the founders that is not true about me?

    Accelerator network, existing audience, domain employment, technical depth, capital, geography. Write the list down before taking any advice from the case.

  6. 6
    What was true about the year that is not true about this year?

    Platform gaps, channel saturation, acquisition cost, regulatory posture. Several moves in this library are not merely harder now, they are closed.

  7. 7
    What is the source, and is it the founder or a retelling?

    A founder’s retrospective is biased toward coherent narrative. A retelling of a retrospective adds a second layer. Check which one you are reading.

  8. 8
    What did they stop doing?

    Almost never recorded, and often the actual decision. Notion stopped maintaining a codebase. Segment stopped trying to differentiate in analytics.

  9. 9
    Who else did the same thing and failed?

    The one question no single case study can answer, and the one that determines whether you are looking at a method or at luck with a narrative attached.

The ninth question is the one that matters and the one no single case study can answer. It is also the reason this page counts its patterns instead of asserting them, and the reason four of the fourteen cases are failures. A method is something that works across a distribution. A case study is one draw from it.

For the operational version of these questions applied to your own company rather than to someone else’s, the founder rulebook has the ICP and validation chapters, and the free tools cover the stack and cost side.

13. Frequently asked questions

Direct answers, each one standing on the sources listed below.

What is a startup case study?

A startup case study is a structured account of how a specific company handled a specific problem: finding its first customers, validating demand, pricing, pivoting, or failing. A useful one names the customers or their roles, names the channel, gives the timeline, cites where the information came from, and separates what is replicable from what depended on circumstances the reader does not share. An account that gives only the outcome and a lesson is a story, not a case study.

What percentage of startups actually fail?

Far fewer than the commonly repeated ninety percent. US Bureau of Labor Statistics data on private sector establishments shows about 77.9 percent survive their first year, and roughly 34.7 percent of the establishments that opened in March 2015 were still operating ten years later. Around half are gone by year five. The ninety percent figure is usually traced to venture-backed startups failing to deliver a venture-scale return, which is a different question from whether a business closes.

What is the most common reason startups fail?

CB Insights analysed 431 venture-backed companies that shut down since 2023, published in March 2026, and found running out of capital cited in about 70 percent of cases, poor product-market fit in 43 percent, bad timing or macro conditions in 29 percent, and unsustainable unit economics in 19 percent. CB Insights itself notes that running out of capital is almost always the final cause rather than the root one. Note that most startup content still quotes the older figures from a 101-company sample, where "no market need" led at 42 percent.

How do most startups get their first customers?

Manually, and from places where the problem was already being described in public. Across the fourteen teardowns on this page, six acquired their first customers through work that could not have scaled: hand installations, setup calls, door knocking, per-user onboarding. Four found them in forums, communities or observed behaviour rather than through a channel strategy. None of the fourteen traces its first customers to press coverage.

How did Stripe get its first customers?

By integrating the product for people on the spot. Rather than sending documentation to developers who expressed interest, the Collison brothers would ask for the prospect’s laptop and install Stripe there and then. Paul Graham named this the "Collison installation" and used it as the main example in his essay on doing things that do not scale. The early users came largely from the Y Combinator network, which is the part of the story that does not transfer.

How did Dropbox validate its idea before building the product?

With a screencast. Drew Houston recorded a short video of the prototype working and seeded roughly a dozen in-jokes aimed specifically at the Digg community, including Chocolate Rain and Office Space references. The video, not the software, was the minimum viable product. It passed ten thousand Diggs in a day and the beta waiting list went from about five thousand to seventy-five thousand overnight. The transferable part is tailoring the artefact to one named community; the Digg front page itself no longer exists.

What is the 40% product-market fit test?

It is Sean Ellis’s survey question: ask users how they would feel if they could no longer use the product, offering "very disappointed", "somewhat disappointed" and "not disappointed". The share answering "very disappointed" is the score, and roughly forty percent is the threshold above which products have tended to grow. Rahul Vohra built Superhuman’s development process around it and published the full method, taking the company from twenty-two percent in 2017 to fifty-eight percent by segmenting respondents and building only for the group that already loved the product.

How many customers do you need to validate a startup idea?

Fewer than most founders assume, and the number matters less than whether you spoke to them. Buffer launched on about 120 signups gathered over seven weeks, and the validation came from conversations with a large share of them rather than from the count. For the forty percent product-market fit survey, roughly forty responses is usually enough for the score to mean something. What does not work is a large number of signups nobody has talked to.

Can you copy a startup case study directly?

Parts of some of them. The portable moves in this library are the priced landing page, the product-market fit survey, answering public descriptions of the problem one by one, and founder-run onboarding. The moves that are closed or badly degraded are the Craigslist cross-posting arbitrage, the 2012 Hacker News front page, the 2007 Digg front page, and Product Hunt as a primary acquisition channel. Every case here states which of its moves fall into which category.

Why are startup case studies unreliable?

Three reasons, all structural. Selection: only companies that survived long enough to be written about get written about, so the same decisions taken by companies that failed are invisible. Retrospective coherence: founders recounting events years later produce a cleaner causal chain than the one they experienced. Era effects: a channel that worked in 2012 may be saturated, closed or illegal now. None of these make case studies useless, but they do mean a single case can never establish that a method works.

What is the difference between a primary and a reported startup case study source?

A primary source is the founder’s own written or recorded account: Joel Gascoigne on Buffer’s landing page, Rahul Vohra on Superhuman’s survey, Sahil Lavingia on Gumroad. A reported source is journalism or third-party analysis drawing on interviews. Primary sources are more specific and more biased toward a flattering narrative; reported sources add distance and sometimes error. Of the fourteen cases here, seven rest on primary sources and seven on reported ones, and each card says which.

How long did these companies take to get traction?

Longer than the retellings imply. Zapier spent roughly eight months on outbound before a single unprompted inbound request. Segment spent about a year failing to differentiate in analytics before the library that worked. Notion spent three years before a rewrite and a version 1.0. Figma spent four years before general release. Loom made six hundred dollars in seven months before the pivot that became the company.

14. Sources

Every source behind every claim on this page, in order of first appearance. 7 of the 14 cases rest on a founder’s own account.

If you find an error in any of these, or a primary source that supersedes a reported one, tell us and the page gets corrected with the date changed. That is the only maintenance promise worth making about a page like this.

Related reading on this site: the startup guide hub for all seven playbooks, the founder rulebook for the end-to-end sequence, the launch guide for channel mechanics, and the go-to-market strategy post for the one-page version of the planning work.