Skip to content

Is Our Industry Built on Data or on Selection Bias?

A confident retool claim at an industry event and what the numbers look like when I went looking for data.

Is Our Industry Built on Data or on Selection Bias?

At an industry event in June, I stood at the edge of a small group and listened. Someone with industry experience was talking about retooling. Revenue would go up a significant amount, the individual said and it’s worth doing it.

The group around him was mixed. A few were people trying to buy their first laundromat, newer owner/operators running stores they'd bought, with older equipment they were figuring out what to do with.

The individual didn't own a laundromat or laundry service. Their connections ran to financing and equipment distribution. They weren't pushing a specific brand, just describing what retooling did to a store's performance as if it were gospel in the industry.

I stood there and my thinking was split. I heard what they were saying, on the other I watched the group hearing it. The numbers came out clean, no specific store's utilities, no comparison against a specific machine's consumption, etc. Just retool, and revenue goes up a a lot amount.

That night I couldn't stop thinking about what I'd watched. It wasn't what they said that stayed with me. It was what was underneath what they said. Where did the numbers came from? What was the pool of stores that were represented in his data?

The sample you never saw

Selection bias is a specific kind of trouble. It's different from the general category of sampling problems. If three clients ask us about a service and we add it, not realizing those three are the ones who spoke up because the service matters to them, that's a sampling problem. Selection bias isn't an accident, it’s when someone with an interest in the outcome chose which data points made it to be shown before we ever saw them all.

In our industry this operates at two altitudes. At the owner/operator scale, three loud clients shape a policy or service for the others. At the industry scale, a lender publishes a "success rate" built from its own loan performance data, drawn from loans issued to borrowers the lender pre-qualified, and that number travels across marketing materials, courses, and podcasts as if it measured the average laundromat's outcome.

There's a second mechanism that keeps a filtered belief alive once it's in circulation. Confirmation bias. Once we've absorbed a number and made a decision around it, our minds do the rest of the work. We remember the confirmations and forget the disconfirmations. The three clients who loved the new service get repeated in memory, the two who quietly stopped coming get filed under something else in our minds versus the true reason they stopped coming.

Selection bias assembles the belief. Confirmation bias keeps it alive.

At our store it played out with WDF

I've watched this operate in my own thinking. When our wash and fold clients started asking about 24-hour turnaround, I could feel the urge to say yes. It's what parts of our industry describe as best practice. It's what a fair number of pickup and delivery operations promise on their websites. I paused, and I asked myself, how many clients are requesting this compared to the entire client base? Then I looked at the numbers.

The math didn't work at 24 hours without straining production or compromising quality. So we stayed with our 48-hour standard, with a premium option for genuine emergencies. It wasn't a policy built to satisfy the industry practices or those few clients. It fit our operations and our larger client base.

Months later I started hearing from other owner/operators who'd gone the other direction. They'd offered 24-hour service, then extended it out to 48, and their business hadn't dropped in any meaningful way. A few clients complained and almost no one left. The owner/operators reported feeling relief. Due to more time for the quality of the work, which was the thing keeping clients coming back in their experiences.

I wrote about this in April 2025 in Speed vs. Quality, though I didn't name the mechanism then. What I was really asking was a selection bias question. Whose story about 24-hour turnaround were we operating on? The ones who'd tried it and reported success? Where were the owner/operators who'd tried it and quietly rolled it back? The pull to add 24-hour service came from a filtered sample of loud successes.

There was another one I remember myself in. A while back I implemented a price change on one of our services. Two or three clients had said something. Not a complaint, more like a mention they noticed an increase. I would catch myself thinking maybe I shouldn’t have raised it. I waited to see if it would come up again, with any other clients. It never did and the price stayed. If I would have moved fast on those comments, I would've been reshaping a service around what turned out to be a subset of a subset. That's how quickly small samples can move an owner/operator's hand.

Every industry has this story

The mechanism isn't unique to our industry. It's one of the most studied problems in statistics.

In 1943, at the Statistical Research Group at Columbia University, Abraham Wald was asked to help the U.S. military decide where to add armor to bombers coming back from missions in Europe. The military had catalogued damage patterns on returning aircraft. Their instinct was to reinforce the areas that showed the most hits.

Wald pointed out that the aircraft available for study were a filtered sample. They were the ones that came back. The planes hit in the engines and the cockpit never returned to be examined. The damage patterns on survivors weren't a map of vulnerability. They were a map of what a plane could take without going down.¹ The armor went to the places where returning planes showed no damage.

Sit with that for a moment. The correct answer looked like the opposite of the obvious answer, and the reason was the make up of the sample.

The finance industry ran into the same problem several decades later. Researchers noticed that reported mutual fund performance systematically overstated actual investor returns. Funds that failed disappeared from the databases used to calculate historical averages. Hmmm, where is the data on all the failed laundromats? Sorry, I digress.

Peer-reviewed research put the size of that distortion between 1.4% and 2% annually.² Once academics corrected for it, most of what had looked like persistent fund outperformance disappeared.³ What looked like the average fund's performance was actually the average surviving fund's performance.

Franchising has a version that runs parallel to ours. For decades the industry's most cited number was that franchises have a 95% success rate. The claim traced back to a single 1991 study that got widely misinterpreted.

When Timothy Bates at Wayne State University ran a proper comparison of more than 20,500 small businesses in 1994, he found that franchises had lower survival rates than independent businesses at the four-year mark. 65% versus 72%. Retail franchises fared worse still.⁴ Later work using SBA loan data showed franchise default rates around 10% across the industry, with individual brands ranging from under 5% to over 40%.⁵

Then there are vending machines. For most of the 2000s, vending machine business opportunity sellers were making earnings claims that couldn't be substantiated. Pitches about $50,000 a year in passive income from a small route, tied to promised locations that turned out not to exist.

The FTC eventually launched joint enforcement actions with state attorneys general. Operation Vend Up Broke in 1998. Project Busted Opportunity in 2001 and 2002. Multiple settlements and permanent bans followed.⁶ Sellers were drawing earnings claims from a pre-filtered sample of successful operators, or from nothing at all, and the numbers traveled to prospective buyers stripped of methodology. The FTC eventually built the Business Opportunity Rule and the Franchise Rule to require pre-sale disclosure of earnings claim methodology.⁷ The vending industry couldn't self-correct until disclosure was made mandatory.

The sample our industry inherits

Consider the retooling narrative in our own industry. Alliance Laundry Systems Distribution defines its "success rate" openly in the footnote of its investment materials, as the number of charged-off loans per laundry divided by the total number of loans originated between 2007 and 2022, based on Alliance's own research.⁸ The methodology is right there at the source. The source isn't the problem.