→ Proof Points
Hello fall!
The girls are back to school here in England, and I've spent the week representing the Red Sox. I'm prepping for my trip to Boston next week, excited to return where I lived for 6 years!

This week we are talking about post-data accumulation. Layer 2: Analysis. I'll share about a former US Secretary of Defense who, entirely by accident, gave us the best framework for it I know…
Previous edition: L1 Data
DEEP DIVE
Data, Data Everywhere

In 2002, Donald Rumsfeld was asked about evidence linking Iraq to weapons of mass destruction. He said this:
"As we know, there are known knowns. There are things we know we know. We also know there are known unknowns. That is to say, we know there are some things we do not know. But there are also unknown unknowns — the ones we don't know we don't know."
He was mocked relentlessly for it. It even became the title of his memoir, and a lot of memes.

Who doesn't love a good politics meme?
He also didn't write it. NASA engineers and defence contractors had been using this phrase since the 1960s. Rumsfeld picked up his version from a NASA administrator on a missile threat commission in the late 1990s and then delivered it to a room full of cameras with none of that context attached.
He was still right. On the epistemology, at least. And I have used those three categories to structure data analysis for digital health companies ever since.
Here is the problem the framework solves.
When a company finally sits down with its data — the mental health app, the clinical trial, the medical device, the genetic cohort — it almost always starts with the most difficult question in the room. The interesting one. The one that would look magnificent on a slide. Why did all the left-handed people start unicycling on the third of September on the north side of the Danube?
And the honest answer is you don't know. You have almost no data on that very small group of people. You may not have checked whether your data is even correct yet. But the exciting question got asked first, and now everyone is committed to answering it.
The Rumsfeld analysis exists to slow you down.
Step 1: the known knowns.
Start with a quality check. Not a glamorous one. Look at your data and find the things that cannot be true.
If you are working with children in a mental health app, how old are they? Are you finding any children aged 99? That is a data entry issue. Are you finding any aged 0? Possibly you're looking at infants, and a parent or a paediatrician didn't know how to enter six months, so they typed 0.6.
Then move to what I'd call your Table 1 variables. Nearly every paper in clinical research has a table near the front that describes the population — what proportion were female, how old people were, the youngest and the oldest, the standard deviation. Race and ethnicity, sometimes. Education, sometimes. Severity of condition, sometimes.
Table 1 lets anyone reading your work decide whether to believe the rest of it.
Now compare that table to everybody else's. Go and look at recent systematic reviews, meta-analyses, large trials in your field, posters from the last two conferences. What sample sizes do they report? How engaged were participants? What was the dropout over time?
You are doing this for two reasons: to catch errors, and to understand your bias.
Bias is just part of research. Part of life, too. If you ran a questionnaire online and recruited through Facebook, you will probably skew female. If you recruited through social media at all, your participants are likely more educated than the general population, because they have a computer, a connection, data, and spare time.
None of that invalidates your work. But if you don't discover it until after you've collected everything, and your condition predominantly affects men, you have some explaining to do at exactly the wrong moment.
Step 2: the known unknowns.
These are the big-ticket questions. The ones you want answered about your product, and the ones the field wants answered.
Systematic reviews are unusually useful here, because they tend to end by listing what nobody has established yet. Whether single-session interventions in mental health work as well as repeated weekly ones, for instance, is a question to which the current answer is: we don't have enough evidence. If your study can contribute to that, it matters enormously that you understand how the rest of the literature framed the question, so your work slots into the conversation rather than sitting beside it.
This is also where digital health has an advantage. You may simply have more data than everyone else.
In one study we ran with Woebot, looking at therapeutic alliance (the bond between the user and the product), we reviewed the published literature on that instrument and found that this one company held roughly twenty times more data on it than the entire scientific literature to that point.
That is worth knowing before you design your analysis. It might be the largest study ever conducted on that measure.
Then, wherever you can, pre-register your statistical analysis plan.
The gold standard is straightforward. Write down what you intend to test before you look at the data. Seal it in an envelope. Collect everything. Then open the envelope and do what you said you would do. When a peer reviewer or a competitor comes for your analysis later — and they will — you can show them the envelope.
HARKing: Hypothesising After the Results are Known. This is when you gather a pile of data with no particular idea in mind, notice on a chart that left-handed people appear to be better at unicycling, construct a theory about why, test it, and discover that yes, they are.
The problem is that you cherry-picked a result you had already seen and then built a story around it. It might be a real effect. It might equally be an error. Perhaps people aren't accurate about their handedness, or so few people are good at unicycling that it’s a quirk of having a small number of people within your sample who that describes. Perhaps you put the "strongly agree" button on the left of the screen and left-handed thumbs got there more often. You cannot rule anything out, because you found a pattern before you had a reason to look for it.
Its close relative is p-hacking. You have 100 people and 100 questions, so you compare male against female, old against young, severe against mild, and keep going. Every additional comparison raises the odds of finding a difference by pure chance. Run enough of them and you are guaranteed a result, especially if you wiggle the threshold around with terms like “approaching statistical significance”. But the finding just won't replicate.
Step three: the unknown unknowns.
This is the fun part, and it is genuinely one of the great advantages of working in digital health. Our datasets are large and our barriers are low, which means we can answer questions traditional grant-funded research never could.
At PatientsLikeMe, we ran a study on the frequency of uncontrolled yawning in ALS. Four weeks, 400 patients, and it was the largest of its kind in the literature, by a lot. Still is. We could do that only because the cost of asking was so small.
We also looked at how the menopause affects women with multiple sclerosis — a serious question that the academic we collaborated with at UCSF has since built much of her research career around. At the time, she could not get it funded, because it simply wasn't a fundable area.
So yes: go exploring. But only once the first two steps are solid. Clean your data, understand your biases, answer what the field is actually asking. Then go fishing for interesting stories.
Do it in the other order and you might find something remarkable — but spend the next two years explaining why nobody can reproduce it.
Where ProofStack exists
To be clear, the analysis belongs to the client. ProofStack advises on its shape. We don't run the statistics ourselves. We work alongside statisticians, data scientists and epidemiologists who do.
What we do is help decide which analyses actually matter.
One more thing: where did the data come from?
There are two kinds of analysis we're asked to help with.
Prospective, where you made a plan, collected the data accordingly, and now get to look at it carefully. That is the minority.
Retrospective, where someone wants to go digging through what the product has been capturing all along. That is most of it.
Retrospective work needs an audit before it needs an analysis. Why was this data captured? How? And, the one that catches everybody, was it captured consistently?
Products change every few weeks. If you used to ask users to rate their sleep from 0 to 10, and later added some emojis, those are two different measures. You cannot pool them. They look poolable in a spreadsheet, which is precisely the danger.
When historical data has drifted like that, our usual recommendation is uncomfortable. Take the most recent clean snapshot, even if it means discarding a great deal of sample size.
It sounds like a loss. It often isn't. Digital health companies tend to work in areas that traditional clinical research has barely touched, so 10% of your data may still be enormous relative to anything published in your field. A clean, consistent subset beats a large, incoherent one every time.
Which brings us back to the known knowns. You cannot skip them. They are the boring part, and they are the part that holds everything else up.
_ TIP
Short answer: no. _ Longer answer, courtesy of Professor Adam Kucharski, who has been documenting this carefully: Off-the-shelf LLMs are not reliable at generating truly random numbers. They do not perform independent reviews of a dataset the way a statistical package does. They can give you different answers depending on how you phrase the question. They are frequently overconfident about the quality of their own output. And they struggle with probability. Kucharski's demonstration is beautifully simple: ask several models to multiply 9,321 by 1,789 (something any calculator handles) and you do not reliably get a consistent answer. He has also found that people massively overestimate how accurate these models are at statistical and mathematical work, which is the more dangerous finding of the two. Use an LLM as a thought partner to structure the work. Execute in a real statistical package: R, MATLAB, or—if you are being held hostage in a psychology department—SPSS. _ |
Where we do use AI
We've built a set of Claude skills at ProofStack that we share and carefully deploy with clients. They contextualise the known knowns fast, pull the papers that tell you whether your data looks strange, and — the one clients use most — calibrate a draft Table 1 against what's already been published.
The second area where they're genuinely good is visualisation. Turning findings into interactive graphs is something clients ask us for more and more, and the tooling has become very capable.
We keep an eye on whether that changes (So far, it hasn't).
Worth reading: Professor Adam Kucharski's piece on probability here.
Also reading: the Wholesum newsletter (Disclosure: I'm an investor).
Next issue: L3 — Proof Assets.
FROM OUR DESK
The ROI of Evidence
A question I get asked constantly, usually by founders deciding whether to spend money on evidence: all other things being equal, what is good evidence actually worth?
There is now a number attached to it (I wrote about this on LinkedIn).
A group called Mentalium assembled the Mental Health Startup Graveyard — 542 digital mental health organisations that left the market between 2000 and 2026, each coded on business model, who paid, funding raised, clinical evidence, medical co-founder, and more. "Leaving the market" covers everything from acquisition to shutdown, and the analysis looks at which way companies went.
Among the companies where they could establish evidence status, those with no clinical evidence ended in shutdown or bankruptcy 44% of the time. Those with published clinical evidence: 34%.
A ten-point difference. A stronger signal, in fact, than whether the company had a medical co-founder.
It sits in the top handful of factors, behind the structural ones, whether an institution paid rather than a consumer, subscription versus one-time purchase, funding raised. Evidence isn't the largest lever in that dataset. But it is one of the few on the list you can still pull after the company is built.
One detail I appreciated. Where Mentalium couldn't establish whether a company had clinical evidence, they kept those records in a separate "not established" group rather than folding them into "no." That would have moved the numbers in a flattering direction. They didn't do it.
That's what good known-knowns hygiene looks like.

Thanks for reading,
Paul Wicks, PhD
Founder & CEO, ProofStack Health
Move Fast. Prove Things.
P.S. Here are 3 ways I can help you:
How to Turn P-Values into Financial Values. One of our most popular freebies for digital health founders.
Follow or connect with me on LinkedIn. I publish top resources and in-depth insights related to building your evidence stack.
Book a strategy session. Uncover the gaps in your evidence and marketing in your Digital Health/MedTech startup.
