Small Data, Big Decisions: What MRI Research Taught Me About Deciding With Incomplete Information

Most of my studies are built on thirty to forty people. Not thirty thousand, not three hundred — thirty-something individuals who agreed to lie still inside a loud magnet for the better part of an hour so that I could measure something about how their brains are changing. From that thin slice of humanity, my field is expected to say something true about the brain in general. Every scan is expensive. Every participant is hard-won. And at the end of it, a decision has to be made: does this result mean anything, or doesn’t it?

I have come to think this is not a quirk of neuroscience. It is the ordinary condition of almost every important decision a person makes at work. You rarely have the sample you would like. You have a handful of data points, some of them noisy, most of them collected for a different purpose, and a choice that cannot wait for the study to be properly powered. The question is never “what does the complete evidence say?” It is “what can I responsibly conclude from the incomplete evidence I actually have?”

That is a skill, not a temperament, and it is one I now believe matters more the further AI advances — because the machine can process the data you have far faster than it can tell you whether that data supports the decision at all. I have argued the broader case for this elsewhere, in how AI makes critical thinkers more valuable, not less; here I want to be narrow and practical, and show you the specific habits a small-data field uses to reach defensible conclusions without deceiving itself.

Why I refuse the tempting shortcut

Let me start with a confession that surprises people, because I build machine-learning pipelines for a living. When I have a study of thirty-five participants, I deliberately do not reach for the most powerful modelling tools available to me. No deep learning. No sprawling model with thousands of free parameters hunting for structure in the scans.

The reason is not fashion or caution for its own sake. A flexible enough model, pointed at a small enough dataset, will always find a pattern. It will fit those thirty-five people beautifully — their noise, their quirks, the accident of who happened to enrol — and it will tell me nothing that generalises to the thirty-sixth person I have never met. The pattern is real in the sense that it is *there in the data*. It is false in the only sense that matters: it will not hold up.

EVIDENCE GRADE: STRONG

That overfitting risk on small samples is not a matter of opinion; it is one of the best-established results in statistics and machine learning. The more freedom you give a model relative to the amount of data you feed it, the more reliably it will memorise the sample instead of learning the world. So we use simpler, interpretable methods matched to the sample size, and we spend our real effort somewhere the software cannot help: on the design of the study and the honesty of the checks.

The translation to your work is direct. When someone hands you a slick dashboard, a confident forecast, or an AI-generated analysis built on a thin trickle of data, the sophistication of the tool is not evidence. A powerful method applied to weak data produces a confident answer, not a correct one. The first question is never “how good is the model?” It is “how much did it actually have to learn from?”

The denominator is the whole story

The single most common way to be fooled by incomplete data is to look at the numerator and forget the denominator. Three complaints sounds alarming until you learn there were three thousand transactions. A feature “everyone is asking for” was, on inspection, asked for by four of your most vocal customers out of eight hundred.

In imaging, I live and die by denominators. “We saw a difference in this brain region” is meaningless until you know: a difference across how many people, of what expected size, and against what natural variation? A number without its denominator is not a small piece of evidence. It is a rhetorical device.

So the first move, always, is to reconstruct the denominator before you react to the numerator. Out of how many? Compared to what base rate? The instinct to ask this is precisely the instinct that fails us under pressure, because a vivid numerator hijacks attention — a mechanism I unpack in the cognitive biases that distort decisions at work, where thin data is exactly the soil these errors grow best in. Small samples do not merely give you less information. They give bias more room to operate, because there is less signal to anchor you and more noise to read a story into.

Baseline before verdict

Closely related, and just as often skipped: you cannot interpret a result without knowing what you would have seen anyway.

This is why my field almost never runs a study without a matched control group. If I scan a group of patients with a chronic inflammatory condition and find that a certain brain measure sits at a particular value, that number alone tells me nothing. I need a comparison group — matched as closely as I can manage on age, on sex, on the things I know move the measure — so that I can ask the only useful question: is this different from what I would expect in people *without* the condition? The finding is not the patient value. The finding is the gap between the patient value and the honest baseline.

EVIDENCE GRADE: STRONG

The need for a comparison or baseline to make a raw number interpretable is foundational across the experimental sciences; it is not a stylistic preference. At work, the equivalent question is almost always available and almost always ignored. Sales rose after the campaign — compared to what they would have done in that season anyway? The new hire’s team shipped faster — faster than a comparable team, or faster than a slow quarter? Before you credit a cause, establish the baseline. A change measured against nothing is not a result.

Selection effects: who is missing from your data

Here is the failure that most rewards a researcher’s paranoia, and the one professionals are least trained to see. It is not about the data in front of you. It is about the data that never reached you.

Every dataset is a survivor. My scans only include the participants who volunteered, who fit the scanner, who could tolerate lying still, whose data survived quality control after I discarded the runs corrupted by head motion. Each of those filters is sensible. Together they mean my sample is not a random slice of patients — it is a specific, self-selected, quality-filtered subset, and I am obligated to say so plainly and to reason about which way it might tilt my conclusions.

EVIDENCE GRADE: MODERATE

I grade this Moderate not because the mechanism is uncertain — selection effects are real and pervasive — but because their *direction and size* in any specific case are usually a matter of careful judgment rather than measurement. You often cannot quantify exactly how your sample is skewed. You can only stay alert to the fact that it is.

Your work data is a survivor too. Customer surveys are answered by the unusually delighted and the unusually furious; the vast indifferent middle stays silent. Exit interviews are given by people polite enough to give them. The users in your analytics are the ones who did not churn before you started measuring. Before you conclude anything from a dataset, ask the uncomfortable question: who is systematically not in here, and would including them change the story? Often it would reverse it.

FIELD NOTE — ERLANGEN

The cleanest illustration of small-data honesty I know is head motion. A brain scan measures the brain, but it also faithfully records every tiny movement of the head inside the scanner — and those movements can masquerade as a genuine biological signal, one that correlates with age, with pain, with exactly the things a patient study is trying to measure. Early in this work I have watched a promising group difference shrink toward nothing the moment we properly modelled and corrected for motion. The lesson stuck harder than any positive result could have: on a sample of thirty-five, the difference between a finding and an artefact is not the size of the effect but the rigour of your accounting for what else could produce it. Before I pre-register a study I now write down, in advance, the confounds that could fake my result and the checks that would expose each one. Deciding what would fool you — before you look — is the most reliable protection against being fooled. It costs an afternoon of imagining ways to be wrong, and it is the cheapest insurance in science.

How the lab reaches a defensible conclusion anyway

None of this is an argument for paralysis. My field publishes conclusions from small samples constantly, and good ones. The trick is that we do not pretend the data is stronger than it is; we build a scaffold of practices that let thin evidence bear a modest, honest weight.

Four of those practices translate almost without alteration to a desk:

Decide the question before you see the answer. We pre-register — we commit, in writing and in advance, to what we are testing and how we will judge it. This is not bureaucracy. It is the single strongest defence against the mind’s talent for finding a story after the fact and mistaking it for a prediction. At work: write down what result would change your decision *before* you pull the report, not after. – Match your comparison. No verdict without a baseline, and the more honest the baseline, the more the verdict is worth. – List what could fool you. Name the confounds out loud — the alternative explanations, the selection effects — and check each one deliberately rather than hoping it away. – Calibrate the claim to the evidence. State your confidence as a range, not a point, and make the range honest. A finding from thirty-five people is a reason to look further, not a law of nature, and saying so is a strength, not a weakness.

The instinctWhat it protects againstThe question to ask
Find the denominatorA vivid numerator with no scaleOut of how many?
Establish the baselineCrediting a change to the wrong causeCompared to what I’d see anyway?
Hunt the selection effectThe data that never reached youWho is missing from this?
Calibrate the claimFalse confidence from thin dataHow sure can I honestly be?

Act reversibly, and let evidence catch up

There is one more habit, and it is the one that resolves the whole tension between “the data is incomplete” and “the decision cannot wait.”

Match the reversibility of your action to the strength of your evidence. Weak evidence does not forbid action — it forbids *irreversible* action. In the lab, a suggestive result from a small study does not get announced as a discovery; it earns a bigger, more careful follow-up. The finding buys the next experiment, nothing more. The stakes of the action are kept proportional to the confidence behind it.

EVIDENCE GRADE: MODERATE

Framing decisions in terms of reversibility and proportional commitment is a well-reasoned principle drawn from decision theory and forecasting practice rather than a single measured finding, which is why I grade it Moderate — but it is one of the most useful rules I know. On thin evidence, prefer the choice you can undo: the pilot over the rollout, the reversible hire over the restructure, the test that buys information over the bet that spends it. Then let the evidence accumulate and let your commitment grow with it. You are not choosing between acting and waiting. You are choosing an action whose cost, if you are wrong, you can afford.

Try this today

Take a decision you are currently making on incomplete data. Write three sentences. First: the denominator — the *out of how many*, and the baseline you are comparing against, even if you have to estimate both. Second: who is systematically missing from the data you do have, and which way that likely tilts it. Third: the most reversible version of the action that still moves you forward. If you cannot fill in the first two honestly, the third just became more important — act small, keep the option open, and let the evidence come to you.

The confidence to decide anyway

Small samples taught me something I did not expect: that rigour and decisiveness are allies, not opposites. The people who are paralysed by incomplete information and the people who charge ahead pretending it is complete are making the same mistake in opposite directions — both are refusing to hold uncertainty and action in the same hand at once.

The discipline of a small-data field is precisely the discipline of doing both. You find the denominator, you build the baseline, you name what is missing, you grade your own confidence honestly — and then, with all of that in view, you make the call, sized to what you actually know. That is not a compromise between rigour and speed. It is what mature judgment looks like, and it is buildable, one habit at a time.

If you want the small set of thinking tools I would install first for exactly this kind of decision, I have collected them in the five mental models every young professional needs first — the ones that do their work before you see the data and after you have to act on it, which is the ground no amount of computing power will ever cover for you.

If this way of thinking is useful to you, the natural next step is the free guide this site is built around: [5 Mental Models to Future-Proof Your Career](/newsletter/) — five models chosen and stress-tested by a brain researcher, each with an honest grade of the evidence behind it. You’ll also get *Signal*, my monthly email: one idea from neuroscience you can use at work. No productivity spam, no AI panic.

*Mageshwar Selvakumar is a doctoral researcher in neuroscience in Erlangen, Germany, studying how chronic pain reshapes the brain using multi-parametric MRI.*

Similar Posts