How AI Makes Critical Thinkers More Valuable, Not Less

I build machine-learning pipelines for a living. Not as a hobby, and not as a side interest — a meaningful share of my working week is spent writing the code that turns brain scans into numbers a study can reason about. I know what these systems can do, because I assemble them. I also know what they cannot do, for the same reason.

So I want to say something that runs against most of what you have read this year: I am not worried about AI, and I do not think you should be either — provided you understand what it is actually pricing. The common fear is that AI replaces thinkers. What I see in my own field is closer to the opposite. AI is making a specific kind of thinking cheaper, and in doing so it is quietly raising the price of everything the machine still cannot do.

That second category — judgment, problem framing, honest handling of evidence — is not shrinking. It is becoming the scarce input. This post is about why, and about the thinking skills that are worth building now precisely because they appreciate as the tools improve.

What AI actually made cheap

It helps to be precise about what changed. A large language model is, at its core, a very good predictor of what text should come next. Trained on an enormous corpus, it produces fluent, plausible answers to almost any question you can phrase. That is a genuine and remarkable capability. It is also a specific one.

What became cheap, almost overnight, is *generating a plausible answer to a well-posed question*. Drafting, summarizing, translating between formats, producing a first pass at code — the tasks that used to cost you an afternoon now cost a paragraph of instruction. This is real, and pretending otherwise helps no one.

EVIDENCE GRADE: STRONG

The strong-grade claim is narrow and worth stating cleanly: on well-specified tasks with abundant training examples, these systems produce competent output at near-zero marginal cost. That is documented across coding, drafting, and summarization, and it matches what I observe when I use them daily.

But notice the shape of that sentence. *Well-posed question. Well-specified task.* Everything expensive in real work lives in the words the model quietly assumes are already handled: deciding which question is worth asking, judging whether the fluent answer is actually correct, and knowing what the answer would cost if it were wrong. The machine runs the middle of the problem. It does not choose the problem, and it does not check itself.

Fluent is not the same as right

Here is the failure mode I watch for most, because it is the one my field is built to resist. A language model will produce a confident, well-structured, entirely wrong answer with exactly the same fluency as a correct one. It has no internal signal that distinguishes the two. Fluency is not calibration.

In neuroimaging this distinction is survival. If I ask a model — or a data-hungry method — to find a pattern in a study of thirty-five patients, it will find one. It will always find one. Whether that pattern is a fact about the brain or an accident of thirty-five particular people is a question the method cannot answer about itself. That is the human’s job, and it is the whole job. I have written before about the five-step procedure I use to separate the two, in how to think like a scientist — and the reason it matters more now, not less, is that it is precisely the method AI can’t run on its own output.

This is why calibration — matching your confidence to your evidence — is becoming the premium skill rather than a niche academic virtue. When answers were expensive to produce, the bottleneck was production. When answers are nearly free, the bottleneck moves to *evaluation*: which of these plausible outputs do I trust, and how much? A tool that generates ten confident answers has not solved your problem. It has handed you a new one — choosing among them — and that problem is pure judgment.

EVIDENCE GRADE: MODERATE

I grade this Moderate rather than Strong, honestly. The claim that automating a task shifts the human’s value toward oversight and judgment is well supported by the broader history of automation and by early workplace studies of these tools. But direct, long-run evidence about knowledge work specifically is still accumulating. The mechanism is sound; the decade of data is not yet in. I would rather tell you that plainly than dress an emerging trend as a settled law.

The engineer’s lesson: fundamentals survive the tool

I have changed fields once already, and it taught me how durability actually works. I trained as an electronics engineer in Chennai, then spent years on speech-enhancement algorithms for hearing aids at Fraunhofer, and now I study the brain with MRI. Almost every specific tool from the first chapter is obsolete or irrelevant to the third. The programming languages changed. The hardware changed. The domain changed entirely.

What transferred was everything underneath the tools: how to decompose a messy problem into parts you can actually test, how to reason about noise and signal, how to design a comparison that could prove you wrong. Those did not survive because I protected them. They survived because they operate at a level the tool-churn never reaches. A technique is a tool. A way of thinking is a foundation, and foundations outlast the buildings put on them.

This is the quiet argument against AI panic. If your value is a set of tasks a tool can now do, a better tool is a threat — and there will always be a better tool. If your value is the judgment about *which* tasks are worth doing and *whether* they were done correctly, a better tool is leverage. It does more of the work you were going to delegate anyway, and it makes your judgment the binding constraint on the whole output. The five thinking tools I would build first for exactly this reason are the models to build, laid out in the five mental models every young professional needs first.

FIELD NOTE — ERLANGEN

I deliberately avoid deep learning on my own studies, and people are often surprised to hear it from someone who builds ML pipelines. The reason is discipline, not distrust of the method. With thirty to forty participants, a flexible model has more than enough freedom to fit the noise perfectly and tell me nothing about the brain. So we use simpler, interpretable methods matched to the sample, and we spend our real effort on the design and the sanity checks. The instinct that keeps that work honest — *what else could produce this result, and does my confidence match my evidence?* — is exactly the instinct no model applies to its own output. Building the tools is what convinced me the judgment around them is the scarce part.

What “critical thinking” concretely means here

“Critical thinking” is a phrase worn smooth by overuse, so let me replace it with three specific capabilities — the ones I watch appreciate as the tools improve.

Framing the question. In research, a well-posed question is most of the work; a badly framed one wastes months no matter how good your analysis is. AI inherits this completely. It answers what you ask with unnerving literalism, which means the value migrates upstream, to the person who decides what to ask. The machine optimizes the answer. Only you can choose the question.

Evaluating the answer. This is calibration again: knowing the denominator, spotting the confound, asking what evidence would change your mind. A fluent answer triggers none of these reflexes on its own. You have to bring them.

Anticipating consequences. What happens after you act on the answer — the second-order effects, the failure modes the confident draft never mentioned. A model predicts the next word. It does not price the downstream cost of being wrong.

AI does this cheaply nowThe human still owns this
Choosing the problemFraming the right question
Producing an answerDrafting, summarizing, first-pass code
Judging the answerCalibration: is it actually true?
Acting on itSecond-order consequences, failure modes

Read the table as a map of where your attention should go. The middle column is being commoditized. The right column is where price is rising. Every hour you invest in the right column compounds; every hour spent competing with the middle column is an hour spent racing a machine at the one thing it is genuinely built to win.

Try this today

Take the next task you are tempted to hand to an AI tool, and split it in two before you start. Write the question you are actually asking in one sentence — sharp enough that a wrong answer would be *obviously* wrong. Then, before you accept whatever comes back, write the single check that would tell you the answer is untrustworthy: the number you would look up, the case that would break it, the person who would know. Ten minutes, two sentences. You have just done the two things the tool cannot do for you — and they are the two things your job is increasingly paying you for.

Match your confidence to your moment

The honest grade on the whole thesis is this: the mechanism is Strong, the decade of workplace data is Emerging, and anyone selling you certainty in either direction — that AI changes nothing, or that it changes everything — is overselling their evidence. What I can tell you, from inside a field that runs on small data and self-suspicion, is that the skills which make thin evidence usable are the same skills a room full of fluent answers now demands. That is not a coincidence. It is the same skill, and it was undervalued before the tools made it scarce.

I am not worried, then, but I am not passive either. The tools are real, the shift is real, and the people who treat judgment as something to practice — not something they already possess — are the ones for whom every improvement in AI is a raise rather than a threat.

If you want the concrete starting kit, the five thinking tools I would build first are laid out in the five mental models every young professional needs first — the ones that operate before the question is asked and after the answer arrives, which is exactly the ground the machine does not touch.

If this way of thinking is useful to you, the natural next step is the free guide this site is built around: [5 Mental Models to Future-Proof Your Career](/newsletter/) — five models chosen and stress-tested by a brain researcher, each with an honest grade of the evidence behind it. You’ll also get *Signal*, my monthly email: one idea from neuroscience you can use at work. No productivity spam, no AI panic.

*Mageshwar Selvakumar is a doctoral researcher in neuroscience in Erlangen, Germany, studying how chronic pain reshapes the brain using multi-parametric MRI.*

Similar Posts