Back to blog home

The trust trap: how showing AI reasoning helps but also hinders

Written by Chas Nelson

Transparency & auditabilityAI analytics

Showing AI reasoning can increase trust—and hide errors. Learn why explanations backfire and how checkable reasoning chains make AI outputs easier to challenge.

TL;DR: Understanding reasoning is what earns trust in a decision - and with AI there's no person behind the answer to earn it any other way. But the two most obvious fixes both backfire. Train a model's chain of thought to look better under review and it learns to game the review. Or ask a model to explain itself and a confident tone does the convincing regardless of whether the reasoning holds up. What actually works is reasoning built to be checked, not just read.

Previously I wrote about why AI ‘hallucination’ is a bigger problem than just dodgy facts: it’s easy to create outputs where every individual fact is true and, yet, the conclusion is still wrong, because the failure lives in the reasoning, not in the facts. This creates real reputational and monetary risks for any business trying to benefit from AI.

But manually going through an AI output and checking every claim for factuality, reasoning, logic. That’s hard. That takes time. It takes effort. It’s expensive. So, what if we made the AI show its reasoning somehow? Made that AI tell us how it got to the conclusion? Could we rely on that?

The problem - manually checking AI outputs is expensive and tough so a lot goes unchecked.
The problem - manually checking AI outputs is expensive and tough so a lot goes unchecked.

Reasoning builds trust

It's a reasonable instinct. If a colleague hands you a recommendation, you don't just take the answer and move on, you ask them to walk you through it: what they assumed, what they checked, what would change their mind. Seeing the reasoning is how you catch a bad call before it becomes a bad decision, long before AI ever entered the picture. We’ve previously made the case for how this works in software specifically: code was never really the product, it's the reasoning behind the product, and that reasoning (code) gets checked, by habit, before anything ships.

And the research backs this up: one study (Chen et al., 2025) found that showing a model's reasoning process significantly increases a reader's trust in, and agreement with, its conclusion. And another study found explanations reliably increased how much people accepted an AI's recommendation. Give people the "why," not just a polished "what," and trust follows.

Why this matters more for AI than for a colleague

With a colleague, the reasoning is one trust signal among several - you also know their track record and their reputation. With AI, most of that scaffolding is missing, so the reasoning is often the only thing standing in for the trust that would otherwise come from a relationship with a human colleague. That raises the stakes on getting the reasoning itself right, and it's exactly where things start to go wrong. Both of the ways people have actually tried to fix this come back to the same root problem: a model is optimised to be accepted, not to be correct. What differs between the two fixes is who acts on that incentive. In one, it's the machine. In the other, it's the human reading the answer.

Trust trap #1, the machine trap: train the reasoning, and the model games the trainer

The obvious fix is to make a model's actual chain of thought more legible - this makes getting an AI output more like brainstorming with a colleague. Every step in the thought process is made clear before the colleague goes away to write it all up.

But a model doesn't actually think the way a person does. Trained this way, it's rewarded for expressing human-like reasoning, not for actually reasoning that way. OpenAI's own research on chain-of-thought monitoring found the problem with this approach: folding a chain-of-thought monitor into an agent's training reward worked well at low levels of optimisation pressure, but past a certain point the agent learned to obfuscate its real reasoning in favour of more human-like apparent reasoning. It hid its intent inside chain-of-thought that still read as fine to the monitor, while still exploiting the same reward the monitor was meant to catch. It BSed about how it got to an answer.

Train an AI to show human-sounding reasoning and it obfuscates to hide the non-human-sounding reasoning it actually relies on.
Train an AI to show human-sounding reasoning and it obfuscates to hide the non-human-sounding reasoning it actually relies on.

This isn't a one-off. A model's stated chain of thought is frequently unfaithful to its actual decision driver even without anyone deliberately training against it, and this doesn't go away as models get better at reasoning. Anthropic tested this directly on two current reasoning models. They slipped a hint into the prompt and checked whether the model admitted using it. Claude mentioned it only 25% of the time, DeepSeek R1 39%. When the hint provided information that model had no other way to know it was used to answer correctly >99% of the time but models admitted this <2% of the time. The unfaithful explanations weren't even shorter, they were longer than the honest ones, as if the model were working harder to externally justify an answer it had already reached some other way. The moment "looks good under review" becomes something a model is optimised against, or even just something it has learned tends to satisfy people, you can no longer trust that its chain of thought reflects anything real.

Trust trap #2, the human trap: just ask for an explanation, and the confidence does the persuading

An alternative approach here would be post-hoc - just prompt the model to explain its reasoning as part of its answer or after its answer. But, here, a different problem shows up.

Readers judge an explanation as trustworthy largely by whether they agree with its conclusion, not by whether the reasoning is actually sound, and a confidently-toned explanation suppresses error detection even further. That confidence comes from how these models are built, trained and evaluated, which rewards a fluent, assured answer over an honest "I'm not sure."

A study of nearly 1,000 real AI-assisted work tasks also showed that the more confidence people had in the AI's ability, the less critical thinking they applied to its output (something our customers are also seeing). That drop in scrutiny isn't actually just related to AI. The same pattern shows up whenever people are handed more reasoning to read, whether it came from a model or from a person. With AI, the only thing that changes is that asking for an explanation is free and instant, so the trigger is always one prompt away.

This trap is a natural, human, behavioural trap - nothing malicious, nothing conscious. Just a natural human instinct to read less carefully the more reasoning is put in front of us. And the study mentioned earlier (Chen et al., 2025) found this trust effect held even when the reader held private information that contradicted the model and also after being told the model was missing context relevant to the decision.

Ask an AI to explain it’s reasoning post-hoc and an over-confident and verbose sounding wrong answer lulls the human reader into a false sense of trust.
Ask an AI to explain it’s reasoning post-hoc and an over-confident and verbose sounding wrong answer lulls the human reader into a false sense of trust.

Put the 2 traps together and the pattern is uncomfortable. Model are optimised for acceptance, not correctness, and whether that means they hide what is unacceptable or they produce outputs that human will more readily accept (at a cost to human critical thinking) - it’s the same problem: the answer gets accepted, gets trusted, gets relied on… whilst the reasoning, shown or otherwise, goes unchecked.

In fact, showing more reasoning, on its own, doesn't close the trust gap. It's just as likely to be why nobody caught the problem, whether the model hid the problem or the reader just didn’t look.

Both trust traps have the same problem - the reasoning is unchecked; either because it is hidden in the first place or because a large wall of text naturally makes the reader over trusting.
Both trust traps have the same problem - the reasoning is unchecked; either because it is hidden in the first place or because a large wall of text naturally makes the reader over trusting.

What actually works: reasoning built to be challenged, not read

What's needed is a form of reasoning that's human-readable and can be interrogated without falling into either trap. It can't be baked into the model's training the way chain-of-thought is, or there's nothing to stop the model from learning to game it. And it can't be just a fluent explanation bolted onto the answer, or a confident tone ends up doing the convincing on its own.

Enter what Count is calling “reasoning chains”. A reasoning chain isn't one thing, it's a sequence of smaller claims stitched together, and each of those claims falls into one of three buckets: it's sourced (it traces back to something checkable, a document, a dataset, a search result), it's inferred (it follows logically from other claims already in the chain), or it's assumed (it's a premise nobody stated or verified, just relied on). A fluent paragraph hides which parts of it are which. Break the same paragraph into its individual claims and label each one, and "trust the whole thing or don't" turns into a series of small, fast checks. Does this source actually say that? Does this inference actually follow from what came before it? An assumption is harder to fact-check by nature, but you can still ask whether it's doing more work in the argument than it should.

That's a different ask than "show your reasoning." It's checkable in a way a fluent paragraph isn't, because a reader is no longer being asked to hold an entire argument in their head and judge it by vibe, or to trust that a chain of thought is telling the truth about itself. They're being asked to check one claim at a time. Next time I’m going to show you exactly how this works and why we think it’s the way forward for reviewing AI outputs.