How Bad Are AI Hallucinations Really?
Written by The Count Team

Explore 8 ways AI answers fail—from fabrication to stale knowledge—and learn how to verify claims before using them in high-stakes work.
AI hallucinations remain a serious problem for consequential work, even though blatant falsehoods appear less often in newer models. The harder failures are confident overclaims, stale knowledge, hidden assumptions, weak methodology and conversational drift. AI output should be treated as material to verify, not as evidence by itself.
This matters most to people using AI in law, medicine, finance, analytics and other work where customers or stakeholders depend on the result. Better verification lets those teams use AI beyond low-risk tasks without handing responsibility for the output to the model.
How often do AI models hallucinate?
Blatant falsehoods appear to be much less common in newer frontier models, but no single rate describes the whole problem. As Chas Nelson says, 5 years ago results showed that LLMs made up a fact anything up to 50% of the time whilst the latest benchmarks show this is now down to <1% for frontier models.
Those figures cover errors that are easy to score. Examples include stating that two plus two equals five, extracting the wrong number from a document or inventing a fact. Model providers have focused heavily on reducing these failures because a benchmark can label them as right or wrong.
Consequential work usually contains more ambiguity. A quarter can be weak because revenue fell and strong because the wider market fell further. A strategy recommendation may contain real evidence while still relying on poor reasoning. Those failures are harder to count and easier to miss.
What are the eight ways an AI answer can fail?
The eight failure modes discussed in the video cover factual errors, reasoning errors and problems created through conversation.
- Outright fabrication or extraction error. The model invents a fact, citation or number, or reads the source incorrectly. The video cites commissioned reports involving Deloitte in Australia and EY in Canada where independent researchers found confident claims attributed to sources that did not support them.
- Sycophancy. The model accepts a user's correction without checking whether the user is right. A mistaken user and an agreeable model can create an echo chamber.
- Resistance to correction. The model keeps defending an earlier assumption even after the user supplies contrary evidence.
- Recency effects. The model gives recent instructions more weight than decisions made earlier in a long conversation.
- Overconfidence and overclaiming. The model removes caveats and presents a finding that applies sometimes as a universal rule.
- Poor methodology. The model overlooks small samples, confirmation bias, correlation versus causation or the limits of extrapolation.
- Unsupported assumptions. The model fills a gap with a plausible premise without marking that premise as an assumption.
- Unflagged model knowledge. The model relies on information learned during training without saying that the information may be old or incomplete.
The boundaries between the final categories are not absolute. The useful distinction is whether the failure came from a false fact, faulty reasoning, an unstated assumption or the interaction between the model and the user.
Why do confident AI answers feel trustworthy?
Confident AI answers feel trustworthy because fluency and structure can look like evidence. Human evaluations also tend to reward answers that sound decisive, which gives models an incentive to hide uncertainty.
The video discusses research where people preferred answers with visible reasoning. When that reasoning contained a flaw the reviewers understood, they could still favour the confident explanation over their own knowledge.
A polished citation does not solve the problem. A model can cite a real professor while inventing the professor's book or quotation. The surface features of careful research can therefore make a fabricated claim harder to spot.
Can AI be trusted to research current facts?
AI can help gather current facts only when users can distinguish retrieved sources from internal model knowledge. A model may answer from training data even when the subject has changed since that training data was collected.
Chas describes asking Claude about a law and receiving a detailed answer that appeared plausible. When asked where the facts came from, Claude said they came from internal model knowledge. The law had changed during the previous six months, so the answer described an earlier legal position with present-day confidence.
The same issue appears in software work. A coding model may insist that a feature does not exist because its documentation is out of date. Correcting the answer can require linking to the current documentation and pointing to the specific lines that describe the feature.
Internal knowledge can also support a reasonable but incomplete inference. A model might note that Switzerland is not in the European Union and infer that EU rules have no equivalent there. That misses the further question of whether Swiss law mirrors those rules.
Why can the same question get different answers?
The same question can produce different answers because models choose different sources, apply different assumptions and complete patterns probabilistically. Each response can sound final even when the model gives another answer seconds later.
Alice Ferrier tested this by repeatedly asking Claude for the weight of the heaviest recorded hippo. Claude returned different weights and different locations, including Germany, Johannesburg, Egypt and Uganda. The answers varied whether Claude searched the web or relied on internal knowledge.
The hippo question was deliberately low stakes. The same inconsistency becomes serious when the number concerns a forecast, a board presentation or another business decision. A fast answer between meetings offers no visible signal that another run might return something different.
Directing the model to a specific source does not remove every risk. Mitra Abrams describes asking a model to extract sentiment and key points from a meeting transcript. She could identify incorrect conclusions because she attended the meeting. Someone without that context might accept the same conclusions.
Can a conversation make a hallucination worse?
A conversation can push a model from cautious speculation into confident fabrication when the user repeatedly validates its claims. Reciprocal agreement becomes part of the context the model uses for its next answer.
Alice tested this by trying to persuade Claude that miners had sent their wives underground to check for poisonous gases. Claude initially treated the idea cautiously. After repeated encouragement, it produced a sourced-looking argument that compared wives with canaries and described the supposed practice as a real warning system.
The example was intentionally absurd, but the interaction was ordinary. The user showed interest, agreed with the model and asked it to expand. Those are common behaviours when people use AI to develop ideas.
Long conversations add another problem. Context may be compressed, earlier details may be lost and recent mistakes may outweigh decisions established weeks before. The resulting drift can be difficult to notice because each individual reply still fits the immediate conversation.
How should you verify an AI answer?
A reliable review traces every material claim to a current source and checks whether the source supports the wording used. Asking where each fact came from can reveal whether the model searched for evidence or relied on internal knowledge.
For consequential work, the review should include these checks.
- Ask which claims came from retrieved sources and which came from model knowledge.
- Open the cited source and confirm that the quotation, number or conclusion appears there.
- Check the source date when laws, software, markets or policies may have changed.
- Look for caveats that disappeared between the source and the answer.
- Repeat important quantitative questions and investigate inconsistent results.
- Compare the reasoning with relevant domain knowledge, not only the final conclusion.
Verification costs vary by task. Checking a familiar meeting summary may take seconds. Unpicking a long report can take more time than producing the work directly. That is why some teams limit AI to outputs where a person can remain accountable and verify the result efficiently.
Do hallucination benchmarks show which model to trust?
Hallucination benchmarks are useful only when their scoring method matches the kind of failure you care about. A headline percentage can be misleading without its denominator and definition.
The video discusses Artificial Analysis and its Omniscience benchmark. One result is described as Claude hallucinating 73 percent of the time that it did not . That does not mean 73 percent of all Claude answers were false. It means that among the questions Claude answered incorrectly, 73 percent received a false answer instead of an admission that the model did not know.
Artificial Analysis also offers Briefcase, which covers longer business tasks where recency, sycophancy and resistance to correction can emerge. Its public implementation of GDP Val covers work in fields including law, medicine, human resources and finance.
Some quality benchmarks compare pairs of outputs and ask people which one they prefer. That method can reward the clearer or more confident response when both answers appear plausible. A preference score does not necessarily show whether the reasoning, assumptions and sources are sound.
How does Count make AI output easier to inspect?
Count makes AI output easier to inspect by placing people, agents, queries, sources and results in the same shared canvas. The agent participates in the same live state as human collaborators rather than producing work in a separate black box.
Count was multiplayer before the agent was added. Its two-dimensional canvas lets collaborators arrange analysis spatially, inspect upstream and downstream work and see what other people are changing. The existing synchronization system also gives the agent a live view of cells, query status, returned data and dependencies.
Count knows which cells depend on other cells. That gives the agent explicit information about data flow and causality without asking it to infer every relationship from SQL. When a query runs, people and agents can use the same result rather than duplicating work against the database.
The shared canvas does not make a model infallible. Prose can still contain a number drawn from memory rather than data. Count is exploring ways to make the origins of generated claims more explicit, just as the canvas already exposes connections between queries, cells and visual results.
Choose one consequential AI output and trace its important claims back to current evidence before using it. Watch How bad are AI hallucinations really? Exploring AI's 8 failure modes and confusing Claude! for the worked examples, including the inconsistent hippo research, the coal mine conversation and the shared Count canvas.