Back to blog home

How We Talked AI Into a UFO Cover-Up

Written by The Count Team

AI analyticsTransparency & auditability

See how loaded prompts create AI feedback loops—and learn to verify sources, inspect reasoning and prevent polished hallucinations from driving decisions.

Why can a chatbot turn a hunch into an extraterrestrial finding? 👽

A chatbot can turn a hunch into a finding by developing the premise in each prompt instead of testing whether the premise is sound. The person then treats the chatbot's response as fresh evidence and uses it to make the next prompt more certain.

In the CountCast experiment, Alice Ferrier began with reports of UFO sightings in Cheshire in the 1990s. She then asked whether the Ministry of Defence was hiding something. The conversation eventually moved from possible sightings to an "almost certain" Belgian government cover-up involving NATO and GCHQ.

The chatbot did not uncover evidence that established this conclusion. It followed the direction set by the conversation and connected details that appeared to support it.

What is an epistemic feedback loop?

An epistemic feedback loop forms when a person supplies a belief, the chatbot develops it, and the chatbot's answer strengthens the person's next prompt. Each exchange appears to add knowledge even when the conversation is recycling its original assumptions.

Epistemic questions ask for more than an answer. They also ask how the answer is known, where the information came from and why it should be trusted.

Humans make those judgements using what they know about a speaker. A physicist and a stranger in a pub can make the same claim about extraterrestrial life, but their expertise and condition affect how much weight the claim receives. A chatbot may not know whether the person prompting it is an expert or an enthusiastic amateur. The person may know just as little about the sources behind the chatbot's answer.

How do loaded questions affect an AI answer?

A loaded question pushes a chatbot toward accepting its premise before it begins answering. Asking whether a government is containing a leak already assumes there may be a leak to contain.

The UFO experiment began with a request to confirm an assertion about sightings. Later prompts asked whether missing information was deliberate. Alice also told the chatbot that it had mentioned triangular craft when it had not.

The chatbot initially recognised that it had not mentioned triangles. It then tried to reconcile the false recollection with earlier references to wedge shapes. That accommodation gave the conversation another route towards confirmation.

A better check separates the observation from the proposed explanation. Ask what evidence supports the explanation, what evidence weakens it and what information is missing.

Does narrated reasoning prove that an AI answer is sound?

Narrated reasoning does not prove that a chatbot reached its answer through valid logic. Language models can say “wait,” “I've got it” or “absolutely” while producing an explanation that only sounds like careful thought.

The episode opens with a riddle discussed by Dr Melanie Mitchell in the Science article Artificial Intelligence Learns to Reason. Julia has two sisters and one brother. Julia's brother has three sisters, including Julia. Earlier language models struggled with that relationship even while producing plausible explanations.

Models can also hand calculations to Python in the way a person reaches for a calculator. That can improve the mechanical part of an analysis. It does not establish that the model chose the right calculation, read every input correctly or interpreted the result soundly.

Why is a polished AI answer still risky?

A polished answer can make unsupported information look ready for use. Formatting, confident prose and official presentation are not evidence that the underlying claims were checked.

The episode gives several examples.

  • Damien Charlotin maintains a database of more than 2,000 court decisions involving lawyers caught using AI hallucinations, including fabricated case law and misquoted justices.
  • The Chicago Sun-Times published a list of 15 summer books in which 10 titles did not exist.
  • Deloitte produced a report for the Australian government that contained nonexistent academic reports and invented court references. The report had cost nearly half a million dollars, and Deloitte had to make a partial repayment.
  • In the episode's military example, an analyst used the same chatbot to format invented intelligence as an official-looking summary. The presentation helped the false claim travel through command channels.

The failure in each case was not a lack of confident language. The evidence underneath the presentation had not been adequately verified.

How should you verify a consequential AI answer?

A consequential AI answer should be broken into claims that another person can trace back to evidence. The reviewer should inspect the connection between those claims rather than accepting the conclusion as a single polished block.

Ask the questions a detective would ask.

  • What evidence is actually on the board?
  • Where did each claim come from?
  • Does the cited material support the claim being made?
  • What evidence or context is missing?
  • Did the conclusion become more certain without new evidence?

An absence of information does not establish a cover-up. A real source also does not automatically support the conclusion attached to it. Verification requires checking both the source and the reasoning that connects it to the answer.

How does visible file analysis make AI work easier to inspect?

Visible file analysis lets a reviewer see the source files, the calculations and the resulting output in the same workspace. Count can upload files such as PDFs, spreadsheets and Microsoft Office documents to a canvas, then let an agent read or modify them through Python.

The CountCast demonstrates this with fictional expenses, an Excel budget and a vendor renewal letter. The agent reads the files, writes Python and shows the analysis on the canvas. A person can inspect the code, review the charts and revise the work without relying only on the final written answer.

Files introduce their own risks. An agent working through a large document may find one relevant passage and stop before reaching later information that would change the conclusion. A Python package may also modify a file imperfectly. Visible work makes those failures easier to investigate, but it does not remove the need for review.

Who should review AI reasoning before a decision?

A person who was not part of the original feedback loop should review the evidence and reasoning before the answer drives a consequential decision. A third-party reviewer can challenge assumptions that the user and chatbot have reinforced together.

Human meetings often have spectators who can notice when everyone is agreeing with the most senior person in the room. A private chatbot conversation removes that outside perspective. Sharing the inputs, calculations and reasoning gives a colleague something concrete to inspect instead of asking them to trust a summary.

Where can I see the worked examples?

The worked UFO conversation, hallucination examples and file analysis demonstration are in How We Talked AI Into a UFO Cover-Up | The CountCast. Watch the video to see how possibility becomes apparent certainty and how visible evidence changes the review process.