Why AI Chatbots Make Things Up, and How Cited Answers Help

AI hallucination is a business risk: chatbots guess when unsure. Why it happens, what Mata v. Avianca showed, and why cited answers help but need checking.

Verika Editorial··8 min read

Short answer: AI chatbots make things up because they generate likely-sounding text, and their training rewards a confident guess over "I don't know", as OpenAI researchers explained in 2025. Answers that cite your own documents make errors easier to catch, but they do not remove them. The reader still has to open the source and check.

For a business, the risk is not that AI is sometimes wrong. People are sometimes wrong too. The risk is that a wrong AI answer looks exactly like a right one, and it arrives fast enough that nobody checks. This guide explains why that happens, shows a real case where it went badly, and sets out what cited answers do and do not fix.

What is an AI hallucination?

It is a false or unsupported statement presented as fact. The US National Institute of Standards and Technology uses the term confabulation and defines it as "the production of confidently stated but erroneous or false content (known colloquially as 'hallucinations' or 'fabrications') by which users may be misled or deceived" (NIST AI 600-1, July 2024).

The important word is "confidently". A chatbot does not usually signal when it is unsure. An invented policy number, a made-up court case and a correct answer all come out in the same calm tone.

Common forms in business use:

What you seeExampleRisk
A fact with no source"Your return window is 30 days"High: nothing to check against
A source that doesn't existA report, case or regulation with a plausible titleHigh: looks verified but isn't
A real source that doesn't say thatA genuine policy cited for a rule it doesn't containHigh: passes a quick glance
A real source, wrong versionLast year's price list or a superseded regulationMedium: correct once, wrong now
A real source, correctly quotedThe passage says what the answer saysLower: still read the context
"I couldn't find that"No answer givenLowest: go and ask a person

Why do chatbots make things up?

Because they are built to produce plausible text, and the way they are scored rewards guessing.

NIST explains the first part: generative models "generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase." That can produce accurate output, but it "can also produce outputs that are factually inaccurate or internally inconsistent", especially for open-ended questions and in domains that need "highly contextual and/or domain expertise" (NIST AI 600-1, section 2.2). Questions about your company's rules are about as contextual as it gets.

The second part comes from OpenAI's own researchers. In a September 2025 paper, Kalai, Nachum, Vempala and Zhang wrote: "Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty." They argue that models hallucinate "because the training and evaluation procedures reward guessing over acknowledging uncertainty", and that "language models are optimized to be good test-takers, and guessing when uncertain improves test performance" (Kalai et al., "Why Language Models Hallucinate", arXiv:2509.04664).

Their opening example is simple. They asked an open-source model for one of the authors' birthdays, "if you know". On three attempts it gave three different wrong dates, even though it had been told to answer only if it knew.

Why your business questions are especially exposed

A general chatbot has never seen your staff handbook, your price list or your client contracts. Ask it "what's our notice period?" and it can only produce a typical answer for a typical company. It will often do so without saying that it is guessing. That is the gap a tool that answers from your own documents tries to close; see what an AI knowledge assistant is for how that works.

What happened when lawyers trusted a chatbot?

They were sanctioned. In Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), in the US District Court for the Southern District of New York, Judge P. Kevin Castel issued an Opinion and Order on Sanctions on 22 June 2023.

The facts, from the opinion itself (Mata v. Avianca, ECF 54):

  • Roberto Mata sued the airline Avianca over a knee injury from a metal serving cart on a 2019 flight. Avianca moved to dismiss the claim as time-barred.
  • On 1 March 2023 Mata's lawyers filed an opposition that cited court decisions that did not exist. The court found the lawyers had "submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT".
  • On 15 March, Avianca's lawyers told the court they could not locate most of the cases cited.
  • The lawyers then filed what they presented as copies of the decisions. These too were invented.
  • One lawyer had asked ChatGPT whether one of the cases was real. According to the opinion, it replied that it had supplied "real" authorities that could be found on Westlaw and LexisNexis.

The judge was clear that the tool was not the core problem: "Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance. But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings." The lawyers, he wrote, "abandoned their responsibilities" and "continued to stand by the fake opinions after judicial orders called their existence into question."

The court imposed a $5,000 penalty, jointly and severally, on the two lawyers and their firm, and ordered them to send letters to their client and to each real judge falsely named as the author of a fake opinion. In a separate order the same day, the court granted Avianca's motion to dismiss because the claim was filed after the Montreal Convention's two-year limit (Mata v. Avianca, ECF 55).

Three lessons carry over to any business:

  1. Asking the chatbot to check itself is not checking. The same system that invented the cases confirmed they were real.
  2. The person who signs is responsible. The court did not sanction ChatGPT.
  3. The cover-up was worse than the mistake. The opinion says things "would look quite different" if the lawyers had come clean soon after their sources were questioned.

Do cited answers fix the problem?

They reduce it and make the remaining errors easier to catch. They do not remove it.

The best evidence comes from legal research, where tools that search real case law and cite it have been tested independently. A Stanford-led team (Magesh, Surani, Dahl, Suzgun, Manning and Ho) ran what they describe as the first preregistered evaluation of AI legal research tools. They found that hallucinations "are reduced relative to general-purpose chatbots (GPT-4)", but that tools from LexisNexis and Thomson Reuters "each hallucinate between 17% and 33% of the time" (Magesh et al., Journal of Empirical Legal Studies, 2025).

Their definition is useful for any business. A response counts as hallucinated "if it is either incorrect or misgrounded", and "if a model makes a false statement or falsely asserts that a source supports a statement, that constitutes a hallucination." In other words, a citation can itself be the error. A real document cited for something it does not say is a hallucination with a nice footnote.

AI developers say the same about their own models. Anthropic's guidance for developers recommends letting the model say "I don't know", grounding answers in direct quotes, and restricting it to supplied documents, then adds that "while these techniques significantly reduce hallucinations, they don't eliminate them entirely. Always validate critical information, especially for high-stakes decisions" (Anthropic, Reduce hallucinations).

So the honest case for citations is this: they turn "trust me" into "here's where to look". That only helps if someone looks.

How should a business use AI answers safely?

Match the checking to the stakes, and make checking easy.

  1. Decide which questions need a source. Anything about money, safety, legal duties, client commitments or health needs one. Drafting an email does not.
  2. Prefer tools that answer from your documents rather than from general training, and that refuse when they find nothing.
  3. Open the source for anything that matters. Read the cited passage. Check it says what the answer says and that it is the current version.
  4. Watch for a citation you can't open. If the link or document doesn't exist, treat the whole answer as invented.
  5. Don't ask the AI to confirm itself. Check against the source or a person, as Mata showed.
  6. Keep sources current. An assistant reading an out-of-date document gives an out-of-date answer with a real citation.
  7. Log the questions it couldn't answer. Those are gaps in your documents, and filling them once helps everyone.
  8. Make one person accountable for each important answer, the way a signature makes a lawyer accountable for a filing.
  9. Tell staff it's fine to say "I checked and it was wrong". Errors reported early are cheap.

How does Verika handle this?

Verika is our product, so read this with that in mind. It answers your team's questions only from your company's documents and from answers your experts have written and the owner has approved, and every answer shows the document or person it came from. If nothing in your sources supports an answer, it says so instead of guessing and logs the question as a gap. That design targets the failure in Mata, but it does not make checking unnecessary: if a manager asks "what discount can I offer a repeat client?" and Verika answers from your pricing policy, the manager should still open that policy before putting a number in writing. More on the Verika home page, or try it on your own documents.

Sources

  1. National Institute of Standards and Technology, NIST AI 600-1, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile", July 2024, section 2.2 "Confabulation". https://doi.org/10.6028/NIST.AI.600-1
  2. Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E., "Why Language Models Hallucinate", OpenAI and Georgia Tech, 4 September 2025, arXiv:2509.04664. https://arxiv.org/abs/2509.04664
  3. Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order on Sanctions, ECF 54 (S.D.N.Y. 22 June 2023). https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0.pdf
  4. Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order granting motion to dismiss, ECF 55 (S.D.N.Y. 22 June 2023). https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.55.0_2.pdf
  5. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., and Ho, D. E., "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools", Journal of Empirical Legal Studies, 2025 (accepted 14 March 2025), doi:10.1111/jels.12413. https://doi.org/10.1111/jels.12413
  6. Anthropic, "Reduce hallucinations", Claude developer documentation (accessed 7 October 2026). https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations

Frequently asked questions

›What is an AI hallucination?

It is when an AI system states something false or unsupported as if it were fact. NIST calls it confabulation: confidently stated but erroneous or false content.

›Why do AI chatbots guess instead of saying they don't know?

OpenAI researchers argued in 2025 that standard training and evaluation reward guessing over acknowledging uncertainty, much like a multiple-choice exam with no penalty for wrong answers.

›Do citations stop AI hallucinations?

No. A 2025 Stanford-led study found legal research tools that cite sources still hallucinated between 17% and 33% of the time. Citations make errors easier to catch; you still have to open the source.

›What happened in Mata v. Avianca?

In June 2023 a New York federal judge fined two lawyers and their firm $5,000 after they filed a brief citing court decisions that ChatGPT had invented, then stood by them after the court questioned their existence.

Keep reading