navigation

AI hallucinations: why ChatGPT makes things up and how to protect your business

Guide: AI in business › AI law, risk and ethics

ARTIFICIAL INTELLIGENCE IN MARKETING

AI hallucinations are answers in which a language model confidently states something untrue: an invented fact, a non-existent source, a false quote or a wrong figure. They will not disappear entirely. They can, however, be greatly reduced and caught before they reach a customer, a journalist or the board, and this article shows where they come from, how often they occur and what simple process protects your company.

11 min readUpdated: 24 September 2026Author: Sebastian Kopiej
75%of answers from the o4-mini model were wrong in a factual-questions test; a newer model that says ‘I don't know’ more often was wrong 26% of the timeBENCHMARKOpenAI, ‘Why language models hallucinate’, 5 September 2025
17–34%of answers from specialised AI legal research tools contained hallucinationsBENCHMARKStanford RegLab and HAI, 23 May 2024
1,8–24,2%share of summaries containing hallucinations, depending on the modelBENCHMARKVectara Hallucination Leaderboard, 11 May 2026

In 60 seconds Hallucinations in a nutshell

  • What it is: a confident but false answer from a model, from an invented date to a non-existent publication.
  • Who it applies to: anyone who releases AI-assisted text externally, and in practice most often people in marketing, PR and customer service, who write a lot and under time pressure.
  • The first habit: anything that can be checked (a figure, a name, a quote, a source) is checked by a human before the text leaves the company.
  • The most common mistake: trusting a confident tone. The model sounds the same whether it is right or making things up.
  • Who pays for the mistake: the company, not “the system”. The 2024 Air Canada case proved it.

What are AI hallucinations?

Artificial intelligence hallucinations (also called AI hallucination, and for language models also LLM hallucination) are situations in which a model produces a plausible-sounding but untrue answer. OpenAI defines them as “plausible but false statements generated by language models”. An LLM, or large language model, is the engine behind ChatGPT, Gemini and Copilot. We have gathered the wider picture, from choosing tools to rolling them out across a team, in guide to artificial intelligence in business.

A hallucination is not a lie in the human sense. The model does not know it is wrong. It predicts the most probable sequence of words, and that sequence sometimes adds up to a report that does not exist, a launch date that never happened or a CEO quote nobody ever said. The worst part? A wrong answer sounds exactly as confident as a correct one.

  • Invented facts. A date, figure, name or financial result that never existed.
  • Fabricated AI sources. The title of an article, report or ruling that does not exist, often with a credible-looking web address.
  • Misattribution. The source is real, it just doesn’t say what the model claims.
  • Outdated information. Information from a few months ago presented as current.

Why AI makes things up

In September 2025 OpenAI published an analysis of why language models hallucinate. The conclusion is simple and slightly uncomfortable. Models are trained and evaluated in a way that makes guessing pay off: if a test rewards only correct answers, “I don’t know” always scores zero, while a guess at least has some chance. So the model learns to guess.

The second reason is the way a model “remembers”. It masters general language patterns well: grammar, style, typical associations. Rare individual facts, such as the founding date of a mid-sized company from Podkarpacie or the name of its sales director, it has seen in its data once or not at all, so this is exactly where hallucinations in ChatGPT and other chatbots happen most often.

The third reason lies with the person asking. A prompt with a hidden thesis (“give me three studies that confirm…”) encourages the model to fit its answer rather than push back. We show how to write prompts that don’t provoke this in our guide to prompts for marketing.

How often ChatGPT and other models hallucinate

There is no single figure. The hallucination rate depends on the model, the type of task and whether the model has source documents in front of it. Below are three independent measurements that illustrate the scale well.

MeasurementWhat was testedResult
OpenAI, September 2025answers to short factual questions (SimpleQA test)o4-mini: 75% errors with 1% refusals; gpt-5-thinking-mini: 26% errors with 52% refusals
Stanford RegLab and HAI, May 2024specialised AI legal research toolsfrom over 17% to over 34% of answers with hallucinations
Vectara, May 2026summarising a supplied documentfrom 1.8% to 24.2% of summaries with hallucinations, depending on the model

Two high-profile cases where a hallucination cost a company

In February 2024 a Canadian tribunal ruled in Moffatt v. Air Canada. The chatbot on the airline’s website had wrongly told a customer that he could apply for a bereavement fare discount within 90 days of travel. The airline argued that the chatbot was a “separate legal entity” responsible for its own words. The tribunal rejected this: the company is responsible for information on its website, regardless of whether it was provided by a person or a program. The compensation was small, around 650 Canadian dollars. Even so, the case made headlines around the world.

In autumn 2025 Deloitte agreed to refund the final instalment of its fee for a report prepared for the Australian government. The document was found to contain, among other things, non-existent academic publications and a fabricated quote from a court ruling, and the firm admitted it had used the GPT-4o model in its work. The errors were spotted by an academic at the University of Sydney.

What do the two stories have in common? A hallucination stops being a technical problem the moment it reaches a customer. From then on it is a reputational matter and has to be handled as one, with a prepared statement and a person to deliver it. How to prepare for this before anything happens is covered in our guide to crisis PR management.

Who hallucinations are a problem for, and in what situations

The risk rises in two places: where AI content leaves the company unchecked, and where someone makes a decision based on it. Before publishing, ask yourself one question. Is there anything in the text that a journalist or customer could check in five minutes? If so, check it yourself first.

SituationHallucination riskWhat to do
Press release with market datahigh (figures and sources)check every figure against the original, give the source with a date
Customer service chatbot answering about prices and termshigh, because the company is liable for the answerthe chatbot answers only from an approved knowledge base
Summary of a long reportmedium (omissions and distortions)compare the main points with the original
Draft post or campaign headlineslow, few factsstandard editing
Competitor analysis from the model's memoryhigh (rare facts about companies)only from supplied sources, with references

How to reduce AI hallucinations step by step

There is no “hallucination-free” switch. What works is a set of simple habits. Each one on its own changes little, but together they bring the risk down to a level a team can live with day to day, without feeling that every model answer has to be defused like a mine.

  1. STEP 01Give the model sources

    Instead of asking the model from memory, paste in or connect documents: a report, a price list, a memo. Companies that do this systematically build an assistant based on their own documents, and we explain how this works and when it pays off in our article on RAG, or AI built on company knowledge.

  2. STEP 02Allow it to say ‘I don't know’

    Add one sentence to your prompt: “if you are not sure, or the information is not in the documents, say that you don’t know”. This responds directly to the mechanism described by OpenAI.

  3. STEP 03Ask for references and open them

    Open every source the model gives you and check two things: whether it exists at all and whether it really says what the model claims, because a credible-looking web address proves nothing on its own.

  4. STEP 04Narrow the task

    One specific task per prompt. “Summarise chapter 3” produces fewer errors than “write everything about the market”.

  5. STEP 05Always verify four things

    Figures, names, quotes, dates. Checking an AI answer on these four points takes a few minutes and catches the most dangerous errors.

  6. STEP 06A person approves and signs off

    Nothing goes out without sign-off from a specific, named person. This is also a condition for being exempt from labelling AI-generated texts, which you can read more about in our article on labelling AI content under the AI Act.

How much protection against hallucinations costs

The cheapest protection is discipline. A few minutes of checking for every text containing facts costs less than a single correction. The budget only grows when a company launches a customer chatbot or an assistant built on its own documents, because it needs to prepare a knowledge base, test answers to typical questions and appoint someone to supervise it on an ongoing basis.

ItemWhat drives the costWhen it's needed
Content verification by an editornumber of publications with data, how specialised the topic isalways, when content goes out of the company
Team trainingnumber of people, level of AI knowledgeat the first rollout and when new tools are introduced
Assistant based on company documents (RAG)number and quality of documents, integrations, testingwhen AI is to answer customers or employees
Testing chatbot answersnumber of scenarios, how often the offer changesbefore launch and after every change to the price list or terms and conditions
Monitoring what AI says about the brandnumber of questions and models to checkwhen customers look for companies via AI chatbots

When a model makes things up about your brand

Hallucinations affect not only the texts a company writes, but also what models say about the company itself. A customer asks a chatbot about suppliers in the sector and gets an invented offer, an outdated head office or the wrong CEO’s name. A Commplace measurement from 21 September 2026 shows that 28 of the 40 phrases for which the Commplace blog ranks in Google’s top twenty have an AI answer in the results. So people increasingly read the model’s version before they see the company’s website.

Models make things up less often about brands that have plenty of credible, consistent publications, which is why protecting a company against hallucinations is largely PR work: media coverage, up-to-date information on the website, consistent descriptions in directories. We explain the mechanism in more detail in our article on why media coverage determines what ChatGPT says about your company. Start with measurement. AI visibility audit means asking models the questions your customers ask and recording what they answer.

When you can trust AI more — and when you can't

AI works well where you supply the material and judge the result yourself: editing your own text, summarising meeting notes, headline variants, translation followed by a proofread. Hallucinations are rarer here and easy to spot, because you know the original.

Caution is needed with rare facts, regulations, current events, market figures and specific people. Where money, health, the law or reputation depend on the answer, the model can prepare a draft, but a human makes the decision after checking the sources. You will find a fuller map of the risks in our piece on the risks and ethics of AI in business, and the legal obligations are covered in our piece on The AI Act in Poland.

Choosing a tool for your team? You will find a comparison of plans in our overview of ChatGPT, Gemini or Copilot for business. Planning a chatbot for customers? First read how to design chatbots in customer service. We teach marketing and communications teams to treat verification as a habit, not an exception, in AI training for companies.

Habits that let AI fabrications slip into publications

1

Asking the model for sources and not checking them

A list of sources with web addresses looks professional, yet some of the entries may not exist. All it takes is for a journalist or a customer to click one link, and the credibility of the entire piece collapses.

2

Treating a confident tone as proof

A model sounds just as sure when it is right as when it is wrong. Anyone who judges an answer by its style rather than by checked facts lets through exactly the errors that hurt most later, because they end up in texts signed with the company’s name.

3

A customer chatbot without a closed knowledge base

A chatbot answering from the model’s general knowledge can promise a discount or a condition that is not in your terms and conditions. The company is liable for that promise. Air Canada found this out before a tribunal.

4

Questions with a hidden premise

A prompt like “give me studies confirming that…” encourages the model to bend facts to fit the thesis. The result is an argument built on data that does not exist.

5

No one responsible for the publication

“Written by AI, checked by the team” often means nobody checked the figures. Without a name on the sign-off, an error has no owner. And with no owner, the process never improves.

Questions about hallucinations we hear from marketing teams

What are AI hallucinations?

They are language model answers that sound credible but are untrue: invented facts, figures, quotes or sources. The model does not know it is wrong. It predicts probable text rather than checking the truth.

Why does ChatGPT hallucinate?

According to an OpenAI analysis from September 2025, models are trained and evaluated in a way that rewards guessing more than admitting they don't know. They are most often wrong about rare facts that appeared little in the training data.

Can AI hallucinations be eliminated completely?

No. They can, however, be significantly reduced: give the model source documents, allow it to answer “I don’t know”, narrow the task, and review the facts (figures, names, quotes, dates) before publication.

How can you check whether an AI answer is true?

Open every source provided and check whether it exists and whether it says what the model claims. Compare figures and quotes with the original. Can’t find the source? Treat the information as unconfirmed.

Who is liable for a chatbot's mistake on a company website?

The company. In Moffatt v. Air Canada in February 2024, a Canadian tribunal rejected the argument that the chatbot was a separate entity and ordered the airline to pay damages for the incorrect information.

Which tasks are most prone to hallucinations?

Questions about rare facts, specific people and companies, regulations, current events and numerical data, especially when the model answers from memory. The least risky work is on material you provide, for example editing or summarising your own text.

Do paid versions of models hallucinate less often?

Newer models with a reasoning mode and search access are often wrong less frequently, but measurements show large differences depending on the task, so no version exempts you from verifying content that goes out externally.

Sources

Read next

Does your team know how to catch an AI error before a client sees it?

In our training sessions we work on materials from the team's day-to-day work: we show where models make things up most often and how to build a simple verification process.

Sebastian Kopiej, CEO of Commplace®. In public relations since 1996. Written with the help of AI tools and editorially verified by the author. Data current as of 24 September 2026.

Call me Customer panel Contact