Guide: AI in business › AI law, risk and ethics
AI hallucinations are answers in which a language model confidently states something untrue: an invented fact, a non-existent source, a false quote or a wrong figure. They will not disappear entirely. They can, however, be greatly reduced and caught before they reach a customer, a journalist or the board, and this article shows where they come from, how often they occur and what simple process protects your company.
In 60 seconds Hallucinations in a nutshell
- What it is: a confident but false answer from a model, from an invented date to a non-existent publication.
- Who it applies to: anyone who releases AI-assisted text externally, and in practice most often people in marketing, PR and customer service, who write a lot and under time pressure.
- The first habit: anything that can be checked (a figure, a name, a quote, a source) is checked by a human before the text leaves the company.
- The most common mistake: trusting a confident tone. The model sounds the same whether it is right or making things up.
- Who pays for the mistake: the company, not “the system”. The 2024 Air Canada case proved it.
What are AI hallucinations?
Artificial intelligence hallucinations (also called AI hallucination, and for language models also LLM hallucination) are situations in which a model produces a plausible-sounding but untrue answer. OpenAI defines them as “plausible but false statements generated by language models”. An LLM, or large language model, is the engine behind ChatGPT, Gemini and Copilot. We have gathered the wider picture, from choosing tools to rolling them out across a team, in guide to artificial intelligence in business.
A hallucination is not a lie in the human sense. The model does not know it is wrong. It predicts the most probable sequence of words, and that sequence sometimes adds up to a report that does not exist, a launch date that never happened or a CEO quote nobody ever said. The worst part? A wrong answer sounds exactly as confident as a correct one.
- Invented facts. A date, figure, name or financial result that never existed.
- Fabricated AI sources. The title of an article, report or ruling that does not exist, often with a credible-looking web address.
- Misattribution. The source is real, it just doesn’t say what the model claims.
- Outdated information. Information from a few months ago presented as current.
Why AI makes things up
In September 2025 OpenAI published an analysis of why language models hallucinate. The conclusion is simple and slightly uncomfortable. Models are trained and evaluated in a way that makes guessing pay off: if a test rewards only correct answers, “I don’t know” always scores zero, while a guess at least has some chance. So the model learns to guess.
The second reason is the way a model “remembers”. It masters general language patterns well: grammar, style, typical associations. Rare individual facts, such as the founding date of a mid-sized company from Podkarpacie or the name of its sales director, it has seen in its data once or not at all, so this is exactly where hallucinations in ChatGPT and other chatbots happen most often.
The third reason lies with the person asking. A prompt with a hidden thesis (“give me three studies that confirm…”) encourages the model to fit its answer rather than push back. We show how to write prompts that don’t provoke this in our guide to prompts for marketing.
How often ChatGPT and other models hallucinate
There is no single figure. The hallucination rate depends on the model, the type of task and whether the model has source documents in front of it. Below are three independent measurements that illustrate the scale well.
| Measurement | What was tested | Result |
|---|---|---|
| OpenAI, September 2025 | answers to short factual questions (SimpleQA test) | o4-mini: 75% errors with 1% refusals; gpt-5-thinking-mini: 26% errors with 52% refusals |
| Stanford RegLab and HAI, May 2024 | specialised AI legal research tools | from over 17% to over 34% of answers with hallucinations |
| Vectara, May 2026 | summarising a supplied document | from 1.8% to 24.2% of summaries with hallucinations, depending on the model |
Two high-profile cases where a hallucination cost a company
In February 2024 a Canadian tribunal ruled in Moffatt v. Air Canada. The chatbot on the airline’s website had wrongly told a customer that he could apply for a bereavement fare discount within 90 days of travel. The airline argued that the chatbot was a “separate legal entity” responsible for its own words. The tribunal rejected this: the company is responsible for information on its website, regardless of whether it was provided by a person or a program. The compensation was small, around 650 Canadian dollars. Even so, the case made headlines around the world.
In autumn 2025 Deloitte agreed to refund the final instalment of its fee for a report prepared for the Australian government. The document was found to contain, among other things, non-existent academic publications and a fabricated quote from a court ruling, and the firm admitted it had used the GPT-4o model in its work. The errors were spotted by an academic at the University of Sydney.
What do the two stories have in common? A hallucination stops being a technical problem the moment it reaches a customer. From then on it is a reputational matter and has to be handled as one, with a prepared statement and a person to deliver it. How to prepare for this before anything happens is covered in our guide to crisis PR management.
Who hallucinations are a problem for, and in what situations
The risk rises in two places: where AI content leaves the company unchecked, and where someone makes a decision based on it. Before publishing, ask yourself one question. Is there anything in the text that a journalist or customer could check in five minutes? If so, check it yourself first.
| Situation | Hallucination risk | What to do |
|---|---|---|
| Press release with market data | high (figures and sources) | check every figure against the original, give the source with a date |
| Customer service chatbot answering about prices and terms | high, because the company is liable for the answer | the chatbot answers only from an approved knowledge base |
| Summary of a long report | medium (omissions and distortions) | compare the main points with the original |
| Draft post or campaign headlines | low, few facts | standard editing |
| Competitor analysis from the model's memory | high (rare facts about companies) | only from supplied sources, with references |
How to reduce AI hallucinations step by step
There is no “hallucination-free” switch. What works is a set of simple habits. Each one on its own changes little, but together they bring the risk down to a level a team can live with day to day, without feeling that every model answer has to be defused like a mine.
- STEP 01Give the model sources
Instead of asking the model from memory, paste in or connect documents: a report, a price list, a memo. Companies that do this systematically build an assistant based on their own documents, and we explain how this works and when it pays off in our article on RAG, or AI built on company knowledge.
- STEP 02Allow it to say ‘I don't know’
Add one sentence to your prompt: “if you are not sure, or the information is not in the documents, say that you don’t know”. This responds directly to the mechanism described by OpenAI.
- STEP 03Ask for references and open them
Open every source the model gives you and check two things: whether it exists at all and whether it really says what the model claims, because a credible-looking web address proves nothing on its own.
- STEP 04Narrow the task
One specific task per prompt. “Summarise chapter 3” produces fewer errors than “write everything about the market”.
- STEP 05Always verify four things
Figures, names, quotes, dates. Checking an AI answer on these four points takes a few minutes and catches the most dangerous errors.
- STEP 06A person approves and signs off
Nothing goes out without sign-off from a specific, named person. This is also a condition for being exempt from labelling AI-generated texts, which you can read more about in our article on labelling AI content under the AI Act.
How much protection against hallucinations costs
The cheapest protection is discipline. A few minutes of checking for every text containing facts costs less than a single correction. The budget only grows when a company launches a customer chatbot or an assistant built on its own documents, because it needs to prepare a knowledge base, test answers to typical questions and appoint someone to supervise it on an ongoing basis.
| Item | What drives the cost | When it's needed |
|---|---|---|
| Content verification by an editor | number of publications with data, how specialised the topic is | always, when content goes out of the company |
| Team training | number of people, level of AI knowledge | at the first rollout and when new tools are introduced |
| Assistant based on company documents (RAG) | number and quality of documents, integrations, testing | when AI is to answer customers or employees |
| Testing chatbot answers | number of scenarios, how often the offer changes | before launch and after every change to the price list or terms and conditions |
| Monitoring what AI says about the brand | number of questions and models to check | when customers look for companies via AI chatbots |
When a model makes things up about your brand
Hallucinations affect not only the texts a company writes, but also what models say about the company itself. A customer asks a chatbot about suppliers in the sector and gets an invented offer, an outdated head office or the wrong CEO’s name. A Commplace measurement from 21 September 2026 shows that 28 of the 40 phrases for which the Commplace blog ranks in Google’s top twenty have an AI answer in the results. So people increasingly read the model’s version before they see the company’s website.
Models make things up less often about brands that have plenty of credible, consistent publications, which is why protecting a company against hallucinations is largely PR work: media coverage, up-to-date information on the website, consistent descriptions in directories. We explain the mechanism in more detail in our article on why media coverage determines what ChatGPT says about your company. Start with measurement. AI visibility audit means asking models the questions your customers ask and recording what they answer.
When you can trust AI more — and when you can't
AI works well where you supply the material and judge the result yourself: editing your own text, summarising meeting notes, headline variants, translation followed by a proofread. Hallucinations are rarer here and easy to spot, because you know the original.
Caution is needed with rare facts, regulations, current events, market figures and specific people. Where money, health, the law or reputation depend on the answer, the model can prepare a draft, but a human makes the decision after checking the sources. You will find a fuller map of the risks in our piece on the risks and ethics of AI in business, and the legal obligations are covered in our piece on The AI Act in Poland.
Choosing a tool for your team? You will find a comparison of plans in our overview of ChatGPT, Gemini or Copilot for business. Planning a chatbot for customers? First read how to design chatbots in customer service. We teach marketing and communications teams to treat verification as a habit, not an exception, in AI training for companies.
Habits that let AI fabrications slip into publications
A list of sources with web addresses looks professional, yet some of the entries may not exist. All it takes is for a journalist or a customer to click one link, and the credibility of the entire piece collapses.
A model sounds just as sure when it is right as when it is wrong. Anyone who judges an answer by its style rather than by checked facts lets through exactly the errors that hurt most later, because they end up in texts signed with the company’s name.
A chatbot answering from the model’s general knowledge can promise a discount or a condition that is not in your terms and conditions. The company is liable for that promise. Air Canada found this out before a tribunal.
A prompt like “give me studies confirming that…” encourages the model to bend facts to fit the thesis. The result is an argument built on data that does not exist.
“Written by AI, checked by the team” often means nobody checked the figures. Without a name on the sign-off, an error has no owner. And with no owner, the process never improves.
Questions about hallucinations we hear from marketing teams
What are AI hallucinations?
They are language model answers that sound credible but are untrue: invented facts, figures, quotes or sources. The model does not know it is wrong. It predicts probable text rather than checking the truth.
Why does ChatGPT hallucinate?
According to an OpenAI analysis from September 2025, models are trained and evaluated in a way that rewards guessing more than admitting they don't know. They are most often wrong about rare facts that appeared little in the training data.
Can AI hallucinations be eliminated completely?
No. They can, however, be significantly reduced: give the model source documents, allow it to answer “I don’t know”, narrow the task, and review the facts (figures, names, quotes, dates) before publication.
How can you check whether an AI answer is true?
Open every source provided and check whether it exists and whether it says what the model claims. Compare figures and quotes with the original. Can’t find the source? Treat the information as unconfirmed.
Who is liable for a chatbot's mistake on a company website?
The company. In Moffatt v. Air Canada in February 2024, a Canadian tribunal rejected the argument that the chatbot was a separate entity and ordered the airline to pay damages for the incorrect information.
Which tasks are most prone to hallucinations?
Questions about rare facts, specific people and companies, regulations, current events and numerical data, especially when the model answers from memory. The least risky work is on material you provide, for example editing or summarising your own text.
Do paid versions of models hallucinate less often?
Newer models with a reasoning mode and search access are often wrong less frequently, but measurements show large differences depending on the task, so no version exempts you from verifying content that goes out externally.
Sources
- SOURCEOpenAI: Why language models hallucinate · 5 September 2025
- SOURCE
- SOURCEVectara: Hallucination Leaderboard · updated 11 May 2026
- SOURCE
- SOURCE
- SOURCE
Read next
In our training sessions we work on materials from the team's day-to-day work: we show where models make things up most often and how to build a simple verification process.
Sebastian Kopiej, CEO of Commplace®. In public relations since 1996. Written with the help of AI tools and editorially verified by the author. Data current as of 24 September 2026.