Home / Blog / AI Agent Hallucinations in Contact Centres: What They Actually Look Like and How to Stop Them
AI Voice

AI Agent Hallucinations in Contact Centres: What They Actually Look Like and How to Stop Them

Your AI voice agent just told a customer their warranty covers accidental damage. It doesn't. Here's how to recognise hallucinations before they cause real harm.

By Hostcomm

Last month, a UK insurance company's AI agent told a customer their policy covered flood damage. It didn't. The customer made decisions based on that information. The complaint went to the Financial Ombudsman.

This is what hallucination looks like in production. Not the dramatic "AI goes rogue" scenarios from headlines, but quiet, confident errors that slip through because the AI sounds so plausible.

If you're running AI agents in your contact centre — or planning to — you need to understand exactly what hallucinations look like, why they happen, and how to catch them before they reach customers.

What hallucinations actually look like

In contact centre contexts, AI hallucinations fall into predictable patterns. Knowing these helps you spot them in QA reviews and build prevention into your systems.

Invented policies and procedures. The agent creates rules that sound reasonable but don't exist. "Our standard policy is to offer a 20% goodwill discount after three complaints." No such policy exists — the model generated something plausible-sounding based on patterns in its training data.

Confident wrong numbers. "Your account shows a balance of £247.32." The actual balance is £312.50. The agent didn't check — it generated a number that fit the conversational context. This happens more often when integration latency causes the AI to fill gaps rather than wait.

Made-up product features. "Yes, that model includes built-in WiFi connectivity." It doesn't. The agent extrapolated from similar products in its knowledge base — or from its general training data about the product category.

Fictional case histories. "I can see from your previous call on 15th September that you discussed this issue." No such call exists. The model inferred a plausible history from the conversation context.

Process inventions. "I'll escalate this to our specialist fraud team who will call you within 24 hours." Your organisation has no such team, no such escalation path, no such SLA.

The common thread: these errors are delivered with the same confident tone as accurate information. That's the problem. Customers can't tell the difference.

Why contact centre AI hallucinates

Understanding the causes points directly to prevention strategies.

Knowledge base gaps. When the AI can't find a definitive answer in its grounded knowledge, it falls back on its general training data or generates something that fits the pattern. The fix is obvious but labour-intensive: comprehensive, well-structured knowledge bases with explicit coverage of edge cases.

Ambiguous retrieval. The AI pulls the wrong document chunk because the query matched on keywords rather than intent. Customer asks about "cancellation" meaning contract termination; the system retrieves content about appointment cancellation. The AI confidently delivers the wrong answer.

Temporal confusion. The model doesn't inherently understand that a policy changed last month, or that a promotion ended yesterday. Without explicit date handling in your knowledge architecture, it serves outdated information as current fact.

Over-generalisation. The AI sees patterns and extends them beyond their valid scope. Your returns policy allows exchanges within 30 days for clothing. The model extends this to electronics without being explicitly told it doesn't apply.

Integration failures. When real-time data lookups fail or time out, some systems are configured to continue the conversation rather than acknowledge the gap. The AI generates plausible-sounding data to fill the void.

Building prevention into your architecture

The goal isn't zero hallucinations — that's not achievable with current technology. The goal is catching them before they reach customers and minimising the damage when they slip through.

Explicit "I don't know" training. Your AI agent needs permission — and training — to say it doesn't know. This requires careful prompt engineering and ideally fine-tuning on examples where the correct response is uncertainty. Too many implementations optimise for sounding helpful, which creates pressure toward confident invention.

Retrieval validation layers. Before the AI uses retrieved information, verify the match quality. If the retrieval confidence is below threshold, don't serve the answer — escalate to human review or ask a clarifying question. This catches the ambiguous retrieval problem.

Structured fact extraction. For critical information — prices, policy details, account balances — don't let the AI generate these conversationally. Pull them from structured data sources and insert them into the response. The AI handles the conversation; verified systems handle the facts.

Action validation. When an AI agent claims it's going to do something ("I've initiated your refund"), require programmatic confirmation that the action actually succeeded. Don't let the AI confirm completion of actions it can't verify.

Knowledge base versioning. Implement explicit date ranges on all knowledge content. When did this policy start? When does it expire? The AI should refuse to answer questions about policies outside their valid period rather than guess.

Detection and monitoring

You can't prevent what you can't see. Your QA processes need to evolve for AI-specific failure modes.

Factual accuracy sampling. Select a random sample of AI interactions daily and verify every factual claim against source systems. Not spot-checks — systematic verification. This gives you a baseline hallucination rate and catches patterns before they scale.

Customer correction tracking. When customers say "that's not right" or "are you sure?" during AI interactions, flag these automatically for review. Customers often catch hallucinations in real-time.

Comparison testing. Periodically run the same queries through your AI agent and a verified knowledge source. Divergence indicates potential hallucination. Automate this for high-stakes query types.

Confidence scoring. Most modern LLM architectures can provide confidence metrics. Route low-confidence responses to human review rather than serving them directly. The trade-off is speed; the benefit is accuracy.

Post-interaction verification. For high-value transactions or compliance-sensitive interactions, implement automated checks that verify what the AI said against ground truth after the conversation ends. This catches errors that escaped real-time detection.

The human-in-the-loop reality

Here's the uncomfortable truth: for anything with regulatory, financial, or significant customer impact, current AI requires human validation of key decisions.

That doesn't mean humans review every interaction. It means your system architecture recognises which decisions matter and routes those for human confirmation. The AI handles routine queries end-to-end. Anything involving money, policy exceptions, or compliance triggers human review.

This isn't a limitation to work around — it's the architecture that makes AI deployments safe. Contact centres that skip this step, chasing full automation metrics, are the ones generating FCA complaints and brand damage.

Measuring what matters

Track these metrics weekly:

  • Hallucination rate: Percentage of sampled interactions containing factual errors. Baseline target: under 2% for grounded queries.
  • Catch rate: Percentage of hallucinations caught by your prevention systems before customer delivery. Target: over 90%.
  • Customer correction rate: How often customers flag AI errors during conversations. Trending up means your prevention is failing.
  • Confidence distribution: What percentage of queries fall below your confidence threshold? This tells you where your knowledge base needs work.

Getting it right

AI hallucinations aren't going away. They're a fundamental characteristic of how large language models work — statistical pattern completion that sometimes completes in the wrong direction.

The organisations getting value from AI agents aren't the ones claiming zero errors. They're the ones with robust architectures for catching errors, clear escalation paths when uncertainty hits, and monitoring systems that spot problems before they compound.

If you're evaluating AI voice agents or contact centre automation, ask vendors hard questions about hallucination rates in production, not demo conditions. Ask how their systems handle uncertainty. Ask to see their monitoring dashboards.


Hostcomm's CXCortex platform includes automated quality monitoring that flags potential hallucinations and tracks accuracy metrics across your AI interactions. Talk to our team about building prevention into your contact centre AI deployment.