Customer-support chatbot grounding — stop bots from hallucinating product facts
Production support bots hallucinate pricing, release dates, integration details. Add a verify layer on top of RAG over your docs; users see accurate citations or 'I'm not sure'.
The problem
Customer support chatbots hallucinate things customers actually care about: pricing tiers, integration partners, rate limits, supported regions, refund policies, model context windows. A bot telling a paying customer "Sure, your Pro tier supports unlimited API calls" when the real answer is 50k/month creates a billing dispute + trust collapse + maybe a chargeback.
Standard RAG over your help docs catches most factual questions, but the tail of failure modes — wrong number on the right doc, fabricated integration, made-up rate limit — is what dings trust.
The pattern
Three-layer architecture:
- RAG over your docs. Index your help center, API docs, pricing pages. Standard retrieval.
- Check atomic claims in the response. Extract assertions (prices, limits, dates, features) and compare each with an authoritative product-facts catalog. For your product's claims, you maintain the catalog. For AI/ML claims, SourceScore VERITAS can retrieve candidate records, but a separate comparison or human review must decide support.
- Decline unsupported claims. If the bot would emit a claim and the support check fails, the bot says "I'm not sure — let me get a human." The cost of declining is much lower than the cost of being wrong.
Implementation sketch
# Two-catalog setup: your-own + SourceScore VERITAS
# 1. Build your own reviewed catalog of product facts.
# (Pricing, rate limits, feature support, etc.)
# Update it whenever pricing/features change.
# JSON file or simple key-value DB.
PRODUCT_FACTS = {
# Hypothetical placeholders: replace from your current source of truth.
"pro_tier_monthly_calls": "<current documented limit>",
"scale_tier_monthly_calls": "<current documented limit>",
"billing_supported": "<current documented value>",
"free_tier_signup_required": "<current documented value>",
# ...
}
# 2. For AI/ML factual claims (if user asks about Llama 3.1, Claude,
# etc.), use SourceScore VERITAS.
import httpx
def find_aiml_candidate(claim_text: str) -> dict | None:
r = httpx.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": claim_text, "minConfidence": 0.85},
timeout=2.0,
)
result = r.json()
return result.get("bestMatch")
# 3. In the bot's response pipeline:
async def respond(user_question: str):
# Standard RAG
retrieved = await rag.retrieve(user_question, k=5)
draft = await llm.generate(user_question, context=retrieved)
# Extract atomic claims from draft
claims = extract_atomic_claims(draft)
supported = []
needs_review = []
for c in claims:
if c.matches_product_pattern():
ok = verify_against_product_facts(c, PRODUCT_FACTS)
else:
candidate = find_aiml_candidate(c.text)
# Implement entailment or human review here. Similarity alone is
# not proof that candidate.statement supports c.text.
ok = candidate is not None and supports_assertion(c.text, candidate)
(supported if ok else needs_review).append(c)
if needs_review:
# Don't ship the response with claims that still need review
return (
"I'm not 100% certain about one or more facts in my "
"answer. Let me transfer you to a human teammate."
)
return draft # Every claim passed the application's support checkWhat this catches
- Wrong pricing. Bot says "€199/month" when the actual price is "€499/month" — product-facts catalog catches it.
- Hallucinated integrations. Bot says "Yes, we integrate with Zapier" when you don't — catalog catches it.
- Potentially wrong AI/ML facts. VERITAS can surface a nearby cited record for comparison; your support check must decide whether it contradicts or supports the bot.
- Stale info. Bot uses 2-year-old training data for current pricing — catalog (which you update on pricing changes) catches it.
The escape valve: route to human
The bot doesn't need to answer everything. Routing to a human for unverifiable claims is a feature, not a bug. Optimize for supported answers and safe handoffs, not the highest automation percentage. Measure incorrect-answer cost, handoff rate, and time to resolution on your own labeled support conversations.
Free-tier economics
- SourceScore VERITAS public API: free with no account, key, or signup.
- One network request per VERITAS call. Set a timeout and measure in your own stack.
- Your product-facts catalog: cost = engineering time to maintain (small).
- Higher-volume paid access is a demand test only; no paid plan or SLA is live.
Related
- AI agent grounding — broader agent pattern
- RAG pipeline verification — closing the right-doc-wrong-number gap
- LLM grounding concept pillar
- Six grounding strategies blog post