RAG pipeline verification — close the right-doc-wrong-number gap
Your retriever pulls the right document. Your LLM still emits the wrong number on the page. RAG retrieves; it doesn't verify. Add a verify-then-respond layer to close the gap.
The problem
You built RAG. Embedded your corpus, picked a vector DB, tuned top-K, wrote the prompt template. Production users file tickets:
"It told me the model has 32k context. The source it cited literally says 128k."
You read the source. It says 128k. Your retriever found it. Your prompt included it. The model still hallucinated.
This isn't a retrieval failure. It's a verification failure. RAG = Retrieval-Augmented Generation. There's no built-in step that checks the model's output against the context. The inconsistency is invisible.
The pattern: match, review, then respond
Add a third stage to your RAG pipeline:
- Retrieve. Pull top-K from your vector DB. Unchanged.
- Generate. Model produces a response. Unchanged.
- Match and review. Extract atomic assertions, retrieve candidate VERITAS records, then compare statements and cited evidence before labeling support.
Code (Python, ~30 lines)
import re
import httpx
def retrieve_candidates(llm_response: str) -> dict:
# Naive extraction: sentences with "is" / "has" / "released" verbs
sentences = re.split(r'(?<=[.!?])\s+', llm_response)
candidates = [
s for s in sentences
if re.search(r'\b(is|has|released|introduced)\b', s, re.IGNORECASE)
]
candidate_records = []
no_match = []
for claim in candidates:
r = httpx.post(
'https://sourcescore.org/api/v1/verify',
json={'claim': claim, 'minConfidence': 0.85},
timeout=2.0,
)
result = r.json()
if result.get('bestMatch'):
candidate_records.append({
'claim': claim,
'candidate_statement': result['bestMatch']['statement'],
'source_url': result['bestMatch']['detailUrl'],
})
else:
no_match.append(claim)
return {'candidate_records': candidate_records, 'no_match': no_match}
# In your RAG flow:
response = rag_chain.invoke(query)
screen = retrieve_candidates(response)
if screen['no_match']:
response += f"\n\n*Note: {len(screen['no_match'])} claim(s) have no catalog candidate.*"
# Do not attach a verified badge here. Compare each candidate_statement and
# its cited evidence with the submitted claim in a separate review step.What this can surface
Candidate retrieval can expose a dated catalog record, a different number, or a source worth reviewing. It can also return a topically similar record that does not entail the submitted assertion.
- Missing catalog coverage that should route to another retrieval path
- Candidate evidence for dates, attributions, and specifications
- Potential disagreement between generated text and a reviewed record
Measure precision and recall on your own labeled evaluation set before using this in production. Ambiguous, high-stakes, or out-of-catalog assertions should route to primary-source or human review.
Performance
- One network request per verify call; benchmark from your own region
- Public API: free, no signup, no auth, and no account-level meter
- Cached responses (claim → envelope) for repeated assertions
- Batching or parallelization is your implementation choice; respect network abuse controls
Integration guides per framework
- LangChain — retrieve-then-cite + generate-then-verify patterns
- LlamaIndex — custom Retriever + NodePostprocessor
- DSPy — verify-and-flag post-processor module
- OpenAI tools — native function-calling pattern
When this fits
- RAG over AI/ML knowledge bases (papers, model docs, technical content)
- Documentation chatbots
- Research-assistant pipelines
- Any production RAG with hallucination tickets where the source data is correct