Concept · 2026-05-16
RAG vs curated claim retrieval (VERITAS) — when to use each
RAG retrieves prose chunks. VERITAS retrieves typed records from a bounded curated catalog. They are different shapes with complementary uses; neither retrieval method proves an answer true.
TL;DR
RAG is the right tool when your knowledge lives in prose — documentation, articles, customer-support tickets, internal wikis. You vector-index it, retrieve semantically similar chunks at query time, paste them into the prompt as context. Works at any scale; tolerates messy unstructured input.
VERITAS is useful when you need atomic claim records — specific assertions like "GPT-4 was released on 2023-03-14" with sources you can inspect. Coverage is bounded, and candidate retrieval still needs statement/evidence comparison.
They're complementary. RAG covers breadth; VERITAS covers atoms. Most production LLM applications eventually run both.
The shape difference
The two systems retrieve different things:
| Aspect | RAG | VERITAS |
|---|---|---|
| Retrieves | Prose chunks (200-2000 tokens) | Atomic claims (subject + predicate + object) |
| Returns | Retrieved prose chunks | Candidate claim records with cited evidence |
| Trust model | Inspect the indexed corpus and answer support | Inspect the curation method, record, and cited evidence |
| Scale | Depends on corpus and infrastructure | Limited by curation effort (384 records today) |
| Accuracy | Measure on your task and corpus | Measure candidate entailment on your assertions |
| Coverage | Whatever you index | AI/ML catalog only today |
| Citation precision | Chunk-level (paragraph at best) | Fact-level (single assertion) |
| Auditability | Depends on stored chunks and citations | Stable record URLs + cited evidence; HMAC is source-issued only |
| Infra requirement | Vector DB + embedding model | HTTP fetch (zero infra) |
Why RAG alone leaks
RAG works well most of the time and fails in a specific class of cases:
- Semantic-similar but factually-wrong chunks. A chunk about "OpenAI launched ChatGPT in 2022" retrieves on a query about GPT-4. Embeddings see the same topic; the dates are different. The model stitches the retrieved date into the wrong context.
- Multiple chunks contradict. Two retrieved chunks disagree on a fact. The model picks one — sometimes the wrong one — without surfacing the contradiction.
- Chunk boundaries split facts. The relevant fact is at the boundary between two retrieved chunks. The model gets half of it; fabricates the other half.
- Indexed corpus is wrong. RAG retrieves faithfully from a corpus. If the corpus has wrong facts, RAG confidently surfaces them as authoritative.
- Citation drift. The model cites "chunk #4 says X" but actually emitted X from its parametric memory and pasted the chunk citation as plausible cover.
Signed-claim verification doesn't have these failure modes because the unit is the typed fact, not a prose chunk. Either the catalog has the claim (return it, confidence-stamped) or it doesn't (return null, force fallback path).
Why VERITAS alone is insufficient
The signed-claim approach has its own limits:
- Coverage gaps. If the user asks about something the catalog doesn't cover — say, a specific product feature or a niche academic result — VERITAS returns null and your code falls through to whatever else you have. Without RAG as the fallback, your application is silent.
- Curation latency. New facts (a model released yesterday) take time to verify and enter the catalog. RAG over a freshly-indexed news corpus is faster.
- Prose vs claim shape. Some queries genuinely want explanatory prose, not a typed fact. "How does attention work?" isn't an atomic claim — it's a paragraph. RAG over the right corpus answers; VERITAS doesn't.
- Subjective questions. "Is GPT-4 better than Claude 3 for code generation?" isn't a fact — it's an evaluation. VERITAS explicitly doesn't ship performance-comparison claims (see the methodology post). RAG over benchmark reports + community discussion fills this gap.
The hybrid pattern
Three layers, ordered by precision:
# Pseudocode
async def answer(question):
# Layer 1 — VERITAS for atomic facts (high precision)
veritas_claims = await veritas_search(question, limit=3)
# Layer 2 — RAG for explanatory context
rag_chunks = await vector_search(question, limit=5)
# Layer 3 — model prompt with both, ordered by trust
context = ""
if veritas_claims:
context += "Candidate atomic records (cite only after evidence review):\n"
context += "\n".join(f"- {c.statement} [{c.id}]" for c in veritas_claims)
if rag_chunks:
context += "\n\nDocumentation context (cite [chunk_id]):\n"
context += "\n".join(f"- {c.text} [{c.id}]" for c in rag_chunks)
return await llm.generate(
f"Use a record only when its exact statement and evidence support the answer. "
f"Cite every supported assertion.\n\n{context}\n\nQ: {question}"
)The model is instructed to use a candidate only when its exact statement supports the answer. Candidate-record links in the UI distinguish the two evidence paths — clicking [claim_id] opens the canonical SourceScore page; clicking [chunk_id] opens your indexed source.
Cost comparison
The public VERITAS API is currently free with no account-level meter. RAG cost depends on your embedding model, vector store, retrieval pattern, cache hit rate, and hosting. Measure both paths in your own stack instead of relying on a universal per-query estimate. Proposed higher-volume VERITAS prices are a demand test, not purchasable plans.
When to use RAG only (skip VERITAS)
- Your domain isn't AI/ML and isn't covered by any other signed-claim catalog.
- Your queries are explanatory, not factual (how does X work, why is Y, summarize Z).
- You have a high-quality indexed corpus you trust completely.
- Latency is dominant; you can't afford the verification step.
When to use VERITAS only (skip RAG)
- Your application is bounded to AI/ML facts.
- You don't have time to build a RAG pipeline.
- You need zero-infra grounding — VERITAS is one HTTP call, no vector DB.
- You need programmatic auditability of cited facts (signatures).
Is RAG dead?
No. The "RAG is dead" takes you see on Twitter are marketing for the framing-of-the-month, not a methodological shift. RAG remains the default pattern for indexing your own corpus of prose, and that need isn't going away.
What's changed is the recognition that RAG isn't a complete grounding solution — it's one layer in a stack. Signed-claim verification, tool-use grounding, prompt-stuffed invariants, and structured-output schemas all sit alongside RAG in a production-grade pipeline.
Getting started
If you have RAG today: add VERITAS as a parallel retrieval step. Five lines of code change, no infrastructure change. See the 5-line Python tutorial.
If you don't have RAG yet, test the VERITAS catalog against a representative query set before designing around it. Add broader retrieval when your observed questions fall outside the catalog; do not assume a universal coverage percentage.
Further reading
- LLM grounding — the broader concept
- LLM hallucination — what grounding fixes
- 5-line Python verification tutorial
- LangChain integration — including a retrieve-then-cite pattern that mirrors classic RAG flow
- Quickstart