AI agent grounding — retrieve evidence for assertions in tool-using chains
Give agents a bounded AI/ML evidence lookup for release dates, parameter counts, and architectural facts. A candidate match starts review; it is not an automatic truth verdict.
The problem
An AI agent calls tools in a loop. Search, fetch URL, run code, query database — each tool gives the model fresh context. The model decides what to do next based on that context.
The failure mode: the model emits factual assertions that aren't in the tool results. The agent uses its parametric memory to fill in gaps, gets a release date wrong by 6 months, gets a parameter count wrong by 10× — and the rest of the agent loop builds on that false foundation. Multi-step agent failures compound.
The pattern: add a catalog lookup to the agent's tool catalog. Instruct the agent to call it before asserting an AI/ML factual claim. A returned record points to evidence the agent or a reviewer can compare with the assertion.
The pattern
# OpenAI tool catalog excerpt
tools = [
# ... your existing tools ...
{
"type": "function",
"function": {
"name": "find_claim_candidate",
"description": (
"Find a similar SourceScore catalog record for an AI/ML assertion. "
"A result is a candidate for evidence review, not a truth verdict."
),
"parameters": {
"type": "object",
"properties": {
"claim": {"type": "string"},
"min_confidence": {"type": "number", "default": 0.85},
},
"required": ["claim"],
},
},
},
]
def execute_find_claim_candidate(args):
import httpx
r = httpx.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": args["claim"], "minConfidence": args.get("min_confidence", 0.85)},
timeout=5.0,
)
return r.json()System prompt addition
Add to your agent's system prompt:
When you assert any factual claim about AI/ML topics (model releases, paper dates, parameter counts, organization founding), call find_claim_candidate first. Similarity confidence is not truth confidence. Compare the assertion with the returned statement and its cited evidence; cite the detail URL only when that evidence supports the assertion. A null result means only that the bounded catalog supplied no candidate. Never invent a citation.
What to measure in your own evaluation
- Unsupported-assertion rate. Compare a baseline agent with the evidence-review workflow on the same labeled prompts.
- Citation usefulness. Review whether a cited record actually entails the assertion and whether its sources are sufficient for your use case.
- Agent loop cost: at least one network request per lookup. Cache stable responses for repeated claims. The public API is free, needs no signup or key, and has no account-level meter.
Drop-in integration guides
- OpenAI tool-calls — full pattern with chat-completions loop
- Anthropic SDK — Claude tool_use protocol
- Pydantic AI — type-safe variant
- LangChain — @tool decorator
- DSPy — programs-not-prompts variant
When this fits
- Production agent that answers AI/ML factual questions
- Research-assistant agents over technical literature
- Documentation chatbots over AI/ML knowledge bases
- Multi-step agentic flows where one step is fact lookup
When this doesn't fit (yet)
- Agents that primarily answer non-AI/ML factual questions — our catalog is currently bounded to AI/ML.
- Agents that need real-time citations (current news) — VERITAS is a curated historical catalog, not a live-news source.
Try it
- Browser playground — no signup, paste a claim, see the response
- 5-min Quickstart — full pattern in curl + JS + Python
- OpenAPI 3.1 spec