Pydantic AI + SourceScore VERITAS
Type-safe candidate retrieval as a Pydantic AI tool. The model calls find_claim_candidate() with a structured input, gets back a typed retrieval envelope, and the agent loop continues with validated data — not free-text.
Why Pydantic AI fits this well
Pydantic AI's design principle is: tools are typed functions with Pydantic models for inputs and outputs. The model gets a JSON schema; the runtime validates every tool call against the schema before the function runs.
That maps cleanly onto VERITAS's envelope format. We define a CandidateLookupResultPydantic model, the agent emits structured tool calls, and the downstream consumer (your application) gets a typed object — not a free-text claim with maybe-a-link.
Installation
pip install pydantic-ai httpx
Pattern: candidate lookup with the real response shape
Define the Pydantic models for tool input + output, register the tool with the agent, and let the LLM call it when it needs a possible catalog record. The type layer does not prove a fact.
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
import httpx
class ClaimLookupInput(BaseModel):
"""Input for candidate retrieval."""
claim: str = Field(description="Natural-language assertion to look up")
min_confidence: float = Field(default=0.85, ge=0.0, le=1.0)
class CandidateRecord(BaseModel):
id: str
subject: str
predicate: str
object: str
statement: str
confidence: float
detailUrl: str
class CandidateMatch(BaseModel):
claim: CandidateRecord
matchScore: float
rationale: str
class CandidateLookupResult(BaseModel):
"""Typed /verify response. The endpoint name is legacy."""
query: str
method: str
note: str
matches: list[CandidateMatch]
bestMatch: CandidateRecord | None = None
signature: dict | None = None
agent = Agent(
"openai:gpt-4o",
system_prompt=(
"You are a research assistant. When a user makes a factual "
"assertion about AI/ML, call find_claim_candidate() before responding. "
"A bestMatch is retrieval similarity, not a truth verdict. Return its "
"detailUrl for application-side evidence review; do not present it as proof."
),
)
@agent.tool
async def find_claim_candidate(ctx: RunContext, input: ClaimLookupInput) -> CandidateLookupResult:
"""Retrieve a candidate record for a natural-language claim."""
async with httpx.AsyncClient() as client:
r = await client.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": input.claim, "minConfidence": input.min_confidence},
timeout=5.0,
)
data = r.json()
return CandidateLookupResult.model_validate(data)
# Use it:
result = await agent.run("When was Llama 3.1 released?")
print(result.output)
# Agent can call find_claim_candidate with a typed input. A returned candidate
# is not enough to answer: fetch candidate.detailUrl and compare its statement
# and cited evidence with the intended assertion before rendering a citation.Pattern: keep retrieval and reviewed support separate
Ask the agent to emit typed candidate metadata via Pydantic AI'soutput_type parameter. The model can't return a free-text answer; it must populate a structured object. Your application adds a separate reviewed-support decision after inspecting evidence.
class AnsweredQuestion(BaseModel):
"""Retrieval output shape; not an accuracy certificate."""
question: str
answer: str
retrieval_status: str # "candidate_found" | "no_candidate"
candidate_urls: list[str] # fetch and review before asserting truth
record_confidence: float
agent_with_typed_output = Agent(
"openai:gpt-4o",
output_type=AnsweredQuestion,
system_prompt=(
"For AI/ML assertions, call find_claim_candidate() for "
"each factual assertion. Populate the AnsweredQuestion fields "
"with candidate retrieval data; never invent sources or treat a match as proof."
),
)
# Register the same public tool function on this agent; do not copy private internals.
agent_with_typed_output.tool(find_claim_candidate)
result = await agent_with_typed_output.run(
"When was GPT-4 released and how many parameters does it have?"
)
# result.output is now strictly typed:
print(result.output.question) # str
print(result.output.answer) # str
print(result.output.retrieval_status) # "candidate_found" | "no_candidate"
print(result.output.candidate_urls) # list[str], not yet evidence-approved
print(result.output.record_confidence) # legacy record metadataPattern: multi-claim candidate retrieval
For research-assistant agents that return multiple claims, retrieve candidates in parallel, then review each before composing.
from typing import List
import asyncio
class ClaimWithCandidate(BaseModel):
claim_text: str
candidate_found: bool
record_confidence: float
candidate_url: str | None
class MultiClaimResponse(BaseModel):
summary: str
claims: List[ClaimWithCandidate]
candidate_rate: float # % with a similar catalog record; not a truth rate
@agent.tool
async def retrieve_many(ctx: RunContext, claims: list[str]) -> List[ClaimWithCandidate]:
"""Retrieve candidate records in parallel; review their evidence separately."""
async with httpx.AsyncClient() as client:
responses = await asyncio.gather(*[
client.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": c, "minConfidence": 0.85},
timeout=5.0,
)
for c in claims
])
out = []
for claim, response in zip(claims, responses):
data = response.json()
match = data.get("bestMatch")
out.append(ClaimWithCandidate(
claim_text=claim,
candidate_found=match is not None and match["confidence"] >= 0.85,
record_confidence=match["confidence"] if match else 0.0,
candidate_url=match["detailUrl"] if match else None,
))
return out
# In the agent's response composition:
result = await agent.run(
"What can you tell me about Llama 3.1, GPT-4, and Claude 3?"
)
# Agent calls retrieve_many(["Llama 3.1 release", "GPT-4 release", "Claude 3 release"])
# Compares the returned primary sources before composing factual prose; labels
# missing candidates explicitlyWhy this pattern beats free-text
- Structured source handling. The example passes returned source URLs through typed fields, which can make application-side validation easier; the model can still produce unsupported text, so validate final output.
- Explicit confidence handling. Pydantic validates that a returned confidence value fits the expected range; it does not establish that a claim is correct.
- Downstream code is type-safe. Your application that consumes the agent output gets a Pydantic object, not a JSON-shaped string that might be missing fields.
- Validators catch errors early. Add Pydantic
@field_validatordecorators to reject results that don't pass your business rules.
What VERITAS is not (for Pydantic AI agents)
VERITAS today covers AI/ML research — model releases, foundational papers, organizations, datasets, benchmarks. If your agent asks about "the capital of France" the lookup tool will return no bestMatch and your agent should fall through to a different retrieval path.
Catalog: 384 reviewed claim records today. Future coverage is not promised; check the catalog before relying on a topic.
Next steps
- • Browser playground — try /verify before wiring it up
- • OpenAPI 3.1 spec — generate Pydantic models from the spec via
datamodel-code-generator - • DSPy guide — compound-AI-system framework
- • OpenAI tool-calls — the underlying primitive
- • Browse the catalog — 384 reviewed AI/ML claim records
Bug in this guide? Tell us. Pydantic AI's API surface evolves fast; we update this guide on every minor release.