Instructor + SourceScore VERITAS
Instructor (Jason Liu's library) is the canonical pattern for getting typed structured outputs from LLMs. Pair it with VERITAS for structured responses where cited assertions can carry a candidate record. Pydantic validates that the retrieval shape is usable; your application must still compare its primary evidence before making a factual assertion.
Installation
pip install instructor httpx
Pattern: typed claim with candidate-source field
Define a Pydantic model where the LLM populates structured fields including a source_url field populated from a VERITAS lookup. The validator runs at response-parsing time; a match is retrieval similarity, not a truth verdict, so compare the linked primary evidence before use.
from pydantic import BaseModel, field_validator, model_validator
from openai import OpenAI
import instructor
import httpx
client = instructor.from_openai(OpenAI())
class ClaimAnswer(BaseModel):
"""LLM response with a candidate catalog record."""
claim: str
answer: str
source_url: str | None = None
confidence: float = 0.0
@model_validator(mode="after")
def retrieve_candidate(self) -> "ClaimAnswer":
"""Retrieve a candidate record; compare its evidence before asserting truth."""
r = httpx.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": self.claim, "minConfidence": 0.85},
timeout=5.0,
)
result = r.json()
match = result.get("bestMatch")
if match and match["confidence"] >= 0.85:
self.source_url = match["detailUrl"]
self.confidence = match["confidence"]
else:
# Trigger Instructor retry with a different LLM phrasing
raise ValueError(
f"No sufficiently similar catalog candidate for '{self.claim}'. "
"Please rephrase using a more specific fact."
)
return self
# Use it:
result = client.chat.completions.create(
model="gpt-4o",
response_model=ClaimAnswer,
messages=[
{"role": "user", "content": "When was Llama 3.1 released?"},
],
max_retries=3, # Instructor retries on validation failure
)
print(result.claim) # "Llama 3.1 release date"
print(result.answer) # "2024-07-23"
print(result.source_url) # "https://sourcescore.org/api/v1/claims/.../"
print(result.confidence) # 1.0Pattern: list of candidate records
For research-assistant agents that produce multiple claims, extract a list of candidate-record objects, then independently review the linked evidence for each assertion:
from typing import List
from pydantic import BaseModel, field_validator
class CandidateRecord(BaseModel):
statement: str
source_url: str
confidence: float
@field_validator("source_url", mode="before")
@classmethod
def retrieve_candidate(cls, v, info):
statement = info.data.get("statement", "")
r = httpx.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": statement, "minConfidence": 0.85},
timeout=5.0,
)
result = r.json()
match = result.get("bestMatch")
if not match or match["confidence"] < 0.85:
raise ValueError(f"No sufficiently similar catalog candidate: {statement!r}")
return match["detailUrl"]
class ResearchSummary(BaseModel):
topic: str
summary: str
key_claims: List[CandidateRecord]
result = client.chat.completions.create(
model="gpt-4o",
response_model=ResearchSummary,
messages=[
{"role": "user", "content": "Summarize the foundational papers behind modern LLMs."},
],
max_retries=3,
)
# result is a typed ResearchSummary; typing does not prove its assertions
# every key_claims entry has a retrieved catalog candidate; compare primary
# sources yourself before presenting the statement as fact
for c in result.key_claims:
print(f"{c.statement} — {c.source_url} (conf: {c.confidence})")Why this pattern beats free-text + post-hoc retrieval
- Validation happens at parse-time. A missing candidate triggers Instructor's retry mechanism before the user sees a response.
- Type-safety at the application boundary. Downstream code receives a typed Pydantic object and can require an evidence-review step before rendering a factual assertion.
- No regex extraction. Free-text + post-hoc retrieval needs heuristic claim extraction (which fails on multi-clause sentences). Instructor extracts claims at structured-output time.
- Retries are automatic. max_retries=3 means three attempts at a catalog match before failing. Tunable per-call.
When this pattern fits
- Production AI/ML research assistants with a human or programmatic primary-evidence review step
- Documentation chatbots that summarize technical content
- Internal company knowledge tools that need typed candidate records and citations
- Citation-heavy reports or briefs where each assertion is reviewed against its sources
Comparison: Instructor vs Pydantic AI
Both libraries solve the "typed LLM outputs" problem. Differences:
- Instructor — older, larger ecosystem, supports more LLM providers, no built-in tool-calling abstractions. Pair-with-anything design.
- Pydantic AI — newer, agent-loop-aware, native tool registration, type-safety end-to-end. More opinionated; bigger framework.
Pick Instructor for one-shot structured-output use cases. Pick Pydantic AI for agent loops with multiple tool calls. Both work with VERITAS the same way.
Next steps
- • Pydantic AI guide — agent-loop variant
- • Research citation use case — Instructor-shape patterns
- • Playground — try /verify before wiring it up
- • OpenAPI 3.1 spec
- • Catalog — 384 reviewed AI/ML claim records