Anthropic SDK + SourceScore VERITAS
Expose VERITAS as a Claude tool via the Anthropic SDK. When Claude needs catalog evidence, it emits a tool_use block calling find_claim_candidate; you execute the API call and feed the result back as a tool_result; Claude composes a final answer after comparing the returned statement and evidence.
Installation
# Python pip install anthropic httpx # TypeScript / Node npm install @anthropic-ai/sdk
Pattern: tool definition + agent loop (Python)
Claude's tool-use protocol is a multi-turn loop: model emits tool_use, you execute, you reply with tool_result, model composes final response. SourceScore's response envelope drops in as a tool_result content block verbatim.
import anthropic
import httpx
import json
import os
client = anthropic.Anthropic() # picks up ANTHROPIC_API_KEY
model = os.environ["ANTHROPIC_MODEL"] # pin and test the model used in production
tools = [
{
"name": "find_claim_candidate",
"description": (
"Find a similar SourceScore catalog record for an AI/ML assertion. "
"A result is a candidate for evidence review, not a truth verdict."
),
"input_schema": {
"type": "object",
"properties": {
"claim": {
"type": "string",
"description": "Natural-language assertion to look up",
},
"min_confidence": {
"type": "number",
"description": "Minimum confidence threshold (0.0-1.0)",
"default": 0.85,
},
},
"required": ["claim"],
},
},
]
async def execute_find_claim_candidate(claim: str, min_confidence: float = 0.85) -> dict:
"""Retrieve catalog candidates; compare cited evidence before use."""
async with httpx.AsyncClient() as http:
r = await http.post(
"https://sourcescore.org/api/v1/verify",
json={"claim": claim, "minConfidence": min_confidence},
timeout=5.0,
)
return r.json()
async def chat(user_message: str) -> str:
messages = [{"role": "user", "content": user_message}]
while True:
response = client.messages.create(
model=model,
max_tokens=1024,
tools=tools,
messages=messages,
)
# If Claude wants to use a tool, execute it
if response.stop_reason == "tool_use":
tool_use_block = next(
b for b in response.content if b.type == "tool_use"
)
if tool_use_block.name == "find_claim_candidate":
result = await execute_find_claim_candidate(
claim=tool_use_block.input["claim"],
min_confidence=tool_use_block.input.get("min_confidence", 0.85),
)
# Continue the loop with the tool result
messages.append({"role": "assistant", "content": response.content})
messages.append({
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": tool_use_block.id,
"content": json.dumps(result),
}
],
})
continue
# No more tool use; return Claude's final response
return "".join(b.text for b in response.content if b.type == "text")
# Use it, then inspect the returned candidate evidence:
import asyncio
answer = asyncio.run(chat("When was Llama 3.1 released?"))
print(answer)
Pattern: TypeScript with the @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const model = process.env.ANTHROPIC_MODEL;
if (!model) throw new Error("Set ANTHROPIC_MODEL to a pinned, tested model ID.");
const tools: Anthropic.Tool[] = [
{
name: "find_claim_candidate",
description: (
"Find a similar SourceScore catalog record for an AI/ML assertion. " +
"A result is a candidate for evidence review, not a truth verdict."
),
input_schema: {
type: "object",
properties: {
claim: { type: "string" },
min_confidence: { type: "number", default: 0.85 },
},
required: ["claim"],
},
},
];
async function executeFindClaimCandidate(claim: string, minConfidence = 0.85) {
const r = await fetch("https://sourcescore.org/api/v1/verify", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ claim, minConfidence }),
});
return r.json();
}
async function chat(userMessage: string): Promise<string> {
const messages: Anthropic.MessageParam[] = [
{ role: "user", content: userMessage },
];
while (true) {
const response = await client.messages.create({
model,
max_tokens: 1024,
tools,
messages,
});
if (response.stop_reason === "tool_use") {
const toolUseBlock = response.content.find(
(b): b is Anthropic.ToolUseBlock => b.type === "tool_use",
);
if (!toolUseBlock) break;
if (toolUseBlock.name === "find_claim_candidate") {
const result = await executeFindClaimCandidate(
(toolUseBlock.input as { claim: string }).claim,
(toolUseBlock.input as { min_confidence?: number }).min_confidence ?? 0.85,
);
messages.push({ role: "assistant", content: response.content });
messages.push({
role: "user",
content: [
{
type: "tool_result",
tool_use_id: toolUseBlock.id,
content: JSON.stringify(result),
},
],
});
continue;
}
}
return response.content
.filter((b): b is Anthropic.TextBlock => b.type === "text")
.map((b) => b.text)
.join("");
}
return "";
}
const answer = await chat("When was Llama 3.1 released?");
console.log(answer);Require evidence comparison in the agent prompt
One useful pattern: a system prompt that instructs Claude to retrieve a candidate for AI/ML assertions, then compare the candidate statement and cited evidence before including a citation. The tool result alone must never confirm the assertion.
system_prompt = """You are a research assistant for AI/ML topics. CRITICAL: Before making a factual claim about an AI model, paper, release date, parameter count, or architecture decision, call find_claim_candidate to retrieve possible catalog evidence. bestMatch and matchScore describe retrieval, not truth. Compare the exact candidate statement and cited primary evidence with your assertion. Cite the detailUrl only when that evidence supports the assertion. If no candidate is returned, say the bounded catalog supplied no evidence; do not call it false. NEVER assert a release date or parameter count without first calling find_claim_candidate and completing the evidence comparison."""
For higher-stakes uses, enforce the comparison outside the model as well; a prompt is not a security or accuracy boundary.
When this pattern fits
- Conversational AI/ML research assistants — where users ask factual questions and you want the model to ground itself
- Documentation chatbots over AI/ML knowledge — internal team support tools, public-facing FAQs
- Citation-required production systems — papers, technical reports, audit trails
- Multi-step agentic flows where one step is "look up a fact"
Next steps
- • OpenAI tool-calls guide — the parallel pattern for GPT-4
- • Pydantic AI guide — typed-tool pattern with validators
- • Playground — try /verify before wiring it up
- • OpenAPI 3.1 spec — full endpoint reference
- • Catalog — 384 reviewed AI/ML claim records