SourceScore

Anthropic SDK + SourceScore VERITAS

Expose VERITAS as a Claude tool via the Anthropic SDK. When Claude needs catalog evidence, it emits a tool_use block calling find_claim_candidate; you execute the API call and feed the result back as a tool_result; Claude composes a final answer after comparing the returned statement and evidence.

Installation

# Python
pip install anthropic httpx

# TypeScript / Node
npm install @anthropic-ai/sdk

Pattern: tool definition + agent loop (Python)

Claude's tool-use protocol is a multi-turn loop: model emits tool_use, you execute, you reply with tool_result, model composes final response. SourceScore's response envelope drops in as a tool_result content block verbatim.

import anthropic
import httpx
import json
import os

client = anthropic.Anthropic()  # picks up ANTHROPIC_API_KEY
model = os.environ["ANTHROPIC_MODEL"]  # pin and test the model used in production

tools = [
    {
        "name": "find_claim_candidate",
        "description": (
            "Find a similar SourceScore catalog record for an AI/ML assertion. "
            "A result is a candidate for evidence review, not a truth verdict."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "claim": {
                    "type": "string",
                    "description": "Natural-language assertion to look up",
                },
                "min_confidence": {
                    "type": "number",
                    "description": "Minimum confidence threshold (0.0-1.0)",
                    "default": 0.85,
                },
            },
            "required": ["claim"],
        },
    },
]

async def execute_find_claim_candidate(claim: str, min_confidence: float = 0.85) -> dict:
    """Retrieve catalog candidates; compare cited evidence before use."""
    async with httpx.AsyncClient() as http:
        r = await http.post(
            "https://sourcescore.org/api/v1/verify",
            json={"claim": claim, "minConfidence": min_confidence},
            timeout=5.0,
        )
        return r.json()

async def chat(user_message: str) -> str:
    messages = [{"role": "user", "content": user_message}]

    while True:
        response = client.messages.create(
            model=model,
            max_tokens=1024,
            tools=tools,
            messages=messages,
        )

        # If Claude wants to use a tool, execute it
        if response.stop_reason == "tool_use":
            tool_use_block = next(
                b for b in response.content if b.type == "tool_use"
            )

            if tool_use_block.name == "find_claim_candidate":
                result = await execute_find_claim_candidate(
                    claim=tool_use_block.input["claim"],
                    min_confidence=tool_use_block.input.get("min_confidence", 0.85),
                )

                # Continue the loop with the tool result
                messages.append({"role": "assistant", "content": response.content})
                messages.append({
                    "role": "user",
                    "content": [
                        {
                            "type": "tool_result",
                            "tool_use_id": tool_use_block.id,
                            "content": json.dumps(result),
                        }
                    ],
                })
                continue

        # No more tool use; return Claude's final response
        return "".join(b.text for b in response.content if b.type == "text")

# Use it, then inspect the returned candidate evidence:
import asyncio
answer = asyncio.run(chat("When was Llama 3.1 released?"))
print(answer)

Pattern: TypeScript with the @anthropic-ai/sdk

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();
const model = process.env.ANTHROPIC_MODEL;
if (!model) throw new Error("Set ANTHROPIC_MODEL to a pinned, tested model ID.");

const tools: Anthropic.Tool[] = [
  {
    name: "find_claim_candidate",
    description: (
      "Find a similar SourceScore catalog record for an AI/ML assertion. " +
      "A result is a candidate for evidence review, not a truth verdict."
    ),
    input_schema: {
      type: "object",
      properties: {
        claim: { type: "string" },
        min_confidence: { type: "number", default: 0.85 },
      },
      required: ["claim"],
    },
  },
];

async function executeFindClaimCandidate(claim: string, minConfidence = 0.85) {
  const r = await fetch("https://sourcescore.org/api/v1/verify", {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ claim, minConfidence }),
  });
  return r.json();
}

async function chat(userMessage: string): Promise<string> {
  const messages: Anthropic.MessageParam[] = [
    { role: "user", content: userMessage },
  ];

  while (true) {
    const response = await client.messages.create({
      model,
      max_tokens: 1024,
      tools,
      messages,
    });

    if (response.stop_reason === "tool_use") {
      const toolUseBlock = response.content.find(
        (b): b is Anthropic.ToolUseBlock => b.type === "tool_use",
      );
      if (!toolUseBlock) break;

      if (toolUseBlock.name === "find_claim_candidate") {
        const result = await executeFindClaimCandidate(
          (toolUseBlock.input as { claim: string }).claim,
          (toolUseBlock.input as { min_confidence?: number }).min_confidence ?? 0.85,
        );

        messages.push({ role: "assistant", content: response.content });
        messages.push({
          role: "user",
          content: [
            {
              type: "tool_result",
              tool_use_id: toolUseBlock.id,
              content: JSON.stringify(result),
            },
          ],
        });
        continue;
      }
    }

    return response.content
      .filter((b): b is Anthropic.TextBlock => b.type === "text")
      .map((b) => b.text)
      .join("");
  }
  return "";
}

const answer = await chat("When was Llama 3.1 released?");
console.log(answer);

Require evidence comparison in the agent prompt

One useful pattern: a system prompt that instructs Claude to retrieve a candidate for AI/ML assertions, then compare the candidate statement and cited evidence before including a citation. The tool result alone must never confirm the assertion.

system_prompt = """You are a research assistant for AI/ML topics.

CRITICAL: Before making a factual claim about an AI model, paper,
release date, parameter count, or architecture decision, call
find_claim_candidate to retrieve possible catalog evidence.

bestMatch and matchScore describe retrieval, not truth. Compare the exact
candidate statement and cited primary evidence with your assertion. Cite the
detailUrl only when that evidence supports the assertion. If no candidate is
returned, say the bounded catalog supplied no evidence; do not call it false.

NEVER assert a release date or parameter count without first calling
find_claim_candidate and completing the evidence comparison."""

For higher-stakes uses, enforce the comparison outside the model as well; a prompt is not a security or accuracy boundary.

When this pattern fits

  • Conversational AI/ML research assistants — where users ask factual questions and you want the model to ground itself
  • Documentation chatbots over AI/ML knowledge — internal team support tools, public-facing FAQs
  • Citation-required production systems — papers, technical reports, audit trails
  • Multi-step agentic flows where one step is "look up a fact"

Next steps