SourceScore

Tutorial · 2026-05-16

Match AI-generated facts to a cited catalog in 5 lines of Python

Use SourceScore VERITAS as a post-generation screening step. It returns nearby catalog records and canonical citations to compare—not a truth verdict.

The problem

You wired up GPT-4 or Claude to answer questions about AI/ML research. Demo is great. Then a user asks "when was the Transformer architecture introduced and by whom?" and the model invents a plausible-but-wrong attribution. You catch it this time. You won't catch it the next thousand times.

A common mitigation is RAG: retrieve relevant context and provide it to the model. That improves grounding, but it does not guarantee that every generated assertion is supported by the retrieved text.

The 5-line catalog check

A different approach: let the model answer freely, then match each assertion against a catalog of sourced claims. A result is a candidate record to compare with the model output—not proof that the model's wording is true.

import requests

def find_catalog_match(claim: str, threshold: float = 0.85):
    r = requests.post("https://sourcescore.org/api/v1/verify",
        json={"claim": claim, "minConfidence": threshold}, timeout=8)
    r.raise_for_status()
    return r.json().get("bestMatch")  # candidate record, not a truth verdict

That's the whole client. Five lines including the import. Use it to find reviewable AI/ML catalog records before publishing. Your application still needs to compare the returned statement and cited evidence with the assertion it intends to show.

Wire it into a chain

Here's the same function inside a generate-then-verify loop:

from openai import OpenAI

client = OpenAI()

def answer_with_citations(question: str) -> str:
    # Step 1 — model generates one fact per line
    raw = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": f"Answer with one fact per line:\n{question}"}],
        temperature=0,
    ).choices[0].message.content

    # Step 2 — find a candidate record for each line
    out = []
    for line in raw.strip().split("\n"):
        if not line.strip(): continue
        best = find_catalog_match(line)
        if best:
            url   = f"https://sourcescore.org/claims/{best['id']}/"
            out.append(f"Input: {line.strip()}\nCandidate: {best['statement']}\nReview: {url}")
        else:
            out.append(f"Input: {line.strip()}\nNo catalog candidate found")
    return "\n".join(out)

print(answer_with_citations("When was the Transformer architecture introduced and by whom?"))

Sample output:

Input: The Transformer architecture was introduced in 2017.
Candidate: Transformer architecture introduced in paper: Attention Is All You Need (Vaswani et al., 2017).
Review: https://sourcescore.org/claims/ad17e76a8baad7a1/

What you get

  • Screening aid. Assertions with no nearby catalog record can be routed to another retrieval or human-review path.
  • Reviewable citations. Candidate records ship with a canonical URL where users can inspect cited sources, integrity metadata, and the last-reviewed date.
  • Visible request cost. Each assertion you submit is one additional network request. The public v0 endpoints need no key; measure latency and traffic in your own stack.

Scope honesty

VERITAS today is bounded to AI/ML research — 384 hand-verified claims across foundational papers, model releases, organizations, and datasets. If your chain asks about "the capital of France" we return no match and your code falls through to whatever retrieval you'd use anyway.

Catalog expansion is gated by our methodology: every claim must cite primary evidence, show source counts, and not be a performance comparison (benchmark numbers vary by prompt format, version, shot count, and evaluation setup). No date is promised for new verticals.

Going deeper

One question I get a lot

"Why not just put all 384 claims in the prompt as context?"

You can, and for a Day 1 demo you should. The reason to pull via API instead is:

  1. An API lets you avoid inserting the entire catalog into every prompt.
  2. Retrieval ranks claims by relevance to the actual question — you can send only candidate records relevant to the question.
  3. Refetching the canonical API record lets you compare your copy with SourceScore's current copy. The HMAC tag is not independently verifiable by public users because the shared secret is unpublished.

Start with the simplest pattern that fits. Move to API retrieval when measured context size, latency, or maintenance cost justifies it; the small client above shows the request shape.