Integration guide
LangChain + SourceScore VERITAS
Wire curated record retrieval into your LangChain pipeline. Each record links to cited evidence and a stable canonical page; your application still decides whether that evidence supports its answer.
When to use this
Two patterns. Both are drop-in additions to an existing chain.
- Retrieve-then-review: fetch candidate VERITAS records, compare their exact statements and cited evidence with the question, and cite only records that actually support the answer.
- Generate-then-find-candidates: extract atomic assertions, send each to the legacy
/api/v1/verifyroute, then send returned candidates and their evidence to review.
Install
# Python
pip install langchain langchain-openai requests
# JavaScript
npm install @langchain/core @langchain/openaiPattern 1 — Retrieve-then-cite (Python)
The VERITAS catalog acts as a curated retriever. Search returns the top-N matching claims; we render them as context blocks the LLM is instructed to cite from.
import os, requests
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
VERITAS = "https://sourcescore.org/api/v1"
def veritas_retrieve(query: str, k: int = 5) -> str:
"""Fetch top-k VERITAS claims for a query, render as numbered context."""
r = requests.get(f"{VERITAS}/search", params={"q": query, "limit": k}, timeout=8)
r.raise_for_status()
claims = r.json().get("results", [])
if not claims:
return "(no VERITAS claims match this query)"
lines = []
for i, c in enumerate(claims, 1):
lines.append(
f"[{i}] {c['statement']} "
f"(claim_id={c['id']}, confidence={c['confidence']:.2f})"
)
return "\n".join(lines)
prompt = ChatPromptTemplate.from_template("""You are a precise assistant. The records below are
candidates, not truth verdicts. Use one only when its exact statement supports
the answer; cite it with [claim_id]. If the records do not cover the question,
say so explicitly — do not improvise.
Candidate records:
{context}
Question: {question}
Answer (every fact must end with [claim_id]):""")
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
chain = (
{"context": lambda x: veritas_retrieve(x["question"]),
"question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
print(chain.invoke({"question": "When was the Transformer architecture introduced?"}))
Treat a missing citation as a review signal, not proof of a hallucination. Validate the final answer and cited evidence outside the model before publishing it.
Pattern 2 — Generate-then-find-candidates (JavaScript)
Useful when you want free-form output followed by a bounded-catalog check. Each assertion gets either a candidate link for evidence review or a clear no-candidate state.
import { ChatOpenAI } from "@langchain/openai";
import { ChatPromptTemplate } from "@langchain/core/prompts";
const VERITAS = "https://sourcescore.org/api/v1";
async function findClaimCandidate(text) {
const r = await fetch(`${VERITAS}/verify`, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ claim: text, minConfidence: 0.85 }),
});
return r.json();
}
// Step 1 — model generates answer
const llm = new ChatOpenAI({ model: "gpt-4o-mini", temperature: 0 });
const prompt = ChatPromptTemplate.fromTemplate(`
Answer the user's question with one fact per line. Be concise.
Question: {question}
`);
const answer = await prompt.pipe(llm).invoke({
question: "When did OpenAI release GPT-4?",
});
// Step 2 — retrieve a candidate for each line (then review its sources)
const lines = answer.content.split("\n").filter(Boolean);
const candidates = [];
for (const line of lines) {
const v = await findClaimCandidate(line);
candidates.push({
statement: line,
matched: !!v.bestMatch,
confidence: v.bestMatch?.confidence ?? 0,
veritasId: v.bestMatch?.id,
sourceUrl: v.bestMatch ? `https://sourcescore.org/claims/${v.bestMatch.id}/` : null,
});
}
// Step 3 — render candidate links; a match is not a truth verdict
for (const r of candidates) {
const badge = r.matched ? `🔎 candidate [${r.veritasId}] — review sources` : "⚠️ no catalog candidate";
console.log(`${r.statement.trim()} ${badge}`);
}
UI suggestion: render a candidate link only as a prompt to inspect its primary sources, never as a factual-verification badge. Mark absent matches clearly and send disputed results to human review.
Pattern 3 — Refetch the canonical record (defensive)
High-stakes deployments should refetch the canonical HTTPS record and compare the claim content and cited evidence before using it. The HMAC tag is not publicly independently verifiable.
import requests
VERITAS = "https://sourcescore.org/api/v1"
env = requests.get(f"{VERITAS}/claims/<claim_id>.json").json()
canonical = requests.get(f"{VERITAS}/claims/{env['claim']['id']}.json").json()
assert canonical["claim"] == env["claim"], "Canonical record changed — inspect cited evidence"
print("canonical record matches; inspect cited evidence for your use case")
No public or enterprise shared secret is available. Refetching checks the current canonical record, not a cryptographic proof of origin.
Choosing a pattern
| Use case | Pattern | Latency |
|---|---|---|
| Educational Q&A bot | Retrieve-then-cite | One catalog request + LLM time; benchmark locally |
| Search auto-complete | Retrieve only (skip LLM) | One catalog request; benchmark locally |
| Internal research assistant | Generate-then-find-candidates | LLM + N catalog requests |
| High-stakes citation badge | Candidate retrieval + independent evidence review | LLM + N catalog requests + review |
Cost model
The public VERITAS API is free, requires no account or key, and has no account-level meter. Search and verify are separate network requests; cache stable claim records and measure traffic in your own stack.
Higher-volume prices on the pricing page are proposals used to test demand, not plans available for purchase.
What VERITAS is not
We are deliberately not a generic fact-checker. The Day 1 catalog (384 claims today) covers AI/ML research — model releases, foundational papers, organizations, datasets. If your chain asks about "the capital of France" we will return no matches and your code should fall through to whatever retrieval you'd use anyway.
Catalog expansion is gated by the published methodology (cited evidence, source counts, and no unstable performance-comparison claims). No date is promised for additional verticals.
Next steps
- • Full API reference — every endpoint with curl + JS + Python examples
- • Browse the catalog — 384 reviewed AI/ML claim records
- • OpenAPI spec — generate clients in any language
- • Pricing — free API and proposed higher-volume tiers