SourceScore

Integration guide

Haystack + SourceScore VERITAS

Two Haystack 2.x components: a retriever for curated VERITAS records and an annotator for possible matches. A separate evidence or entailment review decides whether an assertion is supported.

Install

pip install haystack-ai requests

A VERITAS retriever component

Haystack 2.x components are plain classes decorated with @component. The run() method declares its outputs via @component.output_types. This one turns VERITAS /search hits into Haystack Documents, carrying the claim id, confidence, and canonical URL in meta.

import requests
from typing import List
from haystack import component, Document

VERITAS = "https://sourcescore.org/api/v1"

@component
class VeritasRetriever:
    def __init__(self, top_k: int = 5):
        self.top_k = top_k

    @component.output_types(documents=List[Document])
    def run(self, query: str):
        r = requests.get(
            f"{VERITAS}/search",
            params={"q": query, "limit": self.top_k},
            timeout=8,
        )
        r.raise_for_status()
        docs = []
        for c in r.json().get("results", []):
            docs.append(
                Document(
                    content=c["statement"],
                    score=c.get("matchScore", c["confidence"]),
                    meta={
                        "claim_id": c["id"],
                        "confidence": c["confidence"],
                        "url": f"https://sourcescore.org/claims/{c['id']}/",
                        "tags": c.get("tags", []),
                    },
                )
            )
        return {"documents": docs}

A candidate annotator (for any existing retriever)

If your pipeline already has a primary retriever (a vector store, say), add this annotator after it. It POSTs each document to /verify and attaches any returned candidate. It never keeps or drops a document based on similarity alone.

import requests
from typing import List
from haystack import component, Document

@component
class VeritasCandidateAnnotator:
    def __init__(self, min_confidence: float = 0.85):
        self.min_confidence = min_confidence

    @component.output_types(documents=List[Document])
    def run(self, documents: List[Document]):
        for d in documents:
            r = requests.post(
                f"{VERITAS}/verify",
                json={"claim": d.content, "minConfidence": self.min_confidence},
                timeout=8,
            ).json()
            best = r.get("bestMatch")
            if best:
                d.meta["veritas_candidate_id"] = best["id"]
                d.meta["veritas_record_confidence"] = best["confidence"]
                d.meta["veritas_candidate_url"] = best.get("detailUrl")
                d.meta["requires_evidence_review"] = True
        return {"documents": documents}

Wire the pipeline

Connect the retriever to a PromptBuilder and an OpenAIGenerator. The prompt instructs the model to use a candidate only when its exact statement supports the answer.

from haystack import Pipeline
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator

template = """The records below are retrieval candidates, not truth verdicts.
Use a record only when its exact statement supports the answer; otherwise say
the supplied evidence does not cover the question.

{% for doc in documents %}
[{{ doc.meta.claim_id }}] {{ doc.content }} (confidence {{ doc.meta.confidence }})
{% endfor %}

Question: {{ query }}
Answer (cite [claim_id] only after exact-statement comparison):"""

pipe = Pipeline()
pipe.add_component("retriever", VeritasRetriever(top_k=5))
pipe.add_component("prompt", PromptBuilder(template=template, required_variables=["query"]))
pipe.add_component("llm", OpenAIGenerator(model="gpt-4o-mini"))

pipe.connect("retriever.documents", "prompt.documents")
pipe.connect("prompt.prompt", "llm.prompt")

question = "Who introduced the Transformer architecture?"
result = pipe.run({
    "retriever": {"query": question},
    "prompt": {"query": question},
})
print(result["llm"]["replies"][0])

Why an evidence-review step

A retriever returns the closest documents; it does not confirm a generated answer is consistent with them. This component only selects similar catalog candidates; it does not close that gap. Compare each linked primary source before asserting a fact. The public API is free with no account, key, or signup; it is the same API the rest of these guides use.

Next steps