SourceScore

Research citation — programmatic citations for AI/ML research tools

Stable claim IDs, cited evidence with short excerpts, and SourceScore-issued HMAC metadata. Candidate retrieval for academic AI assistants and literature-review tools, with evidence review before citation.

The problem

You're building an AI tool for researchers — literature review assistant, citation finder, paper-summary chatbot, academic search. Your users care about citations more than your typical LLM-app user. Wrong dates or fabricated authors aren't a UX bug — they're a credibility-destroying incident.

Standard RAG over arXiv produces fluent summaries but hallucinates dates, authors, and methodology details. Researchers notice. Trust collapses fast.

The pattern

Three properties researchers need that VERITAS provides out-of-the-box:

  1. Stable claim IDs. Every claim has a 16-hex identifier derived from canonical fields. Use the canonical URL to refetch the current record; SourceScore does not promise permanent hosting or byte-identical content for a fixed number of years.
  2. Cited evidence. Every record lists at least one source; many, but not all, source entries include an excerpt. Always inspect and cite the original source for academic work.
  3. HMAC metadata. Every envelope carries a SourceScore-issued HMAC tag. Public users cannot independently recompute it because the shared secret is not published.

Example: literature-review assistant

import httpx

# User asks: "What pretraining methods preceded BERT?"
# Your assistant retrieves relevant papers from arXiv.
# Before responding, retrieve a candidate record for each assertion.

assertions_to_check = [
    "BERT was introduced in 2019 by Devlin et al.",
    "T5 was introduced by Raffel et al. in 2020",
    "RoBERTa was introduced by Liu et al. at Facebook AI in 2019",
]

citations_to_review = []
for claim in assertions_to_check:
    r = httpx.post(
        "https://sourcescore.org/api/v1/verify",
        json={"claim": claim, "minConfidence": 0.85},
    )
    result = r.json()
    if result.get("bestMatch"):
        candidate = result["bestMatch"]
        envelope = httpx.get(candidate["detailUrl"]).json()
        citations_to_review.append({
            "claim": claim,
            "candidate_statement": candidate["statement"],
            "id": candidate["id"],
            "source_urls": [s["url"] for s in envelope["claim"]["sources"]],
            "excerpts": [s.get("excerpt") for s in envelope["claim"]["sources"]],
        })

# A reviewer or entailment step must compare each assertion, candidate
# statement, and original source before your assistant cites it.
# After that review, your assistant may cite:
#   "BERT (Devlin et al., 2019) [^1]"
# Where [^1] resolves to a citation block with:
#   - Stable ID: a1b2c3d4...
#   - Primary source: https://arxiv.org/abs/1810.04805
#   - Verbatim excerpt from the abstract
#   - SourceScore-issued HMAC metadata (not publicly independently verifiable)

Citation export format

For researchers who need machine-readable citations:

# BibTeX-style export for a VERITAS claim
@misc{sourcescore_a1b2c3d4,
  title = {SourceScore VERITAS reviewed claim record a1b2c3d4},
  publisher = {SourceScore},
  year = {2026},
  url = {https://sourcescore.org/claims/a1b2c3d4/},
  note = {SourceScore record with cited evidence: [URL1, URL2]. Public HMAC verification is not available.},
}

What the catalog covers

v0.1 catalog (384 claims spanning 1997-2025) covers AI/ML research:

  • Foundational papers — Transformer, LSTM, BERT, RLHF, RAG, LoRA, etc.
  • Model releases — GPT family, Claude family, Llama family, Gemini, Mistral, DeepSeek, Phi, etc.
  • Benchmarks + datasets — MMLU, GLUE, ImageNet, C4, The Pile, etc.
  • Frameworks + libraries — PyTorch, TensorFlow, JAX, LangChain, LlamaIndex, etc.
  • Organizations — OpenAI, Anthropic, DeepMind, Mistral, Hugging Face, etc.

Out of scope for v0: papers in scientific computing, cybersecurity, and biology. No expansion date is promised. Performance comparisons (see why we don't ship those).

License

The published methodology and claim data are CC-BY 4.0. Cite as: SourceScore Claim <id>, sourcescore.org. You can redistribute, re-publish, derive — under the attribution condition.

For academic submissions

For formal papers, cite the original primary source. If you also need to document the SourceScore record used during review, use:

SourceScore VERITAS (2026). Reviewed claim record <id>. https://sourcescore.org/claims/<id>/

Refetch the record at review time and preserve the primary-source citation in your own research materials. SourceScore does not promise perpetual hosting, and its HMAC tag is not public proof.

Integration guides

  • DSPy — for research workflows with optimizers
  • LangChain — retrieve-then-cite pattern
  • LlamaIndex — for paper-corpus RAG
  • Pydantic AI — type-safe structured citation output

Related