SourceScore

Content moderation — fact-check LLM outputs before publishing

Editorial AI tools, content generation platforms, and publishing assistants ship hallucinated facts to thousands of readers. Add verification between draft and publish.

The problem

AI content tools — newsletter generators, blog assistants, report drafters, automated summary tools — generate fluent text fast. Their failure mode at scale is shipping hallucinated facts to thousands of readers.

Examples of damage at scale:

  • Tech newsletter auto-summarizes a paper, gets the author name wrong.
  • Industry-report tool cites a release date that's 2 years off.
  • Blog assistant attributes a quote to the wrong founder.
  • Marketing copy generator invents a non-existent integration.

The reader doesn't know it's wrong. The error compounds — gets reshared, quoted, indexed. Months later you're searching for "Llama 3 released 2025" and seeing your own incorrect content cited back at you.

The pattern

Pre-publish verification gate. Three steps between LLM draft and publish button:

  1. Extract atomic claims. Parse the draft into discrete assertions (dates, names, numbers, attributions).
  2. Retrieve candidate evidence. Use domain catalogs and primary sources. VERITAS can return nearby AI/ML records, but similarity is not a truth verdict.
  3. Review support before publishing. Route candidate and unmatched assertions to a human or a separate entailment check; never auto-publish from bestMatch alone.

Implementation

# Python — content moderation pipeline
import re
import httpx
from typing import Literal

class ClaimCheck:
    text: str
    status: Literal["candidate", "no_match"]
    source_url: str | None = None
    confidence: float | None = None

def extract_factual_claims(draft: str) -> list[str]:
    # Naive: extract sentences with proper nouns + numbers + dates
    # Production: use a dedicated claim-extraction model
    sentences = re.split(r'(?<=[.!?])\s+', draft)
    return [
        s for s in sentences
        if re.search(r'\b\d{4}\b|\b[A-Z][a-z]+\s+[A-Z][a-z]+\b', s)
    ]

def verify_aiml(claim: str) -> ClaimCheck:
    r = httpx.post(
        "https://sourcescore.org/api/v1/verify",
        json={"claim": claim, "minConfidence": 0.85},
        timeout=2.0,
    )
    result = r.json()
    match = result.get("bestMatch")
    if match and match["confidence"] >= 0.85:
        return ClaimCheck(
            text=claim,
            status="candidate",
            source_url=match["detailUrl"],
            confidence=match["confidence"],
        )
    return ClaimCheck(text=claim, status="no_match")

def moderate(draft: str) -> dict:
    claims = extract_factual_claims(draft)
    checks = [verify_aiml(c) for c in claims]

    return {
        "draft": draft,
        "claims_checked": len(checks),
        "candidate_count": sum(1 for c in checks if c.status == "candidate"),
        "no_match_count": sum(1 for c in checks if c.status == "no_match"),
        "checks": checks,
        "review_required": True,
    }

# In your editorial workflow:
result = moderate(llm_draft)
route_to_human_review(result["draft"], result["checks"])

Use across editorial workflows

  • Newsletter platforms. Pre-flight every AI-generated section. Show editors a list of unverified claims with one-click strike-through.
  • Auto-summary tools. Attach candidate records for review; do not treat record confidence as query entailment.
  • SEO-content platforms. Block publish until each factual assertion is supported by evidence a reviewer or dedicated entailment step has checked.
  • Internal company comms. Verify before sending all-hands or external comms drafted by AI.

What this catches vs misses

Catches well:

  • Wrong dates (Llama 3 released 2025 — wrong)
  • Wrong attributions (Transformer paper by Hinton — wrong)
  • Hallucinated specs (32k context window when source says 128k)
  • Made-up citations

Doesn't catch:

  • Plausible-sounding new claims not in any catalog (genuine ambiguity)
  • Style + tone issues
  • Bias + misleading framing of correct facts
  • Plagiarism / verbatim copy from a source

Free-tier viability

The SourceScore VERITAS public API is free with no account, key, or signup. A newsletter publishing 4 issues per week with 5 verifiable claims per issue would make roughly 80 candidate-lookups per month. Higher-volume paid access is a demand test only; no paid plan or service commitment is live.

Related