A citation that loads is not evidence. Picture a research assistant asked whether water boils at 50 degrees Celsius at sea level. It answers with a confident paragraph and a link. The link works. The page says 100 degrees. The URL was real; the citation was not support for the claim.

That is the gap in hallucination checks that only look at the generated answer. A sentence can be fluent, plausible and attached to a working URL while the cited source says something else. A link is only a pointer. Verifying it means fetching what is behind the pointer, comparing it with the claim, and keeping the evidence the decision rested on.

The service in this post automates that comparison. You give it a claim and the URLs the model cited. It fetches each page with the Crawling API, keeps the passages most relevant to the claim, asks a judge model for a verdict, checks that the judge's evidence quote really appears on the page, and combines the per-citation results into one label: supported, contradicted, not_found or unreachable.

The short version
  • HTTP 200 is not evidence. A page can load and still be a bot wall, an empty shell, a soft 404 or the site's homepage.
  • Send the judge three relevant passages, not the whole page. Less noise, lower cost, better verdicts.
  • Never trust the judge's quote. Accept supported or contradicted only when the quote is a real substring of the fetched page.
  • Aggregate with rules, not another model: one contradiction outweighs every supporting citation.
  • Store each fetch, so a verdict can be replayed after the page changes.
Five stages, and only one of them is a language model. Fetching, retrieval, the quote guard and aggregation are deterministic. The judge interprets the evidence; the code decides whether that interpretation is allowed to count.

The code lives in ScraperHub/llm-citation-verification-service-with-crawling-api, with one checkpoint per stage under steps/ and the complete package under final/citation_verifier/. Snippets below are quoted from final/ unless they are marked otherwise. You need a Crawlbase account; keep its tokens in environment variables, never in the repository.

Fetch the cited page, not just an HTTP 200

Why a plain GET is not enough

A successful request does not mean you retrieved the content a citation points to. A single-page app can return 200 with an empty <div id="root">. A bot wall can return 200 while serving a challenge. A soft 404 returns 200 with a "this page no longer exists" template, and a deep link can redirect to the homepage. The fetch stage has to establish more than reachability: it keeps the original status, notices redirects, and returns text fit to be used as evidence.

Markdown with a readability pass

The Crawling API handles the fetch. The verifier asks for format=md, which returns the page as GitHub Flavored Markdown, and md_readability=true, which strips navigation, sidebars, footers and ads before conversion so retrieval starts from the article itself. It also sets store=true so every fetch can be audited later.

python
# final/citation_verifier/fetch.py
def crawl_options(readability: bool) -> dict[str, str]:
    """Query parameters for one Crawling API call.

    ``md_readability`` is omitted on the fallback request. The docs define it
    only as a switch on ``format=md``, not as a guarantee about thin pages.
    """
    options = {
        "format": "md",
        "store": "true",
        "custom_success_codes": "404,410",
    }
    if readability:
        options["md_readability"] = "true"
    return options

One parameter here does no work. custom_success_codes exists for origin statuses the API would otherwise treat as failures, but 404 and 410 already count as successful crawls: they come back with cb_status 200 and the real original_status, and they are not retried. You can drop it without changing behaviour.

Readability can strip too much on short or unusual pages, so the verifier treats fewer than 200 non-whitespace characters as a failed pass and fetches the same URL again without md_readability. The 200-character threshold is an application rule, not a Crawlbase limit. Both calls are billed, and with store=true both pages are stored, so keep the retry for pages that need it.

Most bad citations are caught before any model runs. Dead links, homepage redirects and PDFs are labelled in the fetch stage, and only pages with usable text move on to retrieval.

Two statuses travel with every response. original_status is what the cited site answered; cb_status is Crawlbase's result for the crawl. The verifier labels any citation whose site answered 404 or 410 as unreachable with a reason such as http_404, and it never reaches the judge. That crawl is still billed, because Crawlbase did fetch an answer: a 404 is a successful crawl of a missing page.

Redirects need their own check. A deep link that lands on / returns a perfectly good page that says nothing about the claim, and generic homepage text could look like support. When the requested path is deeper than / but the final path is /, the citation becomes not_found with reason redirected_to_homepage, and the judge is skipped.

python
# final/citation_verifier/fetch.py
def _normalized_path(url: str) -> str:
    path = urlsplit(url).path or "/"
    if len(path) > 1 and path.endswith("/"):
        path = path[:-1]
    return path or "/"


def is_homepage_redirect(requested: str, final: str | None) -> bool:
    """True when a deep link landed on ``/``.

    The Crawling API returns the post-redirect URL in the ``url`` header.
    A citation that survives only as the site homepage did not survive.
    """
    if not final:
        return False
    return _normalized_path(requested) != "/" and _normalized_path(final) == "/"

The url header carries the final URL after the redirects Crawlbase follows on a normal request. On a JavaScript-rendered request, a redirect the browser performs by itself can leave url as the URL you asked for, so treat the homepage check as a strong signal on the normal path and a best effort on the rendered one.

Soft 404s are handled separately: a 200 page counts as a soft 404 only when its readable text is a short "page does not exist" template under 800 non-whitespace characters, so a long article that merely mentions a 404 is not thrown away.

When to reach for the JavaScript token

Rendering costs more, so the verifier starts every citation on the normal token and escalates only when the first attempt shows the page needs a browser. A JavaScript request is billed at twice the credits of a normal one. You can escalate with a second client on the JavaScript token, as the repository does, or keep a single client on the normal token and add javascript=true to the retry; either way the request is billed as a JavaScript request.

The repository's escalation check, in fetch.py, triggers on cb_status values of 520 or 525. Neither signal arrives in practice. 520 is not a cb_status at all: it is the HTTP status the Crawling API returns when a crawl fails, and the pinned crawlbase 1.0.0 SDK returns such a response with an empty body and no headers, so there is no cb_status to read. A page whose server answered 200 with nothing in it comes back as cb_status 207, and 525 is an internal code unrelated to anti-bot challenges. Escalate on the signals the API actually sends:

python
# Production shape, not from the repository: when to retry on the JavaScript token.
def needs_javascript(response: dict, markdown: str, min_chars: int) -> bool:
    if response.get("status_code") == 520:
        return True  # the crawl failed; failed requests are not billed
    headers = response.get("headers") or {}
    if str(headers.get("pc_status")) == "207":
        return True  # the origin answered 200 with an empty page
    return len("".join(markdown.split())) < min_chars  # still thin after the readability retry

The SDK version the repository pins exposes the status under its legacy name, pc_status; the API sends it as cb_status as well, and new integrations should read cb_status. After the JavaScript attempt, run the same interpretation again, including the readability fallback, so a rendered page goes through exactly the checks a static one does.

Retrieve the evidence before asking the judge

Sending a whole page to the judge is expensive and noisy. A 2,000-word article may hold two relevant sentences, and the rest gives the model more chances to latch onto navigation, disclaimers or unrelated text. So each page is cut into passages first: split on blank lines, pack short paragraphs together up to 200 words, and break longer ones on word boundaries. The 200-word budget is counted with split(), so no tokenizer is needed to keep passages roughly even.

python
# final/citation_verifier/passages.py (excerpt)
def split_passages(markdown: str, max_words: int = DEFAULT_MAX_WORDS) -> list[str]:
    """Split on blank lines, then pack up to ``max_words`` words.

    A paragraph longer than the cap is cut on word boundaries. Shorter
    paragraphs share a passage until the next one would overflow. Blank
    blocks are dropped.
    """
    if max_words < 1:
        raise ValueError("max_words must be at least 1")
    text = markdown.replace("\r\n", "\n").replace("\r", "\n")
    pieces: list[list[str]] = []
    for paragraph in re.split(r"\n\s*\n", text):
        words = paragraph.split()
        if not words:
            continue
        if len(words) > max_words:
            for start in range(0, len(words), max_words):
                pieces.append(words[start : start + max_words])
        else:
            pieces.append(words)

The passages are then ranked against the claim with all-MiniLM-L6-v2 embeddings and cosine similarity, and the top three are kept. The model loads on the first retrieval call, so importing the package never triggers a download.

python
# final/citation_verifier/retrieval.py
def top_k_passages(
    claim: str,
    passages: list[str],
    embedder: Embedder,
    k: int = DEFAULT_TOP_K,
) -> list[str]:
    """Return up to ``k`` passages, highest cosine similarity first."""
    if k < 1 or not passages:
        return []
    vectors = embedder.embed([claim, *passages])
    query = vectors[0]
    ranked = [
        (cosine(query, vectors[index + 1]), index, passage)
        for index, passage in enumerate(passages)
    ]
    ranked.sort(key=lambda item: (-item[0], item[1]))
    return [passage for _, _, passage in ranked[:k]]

Three is a deliberate cap. One passage makes retrieval brittle when the best evidence ranks second; ten brings back most of the noise retrieval was meant to remove. Evidence that falls outside the top three ends as not_found, which is the honest answer for evidence the verifier never showed the judge.

Make the judge return evidence, not just a verdict

The judge interprets evidence; it is not the source of truth. A model can return a confident supported with a quote that never appeared on the page, and strict JSON makes the output predictable without making it honest. So the prompt limits the judge to three labels and requires a contiguous quote copied from the passages for anything but not_found. unreachable never reaches the model; the fetch stage assigns it.

python
# final/citation_verifier/judge.py
SYSTEM_PROMPT = """You verify one claim against the passages in the user message. The passages are the only evidence you may use.

Return JSON with these fields:
- label: supported, contradicted, or not_found
- evidence_quote: a contiguous quote copied from the passages, or null when the label is not_found
- confidence: a number from 0 to 1

supported means the passages state the claim. contradicted means the passages state the opposite. not_found means the passages do not settle the claim. If you cannot copy a quote that appears in the passages, return not_found and null. Do not add words, numbers, or sources that are not in the passages."""


class RawJudgement(BaseModel):
    label: Label
    evidence_quote: str | None = None
    confidence: float = Field(ge=0, le=1)

OpenAI is the default judge, called through httpx with a strict JSON schema in response_format; JUDGE_PROVIDER=anthropic and JUDGE_PROVIDER=ollama run the same prompt elsewhere. The provider changes how a verdict is produced, never the rules that decide whether it is accepted.

Guard the quote against the page

The guard is the step that turns the judge from an oracle into an interpreter. It straightens curly quotes, collapses whitespace, and requires the quote to appear as an exact substring of the retrieved passages. A supported or contradicted verdict without a matching quote is downgraded to not_found with confidence 0 and reason evidence_quote_not_in_source.

A verdict survives only with a quote the page contains. The judge can be right about the claim and still invent its evidence; the substring check catches that before the verdict counts.
python
# final/citation_verifier/judge.py
def guard_judgement(
    raw: RawJudgement,
    passages: list[str],
) -> tuple[Label, str | None, float, str | None]:
    """Drop a quote the passages do not contain.

    Supported and contradicted both need a real quote. A quote that fails the
    check downgrades the citation to not_found with confidence 0.
    """
    quote = raw.evidence_quote
    if raw.label in {Label.supported, Label.contradicted}:
        if quote and quote_in_passages(quote, passages):
            return raw.label, quote, raw.confidence, None
        return Label.not_found, None, 0.0, REASON_QUOTE
    if quote and not quote_in_passages(quote, passages):
        return Label.not_found, None, 0.0, REASON_QUOTE
    return Label.not_found, None, raw.confidence, None

In the boiling-point example the claim says 50 degrees and the page says 100, so the contradicted verdict survives: its quote is on the page. The model's confidence is kept only when the evidence passes; aggregation ignores it and works from labels alone.

Expose it as a /verify API

The finished service has two routes: GET /health and POST /verify. A request carries a claim and one or more cited URLs; the response carries the claim-level label and a verdict per citation with its quote, confidence, final URL, original_status, the storage rid and any reason the verifier assigned.

python
# final/citation_verifier/schema.py
class UrlVerdict(BaseModel):
    url: str
    label: Label
    evidence_quote: str | None = None
    confidence: float = Field(ge=0, le=1)
    final_url: str | None = None
    original_status: int | None = None
    rid: str | None = None
    reason: str | None = None


class VerifyRequest(BaseModel):
    claim: str = Field(min_length=1)
    urls: list[str] = Field(min_length=1)

Aggregation is a fixed rule, not another model call. One contradicted citation outweighs every supported one; with no contradiction, any support carries the claim; the claim is unreachable only when every citation is; everything else is not_found.

python
# final/citation_verifier/aggregate.py
def aggregate(verdicts: list[UrlVerdict]) -> Label:
    """One contradiction outweighs every supporting citation.

    Otherwise any supported citation carries the claim. The claim is
    unreachable only when every citation is unreachable. Every remaining mix,
    including an empty list, is not_found.
    """
    labels = [verdict.label for verdict in verdicts]
    if any(label == Label.contradicted for label in labels):
        return Label.contradicted
    if any(label == Label.supported for label in labels):
        return Label.supported
    if labels and all(label == Label.unreachable for label in labels):
        return Label.unreachable
    return Label.not_found

The HTTP layer stays thin: POST /verify hands the request to the verifier and returns 503 when a required token is missing, naming the environment variable without exposing its value. The repository's dry run, python steps/step_07_api/main.py --dry-run, produces this response. The pages are fixtures on example.com, the two Crawlbase requests were live, and the judge used a fixture sentence because no OpenAI key was set, so read it as a fixture rather than a model trace:

json
{
  "claim": "Water boils at 100 degrees Celsius at sea level.",
  "overall": "supported",
  "citations": [
    {
      "url": "https://example.com/articles/water-boiling-point",
      "label": "supported",
      "evidence_quote": "Water boils at 100 degrees Celsius at sea level.",
      "confidence": 0.96,
      "final_url": "https://example.com/articles/water-boiling-point",
      "original_status": 200,
      "rid": "rid-boil",
      "reason": null
    },
    {
      "url": "https://example.com/old-report",
      "label": "unreachable",
      "evidence_quote": null,
      "confidence": 1.0,
      "final_url": "https://example.com/old-report",
      "original_status": 404,
      "rid": "rid-404",
      "reason": "http_404"
    }
  ]
}

The claim comes out supported: one citation supports it and the other is simply gone. That distinction is the point of keeping four labels. A dead source is not evidence against a claim; a live source that says the opposite is.

Crawlbase Crawling API

Fetch the page a citation points to as clean Markdown, rendered when it needs a browser, with the site's real status code alongside it. Failed requests are not billed. Start with up to 5,000 free requests, no card.

Score citation precision across a batch

One verified claim demonstrates the pipeline; it does not tell you how reliably a model cites. For that, the verifier scores a batch: a JSONL file of records with an id, a claim and its cited URLs, where every claim-URL pair counts as one citation. Citation precision is supported citations divided by all citations. Reachable-only precision keeps the same numerator and drops unreachable citations from the denominator, separating "the model cites dead links" from "the model cites live pages that do not say what it claims".

python
# final/citation_verifier/batch.py
@dataclass(frozen=True)
class Precision:
    """Citation precision over claim-URL pairs.

    citation precision is supported citations divided by every citation.
    Reachable-only precision uses the same numerator and drops unreachable
    citations from the denominator. It is None when nothing was reachable.
    """

    supported: int
    total: int
    reachable: int

Runs are capped by an asyncio.Semaphore (BATCH_CONCURRENCY, default 4), and because the Crawlbase SDK call is synchronous it runs in a worker thread instead of blocking the event loop. On the sample fixture the dry run reports nine citations, two supported and three unreachable:

text
citation_precision=0.2222 (2/9)
reachable_precision=0.3333 (2/6)

Run it against your own output with python -m citation_verifier.batch your_outputs.jsonl --out citation_report.csv once the tokens and a judge key are set. Start with a few dozen claims and read the contradicted rows first: one verified contradiction can expose a problem that an aggregate accuracy score hides.

Keep the evidence with Cloud Storage

A precision score is only as good as your ability to inspect it later, and the cited page may change the day after you check it. With store=true, Crawlbase keeps the Markdown it returned in Cloud Storage for 30 days, and the response carries a storage_url whose query string holds the page's rid. The verifier stores only the rid, because the full URL contains the token, and keeps a small SQLite audit index of rid, claim, URL, verdict, confidence and time. Storing adds half a credit per page to the crawl; reading a stored page back is free.

python
# final/citation_verifier/audit.py (docstring omitted)
def replay_rid(rid: str, token: str | None = None) -> str:
    if not rid:
        raise ValueError("rid is required")
    settings = get_settings()
    used = token if token else settings.crawlbase_token
    if not used:
        raise MissingTokenError(
            "CRAWLBASE_TOKEN is not set. Replay needs the token that stored the page."
        )
    client = StorageAPI({"token": used})
    response = _storage_get(client, rid)

Stored pages belong to your account, not to the token that crawled them, so either token can replay any rid; the repository's comments suggest otherwise, and you can ignore that restriction. In the pinned crawlbase 1.0.0 SDK, StorageAPI.get takes the rid as a positional argument. A citation that never produced a stored page, such as a skipped PDF, still gets an audit row with an empty rid, so the decision is recorded even when there is no body to replay.

Limits: paywalls, PDFs and multi-source claims

Paywalls. A paywalled page can answer 200 with only a subscription prompt. If that is what retrieval finds, the correct verdict is not_found: the publicly accessible page does not support the claim. That is narrower than "the claim is false", and more honest than paywall heuristics that misfire on articles which merely mention subscriptions.

PDFs. A URL ending in .pdf is marked not_found with reason pdf_unsupported. Crawlbase's pdf=true renders a web page into a PDF; it does not extract text from a URL that already is one, so the verifier declines rather than judge text it cannot vouch for.

Multi-source claims. Each URL is judged on its own. A claim that only follows from combining two documents stays not_found, because the substring guard can prove a quote came from a page but cannot validate an inference across pages. The verifier does not establish that a claim is true; it establishes whether the cited, accessible source supported it when it was checked.

Conclusion

Citation verification belongs in the pipeline as a data-validation step, not as another prompt on top of the model. Fetch the cited source, reduce it to relevant evidence, let a model interpret that evidence, and then apply deterministic checks before anything counts. Each failure then has a name and a reason you can test and audit.

Start with the dry-run /verify flow in ScraperHub/llm-citation-verification-service-with-crawling-api, then create a free Crawlbase account, set CRAWLBASE_TOKEN, CRAWLBASE_JS_TOKEN and OPENAI_API_KEY, and run it against citations from your own output.

Frequently Asked Questions (FAQs)

How does the verifier check an LLM's sources?

It fetches each cited URL with the Crawling API as Markdown, keeps the three passages most relevant to the claim, and asks a judge model for a label and a supporting quote. A deterministic guard then accepts the verdict only if the quote is an exact substring of the fetched passages.

What happens when a cited page is gone?

A citation whose site answers 404 or 410 is labelled unreachable and never reaches the judge. A deep link that redirects to the homepage is labelled not_found instead, because the page it pointed to no longer exists at that address.

Why fetch pages as Markdown?

Markdown keeps the structure of the page without the markup, so it splits cleanly into passages and costs far fewer tokens than HTML. With md_readability=true, navigation, sidebars, footers and ads are removed before conversion, so retrieval starts from the content itself.

When does a citation need the JavaScript token?

When the normal-token crawl fails or comes back as an empty shell, which is what pages built in the browser look like without rendering. Retry those on the JavaScript token, or with javascript=true on the normal token. A rendered request is billed at twice the credits, which is why it is a fallback and not the default.

Can this verify any claim an LLM makes?

It verifies whether an accessible cited page supports or contradicts the claim. It does not validate claims that need several documents combined, extract text from PDF URLs, or see content behind a paywall.

Start Building

Crawl any site at scale, without fighting infrastructure.

Crawlbase handles proxies, fingerprints, and CAPTCHAs so your team ships data pipelines instead of maintaining crawl plumbing. Up to 5,000 requests free, no card required.

Self-serve · No sales call required · Enterprise crawl volumes available