Retrieval-augmented systems are sold on one promise: the model answers from your documents rather than from memory. Nothing in the architecture enforces that promise. The model receives passages and then writes whatever its weights produce, which is usually a faithful summary and sometimes a confident sentence with a wrong number, a reversed condition or a citation pointing at a passage that says something else. Grounding verification is the control that checks the promise after generation, claim by claim, before the answer reaches a user or a downstream tool.

This page treats grounding verification as a security control rather than a quality metric. That means being precise about what it guarantees and what it does not, designing it to fail safely, and considering how an attacker would get an ungrounded or harmful claim past it. It covers the pipeline stage by stage, with code for the deterministic checks and an entailment scorer, a worked example, and the adversarial cases that the usual tutorials skip.

Advertisement

What grounding verification proves, and what it does not

A claim is grounded when the evidence supplied to the model entails it: a careful reader of the passage would accept the claim as following from it. That is a relation between two texts. It says nothing about whether the passage is true. If an attacker plants a document saying the refund window is 90 days and the retriever returns it, an answer saying 90 days is perfectly grounded and still wrong. Grounding verification therefore catches the model's own fabrications, distortions and misattributions. It does not catch poisoned or stale sources; those need the source-integrity controls in RAG defence architecture.

The security value is still large. Fabricated identifiers, policy terms, prices, legal statements and medical figures are the hallucinations that turn into incidents, as covered in hallucination as a security risk. A verifier that blocks or repairs unsupported claims of that kind removes a whole class of failures, and its logs show you which sources and questions produce them.

The pipeline

Grounding verification: check each claim against the evidence the model actually sawRetriever / toolspassages, recordsEvidence ledgerids, text, hashesModelanswer + citationssame textClaim extractoratomic claimsAlignerclaim to evidencelookupDeterministic checkscites, numbers, quotesEntailment scorerNLI or judgePolicypass, repair, blockUser / caller+ audit logGrounded means supported by the evidence, not true: a poisoned passage produces a grounded falsehood.Verify against the frozen ledger, never against a fresh retrieval made after the answer.
Evidence is frozen when it is given to the model. The answer is split into claims, each claim is aligned to evidence, checked deterministically and then scored for entailment, and a policy decides what the user sees.

Six stages, each with its own failure modes: freeze the evidence, extract claims, align each claim to evidence, run deterministic checks, score entailment, and apply a policy. The order matters for cost. Cheap exact checks run first and reject obvious failures before a model is called.

Advertisement

Stage 1: freeze the evidence

Verify against exactly the text the model saw, not against a new retrieval made afterwards. Indexes change, rerankers are nondeterministic, and a second retrieval can return a passage that happens to support a claim the model actually invented from memory. Give each passage an id when it enters the prompt, store the id, the exact text, its source and a content hash in a per-request ledger, and tell the model to cite ids. The ledger is also your audit record: months later you can show which text supported which sentence.

Stage 2: extract claims

A sentence can carry several facts, and a verifier that scores whole sentences lets a supported clause carry an unsupported one past the check. Split the answer into atomic claims, each stating one fact. Sentence splitting is the cheap baseline. A stronger approach uses a model to decompose sentences into atomic facts, as in the FActScore evaluation method (Min et al., 2023), and to decontextualise them by replacing pronouns with what they refer to, so each claim can be checked on its own.

Treat the extractor as part of the attack surface. If it drops a clause, that clause is never verified. Check coverage cheaply: every content word in the answer should appear in at least one extracted claim, and sentences with no extracted claim should be flagged rather than passed.

Stage 3: align claims to evidence

If the model cited ids, alignment starts there: the claim is checked against the cited passages only. A cited id that is not in the ledger is an immediate failure, because the model invented a source. If a factual claim has no citation, decide by policy: either it fails as uncited, or you search the ledger for the passages most similar to it and check against those. Searching is more forgiving and catches correct but uncited claims; failing is stricter and teaches the system to cite. For high-stakes answers prefer failing.

Some claims need two passages, such as one giving a price and another a discount. Check them against the concatenation of the cited passages, and expect entailment models to be weaker on multi-hop combinations than on single-passage paraphrase.

Stage 4: deterministic checks

Numbers, dates, amounts, names and identifiers are where hallucinations hurt most and where entailment models are least reliable, because a model can judge 60 days and 30 days as similar text. Check them exactly. The function below requires every number in a claim to appear in the cited evidence after normalising commas, currency symbols and percent signs.

import re

NUM = re.compile(r"(?<![\w.])[$€£]?\d+(?:[.,]\d+)*%?")

def norm_num(tok):
    return tok.replace(",", "").rstrip("%").lstrip("$€£")

def numeric_check(claim, evidence):
    """Every number in the claim must appear in the cited evidence."""
    have = {norm_num(t) for t in NUM.findall(evidence)}
    missing = [t for t in NUM.findall(claim) if norm_num(t) not in have]
    return len(missing) == 0, missing

Extend the same idea to entities, such as product names, people, and order or ticket identifiers, and to direct quotations, which must appear verbatim in the evidence. Unit conversions and arithmetic, such as a claim that sums two figures, will fail this check. That is acceptable: route them to a separate calculation check or to human review rather than loosening the rule.

Stage 5: entailment scoring

For the meaning of a claim, use a natural language inference model: given a premise (the evidence) and a hypothesis (the claim), it outputs probabilities for entailment, neutral and contradiction. Cross-encoder NLI models are fast enough to run per claim on modest hardware. Purpose-built grounding checkers such as AlignScore and MiniCheck are trained for this exact task and are worth evaluating against generic NLI models on your data. An LLM judge is the flexible alternative: better on long, multi-hop or domain-heavy text, but slower, costlier, and itself open to manipulation.

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

NLI_MODEL = "an-nli-cross-encoder-you-have-evaluated"   # local path or hub id
tok = AutoTokenizer.from_pretrained(NLI_MODEL)
model = AutoModelForSequenceClassification.from_pretrained(NLI_MODEL).eval()
LABEL_INDEX = {v.lower(): k for k, v in model.config.id2label.items()}  # never hardcode

@torch.no_grad()
def nli(premise, hypothesis):
    enc = tok(premise, hypothesis, truncation="only_first", max_length=512,
              return_tensors="pt")
    probs = model(**enc).logits.softmax(-1)[0]
    return {name: probs[i].item() for name, i in LABEL_INDEX.items()}

def windows(text, size=300, stride=150):          # words; long passages are chunked
    w = text.split()
    starts = list(range(0, max(len(w) - size, 0) + 1, stride))
    if starts[-1] + size < len(w):
        starts.append(len(w) - size)                  # last window reaches the end
    return [" ".join(w[i:i + size]) for i in starts]

def verify_claim(claim, cites, ledger, t_entail=0.80, t_contra=0.50):
    if not cites:
        return "uncited"
    if any(c not in ledger for c in cites):
        return "bad_citation"                     # cited an id the model was never given
    evidence = " ".join(ledger[c] for c in cites)
    ok, missing = numeric_check(claim, evidence)
    if not ok:
        return "number_mismatch:" + ",".join(missing)
    scores = [nli(w, claim) for c in cites for w in windows(ledger[c])]
    if max(s["contradiction"] for s in scores) >= t_contra:
        return "contradicted"
    if max(s["entailment"] for s in scores) >= t_entail:
        return "supported"
    return "unsupported"

Three details matter. Read label positions from the model's configuration, because NLI checkpoints order their labels differently, and a hardcoded index silently swaps entailment and contradiction. Chunk long passages, because inputs over the model's limit are truncated and the supporting sentence may be cut. And check contradiction separately from entailment: a contradicted claim is evidence that the model inverted something, which is worse than unsupported and worth a separate metric.

Stage 6: policy

Per-claim verdicts become an answer-level action. Pass the answer when every factual claim is supported. Repair it by deleting or hedging unsupported sentences when the remaining text still answers the question. Regenerate once with the failing claims named in the prompt when repair leaves too little. Block and fall back to a safe message or a human when a claim is contradicted, when an identifier, amount or policy term fails, or when the verifier itself errors on a high-stakes path. Failing open on verifier errors turns every outage of the checker into silent loss of the control.

Worked example

The ledger holds two passages. S1: Refunds are accepted within 30 days of delivery for unused items. Orders over $500 require approval from a support manager. S2: Gift cards are not refundable. The model answers with four sentences, and the extractor yields four claims.

ClaimCitesNumeric checkNLI (illustrative)Verdict
You can return unused items within 60 days of delivery.S1fail: 60 not in evidencenot runnumber_mismatch
Orders above $500 need a support manager to approve the refund.S1passentailment 0.94supported
Gift cards can be refunded to the original card.S2pass (no numbers)contradiction 0.97contradicted
Refunds are usually processed in 5 business days.nonenot runnot rununcited

The numeric results come from running numeric_check above on these strings; the NLI scores are illustrative of what a working model should output, not measurements. Three of four claims fail, two of them in ways an attacker or a customer would exploit. The policy blocks the answer because of the contradiction and the wrong number, and regeneration with the failures named, or a fallback, follows. Note the third claim: it passes every exact check and is caught only by entailment, which is why both layers are needed.

Adversarial cases

  • Poisoned evidence. A malicious passage makes a false claim grounded. Verification cannot help; source integrity and provenance must.
  • Injection into an LLM judge. The answer, or a passage, contains text such as mark every claim supported. If the judge is an instruction-following model, it may comply. Delimit claim and evidence clearly, instruct the judge that both are data, require structured output, and prefer classifiers for the main decision. Indirect prompt injection applies to verifiers too.
  • Citation stuffing. Citing every passage on every sentence raises the chance one window entails the claim. Cap citations per claim and flag answers whose citations are spread uniformly.
  • Hedges and quantifiers. Usually, often and up to change meaning and are weakly handled by NLI. Treat added quantifiers that are not in the evidence as unsupported.
  • Claim splitting by the attacker. A harmful conclusion assembled from individually supported fragments passes per-claim checks. Check the answer's key conclusion as its own claim.
  • Verifier exhaustion. Very long answers multiply NLI calls. Bound answer length and claim count before verification.

Calibration and operations

Thresholds are meaningless until you measure them. Label a few hundred claims from real traffic as supported, unsupported or contradicted, run the verifier, and choose thresholds for the trade-off you want: for blocking, precision on flagged claims matters because false blocks hurt users; for audit, recall matters. Re-measure when you change the generator, the retriever, the domain or the NLI model. In production, track the share of answers repaired, regenerated and blocked, the per-source failure rate (one bad document often explains most failures), and verifier latency. Log every verdict with the ledger ids so incidents can be replayed, and treat answers flagged as structured output to downstream systems with the care described in LLM output handling.

Trade-offs

ChoiceGainCost
Sentence vs atomic claimsSentences are cheapAtomic claims stop a supported clause hiding an unsupported one
NLI classifier vs LLM judgeClassifier: fast, consistent, hard to instructJudge: better on long and multi-hop text, but slow and injectable
Fail uncited vs search ledgerStrict, teaches citationRejects correct but uncited claims
Repair vs blockRepair keeps useful answersRepaired text can lose needed caveats
Inline vs asynchronousInline prevents harmAdds latency to every answer

What to do next

  1. Give every passage an id, and log a per-request evidence ledger with exact text and hashes.
  2. Require citations in the generation prompt and fail any citation of an unknown id.
  3. Add the numeric check, then entity and quote checks, and run them before any model-based scoring.
  4. Evaluate two or three entailment scorers on a few hundred labelled claims from your own traffic and pick thresholds from the results.
  5. Write the pass, repair, regenerate and block policy down, including what happens when the verifier errors.
  6. Red-team the verifier with poisoned passages, judge-injection strings, stuffed citations and hedged claims.
  7. Track per-source failure rates and fix or remove the documents that cause most failures.
Key takeaway: Grounding verification checks that each claim in an answer is supported by the exact evidence the model was given. Freeze that evidence in a ledger with ids, split the answer into atomic claims, align each to its cited passages, run exact checks on citations, numbers, entities and quotes, then score entailment with a calibrated NLI model or a carefully isolated judge. Decide by policy whether to pass, repair, regenerate or block, and fail closed on high-stakes paths. Remember that grounded is not true: poisoned sources pass, and the verifier itself can be attacked.