Yes, you can run a useful hallucination guardrail without a GPU, but only for some tasks. On the evidence available, a similarity-plus-NLI ensemble running on CPU catches question-answering hallucinations well (F1 0.792, AUC-ROC 0.873) and dialogue hallucinations moderately (F1 0.694, AUC-ROC 0.749), while every cheap detector tested performs near chance on summarisation in its default configuration. One caveat belongs up front, because it conditions everything below: these numbers come from a single arXiv preprint, a systematic benchmark of lightweight hallucination detection, and the only companion artifact is the authors’ own code repository. No independent replication exists yet. Treat what follows as a well-documented starting framework, not settled fact, and plan to validate on your own traffic before you trust any threshold.
That said, the preprint is unusually useful for a practitioner because it does not just rank detectors. It documents what each detector is allowed to see, what a fix costs when it fails, and where the cheap tier ends.
What was actually benchmarked
The study evaluates four CPU-feasible detectors plus an ensemble across the three tasks of the HaluEval benchmark: question answering, dialogue, and summarisation. The lineup spans three detection paradigms:
- ROUGE-L: lexical overlap between the response and the source. No model inference at all.
- Semantic similarity: embedding cosine similarity using all-MiniLM-L6-v2.
- BERTScore: token-level embedding matching against a reference.
- NLI-based detection: a natural language inference model, DeBERTa-v3-base-mnli-fever-anli at roughly 184M parameters, scoring the response as 1 − P(entailment) against the source.
- An ensemble: a score-level combination of similarity and NLI.
The repository confirms the whole suite runs CPU-only. “CPU-only” deserves a precise reading: it does not mean free. The NLI detector is a 184M-parameter transformer inference per response. It means no accelerator, no second LLM call, and a model small enough to colocate with an application server, which is a very different cost posture from routing every output through an LLM judge.
The task-by-task scorecard
The central finding is that detector quality is a property of the task, not the detector. From the preprint:
| Task | Best cheap option | Reported score | Cost posture | Verdict |
|---|---|---|---|---|
| Question answering | Similarity + NLI ensemble | F1 0.792, AUC-ROC 0.873 | 1 NLI call + embedding per response | Ship it, with local validation |
| Dialogue | Similarity + NLI ensemble (NLI strongest standalone) | F1 0.694, AUC-ROC 0.749 | Same as QA | Usable, tune thresholds on your traffic |
| Summarisation, single-pass | None | AUC-ROC ≤ 0.574 for every method | Same as QA | Do not ship |
| Summarisation, chunk aggregation | NLI with sentence-level chunks | AUC-ROC 0.567 → 0.683 | ~20× the NLI calls | Viable if the latency budget allows |
QA is the comfortable case. An AUC of 0.873 means the ensemble ranks hallucinated answers above faithful ones reliably enough to gate retries or flag for review. Dialogue is middling but workable, with NLI doing the heavy lifting. Summarisation single-pass is the silent failure: at or below 0.574 AUC, the detector is barely better than a coin flip, and a guardrail that cannot distinguish hallucination from faithfulness is worse than no guardrail, because it lends false confidence to whatever it approves.
The summarisation gap is partly recoverable without leaving the CPU. The preprint reports two levers: raising the premise budget recovers about six AUC-ROC points, and sentence-level chunk aggregation, scoring each summary sentence against chunked premises in a SummaC-style pattern and aggregating, recovers about twelve. The repository states the measured outcome: “SummaC-style chunk aggregation lifts summarisation AUC-ROC from 0.567 to 0.683, still on CPU with the same model, at roughly 20× the NLI calls.”
Two things about that 0.683 deserve emphasis. First, it is recovery, not parity: it remains well below the QA number, so even the fixed version is a weaker guardrail. Second, the “20×” counts NLI calls, not measured wall-clock latency. Twenty sequential CPU inferences of a 184M-parameter model per response is a real latency bill, and whether you can pay it depends on your serving path, not on this paper.
Why summarisation fails: an input-budget confound
Before concluding that cheap detectors simply cannot do summarisation, look at what they were allowed to read. Per the repository: “The NLI premise is capped at 800 source characters, BERTScore references at 512 characters,” and “all-MiniLM-L6-v2 truncates at its default 256 tokens; only ROUGE-L sees the full source” (repository input limits).
Summarisation sources are documents, often far longer than 800 characters; the preprint measures a median source length of 3,458 characters. A detector asked “is this summary entailed by the source?” while only seeing the first 800 characters of that source (the preprint’s cap) is being asked a different, easier-to-fail question. Some share of the near-chance performance is likely an artifact of truncation rather than a property of NLI as a method. The preprint’s own premise-budget experiment supports this reading: widening the input window recovers six AUC points without changing the detector at all, but the upside is bounded. The preprint’s sweep peaks at 1,600 characters (AUC-ROC 0.629), then falls back to 0.577 at 3,200: the checkpoint is configured for at most 512 positions, and longer premises degrade its entailment judgments.
The practical consequence cuts both ways. Do not write off NLI for long-document tasks until you have tested with a premise budget that matches your document lengths, but treat widening as a bounded lever: past the encoder’s configured length it degrades performance rather than helping, and for genuinely long sources the paper’s own recommendation is aggregation, not a larger cap. Do not assume the chunk-aggregation fix transfers to your documents either, because the benchmark’s recovery numbers were measured under its own chunking and caps.
Where the cheap tier ends
Two companion studies bound the cheap-detector approach in cost and in access requirements, and both matter for sizing decisions.
The costlier tier. A temporal multi-signal fusion study reaches 0.840 AUC on RAGTruth by fusing multiple signals under temporal modeling, where no single signal exceeded 0.641 alone. The catch is the serving cost: “Feature extraction requires inference passes through DeBERTa (350M parameters) and TinyLlama (1.1B parameters) at test time; in latency-sensitive applications, this overhead may be a limiting factor.” That is the authors’ own caveat. What that 0.840 establishes is internal to the fusion study’s own benchmark: fused temporal modeling versus 0.641 for the best single signal. It cannot be ranked against the CPU benchmark’s 0.873 on HaluEval QA, because the two studies share no dataset and score at different granularities, token-level on RAGTruth versus response-level on HaluEval. What is not in dispute is the cost of reaching it: test-time model inference that starts to resemble the extra-hop economics you were trying to avoid.
The access requirement. A RAG-focused method based on Maximum Mean Discrepancy measures divergence between the distributions of query and response hidden states. It reads model internals, not output text. If your LLM comes through an API that returns only text, this entire detector class is unavailable to you regardless of its accuracy, and that constraint, not benchmarks, decides the question. Output-text detectors like the NLI ensemble are the default for API-served models precisely because they need nothing the API does not already give you.
Budgeting detector cost like an engineer
The 20× figure is a call count, and call counts are not budgets. A separate benchmark, OpenHalDet, proposes the right accounting discipline: “We report Cost@N, the wall-clock time for applying a detector to N samples under a fixed hardware and evaluation protocol, decomposed into feature preparation, training, and inference time.” OpenHalDet spans 17 datasets and full evaluations on 4 backbone LLMs, and its framing is the one to steal: measure wall-clock per thousand responses on your hardware, decomposed so you can see whether the bill is feature extraction, the model forward pass, or orchestration overhead.
Applied here, that means three concrete measurements before you commit. Single-pass ensemble cost per response on your CPU fleet. Chunk-aggregation cost on your longest summarisation inputs, since 20 NLI calls against short premises differ from 20 against long ones. And p99 latency, not the mean, because a guardrail sits in the serving path and its tail becomes your users’ tail.
What a miss costs downstream
Detection quality is not an abstract metric when the detector gates generation. A study of package-name hallucinations in local coding LLMs shows what imperfect detection converts into: across 300 curated prompts, a guarded pipeline produced hallucination-free code on 76% of runs, and the primary model exhausted its retry budget on 28.7%. That is a different domain (slopsquatting risk in generated package names), so treat it as an illustration of mechanism, not a prediction for your pipeline. The mechanism is general: when detection misses, the system either ships the bad output or burns its retry budget and stalls, and a detector near chance on your task pushes you toward both failure modes at once.
This is the real argument against shipping a single-pass summarisation guardrail at 0.574 AUC. You are not adding safety; you are adding latency and retry churn while approving hallucinations at roughly the same rate.
A selection protocol, and what would change this verdict
Putting the evidence together, here is how I would allocate guardrail budget:
- QA pipelines: deploy the similarity + NLI ensemble on CPU. The margin over chance is large, the cost is one small-model inference per response, and the access requirement is just output text plus source.
- Dialogue pipelines: same ensemble, lower expectations. NLI is the strongest standalone method here, so if you must simplify to one detector, keep the NLI half. Tune decision thresholds on your own transcripts.
- Summarisation pipelines: do not ship any single-pass cheap detector. Either budget roughly 20× NLI calls for sentence-level chunk aggregation and accept a 0.683-class detector, or escalate summarisation to a heavier verifier and keep CPU detection for the other tasks. If your documents are long, re-test the premise budget against them, but know the preprint’s sweep peaks at 1,600 characters and degrades past the checkpoint’s configured 512 positions; for long sources the paper’s own recommendation is aggregation.
- Every pipeline: validate on your own traffic before trusting any threshold, and budget with a Cost@N-style wall-clock protocol rather than call counts.
What would change this verdict is straightforward to enumerate. Independent replication of the HaluEval numbers would let you trust the thresholds rather than treating them as priors. Evidence on domain-specific, long-form, or non-English workloads would settle transfer, which this benchmark cannot. And a direct comparison against LLM-as-judge or SelfCheckGPT-style baselines on the same tasks would answer the cost-quality question this article opens with; the benchmark evaluates only the four lightweight detectors plus the ensemble, with no LLM-as-judge or SelfCheckGPT baseline, so the claim that the CPU ensemble is “good enough” rests on its absolute task scores, not on a measured comparison against the expensive alternative.
The economics are worth stating plainly once. If the numbers hold, output guardrails for QA and dialogue stop being priced like an extra inference hop and start being priced like a small CPU sidecar. The bottleneck then moves from detection cost to detector-selection discipline: knowing which task you are serving, what your detector can actually see, and whether anyone has checked its numbers besides its authors.
Frequently Asked Questions
What are the performance scores for the similarity plus NLI ensemble on question answering and dialogue tasks?
On the evidence available, a similarity-plus-NLI ensemble running on CPU catches question-answering hallucinations well (F1 0.792, AUC-ROC 0.873) and dialogue hallucinations moderately (F1 0.694, AUC-ROC 0.749), while every cheap detector tested performs near chance on summarisation in its default configuration.
How does sentence-level chunk aggregation affect summarisation detection accuracy and cost?
The repository states the measured outcome: “SummaC-style chunk aggregation lifts summarisation AUC-ROC from 0.567 to 0.683, still on CPU with the same model, at roughly 20× the NLI calls.”
What are the input length limits for the NLI and BERTScore detectors in the benchmark?
Per the repository: “The NLI premise is capped at 800 source characters, BERTScore references at 512 characters,” and “all-MiniLM-L6-v2 truncates at its default 256 tokens; only ROUGE-L sees the full source” (repository input limits).

Join the discussion
Share a useful perspective or ask a question about this article.