When you self-host an open-weight LLM, the safety layer is yours to build, and the default fix for it, safety-focused preference tuning, has a documented failure mode: post-trained open-weight variants “frequently suffer from over-refusal and degraded general quality,” as the authors of the Suan preprint put it. The practical answer to the title’s question is a release decision built on three paired numbers, not one: a safety benchmark score, an over-refusal rate measured on benign-but-sensitive prompts, and a capability-retention check on general benchmarks. All three are benchmark metrics; none of the cited studies measure production-stack behavior, and that distinction matters, because one finding below shows safeguarded open weights losing much of their protection under simple attacks.
That framing exists now because of a specific paper. Suan, a workshop preprint accepted at the Trustworthy AI for Good workshop at NeurIPS 2026, names over-refusal as the default cost of safety-tuning open weights and claims a fix. Its result is one self-reported evidence point. The more durable contribution is the measurement discipline the wider literature already supports: over-refusal is quantifiable, and it is manufactured by the way teams optimize for safety. None of these papers says anything about whether teams measure it before release; treating it as a first-class release metric is this article’s recommendation, not a documented norm.
Why safety-tuned open weights over-refuse
The causal trap is documented, not speculative. The OR-Bench paper, which benchmarked 32 models from 8 families including black-box and open-source systems, states it directly: “All these benchmarks are designed to evaluate safety of LLMs, so purely optimizing the safety scores within these benchmarks may inadvertently result in over-refusal models.” A model that learns to refuse anything resembling a sensitive topic scores well on the safety suite you trained against. It also refuses your users’ legitimate requests.
The cost shows up where the boundary matters most. Health-ORSC-Bench built 31,920 benign boundary prompts across seven health categories (self-harm and medical misinformation among them), varying intent ambiguity, and evaluated 30 models including GPT-5 and Claude-4. Its finding, per the same paper’s abstract: safety-optimised models “frequently refuse up to 80% of “Hard” benign prompts.” The scope matters: the tier definitions make “Hard” (Hard-1K) the most refusal-prone of the benchmark’s three difficulty levels, so that figure is a ceiling on the hardest boundary cases, not a rate across all 31,920 prompts.
For baseline calibration, SORRY-Bench found 27 of 56 evaluated LLMs sitting in a medium fulfillment band of 20%–50% on borderline prompts, with GPT-4o at 30% and Llama-3-70b at 35%. The same study measured GPT-4o’s fulfillment rate (0.30) well above GPT-3.5-turbo 1106’s (0.11), so refusal strictness swings widely even across generations of one vendor’s line. If you gate releases on safety score alone, you have no idea where on that spread your tuned model landed until a support ticket tells you.
What Suan claims, and what one workshop preprint can establish
Suan proposes a rectified variant of direct preference optimization (DPO), the standard recipe where a model is tuned on pairs of preferred and rejected responses. The authors’ claim is stated in superlatives: “Through extensive evaluations, we demonstrate that Suan achieves state-of-the-art performance, simultaneously enhancing safety and reducing over-refusal without degrading response quality.” The evaluation covered “eight diverse language models” across “a broad spectrum of safety and capability benchmarks.”
Three qualifiers belong next to that claim whenever you repeat it. First, it is self-reported: a single workshop preprint with no independent replication and no production deployment evidence as of 2026-10-06. Second, the paper is specific about its setup: eight model families (Mistral-12B, Falcon3-7B, Llama-3.1-8B, Gemma-2-9B, Qwen-3, Yi-1.5-9B, DeepSeek-7B, OLMo-3-7B), SFT on Alpaca, preference optimization on PKU-SafeRLHF, and a benchmark suite spanning Malicious Instruct, HarmBench, AdvBench and SORRY-Bench for safety, XS-Test and OR-Bench for over-refusal, and AlpacaEval, MT-Bench and ArenaHard for quality, plus ARC and MMLU. What the prose does not give you is the per-benchmark deltas: those sit across Tables C.2, C.6 and Figure 2 of the paper, with the safety results in Table C.3 and the over-refusal results in Table C.4, so any specific number attributed to Suan needs to be read off those tables before it enters a design doc. Third, “state-of-the-art” on a benchmark suite is a claim about that suite. Nothing in the preprint’s scope addresses what happens to the alignment under adversarial pressure or later fine-tuning, and the evidence on those two questions, covered below, is sobering.
None of this makes Suan uninteresting. A method that explicitly optimizes the safety/over-refusal tradeoff rather than safety alone is aimed at the right target, and it joins a small family of DPO variants with the same ambition. It makes Suan a candidate to trial, with per-benchmark verification, not a verdict to adopt.
Capability retention: the parity test DPO variants have to pass
The fear behind “safety tuning degrades quality” is measurable, and at least one method family shows it is not inevitable. On the GSM8K benchmark using Qwen2-7B-Instruct, B-DPO, a balanced DPO approach, reports that “B-DPO (72.71%) maintains performance nearly identical to the base model (71.19%)” (B-DPO paper). A math-reasoning score that survives safety tuning is exactly the kind of paired evidence you want: the same paper reports safety gains while a general-capability benchmark holds at parity.
One caveat keeps this honest: GSM8K parity on Qwen2-7B-Instruct is one model and one benchmark. It establishes that capability loss is not a law of physics for DPO-style alignment. It does not establish that your model family, your capability suite, or your data mix will come through intact. The operational pattern is what transfers: never accept a safety delta without a capability delta measured on the same tuned artifact.
Over-refusal is a product-trust cost, so price it like one
The benchmark literature gives you three complementary instruments, each measuring a different face of the problem:
- OR-Bench: over-refusal on benign prompts that merely look toxic, across 32 models from 8 families. Use it to measure whether your tuning manufactured refusals.
- SORRY-Bench: fulfillment rates on borderline requests, giving you a landscape baseline (that 20%–50% medium band across 27 of 56 models) to position your release against.
- Health-ORSC-Bench: domain-specific boundary prompts at varying intent ambiguity, the closest instrument to “will our users in a sensitive vertical hit walls.”
Why does this deserve first-class metric status rather than a checkbox? Because refusal is user-visible product behavior. A false refusal in a health-adjacent support tool, a coding assistant that declines a security question, or a document pipeline that balks at legal text all read as product defects, and the Health-ORSC results suggest safety-optimised models are exactly where these failures concentrate. Teams that measure only harm benchmarks ship this defect blind. OR-Bench’s warning explains why: the optimization target itself manufactures the failure.
Train-time alignment or an inference-time guardrail?
This is the decision the title implies, and none of the cited studies provides head-to-head data comparing train-time alignment against an inference-time guardrail classifier. There are no guardrail false-positive rates and no latency numbers in any of them. So the honest version of the comparison is about ownership, evaluation burden, and erosion risk, not a measured ranking.
| Axis | Train-time alignment (Suan, B-DPO class) | Inference-time guardrail (external classifier) |
|---|---|---|
| Over-refusal risk | Tuning against safety scores alone is documented to manufacture over-refusal (OR-Bench); rectified variants claim to reduce it (Suan, self-reported) | The model itself may refuse nothing by design; all refusal behavior lives in the checker, an architecture the OVERT benchmark documents for SD-3.5-Large in the text-to-image domain (an analogy, not LLM production evidence) |
| Capability retention | B-DPO held GSM8K at 72.71% vs 71.19% base on Qwen2-7B-Instruct; parity is achievable but must be re-proven per model | Base model capabilities untouched by alignment; the classifier adds its own error modes, which no cited study quantifies |
| Pipeline ownership | You own a training pipeline and every regression it introduces; benign fine-tuning is documented to erode safety (DataShield) | You own a serving component and its coverage; the model weights carry no safety to erode, but also none to rely on if the checker misses |
| Post-deployment robustness | Abliteration and prefilling attacks raised attack success rates from below 10% to 16%–96% against safeguarded open-weight models across BeaverTails, HarmBench, and AdvBench (arXiv:2605.26526) | No cited study measures attack robustness for guardrail stacks; treat as unknown, not safe |
| Evaluation burden | Safety + over-refusal + capability suites per tuning run, repeated after every downstream fine-tune | Same suites for the classifier’s accept/reject behavior, plus the base model’s unmitigated outputs |
The ownership row deserves a moment. Choosing train-time alignment means your security team now owns a training pipeline: preference data, tuning runs, regression suites, and re-validation every time someone fine-tunes the model for a product feature. That is real headcount and real process, and it changes the economics of self-hosting versus renting a vendor API where the safety layer is someone else’s problem. Choosing the guardrail route concentrates the work in a serving-layer component. The OVERT observation about SD-3.5-Large shows what that architecture looks like in the adjacent text-to-image world: an open-sourced model “without integrated safety alignment… does not reject any input by design and its safety mechanism depends solely on an output safety checker.” The design pattern transfers; the production evidence does not, and nothing here tells you that checker’s false-positive rate on your traffic.
If your team cannot staff a training pipeline with regression discipline, the guardrail route fails more gracefully: a checker you can swap and re-test beats a tuned-in behavior you cannot locate. If you can staff it, train-time alignment gives you refusal behavior shaped to your actual boundary cases rather than a generic classifier’s. Either way, I would not sign off on either without the three-number gate below, because both layers have documented erosion paths.
Whatever you choose, assume it erodes
The adversarial evidence is the strongest counterweight to every alignment claim in this article. Across BeaverTails, HarmBench, and AdvBench, simple abliteration and prefilling attacks “increase attack success rates against safeguarded open-weight models from below 10% to a range of 16%–96%,” per the attack study. Read that alongside any “state-of-the-art safety” claim: benchmark safety scores were measured on the model as released, and an attacker with weight access or a prefilled prompt operates on a different model in practice.
The second erosion path needs no attacker at all. DataShield found that “benign fine-tuning does not impair the LLM’s perception of harmfulness; rather, it increases its overall response compliance, which primarily accounts for the observed safety degradation.” The mechanism is almost insulting: ordinary product fine-tuning teaches the model to comply more, and compliance is the problem. A third study on why guardrails collapse after fine-tuning adds a predictive factor: representation similarity between upstream alignment data and downstream fine-tuning tasks is “a critical yet previously overlooked factor in the erosion of LLMs’ safety guardrails.”
Together these reframe the release gate. Safety is not a property you tune in once; it is a regression surface you re-measure after every downstream fine-tune, and the fine-tuning team probably does not know that. If your release process treats safety sign-off as a one-time event, the DataShield mechanism says you are shipping silent degradation on a schedule set by whoever fine-tunes next.
The three-number gate and what to verify before you cite Suan
The discipline that survives every caveat in this article:
- Safety score on the harm benchmarks your threat model cares about, gated as usual.
- Over-refusal rate on benign boundary suites in the OR-Bench and SORRY-Bench class, plus a domain suite like Health-ORSC-Bench if you serve a sensitive vertical, gated as a regression, with a budget (for example, no release that moves the benign-refusal rate by more than an agreed margin over the base model).
- Capability retention on a general suite, using the B-DPO pattern: the tuned artifact must hold parity with the same base model on benchmarks like GSM8K, measured on the same run.
Then instrument over-refusal in production: log refusals, sample and label them, and track the benign-refusal rate per release. What the cited studies document is the mechanism: optimizing safety scores manufactures refusals. None of them documents how teams detect the problem once it ships. If over-refusal is not in your dashboards, the first signal will be a support ticket, and every measurement above is cheap compared to that discovery path.
One verification task before Suan’s numbers enter your internal documents: the models, datasets and benchmark suite are named in the paper’s prose, but the per-benchmark deltas sit across Tables C.2, C.6 (safety in Table C.3, over-refusal in Table C.4), so pull the specific figures from the full paper. The “state-of-the-art” claim remains self-reported with no independent replication as of 2026-10-06. If that check passes, Suan and B-DPO are credible candidates for a train-time trial on work that can absorb a rollback. The gate, not any single method, is what keeps your next safety release from trading one trust problem for another.
Frequently Asked Questions
What three metrics should gate a safety release for open-weight LLMs?
The practical answer to the title’s question is a release decision built on three paired numbers, not one: a safety benchmark score, an over-refusal rate measured on benign-but-sensitive prompts, and a capability-retention check on general benchmarks.
How much do safety-optimized models refuse benign prompts?
Its finding, per the same paper’s abstract: safety-optimised models “frequently refuse up to 80% of “Hard” benign prompts.” The scope matters: the tier definitions make “Hard” (Hard-1K) the most refusal-prone of the benchmark’s three difficulty levels, so that figure is a ceiling on the hardest boundary cases, not a rate across all 31,920 prompts.
Does benign fine-tuning degrade safety alignment?
DataShield found that “benign fine-tuning does not impair the LLM’s perception of harmfulness; rather, it increases its overall response compliance, which primarily accounts for the observed safety degradation.”

Join the discussion
Share a useful perspective or ask a question about this article.