A domain fine-tune that passes its own eval can still be a regression. The most useful finding in the current literature is a simple measurement: a standard supervised fine-tune on domain data costs 3.4–14.5% on general-capability benchmarks, according to a new preprint proposing critical-point routing, and the same preprint reports a mitigation that recovers that drop to at most 0.5%. Every number in that sentence, and in this article, is a reported claim from an unreplicated preprint, not an established result. But the pattern they describe matches a failure many ML engineers have watched happen silently: you tune Llama, Qwen, or GLM on proprietary data, the domain eval goes up, and nobody checks whether the instruction-following and reasoning that justified picking the base model are still there.
The practical answer to the title’s question, then, has two parts. First, gate every adapter on a before/after regression suite covering instruction-following, reasoning, and long-context behavior. Second, pick a mitigation whose overhead envelope fits your budget: replay-class methods at a reported 3–5% wall-clock overhead per the MSSR replay preprint when you can spare replay data, constrained low-rank tuning when domain data is scarce, and decoupled expert routing only when near-zero general-capability loss justifies holding two models in GPU memory. What follows is the procedure, the comparison, and where the evidence runs out.
The silent failure: what a 3.4–14.5% general-capability drop actually erases
The number that anchors this article comes from the critical-point routing preprint (CPR, arXiv:2608.30158), which surfaced on arXiv and measures what plain domain SFT does to a base model’s general skills: a 3.4–14.5% drop across general-capability benchmarks for the SFT expert, per the same paper. The range is wide because the drop varies by setting across the paper’s backbones and domains, so treat it as an order of magnitude, not a precise forecast for your stack.
Why does a drop of that size matter? Because the general capabilities are usually the reason the base model was chosen. A team fine-tuning an open-weight model on legal or clinical text is not buying a blank model with domain knowledge; it is buying a model that follows instructions reliably, reasons through multi-step problems, and handles long documents, plus domain knowledge. If tuning quietly erodes the first three, the team has paid a training run and a serving migration for a model that is worse at the things users notice first. And because standard evals are domain-scoped, the whole point of the run, nobody sees the erosion until a user does.
The CPR paper frames prior methods as compressing both domain and general capabilities into a single weight set, bound to a domain-general trade-off that, citing Lin et al. 2026, no existing method has fully eliminated. That framing is the useful part: it tells you the trade-off is structural, not a hyperparameter bug you can tune away. The response is to measure it, then pick how much of it you are willing to accept and what you will pay to reduce it.
Step zero: build the regression suite before you train
The single most effective move costs no training compute at all: run a general-capability eval on the base model before fine-tuning, and run the same eval on the tuned artifact before shipping. The suite should cover at least three skill areas, because those are the ones the forgetting literature says die first. The MSSR replay paper reports that in continual fine-tuning across multiple backbones (Qwen2.5-7B, LLaMA-3.1-8B, Gemma2-9B), the largest forgetting concentrates on long-context and reasoning benchmarks, “early-task forgetting is most severe” there. So a regression suite that only checks instruction-following on short prompts will miss the damage where it is worst.
A workable procedure:
- Freeze a base-model scorecard. Before any tuning, evaluate the base model on an instruction-following benchmark, a reasoning benchmark, and a long-context benchmark, plus any safety behavior your product depends on. Record exact versions, prompts, and seeds.
- Set a ship gate, not a vibe. Decide in advance what drop is acceptable, for example, no more than one or two points on any general benchmark, given what you know about measurement noise on your harness. The “at most 0.5%” residual drop in the CPR paper is a reported claim from one setup, not a universal target; your gate should reflect your tolerance.
- Re-run after every candidate adapter. Domain gain and general retention get reported side by side. An adapter that wins on domain and loses on reasoning is not a pass; it is a trade that needs an explicit decision.
- Keep the harness stable. The value of the suite is comparability across runs and mitigations. Changing benchmarks mid-project resets the baseline and hides drift.
This is deliberately plain. None of the four papers discussed here packages exactly this procedure; it is an inference from their shared result that forgetting is real, measurable, and concentrated in predictable skill areas. The papers supply the “what to test” and the mitigations; the discipline of testing before shipping is yours.
Replay: the strongest reported retention per unit of compute, at a data cost
Replay is the oldest idea in continual learning: mix samples of the old distribution into the new training data so the model cannot fully drift. The current refinement on the LLM side is MSSR, memory-aware adaptive replay, described in its preprint. The authors report that MSSR “delivers stable gains over fixed, loss-based, and accuracy-based replay on multiple backbones (Qwen2.5-7B, LLaMA-3.1-8B, Gemma2-9B)” across both 3-task and 11-task continual-learning settings, with the largest improvements on long-context and reasoning benchmarks, exactly the skills the regression suite should gate.
The cost envelope is what makes replay worth considering first. The MSSR authors report overhead of “3–5% wall-clock, 4–6% peak memory” relative to a fixed-replay baseline, with only scalar per-sample updates and no additional forward or backward passes. If those figures hold up outside the paper, replay is the cheapest mitigation on the training-compute axis by a wide margin.
The real cost sits elsewhere: replay requires replay data. You need a buffer of general-domain samples that approximates what the base model learned, and you need the rights and the pipeline to mix it in. For a proprietary deployment, that is sometimes easy (public instruction data is abundant) and sometimes awkward (if your tuning pipeline is tightly scoped to licensed domain data, adding a second data stream is real work). The other caveat: MSSR’s evidence comes from continual-learning settings with task sequences, which may differ from single-shot domain adaptation. The mechanism transfers; the exact numbers may not.
Constrained low-rank tuning: the limited-data option, with a hidden failure mode
LoRA is the default fine-tuning path for most teams, and with it comes a comforting folk belief: because you only train a small low-rank adapter, the base model’s behavior is mostly preserved. The evidence says that belief is only partly true. CURLoRA, described in its preprint, exists precisely because vanilla LoRA does drift. Its authors report that CURLoRA “achieves very good and stable task accuracy while maintaining base model’s perplexity scores fixed compared to LoRA upon continual fine-tuning, particularly in scenarios with limited data”, the implication being that LoRA lets that perplexity move. Perplexity drift is a proxy for general-behavior drift, not a direct measure of reasoning or instruction-following, so treat it as a warning light rather than a diagnosis.
CURLoRA’s niche is the limited-data regime. When your domain corpus is small, full fine-tuning overfits quickly and replay buffers can dominate the batch; a stabilized low-rank method with far fewer trainable parameters is a reasonable default. Its reported advantage is stability across tasks with significantly reduced parameter counts.
The deeper problem is that parameter-space constraints do not certify output-space preservation. The function-space analysis of continual adaptation finds that forgetting “concentrates in a small number of old-task NTK eigenmodes” and, in the authors’ words, their results “explain why parameter-space regularizers can miss output-space interference, and motivate a targeted spectral regularizer.” In plain terms: you can restrict how much the weights move and still shift what the model outputs on old tasks, because the damage concentrates in specific functional directions that a parameter-norm penalty does not see. Low trainable-parameter counts are a budget feature, not a preservation guarantee. That is a mechanism-level reason the regression suite in step zero is not optional, no matter how constrained your tuner is.
Critical-point routing: decouple the expert, pay in memory and latency
CPR, the preprint that prompted this article, takes a different architectural stance. Instead of compressing domain and general skills into one weight set, it keeps the base model frozen and trains a separate domain expert, then routes tokens between them at “critical points.” The authors report in the CPR abstract that “CPR achieves state-of-the-art performance across all settings, surpassing SFT expert by 1.4-5.5% in domain performance while recovering its general-capability drop from 3.4-14.5% to at most 0.5%” (the same abstract), with minimal overhead from invoking the expert on only one-third of tokens.
Read the last clause carefully, because it invites a serving-cost misconception. In the CPR paper, the expert firing on roughly 30% of tokens does not mean 30% memory or 30% latency. The paper’s own limitations state that CPR “requires holding both the base model and the domain expert simultaneously, increasing GPU memory overhead relative to single-model approaches,” and the base model must run on every token to produce the hidden states that feed the router. So serving memory is effectively dual-model, and per-step latency stays above single-model inference even with the selective expert invocation. For a deployment already tight on VRAM, that can be the disqualifier; for a batch-oriented internal tool with headroom, it may be fine.
The bigger caveat is evidentiary. This is a single preprint with no independent replication as of October 2026. Its headline numbers come from two open-weight backbones on math and medical data, so the 1.4–5.5% domain gain and the at-most-0.5% residual drop are, per the paper, one group’s measurements on one setup rather than a cross-checked result. “State-of-the-art” here means “best in our own experiments.” That is worth testing on your stack if the dual-model cost is acceptable to you; it is not a settled ranking.
Where forgetting lives, and why it matters for your choice
The function-space paper (arXiv:2606.18024) supplies the theoretical thread connecting these methods. Its finding that forgetting concentrates in a small number of old-task NTK eigenmodes, with a Kronecker scaling rule for the vulnerable rank under frozen linear heads, does two practical things. It explains why some mitigations work: replay and routing both protect old-task function directly rather than constraining parameter movement. And it predicts a failure mode: any mitigation whose mechanism is “keep the weights close to where they started” can still allow targeted functional damage, because closeness in parameter space and closeness in output space are different metrics.
The authors use this to motivate a targeted spectral regularizer. If that line of work holds up, the next generation of constrained tuning methods will regularize the vulnerable eigenmodes directly instead of penalizing all parameter movement equally. For a practitioner today, the actionable takeaway is narrower: when evaluating any mitigation, including ones not covered here, ask whether its mechanism protects outputs on old tasks or only restricts weight movement. The first is evidence-relevant; the second is a proxy that the cited analysis shows can fail.
Side by side
The table below compares the three mitigation families on the decision axes a shipping team actually faces. Every figure is self-reported by the respective preprint, measured on different backbones and settings, and no source runs these methods head-to-head at matched budgets. Use the table to match overhead envelopes to your constraints, not to rank winners.
| Axis | Replay (MSSR-class) | Constrained low-rank (CURLoRA-class) | Critical-point routing (CPR) |
|---|---|---|---|
| Reported general-skill retention | Stable gains over other replay strategies; largest on long-context and reasoning | Base perplexity held fixed vs LoRA in limited-data continual tuning | Drop recovered from 3.4–14.5% to at most 0.5% |
| Reported domain gain | Gains over fixed/loss/accuracy replay baselines | Stable task accuracy reported | 1.4–5.5% over SFT expert |
| Training overhead | Reported 3–5% wall-clock, 4–6% peak memory; no extra forward/backward passes | Fewer trainable parameters; compute roughly LoRA-class | Router plus expert training (excerpts do not quantify) |
| Serving footprint | Single weight set | Single weight set plus small adapter | Base model plus resident expert; dual-model GPU memory |
| Data requirement | Needs a replay buffer approximating general data | Strongest reported fit for limited-data regimes | Domain data for the expert; general skills live in the frozen base |
| Main failure mode | Replay buffer quality and data-pipeline cost | Parameter-space constraints can miss output-space interference | Unreplicated; residual latency above single-model inference |
Fine-tune, RAG, or long context: when skipping the adapter wins
Everything above assumes fine-tuning is worth doing. Sometimes it is not, and the forgetting literature sharpens that decision. If the domain value you need is factual recall over a corpus that changes, documentation, policies, case files, retrieval-augmented generation keeps the base model’s general skills fully intact by construction, because the weights never move. If the domain material fits in the context window, long-context prompting achieves the same. Fine-tuning earns its cost when you need behavior change: a house style, a task format, a domain reasoning pattern, or latency and token costs that retrieval cannot match at your traffic.
No fetched source quantifies the RAG-versus-fine-tune economics, so this is decision logic rather than cited numbers, and the boundaries move with retrieval quality and context-window pricing. But the forgetting evidence changes the calculus in one specific way: fine-tuning’s true cost now includes the regression suite, the mitigation overhead, and the residual risk that some general skill degrades past your gate anyway. When that full cost is on the table, the “just RAG it” option wins more often than fine-tuning advocates tend to assume. If your domain need is knowledge, retrieve. If it is behavior, fine-tune, and budget for verification.
The open question the evidence cannot settle
Which mitigation wins at which budget is genuinely unresolved. The four papers here use different backbones, different task sequences, and different baselines; CPR evaluates two backbones on two domains. Nothing in this evidence runs replay against constrained tuning against routing at matched compute on the same model and domain. Anyone who tells you one method is simply best is extrapolating past the literature, and so would I if I ranked them beyond their overhead envelopes.
What I would do, given this evidence: stand up the pre/post regression suite first, because it is cheap and it makes every later decision measurable. Default to replay if you can source a reasonable buffer, since its reported overhead is a few percent of wall-clock and its retention gains land exactly where forgetting bites hardest. Reach for CURLoRA-class constrained tuning when domain data is scarce, with the explicit understanding that low parameter counts do not certify preserved behavior, the suite is still the gate. And trial CPR only if near-zero general-capability loss is worth dual-model serving memory to you, treating its headline figures as hypotheses to reproduce on your own model, your own data, and your own eval harness before you ship anything based on them. The base model’s general skills are part of what you bought. Measure them like it.
Frequently Asked Questions
What is the reported overhead of using replay for fine-tuning?
The MSSR authors report overhead of “3–5% wall-clock, 4–6% peak memory” relative to a fixed-replay baseline, with only scalar per-sample updates and no additional forward or backward passes.
Does critical-point routing reduce serving memory usage?
The paper’s own limitations state that CPR “requires holding both the base model and the domain expert simultaneously, increasing GPU memory overhead relative to single-model approaches,” and the base model must run on every token to produce the hidden states that feed the router.
When should a team choose RAG over fine-tuning?
If the domain value you need is factual recall over a corpus that changes, documentation, policies, case files, retrieval-augmented generation keeps the base model’s general skills fully intact by construction, because the weights never move.

Join the discussion
Share a useful perspective or ask a question about this article.