A preprint observed on 2026-10-09, Jailbreak Scaling Laws for Large Language Models: Polynomial, Exponential Crossover (arXiv:2603.11331), claims something red-team leads have wanted for years: a quantitative relationship between how many attempts an attacker gets and how often a jailbreak succeeds. The claim is self-reported, recently posted and not yet externally verified, but its shape matters for budgeting even before anyone replicates it. The authors report that attack success rate grows only slowly, polynomially, with the number of inference-time samples when the attacker submits plain harmful questions, and shifts to exponential growth once adversarial prompt injection works. In their words: “Empirically, we find that adversarial prompt-injection attacks can amplify attack success rate from the slow polynomial growth observed without injection to exponential growth with the number of inference-time samples.”
If that crossover holds in your deployment, two familiar budget decisions change. First, a jailbreak evaluation that fixes attacker effort at one setting measures a lower bound, not a safety level. Second, aggregate controls like per-user rate limits stop being sufficient precisely where the exponential regime begins, and spending has to move toward per-attempt hardening and monitoring of repeated adversarial queries. One caveat carries throughout: the exponential direction and its generality come from a single self-reported preprint, and the slope of success against effort has to be measured on your own model and harness rather than imported from its experiments.
What the crossover claim says, and what it does not
The paper offers both an empirical observation and a statistical mechanism. On the theory side, it identifies “a small set of assumptions on the distribution of safe generation across contexts under which both scaling laws follow” and proposes a generative model of proxy language in terms of a spin-glass system in a replica-symmetry-breaking regime that analytically realizes those assumptions. Read loosely, the claim is that without injection each added sample buys the attacker less than a constant increment of success; once injection steers generation toward unsafe behavior, cumulative success compounds with the number of samples.
The experimental setup deserves attention before anyone extrapolates. The authors tested two injection methods, GCG universal adversarial strings and AutoDAN stealthy prompts, against AdvBench and HarmBench harmful questions, with a no-injection baseline. The targets are four frontier models: Claude-Sonnet-4.5 (claude-sonnet-4-5-20250929), Claude-Haiku-3.5 (claude-3-5-haiku-20241022), GPT-Turbo-3.5 (gpt-3.5-turbo-012), and GPT-4 (gpt-4-0613), queried via API at temperature 0 with 512 new tokens and one sample per prompt. Note what those frontier runs can and cannot show: at one sample per prompt they support the paper’s attack and ASR-metric comparisons, not its scaling curves. The multi-sample scaling observations, the polynomial-versus-exponential evidence itself, are reported on open models including Llama-3-8B-Instruct and OLMo-2-0325-32B-Instruct. What is left open is the defended stack: whether the exponential regime survives an output filter, a system prompt, or a detection layer in front of the model is a question for your own deployment.
One detail in the paper cuts in the defender’s favor. The same authors caution that “the refusal string-based detection overestimates ASR compared to LLM-as-a-Judge”. So even within this single study, the headline success numbers depend on how you count. That matters beyond this paper: much of the published jailbreak literature uses refusal-substring matching, and cross-paper ASR comparisons are shakier than the decimal places suggest.
Why a single-setting ASR score is a lower bound
Independent of the crossover preprint, there is good evidence that jailbreak evaluation results move around with attacker budget and setup. The JBDistill benchmark paper documents that LLM-based red-teaming “develops different attack prompts for different models under inconsistent compute budgets, and small changes in its the attack setup (e.g., hyperparameters, chat templates) can lead to large variability in attack success.” A model that scores a low ASR under one red-team harness can score very differently under another with more compute or a tweaked chat template. That is not a rounding issue; it is the difference between comparable and incomparable release gates.
Combined with the scaling claim, the picture is this: a fixed-effort eval samples one point on a curve, and the preprint argues the curve’s shape itself changes once injection works. A pre-release eval that gives the attacker, say, 20 attempts per harmful question tells you about 20-attempt adversaries. It says nothing defensible about a patient adversary with a corporate API budget unless you know which growth regime you are in.
Sizing the red-team budget
The practical output for a red-team lead is not a single attempt count, because the right constants are deployment-specific: they depend on the model, the harness, and whatever defenses sit in front of the endpoint. What the evidence supports is a budgeting procedure:
- Evaluate at multiple effort levels. Run your jailbreak suite at several sample budgets per harmful question (for example, single-shot, tens, and hundreds of attempts) rather than one fixed setting. The goal is to estimate the local slope of success against effort for your model and harness, not to import the paper’s constants, which would not describe your deployment anyway.
- Test with and without injection separately. The preprint’s core claim is that the regime flips when injection works. If your red-team campaign only runs one attack family, you cannot tell whether you sit on the polynomial branch or the exponential one. Budget both a no-injection baseline and at least one strong injection method; GCG and AutoDAN are the paper’s choices and are established baselines.
- Fix the scoring metric before comparing runs. Because refusal-string scoring overestimates ASR relative to an LLM judge, pick one metric (ideally an LLM judge) and hold it constant across effort levels, or every slope you measure mixes attacker progress with metric drift.
- Record the harness. JBDistill’s variability finding means hyperparameters and chat templates belong in the eval report alongside the ASR number.
This is more expensive than a one-shot eval, and it is worth being honest about that: multiplying effort levels multiplies attacker compute. But the alternative is a release gate that certifies only against the least patient attacker you tested.
Where rate limits still work
Rate limits and per-query pricing are the oldest query-budget controls, and the economics are concrete. Classic black-box adversarial work priced attacker queries directly: after the first 2,500 predictions, the Clarifai API cost “upwards of $2.40 per 1000 queries. This makes a 1-million query attack, for example, cost $2400”. There is also a theoretical floor under this lever: query-complexity results prove lower bounds on attacker queries in terms of decision-boundary entropy, suggesting some learners are inherently harder for query-bounded adversaries. Query budgets are a real security parameter, not a compliance ritual.
The crossover claim tells you where that lever loses its force. Against the polynomial regime, which the preprint reports for repeated harmful questions without injection, raising the marginal cost of each query degrades attacker success roughly in proportion to budget. A million-query cap genuinely constrains a naive sampler. Against the exponential regime, if it holds, the same cap constrains far less: the attacker needs fewer samples per unit of success, so the defense gets less safety per dollar of attacker cost imposed. My inference from the preprint’s stated regimes, not a measured result: rate limits are a tax that scales linearly while the attacker’s return on queries scales faster. Taxes lose races against exponentials.
That reasoning produces a conditional rather than a verdict to rip out rate limiting. Keep it; it bites hardest exactly where attacks are unsophisticated, and it costs little. But do not let it substitute for controls that act per attempt or on the attacker’s search process, because those are the controls that still function if the exponential branch is real.
Controls that act on the search loop
The strongest defense-side evidence targets the attacker’s iteration, not individual prompts. ProAct feeds spurious “jailbroken” responses into iterative attackers’ optimization loops, so the attacker’s own search stops early on a false signal. Its authors report ASR reductions “by up to 94% without affecting utility,” and reductions “to below 3% across all four benchmarks” (ProAct paper) against a state-of-the-art multi-turn scheme when paired with an output filter. These are self-reported numbers on the authors’ attack and benchmark mix, not a head-to-head against other defenses. The instructive ablation: on AIR-Bench, “on average 69% of successful defences can be attributed to ProAct misleading the jailbreak evaluator into an early stop”. If success compounds with samples, then a defense that truncates the sample count attacks the exponent itself. That is the mechanistic link between this defense result and the scaling preprint.
Query-level detection is the second axis. Gradient Cuff rejects jailbreak queries using refusal-loss gradient norms at a 5% false-positive rate, and its refusal rate largely held under adaptive PAIR, TAP, and GCG attacks designed against it on LLaMA-2-7B-Chat (GCG: 0.988 refusal without the adaptive attack versus 0.986 with it). That durability is model-dependent, and the paper’s own table supplies the counter-evidence: on Vicuna-7B-V1.5, adaptive PAIR cut Gradient Cuff’s refusal rate from 0.694 to 0.356 (Gradient Cuff paper). A 5% false-positive rate is a real operational cost on high-volume traffic, but the result demonstrates that per-query tripwires can survive attackers who know the defense exists, at least on some models, which is the threat model that matters once your deployment is worth targeting. The Vicuna cell is the reason to test any detector against adaptive attacks on your own model rather than assume the LLaMA result transfers.
Per-attempt response filtering is the third axis, and its ceiling is documented. AutoDefense, a three-agent LLM filter using LLaMA-2-13B agents, cut jailbreak ASR on GPT-3.5 from 55.74% to 7.95% (AutoDefense paper). That is a large reduction, but the residual 7.95% (AutoDefense) is a per-attempt floor: if the attacker gets enough attempts, a constant per-attempt success probability still compounds. Filtering raises the attacker’s required budget; it does not cap cumulative success on its own.
Continuous-update defenses complete the picture but carry infrastructure cost. An online-learning defense against iterative jailbreaks updates a prompt-optimization model during inference with Past-Direction Gradient Damping; its experiments required 8 NVIDIA H100 GPUs. That is a research-scale footprint, not necessarily a production requirement, but it frames the tradeoff: defenses that adapt to the attacker cost ongoing compute, while static defenses add no adaptation cost, though per-response filters still pay inference overhead on every request, and they lose effectiveness against attackers who adapt.
The table below compresses how these controls map onto the two regimes the preprint describes.
| Control | Acts on | Polynomial regime (no working injection) | Exponential regime (injection works, if claim holds) | Evidence base |
|---|---|---|---|---|
| Rate limits, per-query pricing | Attacker’s aggregate query budget | Strong: cost scales linearly with attacker effort (Clarifai pricing analysis; query-complexity lower bounds) | Weakens: attacker needs fewer samples per success, so each dollar of imposed cost buys less safety | Classic adversarial ML, plus the preprint’s stated regimes |
| Per-attempt response filtering (AutoDefense) | Each response independently | Reduces per-attempt ASR (55.74% to 7.95% on GPT-3.5, self-reported by AutoDefense authors) | Residual floor compounds with attempts; insufficient alone | Single defense paper, own benchmark mix |
| Query-level detection (Gradient Cuff) | Each incoming query | Cuts attempts before they reach the model at 5% FPR (Gradient Cuff) | Same mechanism applies; adaptive-attack durability largely held on LLaMA-2-7B-Chat but fell to 0.356 refusal on Vicuna-7B-V1.5 under adaptive PAIR | Single defense paper |
| Loop disruption (ProAct) | Attacker’s optimization loop | Truncates search early; up to 94% ASR reduction claimed (ProAct) | Directly targets the sample-count term in the preprint’s law | Single defense paper, self-reported |
| Online-learning defense | Deployed guardrail over time | Adapts to iterative attacks (online-learning defense) | Same, at 8-H100 research-scale cost | Single defense paper |
None of these numbers are comparable head-to-head; each paper evaluates on its own models, attacks, and benchmarks, and all are self-reported by their authors.
Static guardrail benchmarks as release gates
The contamination evidence gives an independent reason to stop treating static jailbreak suites as gates. In experiments that leak a fraction of the benchmark into fine-tuning, static HumanEval accuracy rises as contamination increases from 25% to 100%, while the dynamic DyCodeEval benchmark holds stable scores. The safety parallel: a static harmful-question suite that has circulated long enough may be partially memorized or fitted by frontier training pipelines, and a passing score increasingly measures exposure rather than robustness. AdvBench, the suite used in the crossover preprint, dates to 2023 and is heavily studied; that does not invalidate the preprint’s internal comparisons (all conditions use the same questions), but it weakens any absolute ASR number as a production-risk estimate.
The alternative pattern is renewable evaluation. JBDistill’s framing, amortizing expensive jailbreak generation into fresh benchmark construction, and DyCodeEval’s dynamic regeneration both point the same direction: the eval should refresh faster than training pipelines can absorb it. For a release gate, that means a generated-per-release attack suite plus a multi-effort-level campaign, rather than a fixed ASR threshold on a fixed question list.
What to change now, and what to wait on
If I were setting the next release’s jailbreak budget on this evidence, I would make three changes immediately and defer one.
Change now: run the eval at multiple attacker effort levels with and without prompt injection, scored by a single held-constant LLM judge, with harness settings recorded. This costs more attacker compute and buys the slope estimate that tells you which regime your model sits in. Change now: instrument production traffic for repeated adversarial query patterns per user or key, because every effective control above, detection, loop disruption, even rate limiting, depends on seeing the attempt sequence rather than isolated requests. Change now: demote static-suite ASR thresholds from release gate to one input among several, per the contamination evidence.
Wait: do not reprice your entire defense stack on the crossover’s exponent. The claim is a single self-reported preprint, unverified externally, with multi-sample scaling curves reported on open models under two injection methods; the frontier-model runs in the same paper used a single sample per prompt. Any numeric attempt budget derived from it today would be invented precision, because the slope of success against effort is a property of your model and harness rather than a constant you can import. The defensible posture is the conditional one: rate limits stay because they are cheap and effective against the polynomial branch; per-attempt hardening, query-level detection, and loop disruption earn new budget because they are the controls that still work if the exponential branch is real; and fixed-effort eval scores get read as lower bounds until someone replicates the curve behind a defended production stack.

Join the discussion
Share a useful perspective or ask a question about this article.