groundy
Open Source

EuroLLM vs Apertus: Choosing a Sovereign Open Model for EU Apps

EuroLLM and Apertus are distinct sovereign open models with different licenses and coverage. Cloudflare's hosting is vendor-reported intent, not shipped capability.

Published 6 references
Two handmade resin dinosaurs face each other on an ivory surface. A squat green dinosaur has three blank speech bubbles beside its mouth; a clear, curled-tail dinosaur exposes internal joints through an open belly panel.
On this page11 sections

Cloudflare’s Birthday Week post says EuroLLM and Apertus, two European open-weight model families, are coming to Workers AI with access requests open today. That availability is vendor-reported intent, not shipped capability, and every headline figure in this article is either self-reported by the model labs or restated by Cloudflare; every head-to-head comparison that exists comes from one of the two labs. What you can decide now is which family’s coverage, size, and license fit your application, and where to serve it.

What Cloudflare actually committed to

According to Cloudflare’s sovereign AI post, published today, both model families are on the way to Workers AI and access requests are open. That is a roadmap statement. The same post reports a network running in more than 335 cities across 125+ countries (Cloudflare’s post), with GPUs for AI inference in more than 230 of them.

What is not vendor-reported: the models themselves. EuroLLM and Apertus both have primary technical reports and published weights independent of Cloudflare, and Apertus also has launch coverage from Swisscom. That is what makes the choice between them decidable today, even while the hosting option is still a promise. The useful frame is to separate the two decisions this announcement bundles together: which model family to standardize on, and which hosting locus satisfies your residency obligations. The first is answerable from the evidence. The second is a contract question, not a model question.

The two families, on the evidence that exists

EuroLLM is a research consortium’s entry, not an EU institutional product. Per Cloudflare’s post, it supports 35 languages including all 24 official EU languages, and was developed with support from Horizon Europe, the European Research Council and EuroHPC by a consortium of universities and companies, including Unbabel and Naver Labs. The primary reports confirm the coverage claim at two sizes: EuroLLM-9B’s technical report describes a model trained from scratch for all 24 official EU languages plus 11 additional ones, and the EuroLLM-22B report makes the same coverage claim at the larger size.

Apertus comes from Swiss public institutions. Per Cloudflare’s post, it was developed by ETH Zurich, EPFL and the Swiss National Supercomputing Centre (CSCS) as part of the Swiss AI Initiative, with architecture, weights, training data and methods all published. Cloudflare describes the training as taking place on CSCS’s Alps supercomputer with more than 10,000 GH200 GPUs, but that figure describes the machine, not the run: the Apertus technical report states that the training runs used a dedicated vCluster of roughly 1,500 GH200 nodes (four GPUs each) of the Alps system. Swisscom’s launch coverage states the licence terms: two sizes, 8 billion and 70 billion parameters, both released under the Apache 2.0 licence, developed with due consideration to Swiss data protection law, Swiss copyright law, and the transparency obligations of the EU AI Act.

AxisEuroLLMApertus
Sizes9B, 22B8B, 70B, plus distilled Apertus-v1.1 up to 4B
Language coverageAll 24 official EU languages + 11 more (35 total)1,000+ (Swisscom) or 1,500+ (Cloudflare); sources disagree
Training tokens~4T (EuroLLM-22B)15T (Apertus-70B), 40% non-English
LicenseVerify on the model card; not stated in Cloudflare’s postApache 2.0
Training data publishedSources and filtering described; 9B report releases the EuroFilter filter and synthetic SFT set but no pre-training corpus; 22B report announces a public dataset release (EuroWeb)Yes, along with architecture, weights and methods
ProvenanceUniversity-company consortium backed by Horizon Europe, ERC, EuroHPCETH Zurich, EPFL, CSCS; trained on Alps

Two cells in that table deserve scrutiny before you rely on them.

How much weight to give the numbers

Cross-family figures do exist; the problem is that every one of them comes from one of the two labs. EuroLLM-22B’s report scores EuroLLM-9B and EuroLLM-22B directly against Apertus-8B and Apertus-70B on English and multilingual suites, reporting that EuroLLM-9B consistently improves over Apertus-8B and that EuroLLM-22B, at roughly a third of Apertus-70B’s parameters, is frequently competitive and in several settings scores higher. The same report states that EuroLLM-22B trains on approximately 4T tokens while Apertus-70B reports 15T tokens at a much larger parameter scale. That is EuroLLM’s own team characterizing a competitor. The Swiss labs run their own version of the matchup and get the opposite ordering at the smaller size: the Apertus technical report includes EuroLLM among its baselines, and the distillation paper reports multilingual instruction-tuned averages of 0.534 for Apertus-8B-Instruct-2509 against 0.480 for EuroLLM-9B-Instruct, the 8B Apertus model outscoring the 9B EuroLLM one on that suite. The numbers are likely accurate as self-reports, but token totals are not quality measurements, the benchmark suites and settings are chosen by the evaluator, and nothing in either account independently checks the other side’s figures. Treat 4T versus 15T as a statement about training investment, and the score tables as each side’s account of the matchup rather than a settled answer on which model serves Polish or Portuguese better.

The language counts conflict outright. Swisscom’s coverage says Apertus was trained on 15 trillion tokens across more than 1,000 languages; Cloudflare’s post says more than 1,500. The difference probably reflects how you count languages versus dialects or macro-languages, but neither source explains the discrepancy. More important: neither count tells you per-language quality. A model that nominally covers 1,500 languages may be weak in 1,400 of them. Coverage is a claim about what the training data included, and per-language quality is lab-measured rather than independently checked.

The third gap is EuroLLM’s licence. Swisscom’s launch coverage and Cloudflare’s post state Apertus’s Apache 2.0 terms plainly. Cloudflare’s announcement doesn’t name EuroLLM’s licence terms, so confirm them on the model card before standardizing on EuroLLM for a commercial product. This is the same discipline that applies to any open-weight procurement: verify the re-hosting terms before you need them. Groundy’s coverage of Mistral’s raise makes the adjacent point that a European vendor’s positioning establishes capital and intent, not compliance or hosting cost.

Where to serve it: three loci, one compliance surface

The model choice and the hosting choice interact, but they are not the same decision. Three options cover the realistic range.

Self-host. Both families publish weights, so this is available regardless of what Workers AI ships. Apertus-v1.1 matters here: per the distillation paper, it is a family of models up to 4B parameters built on the open-recipe Apertus 8B and trained on 1.7T permissive-license tokens, which puts a credible model within reach of modest hardware. EuroLLM-9B sits above the distilled range, though it remains far smaller than the 70B class. Self-hosting gives you the cleanest residency story, since prompts never leave infrastructure you control, and the auditability argument extends to the weights themselves: Groundy’s Soofi S analysis made the point that open weights plus published data accounting plus in-country compute is an audit you can walk a regulator through. Apertus’s published training data gives you the same walk-through potential.

Workers AI. The appeal is operational: no GPU procurement, inference close to users across a large edge network. The catch is structural. Cloudflare is a US company, and nothing in the announcement shows that serving European-sovereign weights on a US-owned edge network satisfies any specific EU residency regime. “Sovereign” describes where the model came from, not where your prompts are processed. If you go this route, residency obligations move from the model choice into the hosting contract: the data processing agreement, the sub-processor list, the regional-inference terms. Evaluate Workers AI the way you would evaluate any US cloud’s EU region, on the paperwork, not on the provenance of the weights it happens to serve. And because availability is stated intent as of today, do not write it into a shipping plan yet.

Hugging Face hosted inference. A middle path with no Cloudflare dependency and no hardware to buy. Cloudflare’s post and the Apertus launch coverage are silent on pricing and residency terms for this route, so treat it as a line item to evaluate rather than a recommendation. Its main virtue is that the weights’ presence on Hugging Face is the one availability fact you can confirm yourself, today, by downloading them.

The security harness changes what switching costs

One piece of the Cloudflare announcement survives skepticism about the rest. Per the same post, Cloudflare built a security harness, an orchestration layer that coordinates multiple AI models working in parallel to hunt for vulnerabilities, verify findings and prioritize threats, and open-sourced it so any organization can run it with the models of its choice.

The phrase “models of its choice” is the load-bearing part. A recurring hidden cost of standardizing on one model family is that your defensive tooling quietly becomes coupled to it: prompt-injection filters tuned to one tokenizer, eval suites written against one model’s failure modes, red-team harnesses wired to one API. A model-agnostic harness means the security layer is portable. If you standardize on EuroLLM now and switch to Apertus later, or run both against different workloads, the orchestration and verification tooling does not have to be rebuilt. That lowers the cost of being wrong about the model choice, which matters precisely because the quality evidence is thin. You are making a bet under uncertainty; portable defenses reduce the stakes of the bet.

The caveat: this is one vendor’s description of its own tool. The harness exists and is open-sourced, per the post, but its effectiveness with EuroLLM or Apertus specifically, as opposed to the frontier and open models Cloudflare tested with, is unverified.

Decision rules by team type

The evidence supports a conditional verdict, not a universal one.

Standardize on Apertus (8B or 70B) when license auditability and transparency dominate. Apache 2.0 is unambiguous, the training data is published, and the compliance framing (Swiss data protection, Swiss copyright, EU AI Act transparency obligations) is already documented. Public-sector platforms and regulated industries that must show their work to auditors get the most from this family. Add Apertus-v1.1 to the evaluation if serving budget is the binding constraint.

Standardize on EuroLLM (9B or 22B) when stated coverage of all 24 official EU languages at smaller serving sizes dominates. If your product must work in Maltese, Irish and Estonian on day one, a model trained from scratch for exactly that mandate is a more direct fit than a 1,000-language model whose per-language depth only its own developers have measured. Verify the licence terms on the model card first; Cloudflare’s post doesn’t state them.

Either way, treat Workers AI availability as vendor-reported intent, put residency obligations in the hosting contract rather than assuming sovereign weights confer sovereign processing, and keep the model-agnostic harness in mind as insurance against a future switch.

What would change this verdict

Three findings would rewrite this comparison. First, an independent multilingual benchmark covering both families across the EU languages would convert the choice from a criteria-ranking exercise into a quality comparison; every quality figure in this comparison was produced by one of the two labs. Second, published EuroLLM licence terms could remove the one asymmetry Apertus currently holds unopposed. Third, shipped Workers AI availability with documented regional-inference terms would make the hosting leg of this decision testable instead of theoretical. Until then, the strongest honest claim is narrower than the announcement implies: both families are real, downloadable and distinct enough that the right choice depends on whether your binding constraint is licence auditability, EU-language coverage at small sizes, or serving footprint, and the sovereignty of your deployment will be decided by the contract you sign, not the weights you download.

Frequently Asked Questions

What license does Apertus use?

Swisscom’s launch coverage states the licence terms: two sizes, 8 billion and 70 billion parameters, both released under the Apache 2.0 licence, developed with due consideration to Swiss data protection law, Swiss copyright law, and the transparency obligations of the EU AI Act.

Does Cloudflare’s post state EuroLLM’s license terms?

Cloudflare’s announcement doesn’t name EuroLLM’s licence terms, so confirm them on the model card before standardizing on EuroLLM for a commercial product.

Is Workers AI availability for these models confirmed?

According to Cloudflare’s sovereign AI post, published today, both model families are on the way to Workers AI and access requests are open. That is a roadmap statement.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Cloudflare's sovereign AI postblog.cloudflare.comAccessed
  2. EuroLLM-9B's technical reportarxiv.orgAccessed
  3. EuroLLM-22B reportarxiv.orgAccessed
  4. Apertus technical reportarxiv.orgAccessed
  5. Swisscom's launch coverageswisscom.chAccessed
  6. the distillation paperarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy