A product team evaluating audio language models for non-English markets faces a measurement problem that public leaderboards do not solve: the scores that dominate vendor pages are mostly English, mostly text or image, and rarely test the thing a contact-center bot or media indexer actually does, which is listen to mixed audio and reason about it. A new preprint, EXAM2: Extending Audio Understanding in Multilingual and Multimodal Analysis, puts numbers on that gap. Its authors evaluated current open-source and proprietary audio models across six languages and mixed speech, sound, and music scenes, and report what they call “substantial performance gaps in multilingual and cross-modal understanding.”
The practical consequence: if you are procuring an audio model for deployment outside English, the evaluation burden sits with you. The useful contribution of EXAM2 is not its leaderboard, it is its structure, which you can turn into an acceptance test. But one of its headline numbers deserves skepticism up front, because the biggest reported improvement comes from a model the authors fine-tuned on the benchmark’s own training data.
The deployment gap: English scores do not transfer
The demand for non-English audio understanding is no longer speculative. Citing Luminate’s 2026 midyear report, Billboard reports that “nearly one in 10 streams is of music in Spanish” in the United States, with casual listenership of Spanish-language music reaching 54% in the first quarter, the highest on record. Bad Bunny’s DTMF, a primarily Spanish-language album, won album of the year at the 2026 Grammys, only the second primarily non-English album to take that award in its nearly 70-year history.
That is music-consumption data, not voice-assistant usage data, and it does not tell you how a model will handle a Malay-language support call. What it does establish is that audio products now operate in markets where English is not the default, and the evaluation infrastructure has not followed. The EXAM2 authors put it directly: their benchmark exists because multilingual audio understanding needed measuring across “six languages and multiple modalities, including speech, sound, music, mixed-audio settings, and visual images” (arXiv:2608.23758).
Inside EXAM2: six languages, mixed scenes, images as answers
EXAM2 covers German, Spanish, Japanese, Malay, Chinese, and English (DE, ES, JA, MS, ZH, EN). It comprises 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations, according to the paper.
The design detail that matters most for practitioners is the question format. Each instance pairs a raw audio waveform with a set of candidate images (one per answer choice), a natural-language question, candidate answers, and a gold answer. The model is not transcribing; it is listening to a scene and selecting the image that correctly answers a question about it. That is a different capability from speech-to-text accuracy. A model can transcribe a German sentence flawlessly and still fail to reason about the acoustic scene the sentence describes.
The benchmark is split in two:
| EXAM2-train | EXAM2-test | |
|---|---|---|
| Questions | 4,669 | 998 |
| Visual choices | 18,676 | 3,938 |
| Audio domains | speech, sound, music, mixed | speech, sound, music |
| Source data | MMAR and Clotho, recreated and filtered | MMAU-test-mini, restructured and annotated |
This provenance matters twice. First, it tells you the test split derives from MMAU-test-mini, an existing benchmark family, so any candidate model that trained on MMAU-family data may have seen related material. Second, as the next section shows, the train split is the source of the paper’s biggest improvement claim.
On annotation quality, the paper reports manual edit rates on machine-translated material of 1.10% (Malay), 2.30% (Spanish), 2.81% (Japanese), 3.01% (German), and 5.31% (Chinese) (arXiv:2608.23758), review by native speakers or experienced users with more than five years of advanced proficiency, and inter-annotator agreement at a Cohen’s kappa of 0.9830. These are the authors’ own QC figures, but they indicate the translations were checked rather than shipped raw.
Reading the improvement claims correctly
Here is the circularity trap, and it is the single most important thing to understand before citing this paper. The abstract reports that “Gemma3n-EXAM2, a lightweight fusion-model fine-tuned on EXAM2-train, achieves up to 15.8% improvement in multilingual settings and 16.5% gains in multimodal evaluation over a strong baseline” (arXiv:2608.23758).
Gemma3n-EXAM2 was fine-tuned on EXAM2-train, the benchmark’s own training split. Those 4,669 questions come from different source datasets than the 998-question test set (MMAR and Clotho for train, MMAU-test-mini for test), but they were built through the same construction pipeline, question format, and translation process, so a model tuned on them should be expected to improve on that held-out test. The gain partly measures fit to the benchmark’s task, not deployable superiority on audio the model has never seen. This is not an accusation of misconduct; training on a benchmark’s train split and evaluating on its held-out test split is standard practice. The trap is in how such numbers travel: stripped of their provenance, “up to 15.8% improvement” (arXiv:2608.23758) becomes a marketing line about a research checkpoint that is not a shipped product, evaluated in a way that overstates what a general deployment would see.
The paper also introduces OmniLoRA, an omni-language uniform tuning method based on LoRA, reported as effective across the authors’ experiments. Same caveat applies: these are the authors’ measurements on their own benchmark.
The generalizable rule is worth keeping after EXAM2 leaves the feed: when a benchmark’s headline improvement comes from a model fine-tuned on that benchmark’s training data, treat the number as evidence the benchmark is learnable, not as evidence of a better product.
What vendor scorecards show, and what they leave out
The contrast with vendor-published evaluations is instructive. Google’s Gemini 3.1 Pro page lists the model in Preview with text, image, video, audio, and PDF inputs. Audio is on the input list. But the published benchmark table demonstrates multimodal and multilingual ability through text and image tests: MMMU-Pro multimodal understanding at 80.5%, and MMMLU multilingual Q&A at 92.6%. No scene-mixed audio comprehension benchmark appears in the published table.
Read carefully, that 92.6% MMMLU score is exactly the number that gets misread as “multilingual is solved.” The row’s own label is “MMMLU Multilingual Q&A,” and no audio comprehension benchmark appears anywhere in the published table. EXAM2 measures whether a model can reason about audio scenes across languages, and its authors report “substantial performance gaps” there. A vendor page that lists audio as an input modality is making a capability claim, not presenting audio evaluation results.
One vendor page is an anecdote rather than a survey, but the paper’s introduction reports the same gap from the research side: existing audio benchmarks “remain primarily audio-only” and “largely rely on English-centric question-answer settings” (arXiv:2608.23758). Either way, a deployer who relies on published scorecards is comparing models on dimensions adjacent to the one that matters.
A runnable acceptance-test checklist
EXAM2’s structure converts directly into a procurement test harness. Before committing to an audio model for non-English deployment, I would build an acceptance test along these axes:
- Language coverage. EXAM2 covers DE, EN, ES, JA, MS, and ZH. Map that against your actual target markets. If you are deploying into Portuguese, Korean, or Hindi, EXAM2’s six-language results tell you nothing directly, and you need your own test set regardless.
- Scene coverage. Include speech, sound, music, and mixed audio, not clean read speech only. Contact-center audio is speech over hold music and background noise; media indexing is music, effects, and dialogue layered together. EXAM2’s train split covers all four domains; its test split covers three.
- Comprehension, not transcription. Include questions where the answer requires reasoning about the audio scene, with image or structured answer choices, not just word-error-rate checks on transcripts.
- Score provenance. For every number a vendor gives you, ask whether the evaluation was independent, self-reported, or produced by a model variant fine-tuned on the benchmark’s own training data. Discount the third category heavily.
- Contamination check. EXAM2-test derives from MMAU-test-mini, and EXAM2-train from MMAR and Clotho. Check whether candidate models trained on those datasets before comparing scores. The authors state that all code, data, and models are available in the repository the paper links, so you can inspect the released splits yourself rather than take the paper’s word for their construction.
- Translation quality. If you build your own multilingual test material, budget for human review. EXAM2’s reported edit rates (1.10% to 5.31% across languages, kappa 0.9830 agreement) are a reasonable reference point for what checked translations look like.
The first and last items are where teams tend to cut corners, because translation review is slow and the target-language list is where scope creep lives. Both are also where a procurement decision quietly fails.
Caveats and what to verify yourself
This article rests on a single preprint with the authors’ own measurements, and it treats those figures as one lab’s findings until independent reproduction appears. The full results are in the paper and are worth reading directly. Table 3 of the paper reports GPT-4o-audio as the strongest closed-source model at a 73.83% overall average across speech, sound, and music in six languages, with Gemma3n-EXAM2 the best open-source model at 61.24% (paper); Table 4 breaks those averages down by language and domain. One pattern in that breakdown matters for deployment planning: the authors report that western languages (German, English, Spanish) generally score higher than Japanese and Chinese across most model families, with the gap most pronounced on sound and music understanding, where degradation in Japanese and Chinese remains substantial even among frontier proprietary models. If your target market is Japanese or Chinese, assume the low end of any reported range and test accordingly.
Two further limits matter for deployment decisions. Multiple-choice comprehension accuracy does not measure production metrics like transcription accuracy, latency, or dialogue-state handling; the authors write that they “have prioritized accuracy over efficiency” and leave inference speed and deployment constraints to future evaluations (arXiv:2608.23758), so EXAM2-style testing supplements rather than replaces your operational evaluation. And the market evidence cited here is streaming data, which establishes demand for non-English audio products without saying anything about assistant usage patterns in any specific language.
The durable takeaway survives all of these caveats: English benchmark scores do not transfer to non-English audio deployment, vendor scorecards currently do not measure scene-mixed multilingual audio comprehension, and the strongest reported gains on EXAM2, which its authors describe as “to our best knowledge, the first to address multilingual and multimodal audio understanding,” come from a model tuned on its own training data. The evaluation you need is the one you build.
Frequently Asked Questions
Which languages does the EXAM2 benchmark cover?
EXAM2 covers German, Spanish, Japanese, Malay, Chinese, and English (DE, ES, JA, MS, ZH, EN).

Join the discussion
Share a useful perspective or ask a question about this article.