Qwen-Audio-3.1-Realtime, described in an Alibaba-authored arXiv report, claims a task-success gain from 78.4% to 82.0% on an agentic voice benchmark, but every number in it is vendor-reported and unverified until independently replicated. The practical consequence for developers weighing Qwen’s vendor-run realtime line against hosted realtime APIs: neither side of that comparison is settled by the available evidence, so the deciding work is a pre-launch test of action completion, not a reading of headline scores.
The claim, read line by line
The report, Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction, frames itself around exactly the problem production teams hit: voice demos that sound fine and then fail the moment the agent has to do something. Its central quantitative claim is that version 3.1 “raises overall task success from 78.4% to 82.0%” over Qwen-Audio-3.0-Realtime, measured on what the authors describe as a “half-duplex speech-to-text adaptation of τ-Voice.”
Two parts of that sentence deserve attention before the number does. First, the benchmark is an adaptation: τ-Voice was not natively run in a full-duplex realtime streaming setting, so the result says nothing directly about barge-in behavior, overlapping speech, or mid-utterance interruption handling. Second, the comparison point is the vendor’s own previous model. A 3.6-point gain over your own predecessor, on your own adapted benchmark, is a training-progress claim, not a market-position claim.
The timing matters for context. Alibaba released the open-weights Qwen3.5 and the proprietary Qwen3.5-Plus in February 2026, and this realtime line extends that cadence into audio. Whether Qwen-Audio-3.1-Realtime itself ships as downloadable weights is not stated in the report, whose release-scope note says only that the authors “compare the released Qwen-Audio-3.0-Realtime and Qwen-Audio-3.1-Realtime systems” (report). Cadence is not evidence of reliability either, and Groundy has been burned by this pattern before: with Qwen-Agent’s advertised 1M context, the vendor’s own retrievable benchmark material never actually tested accuracy at the advertised window. The healthy default with vendor-run agentic evaluations is to read the evaluation before the score.
Thirteen sessions that did not count
The most revealing passage in the report is not the headline number. It is a coverage note attached to a different benchmark, EVA-A, the Accuracy composite of EVA-Bench. There the report says the 3.1 result “covers 200 of 213 sessions”; the remaining sessions “were excluded after inference failures or repeated scoring timeouts.” Under that configuration, 3.1 obtains 47.74% Pass and 66.26% Mean, against 43.10% and 70.50% for 3.0 (report), a comparison the authors themselves call descriptive because identical session coverage is not established.
That is 6.1% of EVA-A’s sessions removed from the denominator (report), for reasons that are themselves reliability failures. Inference failures and scoring timeouts are precisely the events a reliability-focused evaluation exists to surface. If those 13 sessions had been scored as failures, the EVA-A Pass rate would land below the reported figure. The report does not state what the score would be with them included, and that question should be the first one asked of any vendor benchmark claiming reliability.
A separation worth making explicit: the exclusion note belongs to EVA-A. The τ-Voice comparison behind the headline task-success figure is presented without any coverage caveat. That limits how far the caveat travels, but it does not remove it.
This is not an accusation of manipulation; exclusions for harness failures are common and sometimes legitimate. It is a reminder that a reliability claim computed after removing the least reliable runs measures something narrower than reliability. A pattern worth remembering from Groundy’s earlier analysis: Qwen-Image-3.0’s entire claimed existence could not be confirmed from primary sources at the time. Here the model and report are real and on arXiv, which is further than that claim ever got, but the verification posture should be the same until independent results exist.
The GPT-Realtime-2 head-to-heads are all safety metrics, and most favor Qwen
Every comparison the report makes against GPT-Realtime-2 sits in its safety and reliability evaluations, and no task-success figure for it appears in the report’s text. The report does not state how GPT-Realtime-2 is served or by whom, noting only that it is “evaluated with the low-effort setting” (report):
| Axis | Qwen-Audio-3.1-Realtime | GPT-Realtime-2 | Source of figure |
|---|---|---|---|
| Task success (agentic voice) | 82.0%, up from 78.4% on 3.0 | No task-success figure appears in the report’s text | Vendor’s own half-duplex τ-Voice adaptation |
| EVA-A coverage | 200 of 213 sessions; 13 excluded for inference failures or scoring timeouts (EVA-A only, not the τ-Voice figure) | Not applicable | Vendor’s own report |
| S2T safety suite | Leads six of the seven displayed metrics | Highest TruthfulQA score, 83.92% | Vendor-run evaluation |
| Fine-grained safety categories | Higher in 9 of 10 risk categories; adult content lower by 0.90 points | Lower in 9 of 10 | Vendor-run evaluation |
| Session-level safety pass rate | 92.00% | 96.00% | Vendor-run human red-team comparison |
| Duplex setting tested | Half-duplex speech-to-text adaptation | Not stated in the report | Vendor’s own report |
| Who ran it | Vendor throughout; no third-party run accompanies these results in the report | Appears only in the vendor’s safety evaluations | Report’s own scope |
According to the Qwen report, “Qwen-Audio-3.1-Realtime achieves a session-level safety pass rate of 92.00%, compared with 96.00% for GPT-Realtime-2.” GPT-Realtime-2 also holds the highest TruthfulQA score at 83.92% (report). Those are the two comparisons GPT-Realtime-2 wins. The report’s own summaries of the same evaluations run the other way: “3.1 leads six of the seven metrics” across the displayed systems, and across the ten fine-grained risk categories “Qwen-Audio-3.1-Realtime scores above GPT-Realtime-2 in nine of the ten,” with adult content “the only category in which its score is lower, by 0.90 points” (report). Credit to the authors for publishing comparisons they lose; some vendor reports omit those. The evidenced cross-stack picture is mixed and confined to safety: the report contains no function-calling comparison against GPT-Realtime-2 at all.
No public function-calling measurement of a hosted realtime API accompanies these results. On the hosted-API side, function-calling success rates, latency under tool-call load, and behavior when a tool errors go unmeasured. Any claim about how GPT-Realtime-2 handles tool calls in production would go beyond what the report measured. The honest narrowing: this article can compare what the vendor measured, and it cannot rank the stacks.
What one community repo shows, and what it does not
The only hosted-realtime artifact here that is not a vendor report is a community project, an Inworld Realtime AI voice-agent desktop companion with barge-in support whose configuration default is INWORLD_MODEL=inworld/models/deepseek-v4-flash.
That tells you something real but narrow: hosted realtime APIs are configurable enough that a hobby-scale project can swap the underlying LLM, and at least one community builder defaults to a DeepSeek model behind Inworld’s realtime layer. It says nothing about function-calling reliability, because the repository documents configuration, not measured behavior. Repository activity, popularity, or a working demo would not change that; a demo that completes one tool call on camera is not evidence about the ninety-ninth call.
This gap is the actual state of the comparison. The vendor report quotes task-success numbers for its own two versions and stops its comparisons against GPT-Realtime-2 at safety metrics; the community repo documents configuration, not behavior. Neither answers the question a launch decision needs.
Why the tool call is where voice agents die
The historical pattern explains why. Deployment of LLM agents began to accelerate in late 2023, after OpenAI’s function-calling API became available: structured tool invocation is what turned chat models into agents, and it has been the fragile seam ever since. In text agents, a failed tool call produces a wrong or stalled response. In a voice agent, it produces something worse: dead air, a hallucinated confirmation (“your table is booked” when the booking API returned an error), or a recovery attempt spoken over a user who has already lost patience.
Three failure modes do most of the damage, and none of them appear in transcript-quality metrics:
- Pending-tool behavior. What does the agent say and do while a lookup or booking call is in flight? Stall phrasing, premature confirmation, and silence are all failure states with different user costs.
- Stale context after interruption. If the user barges in mid-tool-call (“actually, make it Friday”), does the agent cancel the in-flight call, carry the correction into the retry, or complete the original action anyway?
- Failed-call recovery. When the tool errors, times out, or returns empty, does the agent say so, retry sensibly, and avoid claiming success?
Two things in the Qwen report deserve credit here, because they score exactly these behaviors. The evaluator “waits for a pending action’s deadline and never rewards a claim of completion before it is true,” and “premature promises, repeated filler, and unnecessary internal details are penalized” (report). The report also evaluates Full-Duplex-Bench v3.0’s tool-selection and argument-accuracy metrics alongside its τ-Voice numbers. What is missing is not the metric; it is any run of these tests by someone other than the vendor.
A voice stack that scores well on speech recognition and synthesis can still fail all three. That is the sense in which the bottleneck has moved from speech quality to tool-layer reliability, and why a benchmark like τ-Voice, which at least attempts task-success measurement, is directionally the right instrument, whoever runs it.
Reliability is also a training problem
Two other recent arXiv papers explain why multi-turn agentic reliability resists quick fixes, even though neither studies voice agents directly.
The first, From Self-Distillation to Self-Practice, reports a counterintuitive training result: a distillation variant trained with privileged information (giving the teacher extra context the student will not have at inference) can end up below the untrained base model in multi-turn agents. The proposed alternative, Privileged Self-Practice, improves “task-goal completion by up to 65% on AppWorld and the resolved rate by up to 61% on SWE-bench Verified” (paper). The caution for voice-agent builders: if you plan to fine-tune a realtime model whose weights you can run on tool-call traces, the naive recipe can make multi-turn behavior worse. Training-side regressions are a deployment risk, not an academic curiosity.
The second, a survival analysis of multi-turn consistency, offers a method rather than a model score. It applies Kaplan-Meier survival curves to agent episodes of T=20 steps, pooling N=84,540 trajectories across 8 models (paper), to measure how long an agent sustains coherent behavior before degrading. Critically, the authors state they “do not evaluate adaptive replanning, tool use in open-ended environments, or domain expertise.” So the study cannot be cited as evidence about voice agents or tool calling at all. What it contributes is a better instrument: instead of asking “did the session succeed,” ask “at which turn do sessions like mine start dying.” That question transfers directly to voice-agent evaluation even though the paper’s results do not.
A third data point sharpens the recovery axis. In a study of safe skill retirement for physical agents, different models failed differently at the moment of action: “GPT-5.6-sol refused before proposing, while the sink denied Gemma4:31b.” Failure style varies by model and by safety architecture, which means your failed-call testing has to run against the specific model you ship, not a stand-in from the same family.
A pre-launch action-completion checklist
Since the vendor report settles only its own side of the comparison, the deciding artifact is a test suite you run yourself. Build it around the three failure modes above, against your actual tools, before launch:
- Pending-tool behavior. Issue requests whose tools take 2, 5, and 15 seconds. Check what the agent says during the wait, whether it ever confirms before the result returns, and whether latency degrades speech quality.
- Interruption with a pending call. Barge in mid-tool-call with a correction. Verify the in-flight call is cancelled or superseded, the correction is carried into the retry, and the original action does not silently complete.
- Failed, empty, and timed-out calls. Force tool errors. The agent must say the action failed, must not claim success, and should retry at most a bounded number of times before escalating to the user.
- Session-level survival. Borrow the survival-analysis framing: run multi-turn episodes, log the turn index of the first unrecoverable failure, and compare stacks on where sessions die rather than on a single pass rate. The survival-analysis paper shows the method at T=20 steps; your episodes should reflect your real conversation lengths.
- Safety at the action boundary. Red-team the moment of commitment, not just the dialogue. The Qwen report’s 92.00% vs 96.00% session-level pass rates show vendors measure this; you should too, on your action set.
Run the identical suite against both candidate stacks. The stack that survives your tool layer wins your decision, regardless of anyone’s benchmark adaptation.
So, Qwen’s realtime line or a hosted API?
On the available evidence, the verdict is deliberately unsatisfying: unproven on both sides. Qwen-Audio-3.1-Realtime is a promising entry in Qwen’s realtime line whose reliability claims are self-reported and measured on the vendor’s own half-duplex adaptation. Its comparisons against GPT-Realtime-2 are safety metrics that split: GPT-Realtime-2 wins the red-team pass rate, 96.00% to 92.00% (report), and TruthfulQA, while Qwen-Audio-3.1-Realtime leads most of the rest. The τ-Voice and EVA-A results are vendor-run, and no third-party run accompanies them in the report; the report documents no function-calling behavior for GPT-Realtime-2 beyond its safety numbers, and no public function-calling measurement of a hosted realtime API accompanies these results. A community repo’s model default is the nearest non-vendor data point, and it records configuration, not behavior.
The narrowing matters. If your decision must be made on this evidence alone, it rests on one vendor’s grading of its own models. If hosted APIs are genuinely on the table, source their function-calling documentation and run the checklist above before treating the comparison as real. As Groundy’s route-by-axis model comparison concluded for Chinese flagships generally: vendor-graded leads against fields of one do not route workloads; measured behavior on your workload does.
The durable takeaway outlasts this release cycle. Voice-agent evaluation is shifting from transcript quality to action completion, and that shift changes the economics of shipping: a demo that sounds good is cheap, and an agent that survives pending calls, interruptions, and failed tools is the expensive part. Budget accordingly, and treat every reliability number you did not produce as a hypothesis about your own test suite.
Frequently Asked Questions
Does the Qwen report include function-calling comparisons against GPT-Realtime-2?
No public function-calling measurement of a hosted realtime API accompanies these results. On the hosted-API side, function-calling success rates, latency under tool-call load, and behavior when a tool errors go unmeasured. Any claim about how GPT-Realtime-2 handles tool calls in production would go beyond what the report measured. The honest narrowing: this article can compare what the vendor measured, and it cannot rank the stacks.

Join the discussion
Share a useful perspective or ask a question about this article.