
GLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
Zhipu published a full benchmark suite for GLM-5.2 on June 19, 2026. Each score targets a different skill domain, and each carries distinct contamination or harness caveats.
The archive · Page 5 of 6
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
97–120 of 129 articles · Newest first

Zhipu published a full benchmark suite for GLM-5.2 on June 19, 2026. Each score targets a different skill domain, and each carries distinct contamination or harness caveats.

GLM-5.2 scores 81.0 on Terminal-Bench 2.1, trails Claude Opus 4.8 (85.0) on shell tasks, but wins on 1M-token monorepo context. Here is how to route tasks.
GLM-5.2 posts 62.1% on SWE-bench Pro and 81.0 on Terminal-Bench 2.1, four points behind Opus 4.8. MIT weights are self-hostable; flat plan starts at $18/month.
GLM-5.2 ships 753B parameters under MIT with strong coding benchmarks, but the full MoE weight load makes self-hosting far heavier than the license terms imply.
STAR uses attention-derived spatial maps to replace uniform scalar reward in diffusion RL, shifting the bottleneck from preference pair volume to reward localization.
A June 2026 preprint localizes Gemma 4 repetition loops to a few MLP neurons and removes them with a one-time weight edit while benchmark scores hold.
Claude Fable 5 claims top benchmark scores. Verified data shows every model below 14% on FrontierCode Diamond, and partner scores lack public methodology.
A June 2026 paper traces attribution patching's errors to downstream non-linearities and proposes a Hessian-vector-product correction that costs one extra backward pass.
Task-aware layer pruning removes distortion-amplifying layers and improves out-of-distribution accuracy, which means standard in-distribution benchmarks miss the real effect.
IMUG-Bench tests whether unified multimodal models can alternate between understanding and generation in one context, exposing gaps that separate benchmarks conceal.
New research isolates a compact attention-head circuit for entity rebinding in Gemma and Llama, showing tracking failures stem from a binding step, not context length.
Claude Fable 5 prices at $10/$50 per million tokens, 2x Opus 4.8. Frontier research, long-context agents, and molecule design clear the bar. Standard coding does not.
Claude Mythos 5 shares Fable 5's architecture but with safeguards lifted in select areas. Access requires Project Glasswing approval or a biology research designation.
Claude Fable 5 ships with distillation protection to prevent capability extraction. A first-principles look at what it is, how it works, and why API consumers should care.
A new study claims LLMs write 'appropriate' research titles, but the evidence rests on similarity metrics that measure pattern matching, not whether titles actually serve.
Standard concept bottleneck model benchmarks confound genuine concept learning with dataset shortcuts. Synthetic benchmarks from Skirzynski et al. expose the gap.
A 43-model audit finds that geographic and role framing in LLM prompts systematically shifts which scholars get recommended as experts, with no neutral default.

Opus 4.8 has a 1M token context window (200k on Foundry), 128k standard output, and 300k output via Batch API beta. January 2026 cutoff. Batch design and quota allocation.
An arXiv paper shows the embedding learning rate accounts for most of μP's advantage over standard parameterization, and a single scaling fix recovers the bulk of the benefit.
Alibaba's dense Qwen3.6-27B outperforms its MoE sibling on coding benchmarks, trading predictable inference latency for a larger memory footprint than sparse alternatives.

Compare DeepSeek, Qwen, Kimi, Doubao, ERNIE and GLM through dated benchmarks, license terms, context windows, API pricing and practical workload fit.

Running DeepSeek R1 locally on real hardware: per-quant VRAM budgets, computed decode ceilings for a 24 GB card, measured quality loss, and sourced throughput figures.
Fish Audio's open-source S2 TTS hit SOTA in March 2026 with sub-100ms latency and 80+ languages, challenging ElevenLabs' commercial pricing and exposing licensing gaps.

The supply of high-quality human text for AI training is running out. Synthetic data fills the gap, but model collapse and provenance drift set the limits.