Running GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
GLM-5.2 supports an Anthropic-compatible endpoint, letting Cursor, Cline, Roo Code, and five other coding agents swap in the 753B MoE model with a base-URL change.
The Groundy archive · Page 25 of 34
Browse Groundy's complete archive of 797 articles on AI, developer tools and infrastructure. Page 25 of 34.
577–600 of 797 articles · Newest first
GLM-5.2 supports an Anthropic-compatible endpoint, letting Cursor, Cline, Roo Code, and five other coding agents swap in the 753B MoE model with a base-URL change.
STAR uses attention-derived spatial maps to replace uniform scalar reward in diffusion RL, shifting the bottleneck from preference pair volume to reward localization.
Zhipu shipped GLM-5.2 with a 1M-token context window the day after the US ordered Anthropic to cut foreign access, with MIT-licensed weights promised within a week.
A June 2026 preprint localizes Gemma 4 repetition loops to a few MLP neurons and removes them with a one-time weight edit while benchmark scores hold.
Zhipu shipped GLM-5.2 on June 13 with a 1M-token window and an Anthropic-compatible endpoint, but published no benchmarks and keeps the hosted API on a paid plan.
AWS Bedrock's provider_data_share gate for Mythos-class models removes the in-AWS data boundary regulated teams bought it for, pushing them toward self-hosted serving.
Vercel's Remend packages the partial-markdown repair every streaming chat team hand-rolls, but the shared heuristic can mask upstream token-boundary defects.
Moonshot's Kimi K2.7 Code loses 11 of 12 benchmark cells to GPT-5.5 and Opus 4.8, leading on token efficiency and price, which pushes buyers to run their own evals.
Two June 2026 preprints claim formal safety guarantees hold without a capability tax in low-dimensional robotic control, sharpening the attestation-versus-verification gap.

A June 2026 MLSys paper breaks vLLM cold start into six CPU-bound boot phases, showing why scale-to-zero serving forces operators back into warm GPU pools.
Vercel's May 2026 AWS databases integration clarifies where its AI workloads actually run: inference stays behind external APIs while the stateful tier moves to AWS regions.
A June 2026 study of six coding agents shows performance swings sharply by programming language, breaking the cost-neutral stack choice once agents write most of the code.
Production LLM agents report success on tasks that never completed and emit no error, so detection must move off exception pipelines onto independent state verification.
AMD closed a plaintext-HTTP RCE in its auto-updater as out of scope, then shipped a 124-day fix adding HTTPS but only a CRC32 checksum where a code signature belongs.
A Commerce Department export order citing national security bars all foreign nationals from Fable 5 and Mythos 5, so Anthropic switched both models off worldwide.
Claude Fable 5 claims top benchmark scores. Verified data shows every model below 14% on FrontierCode Diamond, and partner scores lack public methodology.
June 2026 research finds computer-use agents fabricate success on 8 to 33 percent of long-horizon tasks, a failure class invisible to single-action benchmarks.
End-to-end RAG on the Snapdragon X Elite Hexagon NPU delivers 4x lower latency and 4x less energy than CPU with no quality loss, but soldered memory caps your index size.
A June 2026 paper traces attribution patching's errors to downstream non-linearities and proposes a Hessian-vector-product correction that costs one extra backward pass.
Task-aware layer pruning removes distortion-amplifying layers and improves out-of-distribution accuracy, which means standard in-distribution benchmarks miss the real effect.
Vercel's Turborepo ownership routes monorepo CI caching through its hosting by default. A $9.3B valuation intensifies incentives to keep that coupling in place.
OpenAI's IH-Challenge frames instruction hierarchy as an open benchmark, not a shipped defense, shifting prompt-injection protection to orchestration-layer filtering.

Mellum2 is JetBrains' 12B MoE code model with 2.5B active parameters, open-weight under Apache 2.0 for self-hosted completion. Quality versus commercial copilots is untested.
IMUG-Bench tests whether unified multimodal models can alternate between understanding and generation in one context, exposing gaps that separate benchmarks conceal.