models
models & research
more in this beat
- jun 25modelsMeituan's General 365 Benchmark: Top Models All Score Under 63%
- jun 25modelsLLM Surrogates in A/B Tests: The 39% Recovery Gap and the Silent Bias Risk
- jun 25modelsLLM Token Pricing vs Compute Cost: What the Tokenomics Math Shows
- jun 25modelsDo LLM Judges Favor Their Own Output? A Sanity Check on Self-Preference
- jun 24modelsCan AI Write CAD Programs? CADBench Measures the Gap
- jun 24modelsByteDance's Doubao 2.1 Pro vs GPT-5.5: Reading Self-Reported Benchmarks
- jun 20modelsCan RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
- jun 20modelsPruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
- jun 20modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
- jun 20modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
- jun 19modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- jun 19modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
- jun 19modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
- jun 19modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
- jun 18modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
- jun 16modelsCan Editing One Neuron Fix LLM Repetition Loops?
- jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
- jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
- jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
- jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
- jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
- jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
- jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
- jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
- jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
- jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
- may 29modelsPersona Prompts Change Who an LLM Recommends as an Expert
- may 28modelsOpus 4.8 Batch API: 1M Context, 300k Output, and Team Cost Controls
- may 25modelsμP Hyperparameter Transfer Has an Embedding Layer Hole, New arXiv Paper Says
- apr 23modelsQwen3.6-27B's Dense Architecture Challenges the MoE-Only Playbook for Flagship-Class Coding Models
- mar 24modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
- mar 24modelsRunning DeepSeek R1 Locally: Hardware Requirements, Quantization, and Real Throughput
- mar 15modelsFish-Speech: The Open-Source TTS Model That's Threatening ElevenLabs
- feb 27modelsSynthetic Data Is Eating AI Training
- feb 27modelsGoogle's TimesFM: A Foundation Model for Time Series
- feb 27modelsGemini 2.0 Pro's 2 Million Token Context: What Can You Actually Do With It?
- feb 27modelsDeepSeek V3/R1: How Chinese Engineers Matched GPT-4 for $6 Million
- feb 27modelsClaude's Web Search Changes Everything for AI Research
- feb 27modelsThe Million-Token Context Window: What Can You Actually Do?
- feb 18modelsKimi Claw: Moonshot AI's Answer to Claude and ChatGPT
- feb 18modelsWiFi DensePose: Full-Body Tracking Through Walls Using Your Router
- feb 15modelsAI Code Generation Benchmarks 2026: Which Model Actually Writes Better Code?
- feb 11modelsThe Best AI Models for OpenClaw in 2026