models
models & research
archive
- InduceKV Tests Continual Learning for Multimodal LLMs Without Expanding Cache
- Sonnet 5 vs GPT-5.5: Pricing, Benchmarks, and the Switching Math
- How LLMs Fuse Conflicting Facts: Single-Source vs Multi-Source Truth
- Linear Transformers Get a Learnable Kernel: Does Flexformer Change the Efficiency Tradeoff?
- Huawei Ships CUDA-Free AI Compute On-Device, but Ascend Quantization Accuracy Is Unverified
- Do Multimodal RAG Models Ignore Late Evidence? A Primacy Bias Test
- Can Deep Learning Design RF Power Amplifiers Without Full EM Simulation?
- Synthetic Clinical Notes from LLMs: Believable Prose Is Not Clinical Validity
- Doubao vs Qwen 3.7 vs GLM-5.2: Route by Axis, Not Leaderboard
- Can Dynamic Experts Fix Catastrophic Forgetting in Robot Manipulation?
- Error-Conditioned Neural Solvers vs Iterative Refinement: When Does Learned Correction Win?
- Vision-Language Models Move Past Object Detection: The MLLM Perception Shift
- Can Autoregressive Boltzmann Generators Replace MCMC in Simulation?
- Look-Before-Move Plans Observation Before Motion in Dynamic 3D Story Worlds
- GLM 5.2, Qwen 3.7, and DeepSeek in 2026: A Routing Map by Workload, Not by Rank
- SLM Pipeline Catches 10% of Papers Human Reviewers Missed, but No Model Matched Human Accuracy
- MiniMax M3 vs GLM-5.2: Whose 1M-Context Claim Holds Up?
- DeepSeek V4.1 Flash vs Qwen 3.7 vs Llama 4.5: June 2026 HF Trending Ranks Velocity, Not Installs
- 125 Targeted Wikipedia Edits Left a Detectable Signal in Llama Pretraining
- Can SAE Features Stop LLMs From Forgetting During Continual Learning?
- Can a 30B Model Post-Train Itself? A-Evolve-Training Tests Autonomous RL
- Open-Weight LLM Leaderboards 2026: Where DeepSeek, Qwen, and GLM Rank
- Qwen3.7-Max's Top-Ranked Claim vs the Artificial Analysis Index
- Does Tree-of-Thought Reasoning Scale to Billion-User Modeling?
- Can LLMs Debug Verilog? VeriPilot Puts an Agent on RTL Errors
- Task Decomposition Helps LLMs by Shrinking Output Space, Not by Cutting Labeling Cost
- Flow Matching vs U-Net: A Skip-Free Backbone for Speech Models
- A Per-Neuron Sequence Model Was Withdrawn From arXiv as Coverage Hailed It
- PV-TAM Corrects Decoding Drift and Boundary-Marker Bias in VLM Localization Scoring
- Meituan's General 365 Benchmark: Top Models All Score Under 63%
- LLM Surrogates in A/B Tests: The 39% Recovery Gap and the Silent Bias Risk
- LLM Token Pricing vs Compute Cost: What the Tokenomics Math Shows
- Do LLM Judges Favor Their Own Output? A Sanity Check on Self-Preference
- Can AI Write CAD Programs? CADBench Measures the Gap
- ByteDance's Doubao 2.1 Pro vs GPT-5.5: Reading Self-Reported Benchmarks
- Can RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
- Pruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
- GLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
- How Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
- GLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- GLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
- GLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
- GLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
- STAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
- Can Editing One Neuron Fix LLM Repetition Loops?
- Claude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
- Does Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- Can You Make a Multimodal Model Unlearn With Activation Steering?
- Why Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- Do Unified Multimodal Models Actually Interleave Understanding and Generation?