groundy

models & research

  1. jun 25modelsMeituan's General 365 Benchmark: Top Models All Score Under 63%
  2. jun 25modelsLLM Surrogates in A/B Tests: The 39% Recovery Gap and the Silent Bias Risk
  3. jun 25modelsLLM Token Pricing vs Compute Cost: What the Tokenomics Math Shows
  4. jun 25modelsDo LLM Judges Favor Their Own Output? A Sanity Check on Self-Preference
  5. jun 24modelsCan AI Write CAD Programs? CADBench Measures the Gap
  6. jun 24modelsByteDance's Doubao 2.1 Pro vs GPT-5.5: Reading Self-Reported Benchmarks
  7. jun 20modelsCan RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
  8. jun 20modelsPruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
  9. jun 20modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
  10. jun 20modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
  11. jun 19modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
  12. jun 19modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
  13. jun 19modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
  14. jun 19modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
  15. jun 18modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
  16. jun 16modelsCan Editing One Neuron Fix LLM Repetition Loops?
  17. jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
  18. jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
  19. jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
  20. jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
  21. jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
  22. jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
  23. jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
  24. jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
  25. jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
  26. jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
  27. jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
  28. jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
  29. may 29modelsPersona Prompts Change Who an LLM Recommends as an Expert
  30. may 28modelsOpus 4.8 Batch API: 1M Context, 300k Output, and Team Cost Controls
  31. may 25modelsμP Hyperparameter Transfer Has an Embedding Layer Hole, New arXiv Paper Says
  32. apr 23modelsQwen3.6-27B's Dense Architecture Challenges the MoE-Only Playbook for Flagship-Class Coding Models
  33. mar 24modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
  34. mar 24modelsRunning DeepSeek R1 Locally: Hardware Requirements, Quantization, and Real Throughput
  35. mar 15modelsFish-Speech: The Open-Source TTS Model That's Threatening ElevenLabs
  36. feb 27modelsSynthetic Data Is Eating AI Training
  37. feb 27modelsGoogle's TimesFM: A Foundation Model for Time Series
  38. feb 27modelsGemini 2.0 Pro's 2 Million Token Context: What Can You Actually Do With It?
  39. feb 27modelsDeepSeek V3/R1: How Chinese Engineers Matched GPT-4 for $6 Million
  40. feb 27modelsClaude's Web Search Changes Everything for AI Research
  41. feb 27modelsThe Million-Token Context Window: What Can You Actually Do?
  42. feb 18modelsKimi Claw: Moonshot AI's Answer to Claude and ChatGPT
  43. feb 18modelsWiFi DensePose: Full-Body Tracking Through Walls Using Your Router
  44. feb 15modelsAI Code Generation Benchmarks 2026: Which Model Actually Writes Better Code?
  45. feb 11modelsThe Best AI Models for OpenClaw in 2026