groundy
articlessearch

models & research

  1. InduceKV Tests Continual Learning for Multimodal LLMs Without Expanding Cache
  2. Sonnet 5 vs GPT-5.5: Pricing, Benchmarks, and the Switching Math
  3. How LLMs Fuse Conflicting Facts: Single-Source vs Multi-Source Truth
  4. Linear Transformers Get a Learnable Kernel: Does Flexformer Change the Efficiency Tradeoff?
  5. Huawei Ships CUDA-Free AI Compute On-Device, but Ascend Quantization Accuracy Is Unverified
  6. Do Multimodal RAG Models Ignore Late Evidence? A Primacy Bias Test
  7. Can Deep Learning Design RF Power Amplifiers Without Full EM Simulation?
  8. Synthetic Clinical Notes from LLMs: Believable Prose Is Not Clinical Validity
  9. Doubao vs Qwen 3.7 vs GLM-5.2: Route by Axis, Not Leaderboard
  10. Can Dynamic Experts Fix Catastrophic Forgetting in Robot Manipulation?
  11. Error-Conditioned Neural Solvers vs Iterative Refinement: When Does Learned Correction Win?
  12. Vision-Language Models Move Past Object Detection: The MLLM Perception Shift
  13. Can Autoregressive Boltzmann Generators Replace MCMC in Simulation?
  14. Look-Before-Move Plans Observation Before Motion in Dynamic 3D Story Worlds
  15. GLM 5.2, Qwen 3.7, and DeepSeek in 2026: A Routing Map by Workload, Not by Rank
  16. SLM Pipeline Catches 10% of Papers Human Reviewers Missed, but No Model Matched Human Accuracy
  17. MiniMax M3 vs GLM-5.2: Whose 1M-Context Claim Holds Up?
  18. DeepSeek V4.1 Flash vs Qwen 3.7 vs Llama 4.5: June 2026 HF Trending Ranks Velocity, Not Installs
  19. 125 Targeted Wikipedia Edits Left a Detectable Signal in Llama Pretraining
  20. Can SAE Features Stop LLMs From Forgetting During Continual Learning?
  21. Can a 30B Model Post-Train Itself? A-Evolve-Training Tests Autonomous RL
  22. Open-Weight LLM Leaderboards 2026: Where DeepSeek, Qwen, and GLM Rank
  23. Qwen3.7-Max's Top-Ranked Claim vs the Artificial Analysis Index
  24. Does Tree-of-Thought Reasoning Scale to Billion-User Modeling?
  25. Can LLMs Debug Verilog? VeriPilot Puts an Agent on RTL Errors
  26. Task Decomposition Helps LLMs by Shrinking Output Space, Not by Cutting Labeling Cost
  27. Flow Matching vs U-Net: A Skip-Free Backbone for Speech Models
  28. A Per-Neuron Sequence Model Was Withdrawn From arXiv as Coverage Hailed It
  29. PV-TAM Corrects Decoding Drift and Boundary-Marker Bias in VLM Localization Scoring
  30. Meituan's General 365 Benchmark: Top Models All Score Under 63%
  31. LLM Surrogates in A/B Tests: The 39% Recovery Gap and the Silent Bias Risk
  32. LLM Token Pricing vs Compute Cost: What the Tokenomics Math Shows
  33. Do LLM Judges Favor Their Own Output? A Sanity Check on Self-Preference
  34. Can AI Write CAD Programs? CADBench Measures the Gap
  35. ByteDance's Doubao 2.1 Pro vs GPT-5.5: Reading Self-Reported Benchmarks
  36. Can RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
  37. Pruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
  38. GLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
  39. How Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
  40. GLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
  41. GLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
  42. GLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
  43. GLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
  44. STAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
  45. Can Editing One Neuron Fix LLM Repetition Loops?
  46. Claude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
  47. Does Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
  48. Can You Make a Multimodal Model Unlearn With Activation Steering?
  49. Why Pruning a Model Can Raise Its Out-of-Distribution Accuracy
  50. Do Unified Multimodal Models Actually Interleave Understanding and Generation?