models & research
Top in models & research
Operator-Level Triage for Silent Mixed-Precision Instability
arXiv 2607.25494 introduces a single-pass tool to localize silent numerical instability in bf16 and fp16 training. Treat it as a triage layer for unexplained loss spikes, not.
modelsBeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration
arXiv:2607.25600 shows verbalized confidence is a weak but real routing signal for RAG. It saves 20.4% retrieval calls for 28.2% token overhead. The probe is poorly.
Kimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
Running Kimi K3 on an M1 Max proves local MoE feasibility via expert offloading, but unified memory bandwidth caps throughput. Treat this as a prototyping tool, not a serving.
modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
Kimi Linear cuts KV cache 75% and boosts decode 6x for 1M-token contexts, but recall stays the binding constraint. Route memory-bound jobs only after benchmarking in-context.
modelsWhy Chat Leaderboards Do Not Predict Image Quality
A July 22 visual test comparing GPT-5.6, Claude, Gemini, and Grok highlights a structural gap: chat models lack dedicated image backends. Teams must route by subtask using.
modelsKimi K3 Procurement: Governance Review Over Phantom Government Assessments
Moonshot AI's Kimi K3 release triggers governance review for regulated teams. Beijing jurisdiction, Anthropic accusations, and missing government assessment require internal.
modelsDeepSeek Compute Leak: Why Open-Weight Routing Needs a Swap Path
A leaked transcript claims DeepSeek faces a compute gap, but V4 shipped in April. Teams should treat this as a resilience prompt to pin Qwen as a swap-in spine, not panic.
modelsContext Ordering Beats Window Size for Long-Context Agents
ARBIGRAPH shows tool agents lose 33.3% accuracy on dependent chains despite fitting context. Context ordering, truncation, and caching now beat raw window size for agent.
- jul 24modelsDiffusion LLMs: Training Cost, Not Parallel Decoding, Drives Deployment
- jul 24modelsOpen-Weight Routers vs Fable 5: The Routing Math That Actually Matters
- jul 23modelsDeepSeek-V4 1M Context vs RAG: Why Retrieval Stays
- jul 23modelsQwen-Image-3.0 Does Not Exist: Why Self-Hosting Image Models Is Premature
- jul 20modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- jul 20modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- jul 20modelsQwen3.8 Max Preview: Missing Benchmarks, Weights, and Pricing
- jul 19modelsHuggingFace 100x Inference: Generalizable vs Platform-Locked Optimizations
- jul 19modelsKimi K3 Code Arena Rank: Self-Hosting Cost Math for Coding Agents
- jul 14modelsCan Tool-Adaptive LLM Rerankers Improve RAG Without Always Calling Tools?
- jul 14modelsDoes Speculative Decoding with Progressive Tree Drafting Cut LLM Latency?
- jul 10modelsAnalytic Inference Cuts Bayesian Deep Ensemble Serving Cost, But Leaves Training as the Bottleneck
- jul 10modelsTree-of-Thoughts Improves Text-to-Image Prompting by Reasoning Over Hypotheses, Not Pixels
- jul 10modelsFourierQK's spectral Q/K filter cuts TinyShakespeare loss by 79%, but long-context proof is missing
- jul 10modelsTencent Hunyuan 3's Agent Push Has No Public DeepSeek or Qwen Benchmarks Yet
- jul 10modelsWhen Does Memory, Not Compute, Decide Who Can Profitably Serve LLMs?
- jul 10modelsCan We Trust LLM Logic? A Graph-Based Stress Test Finds Three Failure Modes
- jul 09modelsLongCat-2.0 Hits Claude Opus 4.6 Class on Agents From a 50K-GPU Cluster
- jul 08modelsWhen Do Time Series Foundation Models Pay Off? The Break-Even Threshold
- jul 07modelsVLA Grounder Tests Language Conditioning to Optimize Black-Box Vision-Action Models
- jul 07modelsBlack-Box LLM Architecture Inference: What API Restrictions Reveal About Hidden Model Structure
- jul 07modelsInduceKV Tests Continual Learning for Multimodal LLMs Without Expanding Cache
- jul 03modelsSonnet 5 vs GPT-5.5: Pricing, Benchmarks, and the Switching Math
- jun 30modelsHow LLMs Fuse Conflicting Facts: Single-Source vs Multi-Source Truth
- jun 29modelsLinear Transformers Get a Learnable Kernel: Does Flexformer Change the Efficiency Tradeoff?
- jun 29modelsHuawei Ships CUDA-Free AI Compute On-Device, but Ascend Quantization Accuracy Is Unverified
- jun 29modelsDo Multimodal RAG Models Ignore Late Evidence? A Primacy Bias Test
- jun 29modelsCan Deep Learning Design RF Power Amplifiers Without Full EM Simulation?
- jun 28modelsSynthetic Clinical Notes from LLMs: Believable Prose Is Not Clinical Validity
- jun 28modelsDoubao vs Qwen 3.7 vs GLM-5.2: Route by Axis, Not Leaderboard
- jun 28modelsCan Dynamic Experts Fix Catastrophic Forgetting in Robot Manipulation?
- jun 28modelsError-Conditioned Neural Solvers vs Iterative Refinement: When Does Learned Correction Win?
- jun 28modelsVision-Language Models Move Past Object Detection: The MLLM Perception Shift
- jun 28modelsCan Autoregressive Boltzmann Generators Replace MCMC in Simulation?
- jun 28modelsLook-Before-Move Plans Observation Before Motion in Dynamic 3D Story Worlds
- jun 27modelsGLM 5.2, Qwen 3.7, and DeepSeek in 2026: A Routing Map by Workload, Not by Rank
- jun 27modelsSLM Pipeline Catches 10% of Papers Human Reviewers Missed, but No Model Matched Human Accuracy
- jun 27modelsMiniMax M3 vs GLM-5.2: Whose 1M-Context Claim Holds Up?
- jun 27modelsDeepSeek V4.1 Flash vs Qwen 3.7 vs Llama 4.5: June 2026 HF Trending Ranks Velocity, Not Installs
- jun 27models125 Targeted Wikipedia Edits Left a Detectable Signal in Llama Pretraining
- jun 27modelsCan SAE Features Stop LLMs From Forgetting During Continual Learning?
- jun 27modelsCan a 30B Model Post-Train Itself? A-Evolve-Training Tests Autonomous RL
- jun 26modelsOpen-Weight LLM Leaderboards 2026: Where DeepSeek, Qwen, and GLM Rank
- jun 26modelsQwen3.7-Max's Top-Ranked Claim vs the Artificial Analysis Index
- jun 26modelsDoes Tree-of-Thought Reasoning Scale to Billion-User Modeling?
- jun 26modelsCan LLMs Debug Verilog? VeriPilot Puts an Agent on RTL Errors
- jun 25modelsTask Decomposition Helps LLMs by Shrinking Output Space, Not by Cutting Labeling Cost
- jun 25modelsFlow Matching vs U-Net: A Skip-Free Backbone for Speech Models
- jun 25modelsA Per-Neuron Sequence Model Was Withdrawn From arXiv as Coverage Hailed It
- jun 25modelsPV-TAM Corrects Decoding Drift and Boundary-Marker Bias in VLM Localization Scoring
Every new foundation model arrives wrapped in a benchmark chart and a press release. This beat exists because the chart almost never tells you what the model actually does — and the press release tells you even less. The interesting work is one layer down: the attention variant that changes the cost curve, the parameterization fix that makes scaling laws actually transfer, the eval protocol whose failure modes flip the standings once you change a decoding parameter or a codec. We cover that layer.
We treat open-weight and closed releases as the same story told from opposite ends of a distribution shift. A trillion-parameter mixture beating a dense competitor on one harness and losing on another is not a contradiction; it is evidence about what the harness measures. Training-efficiency claims, context-window claims, and reasoning-benchmark claims all get the same treatment — read the method section, find the assumption that was load-bearing, and report whether removing it changes the result. When a paper proposes a new safety mechanism, a new compression trick, or a new continual-learning split, we are interested in the part the authors did not want to highlight.
The throughline is comparative and skeptical without being contrarian. Foundation-model research has become the field where the largest gap between published numbers and deployed behavior tends to live. Closing that gap — with reproduction notes, ablation reading, and honest accounting of what generalizes versus what was tuned into the eval — is the beat.