models
models & research
more in this beat
- jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
- jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
- jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
- jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
- jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
- jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
- jun 10modelsSkip Fable 5 or Upgrade? When Opus 4.8 and Sonnet 4.6 Are Still Enough
- jun 09modelsLLM Steganography: Can Defenders Detect Payloads Hidden in Model Output?
- jun 09modelsDo Privacy Defenses Actually Protect Fine-Tuned LLMs? A New Benchmark
- jun 09modelsCan You Reconstruct an LLM's System Prompt From Its Activations?
- jun 09modelsDoes Softmax Normalization Limit What Attention Can Represent?
- jun 08modelsCan an Attacker Steal Your Model's Last Layer From Its Outputs?
- jun 07modelsCan LLMs Leak Training Data? A New Test Splits Capacity From Intent
- jun 07modelsWhen an AI Agent's Tools Break, Can It Recover? A New Benchmark
- jun 06modelsMiniMax M3 Bets on Sparse Attention for 1M Context. Does the Math Hold?
- jun 06modelsCan One Model Handle Every CAD Task? UniCAD Tests It
- jun 06modelsDo Foundation Models Actually Learn Relational Structure In-Context?
- jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
- jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
- jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
- jun 06modelsContinuous Bit-Width Quantization vs Fixed INT4: Does LiftQuant Beat Discrete?
- jun 05modelsFederated Learning for Industrial IoT Anomaly Detection: The Data-Locality Tradeoff
- jun 05modelsReading Failed LLM Reasoning Traces Won't Tell You Which Ones RL Can Fix
- jun 05modelsCan You Stitch Two Foundation Models Together Without Retraining?
- jun 05modelsDo Reasoning LLMs Waste Tokens? OckBench Tries to Measure It
- jun 04modelsWhich Layer Detects LLM Hallucinations Best? The Case Against Fixed-Layer Probes
- jun 03modelsCross-Domain RL Training Degrades Capabilities. CARE-RL Reweights to Fix It
- jun 03modelsLLM Watermarking Without Quality Loss: The Non-Distortionary Approach
- jun 02modelsTreating LLM Agent Memory as a Database: The VikingMem Approach
- jun 02modelsCan a Language Model Work Without a Neural Network? A New arXiv Paper Says Yes
- jun 02modelsCan Code-Generating LLMs Do Engineering Math? FEM-Bench Tests Them
- jun 02modelsUnlearning Isn't Deletion: arXiv 2505.16831 Shows Machine Unlearning in LLMs Is Reversible
- jun 01modelsWhy LLMs Fail at Spatial Reasoning When Planning Navigation
- jun 01modelsDoes Giving AI Agents More Skills Help? A Controlled SkillsBench Study
- may 31modelsCan an LLM Peer-Review Your Paper? A New Behavior Benchmark
- may 31modelsAnthropic Scaled Sparse Autoencoders to Claude 3 Sonnet. Interpretability Now Costs Compute
- may 29modelsTracing Why LLM Agent Memory Fails: A Method for Attributing Errors
- may 29modelsPersona Prompts Change Who an LLM Recommends as an Expert
- may 28modelsOpus 4.8 vs Opus 4.7: What Changed and What Did Not
- may 28modelsOpus 4.8 Batch API: 1M Context, 300k Output, and Team Cost Controls
- may 27modelsScale Vectors: Tiny Parameter Subsets That Disproportionately Steer LLM Behavior
- may 27modelsOne Learning Rate Doesn't Fit All: Heavy-Tail Layerwise LR Schedules for LLM Pretraining
- may 26modelsAudio LLMs Break When the Codec Changes: A Robustness Vector Voice-AI Teams Haven't Tested
- may 26modelsDo LLMs Know What Not to Say? Causal Evidence for Statistical Preemption
- may 25modelsEmbedding Compression at Training Time: DIVE's Gradient Trick vs Post-Hoc Quantization for Vector DBs
- may 25modelsμP Hyperparameter Transfer Has an Embedding Layer Hole, New arXiv Paper Says
- may 24modelsProject Glasswing One Month In: AI Bug Discovery Has Outpaced the Patch Pipeline
- may 23modelsarXiv 2605.16428 Measures AI Search's Drag on Publisher Traffic Using Paired Google and Reddit Data