articles
all articles
feed
- jun 21ossAdam's Open-Source AI CAD Claim Lacks a Confirmed Repo or Accuracy Benchmark
- jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
- jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
- jun 21ossEpic Open-Sources Lore, a VCS Pitched at Git's Scaling Ceiling
- jun 21infraRunning Long-Context Agents on a 4-Bit KV Cache: Where Accuracy Breaks
- jun 21securityDefending Agentic AI With Deception: Misdirecting Model-Guided Attacks
- jun 21securityThe Autonomy Tax: Why RL Rewards the Wrong Behavior in Agents
- jun 21securityAnthropic's Procurement Risk Is Policy Refusal, Not Jailbreaks
- jun 20industryCan You Predict a Fine-Tune's Payoff Before Training Finishes?
- jun 20cultureWhen an Algorithm Sequences Gig Hiring, Whose Objective Does It Optimize?
- jun 20infraWhen LLM-Generated CUDA Kernels Pass Tests but Get the Math Wrong
- jun 20modelsCan RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
- jun 20modelsPruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
- jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
- jun 20modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
- jun 20modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
- jun 20devtoolsCursor Goes to SpaceX, Windsurf to Cognition: What Changes for Dev Teams
- jun 19cultureAI Essay Grading: What a Probe of LLM Internals Reveals About Scoring
- jun 19modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- jun 19policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
- jun 19modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
- jun 19modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
- jun 19modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
- jun 19infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
- jun 19devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
- jun 18modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
- jun 16ossZhipu Open-Sources GLM-5.2 Under MIT While Anthropic Tightens Model Access
- jun 16modelsCan Editing One Neuron Fix LLM Repetition Loops?
- jun 16industryZhipu Ships GLM-5.2 With 1M Context and MIT Weights, but Zero Benchmarks at Launch
- jun 16infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
- jun 16devtoolsVercel's Remend Turns Streaming-Markdown Repair Into a Dependency
- jun 16industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
- jun 15policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
- jun 15infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
- jun 15infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
- jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
- jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
- jun 14securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
- jun 13policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
- jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
- jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
- jun 11infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
- jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
- jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- jun 12industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
- jun 11securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
- jun 11devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
- jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
- jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
- jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
- jun 10infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
- jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
- jun 10infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
- jun 10devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
- jun 10infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
- jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
- jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
- jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
- jun 10policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
- jun 10industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
- jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
- jun 09securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
- jun 09policyWho Gets to Audit Your Health Chatbot? Almost No One
- jun 09policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
- jun 09infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
- jun 09industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
- jun 09policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
- jun 09policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
- jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
- jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
- jun 09cultureDoes Debate Quality Survive When LLMs Argue Outside English?
- jun 09securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
- jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
- jun 09policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
- jun 09industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
- jun 09devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
- jun 09ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
- jun 08policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
- jun 08policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
- jun 08securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
- jun 08securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
- jun 08securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
- jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
- jun 08industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
- jun 08industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
- jun 07policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
- jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
- jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
- jun 07cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
- jun 07infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
- jun 07policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
- jun 07industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
- jun 06infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
- jun 06securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
- jun 06securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
- jun 06cultureDo LLMs Understand Idioms in Low-Resource Languages?
- jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
- jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
- jun 06infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack