articles
all articles
feed
- jun 16industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
- jun 15policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
- jun 15infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
- jun 15infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
- jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
- jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
- jun 14securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
- jun 13policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
- jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
- jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
- jun 11infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
- jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
- jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- jun 12industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
- jun 11securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
- jun 11devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
- jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
- jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
- jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
- jun 10infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
- jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
- jun 10infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
- jun 10devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
- jun 10infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
- jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
- jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
- jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
- jun 10policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
- jun 10industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
- jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
- jun 09securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
- jun 09policyWho Gets to Audit Your Health Chatbot? Almost No One
- jun 09policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
- jun 09infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
- jun 09industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
- jun 09policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
- jun 09policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
- jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
- jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
- jun 09cultureDoes Debate Quality Survive When LLMs Argue Outside English?
- jun 09securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
- jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
- jun 09policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
- jun 09industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
- jun 09devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
- jun 09ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
- jun 08policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
- jun 08policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
- jun 08securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
- jun 08securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
- jun 08securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
- jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
- jun 08industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
- jun 08industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
- jun 07policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
- jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
- jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
- jun 07cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
- jun 07infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
- jun 07policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
- jun 07industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
- jun 06infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
- jun 06securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
- jun 06securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
- jun 06cultureDo LLMs Understand Idioms in Low-Resource Languages?
- jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
- jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
- jun 06infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
- jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
- jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
- jun 06securityVercel's Flags SDK Exposed Feature-Flag Definitions via CVE-2025-46332
- jun 05infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
- jun 05devtoolsAlibaba's Open Code Review Moves AI Review Into the CLI, Not the PR
- jun 05infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
- jun 05infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
- jun 05securityJailbreak Suffixes Hit Harder at Specific Token Positions, New GCG Variant Shows
- jun 05policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
- jun 05policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
- jun 05securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
- jun 05cultureCan Teaching Logical Fallacies Inoculate People Against AI Misinformation?
- jun 05devtoolsVercel Ships Experimental Native CLI Binaries to Cut the Node Startup Tax
- jun 05policyRefusal Steering Targets Individual Experts in MoE LLMs
- jun 05infraPutting a Datacenter V100 in a Gaming PC: The Local LLM Math
- jun 05devtoolsVercel Rebuilds Its Marketplace CLI for Agents Instead of Humans
- jun 05securityThe 2026 npm Attacks Proved AI Coding Assistants Are a Supply-Chain Target
- jun 04securityChatGPT's New Lockdown Mode Borrows Apple's Name for a Prompt-Injection Kill Switch
- jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
- jun 04policyStacked Org Policies in LLM Chatbots Break Where Rules Collide
- jun 04securityStored Prompt Injection Now Persists Across AI Agent Sessions
- jun 04industryMiniMax M3 Bundles 1M Context and Native Multimodal Into One Open-Weight Model
- jun 04securityLLM Data Poisoning Survives the Data-Cleaning Defenses Built to Stop It
- jun 04devtoolsOpenAI Upgrades Codex Right as Teams Weigh Leaving Claude Code
- jun 03industryMorningstar's $780B SpaceX Mark Undercuts the IPO Target by Half
- jun 02ossAn Open-Source Home Camera That Encrypts End-to-End Instead of Trusting Ring
- jun 02devtoolsJetBrains Ships Codex Natively, Making Its IDE the Multi-Vendor AI Surface
- jun 01securityWhy Attack Success Rate Misleads LLM Jailbreak Benchmarks
- jun 01devtoolsTransformers.js v4 Moves Transformer Inference Into the Browser
- may 30ossAn Open-Source 80386 Rebuilt Around Intel's Original Microcode
- may 29cultureWikipedia's Foundation Is Running Big Tech's Anti-Labor Playbook, an Editor Argues