articles
all articles
feed
- modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
- modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
- devtoolsCursor Goes to SpaceX, Windsurf to Cognition: What Changes for Dev Teams
- cultureAI Essay Grading: What a Probe of LLM Internals Reveals About Scoring
- modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
- modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
- modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
- modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
- infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
- devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
- modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
- ossZhipu Open-Sources GLM-5.2 Under MIT While Anthropic Tightens Model Access
- modelsCan Editing One Neuron Fix LLM Repetition Loops?
- industryZhipu Ships GLM-5.2 With 1M Context and MIT Weights, but Zero Benchmarks at Launch
- infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
- devtoolsVercel's Remend Turns Streaming-Markdown Repair Into a Dependency
- industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
- policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
- infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
- infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
- agentsDo Programming Languages Still Matter to Your AI Coding Agent?
- agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
- securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
- policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
- modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
- agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
- infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
- modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
- modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
- modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
- industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
- securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
- devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
- modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
- agentsCan AI Agents Share Context Without a Central Coordinator?
- agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
- infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
- modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
- infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
- devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
- infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
- agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
- modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
- modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
- policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
- industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
- modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
- securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
- policyWho Gets to Audit Your Health Chatbot? Almost No One
- policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
- infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
- industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
- policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
- policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
- agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
- agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
- cultureDoes Debate Quality Survive When LLMs Argue Outside English?
- securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
- agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
- policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
- industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
- devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
- ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
- policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
- policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
- securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
- securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
- securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
- agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
- industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
- industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
- policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
- agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
- agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
- cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
- infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
- policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
- industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
- infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
- securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
- securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
- cultureDo LLMs Understand Idioms in Low-Resource Languages?
- modelsCan LLMs Write Better Research Paper Titles Than Authors?
- modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
- infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
- modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
- agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
- securityVercel's Flags SDK Exposed Feature-Flag Definitions via CVE-2025-46332
- infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
- devtoolsAlibaba's Open Code Review Moves AI Review Into the CLI, Not the PR
- infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
- infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
- securityJailbreak Suffixes Hit Harder at Specific Token Positions, New GCG Variant Shows
- policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
- policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
- securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
- cultureCan Teaching Logical Fallacies Inoculate People Against AI Misinformation?
- devtoolsVercel Ships Experimental Native CLI Binaries to Cut the Node Startup Tax
- policyRefusal Steering Targets Individual Experts in MoE LLMs