
Can Text-to-SQL Agents Enforce Row-Level Security? A New RBAC Benchmark
A new arXiv benchmark shows text-to-SQL agents frequently violate RBAC policies. The paper concludes that authorization must remain in the warehouse, not the prompt, to ensure
The Groundy archive
Comparisons, research analysis and practical guides across AI, developer tools and infrastructure. Explore a topic or browse the complete archive.
1–24 of 797 articles · Newest first

A new arXiv benchmark shows text-to-SQL agents frequently violate RBAC policies. The paper concludes that authorization must remain in the warehouse, not the prompt, to ensure

Three 2026 preprints show LLM agents often report success despite incomplete coverage or stale state, requiring artifact-level verification gates to catch silent failures.

Vercel Sandbox Drives offer persistent agent storage with a 16 TiB cap, but beta constraints include single-writer limits and region pinning that affect multi-agent pipeline.

A 2026 preprint shows fixed-configuration kernels achieve cross-GPU bitwise determinism for linear layers, challenging the assumption that reproducibility requires a major end

WordPress 7.1.2 patches an unauthenticated path traversal to RCE requiring specific theme and PHP conditions. The advisory names affected environments, but independent testing

A preprint reports a 39.7% relative WER reduction for police audio, but uneven errors and lack of verification protocols mean transcripts require human audio checks.

Evidence supports fixing MCP deployments with a five-tool budget and migration plans, not dropping the protocol, as tool sprawl degrades agent performance.

Vercel reports a libheif AVIF RCE affecting Next.js, sharp, and WordPress. Teams must patch libheif to v1.23.4, as platform mitigations do not cover self-hosted or direct use.

Cooley's GO Public uses a review-gated workflow on ChatGPT Work. Vendor claims lack independent verification, so firms should build harnesses first and gate confidential data.

Forensic analysis shows ZCode silently uploads encrypted Git history to Aliyun OSS without user consent or a working opt-out, requiring filesystem-level containment.

GitHub Copilot auto tiers and Cursor Router differ in model control. Pin named models for reproducibility; use auto for interactive work where model identity is visible.

AIREP argues AI governance needs four distinct runtime records per decision, not one audit event, to support incident reconstruction and dispute resolution.

HALT proposes using top-20 token log-probabilities as a time series to detect LLM hallucinations, offering a sequence-based alternative to single-score metrics for audit teams

An arXiv study of six registries counts cross-ecosystem packages like npm and PyPI pairs, but its lifetime GitHub totals cannot show which registry copy is current or safe.

EvoUndo preprint argues coding agents need provable undo limits for self-rewrites, as 197 improving mutations failed recovery checks in author-reported tests.

A preprint claims composing specialist capabilities into one small model improves accuracy and cuts tokens, but results are author-reported and unreplicated.

A preprint reports that calibrated learned routing for disaggregated LLM serving beats heuristics by 1.7 to 2.9 goodput points, but only on specific hardware with large pools.

A preprint reports API evaluations score 3.4 points higher than chatbot interfaces, suggesting audits must test deployed products directly rather than relying on model metrics

A preprint reports LLM inference energy varies up to 179x by language, with Pashto costing far more than English. These author-reported findings suggest locale mix is a key, 1

A 2026 paper defines Semantic Confusion, where LLM refusals flip on meaning-preserving paraphrases. Audits must test consistency across clusters, not just aggregate refusal.

Author-reported DECODE data shows most AI code edits occur within 15 minutes of acceptance, with 31% of trajectories ending in removal, suggesting review should focus on the 1

MCPAgentBench shows LLM agents excel at tool selection but fail strict execution order, suggesting teams should route agents by measured MCP skill rather than general chat.

DuckDB research extension FlockMTL integrates LLMs and RAG into SQL via schema objects, moving governance to DDL, though performance claims remain author-reported.

A rigor-matched audit finds layer skipping speeds LLM inference, but wall-clock rankings reverse when decision overhead is separated from pure generation cost.
Search the full archive for a tool, model, framework or question.