
Does Committing an AGENTS.md File Lower AI Coding Defect Costs?
A 441-repo preprint links committed AI config to lower defect costs, though authors note it is correlational and hypothesis-generating rather than proven causality.
The Groundy archive · Page 2 of 34
Browse Groundy's complete archive of 797 articles on AI, developer tools and infrastructure. Page 2 of 34.
25–48 of 797 articles · Newest first

A 441-repo preprint links committed AI config to lower defect costs, though authors note it is correlational and hypothesis-generating rather than proven causality.

A preprint argues RL alignment yields conditional compliance, suggesting governance shift from eval scores to architectural constraints and deployment monitoring.

A preprint argues GPU utilization misleads LLM capacity planning by conflating memory-bound decode with compute saturation, urging per-phase metrics.

A 2026 preprint finds LLMs generally track human legal reasonableness judgments but show homogeneity, stakeholder, and demographic skew, limiting autonomous use.

An arXiv preprint argues LLM metacognition is coarse and context-dependent, suggesting self-amendment requires external calibration and human approval gates.

A preprint argues zero-shot accuracy misses distribution drift in quantized LLMs, recommending divergence metrics like JSD and TV against BF16 bases for safer deployment.

Preprints show forced evidence gathering raised clinical agent error from 28.3% to 34.3%, while curated tool menus lifted ToolBench success to 0.898, suggesting gating is key.

FluxMoE streams MoE weights from host DRAM to save VRAM, trading capacity for bandwidth. Author-reported gains on multi-GPU setups lack independent replication.

A preprint shows prompted LLMs lag compact domain models in PET/CT report error detection, suggesting hospitals should prioritize specialized tools over general chatbots.

Deltafin streams Kimi K3 weights from SSDs to run on 64 GB Macs, but self-reported benchmarks vary widely, limiting it to batch workloads rather than interactive use.

A preprint reports that verbalized LLM confidence weakly tracks logprob signals, with instruction-tuned models showing worse calibration. Gate on verifiable internal signals.

vLLM's ROCm speculative decoding speedup is self-reported and unreplicated. AMD's gigawatt deals de-risk the platform, but operators must benchmark acceptance rates on MI300X.

A Hacker News claim that 9 in 10 European CDN users rely on Cloudflare lacks independent verification. Operators must audit failover paths to avoid correlated failure and meet

JetBrains' Mellum2 ships as Apache 2.0 weights with 12B total and 2.5B active parameters; here is what the release settles before you self-host.

Mistral's €3B Series D buys staying power, not proof of quality or compliance, so compare its regional API against a customer-controlled deployment before committing.

IndicSafeEval shows English refusal rates do not transfer to Hindi, Bengali, Marathi, or Punjabi. Teams need native-language persuasive probes and per-category baselines for a

SPD claims 28 ms listwise LLM reranking via single-pass decoding, but the evidence is one unreplicated preprint. Keep cross-encoders in production until independent benchmarks

CUA-Universe preprint shows hybrid GUI+CLI agents cut steps 37% and tokens 60% on 16 apps. Audit your eval stack: screenshot-only benchmarks mismeasure tasks with terminal.

A new preprint shows LLMs perform hidden computation invisible in chain-of-thought. This breaks audit assumptions, forcing a shift from transcript review to behavioral evals.

Cloudflare claims MCP traffic classification, but stdio servers remain invisible to edge proxies. Build a filter policy that gates remote endpoints and treats local tool on as

An author-reported 44% on ARC-AGI-1 public eval for about $0.67 of rented compute shows what small budgets can justify, and what still needs replication.

A new NCCL shim recovers 13-38% bandwidth on shared GPU clusters by tuning collective patterns. Test for cross-tenant interference before buying more fabric.

ChatGPT ads target Free and Go users only, leaving paid subscribers untouched. This split redefines AI search economics, forcing teams to treat organic citations as the new mo

FP8 and MXFP4 are umbrella specs, not single formats. A new preprint offers bit-exact conformance vectors to test quantized LLM portability across GPUs, exposing hidden format