
Why the Same LLM Prompt Returns Different Outputs Across GPUs
A 2026 preprint shows fixed-configuration kernels achieve cross-GPU bitwise determinism for linear layers, challenging the assumption that reproducibility requires a major end
A publication by Berry Mingus
Groundy is Berry Mingus's publication about AI and large language models, developer tools, infrastructure, and software culture.

A 2026 preprint shows fixed-configuration kernels achieve cross-GPU bitwise determinism for linear layers, challenging the assumption that reproducibility requires a major end
Popular with Groundy readers.

MLX and llama.cpp both run quantized LLMs on Apple Silicon unified memory. Here is what each project documents, and how to measure which wins on your Mac.

Compare DeepSeek, Qwen, Kimi, Doubao, ERNIE and GLM through dated benchmarks, license terms, context windows, API pricing and practical workload fit.

The EU's 2027 battery mandate is confirmed. Here's what 'user-replaceable' legally means, which phones comply now, and how to buy smart before the rules change.

DataLearner's June 2026 snapshot ranks GLM-5.2 seventh by HLE at 54.70 and places no Chinese flagship in the overall top three, undercutting launch-day claims.

Chinese models hit 41% of Hugging Face downloads, overtaking the US, while independents hit 39%. Top 200 models capture half of all downloads, forcing Western procurement.

A used V100 looks like cheap VRAM for local inference, but no bf16, no FlashAttention, and CUDA 13 deprecation lock buyers into a software stack that is actively contracting.
Guides, comparisons and analysis, organized by topic.
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
The economics, interop standards, and workflow tradeoffs reshaping how code gets written, reviewed, and shipped when AI agents share the editor with the engineer.
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.

Vercel reports a libheif AVIF RCE affecting Next.js, sharp, and WordPress. Teams must patch libheif to v1.23.4, as platform mitigations do not cover self-hosted or direct use.

Cooley's GO Public uses a review-gated workflow on ChatGPT Work. Vendor claims lack independent verification, so firms should build harnesses first and gate confidential data.

AIREP argues AI governance needs four distinct runtime records per decision, not one audit event, to support incident reconstruction and dispute resolution.

HALT proposes using top-20 token log-probabilities as a time series to detect LLM hallucinations, offering a sequence-based alternative to single-score metrics for audit teams

A preprint claims composing specialist capabilities into one small model improves accuracy and cuts tokens, but results are author-reported and unreplicated.

A preprint reports API evaluations score 3.4 points higher than chatbot interfaces, suggesting audits must test deployed products directly rather than relying on model metrics

A 2026 paper defines Semantic Confusion, where LLM refusals flip on meaning-preserving paraphrases. Audits must test consistency across clusters, not just aggregate refusal.

MCPAgentBench shows LLM agents excel at tool selection but fail strict execution order, suggesting teams should route agents by measured MCP skill rather than general chat.

A rigor-matched audit finds layer skipping speeds LLM inference, but wall-clock rankings reverse when decision overhead is separated from pure generation cost.

A 441-repo preprint links committed AI config to lower defect costs, though authors note it is correlational and hypothesis-generating rather than proven causality.

A preprint argues RL alignment yields conditional compliance, suggesting governance shift from eval scores to architectural constraints and deployment monitoring.

A preprint argues GPU utilization misleads LLM capacity planning by conflating memory-bound decode with compute saturation, urging per-phase metrics.
Featured analysis and deeper reads.