
How Log-Probability Signals Catch LLM Hallucinations: What HALT Adds
HALT proposes using top-20 token log-probabilities as a time series to detect LLM hallucinations, offering a sequence-based alternative to single-score metrics for audit teams
A publication by Berry Mingus
Groundy is Berry Mingus's publication about AI and large language models, developer tools, infrastructure, and software culture.

HALT proposes using top-20 token log-probabilities as a time series to detect LLM hallucinations, offering a sequence-based alternative to single-score metrics for audit teams
Popular with Groundy readers.

MLX delivers 20-87% faster generation on Apple Silicon for models under 14B parameters. llama.cpp wins for cross-platform use and long contexts.

DeepSeek isn't China's only frontier AI. Compare DeepSeek, Qwen, Kimi, Doubao, and Ernie on benchmarks, licensing, API access, and use-case fit.

The EU's 2027 battery mandate is confirmed. Here's what 'user-replaceable' legally means, which phones comply now, and how to buy smart before the rules change.

Cursor hit $300M ARR in April 2025 by forking VS Code and baking AI into the editor's core. By June 2026 it was at $4B annualized and agreed to a $60B SpaceX acquisition. Here's how it happened and what it signals.

DataLearner's June 2026 snapshot ranks GLM-5.2 seventh by HLE at 54.70 and places no Chinese flagship in the overall top three, undercutting launch-day claims.

OpenAI's Stargate is a $500B joint venture to build U.S. AI data centers. Here's what's being built, who's paying, and what it means for compute markets.
Guides, comparisons and analysis, organized by topic.
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
The economics, interop standards, and workflow tradeoffs reshaping how code gets written, reviewed, and shipped when AI agents share the editor with the engineer.
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.

A preprint reports API evaluations score 3.4 points higher than chatbot interfaces, suggesting audits must test deployed products directly rather than relying on model metrics

A 2026 paper defines Semantic Confusion, where LLM refusals flip on meaning-preserving paraphrases. Audits must test consistency across clusters, not just aggregate refusal.

A rigor-matched audit finds layer skipping speeds LLM inference, but wall-clock rankings reverse when decision overhead is separated from pure generation cost.

A preprint argues RL alignment yields conditional compliance, suggesting governance shift from eval scores to architectural constraints and deployment monitoring.

A preprint argues GPU utilization misleads LLM capacity planning by conflating memory-bound decode with compute saturation, urging per-phase metrics.

A 2026 preprint finds LLMs generally track human legal reasonableness judgments but show homogeneity, stakeholder, and demographic skew, limiting autonomous use.

An arXiv preprint argues LLM metacognition is coarse and context-dependent, suggesting self-amendment requires external calibration and human approval gates.

Preprints show forced evidence gathering raised clinical agent error from 28.3% to 34.3%, while curated tool menus lifted ToolBench success to 0.898, suggesting gating is key.

FluxMoE streams MoE weights from host DRAM to save VRAM, trading capacity for bandwidth. Author-reported gains on multi-GPU setups lack independent replication.

A preprint shows prompted LLMs lag compact domain models in PET/CT report error detection, suggesting hospitals should prioritize specialized tools over general chatbots.

Deltafin streams Kimi K3 weights from SSDs to run on 64 GB Macs, but self-reported benchmarks vary widely, limiting it to batch workloads rather than interactive use.

A preprint reports that verbalized LLM confidence weakly tracks logprob signals, with instruction-tuned models showing worse calibration. Gate on verifiable internal signals.
Featured analysis and deeper reads.