
How to Evaluate an LLM for Mental Health Conversations: MentalHealthBench
OpenAI's MentalHealthBench offers a rubric-based alternative to refusal audits, though vendor-reported scores and grader bias limit independent verification.
A publication by Berry Mingus
Groundy is Berry Mingus's publication about AI and large language models, developer tools, infrastructure, and software culture.

OpenAI's MentalHealthBench offers a rubric-based alternative to refusal audits, though vendor-reported scores and grader bias limit independent verification.
Popular with Groundy readers.

MLX and llama.cpp both run quantized LLMs on Apple Silicon unified memory. Here is what each project documents, and how to measure which wins on your Mac.

Compare DeepSeek, Qwen, Kimi, Doubao, ERNIE and GLM through dated benchmarks, license terms, context windows, API pricing and practical workload fit.

DataLearner's June 2026 snapshot ranks GLM-5.2 seventh by HLE at 54.70 and places no Chinese flagship in the overall top three, undercutting launch-day claims.

The EU's 2027 battery mandate is confirmed. Here's what 'user-replaceable' legally means, which phones comply now, and how to buy smart before the rules change.

F-Droid, the open-source Android app repository, is leading a global campaign against Google's mandatory developer verification program, a policy set to take effect in September 2026 that critics say will end alternative app distribution and hand Google total control over what software can run on Android devices.

Cursor hit $300M ARR in April 2025 by forking VS Code and baking AI into the editor's core. By June 2026 it was at $4B annualized and agreed to a $60B SpaceX acquisition. Here's how it happened and what it signals.
Guides, comparisons and analysis, organized by topic.
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
The economics, interop standards, and workflow tradeoffs reshaping how code gets written, reviewed, and shipped when AI agents share the editor with the engineer.

WordPress 7.1.2 patches an unauthenticated path traversal to RCE requiring specific theme and PHP conditions. The advisory names affected environments, but independent testing

A preprint reports a 39.7% relative WER reduction for police audio, but uneven errors and lack of verification protocols mean transcripts require human audio checks.

Evidence supports fixing MCP deployments with a five-tool budget and migration plans, not dropping the protocol, as tool sprawl degrades agent performance.

Vercel reports a libheif AVIF RCE affecting Next.js, sharp, and WordPress. Teams must patch libheif to v1.23.4, as platform mitigations do not cover self-hosted or direct use.

Cooley's GO Public uses a review-gated workflow on ChatGPT Work. Vendor claims lack independent verification, so firms should build harnesses first and gate confidential data.

Forensic analysis shows ZCode silently uploads encrypted Git history to Aliyun OSS without user consent or a working opt-out, requiring filesystem-level containment.

AIREP argues AI governance needs four distinct runtime records per decision, not one audit event, to support incident reconstruction and dispute resolution.

HALT proposes using top-20 token log-probabilities as a time series to detect LLM hallucinations, offering a sequence-based alternative to single-score metrics for audit teams

EvoUndo preprint argues coding agents need provable undo limits for self-rewrites, as 197 improving mutations failed recovery checks in author-reported tests.

A preprint claims composing specialist capabilities into one small model improves accuracy and cuts tokens, but results are author-reported and unreplicated.

A preprint reports API evaluations score 3.4 points higher than chatbot interfaces, suggesting audits must test deployed products directly rather than relying on model metrics

A preprint reports LLM inference energy varies up to 179x by language, with Pashto costing far more than English. These author-reported findings suggest locale mix is a key, 1
Featured analysis and deeper reads.