
How to Evaluate LLM Long-Term Memory: Repeated Judging, Negative Controls
A September 2026 preprint argues that LLM memory benchmarks require repeated judging, negative controls, and ablation arms to distinguish real gains from judge noise.
A publication by Berry Mingus
Groundy is Berry Mingus's publication about AI and large language models, developer tools, infrastructure, and software culture.

A September 2026 preprint argues that LLM memory benchmarks require repeated judging, negative controls, and ablation arms to distinguish real gains from judge noise.
Popular with Groundy readers.

MLX and llama.cpp both run quantized LLMs on Apple Silicon unified memory. Here is what each project documents, and how to measure which wins on your Mac.

The EU's 2027 battery mandate is confirmed. Here's what 'user-replaceable' legally means, which phones comply now, and how to buy smart before the rules change.

Compare DeepSeek, Qwen, Kimi, Doubao, ERNIE and GLM through dated benchmarks, license terms, context windows, API pricing and practical workload fit.

Self-improvement claims require held-out benchmarks, lineage logs, and calibrated judges to distinguish genuine capability gains from memorized feedback in agent loops.

DataLearner's June 2026 snapshot ranks GLM-5.2 seventh by HLE at 54.70 and places no Chinese flagship in the overall top three, undercutting launch-day claims.

A hosted instance saves you some setup work, while those choices remain yours.
Guides, comparisons and analysis, organized by topic.
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
The economics, interop standards, and workflow tradeoffs reshaping how code gets written, reviewed, and shipped when AI agents share the editor with the engineer.

EuroLLM and Apertus are distinct sovereign open models with different licenses and coverage. Cloudflare's hosting is vendor-reported intent, not shipped capability.

A preprint proposes canary tokens to track prompt injection across agent stages, showing that outcome-only metrics hide distinct failure modes in LLM pipelines.

A 2026 preprint reports AI homogenizes code syntax but not intent, while algorithmic monoculture models show bounded welfare harm, requiring distinct policy levers.

An 815-participant experiment found anthropomorphic AI verbs barely shift perception, while explicit danger framing does. Style guides should prioritize risk accuracy over ban

A new preprint proposes pricing LLM training data by measured gain rather than flat token rates, offering a hybrid contract structure for buyers and sellers.

MetaPermit proposes LLM-inferred attributes for auditable agent access control, reporting 31% consistency gains in a self-reported preprint.

A preprint proposes Counterfactual Rollout Replay to derive step-level rewards for coding agents by forking environments, trading labeling costs for compute without learned PR

Research shows MCP error wording shifts agent recovery from 6% to 88%. Ship stable codes, offending parameters, and one named action to enable deterministic recovery.

TraceML pairs human and agent Kaggle logs to show agents collapse into narrow loops, supporting trajectory audits over leaderboard scores for autonomy decisions.

Pooled faithfulness checkers miss cross-source conflation in MCP agents. ProvenanceGuard adds per-claim source verification, though author-reported results remain unreplicated

Cloudflare Forge is a new open-source SDK pipeline for AI agents. Evidence is limited to vendor claims, so pilot one artifact before migrating from incumbent generators.

Cloudflare's cf CLI expands agent access to 3,000 operations. Grant read-only scopes by default and gate mutations behind human checkpoints to limit blast radius.
Featured analysis and deeper reads.