<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Groundy · Models &amp; Research</title><description>Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.</description><link>https://groundy.com/</link><language>en-us</language><ttl>60</ttl><lastBuildDate>Mon, 14 Sep 2026 04:24:51 GMT</lastBuildDate><atom:link href="https://groundy.com/category/models-research/rss.xml" rel="self" type="application/rss+xml"/><atom:link href="https://groundy.com/feeds/" rel="alternate" type="text/html"/><image><url>https://groundy.com/rss-icon.png</url><title>Groundy · Models &amp; Research</title><link>https://groundy.com/</link><width>144</width><height>144</height></image><item><title>Can You Trust LLM Confidence Scores? Verbalized vs Logprob Signals</title><link>https://groundy.com/articles/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/</guid><description>A preprint reports that verbalized LLM confidence weakly tracks logprob signals, with instruction-tuned models showing worse calibration. Gate on verifiable internal signals.</description><pubDate>Wed, 09 Sep 2026 04:05:47 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;&lt;img src=&quot;https://groundy.com/feed-images/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/image.jpg?v=aa87e086763c7b7e&quot; width=&quot;960&quot; height=&quot;540&quot; alt=&quot;A copper speaking horn projects toward a glass-covered precision gauge whose internal mechanism and needle disagree, with calibration weights below.&quot; style=&quot;max-width:100%;height:auto&quot; /&gt;&lt;/a&gt;&lt;/p&gt;&lt;p&gt;A preprint reports that verbalized LLM confidence weakly tracks logprob signals, with instruction-tuned models showing worse calibration. Gate on verifiable internal signals.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 13 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><media:content url="https://groundy.com/feed-images/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/image.jpg?v=aa87e086763c7b7e" type="image/jpeg" medium="image" width="960" height="540" fileSize="109781"><media:description type="plain">A copper speaking horn projects toward a glass-covered precision gauge whose internal mechanism and needle disagree, with calibration weights below.</media:description></media:content><media:thumbnail url="https://groundy.com/feed-images/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/thumb.jpg?v=aa87e086763c7b7e" width="640" height="360"/><category>Models &amp; Research</category><category>llm-confidence</category><category>logprob-vs-verbalized</category><category>model-calibration</category><category>instruction-tuning</category><category>reliability-gating</category><enclosure url="https://groundy.com/feed-images/can-you-trust-llm-confidence-scores-verbalized-vs-logprob-signals/image.jpg?v=aa87e086763c7b7e" length="109781" type="image/jpeg"/></item><item><title>Generative vs Cross-Encoder LLM Rerankers: Can Single-Pass Decoding Close the Gap?</title><link>https://groundy.com/articles/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/</guid><description>SPD claims 28 ms listwise LLM reranking via single-pass decoding, but the evidence is one unreplicated preprint. Keep cross-encoders in production until independent benchmarks</description><pubDate>Tue, 08 Sep 2026 08:09:56 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;&lt;img src=&quot;https://groundy.com/feed-images/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/image.jpg?v=79f399ffb795dcd4&quot; width=&quot;960&quot; height=&quot;540&quot; alt=&quot;A global brass sorting drum orders blank document cards in one pass beside a smaller comparator that cycles cards repeatedly.&quot; style=&quot;max-width:100%;height:auto&quot; /&gt;&lt;/a&gt;&lt;/p&gt;&lt;p&gt;SPD claims 28 ms listwise LLM reranking via single-pass decoding, but the evidence is one unreplicated preprint. Keep cross-encoders in production until independent benchmarks&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><media:content url="https://groundy.com/feed-images/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/image.jpg?v=79f399ffb795dcd4" type="image/jpeg" medium="image" width="960" height="540" fileSize="139525"><media:description type="plain">A global brass sorting drum orders blank document cards in one pass beside a smaller comparator that cycles cards repeatedly.</media:description></media:content><media:thumbnail url="https://groundy.com/feed-images/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/thumb.jpg?v=79f399ffb795dcd4" width="640" height="360"/><category>Models &amp; Research</category><category>rag</category><category>reranking</category><category>llm-inference</category><category>arxiv</category><category>cross-encoder</category><category>generative-ai</category><enclosure url="https://groundy.com/feed-images/generative-vs-cross-encoder-llm-rerankers-can-single-pass-decoding-close-the-gap/image.jpg?v=79f399ffb795dcd4" length="139525" type="image/jpeg"/></item><item><title>mdlARC&apos;s 44% ARC-AGI-1 Claim: What the Small Budget Leaves Out</title><link>https://groundy.com/articles/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/</guid><description>An author-reported 44% on ARC-AGI-1 public eval for about $0.67 of rented compute shows what small budgets can justify, and what still needs replication.</description><pubDate>Mon, 07 Sep 2026 08:37:22 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;&lt;img src=&quot;https://groundy.com/feed-images/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/image.jpg?v=fdbe31d690370b2b&quot; width=&quot;960&quot; height=&quot;540&quot; alt=&quot;A compact copper furnace makes a mosaic key for one abstract lock while enormous general engines stand idle behind it.&quot; style=&quot;max-width:100%;height:auto&quot; /&gt;&lt;/a&gt;&lt;/p&gt;&lt;p&gt;An author-reported 44% on ARC-AGI-1 public eval for about $0.67 of rented compute shows what small budgets can justify, and what still needs replication.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 7 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><atom:updated>2026-09-08T00:00:00.000Z</atom:updated><media:content url="https://groundy.com/feed-images/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/image.jpg?v=fdbe31d690370b2b" type="image/jpeg" medium="image" width="960" height="540" fileSize="131057"><media:description type="plain">A compact copper furnace makes a mosaic key for one abstract lock while enormous general engines stand idle behind it.</media:description></media:content><media:thumbnail url="https://groundy.com/feed-images/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/thumb.jpg?v=fdbe31d690370b2b" width="640" height="360"/><category>Models &amp; Research</category><category>test-time-training</category><category>arc-agi</category><category>llm-evaluation</category><category>cost-optimization</category><category>small-models</category><category>inference-cost</category><enclosure url="https://groundy.com/feed-images/a-1-5-hour-transformer-beats-frontier-llms-on-arc-agi-when-to-train-small/image.jpg?v=fdbe31d690370b2b" length="131057" type="image/jpeg"/></item><item><title>FP8 vs MXFP4 vs BF16: Why Your Quantized LLM Disagrees Across GPUs</title><link>https://groundy.com/articles/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/</guid><description>FP8 and MXFP4 are umbrella specs, not single formats. A new preprint offers bit-exact conformance vectors to test quantized LLM portability across GPUs, exposing hidden format</description><pubDate>Mon, 07 Sep 2026 07:15:12 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;&lt;img src=&quot;https://groundy.com/feed-images/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/image.jpg?v=9ac8aca899d82bb1&quot; width=&quot;960&quot; height=&quot;540&quot; alt=&quot;Three presses shape identical copper blanks into subtly different gears beside one shared brass reference gauge.&quot; style=&quot;max-width:100%;height:auto&quot; /&gt;&lt;/a&gt;&lt;/p&gt;&lt;p&gt;FP8 and MXFP4 are umbrella specs, not single formats. A new preprint offers bit-exact conformance vectors to test quantized LLM portability across GPUs, exposing hidden format&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 13 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><media:content url="https://groundy.com/feed-images/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/image.jpg?v=9ac8aca899d82bb1" type="image/jpeg" medium="image" width="960" height="540" fileSize="141208"><media:description type="plain">Three presses shape identical copper blanks into subtly different gears beside one shared brass reference gauge.</media:description></media:content><media:thumbnail url="https://groundy.com/feed-images/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/thumb.jpg?v=9ac8aca899d82bb1" width="640" height="360"/><category>Models &amp; Research</category><category>quantization</category><category>fp8</category><category>mxfp4</category><category>gpu-compatibility</category><category>llm-inference</category><category>conformance-testing</category><enclosure url="https://groundy.com/feed-images/fp8-vs-mxfp4-vs-bf16-why-your-quantized-llm-disagrees-across-gpus/image.jpg?v=9ac8aca899d82bb1" length="141208" type="image/jpeg"/></item><item><title>Running a 104GB LLM on a 48GB Mac: What Expert Streaming Costs</title><link>https://groundy.com/articles/running-a-104gb-llm-on-a-48gb-mac-what-expert-streaming-costs/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/running-a-104gb-llm-on-a-48gb-mac-what-expert-streaming-costs/</guid><description>Expert streaming lets 48GB Macs run 104GB MoE models by paging weights from SSD. This guide compares streaming to quantization and offload, analyzing workload fit and hardware</description><pubDate>Sun, 06 Sep 2026 06:56:19 GMT</pubDate><content:encoded>&lt;p&gt;Expert streaming lets 48GB Macs run 104GB MoE models by paging weights from SSD. This guide compares streaming to quantization and offload, analyzing workload fit and hardware&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/running-a-104gb-llm-on-a-48gb-mac-what-expert-streaming-costs/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>mixture-of-experts</category><category>local-llm-inference</category><category>apple-silicon</category><category>llama-cpp</category><category>model-optimization</category><category>storage-bandwidth</category></item><item><title>GPT-6 Astra on ARC-AGI-3: What the Agentic Score Actually Measures</title><link>https://groundy.com/articles/gpt-6-astra-on-arc-agi-3-what-the-agentic-score-actually-measures/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/gpt-6-astra-on-arc-agi-3-what-the-agentic-score-actually-measures/</guid><description>GPT-6 Astra&apos;s ARC-AGI-3 scores vary by 97 points based on effort settings. Route agents by cost per solved task, not launch-day leaderboards.</description><pubDate>Sat, 05 Sep 2026 19:22:46 GMT</pubDate><content:encoded>&lt;p&gt;GPT-6 Astra&amp;apos;s ARC-AGI-3 scores vary by 97 points based on effort settings. Route agents by cost per solved task, not launch-day leaderboards.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 14 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/gpt-6-astra-on-arc-agi-3-what-the-agentic-score-actually-measures/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>arc-agi-3</category><category>gpt-6-astra</category><category>agentic-evaluation</category><category>model-selection</category><category>benchmark-interpretation</category><category>reasoning-effort</category></item><item><title>Deepfake KYC Fraud: Tamper-Resilient Watermarks That Recover the Original Face</title><link>https://groundy.com/articles/deepfake-kyc-fraud-tamper-resilient-watermarks-that-recover-the-original-face/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/deepfake-kyc-fraud-tamper-resilient-watermarks-that-recover-the-original-face/</guid><description>VeriFi preprint claims watermarks recover original faces after tampering, not just flag fakes. Robustness is untested against commercial tools and compression, so treat as a R</description><pubDate>Mon, 31 Aug 2026 15:47:10 GMT</pubDate><content:encoded>&lt;p&gt;VeriFi preprint claims watermarks recover original faces after tampering, not just flag fakes. Robustness is untested against commercial tools and compression, so treat as a R&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/deepfake-kyc-fraud-tamper-resilient-watermarks-that-recover-the-original-face/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>deepfake-detection</category><category>kyc-security</category><category>image-watermarking</category><category>face-recovery</category><category>media-provenance</category><category>arxiv-preprint</category></item><item><title>Detecting AI-Generated Audio: Why Decay Tails Betray Voice Clones</title><link>https://groundy.com/articles/detecting-ai-generated-audio-why-decay-tails-betray-voice-clones/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/detecting-ai-generated-audio-why-decay-tails-betray-voice-clones/</guid><description>A new preprint shows AI audio leaks in decay tails via group delay. It offers a cheap, watermark-free filter for fraud screening, though hold-out accuracy is only 66.7%.</description><pubDate>Mon, 31 Aug 2026 04:14:43 GMT</pubDate><content:encoded>&lt;p&gt;A new preprint shows AI audio leaks in decay tails via group delay. It offers a cheap, watermark-free filter for fraud screening, though hold-out accuracy is only 66.7%.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/detecting-ai-generated-audio-why-decay-tails-betray-voice-clones/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>audio-forensics</category><category>ai-detection</category><category>signal-processing</category><category>group-delay</category><category>voice-cloning</category><category>deepfake-detection</category></item><item><title>Reward Hacking Starts in the Verifier: Rule Checks vs LLM Judges for Math RL</title><link>https://groundy.com/articles/reward-hacking-starts-in-the-verifier-rule-checks-vs-llm-judges-for-math/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/reward-hacking-starts-in-the-verifier-rule-checks-vs-llm-judges-for-math/</guid><description>Rule checkers miss format variants while LLM judges get hacked during RL. This failure map from arXiv 2505.22203 shows why hybrid verifiers and reward audits are now essential</description><pubDate>Sun, 30 Aug 2026 16:01:22 GMT</pubDate><content:encoded>&lt;p&gt;Rule checkers miss format variants while LLM judges get hacked during RL. This failure map from arXiv 2505.22203 shows why hybrid verifiers and reward audits are now essential&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 15 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/reward-hacking-starts-in-the-verifier-rule-checks-vs-llm-judges-for-math/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>rlvr</category><category>reward-hacking</category><category>verifier-design</category><category>math-reasoning</category><category>reinforcement-learning</category><category>llm-judges</category></item><item><title>Can LLMs Train on Their Own Problems? What Zero-Data Self-Play Changes</title><link>https://groundy.com/articles/can-llms-train-on-their-own-problems-what-zero-data-self-play-changes/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-llms-train-on-their-own-problems-what-zero-data-self-play-changes/</guid><description>J-Zero claims zero-data self-play beats baselines by 8.0 points on unverifiable tasks, but judge verdict flips of 5.3 to 48.4 percent suggest the gains may be self-validated.</description><pubDate>Sun, 30 Aug 2026 03:19:29 GMT</pubDate><content:encoded>&lt;p&gt;J-Zero claims zero-data self-play beats baselines by 8.0 points on unverifiable tasks, but judge verdict flips of 5.3 to 48.4 percent suggest the gains may be self-validated.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-llms-train-on-their-own-problems-what-zero-data-self-play-changes/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-self-play</category><category>reinforcement-learning</category><category>judge-reliability</category><category>post-training</category><category>arxiv-preprint</category><category>verifier-audit</category></item><item><title>Why LLM Agent Benchmarks Move When the Harness Changes</title><link>https://groundy.com/articles/why-llm-agent-benchmarks-move-when-the-harness-changes/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/why-llm-agent-benchmarks-move-when-the-harness-changes/</guid><description>A new preprint argues agent benchmark scores track the evaluation harness as much as the model. Teams should pin harness versions and ablate scaffold changes to avoid misat.</description><pubDate>Sat, 29 Aug 2026 15:09:10 GMT</pubDate><content:encoded>&lt;p&gt;A new preprint argues agent benchmark scores track the evaluation harness as much as the model. Teams should pin harness versions and ablate scaffold changes to avoid misat.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/why-llm-agent-benchmarks-move-when-the-harness-changes/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-agents</category><category>benchmarking</category><category>evaluation-protocols</category><category>harness-evolution</category><category>model-research</category><category>agent-evals</category></item><item><title>Can You Serve LLMs on 2-Bit Weights? What Ultra-Low-Bit Quantization Costs</title><link>https://groundy.com/articles/can-you-serve-llms-on-2-bit-weights-what-ultra-low-bit-quantization-costs/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-you-serve-llms-on-2-bit-weights-what-ultra-low-bit-quantization-costs/</guid><description>arXiv 2508.06753 reports 2-bit LLM serving with up to 7x speedups on Intel Xe2. For self-hosters, the binding cost is not memory but task-specific eval coverage to verify QAT-</description><pubDate>Fri, 28 Aug 2026 16:09:46 GMT</pubDate><content:encoded>&lt;p&gt;arXiv 2508.06753 reports 2-bit LLM serving with up to 7x speedups on Intel Xe2. For self-hosters, the binding cost is not memory but task-specific eval coverage to verify QAT-&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-you-serve-llms-on-2-bit-weights-what-ultra-low-bit-quantization-costs/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-quantization</category><category>int2-weights</category><category>vllm</category><category>self-hosting</category><category>inference-optimization</category><category>intel-xe2</category></item><item><title>Do LLMs Still Need BPE? What RL-Trained Tokenizers Change</title><link>https://groundy.com/articles/do-llms-still-need-bpe-what-rl-trained-tokenizers-change/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/do-llms-still-need-bpe-what-rl-trained-tokenizers-change/</guid><description>An ICML 2026 paper shows RL can learn token boundaries end-to-end, beating straight-through baselines at 100M parameters. For serving or fine-tuning, keep your frozen BPE or S</description><pubDate>Thu, 27 Aug 2026 15:26:53 GMT</pubDate><content:encoded>&lt;p&gt;An ICML 2026 paper shows RL can learn token boundaries end-to-end, beating straight-through baselines at 100M parameters. For serving or fine-tuning, keep your frozen BPE or S&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/do-llms-still-need-bpe-what-rl-trained-tokenizers-change/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>tokenization</category><category>reinforcement-learning</category><category>llm-pretraining</category><category>bpe</category><category>sentencepiece</category><category>model-architecture</category></item><item><title>GLM-5.3-Flash vs Qwen3.8-Flash-Next: Which Budget LLM to Route To</title><link>https://groundy.com/articles/glm-5-3-flash-vs-qwen3-8-flash-next-which-budget-llm-to-route/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/glm-5-3-flash-vs-qwen3-8-flash-next-which-budget-llm-to-route/</guid><description>GLM-5.3-Flash and Qwen3.8-Flash-Next lack confirmed pricing and independent evals. Hold your router, verify per-token costs, and run a 50-case domain holdout before switching.</description><pubDate>Thu, 27 Aug 2026 14:27:49 GMT</pubDate><content:encoded>&lt;p&gt;GLM-5.3-Flash and Qwen3.8-Flash-Next lack confirmed pricing and independent evals. Hold your router, verify per-token costs, and run a 50-case domain holdout before switching.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/glm-5-3-flash-vs-qwen3-8-flash-next-which-budget-llm-to-route/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-routing</category><category>agent-economics</category><category>glm-5-3</category><category>qwen3</category><category>cost-optimization</category><category>model-evaluation</category></item><item><title>Why Prompt Caching Can Change Model Outputs: A Prefix Invariance Audit</title><link>https://groundy.com/articles/why-prompt-caching-can-change-model-outputs-a-prefix-invariance-audit/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/why-prompt-caching-can-change-model-outputs-a-prefix-invariance-audit/</guid><description>A new arXiv audit shows attention masks miss causality leaks in state-space models. Run a two-pass check before trusting prefix caches on hybrid stacks.</description><pubDate>Wed, 26 Aug 2026 13:34:03 GMT</pubDate><content:encoded>&lt;p&gt;A new arXiv audit shows attention masks miss causality leaks in state-space models. Run a two-pass check before trusting prefix caches on hybrid stacks.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 15 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/why-prompt-caching-can-change-model-outputs-a-prefix-invariance-audit/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>prefix-caching</category><category>llm-serving</category><category>state-space-models</category><category>model-auditing</category><category>inference-optimization</category></item><item><title>GDPR Deletion Requests vs LLM Weights: What Machine Unlearning Actually Removes</title><link>https://groundy.com/articles/gdpr-deletion-requests-vs-llm-weights-what-machine-unlearning-actually-removes/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/gdpr-deletion-requests-vs-llm-weights-what-machine-unlearning-actually-removes/</guid><description>GDPR erasure demands hit LLM weights, but retraining is the only defensible fix. Behavior-level unlearning passes probes without proving data is gone, leaving compliance teams</description><pubDate>Tue, 25 Aug 2026 13:56:47 GMT</pubDate><content:encoded>&lt;p&gt;GDPR erasure demands hit LLM weights, but retraining is the only defensible fix. Behavior-level unlearning passes probes without proving data is gone, leaving compliance teams&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/gdpr-deletion-requests-vs-llm-weights-what-machine-unlearning-actually-removes/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>gdpr-compliance</category><category>machine-unlearning</category><category>llm-privacy</category><category>data-erasure</category><category>model-retraining</category><category>ai-governance</category></item><item><title>Why Drift Monitors Confuse Covariate Shift With Concept Drift</title><link>https://groundy.com/articles/why-drift-monitors-confuse-covariate-shift-with-concept-drift/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/why-drift-monitors-confuse-covariate-shift-with-concept-drift/</guid><description>A new preprint proposes CJSD, a two-discriminator test that separates covariate shift from concept drift. This distinction determines whether to retrain or recalibrate, fixing</description><pubDate>Tue, 25 Aug 2026 02:06:11 GMT</pubDate><content:encoded>&lt;p&gt;A new preprint proposes CJSD, a two-discriminator test that separates covariate shift from concept drift. This distinction determines whether to retrain or recalibrate, fixing&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 13 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/why-drift-monitors-confuse-covariate-shift-with-concept-drift/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>drift-detection</category><category>covariate-shift</category><category>concept-drift</category><category>model-monitoring</category><category>machine-learning</category><category>preprint-analysis</category></item><item><title>AI-Generated Apps Look Right, but Do They Actually Work?</title><link>https://groundy.com/articles/ai-generated-apps-look-right-but-do-they-actually-work/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/ai-generated-apps-look-right-but-do-they-actually-work/</guid><description>MobileForge shows AI apps compile but fail navigation. Shift review from screenshots to project-level state tests to ship reliable generated frontends.</description><pubDate>Mon, 24 Aug 2026 14:09:46 GMT</pubDate><content:encoded>&lt;p&gt;MobileForge shows AI apps compile but fail navigation. Shift review from screenshots to project-level state tests to ship reliable generated frontends.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 14 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/ai-generated-apps-look-right-but-do-they-actually-work/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>ai-generated-apps</category><category>mobile-development</category><category>software-testing</category><category>llm-evaluation</category><category>frontend-engineering</category><category>acceptance-testing</category></item><item><title>Prompt Injection in 3D Scenes: The Attack Surface Multimodal Agents Ignore</title><link>https://groundy.com/articles/prompt-injection-in-3d-scenes-the-attack-surface-multimodal-agents-ignore/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/prompt-injection-in-3d-scenes-the-attack-surface-multimodal-agents-ignore/</guid><description>arXiv 2602.07104 shows 3D object placement injects instructions into multimodal agents, bypassing text and image filters. Teams must gate modalities and scope actions to.</description><pubDate>Sun, 23 Aug 2026 04:43:40 GMT</pubDate><content:encoded>&lt;p&gt;arXiv 2602.07104 shows 3D object placement injects instructions into multimodal agents, bypassing text and image filters. Teams must gate modalities and scope actions to.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 15 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/prompt-injection-in-3d-scenes-the-attack-surface-multimodal-agents-ignore/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>prompt-injection</category><category>multimodal-llm</category><category>embodied-ai</category><category>3d-scene-security</category><category>agent-safety</category><category>adversarial-ml</category></item><item><title>Why LLM Log Anomaly Detection Pages You for Nothing</title><link>https://groundy.com/articles/why-llm-log-anomaly-detection-pages-you-for-nothing/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/why-llm-log-anomaly-detection-pages-you-for-nothing/</guid><description>arXiv 2608.17965 shows LLM log detectors are overconfident in wrong verdicts. Gate paging on calibrated confidence, not raw F1, to prevent alert fatigue and ignored critical.</description><pubDate>Sat, 22 Aug 2026 05:10:14 GMT</pubDate><content:encoded>&lt;p&gt;arXiv 2608.17965 shows LLM log detectors are overconfident in wrong verdicts. Gate paging on calibrated confidence, not raw F1, to prevent alert fatigue and ignored critical.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/why-llm-log-anomaly-detection-pages-you-for-nothing/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-calibration</category><category>log-anomaly-detection</category><category>aiops</category><category>alert-fatigue</category><category>on-call</category><category>reliability-engineering</category></item><item><title>DeepSeek v4 Flash Vision: Routing Images Without Verified Pricing</title><link>https://groundy.com/articles/deepseek-v4-flash-vision-routing-images-without-verified-pricing/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/deepseek-v4-flash-vision-routing-images-without-verified-pricing/</guid><description>DeepSeek v4-flash-vision-exp lacks verified pricing and benchmarks. This guide prices Qwen-VL alternatives and outlines a fallback protocol for experimental vision endpoints.</description><pubDate>Sat, 22 Aug 2026 04:32:14 GMT</pubDate><content:encoded>&lt;p&gt;DeepSeek v4-flash-vision-exp lacks verified pricing and benchmarks. This guide prices Qwen-VL alternatives and outlines a fallback protocol for experimental vision endpoints.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/deepseek-v4-flash-vision-routing-images-without-verified-pricing/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>deepseek</category><category>vision-models</category><category>image-routing</category><category>qwen-vl</category><category>api-pricing</category><category>agent-pipelines</category></item><item><title>DeepSeek 32B on RTX 3090: Tokens per Second by Quant and Context</title><link>https://groundy.com/articles/deepseek-32b-on-rtx-3090-tokens-per-second-by-quant-and-context/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/deepseek-32b-on-rtx-3090-tokens-per-second-by-quant-and-context/</guid><description>Bandwidth math predicts 32-39 tok/s for Q4_K_M DeepSeek 32B on an RTX 3090 at 8k context. 32k context overflows 24GB VRAM. No measured benchmarks exist, so every figure is.</description><pubDate>Fri, 21 Aug 2026 14:45:48 GMT</pubDate><content:encoded>&lt;p&gt;Bandwidth math predicts 32-39 tok/s for Q4_K_M DeepSeek 32B on an RTX 3090 at 8k context. 32k context overflows 24GB VRAM. No measured benchmarks exist, so every figure is.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 14 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/deepseek-32b-on-rtx-3090-tokens-per-second-by-quant-and-context/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-inference</category><category>rtx-3090</category><category>deepseek</category><category>quantization</category><category>kv-cache</category><category>local-llm</category></item><item><title>PTXBench: LLMs Can Port GPU Kernels, But Not Beat Tuned Libraries</title><link>https://groundy.com/articles/ptxbench-llms-can-port-gpu-kernels-but-not-beat-tuned-libraries/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/ptxbench-llms-can-port-gpu-kernels-but-not-beat-tuned-libraries/</guid><description>PTXBench shows LLMs can generate architecture-specific PTX for H100 and B200, but no model matches frontier libraries. Use it to cut hot-loop porting costs, not to replace.</description><pubDate>Fri, 21 Aug 2026 14:13:25 GMT</pubDate><content:encoded>&lt;p&gt;PTXBench shows LLMs can generate architecture-specific PTX for H100 and B200, but no model matches frontier libraries. Use it to cut hot-loop porting costs, not to replace.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/ptxbench-llms-can-port-gpu-kernels-but-not-beat-tuned-libraries/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>ptxbench</category><category>gpu-kernels</category><category>ptx</category><category>llm-code-generation</category><category>cuda</category><category>hopper-blackwell</category></item><item><title>Can LLMs Reuse Another Model&apos;s KV Cache? What Cross-Model Transfer Shows</title><link>https://groundy.com/articles/can-llms-reuse-another-models-kv-cache-what-cross-model-transfer-shows/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-llms-reuse-another-models-kv-cache-what-cross-model-transfer-shows/</guid><description>A new preprint proposes a lightweight reader to reuse KV caches across models, but evidence suggests caches remain model-bound. Learn when recompute beats adaptation.</description><pubDate>Thu, 20 Aug 2026 13:44:30 GMT</pubDate><content:encoded>&lt;p&gt;A new preprint proposes a lightweight reader to reuse KV caches across models, but evidence suggests caches remain model-bound. Learn when recompute beats adaptation.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-llms-reuse-another-models-kv-cache-what-cross-model-transfer-shows/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kv-cache</category><category>llm-inference</category><category>model-transfer</category><category>vllm</category><category>prefix-caching</category><category>serving-optimization</category></item><item><title>DeCRIM: Decompose Constraints to Stop Silent Drops in Agent Outputs</title><link>https://groundy.com/articles/decrim-decompose-constraints-to-stop-silent-drops-in-agent-outputs/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/decrim-decompose-constraints-to-stop-silent-drops-in-agent-outputs/</guid><description>DeCRIM shows that decomposing multi-constraint instructions into individually checkable units reduces silent drops by 7-8% on benchmarks, shifting reliability work from.</description><pubDate>Sat, 01 Aug 2026 23:23:57 GMT</pubDate><content:encoded>&lt;p&gt;DeCRIM shows that decomposing multi-constraint instructions into individually checkable units reduces silent drops by 7-8% on benchmarks, shifting reliability work from.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/decrim-decompose-constraints-to-stop-silent-drops-in-agent-outputs/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-reliability</category><category>instruction-following</category><category>agent-architecture</category><category>self-correction</category><category>prompt-engineering</category><category>decrim</category><category>models-research</category></item><item><title>Operator-Level Triage for Silent Mixed-Precision Instability</title><link>https://groundy.com/articles/operator-level-triage-for-silent-mixed-precision-instability/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/operator-level-triage-for-silent-mixed-precision-instability/</guid><description>arXiv 2607.25494 introduces a single-pass tool to localize silent numerical instability in bf16 and fp16 training. Treat it as a triage layer for unexplained loss spikes, not.</description><pubDate>Fri, 31 Jul 2026 17:38:42 GMT</pubDate><content:encoded>&lt;p&gt;arXiv 2607.25494 introduces a single-pass tool to localize silent numerical instability in bf16 and fp16 training. Treat it as a triage layer for unexplained loss spikes, not.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/operator-level-triage-for-silent-mixed-precision-instability/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>mixed-precision</category><category>numerical-stability</category><category>llm-training</category><category>debugging</category><category>pytorch</category><category>cestat</category><category>models-research</category></item><item><title>BeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration</title><link>https://groundy.com/articles/beyonduncertainty-weak-confidence-signal-for-rag-routing-not-calibration/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/beyonduncertainty-weak-confidence-signal-for-rag-routing-not-calibration/</guid><description>arXiv:2607.25600 shows verbalized confidence is a weak but real routing signal for RAG. It saves 20.4% retrieval calls for 28.2% token overhead. The probe is poorly.</description><pubDate>Fri, 31 Jul 2026 17:14:36 GMT</pubDate><content:encoded>&lt;p&gt;arXiv:2607.25600 shows verbalized confidence is a weak but real routing signal for RAG. It saves 20.4% retrieval calls for 28.2% token overhead. The probe is poorly.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/beyonduncertainty-weak-confidence-signal-for-rag-routing-not-calibration/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>rag</category><category>llm-uncertainty</category><category>retrieval-routing</category><category>confidence-calibration</category><category>arxiv-260725600</category><category>beyonduncertainty</category></item><item><title>Kimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference</title><link>https://groundy.com/articles/kimi-k3-on-m1-max-bandwidth-not-capacity-limits-local-moe-inference/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-k3-on-m1-max-bandwidth-not-capacity-limits-local-moe-inference/</guid><description>Running Kimi K3 on an M1 Max proves local MoE feasibility via expert offloading, but unified memory bandwidth caps throughput. Treat this as a prototyping tool, not a serving.</description><pubDate>Thu, 30 Jul 2026 08:31:44 GMT</pubDate><content:encoded>&lt;p&gt;Running Kimi K3 on an M1 Max proves local MoE feasibility via expert offloading, but unified memory bandwidth caps throughput. Treat this as a prototyping tool, not a serving.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-k3-on-m1-max-bandwidth-not-capacity-limits-local-moe-inference/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kimi-k3</category><category>m1-max</category><category>moe-inference</category><category>local-llm</category><category>apple-silicon</category><category>memory-bandwidth</category></item><item><title>Kimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint</title><link>https://groundy.com/articles/kimi-linear-cuts-kv-cache-75-but-recall-remains-the-binding-constraint/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-linear-cuts-kv-cache-75-but-recall-remains-the-binding-constraint/</guid><description>Kimi Linear cuts KV cache 75% and boosts decode 6x for 1M-token contexts, but recall stays the binding constraint. Route memory-bound jobs only after benchmarking in-context.</description><pubDate>Wed, 29 Jul 2026 07:59:50 GMT</pubDate><content:encoded>&lt;p&gt;Kimi Linear cuts KV cache 75% and boosts decode 6x for 1M-token contexts, but recall stays the binding constraint. Route memory-bound jobs only after benchmarking in-context.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-linear-cuts-kv-cache-75-but-recall-remains-the-binding-constraint/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kimi-linear</category><category>linear-attention</category><category>llm-inference</category><category>long-context</category><category>moonshot</category><category>kv-cache</category><category>inference-optimization</category></item><item><title>Why Chat Leaderboards Do Not Predict Image Quality</title><link>https://groundy.com/articles/why-chat-leaderboards-do-not-predict-image-quality/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/why-chat-leaderboards-do-not-predict-image-quality/</guid><description>A July 22 visual test comparing GPT-5.6, Claude, Gemini, and Grok highlights a structural gap: chat models lack dedicated image backends. Teams must route by subtask using.</description><pubDate>Tue, 28 Jul 2026 06:44:08 GMT</pubDate><content:encoded>&lt;p&gt;A July 22 visual test comparing GPT-5.6, Claude, Gemini, and Grok highlights a structural gap: chat models lack dedicated image backends. Teams must route by subtask using.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/why-chat-leaderboards-do-not-predict-image-quality/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>image-generation</category><category>model-routing</category><category>multimodal-eval</category><category>gpt-image-2</category><category>flux-2</category><category>chatbot-arena</category></item><item><title>Kimi K3 Procurement: Governance Review Over Phantom Government Assessments</title><link>https://groundy.com/articles/kimi-k3-procurement-governance-review-over-phantom-government-assessments/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-k3-procurement-governance-review-over-phantom-government-assessments/</guid><description>Moonshot AI&apos;s Kimi K3 release triggers governance review for regulated teams. Beijing jurisdiction, Anthropic accusations, and missing government assessment require internal.</description><pubDate>Mon, 27 Jul 2026 18:37:58 GMT</pubDate><content:encoded>&lt;p&gt;Moonshot AI&amp;apos;s Kimi K3 release triggers governance review for regulated teams. Beijing jurisdiction, Anthropic accusations, and missing government assessment require internal.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-k3-procurement-governance-review-over-phantom-government-assessments/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kimi-k3</category><category>moonshot-ai</category><category>ai-procurement</category><category>vendor-governance</category><category>compliance</category><category>security</category><category>models-research</category></item><item><title>DeepSeek Compute Leak: Why Open-Weight Routing Needs a Swap Path</title><link>https://groundy.com/articles/deepseek-compute-leak-why-open-weight-routing-needs-a-swap-path/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/deepseek-compute-leak-why-open-weight-routing-needs-a-swap-path/</guid><description>A leaked transcript claims DeepSeek faces a compute gap, but V4 shipped in April. Teams should treat this as a resilience prompt to pin Qwen as a swap-in spine, not panic.</description><pubDate>Sun, 26 Jul 2026 18:47:42 GMT</pubDate><content:encoded>&lt;p&gt;A leaked transcript claims DeepSeek faces a compute gap, but V4 shipped in April. Teams should treat this as a resilience prompt to pin Qwen as a swap-in spine, not panic.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/deepseek-compute-leak-why-open-weight-routing-needs-a-swap-path/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>deepseek</category><category>open-weight</category><category>model-routing</category><category>qwen</category><category>supply-chain</category><category>inference</category></item><item><title>Context Ordering Beats Window Size for Long-Context Agents</title><link>https://groundy.com/articles/context-ordering-beats-window-size-for-long-context-agents/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/context-ordering-beats-window-size-for-long-context-agents/</guid><description>ARBIGRAPH shows tool agents lose 33.3% accuracy on dependent chains despite fitting context. Context ordering, truncation, and caching now beat raw window size for agent.</description><pubDate>Sun, 26 Jul 2026 13:05:53 GMT</pubDate><content:encoded>&lt;p&gt;ARBIGRAPH shows tool agents lose 33.3% accuracy on dependent chains despite fitting context. Context ordering, truncation, and caching now beat raw window size for agent.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/context-ordering-beats-window-size-for-long-context-agents/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>context-engineering</category><category>agent-architecture</category><category>llm-inference</category><category>prompt-optimization</category><category>retrieval-augmented-generation</category><category>model-evaluation</category></item><item><title>Diffusion LLMs: Training Cost, Not Parallel Decoding, Drives Deployment</title><link>https://groundy.com/articles/diffusion-llms-training-cost-not-parallel-decoding-drives-deployment/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/diffusion-llms-training-cost-not-parallel-decoding-drives-deployment/</guid><description>Parallel decoding has not cut serving costs for diffusion LLMs. arXiv:2605.13026 closes training gaps by 4x, but KV-cache absence keeps inference slower than autoregressive.</description><pubDate>Fri, 24 Jul 2026 18:52:28 GMT</pubDate><content:encoded>&lt;p&gt;Parallel decoding has not cut serving costs for diffusion LLMs. arXiv:2605.13026 closes training gaps by 4x, but KV-cache absence keeps inference slower than autoregressive.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/diffusion-llms-training-cost-not-parallel-decoding-drives-deployment/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>diffusion-llm</category><category>llada</category><category>inference-cost</category><category>training-efficiency</category><category>llm-architecture</category><category>serving-infrastructure</category></item><item><title>Open-Weight Routers vs Fable 5: The Routing Math That Actually Matters</title><link>https://groundy.com/articles/open-weight-routers-vs-fable-5-the-routing-math-that-actually-matters/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/open-weight-routers-vs-fable-5-the-routing-math-that-actually-matters/</guid><description>Fable 5 pricing and caching set the bar for open-weight routers. The Echo claim lacks verification. Routing wins only for high-volume, cache-unfriendly routine traffic.</description><pubDate>Fri, 24 Jul 2026 06:41:38 GMT</pubDate><content:encoded>&lt;p&gt;Fable 5 pricing and caching set the bar for open-weight routers. The Echo claim lacks verification. Routing wins only for high-volume, cache-unfriendly routine traffic.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/open-weight-routers-vs-fable-5-the-routing-math-that-actually-matters/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-routing</category><category>fable-5</category><category>open-weights</category><category>cost-optimization</category><category>inference-architecture</category><category>prompt-caching</category></item><item><title>DeepSeek-V4 1M Context vs RAG: Why Retrieval Stays</title><link>https://groundy.com/articles/deepseek-v4-1m-context-vs-rag-why-retrieval-stays/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/deepseek-v4-1m-context-vs-rag-why-retrieval-stays/</guid><description>DeepSeek-V4 markets a 1M token window for agents, but no independent benchmarks verify depth recall. CEO-Bench and ProGraph research show structured memory outperforms raw.</description><pubDate>Thu, 23 Jul 2026 17:52:03 GMT</pubDate><content:encoded>&lt;p&gt;DeepSeek-V4 markets a 1M token window for agents, but no independent benchmarks verify depth recall. CEO-Bench and ProGraph research show structured memory outperforms raw.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/deepseek-v4-1m-context-vs-rag-why-retrieval-stays/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>deepseek-v4</category><category>long-context</category><category>rag</category><category>agent-memory</category><category>retrieval</category><category>llm-evaluation</category></item><item><title>Qwen-Image-3.0 Does Not Exist: Why Self-Hosting Image Models Is Premature</title><link>https://groundy.com/articles/qwen-image-3-0-does-not-exist-why-self-hosting-image-models-is-premature/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/qwen-image-3-0-does-not-exist-why-self-hosting-image-models-is-premature/</guid><description>No primary source confirms Qwen-Image-3.0 exists. Independent benchmarks show open-weight models fail 85% of precise image tasks. Self-hosting branded templates remains risky.</description><pubDate>Thu, 23 Jul 2026 06:07:11 GMT</pubDate><content:encoded>&lt;p&gt;No primary source confirms Qwen-Image-3.0 exists. Independent benchmarks show open-weight models fail 85% of precise image tasks. Self-hosting branded templates remains risky.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/qwen-image-3-0-does-not-exist-why-self-hosting-image-models-is-premature/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>qwen</category><category>image-generation</category><category>open-weight</category><category>vector-bench</category><category>self-hosting</category><category>midjourney</category></item><item><title>Kimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality</title><link>https://groundy.com/articles/kimi-k3-architecture-what-2-8t-parameters-change-for-deployment/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-k3-architecture-what-2-8t-parameters-change-for-deployment/</guid><description>Kimi K3 pairs 2.8T parameters with 16-of-896 expert routing and a 1M context window. Hosted API is live, but full weights arrive July 27, 2026. Self-hosting requires.</description><pubDate>Mon, 20 Jul 2026 18:19:20 GMT</pubDate><content:encoded>&lt;p&gt;Kimi K3 pairs 2.8T parameters with 16-of-896 expert routing and a 1M context window. Hosted API is live, but full weights arrive July 27, 2026. Self-hosting requires.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 12 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-k3-architecture-what-2-8t-parameters-change-for-deployment/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kimi-k3</category><category>mixture-of-experts</category><category>llm-inference</category><category>moonshot-ai</category><category>model-deployment</category><category>sparse-attention</category></item><item><title>Kimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026</title><link>https://groundy.com/articles/kimi-k3-vs-qwen3-8-max-which-frontier-model-should-teams-route/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-k3-vs-qwen3-8-max-which-frontier-model-should-teams-route/</guid><description>Kimi K3 offers concrete pricing for production bake-offs while Qwen3.8 Max remains a shadow evaluation. Compare both against Fable 5, GPT-5.6 Sol, and GLM-5.2 by workload.</description><pubDate>Mon, 20 Jul 2026 18:07:18 GMT</pubDate><content:encoded>&lt;p&gt;Kimi K3 offers concrete pricing for production bake-offs while Qwen3.8 Max remains a shadow evaluation. Compare both against Fable 5, GPT-5.6 Sol, and GLM-5.2 by workload.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-k3-vs-qwen3-8-max-which-frontier-model-should-teams-route/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>kimi-k3</category><category>qwen3-8-max</category><category>model-routing</category><category>llm-evaluation</category><category>cost-optimization</category><category>agentic-workflows</category></item><item><title>Qwen3.8 Max Release Audit: API, Open Weights, and the License Catch</title><link>https://groundy.com/articles/qwen3-8-max-preview-what-alibaba-shipped-and-what-is-missing/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/qwen3-8-max-preview-what-alibaba-shipped-and-what-is-missing/</guid><description>Qwen3.8 Max now has a stable API, $2/$6 pricing, 1M context, and downloadable 2.4T weights. The release is real, but the API and open checkpoint are not the same product.</description><pubDate>Mon, 20 Jul 2026 17:45:04 GMT</pubDate><content:encoded>&lt;p&gt;Qwen3.8 Max now has a stable API, $2/$6 pricing, 1M context, and downloadable 2.4T weights. The release is real, but the API and open checkpoint are not the same product.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/qwen3-8-max-preview-what-alibaba-shipped-and-what-is-missing/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><atom:updated>2026-08-18T00:00:00.000Z</atom:updated><category>Models &amp; Research</category><category>qwen3-8-max</category><category>open-weights</category><category>model-routing</category><category>llm-evaluation</category><category>self-hosting</category><category>model-licensing</category></item><item><title>HuggingFace 100x Inference: Generalizable vs Platform-Locked Optimizations</title><link>https://groundy.com/articles/huggingface-100x-inference-generalizable-vs-platform-locked-optimizations/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/huggingface-100x-inference-generalizable-vs-platform-locked-optimizations/</guid><description>HuggingFace claims 100x inference speedup, but independent benchmarks show its TGI engine trails vLLM by 24x on throughput. This article decomposes the claim to separate.</description><pubDate>Sun, 19 Jul 2026 18:50:32 GMT</pubDate><content:encoded>&lt;p&gt;HuggingFace claims 100x inference speedup, but independent benchmarks show its TGI engine trails vLLM by 24x on throughput. This article decomposes the claim to separate.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 13 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/huggingface-100x-inference-generalizable-vs-platform-locked-optimizations/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>inference-optimization</category><category>huggingface</category><category>vllm</category><category>llm-serving</category><category>build-vs-buy</category><category>self-hosted-llm</category></item><item><title>Kimi K3 Code Arena Rank: Self-Hosting Cost Math for Coding Agents</title><link>https://groundy.com/articles/kimi-k3-code-arena-rank-self-hosting-cost-math-for-coding-agents/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/kimi-k3-code-arena-rank-self-hosting-cost-math-for-coding-agents/</guid><description>Kimi K3 is frontier-class, but its 2.8T MXFP4 checkpoint needs 64-plus accelerators. The real deployment math reveals what an arena rank cannot.</description><pubDate>Sun, 19 Jul 2026 17:13:36 GMT</pubDate><content:encoded>&lt;p&gt;Kimi K3 is frontier-class, but its 2.8T MXFP4 checkpoint needs 64-plus accelerators. The real deployment math reveals what an arena rank cannot.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 11 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/kimi-k3-code-arena-rank-self-hosting-cost-math-for-coding-agents/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><atom:updated>2026-07-20T00:00:00.000Z</atom:updated><category>Models &amp; Research</category><category>kimi-k3</category><category>self-hosting</category><category>coding-agents</category><category>gpu-costs</category><category>open-weight</category><category>inference-throughput</category><category>model-economics</category></item><item><title>Can Tool-Adaptive LLM Rerankers Improve RAG Without Always Calling Tools?</title><link>https://groundy.com/articles/can-tool-adaptive-llm-rerankers-improve-rag-without-always-calling-tools/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-tool-adaptive-llm-rerankers-improve-rag-without-always-calling-tools/</guid><description>TALRanker folds the tool-call decision into the reranker&apos;s scoring policy, turning tool latency from a fixed per-query tax into a budget the model spends only when uncertain.</description><pubDate>Tue, 14 Jul 2026 22:48:18 GMT</pubDate><content:encoded>&lt;p&gt;TALRanker folds the tool-call decision into the reranker&amp;apos;s scoring policy, turning tool latency from a fixed per-query tax into a budget the model spends only when uncertain.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 8 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-tool-adaptive-llm-rerankers-improve-rag-without-always-calling-tools/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>rag</category><category>reranking</category><category>tool-calling</category><category>retrieval</category><category>llm-inference</category><category>inference-cost</category></item><item><title>Does Speculative Decoding with Progressive Tree Drafting Cut LLM Latency?</title><link>https://groundy.com/articles/does-speculative-decoding-with-progressive-tree-drafting-cut-llm-latency/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/does-speculative-decoding-with-progressive-tree-drafting-cut-llm-latency/</guid><description>Progressive Tree Drafting grows draft tokens as a pruned tree to claim a 2× speedup, but the gain hinges on verifier acceptance and shrinks on code and reasoning.</description><pubDate>Tue, 14 Jul 2026 21:32:35 GMT</pubDate><content:encoded>&lt;p&gt;Progressive Tree Drafting grows draft tokens as a pruned tree to claim a 2× speedup, but the gain hinges on verifier acceptance and shrinks on code and reasoning.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 8 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/does-speculative-decoding-with-progressive-tree-drafting-cut-llm-latency/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>speculative-decoding</category><category>progressive-tree-drafting</category><category>llm-inference</category><category>inference-latency</category><category>kv-cache</category><category>vllm</category></item><item><title>Analytic Inference Cuts Bayesian Deep Ensemble Serving Cost, But Leaves Training as the Bottleneck</title><link>https://groundy.com/articles/analytic-inference-cuts-bayesian-deep-ensemble-serving-cost-but-leaves-training/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/analytic-inference-cuts-bayesian-deep-ensemble-serving-cost-but-leaves-training/</guid><description>A new arXiv preprint replaces sampling-based averaging in Bayesian deep ensembles with closed-form Bayesian aggregation, cutting per-query inference cost and shifting the.</description><pubDate>Fri, 10 Jul 2026 22:05:25 GMT</pubDate><content:encoded>&lt;p&gt;A new arXiv preprint replaces sampling-based averaging in Bayesian deep ensembles with closed-form Bayesian aggregation, cutting per-query inference cost and shifting the.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 8 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/analytic-inference-cuts-bayesian-deep-ensemble-serving-cost-but-leaves-training/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>bayesian-deep-ensembles</category><category>uncertainty-quantification</category><category>analytic-inference</category><category>bayesian-linear-regression</category><category>inference-cost</category><category>neural-networks</category><category>arxiv-2607-06776</category></item><item><title>Tree-of-Thoughts Improves Text-to-Image Prompting by Reasoning Over Hypotheses, Not Pixels</title><link>https://groundy.com/articles/tree-of-thoughts-improves-text-to-image-prompting-by-reasoning-over-hypotheses/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/tree-of-thoughts-improves-text-to-image-prompting-by-reasoning-over-hypotheses/</guid><description>A new arXiv paper ports Tree-of-Thoughts prompting to text-to-image in-context learning and reports CoBSAT gains, but reasoning runs on prompt hypotheses, not image states.</description><pubDate>Fri, 10 Jul 2026 18:55:21 GMT</pubDate><content:encoded>&lt;p&gt;A new arXiv paper ports Tree-of-Thoughts prompting to text-to-image in-context learning and reports CoBSAT gains, but reasoning runs on prompt hypotheses, not image states.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 8 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/tree-of-thoughts-improves-text-to-image-prompting-by-reasoning-over-hypotheses/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>tree-of-thoughts</category><category>text-to-image</category><category>in-context-learning</category><category>prompt-engineering</category><category>compositional-reasoning</category><category>diffusion-models</category></item><item><title>FourierQK&apos;s spectral Q/K filter cuts TinyShakespeare loss by 79%, but long-context proof is missing</title><link>https://groundy.com/articles/fourierqks-spectral-q-k-filter-cuts-tinyshakespeare-loss-by-79-but-long-context/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/fourierqks-spectral-q-k-filter-cuts-tinyshakespeare-loss-by-79-but-long-context/</guid><description>FourierQK filters query and key projections before attention, cutting TinyShakespeare character-level loss by 79%, but word-level, retrieval and long-context tests are absent.</description><pubDate>Fri, 10 Jul 2026 16:19:03 GMT</pubDate><content:encoded>&lt;p&gt;FourierQK filters query and key projections before attention, cutting TinyShakespeare character-level loss by 79%, but word-level, retrieval and long-context tests are absent.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 7 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/fourierqks-spectral-q-k-filter-cuts-tinyshakespeare-loss-by-79-but-long-context/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>transformer-attention</category><category>spectral-methods</category><category>long-context</category><category>fourierqk</category><category>language-models</category><category>retrieval</category><category>inference-cost</category></item><item><title>Tencent Hunyuan 3&apos;s Agent Push Has No Public DeepSeek or Qwen Benchmarks Yet</title><link>https://groundy.com/articles/tencent-hunyuan-3s-agent-push-has-no-public-deepseek-or-qwen-benchmarks-yet/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/tencent-hunyuan-3s-agent-push-has-no-public-deepseek-or-qwen-benchmarks-yet/</guid><description>Tencent&apos;s Hy3 is billed as a 295B agent model with a 21B active token footprint, but public materials omit benchmark tables and independent DeepSeek or Qwen comparisons.</description><pubDate>Fri, 10 Jul 2026 14:28:52 GMT</pubDate><content:encoded>&lt;p&gt;Tencent&amp;apos;s Hy3 is billed as a 295B agent model with a 21B active token footprint, but public materials omit benchmark tables and independent DeepSeek or Qwen comparisons.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 6 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/tencent-hunyuan-3s-agent-push-has-no-public-deepseek-or-qwen-benchmarks-yet/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>tencent-hunyuan</category><category>hy3</category><category>mixture-of-experts</category><category>agent-models</category><category>model-evaluation</category><category>llm-benchmarks</category><category>routing</category></item><item><title>When Does Memory, Not Compute, Decide Who Can Profitably Serve LLMs?</title><link>https://groundy.com/articles/when-does-memory-not-compute-decide-who-can-profitably-serve-llms/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/when-does-memory-not-compute-decide-who-can-profitably-serve-llms/</guid><description>A July 2026 arXiv paper argues that scarce HBM and DRAM bandwidth, not raw compute, will determine which labs and providers can profitably serve large language models through.</description><pubDate>Fri, 10 Jul 2026 11:38:30 GMT</pubDate><content:encoded>&lt;p&gt;A July 2026 arXiv paper argues that scarce HBM and DRAM bandwidth, not raw compute, will determine which labs and providers can profitably serve large language models through.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 9 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/when-does-memory-not-compute-decide-who-can-profitably-serve-llms/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-economics</category><category>hbm-scarcity</category><category>inference-costs</category><category>ai-infrastructure</category><category>open-weights</category><category>memory-bandwidth</category><category>vintage-solvency</category></item><item><title>Can We Trust LLM Logic? A Graph-Based Stress Test Finds Three Failure Modes</title><link>https://groundy.com/articles/can-we-trust-llm-logic-a-graph-based-stress-test-finds-three-failure-modes/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=models-research</link><guid isPermaLink="true">https://groundy.com/articles/can-we-trust-llm-logic-a-graph-based-stress-test-finds-three-failure-modes/</guid><description>A July 2026 arXiv paper shows Self-Consistency voting can hide contradictory reasoning, and GraphEVAL&apos;s graph-based coherence metrics catch flawed paths output checks miss.</description><pubDate>Fri, 10 Jul 2026 10:18:06 GMT</pubDate><content:encoded>&lt;p&gt;A July 2026 arXiv paper shows Self-Consistency voting can hide contradictory reasoning, and GraphEVAL&amp;apos;s graph-based coherence metrics catch flawed paths output checks miss.&lt;/p&gt;&lt;p&gt;Berry Mingus · Models &amp;amp; Research · 7 min read&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://groundy.com/articles/can-we-trust-llm-logic-a-graph-based-stress-test-finds-three-failure-modes/?utm_source=rss&amp;amp;utm_medium=feed&amp;amp;utm_campaign=models-research&quot;&gt;Read the full article on Groundy →&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Berry Mingus</dc:creator><category>Models &amp; Research</category><category>llm-reasoning</category><category>graph-eval</category><category>self-consistency</category><category>chain-of-thought</category><category>uncertainty-quantification</category><category>verification</category><category>reasoning-agents</category></item></channel></rss>