A Per-Neuron Sequence Model Was Withdrawn From arXiv as Coverage Hailed It
TND proposed per-neuron dynamics as a sequence-modeling primitive, then was withdrawn from arXiv for accuracy errors the day before coverage called it a Transformer.
The Groundy archive · Page 21 of 34
Browse Groundy's complete archive of 799 articles on AI, developer tools and infrastructure. Page 21 of 34.
481–504 of 799 articles · Newest first
TND proposed per-neuron dynamics as a sequence-modeling primitive, then was withdrawn from arXiv for accuracy errors the day before coverage called it a Transformer.
A June 2026 preprint finds refusal decisions are locked at a model's first token, undercutting the safety case for premium reasoning modes billed per thinking token.
Nub ships a Bun-style all-in-one Node toolkit while keeping stock V8, trading Bun's native-module breakage for Node's own startup floor in a single Rust binary.
A 180-million-repo census finds bot-account lookups miss 97% of Claude Code commits, a 30x recall gap that makes prior AI-agent adoption estimates a floor.
Published jailbreak attack-success rates depend on the judges scoring them. A June 2026 audit finds both judge families miscalibrated and manipulable without removing harm.
PV-TAM, proposed in a June 2026 arXiv preprint, corrects two structural biases in VLM localization scoring: decoding drift and modality boundary-marker contamination.
An ETH Zurich benchmark finds AGENTS.md files don't lift task success rates but add over 20% to inference cost. A second study measures runtime savings, not success rate.
A402 binds agent micropayments to verified delivery, but its TEE attests execution rather than product truth, leaving marketplaces to decide who refunds a wrong attribute.
Meituan's General 365 benchmark caps Gemini 3 Pro at 62.8% and most of 26 models below 60%, exposing how saturated public leaderboards inflate reasoning scores.
A 2026 preprint proposes LLMs as A/B test proxies for human subjects. Raw outputs captured 39% of the treatment effect, and surrogacy bias does not average out.
Token pricing decouples from GPU cost via cached discounts, reasoning markups, and per-model multipliers. No single per-token rate works as a budget unit for real workloads.
An LLM judge's self-preference claim holds only if the study fixed generation quality. Without that control, the bias number conflates real preference with a quality gap.
A June 2026 paper proposes the Goal-Oriented Dialogue Runtime, lifting goals, lifecycle state, and invalidation rules into first-class objects teams can version and diff.
PixJail converts text-to-image jailbreak papers into runnable pipelines, reproducing eleven methods with 2.1% error, a fidelity figure, not a real bypass rate.
A certificate can prove an agent output is signed and intact, but not correct. Validity needs a checkable predicate a verifier can re-run, not all tasks have one.
Vercel has been on the AWS Marketplace since 2024. The real shift is that an EDP commit can absorb the spend, coupling the frontend host to AWS and deepening the lock-in.
A June 2026 preprint grounds ODRL's permissions and prohibitions in the UFO-L legal ontology, naming who holds the power to declare a violation in vendor AI usage policies.
Microsoft merged AutoGen into Agent Framework, leaving CrewAI versus MAF as the 2026 choice. The orchestration primitive you pick becomes your trace and policy boundary.
Vercel's 23 June Node server deploy caps a cluster of duration and transport raises that pull realtime workloads back into one Fluid Compute envelope, raising lock-in stakes.
AutoSpec grows LLM agent safety rules from annotated traces and hits 0.98 F1, but readable rules do not prove the rule set is complete. That is the open governance question.
A cascade grader for agentic data analysis hit 100% precision and 97% recall on 153 tasks, but silently returned no verdict on 64% of runs before a nudge fix.
A June 2026 preprint shows LLM agent societies spontaneously form authority hierarchies, so the orchestration topology you specify is not the only coordination layer running.
CrossPool splits MoE FFN weights and KV-cache into separate GPU pools and ships hidden states across the boundary, trading VRAM residency for an interconnect bottleneck.
Vercel BotID emits session telemetry, verdicts, JA4 digests, paths, and verified-bot labels, rich enough to repurpose as a threat feed and flag what the WAF lets through.