groundy
developer tools

Do LLM Users Get Better With Practice? What Longitudinal Chat Logs Show

A preprint tracking 12,000 Bing Copilot users finds LLM habits stay sticky rather than deepening with practice. This suggests AI seat value is front-loaded, shifting budget to

13 min···3 sources ↓

Most LLM users do not appear to get better with practice, according to Adopt ≠ Adapt: Longitudinal Analyses of LLM Conversations in the Wild, a single, unreplicated, author-reported preprint that tracks roughly 12,0001 Bing Copilot users over time and cross-checks them against the WildChat-4.8M corpus. Its central claim: people adopt the tools without adapting how they work. If that holds, the payoff from a Copilot or Claude seat is front-loaded, and AI transformation budgets belong in workflow redesign, not prompt training.

What did the preprint actually measure?

The study reconstructs the conversational trajectories of roughly 12,0001 randomly sampled Microsoft Bing Copilot users and asks whether each individual’s usage deepens the longer they keep using the tool, then checks whether the same patterns appear in the independent WildChat-4.8M dataset.

That design choice matters more than it looks. Most published evidence about “prompting skill” is cross-sectional: a snapshot comparing light users to heavy users at one point in time. The preprint’s own introduction files the prior literature the same way: temporal analyses of LLM usage, it notes, have primarily shown population-level trends (Chatterji et al. 2025) or cohort differences (Massenkoff et al. 2026), which “do not necessarily reflect behavioral changes of individuals,” with individual user trajectories only recently emerging as an object of study (Zhu et al. 2026a, 2026b). Snapshots cannot distinguish a user who improved from a user who was always good. A longitudinal panel can, at least in principle, because it follows the same person across many sessions and asks whether that person’s conversations change. The authors apply exactly that lens to Bing Copilot usage, then repeat the analysis on WildChat-4.8M, a corpus drawn from the public WildChat collection of real-world LLM conversations, to test whether the Bing-specific results are an artifact of one product’s interface or user base.

Two scoping facts belong up front. First, the paper is catalogued as arXiv:2605.29018 in cs.AI, and arXiv does not peer-review submissions: contents are wholly the submitter’s responsibility, presented “as is” without warranty. Every number in this article is therefore author-reported until an independent group replicates it. Second, the abstract names only two data sources, Microsoft Bing Copilot and WildChat-4.8M; nothing here was measured on GitHub Copilot, Claude Code, or any enterprise developer tooling, and the “success” and “complexity” labels attached to conversations are assigned by the study’s own methodology, whose derivation the abstract does not detail.

Because the deepening lives in the population rather than the person: the preprint’s central result is that trends in individual user trajectories are much weaker than population-level trends, a gap the authors summarize by describing user habits as overwhelmingly sticky.

This is the aggregate-versus-panel trap, and it is easy to fall into even with clean data. A population average can drift upward while every individual inside it stays flat. All that requires is compositional movement: heavier users joining the pool, lighter users churning out, or a mix shift toward demographics that were always more intense users. The average improves; nobody improves. When the authors decompose their Bing Copilot panel into individual trajectories, they find the within-user component is weak compared to the between-user structure, and they suggest that existing user behavior is difficult to change.

The second-order consequence lands on every internal AI dashboard. Most organizations tracking “AI maturity” are tracking population statistics: organization-wide session counts, average prompt lengths, aggregate task mix. Those numbers can climb quarter over quarter and produce a reassuring slope on a slide, while the median employee uses the tool exactly the way they did in week one. The preprint’s decomposition suggests this is not a hypothetical failure mode. It is the default reading of the data unless someone explicitly separates within-user change from compositional change. If your AI adoption report cannot answer “did the same people get better,” it is answering a different, easier question.

Are power users’ complex workflows a practice effect?

No, at least not on this evidence: more active users do have more successful conversations and do use the LLM for more complex, professionally oriented tasks, but that pattern is cross-sectional, and the abstract’s summary is that the results “demonstrate the extent of user heterogeneity” and that “existing user behavior is difficult to change.”

The distinction is worth stating precisely because it is the most common misreading this paper will attract. A cross-sectional comparison says: at a given moment, heavy users and light users differ. There are two ways that difference can arise. In the learning story, users start similar and diverge as the motivated ones practice. In the selection story, users were never similar: people with complex professional work gravitate toward heavier use because they have more raw material to feed the tool, and people without that work stay light no matter how long they hold a seat. The preprint’s longitudinal component is what adjudicates between these stories, and it comes down on the side of selection: the population-level deepening is composition, not within-user growth.

QuestionWhat the preprint foundWhat it does not show
Do individual users deepen over time?Individual trajectories are much weaker than population trends; habits are stickyThat no individual anywhere improves
Do active users succeed more?More active users have more successful conversationsThat activity causes success (cross-sectional)
Do active users attempt harder work?Heavier users take on more complex, professionally oriented tasksThat they grew into that work with practice
Does the pattern generalize?Some user trends also appear in WildChat-4.8MAnything about enterprise coding agents
Where should budget go?Implication: workflow redesign over generic prompt trainingA measured ROI comparison of the two
What should be measured?Implication: value per completed taskThat seat counts are worthless as adoption signals

Note what the selection finding does and does not excuse. It does not mean heavy users are doing anything wrong, or that their usage is unsophisticated. It means their sophistication was imported from their jobs, not manufactured by seat time. For a team lead, that reframes the hiring-and-tooling question: the people extracting the most from an LLM seat may be the people whose work was already well-structured for it.

How much weight can one unreplicated preprint carry?

Limited weight: arXiv presents submissions as is, without peer review, so every figure here is a single research group’s classification of chat logs until someone else reproduces it.

The submission history shows one revision, v2 on 31 August 2026, after the original 27 May 2026 posting; the stickiness result survives in the current version’s abstract.

The limitations stack. The sample is large, roughly 12,000 users1 plus the WildChat-4.8M comparison corpus, but size does not fix construct validity. Whether a conversation “succeeded” and whether a task was “complex” or “professionally oriented” are labels assigned by the study’s own methodology, whose derivation the abstract does not detail; different labeling choices could soften or sharpen the stickiness result. The populations are Bing Copilot users and WildChat-4.8M users, not developers inside a structured engineering organization. A company with deliberate onboarding, shared prompt libraries, code-review norms around agent output, and task routing built into its tools is a genuinely different environment, and the preprint’s data says nothing about whether stickiness survives that environment.

Context on the venue itself: arXiv now hosts more than three million scholarly articles across eight subject areas, and it is going through an institutional transition of its own: after decades of partnership with Cornell University, arXiv is establishing itself as an independent nonprofit. None of that bears on this paper’s correctness. It bears on how the finding should be filed: as a well-instrumented hypothesis from one team, entering a literature that has not yet tested it.

None of this makes the result dismissible. The individual-versus-population decomposition, documented in the preprint’s full text, is methodologically the right way to ask the question, and the conclusion aligns with what many practitioners suspect anecdotally. The replication inside the paper is partial, and the authors attach the qualifier themselves: only some user trends also appear in WildChat-4.8M, that corpus is significantly skewed toward highly proficient power users, and the paper warns it does not represent typical user-AI interactions. The skew cuts in one useful direction. A corpus that overrepresents proficient users is where a practice effect should be easiest to find, so partial agreement there is easier to square with selection than with practice. The correct posture is to treat it as the current best evidence on a question where the alternative evidence is mostly vibes, while holding the update loosely.

What does a front-loaded payoff do to the AI budget?

It moves the intervention lever from training budgets to workflow redesign: if usage depth does not grow with seat time, generic prompt training has nothing to compound, and the returns come from restructuring the work itself.

Most AI transformation programs are built on an implicit learning curve. Buy seats, run a prompt-engineering workshop, and let familiarity do the rest; competence, the assumption goes, accrues with exposure. The preprint attacks that assumption at its foundation. If individual trajectories are flat, then a user’s hundredth session looks like their tenth, and the marginal value of another month of seat time is roughly zero once the initial onboarding bump is captured. The seat’s payoff is front-loaded: you get whatever you get early, and then you get it again, repeatedly, at the same depth.

That reframing changes where the money should go. If people will not adapt their behavior to the tool, the remaining lever is adapting the work to the tool: task routing that sends LLM-tractable problems to the model by default, templates and scaffolds embedded in the interface rather than taught in a workshop, guardrails that make the shallow usage pattern harmless, and workflows decomposed so that the model receives well-formed inputs regardless of how little the user has learned to provide them. This is a harder and more expensive program than a training budget line. It is also, on this evidence, the only one with a mechanism. Training asks users to change; the authors suggest that existing user behavior is difficult to change. Workflow design does not ask.

There is a budget-politics wrinkle worth naming. Prompt training is easy to buy and easy to report: hours delivered, seats trained, satisfaction scores collected. Workflow redesign produces no such artifacts and threatens teams whose output is the training itself. Expect the seat-time model to persist in purchasing decisions well after the evidence turns, because the alternative requires admitting that a visible, billable intervention was treating a symptom.

What should you measure instead of seats activated?

Value per completed task: conversation success rate and the task-complexity mix, the same dimensions the preprint operationalized, are better instruments than seat counts or raw session volume.

Seats activated measures adoption, and adoption is precisely the thing the preprint says is happening. The title’s asymmetry is the whole story: adopt, yes; adapt, no. A metric that counts logins, sessions, or license utilization will show a healthy program in exactly the world where the preprint is right and nobody is getting better. It is instrumentation tuned to the wrong half of the phenomenon.

The alternative instruments fall out of the paper’s own design. Conversation success rate, however your organization operationally defines a completed, useful interaction, tracks whether the tool is resolving work rather than absorbing it. Task-complexity mix tracks whether the tool is being applied to the professionally oriented problems where its value concentrates, or to the shallow queries that dominate sticky usage. Both metrics can be computed per person and tracked longitudinally, which gives you the panel structure the preprint used and lets you detect within-user deepening directly rather than inferring it from population drift. If your power users’ complexity edge is selection, this instrumentation will show it: their mix will be sophisticated from week one rather than trending upward.

One caveat applies to the instruments themselves: in the preprint, success and complexity are labels assigned by the study’s own methodology, and your organization’s labels will be your own. That is fine. The goal is not to reproduce their taxonomy but to measure the same dimensions with definitions calibrated to your work, and to resist the temptation to report the population average when the panel decomposition is the informative cut.

What evidence would falsify the claim?

A longitudinal study demonstrating within-user prompting-skill gains over time would do it, and no such study appears in the material reviewed for this article.

That absence is worth stating plainly rather than papering over. The strongest possible counter to a stickiness finding is a panel study on a different population, ideally developers using enterprise tooling under structured onboarding, showing individual trajectories that slope upward: prompts becoming better specified, task complexity rising within the same user, success rates climbing with tenure. The counter-evidence inside the preprint itself is weaker than it looks. The cross-sectional result, that active users succeed more and attempt harder work, is consistent with learning, but the paper’s own longitudinal analysis finds individual trajectories much weaker than population-level trends, so it cannot be recruited as independent support for the practice effect.

There is also a middle position the current evidence cannot exclude. Stickiness may be a property of unsupported use: users left alone with a general-purpose chat interface settle into a local optimum and stay there, while users embedded in a deliberately redesigned workflow might show real within-user growth. The preprint does not test that condition, because the corpora contain no such condition to test. Until a longitudinal study of structured enterprise use exists, the honest summary is that the seat-time model of AI skill growth has taken a serious hit from one well-instrumented, unreplicated preprint, and the burden of proof has moved to whoever wants to keep funding it.

Where does this leave the seat-time assumption?

Budget as if the payoff is front-loaded and the habits are sticky, because that is what the current evidence indicates, while treating the finding as provisional enough to revisit the moment a replication or a structured-onboarding counterexample lands.

The practical program is concrete. Fund workflow redesign over generic prompt training, because training presumes a learning curve this preprint fails to find. Instrument value per completed task, conversation success rate, and task-complexity mix, because seat counts measure the adoption half of the story and stay silent on the adaptation half. Decompose your own usage data into individual trajectories before celebrating population trends, because the aggregate slope is exactly where composition effects hide. And keep the caveat pinned to every slide derived from this paper: one research group, two chat corpora, labels assigned by the study’s own methodology, no peer review, and nothing measured on developer tooling. The claim that users adopt without adapting is the best current answer to whether practice makes better LLM users. It is not yet a settled one.

Frequently Asked Questions

Does the stickiness finding apply to enterprise coding agents like GitHub Copilot?

No, the preprint contains zero measurements of developer tooling. Its data is limited to Bing Copilot and WildChat-4.8M, which the authors explicitly warn does not represent typical user-AI interactions. Extrapolating these consumer chat habits to structured engineering workflows with code-review norms is unsupported by the evidence.

How does the WildChat-4.8M comparison affect the validity of the stickiness claim?

WildChat is significantly skewed toward highly proficient power users, making it the ideal environment to detect a practice effect. The fact that individual trajectories remained flat in this high-skill corpus strengthens the selection argument, as a learning curve should be most visible among users already capable of complex prompting.

What specific operational change does the front-loaded payoff model require for budgeting?

Teams must shift spend from recurring prompt-training workshops to one-time workflow redesign. Since individual usage depth does not compound with seat time, the marginal value of additional training hours is near zero. Budgets should instead fund task routing, interface scaffolds, and guardrails that make shallow usage patterns harmless without requiring user behavior change.

What metric should replace seat activation counts to detect actual value?

Track conversation success rate and task-complexity mix per user over time. These longitudinal panel metrics reveal whether the same individuals are deepening their usage or if population averages are rising due to compositional shifts. Seat counts only measure adoption, which the preprint confirms is occurring without corresponding adaptation.

What type of study would falsify the claim that LLM users do not improve with practice?

A longitudinal panel study on a different population, specifically developers using enterprise tooling under structured onboarding, showing individual trajectories that slope upward. Evidence of prompts becoming better specified and success rates climbing with tenure within the same user would contradict the stickiness finding, but no such study currently exists in the reviewed literature.

sources · 3 cited

  1. About arXiv - arXiv infoinfo.arxiv.orgprimaryaccessed 2026-09-05
  2. ArXiven.wikipedia.orgcommunityaccessed 2026-09-05