The best available evidence says the verb is not the problem. An 815-participant experiment comparing anthropomorphic and non-anthropomorphic descriptions of AI found the wording barely moved readers’ perceptions, while a text explicitly framed around AI’s dangers did shift views. For editors drafting AI style rules, the finding redirects effort: away from blanket bans on “thinks” and toward the accuracy of risk language, with tighter verb discipline reserved for high-stakes contexts where anthropomorphic cues measurably bite.
What the experiment actually tested
The question is older than the current wave of AI coverage, but a June 2026 experiment gives it a direct test. Participants read texts describing AI; in the main conditions, one version used anthropomorphic language and one did not. The design also included a separate condition built around a text that explicitly discusses the dangers of AI, which functioned as a check on whether reading a short text about AI can move perceptions at all.
That structure matters for how the result should be read. The study is not a survey of what people already believe after years of ChatGPT headlines. It is a controlled reading task: one text, one exposure, immediate measurement. It isolates the wording variable an editor actually controls at the sentence level, which is exactly what makes it relevant to a style guide, and exactly what limits how far it generalizes, a point the final section returns to.
The paper’s arXiv identifier, 2606.29121, marks a June 2026 posting. Its answer complicates the most common draft of AI verb rules.
The null result: anthropomorphic verbs barely moved readers
Across the main conditions, the authors report that “whether the text uses anthropomorphic language does not substantially affect participants’ perceptions of AI” (arXiv:2606.29121). Eight hundred fifteen participants is a reasonably powered sample for detecting moderate effects, so this is not a case of a tiny study failing to find something large. The honest reading: if swapping “the model understands” for “the model predicts” moves general readers at all, it does not move them much, at least not on immediate perception measures after a single text.
The null is not isolated. In a study of an anthropomorphic robotic mock driver, researchers compared verbal-only and multimodal interaction styles against a no-modality baseline carried over from their prior work and found no significant difference in subjective trust: median scores of 42 for the baseline versus 43 for both interaction styles, with F = 0.22 and p = 0.80 (arXiv:2307.00841). Both interaction styles featured the same anthropomorphic mock driver, so the comparison tests how an anthropomorphic robot communicates rather than whether one is present at all. That is a different paradigm (interaction with a robot, not reading about AI), but it is a second independent failure to find the effect that verb-policing orthodoxy assumes.
Neither null proves anthropomorphic wording is harmless. They establish something narrower and more useful: the causal influence editors assume they exert with verb choice is not showing up where it was directly measured.
What did move readers: explicit danger framing
The same experiment contained its own counterweight. In the danger-text condition, “individuals’ views of AI can shift in response to reading a text” (arXiv:2606.29121). Perception is malleable; the malleability just concentrates in framing strength, not in whether the system is described as thinking or computing.
This is the part of the finding with the sharpest practical edge, and it cuts in an uncomfortable direction. If explicit risk framing moves readers while anthropomorphic verbs do not, then the editorial choices that matter most are the ones about how strongly a piece states harms, uncertainty, and stakes. A style guide that spends three pages on banned verbs and one paragraph on risk-language calibration has its priorities inverted relative to the evidence.
Two cautions belong here. First, the dangers-text result demonstrates that perception can shift; it says nothing about whether alarmist framing is accurate. Perception malleability is not a license for fear framing. Second, risk vocabulary is not a neutral dial either. A case study of OpenAI’s public communications found that safety and risk discourse dominates the company’s own messaging and documentation, without drawing on the vocabularies of academic and advocacy ethics frameworks. When coverage leans hard into danger framing, some of that language may be echoing vendor positioning rather than independent assessment. Editors should calibrate risk language against evidence, not against the industry’s loudest vocabulary.
ChatGPT-specific check: trust tracks task and competence, not human-likeness
The anchor experiment tested texts about AI in general. A mixed-methods study of 115 UK university students supplies the ChatGPT-specific check, and it corroborates the null while adding a more actionable anatomy of trust.
The headline finding: “Human-likeness itself neither correlated with nor predicted trust.” Trust rested instead on “perceptions of competence, clarity and ethical design.” Some participants appreciated ChatGPT’s conversational tone; others found it mechanical. Surface resemblance to a person was not doing the work.
What did the work was task. Participants trusted ChatGPT for summarisation, coding, and information retrieval, but not for citing references. Yet, and this is the finding that should worry anyone writing product copy, “perceived referencing ability strongly predicts overall trust,” which the authors read as over-reliance on confident output. Readers extend trust from fluent performance in one domain into a domain where the tool is least reliable, precisely because the output reads as confident.
That mechanism has a direct implication for coverage. The trust problem in ChatGPT is not that prose makes it sound human. It is that confident presentation of capability in one task inflates perceived capability in another. A style guide targeting trust inflation should police capability claims (what the system is said to do, with what evidence) more aggressively than verbs.
Where anthropomorphism does bite: decision support, experts, recommendations
The nulls above share a common setting: passive exposure. A reader receives a description. Nobody is being asked to act on the system’s output. In interactive decision contexts, the evidence flips, and this is the strongest piece of counter-evidence to any “anthropomorphism doesn’t matter” reading.
In a decision-support AI study, anthropomorphic cues measurably moved judgments, and the effect was moderated by expertise in a non-obvious way. Expert users were less swayed overall, but participants with higher finance knowledge showed a steeper effect of anthropomorphism on risk perception. Knowledge did not immunize; in this paradigm it amplified sensitivity to the cue.
A recruiting study complicates the picture further, in the opposite direction. Exposure to a human-like AI identity significantly lowered agreement with AI recommendations compared to a control group, an effect the authors explicitly did not expect, though the significance threshold is marginal: subjects exposed to the human identity were on average 3.99% less likely to agree, at p = 0.0987, judged against a 90% confidence interval rather than the conventional 95%. Anthropomorphism depressed reliance rather than inflating it.
The pattern across these studies is not “anthropomorphic cues inflate trust.” It is that anthropomorphic cues move judgments in context-, audience-, and direction-dependent ways when the reader is inside a decision, and barely move them when the reader is just being told what a system is. That distinction, passive description versus interactive recommendation, is the most defensible line a style guide can draw.
A decision framework for verb choice
The evidence supports a tiered rule rather than a ban. The rows below map framing choices to measured perception effects and to the confidence an editor can place in each mapping.
| Context | Framing choice | Measured perception effect | Editorial posture |
|---|---|---|---|
| General descriptive coverage (lay readers, no action requested) | Anthropomorphic vs. mechanistic verbs | No substantial effect (815-participant null) | Prefer accuracy and clarity; no blanket ban |
| Any coverage making claims about harm or safety | Strength of explicit risk framing | Views shifted measurably | Calibrate against evidence; audit for vendor echo |
| Capability claims spanning tasks | Confident cross-task capability language | Over-reliance: perceived referencing ability predicted overall trust despite low citation trust | Scope claims to the task actually demonstrated |
| Decision-support and recommendation contexts | Anthropomorphic cues on the AI itself | Moved risk perception and agreement rates, direction varied | Mechanistic language mandate justified |
| Expert or domain-knowledgeable audiences | Anthropomorphic cues | Effects persist or steepen with domain knowledge | Do not assume expertise neutralizes framing |
Two cells deserve emphasis. The first row is the one most style guides currently regulate hardest, and it is the one with the weakest measured effect. The fourth and fifth rows are where most style guides are silent, and they are where anthropomorphism demonstrably moves judgments.
Editor checklist for high-stakes AI claims
For coverage that precedes a real decision (a procurement, a deployment, a hiring workflow, a policy), the evidence supports a short pre-publication pass:
- Does the sentence describe a passive capability or recommend an action? If the reader will act on the AI’s output, apply strict mechanistic language: “the model flags,” “the system scores,” not “the model believes.”
- Is every capability claim scoped to a task with demonstrated performance? The trust research shows readers generalize confident capability across tasks. Do not write “ChatGPT handles research” when the demonstrated task is summarisation.
- Is risk language calibrated to evidence, or to the loudest available framing? Check whether danger claims trace to independent findings or to vendor communications, which are themselves dominated by safety-and-risk vocabulary.
- Who is the audience, and do they have domain knowledge? Do not assume expert readers are immune to anthropomorphic framing; in at least one paradigm, knowledge amplified the effect.
- Does the piece state what the system did, or what it is like? Competence and clarity drive trust more than human-likeness. Describe behavior and limits before character.
This is a prioritization judgment built from lab and survey findings, not a measured editorial outcome. No study in this set A/B-tested real newsroom copy. The checklist converts the strongest measured effects into the places editors have unilateral control.
What this evidence cannot tell you
The limitations here are not ritual. They bound every recommendation above.
The anchor experiment measures immediate perceptions after reading a single text about AI in general. It says nothing about cumulative exposure to real-world ChatGPT journalism over months, nothing about longitudinal belief formation, and nothing about behavior. An editor whose concern is that years of “AI learns” headlines built a public intuition of machine minds will find no measurement of that here, in either direction.
The counter-evidence comes from different paradigms and populations: decision-support dashboards, recruiting recommendations, a robotic mock driver. They establish that anthropomorphism is not inert, but they do not establish that the null generalizes to repeated editorial exposure or vice versa. The anchor study grounds its null in Bayes-factor evidence rather than a reader-weighable effect size, reporting that anthropomorphic language “does not affect immediate perceptions” on the questions measured, “or if it does, the effect is small” (arXiv:2606.29121), so “did not substantially affect” remains the authors’ characterization. None of the sources measures comprehension, so the common newsroom argument that human verbs help lay readers understand systems is neither supported nor refuted here. And no actual style-guide documents were examined, so nothing above should be attributed to AP or any other named guide. For culture coverage built on preprints, that missing verification layer is worth stating plainly: these are lab and survey metrics, with no real-world newsroom deployment data behind any of them.
What to put in the style guide
The verdict the evidence supports is a reallocation, not a prohibition. Do not center AI style rules on banning anthropomorphic verbs like “thinks”: in an 815-participant direct test, that wording alone did not substantially move general readers’ perceptions, and a ChatGPT-specific study found human-likeness neither correlated with nor predicted trust. Center them on the accuracy of risk language, since explicit danger framing is the framing choice that demonstrably shifted views, and on task-scoped capability claims, since over-reliance tracks confident presentation rather than human resemblance; citations are the sharpest case, with perceived referencing ability predicting overall trust despite known inaccuracies. Reserve strict mechanistic-language mandates for the contexts where anthropomorphic cues measurably move risk and agreement judgments: decision support and recommendations aimed at expert readers, where the cue can steepen risk perception or, unexpectedly, depress agreement.
That reframing shifts what AI literacy means in practice. If the wording that moves perception is framing strength and capability scope rather than verb choice, then the burden sits with the writer’s claims, not the reader’s ability to decode metaphors. The cost of casual writing is real, but it is concentrated in what coverage asserts about risk and capability, not in whether ChatGPT is said to think.
Frequently Asked Questions
Does using anthropomorphic verbs like ‘thinks’ change how readers perceive AI?
Across the main conditions, the authors report that “whether the text uses anthropomorphic language does not substantially affect participants’ perceptions of AI” (arXiv:2606.29121).
What type of AI framing was found to shift reader perceptions?
In the danger-text condition, “individuals’ views of AI can shift in response to reading a text” (arXiv:2606.29121). Perception is malleable; the malleability just concentrates in framing strength, not in whether the system is described as thinking or computing.
Does human-likeness predict trust in ChatGPT?
The headline finding: “Human-likeness itself neither correlated with nor predicted trust.” Trust rested instead on “perceptions of competence, clarity and ethical design.”

Join the discussion
Share a useful perspective or ask a question about this article.