groundy
Ethics, Policy & Safety

Should Therapists Use LLMs on Session Transcripts? What the Evidence Supports

Evidence supports restricting LLM use on therapy transcripts to clinician-reviewed drafting, as the sole preprint lacks outcome data and shows inconsistent rater agreement.

Published 7 references
A hollow ceramic hand suspends an irregular ivory leaf with a copper-colored edge above an empty forest-green document tray, casting hard shadows across a warm ivory background.
On this page10 sections

Therapy clinics can defensibly use LLMs on session transcripts for exactly one job today: drafting case conceptualizations that a clinician then reviews, inside HIPAA and GDPR-compliant data handling. The direct evidence, a single unreviewed preprint, shows the technique works as a proof of concept, but its own raters disagreed sharply on output quality and it reports zero treatment-outcome data. Anything closer to treatment selection waits.

What the proof of concept actually did

The study prompting this question is arXiv:2512.05836, a proof of concept observed in late September 2026. The authors describe it plainly: “we developed an end-to-end pipeline for automatically generating client networks to support case conceptualization and treatment planning.” A client network here means a structured map of a patient’s symptoms and how they relate to each other, the kind of artifact a therapist builds during case formulation. The pitch is that an LLM reading full session transcripts can draft that map automatically, saving clinician hours and surfacing symptom clusters a busy caseload might miss.

The scale of the evaluation matters more than the architecture. The authors “annotated 8,028 utterances from 77 therapy transcripts (N = 6),” per the preprint’s abstract. Read that again: six clients. Seventy-seven sessions is a respectable amount of conversational data, but all of it comes from half a dozen people. Symptom expression, therapist style, language, presenting problems and cultural context vary enormously across a real caseload, and six clients cannot sample that variation. This is a feasibility demonstration, the kind that establishes a pipeline can run end to end, not the kind that estimates how often its outputs are right in the hands of a clinician who did not build it.

That distinction, feasible versus clinically useful, is the entire decision.

The reliability problem: the study’s own raters disagreed

The most important number in the preprint is not a performance score. It is a confession about the scores. The authors report that “interrater agreement on model performance metrics was inconsistent, ranging from poor to substantial.” Inter-rater agreement measures whether independent human evaluators, looking at the same model output, reach the same judgment about its quality. When agreement spans poor to substantial across the evaluation, the headline utility metrics are not settled measurements; they are contested judgments that depend heavily on which rater you believe.

This has a direct operational consequence. If the study’s own raters, evaluating model outputs in a research setting, cannot consistently agree on whether the generated networks are good, a procurement team watching a vendor demo has no reliable way to judge output quality either. Consistency between evaluators is also a weaker property than clinical validity: as Groundy’s coverage of mental-health LLM evaluation notes, judges agreeing with each other says nothing about whether any of them agrees with a clinician. Here the evaluators do not even reliably agree with each other.

And on the question that actually matters, whether any of this helps patients, the authors are explicit: “more research is needed to examine whether these networks improve treatment outcomes.” The clinical benefit is unmeasured, not merely unproven by outside skeptics. No reader should let a coherent-looking network diagram substitute for outcome evidence, because producing a plausible artifact is precisely what LLMs are good at, and plausibility is not the thing being tested.

Preprint, not peer review: what arXiv moderation certifies

Every quantitative claim above is author-reported, and the venue guarantees it. arXiv states that submissions go through “a moderation process that classifies material as topical to the subject area and checks for scholarly value,” but “[m]aterial is not peer-reviewed by arXiv - the contents of arXiv submissions are wholly the responsibility of the submitter and are presented ‘as is’ without any warranty or guarantee,” per arXiv’s own description of its process. Moderation screens out off-topic and non-scholarly submissions. It does not check whether the evaluation was well designed, whether the annotation scheme was sound, or whether the statistics support the conclusions. As Wikipedia’s summary puts it, papers are approved for posting after moderation, but not peer reviewed.

None of this makes the study dishonest. Proof-of-concept preprints are a legitimate way to open a research line, and this one is unusually candid about its limits. But candor from authors is not the same as scrutiny from reviewers, and its abstract reports a single proof-of-concept evaluation. The same discipline Groundy applied to author-reported privacy benchmarks applies here: treat every number as a claim to be tested, not a result to plan a clinic around. The preprint may also be revised or withdrawn; its current status is unreviewed.

The privacy perimeter dominates any pilot design

Even a modest pilot, one clinician, one vendor, a handful of transcripts, sits inside strict legal frameworks, because session transcripts are among the most sensitive records a clinic holds.

In the United States, HIPAA (1996) “establishes federal standards protecting sensitive health information from disclosure without patient’s consent,” and its Privacy Rule “address[es] the use and disclosure of individuals’ protected health information (PHI) by entities subject to the rule,” those “covered entities.” Routing a transcript through an LLM API is a disclosure of PHI, which means the vendor relationship, the data-processing terms, retention behavior and any use of transcripts for model training all fall inside the rule’s scope. The Privacy Rule itself was published in final form on December 28, 2000, after HHS received over 52,000 public comments, per HHS’s regulatory history; it is a mature framework with settled expectations, not a gray area a startup can improvise through.

In the EU, the GDPR has applied in all member states since May 25, 2018, and the financial exposure is severe enough to shape pilot design by itself. According to GDPR.eu’s overview, penalties come in two tiers “which max out at €20 million or 4% of global revenue (whichever is higher),” and the site warns that after a breach “you have 72 hours to tell the data subjects or face penalties,” its own paraphrase of the regulation’s breach-notification window. A misconfigured vendor integration is not an IT incident at those stakes; it can be a reportable event on a three-day notification clock.

A limitation on this article’s scope deserves stating. The discussion above establishes the general rules: HIPAA’s covered-entity PHI handling, GDPR’s applicability, fines and breach window. It does not analyze the provisions most specific to this use case, namely GDPR Article 9’s conditions for processing special-category health data and HIPAA’s distinct treatment of psychotherapy notes, where HHS’s Privacy Rule summary states that “a covered entity must obtain an individual’s authorization to use or disclose psychotherapy notes,” with enumerated exceptions. Both are likely to be load-bearing in any real pilot, and counsel must confirm both against primary regulatory text before transcripts move anywhere. Groundy’s earlier analysis of where privacy law meets engineering makes a related point: compliance obligations get worked out in contracts and code, not on summary pages.

What procurement should demand from vendors

Vendor demos in this category tend to show one sample transcript producing one tidy network. Given the evidence above, that demonstrates almost nothing. A checklist that accepts “works on a sample” misprices both clinical risk and vendor maturity. Before any pilot, procurement teams should demand:

  • Inter-rater reliability data from an independent evaluation. Not internal quality scores. The one public study reviewed here could not get consistent agreement from its own raters; a vendor claiming high utility should show agreement statistics from evaluators with no stake in the sale, on transcripts resembling your caseload.
  • Treatment-outcome evidence, or an honest admission that none exists. This proof of concept reports no treatment-outcome data, and its authors state more research is needed; a vendor implying its tool improves outcomes is ahead of this evidence.
  • A compliance posture anchored to primary text. For US deployments, how PHI flows under the Privacy Rule, including whether transcripts train models and under what data-processing terms. For EU exposure, how Article 9 conditions are met and how the 72-hour breach clock would be honored. Answers citing marketing pages instead of regulatory text are disqualifying.
  • A workflow that keeps a clinician between the model and the chart. The defensible use today is drafting support, so the product should make review, edit and rejection of generated content the default path, not an optional step.

The decision: pilot, restrict, or wait

OptionWhat it meansWhat the evidence supportsExposure
Full pilot with treatment-adjacent useLLM outputs inform treatment planning decisionsNot supported: no outcome data, contested reliability, N = 6Clinical risk plus full HIPAA/GDPR exposure with no demonstrated benefit
Restricted to clinician-reviewed draftingLLM drafts conceptualizations and flags symptom clusters; clinician verifies before anything enters the recordSupported as a bounded inference from the proof of concept: the pipeline generates drafts, humans catch errorsPrivacy obligations still apply in full; drafting utility in production is unmeasured
Wait for outcome and replication evidenceNo transcript processing by LLMsConservative but unnecessary if drafting support has independent value to your cliniciansOpportunity cost of staff time saved (unquantified in the evidence)

The table’s middle row is the verdict: restrict LLM analysis of session transcripts to clinician-reviewed drafting support, inside HIPAA/GDPR-compliant data handling. Do not use outputs for treatment selection until outcome evidence exists, and demand inter-rater reliability and outcome data from vendors. This is an inference drawn from the proof of concept’s own stated limits, not a demonstration that drafting workflows are safe or effective in production; it is the most use the current evidence can carry, and no more.

Three findings would justify revisiting the boundary: peer-reviewed replication with consistent inter-rater agreement, a controlled study showing generated networks improve treatment outcomes, and vendor evaluations conducted by independent raters on representative caseloads. Until then, the strongest limitation governs everything above: it all rests on one author-reported, non-peer-reviewed preprint built on six clients, with internally inconsistent quality ratings, and the regulatory specifics that matter most (Article 9 conditions, psychotherapy-note treatment) remain to be verified against primary text. A demo-grade coherent network is a demo. Price it accordingly.

Frequently Asked Questions

What is the only defensible use of LLMs on session transcripts today?

Therapy clinics can defensibly use LLMs on session transcripts for exactly one job today: drafting case conceptualizations that a clinician then reviews, inside HIPAA and GDPR-compliant data handling.

How many clients were included in the study’s evaluation?

The authors “annotated 8,028 utterances from 77 therapy transcripts (N = 6),” per the preprint’s abstract. Read that again: six clients.

What are the maximum GDPR penalties for a breach?

According to GDPR.eu’s overview, penalties come in two tiers “which max out at €20 million or 4% of global revenue (whichever is higher),” and the site warns that after a breach “you have 72 hours to tell the data subjects or face penalties,” its own paraphrase of the regulation’s breach-notification window.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. arXiv:2512.05836arxiv.orgAccessed
  2. arXiv's own description of its processinfo.arxiv.orgAccessed
  3. Wikipedia's summaryen.wikipedia.orgAccessed
  4. HIPAA (1996)cdc.govAccessed
  5. HHS's regulatory historyhhs.govAccessed
  6. GDPR has applied in all member states since May 25, 2018gdpr-info.euAccessed
  7. GDPR.eu's overviewgdpr.euAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy