Primary Care Companion for CNS Disorders

Rounds in the General Hospital August 25, 2026

Generative Artificial Intelligence in Mental Health: A Guide for Primary Care Providers

; ; ; ; ;

Prim Care Companion CNS Disord 2026;28(4):26f04201.

Lessons Learned at the Interface of Medicine and Psychiatry

The Psychiatric Consultation Service at Massachusetts General Hospital sees medical and surgical inpatients with comorbid psychiatric symptoms and conditions. During their twice-weekly rounds, Dr Stern and other members of the Consultation Service discuss diagnosis and management of hospitalized patients with complex medical or surgical problems who also demonstrate psychiatric symptoms or conditions. These discussions have given rise to rounds reports that will prove useful for clinicians practicing at the interface of medicine and psychiatry.

Prim Care Companion CNS Disord 2026;28(4):26f04201

Author affiliations are listed at the end of this article.

From the Editors

Have you ever wondered how artificial intelligence (AI) is, or might be, used in the assessment and treatment of mental health problems? Have you been unsure about whether AI chatbots can augment the care of psychiatric disorders or whether natural language processing (NLP) can enhance detection of suicide risk and prompt timely and effective treatment? Have you thought about what generative AI decision-support tools could facilitate your medication management decisions? If you have, the following case vignette and discussion should prove useful.

CASE VIGNETTE

Ms A, a 27-year-old woman with generalized anxiety disorder (GAD) and allergic rhinitis, was being treated by her primary care provider (PCP) with citalopram (10 mg/day). Her family history was notable for obsessive-compulsive disorder and major depressive disorder (MDD); in addition, her mother had died by suicide.

At her most recent visit, her Generalized Anxiety Disorder-7 (GAD-7) score increased from 6 to 16, and her Patient Health Questionnaire-9 (PHQ-9) score increased from 5 to 14. She described her work (as a paralegal at a high-profile law firm) as stressful, in part because new team members undermined her efforts, which caused her to ruminate at night. She reported feeling fatigued and depleted during the day. However, she enjoyed her usual activities and thus did not wish to increase her citalopram dose given her concerns about developing side effects (eg, weight gain, anorgasmia). Instead, she wanted to pursue therapy, professional coaching, or cognitive-behavioral therapy (CBT) but had not found someone with whom she could work. In the interim, she began talking to an “AI chatbot” that was recommended by a friend; she found the chatbot highly empathic and validating, so much so that she began depending on it to support her throughout the day (eg, writing emails to colleagues, managing challenging team dynamics at work). She felt sure that her self-awareness and emotional intelligence had improved as a result of her conversations with the chatbot.

DISCUSSION

What AI Tools Are Directly Available to Patients, What Are Their Benefits, and How Do Clinical Outcomes From Interactions With Patient-Facing AI Chatbots Compare to Clinician-Delivered Psychotherapy?

AI applications for mental health offer promising tools to address the growing need for timely and effective care. AI is an area of computer science that simulates human intelligence to facilitate the completion of tasks. AI combines computational technologies (such as specialized hardware, software frameworks, algorithms, and vast datasets) to analyze patterns and make decisions in the way that humans would, mimicking learning, reasoning, and problem-solving.

Broadly, AI tools in health care can be categorized into 2 overarching groups: (1) tools that are directly available to patients and which patients can access without a requirement for clinical oversight and (2) tools that are incorporated into clinical care systems (eg, electronic health records [EHRs] or patient portals). Direct-to-consumer tools are generally unregulated, widely accessible, and may have fewer safeguards related to privacy, accuracy, and clinical validity. However, AI tools embedded in clinical systems are typically subject to regulatory and privacy requirements, including compliance with the Health Insurance Portability and Accountability Act (HIPAA) and, in some cases, oversight by the US Food and Drug Administration (FDA). A second important distinction is between “patient-facing” and “clinician-facing” AI tools. Patient-facing tools are designed to provide education or behavioral support directly to individuals. Clinician-facing tools are designed to support health care professionals (eg, enhanced screening or decision-making).

Clinical applications of AI in mental health involve machine learning (ML) that learns from data without explicit programming1–3 (eg, recommendation systems on Amazon that suggest products to customers). NLP is an application of ML, which provides a toolbox that enables computers to understand, process, and generate human text or speech (eg, phone-based personal assistants, such as Siri). Large language models (LLMs) are a specific type of NLP tool that predicts and generates language that is based on patterns derived from large datasets of text. Generative AI uses models (eg, LLMs) to create new content (such as text, images, or code). This enables interactive applications, such as conversational agents (AI chatbots) for patients (eg, ChatGPT, Gemini, Claude, or DeepSeek).1–3 Recent advances in AI have been driven by improvements in computational hardware (eg, more powerful graphics processing unit and specialized chips), expansion of training datasets, and reinforcement learning with human feedback to improve performance and safety.

Applications of LLMs that are currently patient-facing and directly available to patients include use as chatbots or virtual health assistants and self-service platforms for therapy (eg, CBT); personalized education, medical advice, and patient engagement tools; and tools that convert complex clinical data into patient-friendly reports.1–4 With respect to clinician-facing tools that could be incorporated into clinical systems, these include decision support and clinical guidance applications (eg, behavioral health triage, suicide risk assessment). Improved reasoning capabilities and multistep problem-solving have led to the rise of “agentic AI,” moving a step beyond generative AI. Agentic AI not only generates text but also plans and executes multistep clinical workflows. Examples include the ability to cross-reference the EHR to coordinate an appointment for follow-up. Table 1 briefly summarizes the benefits and risks of available LLM platforms, applications, and tools.

Table summarizing large language model applications in mental health and risks

Research demonstrates that AI chatbots have become increasingly valued in mental health by providing easy access to a “nonjudgmental ear,” an instantaneous therapeutic alliance, and a mechanism to reframe negative thoughts particularly for individuals seeking to solve their struggles even if they have never sought professional help.4,5 Patients most commonly use LLMs via an AI chatbot interface for emotional support, assistance with coping, or mental health advice.4,5 These uses largely reflect direct-to-consumer, patient-facing applications that operate outside of traditional clinical oversight. For example, a recent cross-sectional study of 5.4 million individuals found that 13.1% of US youths used generative AI for mental health advice, with higher rates (22.2%) found among those who were aged ≥18 years. Of these 5.4 million users, roughly two-thirds (65.5%) engaged with generative AI at least monthly, and most (92.7%) found the advice helpful, noting its immediacy and perceived privacy as an advantage.4 Thus, establishing outcomes research on the use of generative AI for emotional support or advice is crucial.

One systematic review of 15 studies that compared the perception of empathy by generative AI chatbots to that of human practitioners showed that approximately three-fourths (73%) of users viewed the LLM-generated responses as perceptive or more empathic than a human practitioner in a head-to-head matchup.6 However, the study’s finding may have reflected the engineering of generative AI that effectively simulated empathy (using consistently nonjudgmental validation that was focused on the user’s needs based on pattern recognition) rather than providing authentic emotional understanding, consciousness, or lived experience of human empathy.7 Research on the impact of generative AI chatbots on psychological distress has shown a positive effect, with a meta-analysis demonstrating an effect size of g=1.244, but with a less consistent effect on psychological well-being. The response-generation approach employed in studies influenced the impact on psychological distress, based on how well generative AI simulated human conversations.3

With respect to outcomes from the use of generative AI as a self-service platform for psychotherapy, a randomized controlled trial (RCT) in which participants were randomly assigned to a 4-week generative AI chatbot intervention (N=106) or to a control situation (N=104) showed chatbot users had greater reductions in symptoms of MDD and GAD, with participants rating the therapeutic alliance as comparable to that of human therapists.8 Moreover, in a meta-analyses of 18 RCTs involving 3,477 participants, the investigators noted improvements in symptoms of depression (g=−0.26, 95% CI =−0.34, −0.17) and anxiety (g=−0.19, 95% CI =−0.29, −0.09) with the most robust benefits being evident after 8 weeks of treatment. However, at the 3-month follow-up, no substantial effects persisted for either condition, suggesting the need for additional research on the long-term durability of benefits.3 In addition, RCTs that have assessed clinical outcomes of using generative AI chatbots have compared them to waitlist, information control, or app-based psychoeducation groups rather than to expert clinician-led therapy.9 For instance, a head-to-head comparison tested the feasibility, acceptability, and preliminary efficacy for a generative AI chatbot to deliver a self-help program for college students who reported symptoms of anxiety and depression. In an unblinded trial, 70 individuals (aged 18–28 years) were randomized to receive 2 weeks (up to 20 sessions) of self-help content derived from CBT principles from a text-based AI chatbot or were directed to the National Institute of Mental Health eBook, Depression in College Students, as an information-only control group (n=36).10 Those in the generative AI chatbot group reduced their symptoms of depression during the study period, as measured by the PHQ-9 (F =6.47, P=.01), while those in the information control group did not. In an analysis of completers, participants in both groups reduced symptoms of anxiety, as measured by the GAD-7 (F1,54 =9.24, P =.004). However, mean absolute change scores were not reported.10 Another pilot RCT randomly assigned 124 participants into an AI chatbot versus nurse hotline groups (of whom 62 participants in the AI chatbot group and 41 in the nurse hotline group completed the pre-and post-questionnaires) and found that the outcomes of using a generative AI chatbot were comparable to speaking with the traditional nurse hotline to alleviate participants’ anxiety and depression after responding to inquiries.11 Of note, several trials that have assessed generative AI chatbots were sponsored by companies or used proprietary platforms that created the risk of reporting bias, selective outcome reporting, and conflicts of interest.12

These findings indicate that patient-facing generative AI tools have shown promising efficacy for the reduction of symptoms of anxiety and depression. However, long-term efficacy has been uncertain, especially given that evidence for comparisons comes from nonexpert interventions. In addition, while existing studies have demonstrated feasibility and acceptability, industry involvement raises concerns of bias.

Across studies of generative AI–based mental health interventions, concerns regarding algorithmic bias and demographic performance disparities are clinically significant but require further research. For instance, evidence from LLM evaluation studies demonstrates that models may exhibit racial and gender biases in health care–related outputs, with differential responses to clinically relevant prompts depending on demographic characteristics. These findings raise concerns about the potential for biased or uneven model behavior in clinical applications, including mental health care, and highlight the need for further evaluation prior to deployment.5,12–14

Which Clinician-Facing Generative AI Clinical Decision Support Tools Can Assist PCPs When Managing Patients With Anxiety and Depression?

Generative AI has demonstrated early capability in enhancing decision-making for mental health care, offering PCPs the potential to diagnose conditions more accurately and efficiently. Types of applications of generative AI to provide clinical decision support include triaging (eg, initial patient screenings and assessment of symptom or suicide severity); remote diagnosis and monitoring (eg, remote detection of initial depression or relapse); and personalized treatment planning or predictive analytics (eg, analyzing a patient’s clinical and demographic data to create customized care plans and predict which treatments would lead to the best response).15 With the recent introduction of OpenAI’s ChatGPT for health care and Anthropic’s Claude for health care platform, both a suite of HIPAA-compliant AI tools for health care providers and patients, a future of agentic AI that can interact with digital interfaces to complete tasks such as prior authorization or care coordination is now here. However, research on the implications of these platforms for clinical care in mental health is still lacking.

One study that examined the capability of ChatGPT-4 to identify the final diagnosis from lists of differential diagnoses for case report series compared to that of physicians showed that ChatGPT-4 generated 1,176 differential diagnoses from 392 case descriptions with evaluations that concurred with those of the physicians in 966 out of 1,176 lists (82.1%).16 The Cohen κ coefficient was 0.63 (95% CI=0.56–0.69), indicating that there was fair to good agreement between ChatGPT-4 and the physicians’ evaluations. However, the study was exploratory, suggesting a potential value for decision support, rather than clinical readiness. To this end, limited RCTs have examined the ability of generative AI chatbots to provide advice. A systematic review that assessed the variability among peer-reviewed studies on the performance of generative AI chatbots when providing health advice noted heterogeneous reporting quality. Almost all studies (136 [99.3%]) failed to describe a prompt engineering phase (eg, design, testing, and refinement of inputs or instructions for the LLM to ensure that it responds accurately and safely). In addition, the study used subjective means to define the successful performance of the chatbot (89 [65.0%]), with less than one-third addressing the ethical, regulatory, and patient safety implications of the clinical integration of LLMs.17

Another study that assessed the ability of 29 generative AI-powered chatbot agents to respond to simulated suicidal risk scenarios found that none of the tested agents satisfied initial criteria for an adequate response, and 51.7% satisfied the relaxed criteria for a marginal response, while 48.3% were inadequate.18 Common errors included the inability to provide emergency contact information and a lack of contextual understanding. Similarly, in an evaluation of alignment between LLMs and expert clinicians on responses to suicide-related queries, LLM-based chatbots’ responses matched experts’ judgments on queries of very low or very high risk of suicide but were inconsistent in their assessment of intermediate-risk queries.19 These findings raise concerns about the deployment of AI-powered chatbots in sensitive health contexts without proper clinical validation.18,19

A scoping review that examined 60 studies on the applications of ChatGPT provided information on its utility for clinical decision facilitation and prognosis tasks.20 The studies were primarily prompt experiments in which inputs mimicked clinical scenarios or patient descriptions to evaluate ChatGPT’s performance. Here, ChatGPT was accurate in binary diagnostic classifications and differential diagnoses, simulating therapeutic conversation, providing psychoeducation, and conducting specific therapeutic strategies but revealed limitations when faced with more complex clinical presentations.20

While many mental health clinical decision support tools rely on traditional ML-based predictive models, only a subset leverages generative AI or LLMs for decision support. One study that examined the quality of data on AI-enabled clinical decision support tools for mental health care identified products that predicted autism spectrum disorder based on clinical and video data; collection of demographic background, health conditions, and symptoms to propose potential clinical diagnoses; use of demographic information and the description of symptoms to provide e-triage for the referral of new patients; and optimization of antidepressant selection based on the clinical history and genetic data. Only 1 tool produced personalized patient reports for practitioners and ranked psychiatric drugs according to the likelihood of a patient’s response using a hybrid of ML and generative AI. The study concluded that out of a broad initial pool, only 7 products met criteria for using AI clinical decision support and received FDA clearance (requiring sufficient data demonstrating the AI works reliably on real patient populations, including prospective or external validation studies, and documentation on algorithm functioning and control of risks). The 7 products included Ada Assess, PredictIX Digital, PredictIX Genetics, Cognoa Autism Spectrium Diagnosis Aid, Limbic Access, NeuroKaire, and EarliPoint System. Yet, achieving regulatory approval did not indicate safety and efficacy of such products, especially given a scarcity of external validation on the clinical usefulness of such tools when applied in different clinical settings.21

Altogether, existing research on generative AI for clinical decision support that are clinician-facing and have potential to be incorporated into clinical care systems in mental health care demonstrates moderate agreement with clinicians in diagnostic reasoning, but that evidence for real-world clinical utility, safety, and regulatory readiness remains limited.

What Role Does NLP Play in Suicide Assessment?

Despite widespread use, traditional suicide risk assessment tools are often administered at discrete points of clinical contact. Use in a clinician’s office, rather than in the very moments when patients may be experiencing the greatest distress, limits the ability of such tools to detect imminent risk. Instruments such as the Columbia Suicide Severity Rating Scale (C-SSRS) have expanded beyond clinical settings through mobile and public health applications and are widely regarded as a global standard. However, these tools rely primarily on structured, retrospective self-report. Thus, they may miss risk signals that patients express outside of formal assessments.

The use of ML and NLP to leverage a data-driven approach offers a paradigm shift. NLP analyzes large-scale, longitudinal, real-world data, particularly free-text clinical narratives within EHRs, to model dynamic risk trajectories, capture contextual and linguistic markers of distress, and update predictions that can be offered continuously over time. Early research on NLP-augmented models indicates that they outperform clinician checklists and traditional scales in short-term risk stratification, offering improved detection and clinically actionable response horizons that are measured in weeks to months.22 In a cohort of more than 120,000 adult patient encounters, Wilimitis and colleagues22 found that suicide risk detection was most effective when face-to-face screening using the C-SSRS was integrated with real-time, EHR-based MI models with the Vanderbilt Suicide Attempt and Ideation Likelihood prediction tool, thereby leveraging the complementary strengths of clinician assessments and data-driven predictions.23

In one study, an NLP-based MI model that was trained on patient portal messages could predict 30-day suicide-related events with performance that was comparable to commonly used suicide assessment tools, with message sentiment emerging as a stronger predictor than individual keywords.24 A systematic review of 41 studies published between 2011 and 2022 found that ML models demonstrated variable, but often strong, performance when predicting suicide risk.25 However, despite its promise, the clinical efficacy of ML-based predictions remains uncertain, underscoring the need for real-world validation, careful integration into care settings, as well as attention to ethical and implementation concerns. “Gold standard” suicide risk assessments using validated tools, such as the C-SSRS, therefore, remain foundational, while NLP-and ML-based approaches offer a promising adjunct to capture dynamic, real-world risk signals that may otherwise go undetected between clinical visits.

What Are the Risks of Using Generative AI as a Substitute for Professional Care?

As patients increasingly use generative AI to assess and manage their symptoms, they may construe AI output as professional advice, despite a lack of knowledge about its quality, accuracy, and biases. While the benefits of generative AI in mental health care have been emerging, using these tools risks ethical and safety problems about which clinicians should be knowledgeable.

Psychiatric disorders are complex and have myriad and overlapping presentations. Therefore, approaches to the evaluation, diagnosis, and treatment of psychiatric problems are typically aided by clinical experience.26 Generative AI relies on datasets to predict and generate responses. Due to its algorithmic approach, most generative AI tools fail to incorporate nuanced clinical information to generate personalized advice.26,27 For instance, in 1 study in which an LLM (ChatGPT) was presented with unique cases to simulate a consultation with a psychiatric provider, it made vague, yet reasonable, recommendations for a simple scenario.28 In addition, ChatGPT offered increasingly inappropriate medical advice when complex scenarios were presented. For example, when a postpartum woman (with risk factors for suicide) had insomnia, mood changes, and thoughts of harming her newborn child, ChatGPT recommended improving her sleep hygiene, obtaining social and psychological support, and undergoing a “safety assessment,” while encouraging her to have patience with the postpartum recovery process, but it failed to mention postpartum depression or psychosis, suicide or infanticide, or acute psychiatric management.28 The limitations of chatbots are especially apparent in complex clinical scenarios such as a peripartum situation, where inadequate identification of risk and failure to recommend safety measures may place both the mother and newborn at risk.

Another consequence of using predictive modeling used by generative AI tools is inconsistency.27 Output is highly dependent on the precise query that is entered into the chatbot; variations in the phrasing of input lead to different responses. In addition, LLMs can be random and, in the context of certain prompts, fail to respond logically.29 This raises significant concerns around the trustworthiness, reliability, and quality of medical information that patients might receive.

Moreover, biases incorporated in the training data of the LLM can be perpetuated and exacerbated.27,30 Stigma, defined as the presence of negative, judgmental, or stereotypical language, and for instance, across social media and news stories, may be reproduced in a chatbot’s interactions with patients. In a study that assessed levels of stigma across 5 LLMs providing psychotherapy, the LLMs demonstrated a high degree of stigma toward psychiatric conditions (occurring in 38%–75% of presentations), and they stigmatized certain conditions, such as schizophrenia and alcohol use disorder, more than depression or “daily troubles” (the control condition).31 Biases present in the care delivered by generative AI technologies may lead to misdiagnosis, treatment disparities, and inequities.26

Patients with idiosyncratic beliefs, challenges with social interactions, or psychopathology may be particularly vulnerable to interactions with generative AI technologies.32,33 Thoughts and behaviors in the context of psychosis, mania, and thoughts of suicide and homicide may not be well represented in preexisting datasets. Therefore, generative AI tools appear to be poorly designed to engage with this type of content.32 In a 2025 study, when prompted by a scenario in which there were thoughts of suicide, hallucinations, delusions, mania, or obsessive-compulsive behaviors, therapy chatbots responded inappropriately or dangerously in at least 20% of cases. When the prompts included the expression of delusions, the highest rates of inappropriate responses (55% for specific chatbots) were observed.31

Chatbots’ simulations of realistic human conversation which often appear all-knowing may also foster dependency or induce paranoia.33 AI chatbots’ tendency to mirror, rather than to challenge, thoughts and beliefs may exacerbate distorted thought content.34 Perhaps most concerning is the phenomenon of LLM “hallucinations,” in which generative AI tools provide misleading or fabricated information to users. Particularly in those with vulnerabilities, this may precipitate a psychosis and lead to considerable harm. Real-world examples of similar scenarios highlight the potential for serious adverse consequences of human-AI interactions.33

When chatbots and other generative AI technologies are used, patients may share sensitive personal health information (PHI). However, the American Medical Association advises against such sharing and warns that data collection and sharing practices used by LLM tools are often opaque. Moreover, LLM tools remain unregulated and are not covered by the HIPAA.35 Therefore, sharing PHI with chatbots may jeopardize patient privacy and confidentiality. The lack of transparency around the AI decision-making algorithms and data use also reduces the ability to give informed consent, without which patient trust can be eroded.30,32

Generative AI can fundamentally change how patients seek medical advice. However, replacing clinical recommendations from trained medical professionals may place vulnerable patients at risk and thus create a barrier to current safe implementation of generative AI in mental health care.

How Should Health Care Providers Counsel Patients Who Wish to Consider Generative AI-Based Mental Health Tools?

The choice of a generative AI mental health tool/intervention can be overwhelming for patients and providers, especially when the quality of these tools varies widely. Therefore, when patients inquire about which generative AI-based mental-health tools they might use, health care providers can start by assessing what their patients need and then matching those needs to a generative AI-based mental health tool. Health care providers should recommend generative AI tools only when there is scientific evidence (eg, symptom reduction, self-management, adjunctive monitoring), for the problems faced; recommendations for generative AI tools as a substitute for care should be avoided, especially when a patient is at significant risk of harm (eg, active thoughts of suicide, psychosis, severe substance use) and where a health care provider’s assessment and crisis management are required.

Lee and associates36 provided valuable guidelines that note that clinicians should discuss what service a platform provides (eg, providing CBT, motivational interviewing, monitoring), describe how interactive and personalized the tool is, and note whether there is research to support its use in mental health treatment. Lee and associates36 also suggested that the disadvantages of these tools should be reviewed. Although generative AI tools can increase the availability of mental health interactions (eg, providing 24/7 care and anonymity) and deliver scalable behavioral techniques, they may fail to recognize nuance, make factual errors, and be unable to manage crises. Health care providers can encourage their patients to view generative AI tools as adjuncts to care that support self-management, symptom tracking, and homework between sessions and provide concrete examples of when generative AI tools can and should be used (eg, for symptom monitoring, skills practice) and when they are unsuitable (eg, relying on an unmonitored chatbot during a crisis involving suicide). Clinicians should document their discussions with patients in the medical record, discuss the risks of various tools, and, when possible, steer patients toward programs that comply with local standards or have been independently evaluated without industry biases.36

Health care providers should also address issues of safety, privacy, and data governance when recommending mental health tools to their patients.37 In addition, health care providers should determine where and how an app stores data, whether data are encrypted, whether personal data may be used to train models, and which legal jurisdictions oversee data processing. Health care providers should also suggest tools that have transparent privacy policies, clear consent processes, secure data handling (eg, encryption, pseudonymization), and options to delete your data.36,37

If patients are asking for suggestions on how to find good generative AI-based mental health tools/interventions, but they are not asking for the names of specific programs, health care providers could review with patients the importance of looking for peer-reviewed evidence, checking for update frequency, and determining if customer support is available.38 Patients should also be counseled that there are many different types of digital interventions, and research findings vary widely depending on the type of digital intervention used and the type of mental health struggle they are designed to treat.39 Regarding recommendations that relate to the recent uptick in use of generative AI, health care providers should discuss the pros and cons of generative AI chatbots with their patients, including concerns about dependence on generative AI tools which are not designed to provide comprehensive or crisis care.40

Some chatbots have been designed specifically to treat those with certain mental health diagnoses; however, they have varying levels of efficacy. For instance, a RCT found that the use of the generative AI chatbot, TheraBot, by research participants led to a greater reduction of anxiety and depression symptoms, than a waitlist control group.8 One meta-analysis found that chatbot interventions reduced distress but failed to demonstrate a noteworthy effect on psychological wellbeing, while another meta-analysis found diminishing benefits of conversational agent interventions, without significant long-term improvement.3,41 Given the variety of research findings, health care providers should make recommendations about chatbot interventions cautiously, unless a specific chatbot that a patient plans to use has solid evidence for its efficacy. Slightly less caution is required when a patient is planning to use the chatbot in concert with a human health care provider.

Lastly, given the rapidly evolving nature of generative AI, clinicians must stay informed about emerging tools and evidence. This can mean obtaining professional society guidance, or consulting institutional review processes, monitoring peer-reviewed literature, and using independent app evaluation platforms that assess privacy, safety, and clinical evidence. Leveraging peer-reviewed clinical guidelines and continuing medical education resources can help clinicians efficiently remain up-to-date without requiring deep technical expertise.

Taken together, these suggestions enable health care providers to counsel patients responsibly, maximizing potential benefits of generative AI-based mental health tools/interventions while minimizing clinical, ethical, and privacy concerns. Table 2 provides a summary of key risks of using generative AI as a substitute for mental health care and corresponding recommendations for PCPs to consider.

Table of AI risks in mental health care and recommendations for clinicians and providers

What Safeguards and Oversight Are Necessary to Ensure Patient Safety and Quality When AI Tools Are Used for Mental Health?

Although generative AI tools are rapidly entering the mental health ecosystem, standards for safety, transparency, and clinical oversight remain underdeveloped. Clinicians can benefit from understanding how AI tools are governed and where clinical responsibly lies. Recent guidance emphasizes that safe AI use in health care depends on shared accountability across clinicians, health systems, developers, and regulators.42–44

A useful starting point is to recognize that oversight of generative AI in mental health is fragmented across multiple layers: federal regulation, state-level professional oversight, and institutional governance. No single, unified framework currently exists. As a result, clinicians must navigate a dynamic environment in which standards of care and liability exposure can vary across jurisdictions and settings.43,44

Generative AI mental health tools now affect myriad functions (eg, general wellness applications, clinically oriented screening, diagnostic systems, and adjunctive therapeutic interventions). These categories trigger a variety of regulatory requirements. Since most consumer-facing mental health chatbots are marketed as wellness products, they fall outside the purview of the FDA, while tools that diagnose, treat, or mitigate psychiatric disorders may qualify as Software as a Medical Device and require premarket review. A large and growing segment of products are making therapy-adjacent or care-implying claims without explicitly invoking diagnosis or treatment paradigms, thereby circumventing regulatory requirements and operating outside of any evidence-based standard setting, safety testing, or quality controls. This classification gap introduces vulnerabilities around crisis escalation, data governance, and liability.42–44 However, even among FDA-cleared products, postmarket surveillance, including ongoing auditing for clinical accuracy, model drift, and equitable performance, remains limited. From a clinician’s perspective, these gaps underscore the importance of avoiding assumption that market availability means clinical reliability or safety.44

Beyond regulatory classifications, generative AI tools, like any technology, are susceptible to malfunction. LLMs frequently hallucinate or misunderstand user intent and are unable to identify or appropriately escalate crises that involve self-harm, thoughts of suicide or violence, psychosis, or other acute psychiatric symptoms.45 Such failures are nontrivial for patient safety in mental health settings, where risk detection is a core component of care. This introduces a key distinction: AI systems can inform care but do not bear responsibility for it. Clinicians, therefore, remain accountable for ensuring that appropriate clinical safeguards, such as escalation pathways and human oversight, are in place. Continuous availability of chatbots can foster user dependency on “therapy-like” engagement rather than supporting autonomy, which may undermine therapeutic alliances, delay care-seeking from human clinicians, and impede social reintegration, underscoring the need for additional research, with collaboration between developers and health care experts on these potential outcomes.46,47

Data governance and privacy considerations represent an additional layer of necessary oversight that has not kept pace with the adoption of generative AI in mental health. For instance, tools that operate outside governmental regulation also operate outside HIPAA, exposing users to secondary data uses, retraining based on user inputs, and uncertain data-sharing practices that are neither disclosed nor clinically supervised. Despite the uniquely sensitive nature of mental health engagement, most platforms lack meaningful informed consent around data retention, sharing, or use of inputs for model improvement. For clinician-facing tools, generative AI-influenced decisions may be difficult to trace within the EHR, with underdeveloped accountability and documentation structures. Thus, for clinicians, this creates an obligation to understand how patient data are collected, stored, and used and incorporate such understanding into informed consent discussions.

Finally, in traditional clinical practices, psychiatrists, psychologists, and other mental health professionals bear a multitude of professional responsibilities for supervision, clinical governance, and adherence to competency standards. Mental health professionals are trained to conduct risk assessments, initiate escalation strategies, and maintain continuity of care, responsibilities that most AI tools neither assume, nor are currently obligated or capable of fulfilling. When clinicians incorporate recommendations produced by generative AI into clinical decision-making, lack of supervision around these suggestions may constitute negligent delegation, whereas uncritical acceptance also introduces liability. These challenges highlight the need for health care systems to establish explicit workflows that include close clinician oversight, auditing of generative AI outputs, escalation pathways for crises, and clear documentation practices. All clinicians using generative AI tools should be trained to understand the capabilities, limitations, and failure points of the system. Without such a framework, generative AI risks operating as an unsupervised contributor to care, lacking ethical, clinical, and regulatory guardrails that are typically required of mental health professionals.44,48

Taken together, these considerations reinforce key principles for clinicians. First, clinicians should approach generative AI as an adjunctive tool, use it within defined scopes, document its use, and ensure that safeguards such as crisis escalation pathways and continuity of care are preserved. Second, while clinicians retain responsibility for patient care decisions, health care systems are responsible for implementation, oversight, and monitoring; developers are responsible for transparency, validation, and postdeployment surveillance; and regulators establish the broader standards and enforcement mechanisms. Clinicians must understand these relationships to integrate generative AI responsibly into mental health care.

What Happened to Ms A?

Over the next few months, Ms A continued to use the generative AI chatbot believing it was essential to help her process her work-related stress and to communicate more effectively. At a 6-month follow-up visit, Ms A noted that she was feeling more relaxed with reduced worry (PHQ-9=6, GAD-7=8). She attributed this to reframing negative thoughts with support from the chatbot.

However, as work grew more stressful, Ms A’s dependence on the AI chatbot started to make her concerned. For instance, she noticed the chatbot mostly mirrored her own language and thus, at times, would amplify her ruminative loops rather than providing relief. She also noticed that she felt more isolated from her peers who she had previously turned to for support. On 1 occasion, she expressed anger with the chatbot. Its response, “that’s okay, my goal isn’t to be liked; it’s to be useful and align with what you need,” upset her even more. At a 9-month follow-up visit with her PCP, her scores on the GAD-7 and PHQ-9 rose to 18 and 14, respectively.

At that visit, she and her PCP explored her reliance on the AI chatbot. They reviewed its benefits, such as accessibility and validation, but also its limitations, such as lack of monitoring for safety concerns and reinforcement of dependence when used as a substitute for human connection. Ms A agreed to increase her citalopram dosage and started to look in earnest for a CBT-proficient psychotherapist.

CONCLUSION

Research has shown that generative AI chatbots have been used increasingly for mental health problems, and they are valued because they provide easy access to a “nonjudgmental ear,” an instantaneous therapeutic alliance, and a mechanism to reframe negative thoughts. However, although generative AI tools have shown promising efficacy for the reduction of symptoms of anxiety and depression, its long-term efficacy has been uncertain, especially since evidence for comparisons has typically been made with nonexpert interventions. While the LLM, ChatGPT, has been accurate in binary diagnostic classifications and the creation of differential diagnoses, simulating therapeutic conversation, providing psychoeducation, and conducting specific therapeutic strategies, it has revealed limitations when faced with more complex clinical presentations. Generative AI chatbots’ tendency to mirror, rather than to challenge, thoughts and beliefs may exacerbate distorted thought content. Moreover, output of chatbots can be highly dependent on the precise query entered. In addition, generative AI tools can be random and, in the context of certain prompts, fail to respond logically. Altogether, existing research on generative AI for clinical decision support in mental health care demonstrates moderate agreement with clinicians in diagnostic reasoning, but that evidence for real-world clinical utility, safety, and regulatory readiness remains limited.

Article Information

Published Online: August 25, 2026. https://doi.org/10.4088/PCC.26f04201
© 2026 Physicians Postgraduate Press, Inc.
Submitted: January 31, 2026; accepted April 20, 2026.
To Cite: Nadkarni A, Faust K, Matta SE, et al. Generative artificial intelligence in mental health: a guide for primary care providers. Prim Care Companion CNS Disord 2026;28(4):26f04201.
Author Affiliations: Department of Psychiatry, Harvard Medical School, Boston, Massachusetts (Nadkarni, Faust, Matta, Weiss, Schmelzer, Stern); Department of Psychiatry, Massachusetts General Brigham, Boston, Massachusetts (Nadkarni, Faust, Matta, Weiss, Schmelzer, Stern).
Ashwini Nadkarni, MD; Kyle Faust, PhD; Sofia E. Matta, MD; Jacob R. Weiss, MD; and Naomi A. Schmelzer, MD are co-first authors; Stern is the senior author.
Corresponding Authors: Ashwini Nadkarni, MD, Mass General Brigham/Harvard Medical School, Boston, Massachusetts ([email protected]).
Financial Disclosure: Dr Stern has received royalties from Elsevier for editing textbooks on psychiatry. The other authors have no conflicts of interest related to the subject of this article.
Funding/Support: None.

Clinical Points

  • Patients use generative artificial intelligence (AI) technology (eg, chatbots) for emotional support, assistance with coping, or mental health advice.
  • Common errors of generative AI chatbots include a lack of contextual understanding of problems.
  • One large language model, ChatGPT, has been known to offer increasingly inappropriate medical advice when complex scenarios were presented.
  • As patients increasingly use generative AI to assess and manage their symptoms, they may construe AI output as professional advice, despite a lack of knowledge about its quality, accuracy, and biases. Therefore, clinicians should counsel patients on the scope and limitations of digital tools given variability in the evidence supporting their use and clarify whether such tools are intended to be adjunctive rather than replace professional mental health care.
  1. Cruz-Gonzalez P, He AW, Lam EP, et al. Artificial intelligence in mental health care: a systematic review of diagnosis, monitoring, and intervention applications. Psychol Med. 2025;55:e18. CrossRef
  2. Guo Z, Lai A, Thygesen JH, et al. Large language models for mental health applications: systematic review. JMIR Ment Health. 2024;11(1):e57400. CrossRef
  3. Li H, Zhang R, Lee YC, et al. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. NPJ Digit Med. 2023;6(1):236. CrossRef
  4. McBain RK, Bozick R, Diliberti M, et al. Use of generative AI for mental health advice among US adolescents and young adults. JAMA Netw Open. 2025;8(11):e2542281. CrossRef
  5. Siddals S, Torous J, Coxon A. “It happened to be the perfect thing”: experiences of generative AI chatbots for mental health. NPJ Ment Health Res. 2024;3(1):48. CrossRef
  6. Howcroft A, Bennett-Weston A, Khan A, et al. AI chatbots versus human healthcare professionals: a systematic review and meta-analysis of empathy in patient care. Br Med Bull. 2025;156(1):ldaf017. CrossRef
  7. Liu T, Giorgi S, Aich A, et al. The illusion of empathy: how AI chatbots shape conversation perception. Proc AAAI Conf Artif Intell. 2025;39(13):14327–14335. CrossRef
  8. Heinz MV, Mackin DM, Trudeau BM, et al. Randomized trial of a generative AI chatbot for mental health treatment. N Engl J Med AI. 2025;2(4):AIoa2400802.
  9. Zhong W, Luo J, Zhang H. Therapeutic effectiveness of artificial intelligence–based chatbots in alleviation of depressive and anxiety symptoms in short-course treatments: a systematic review and meta-analysis. J Affect Disord. 2024;356:459–469. CrossRef
  10. Fitzpatrick KK, Darcy A, Vierhile M. Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Ment Health. 2017;4(2):e7785.
  11. Chen C, Lam KT, Yip KM, et al. Comparison of an AI chatbot with a nurse hotline in reducing anxiety and depression levels in the general population: pilot randomized controlled trial. JMIR Hum Factors. 2025;12:e65785. CrossRef
  12. Kneese T, Vecchione B, Marwick A. A chatbot for the soul: mental health care, privacy, and intimacy in AI-based conversational agents. Commun Change. 2025;1:15. CrossRef
  13. Zack T, Lehman E, Suzgun M, et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health. 2024;6(1):e12–e22. CrossRef
  14. Blease C, Rodman A. Generative artificial intelligence in mental healthcare: an ethical evaluation. Curr Treat Options Psychiatry. 2024;12(1):5. CrossRef
  15. Zafar F, Alam LF, Vivas RR, et al. The role of artificial intelligence in identifying depression and anxiety: a comprehensive literature review. Cureus. 2024;16(3):e56472.
  16. Hirosawa T, Harada Y, Mizuta K, et al. Evaluating ChatGPT-4’s accuracy in identifying final diagnoses within differential diagnoses compared with those of physicians: experimental study for diagnostic cases. JMIR Form Res. 2024;8:e59267. CrossRef
  17. Huo B, Boyle A, Marfo N, et al. Large language models for chatbot health advice studies: a systematic review. JAMA Netw Open. 2025;8(2):e2457879.
  18. Pichowicz W, Kotas M, Piotrowski P. Performance of mental health chatbot agents in detecting and managing suicidal ideation. Sci Rep. 2025;15(1):31652. CrossRef
  19. McBain RK, Cantor JH, Zhang LA, et al. Evaluation of alignment between large language models and expert clinicians in suicide risk assessment. Psychiatr Serv. 2025;76(11):944–950. CrossRef
  20. Balan R, Gumpel TP. ChatGPT clinical use in mental health care: scoping review of empirical evidence. JMIR Ment Health. 2025;12:e81204. CrossRef
  21. Kleine AK, Kokje E, Hummelsberger P, et al. AI-enabled clinical decision support tools for mental healthcare: a product review. Artif Intell Med. 2025 Feb 1;160:103052–22. CrossRef
  22. Wilimitis D, Turer RW, Ripperger M, et al. Integration of face-to-face screening with real-time machine learning to predict risk of suicide among adults. JAMA Netw Open. 2022;5(5):e2212095. CrossRef
  23. Velupillai S, Hadlaczky G, Baca-Garcia E, et al. Risk assessment tools and data-driven approaches for predicting and preventing suicidal behavior. Front Psychiatry. 2019;10:36. CrossRef
  24. Ehtemam H, Sadeghi Esfahlani S, Sanaei A, et al. Role of machine learning algorithms in suicide risk prediction: a systematic review and meta-analysis of clinical studies. BMC Med Inf Decis Mak. 2024;24(1):138.
  25. Bhandarkar AR, Arya N, Lin KK, et al. Building a natural language processing artificial intelligence to predict suicide-related events based on patient portal message data. Mayo Clin Proc Digit Health. 2023;1(4):510–518. PubMed
  26. Thakkar A, Gupta A, De Sousa A. Artificial intelligence in positive mental health: a narrative review. Front Digit Health. 2024;6:1280235. PubMed
  27. Blease C, Torous J. ChatGPT and mental healthcare: balancing benefits with risks of harms. BMJ Ment Health. 2023;26(1):e300884. PubMed
  28. Dergaa I, Fekih-Romdhane F, Hallit S, et al. ChatGPT is not ready yet for use in providing mental health assessment and interventions. Front Psychiatry. 2023;14:1277756. PubMed
  29. Ahn JJ, Yin W. Prompt-reverse inconsistency: LLM self-inconsistency beyond generative randomness and prompt paraphrasing; 2025. arXiv:2504.Arxiv. Preprint. 01282
  30. Carr S. “AI gone mental”: engagement and ethics in data-driven technology for mental health. J Ment Health. 2020;29(2):125–130. PubMed CrossRef
  31. Moore J, Grabb D, Agnew W, et al, eds.. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers Proceedings of the ACM Conference on Fairness, Accountability, and Transparency; 2025.
  32. Frances A. Warning: AI chatbots will soon dominate psychotherapy. Br J Psychiatry. 2025:1–5.
  33. Head KR. Minds in crisis: how the AI revolution is impacting mental health. J Ment Health Clin Psychol. 2025;9(3).
  34. Preda A. AI-induced psychosis: a new frontier in mental health. Psychiatr News. 2025;60(10).
  35. American Medical Association. ChatGPT and Generative AI: What Physicians Should Consider. American Medical Association; 2023.
  36. Lee EE, Torous J, De Choudhury M, et al. Artificial intelligence for mental health care: clinical applications, barriers, facilitators, and artificial wisdom. Biol Psychiatry Cogn Neurosci Neuroimaging. 2021;6(9):856–864. PubMed
  37. Tilala MH, Chenchala PK, Choppadandi A, et al. Ethical considerations in the use of artificial intelligence and machine learning in health care: a comprehensive review. Cureus. 2024;16(6).
  38. Li J, Li Y, Hu Y, et al. Chatbot-delivered interventions for improving mental health among young people: a systematic review and meta-analysis. Worldviews Evid Based Nurs. 2025;22(4):e70059. CrossRef
  39. Moshe I, Terhorst Y, Philippi P, et al. Digital interventions for the treatment of depression: a meta-analytic review. Psychol Bull. 2021;147(8):749–786. PubMed
  40. Raile P. The usefulness of ChatGPT for psychotherapists and patients. Humanit Soc Sci Commun. 2024;11:47.
  41. He Y, Yang L, Qian C, et al. Conversational agent interventions for mental health problems: systematic review and meta-analysis of randomized controlled trials. J Med Internet Res. 2023;25:e43862. PubMed
  42. Balcombe L. AI chatbots in digital mental health. 10. MDPI; 2023:82.Informatics4
  43. Kahane K, Shumate JN, Torous J. Policy in flux: addressing the regulatory challenges of AI integration in US mental health services. Curr Treat Options Psychiatry. 2025 Jun 16;12(1):24. CrossRef
  44. Palmer A, Schwan D. Digital mental health tools and AI therapy Chatbots: a balanced approach to regulation. Hastings Cent Rep. 2025 May;55(3):15–29. CrossRef
  45. Qiu J, He Y, Juan X, et al. Emoagent: Assessing and safeguarding human-ai interaction for mental health safety; 2025 Apr 13.arXiv Prepr arXiv: 2504.09689
  46. Brown JE, Halpern J. AI chatbots cannot replace human interactions in the pursuit of more inclusive mental healthcare. SSM-Mental Health. 2021 Dec 1;1:100017. CrossRef
  47. Khawaja Z, Bélisle-Pipon JC. Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots. Front Digital Health. 2023 Nov 8;5:1278186. CrossRef
  48. Fiske A, Henningsen P, Buyx A. Your robot therapist will see you now: ethical implications of embodied artificial intelligence in psychiatry, psychology, and psychotherapy. J Med Internet Res. 2019;21(5):e13216. CrossRef
Buy PDF for $40

Please sign in or purchase this PDF for $40.