Clinical Summary

Clinical Summary: Predictive Modeling of Suicidal Ideation in US Veterans: A Head-to-Head Comparison of the Patient Health Questionnaire-2 Versus Complex Machine Learning Approaches

Veterans face disproportionately high rates of suicidal ideation and behavior, yet screening tools must balance missing high-risk patients against overwhelming systems with false positives. This study asks a practical question for clinicians: do complex machine learning models meaningfully outperform a simple PHQ-2 sum score when predicting past-year suicidal ideation?

Design a longitudinal study of a nationally representative sample of US military veterans
N 3,078
Population US military veterans
Duration 1 year apart

Key Findings

  • In the “comprehensive variables” model set, AUC values ranged from 0.85 to 0.88, all of which were significantly higher than the PHQ-2 AUC of 0.79.
  • In the “clinically available variables” model set, AUCs ranged from 0.78 to 0.83, compared with an AUC of 0.80 for the PHQ-2 sum score; with the exception of the RF and GBM models, the machine learning models in this set demonstrated significantly higher AUCs than the PHQ-2.
  • Negative predictive values were high across both model sets, ranging from 0.97 to 0.98 for the “comprehensive variables” models and from 0.95 to 0.98 for the “clinically available variables” models.
  • Positive predictive values were markedly lower, ranging from 0.23 to 0.27 for the “comprehensive variables” model set and from 0.16 to 0.30 for the “clinically available variables” model set.
  • Greater model complexity came with more missing-data loss: the “comprehensive variables” model set included n=2,666 participants at wave 1 and n=2,659 at wave 2, whereas the “clinically available variables” model set included n=3,026 participants at wave 1 and n=3,052 at wave 2.
Clinical Bottom Line

For predicting suicidal ideation in veterans, complex machine learning models improved AUC only modestly over the PHQ-2 and still produced low PPVs. In routine practice, a simple PHQ-2-based approach is often the more feasible choice unless a care setting has the infrastructure and follow-up capacity to support more complex models.

Practice Implications

  • A PHQ-2 sum score remains a reasonable front-line tool for suicide risk screening in veterans, given AUCs of 0.79 to 0.80 and the limited incremental gain from more complex models.
  • If a program adopts multipredictor models, plan for many false positives: PPVs were only 0.23 to 0.27 in the “comprehensive variables” set and 0.16 to 0.30 in the “clinically available variables” set.
  • Use these models primarily to help rule out risk rather than confirm it, because NPVs were 0.95 to 0.98 across model sets.
  • Match the prediction approach to local workflow and data capacity, since broader predictor panels reduced analyzable sample size from 3,026/3,052 in the “clinically available variables” set to 2,666/2,659 in the “comprehensive variables” set.
Read full article
Physicians Postgraduate Press, Inc. (PPP) makes no warranties about the accuracy or completeness of any information published in The Journal of Clinical Psychiatry or other PPP materials, and disclaims liability for any use or non-use of that information. Clinicians should not rely solely on these materials and should exercise their own professional judgment when making patient care decisions on an individualized basis.