Clinical Summary
Clinical Summary: Predictive Modeling of Suicidal Ideation in US Veterans: A Head-to-Head Comparison of the Patient Health Questionnaire-2 Versus Complex Machine Learning Approaches
Veterans face disproportionately high rates of suicidal ideation and behavior, yet screening tools must balance missing high-risk patients against overwhelming systems with false positives. This study asks a practical question for clinicians: do complex machine learning models meaningfully outperform a simple PHQ-2 sum score when predicting past-year suicidal ideation?
Design
a longitudinal study of a nationally representative sample of US military veterans
N
3,078
Population
US military veterans
Duration
1 year apart
Key Findings
- In the “comprehensive variables” model set, AUC values ranged from 0.85 to 0.88, all of which were significantly higher than the PHQ-2 AUC of 0.79.
- In the “clinically available variables” model set, AUCs ranged from 0.78 to 0.83, compared with an AUC of 0.80 for the PHQ-2 sum score; with the exception of the RF and GBM models, the machine learning models in this set demonstrated significantly higher AUCs than the PHQ-2.
- Negative predictive values were high across both model sets, ranging from 0.97 to 0.98 for the “comprehensive variables” models and from 0.95 to 0.98 for the “clinically available variables” models.
- Positive predictive values were markedly lower, ranging from 0.23 to 0.27 for the “comprehensive variables” model set and from 0.16 to 0.30 for the “clinically available variables” model set.
- Greater model complexity came with more missing-data loss: the “comprehensive variables” model set included n=2,666 participants at wave 1 and n=2,659 at wave 2, whereas the “clinically available variables” model set included n=3,026 participants at wave 1 and n=3,052 at wave 2.
Clinical Bottom Line
For predicting suicidal ideation in veterans, complex machine learning models improved AUC only modestly over the PHQ-2 and still produced low PPVs. In routine practice, a simple PHQ-2-based approach is often the more feasible choice unless a care setting has the infrastructure and follow-up capacity to support more complex models.
Practice Implications
- A PHQ-2 sum score remains a reasonable front-line tool for suicide risk screening in veterans, given AUCs of 0.79 to 0.80 and the limited incremental gain from more complex models.
- If a program adopts multipredictor models, plan for many false positives: PPVs were only 0.23 to 0.27 in the “comprehensive variables” set and 0.16 to 0.30 in the “clinically available variables” set.
- Use these models primarily to help rule out risk rather than confirm it, because NPVs were 0.95 to 0.98 across model sets.
- Match the prediction approach to local workflow and data capacity, since broader predictor panels reduced analyzable sample size from 3,026/3,052 in the “clinically available variables” set to 2,666/2,659 in the “comprehensive variables” set.