HOW-TO GUIDES 1 guide
Frequently Asked Questions
10 questions-
Yes, but the improvement was modest. In the 32-variable comprehensive model set, machine learning model AUCs ranged from 0.85 to 0.88, compared with 0.79 for the PHQ-2 sum score. In the 9-predictor clinically available set, AUCs ranged from 0.78 to 0.83, compared with 0.80 for the PHQ-2, and only some of those models significantly outperformed the PHQ-2.
The authors concluded that these statistically significant gains were clinically marginal and may not justify the added complexity of multipredictor models.
-
The PHQ-2 alone performed competitively. Its AUC was 0.79 when compared with the comprehensive model set and 0.80 when compared with the clinically available model set, which was close to the performance of several more complex models.
The study used the PHQ-2 as a reference because it is simple, widely used in clinical settings, and does not directly ask about suicidality.
-
These models were better at ruling out suicidal ideation than confirming it. Negative predictive values were consistently high, ranging from 0.97 to 0.98 in the comprehensive models and from 0.95 to 0.98 in the clinically available models, while positive predictive values were much lower at 0.23 to 0.27 and 0.16 to 0.30, respectively.
That means a negative result was usually reassuring, but a positive result often did not correspond to true past-year suicidal ideation.
-
In the comprehensive models, the most influential predictors were consistently the Mental Component Summary score, Purpose in Life, and PHQ-2 items 1 and 2 across the GLM, random forest, gradient boosting machine, and neural network models.
The authors noted that the persistent importance of PHQ-2 items suggests depressive symptom burden captured much of the core variance related to suicidal ideation, even in more complex algorithms.
-
The simple reference approach used the PHQ-2 sum score with different severity cutoffs to generate an ROC curve. The comprehensive model set used 32 variables, including sociodemographic factors, symptoms of mental disorders, history of suicide attempts, questionnaires, and personality traits. The clinically available model set used 9 predictors that are typically available in clinical practice, including basic sociodemographic data and items from the PHQ-2 and GAD-2.
-
Suicidal ideation was defined as any past-year suicidal thoughts reported on item 2 of the Suicide Behaviors Questionnaire-Revised. Participants were classified as positive if they reported thinking about killing themselves at least once in the past year, meaning any response greater than 0 on that item.
-
This was a longitudinal analysis of a nationally representative sample of US veterans assessed at 2 survey waves approximately 1 year apart. Models were trained on wave 1 data and tested on wave 2 data, and the final analytic samples differed by model set because missing data were handled with listwise deletion.
This design provides independent testing across time, but the authors noted that using the same individuals at both time points and relying on retrospective self-report may have inflated performance relative to what would be expected in a fully prospective study or an external validation cohort.
-
A total of 3,078 veterans completed both survey waves, but the analytic sample size varied by model set because of missing data. The comprehensive 32-variable models included 2,666 participants at wave 1 and 2,659 at wave 2, while the clinically available 9-predictor models included 3,026 participants at wave 1 and 3,052 at wave 2.
-
The study suggests that more complex machine learning models did not provide clinically meaningful improvement over a simple PHQ-2-based approach for predicting suicidal ideation in veterans. Although some complex models had higher AUCs, they still produced low positive predictive values and required more data, infrastructure, and implementation effort.
The authors emphasized that the most appropriate model depends on the clinical setting, the intended use of the tool, and whether the system has the capacity to follow up on the large number of patients who may screen positive.
-
The main limitations were that wave 2 data were collected during the early COVID-19 period, when suicidal ideation prevalence declined; suicidal ideation does not fully capture more severe suicidal behavior; and the study relied on retrospective self-report, which may introduce recall bias and underreporting.
Additional limitations were that the PHQ-2 was used as the reference model even though other PHQ-9 items may have stronger predictive value, the sample was specific to veterans, and the repeated-measures design using the same individuals across waves may limit generalizability and may have inflated model performance.