Key Takeaways
Extended Takeaways
- In the 32-variable models, discrimination improved to 0.85–0.88 versus 0.79 for the PHQ-2, but this came with low PPVs of 0.23 to 0.27, indicating that better AUC did not translate into efficient identification of true positives.
- The 9-predictor clinically available models yielded AUCs of 0.78 to 0.83 compared with 0.80 for the PHQ-2 sum score, suggesting that a modest expansion beyond depression screening items may offer little practical gain in routine care.
- Across both model sets, NPVs were consistently high at 0.95 to 0.98, so these approaches were better at ruling out past-year suicidal ideation than confirming it when a screen was positive.
- PHQ-2 items 1 and 2 remained among the most influential predictors even in multivariable models, alongside Mental Component Summary score and Purpose in Life in the comprehensive set, implying that depressive symptom burden captured much of the signal driving prediction.
- Missing-data burden increased as model complexity increased: the comprehensive set included n=2,666 at wave 1 and n=2,659 at wave 2, whereas the clinically available set retained n=3,026 and n=3,052, highlighting a practical implementation cost of using broader predictor panels.
- Because sensitivity and specificity were calculated from default decision thresholds, typically probability=0.50, the reported operating characteristics may not reflect the threshold needed for a screening program that prioritizes minimizing missed cases.