María Teresa García-Ordás, Sergio Rubio-Martín, Alicia Merayo Corcoba, Arturo Crespo Álvaro, Óscar García-Olalla, Isaías García-Rodríguez
Depression affects over 280 million people worldwide, yet early detection remains challenging due to limited access to mental health services and the subjective nature of traditional assessment methods. Automated systems that can accurately and transparently assess depression from natural language could significantly improve early intervention and treatment outcomes. However, existing approaches using traditional natural language processing techniques (e.g., TF-IDF, Word2Vec) produce opaque feature representations that lack clinical interpretability, limiting their adoption in healthcare settings. This paper introduces a novel two-stage framework for depression prediction that leverages large language models (LLMs) to automatically generate interpretable, clinically relevant features from interview transcripts. The framework first employs few-shot prompting to generate a set of clinically grounded questions aligned with PHQ-8 diagnostic criteria, then uses the same LLMs to extract numerical features by answering these questions for each transcript. We evaluated three state-of-the-art LLMs (GPT-4o-mini, Claude-3.5-haiku, Gemini-Flash-1.5-8B) and five regression models on the E-DAIC dataset. Our results demonstrate that the proposed LLM-based feature extraction significantly outperforms traditional methods, with GPT-4o-mini achieving the best performance: RMSE of 4.19, MAE of 3.20, and of 0.50 when combined with Ridge regression, representing improvements of up to 30% over TF-IDF and Word2Vec baselines. Furthermore, SHAP (SHapley Additive exPlanations) analysis revealed that features related to guilt and suicidal ideation are the most predictive, aligning with established clinical knowledge and validating the clinical relevance of the extracted features. This work bridges the critical gap between predictive accuracy and interpretability in mental health assessment, providing a transparent, explainable system that clinicians can trust and understand, thereby laying the foundation for more reliable and clinically adoptable diagnostic tools.