Caleb J Siefert, Barry Dauphin, Jenelle Slavin-Mulford, Sai Dileep Kumar Mukkamala, Venkat Akhila Reddy Tatipally, Areen Alsaid, Abdallah Chehade
Fine-tuning effectively adapted AI models into raters capable of reliably scoring complex psychological constructs from narrative. Ensemble-based AI raters can automate AFF and EIR ratings in research settings and, with further validation, may support clinical use via a human-in-the-loop approach. Fine-tuning adaptable AI models offers considerable potential to increase the feasibility and accessibility of sophisticated, multi-method assessment in research and clinical settings alike.
BACKGROUND: Narrative assessment offers unique insight into psychological functioning but is resource-intensive, requiring extensive expert training and time. Recent work suggests Artificial Intelligence (AI), including moderately sized, adaptable large language models (LLMs), can master narrative assessment's complex, multi-step scoring rules. This study examined whether fine-tuning could enable AI raters to accurately assess narratives.
METHOD: We fine-tuned five diverse LLMs (3-7 billion parameters) on two datasets, one for the SCORS-G Affective Quality of Representations (AFF) scale and one for Emotional Investment in Relationships (EIR). All narratives were rated by expert human raters. Each dataset was split into training and testing samples to evaluate AI-human agreement.
RESULTS: Fine-tuned models showed good to excellent reliability with human experts for both AFF and EIR. Ensemble ratings, averaged across all five models, yielded excellent single-rater and average-rater ICCs and low error rates, with 97% of AFF and 92% of EIR ratings falling within accepted discrepancy limits.
CONCLUSION: Fine-tuning effectively adapted AI models into raters capable of reliably scoring complex psychological constructs from narrative. Ensemble-based AI raters can automate AFF and EIR ratings in research settings and, with further validation, may support clinical use via a human-in-the-loop approach. Fine-tuning adaptable AI models offers considerable potential to increase the feasibility and accessibility of sophisticated, multi-method assessment in research and clinical settings alike.