Rebecka Maria Norman, Lilja Charlotte Storset, Petter Mæhlum, Hilde Hestad Iversen, Elma Jelin, Erik Velldal, Lilja Øvrelid, Øyvind Bjertnæs
Beyond good classification performance, the findings support automated sentiment analysis as a valid patient experience measure in general practice. The sentiment classifications may complement traditional patient experience measures and help identify areas for quality improvement.
BACKGROUND: Automated sentiment analysis can be used to analyse free-text patient experience comments, but in health services research, it requires technical performance and methodological validity. Traditional models lack contextual understanding, whereas masked language models (MLMs) enable nuanced sentiment classification. This study is the first to fine-tune an MLM (NorBERT3large) on human-annotated data and evaluate the validity of four-category sentiment classifications across survey and online feedback in general practice.
OBJECTIVE(S): To evaluate a fine-tuned MLM against human-annotated sentiment labels, compare its performance with traditional methods and a large language model, and assess convergent and known-groups validity using survey and online comments.
METHODS: NorBERT3large was fine-tuned on human-annotated free-text survey comments. Performance was evaluated using F1-scores. The model classified unseen online data. Construct validity was assessed per COSMIN guidelines through correlations with patient experience scores and subgroup analyses.
RESULTS: The fine-tuned NorBERT3large achieved high F1-scores on held-out survey test data and outperformed traditional models used in prior patient experience research, as well as a zero-shot large language model reference. On unseen online reviews, performance was high in the random sample and strongest for mixed and positive sentiment in the balanced evaluation set, while the rare neutral comments were difficult to classify reliably. Construct validity was supported by expected correlations with patient experience scores and known subgroup differences.
CONCLUSION: Beyond good classification performance, the findings support automated sentiment analysis as a valid patient experience measure in general practice. The sentiment classifications may complement traditional patient experience measures and help identify areas for quality improvement.