Ainhoa Serna, Jon Kepa Gerrikagoitia, Juan de Oña
This study evaluates zero-shot domain transfer for multilingual sentiment analysis in sustainable urban mobility using XLM-RoBERTa, a transformer pre-trained on social media data and applied to transport reviews without task- or domain-specific fine-tuning. Starting from a manually annotated English corpus of 375 transport-related user reviews, we created sentence-aligned translations in Spanish, French, German, and Italian, yielding a multilingual evaluation dataset of 1875 instances. Results show that the model assigns consistently high confidence to polarized content (mean: 0.76–0.85) and lower confidence to neutral or ambiguous expressions (0.58–0.65), with visible but preliminary cross-lingual variations that require further linguistic validation. Confidence scores are treated as diagnostic indicators of model certainty, not as evidence of correctness or calibration. A qualitative analysis of 113 categorized low-confidence predictions identifies six recurring linguistic patterns associated with model uncertainty (led by translation drift, mixed sentiment, and idiomatic expressions) with substantial inter-annotator agreement (κ = 0.664). By releasing the annotated multilingual dataset and code publicly, this work provides a reproducible exploratory evaluation framework for annotation-scarce, domain-specific multilingual NLP.