科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Information Processing & Management2026-01-13· Computer science

Performance unfairness of large language models in cross-language fact-checking

Dandan Wang, Stephanie Jean Tsang, Yadong Zhou

原始摘要(英文原文)· Original abstract
• Scalable computable evaluation system for LLM performance in fact-checking. • Inequality quantification across languages. • Checking-worthiness scoring and checking-authenticity verification. • Prompts in different types of language combination with different effects. • Role-restricted prompt engineering and model fine-tuning alleviate unfairness. Large language models (LLMs) are increasingly used for automated fact-checking, yet their performance often varies across languages, raising global fairness concerns. This study evaluated cross-language inequality in LLM-based fact-checking using 4,500 claims spanning nine languages across six language families. Besides building a systematic performance-evaluation pipeline covering instruction following, authenticity classification, evidence generation, and checking-worthiness scoring, we quantified inequality using standard deviation, coefficient of variation, Gini coefficient, and Theil index. Results showed substantial cross-language disparities, with higher performance on claims from rich-resource languages. To mitigate inequality, we tested two interventions, role-restricted prompt engineering and model fine-tuning. Both approaches reduced disparities, with fine-tuning achieving the largest and most consistent improvement across languages, particularly in checking-worthiness scoring. This study provides a reproducible framework for quantifying multilingual performance and fairness in LLM-based fact-checking and offers practical guidance for developing more equitable verification systems across diverse linguistic contexts.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance unfairness of large language models in cross-language fact-checking — 科研速览 Science Skim