Pengfei Bao
The advent of open-source Large Language Models (LLMs) like LLaMA and BLOOM promises a democratization of artificial intelligence. However, this narrative of openness often obscures the inherent linguistic governance embedded within their architectures. This study posits that the technical construction of open-source LLMs constitutes a potent form of implicit, meso-level language policy. Through a mixed-methods analysis combining corpus linguistics of training data disclosures and computational text analysis of model outputs, we interrogate the language power structures and systemic biases these models perpetuate. Our empirical validation, involving structured prompts in English, Chinese, Spanish, and Swahili, reveals a consistent pattern of linguistic hegemony: models excel in tasks for dominant languages while producing impoverished, inaccurate, or culturally void content for lower-resource languages. The findings demonstrate that, far from being neutral platforms, open-source LLMs actively enact a form of algorithmic language planning, systematically privileging certain languages and worldviews while marginalizing others. This process silently shapes the digital future of languages, raising critical questions about equity and the need for explicit policy interventions in AI development.