Jan Schwalbach, Lukas Hetzer, Sven-Oliver Proksch, Christian Rauh, Miklós Sebők
Parliamentary texts are central sources for studying democratic representation, policy-making, and institutional decision-making, and they also constitute a highly valuable domain for natural language processing research. Despite their importance, access to comprehensive, clean, and machine-readable parliamentary corpora remains limited. This paper introduces the ParlLawSpeech (PLS) collection, a large-scale, curated corpus comprising more than three million documents from eight European parliaments, each covering over two decades. PLS extends existing parliamentary text resources in two key ways. First, it integrates full texts and metadata not only for parliamentary speeches but also for legislative bills and finally adopted laws. Second, it provides systematic data linkage across these document types, enabling researchers to trace the full parliamentary decision-making process from draft legislation through debate to adoption. Covering seven national parliaments and the European Parliament, PLS offers a uniquely rich resource for political science and advanced NLP research.