Mathias Angermaier, Elisabeth Höldrich, João Pinheiro Neto, Jana Lasser
Sociality borne by language, as is the predominant digital trace on text-based social media platforms, harbours the raw material for exploring a multitude of social phenomena. Distinctively, the messaging service Telegram provides functionalities that allow for socially interactive as well as one-to-many communication. The Telegram dataset presented here contains over 5800 groups and channels discussing conspiracy-related topics with 63 million messages, originating from a data-hoarding initiative named the "Schwurbelarchiv" (from German schwurbeln: speaking nonsense). Uniquely, it includes the transcriptions of 2.5 million audio and video files. Our contribution is a processed, research-ready version of this data hoard: we parse, clean, and validate the raw archive, pseudonymise user data, and transcribe roughly 126,000 hours of audio and video content. In its original form the archive was stored in a format that is difficult to process and largely inaccessible for systematic research. This dataset publication details the structure, scope, and methodological specifics of the Schwurbelarchiv, emphasising its relevance for further research on the German-language conspiracy-related discourse. We validate its predominantly German origin by linguistic and temporal markers and situate it within the context of similar datasets. We describe process and extent of the transcription of multimedia files. Thanks to this effort the dataset uniquely supports analysis of text from originally multimodal sources like voice messages and videos to investigate online social dynamics and content dissemination. Researchers can employ this resource to explore societal dynamics for example related to conspiracy theories, misinformation, political extremism, and social network structures.