科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Array2026-04-01· Computer science

Scalable conversion of Stack Exchange network-based question answering sites’ data dump to SQL Server database – A case of Stack Overflow data dump

Arjumand Fatima, Onaiza Maqbool

原始摘要(英文原文)· Original abstract
The Stack Exchange Network currently hosts 182 question answering sites (including Stack Overflow) which are broadly categorized into 77 technology and 105 other sites. Majority of these sites have greater than 10 years of age containing thousands to millions of users, questions and answers. The data of each site (since its inception till date) is available under the CC-by-SA license and can be accessed either as quarterly updated data dump (complete data) or through the weekly updated Stack Exchange Data Explorer (SEDE) (queryable interface) or the Stack Exchange API. Both SEDE and API suffer from usage limitations while the data dump can be difficult to preprocess based on its size which keeps on increasing, becoming a bottleneck to explore/utilize the wisdom of crowds hidden in these sites. Although the underlying schema for data dumps and SEDE is publicly available, currently no step by step, detailed procedure is available to scalably preprocess the continuously growing data dumps of any Stack Exchange site on commodity hardware. Existing efforts or discussion on methods of processing data dumps have been confined to Stack Overflow data dump only whereas previous research has primarily revolved around a few popular Stack Exchange sites with most studies based on small subsets of data chosen as sample preferably using SEDE or API. The procedure and codebase presented here were originally developed as part of a large-scale data curation project based on Stack Overflow data dump. However, the procedure and codebase should be useful more broadly for preprocessing the data dumps of any of the 182 existing Stack Exchange sites as they follow the same database schema, paving the ways for exploring the data of many such sites which have never been previously considered for any sort of analysis and enabling large-scale longitudinal studies which are otherwise not possible using SEDE or API. To be beneficial for a larger audience, we have specifically designed the solution for Windows based systems.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Scalable conversion of Stack Exchange network-based question answering sites’ data dump to SQL Server database – A case of Stack Overflow data dump — 科研速览 Science Skim