科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Parallel Programming2026-06-29· Computer science

Mneme: A Parallel Preprocessing Framework for Large Tabular Datasets

Argiris Sofotasios, Dimitris Metaxakis, Panagiotis Hadjidoukas

原始摘要(英文原文)· Original abstract
Abstract The rapid expansion of Machine Learning (ML) applications, especially within its subfield of Deep Learning (DL), has created an increasing demand for efficient preprocessing of large tabular datasets that surpass the available memory capacity of single-node systems. This paper introduces a parallel framework, developed as a Python library, designed to efficiently preprocess large-scale tabular datasets for training Deep Neural Networks (DNNs). The library supports various data transformations, including normalization, categorical encoding, and missing value imputation, leveraging parallel computing and chunk-based processing to efficiently handle massive datasets. By distributing preprocessing tasks across multiple cores and facilitating the parallel loading and processing of data chunks without altering the original data file, the proposed library significantly reduces the time required for data preparation, which often represents a critical bottleneck in modern ML pipelines. Experimental evaluation demonstrates substantial performance gains over conventional sequential approaches and state-of-the-art (SOTA) solutions.Furthermore, the library integrates seamlessly with widely adopted DL frameworks, providing a scalable and flexible High-Performance Computing (HPC) tool for data preprocessing in contemporary ML workflows.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Mneme: A Parallel Preprocessing Framework for Large Tabular Datasets — 科研速览 Science Skim