科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Intelligent Decision Making and Information Science2026-07-31· Computer science

Reproducible Performance Evaluation of a Hybrid Hadoop–Spark Big Data Architecture Using Urban Mobility and Global Event Datasets

Serik Aliaskarov

原始摘要(英文原文)· Original abstract
This study presents a reproducible framework for evaluating the performance of Big Data workflows inspired by Hadoop–Spark processing logic using two open datasets with distinct workload characteristics. The first dataset, the NYC TLC Yellow Taxi Trip Record Data for January 2023, represents an urban mobility scenario and contains temporal and spatial trip attributes. The second dataset, the GDELT Event Daily Export for 1 January 2023, represents global event data with temporal, geographic, categorical, and quantitative characteristics. Three reproducible processing configurations are compared: a disk-based batch baseline, a Spark-only workflow, and a Prepared-storage + Spark workflow, which serves as a storage-oriented approximation of a hybrid Hadoop–Spark processing pipeline. The experimental design includes batch aggregation, spatial–temporal analysis, stream-like processing, scalability testing, and task restart-based fault recovery. Performance is evaluated using execution time, throughput, latency, scalability behavior, CPU utilization, RAM delta, I/O-related behavior, and recovery time metrics. To ensure reproducibility, a unified preprocessing procedure was applied across all configurations. The TLC dataset was reduced from 3,066,766 to 2,883,546 valid records, while the GDELT dataset was reduced from 43,575 to 42,053 valid records. The study does not claim universal superiority of the storage-oriented workflow. Instead, its primary contribution lies in providing a transparent and reproducible methodology for comparing storage-oriented and Spark-based data processing workflows using heterogeneous open datasets under identical preprocessing rules and comparable workload execution conditions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reproducible Performance Evaluation of a Hybrid Hadoop–Spark Big Data Architecture Using Urban Mobility and Global Event Datasets — 科研速览 Science Skim