Serik Aliaskarov
This study presents a reproducible framework for evaluating the performance of Big Data workflows inspired by Hadoop–Spark processing logic using two open datasets with distinct workload characteristics. The first dataset, the NYC TLC Yellow Taxi Trip Record Data for January 2023, represents an urban mobility scenario and contains temporal and spatial trip attributes. The second dataset, the GDELT Event Daily Export for 1 January 2023, represents global event data with temporal, geographic, categorical, and quantitative characteristics. Three reproducible processing configurations are compared: a disk-based batch baseline, a Spark-only workflow, and a Prepared-storage + Spark workflow, which serves as a storage-oriented approximation of a hybrid Hadoop–Spark processing pipeline. The experimental design includes batch aggregation, spatial–temporal analysis, stream-like processing, scalability testing, and task restart-based fault recovery. Performance is evaluated using execution time, throughput, latency, scalability behavior, CPU utilization, RAM delta, I/O-related behavior, and recovery time metrics. To ensure reproducibility, a unified preprocessing procedure was applied across all configurations. The TLC dataset was reduced from 3,066,766 to 2,883,546 valid records, while the GDELT dataset was reduced from 43,575 to 42,053 valid records. The study does not claim universal superiority of the storage-oriented workflow. Instead, its primary contribution lies in providing a transparent and reproducible methodology for comparing storage-oriented and Spark-based data processing workflows using heterogeneous open datasets under identical preprocessing rules and comparable workload execution conditions.