科研速览 · Science Skim继续刷下去 · Keep skimming →
2026-07-31· Computer science

HORATIO: Bridging Management and Analysis of Traces at Scale

Ray A. O. Sinurat, William Nixon, Haryadi S. Gunawi, Nikoli Dryden, Hariharan Devarajan

原始摘要(英文原文)· Original abstract
Modern scientific and deep learning workloads on HPC systems rely on profiling and tracing across multiple software and hardware layers, generating diagnostic traces that often reach terabyte scale. Existing approaches manage these traces along two axes: trace formats and analysis tooling. Practitioners often convert raw traces into queryable formats, but doing so nearly doubles storage when raw files are retained for compatibility and requires upfront schema discovery that profiling tools cannot guarantee. A cleaner path is to make raw traces efficient in place, but this requires overcoming three limitations: lack of selective querying, analysis throughput bottlenecks, and lack of physical clustering. To address these limitations jointly, we developed Horatio, a raw trace management framework that indexes, analyzes, and physically clusters raw traces directly. Three findings emerge from our work. First, Horatio stores a lightweight RocksDB-backed auxiliary index alongside the raw trace, including gzip checkpoints, per-chunk bloom filters, and chunk-level statistics, delivering selective queries up to 75 × faster than naive Parquet at ∼ 1.01 × raw storage, with the highest cross-query mean throughput (232 M events/s) among state-of-the-art formats. Second, offloading event-level computation to a native C++ backend while keeping Dask for orchestration yields 80–83 × speedup over the original Dask-based DFAnalyzer. Third, lossless trace clustering that preserves the same input format yields a further 1.8–5.5 × end-to-end speedup across four AI and scientific workloads, and up to 230 × on h5bench where the preset aligns tightly with the cluster boundary, all with original layouts reconstructible on demand. Across five AI and scientific workloads, Horatio completes pipelines that DFAnalyzer cannot finish within 8 hours. On a 2.2 TB uncompressed trace, Horatio’s MPI-based mode scales to 16 × at 32 nodes on the full event set and its Dask-based path peaks at 4 × on a preset-filtered workload, both completing where DFAnalyzer hits OOM at every scale.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

HORATIO: Bridging Management and Analysis of Traces at Scale — 科研速览 Science Skim