Ray A. O. Sinurat

dblp:327/4195 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-7255-7647ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoDL: A Framework for Studying Cross-Component Interference in Deep Learning Training Pipelines
abstract
Deep learning training on HPC systems exhibits significant step-time variability—data stalls can dominate 50–85% of training time [12, 31]—that limits efficiency. Existing benchmarks and profilers isolate individual components: I/O benchmarks replace compute with sleep, compute profilers ignore I/O, and GPU tracers operate on separate timescales. This isolation prevents understanding real variability, which stems from interactions among data loading, compute, and communication competing for shared CPU resources. We present CoDL, a framework that combines component-isolation benchmarking with multi-level tracing on a unified microsecond timeline, enabling diagnosis that reframes optimization from isolated component throughput to CPU–GPU scheduling coordination. On a single 4-GPU node, evaluation of ResNet-50 shows compute duration more than doubles (+127.4%) under active data loading and NCCL AllReduce degrades 2.8 ×, despite zero direct GPU-kernel overlap with data loading; the dominant mechanism is data-loading workers preempting GPU-launch threads on the CPU, inflating CUDA runtime API latency by 3.3 ×.
Druva Dhakshinamoorthy, Ray A. O. Sinurat, Nikoli Dryden, Arnab Kumar Paul, Hariharan Devarajan
HPDC2
2026 GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale
abstract
Scientific deep learning (DL) at scale typically trains on terabyte-scale datasets across thousands of accelerators, placing immense pressure on storage systems to keep pace with computation. Existing solutions respond to this demand by tuning individual I/O parameters to accelerate training performance. However, these techniques are limited by costly experiments, configuration space explosion, and inability to generalize application-specific optimizations. This leads to applications running with suboptimal configurations that reduce training efficiency, system utilization, or both. To address the challenge of finding the optimal configuration efficiently, we developed GLANCED-IO, a cross-layer I/O optimization framework that optimizes DL pipelines with high-fidelity approximation and efficient configuration space exploration. Through this work, we identified the following three key findings. First, independently optimizing either the application or system configurations leaves up to 2.4 × performance on the table for scientists to efficiently run DL pipelines on HPC systems. Second, GLANCED-IO’s one-factor-at-a-time (OFAT)-guided greedy exploration strategy achieved results comparable to more-expensive autotuning techniques while removing the pre-training required by ML-based approaches. Third, GLANCED-IO avoids executing the full application during optimization by operating on representative data subsets without GPUs, yet preserves 93% performance fidelity on average when deployed in DL pipelines. We demonstrate the efficacy of GLANCED-IO by optimizing large-scale global weather forecasting DL workloads, achieving up to 1.57 × better performance than state-of-the-art with 2.3 × fewer configuration evaluations than AIIO and 3.3 × faster optimization than DeepHyper.
Ray A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng, Sandeep Madireddy, Sam Foreman, Troy Arcomano, Robert B. Ross, Haryadi S. Gunawi, Hariharan Devarajan
HPDC1
2026 HORATIO: Bridging Management and Analysis of Traces at Scale
abstract
Modern scientific and deep learning workloads on HPC systems rely on profiling and tracing across multiple software and hardware layers, generating diagnostic traces that often reach terabyte scale. Existing approaches manage these traces along two axes: trace formats and analysis tooling. Practitioners often convert raw traces into queryable formats, but doing so nearly doubles storage when raw files are retained for compatibility and requires upfront schema discovery that profiling tools cannot guarantee. A cleaner path is to make raw traces efficient in place, but this requires overcoming three limitations: lack of selective querying, analysis throughput bottlenecks, and lack of physical clustering. To address these limitations jointly, we developed Horatio, a raw trace management framework that indexes, analyzes, and physically clusters raw traces directly. Three findings emerge from our work. First, Horatio stores a lightweight RocksDB-backed auxiliary index alongside the raw trace, including gzip checkpoints, per-chunk bloom filters, and chunk-level statistics, delivering selective queries up to 75 × faster than naive Parquet at ∼ 1.01 × raw storage, with the highest cross-query mean throughput (232 M events/s) among state-of-the-art formats. Second, offloading event-level computation to a native C++ backend while keeping Dask for orchestration yields 80–83 × speedup over the original Dask-based DFAnalyzer. Third, lossless trace clustering that preserves the same input format yields a further 1.8–5.5 × end-to-end speedup across four AI and scientific workloads, and up to 230 × on h5bench where the preset aligns tightly with the cluster boundary, all with original layouts reconstructible on demand. Across five AI and scientific workloads, Horatio completes pipelines that DFAnalyzer cannot finish within 8 hours. On a 2.2 TB uncompressed trace, Horatio’s MPI-based mode scales to 16 × at 32 nodes on the full event set and its Dask-based path peaks at 4 × on a preset-filtered workload, both completing where DFAnalyzer hits OOM at every scale.
Ray A. O. Sinurat, William Nixon, Haryadi S. Gunawi, Nikoli Dryden, Hariharan Devarajan
SSDBM1
2025 Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning Pipeline
abstract
This paper introduces Heimdall, a highly accurate and efficient machine learning-powered I/O admission policy for flash storage, designed to operate in a black-box manner. We make domain-specific innovations in various ML stages by introducing accurate period-based labeling, 3-stage noise filtering, in-depth feature engineering, and fine-grained tuning, which together improve the decision accuracy from 67% up to 93%. We perform various deployment optimizations to reach a sub-μs inference latency and a small, 28KB, memory overhead. With 500 unbiased random experiments derived from production traces, we show Heimdall delivers 15-35% lower average I/O latency compared to the state of the art and up to 2x faster to a baseline. Heimdall is ready for user-level, in-kernel, and distributed deployments.
Daniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli, Ray A. O. Sinurat, Janki Bhimani, Sandeep Madireddy, Achmad I. Kistijantoro, Haryadi S. Gunawi
EuroSys5
2025 AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
abstract
Generative machine learning offers new opportunities to better understand complex Earth system dynamics. Recent diffusion-based methods address spectral biases and improve ensemble calibration in weather forecasting compared to deterministic methods, yet have so far proven difficult to scale stably at high resolutions. We introduce AERIS, a 1.3 to 80B parameter pixel-level Swin diffusion transformer to address this gap, and SWiPe, a generalizable technique that composes window parallelism with sequence and pipeline parallelism to shard window-based transformers without added communication cost or increased global batch size. On Aurora (10,080 nodes), AERIS sustains 10.21 ExaFLOPS (mixed precision) and a peak performance of 11.21 ExaFLOPS with 1 × 1 patch size on the 0.25° ERA5 dataset, achieving 95.5% weak scaling efficiency, and 81.6% strong scaling efficiency. AERIS outperforms the IFS ENS and remains stable on seasonal scales to 90 days, highlighting the potential of billion-parameter diffusion models for weather and climate prediction.
Väinö Hatanpää, Eugene Ku, Jason Stock, Murali Emani, Sam Foreman, Chunyong Jung, Sandeep Madireddy, Varuni Sastry 0001, Ray A. O. Sinurat, Huihuo Zheng, Sam Wheeler, Troy Arcomano, Venkatram Vishwanath, Rao Kotamarthi
SC10
2022 Layered Contention Mitigation for Cloud Storage
abstract
We introduce an ecosystem of contention mitigation supports within the operating system, runtime and library layers. This ecosystem provides an end-to-end request abstraction that enables a uniform type of contention mitigation capabilities, namely request cancellation and delay prediction, that can be stackable together across multiple resource layers. Our evaluation shows that in our ecosystem, multi-resource storage applications are faster by 5-70% starting at 90P (the 90thpercentile) compared to popular practices such as speculative execution and is only 3% slower on average compared to a best-case (no contention) scenario.
Meng Wang 0056, Cesar A. Stuardo, Daniar Heri Kurniawan, Ray A. O. Sinurat, Haryadi S. Gunawi
CLOUD4