VLDB 2026 Research / reviewers in the wild / expert
Hariharan Devarajan
dblp:220/7505
· DBLP profile ↗
33ranked-venue papers
15as first author
20since 2021 · last 2026
0000-0001-5625-3494ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 12 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoDL: A Framework for Studying Cross-Component Interference in Deep Learning Training PipelinesabstractDeep learning training on HPC systems exhibits significant step-time variability—data stalls can dominate 50–85% of training time [12, 31]—that limits efficiency. Existing benchmarks and profilers isolate individual components: I/O benchmarks replace compute with sleep, compute profilers ignore I/O, and GPU tracers operate on separate timescales. This isolation prevents understanding real variability, which stems from interactions among data loading, compute, and communication competing for shared CPU resources. We present CoDL, a framework that combines component-isolation benchmarking with multi-level tracing on a unified microsecond timeline, enabling diagnosis that reframes optimization from isolated component throughput to CPU–GPU scheduling coordination. On a single 4-GPU node, evaluation of ResNet-50 shows compute duration more than doubles (+127.4%) under active data loading and NCCL AllReduce degrades 2.8 ×, despite zero direct GPU-kernel overlap with data loading; the dominant mechanism is data-loading workers preempting GPU-launch threads on the CPU, inflating CUDA runtime API latency by 3.3 ×. Druva Dhakshinamoorthy, Ray A. O. Sinurat, Nikoli Dryden, Arnab Kumar Paul, Hariharan Devarajan |
HPDC | 5 |
| 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at ScaleabstractScientific deep learning (DL) at scale typically trains on terabyte-scale datasets across thousands of accelerators, placing immense pressure on storage systems to keep pace with computation. Existing solutions respond to this demand by tuning individual I/O parameters to accelerate training performance. However, these techniques are limited by costly experiments, configuration space explosion, and inability to generalize application-specific optimizations. This leads to applications running with suboptimal configurations that reduce training efficiency, system utilization, or both. To address the challenge of finding the optimal configuration efficiently, we developed GLANCED-IO, a cross-layer I/O optimization framework that optimizes DL pipelines with high-fidelity approximation and efficient configuration space exploration. Through this work, we identified the following three key findings. First, independently optimizing either the application or system configurations leaves up to 2.4 × performance on the table for scientists to efficiently run DL pipelines on HPC systems. Second, GLANCED-IO’s one-factor-at-a-time (OFAT)-guided greedy exploration strategy achieved results comparable to more-expensive autotuning techniques while removing the pre-training required by ML-based approaches. Third, GLANCED-IO avoids executing the full application during optimization by operating on representative data subsets without GPUs, yet preserves 93% performance fidelity on average when deployed in DL pipelines. We demonstrate the efficacy of GLANCED-IO by optimizing large-scale global weather forecasting DL workloads, achieving up to 1.57 × better performance than state-of-the-art with 2.3 × fewer configuration evaluations than AIIO and 3.3 × faster optimization than DeepHyper. Ray A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng, Sandeep Madireddy, Sam Foreman, Troy Arcomano, Robert B. Ross, Haryadi S. Gunawi, Hariharan Devarajan |
HPDC | 10 |
| 2026 | Metis: Agentic Knowledge Synthesis for Explainable I/O Performance in HPC SystemsabstractI/O performance explainability in HPC requires contextual characterization across the full software and system stack. Contextual characterization identifies the semantics and runtime role of I/O functions. State-of-the-art contextual characterization is still largely manual, but remains highly valuable for explaining bottlenecks and guiding optimization. However, manual contextual characterization is difficult to scale and hard to reproduce as I/O libraries and cross-layer interactions grow in complexity. We present Metis, a framework for systematic characterization of HPC I/O functions that uses agentic LLMs to integrate heterogeneous evidence sources and quantify agents agreement. Across evaluation, Metis improves held-out-category generalization over an MCP Tool baseline (0.90 vs. 0.35), reduces runtime (27.7 s vs. 84.5 s) while increasing throughput (41.8 vs. 14.28 functions/min), and sustains high verifier throughput under federated scaling (330K–1.18M functions/s). These results demonstrate that Metis is an effective and practical approach for explainable characterization of complex HDF5 behavior, enabling more trustworthy and reproducible HPC I/O analysis. Karim Youssef, Sarah Neuwirth, Neeraj Rajesh, Hariharan Devarajan |
HPDC | 4 |
| 2026 | HORATIO: Bridging Management and Analysis of Traces at ScaleabstractModern scientific and deep learning workloads on HPC systems rely on profiling and tracing across multiple software and hardware layers, generating diagnostic traces that often reach terabyte scale. Existing approaches manage these traces along two axes: trace formats and analysis tooling. Practitioners often convert raw traces into queryable formats, but doing so nearly doubles storage when raw files are retained for compatibility and requires upfront schema discovery that profiling tools cannot guarantee. A cleaner path is to make raw traces efficient in place, but this requires overcoming three limitations: lack of selective querying, analysis throughput bottlenecks, and lack of physical clustering. To address these limitations jointly, we developed Horatio, a raw trace management framework that indexes, analyzes, and physically clusters raw traces directly. Three findings emerge from our work. First, Horatio stores a lightweight RocksDB-backed auxiliary index alongside the raw trace, including gzip checkpoints, per-chunk bloom filters, and chunk-level statistics, delivering selective queries up to 75 × faster than naive Parquet at ∼ 1.01 × raw storage, with the highest cross-query mean throughput (232 M events/s) among state-of-the-art formats. Second, offloading event-level computation to a native C++ backend while keeping Dask for orchestration yields 80–83 × speedup over the original Dask-based DFAnalyzer. Third, lossless trace clustering that preserves the same input format yields a further 1.8–5.5 × end-to-end speedup across four AI and scientific workloads, and up to 230 × on h5bench where the preset aligns tightly with the cluster boundary, all with original layouts reconstructible on demand. Across five AI and scientific workloads, Horatio completes pipelines that DFAnalyzer cannot finish within 8 hours. On a 2.2 TB uncompressed trace, Horatio’s MPI-based mode scales to 16 × at 32 nodes on the full event set and its Dask-based path peaks at 4 × on a preset-filtered workload, both completing where DFAnalyzer hits OOM at every scale. Ray A. O. Sinurat, William Nixon, Haryadi S. Gunawi, Nikoli Dryden, Hariharan Devarajan |
SSDBM | 5 |
| 2026 | WADO: A Distributed WORM Storage Service for Asynchronous Data OperationsabstractAI-driven scientific workloads increasingly depend on data-intensive input pipelines, where deep learning frameworks must ingest and transform large datasets from hierarchical HPC storage. Existing system-centric data services improve movement and locality between the parallel file system (PFS), node-local storage, and memory. However, they do not directly optimize how input pipeline operations execute across scopes, stage overlap, and resource-specific parallelism. As scale grows, this gap causes worker stalls, contention, and poor hardware utilization. We present WADO, a distributed write-once-read-many (WORM) object-store runtime for data-centric workloads that closes this gap through three coordinated mechanisms: scope-centric processing, explicit pipeline decomposition, and interference-aware explicit parallelism. WADO dynamically maps operations to execution scopes, overlaps stages such as I/O, communication, and transformations, and applies contention-aware concurrency control to match hardware behavior at runtime. Our evaluation shows three main findings: (1) scope-centric processing preserves throughput under scale, improving mixed-operation throughput by up to 1.65 × ; (2) explicit pipeline decomposition converts serialized wait into overlapped progress, delivering up to 2.16 × higher sustained bandwidth; and (3) interference-aware explicit parallelism improves effective bandwidth by up to 4.4 × by avoiding oversubscription collapse. On Unet3D model training, these mechanisms translate to end-to-end gains, improving data loading performance by 4.1 × compared to baseline PyTorch on Lustre, and 1.51 × compared to DYAD, enabled by deeper pipelining, adaptive parallelism, and near-data transformation offloading. Karim Youssef, Hariharan Devarajan, Nikoli Dryden, Roger A. Pearce |
SSDBM | 2 |
| 2026 | IFlux: Intent-Aware Storage Tiers & Software Scheduling for HPC Systems
Hariharan Devarajan, Vanessa Sochat, Daniel Milroy, Thomas Scogland |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El CapitanabstractWe present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures. Ian Lumsden, Stefano Markidis, Andong Hu, Ivy Bo Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Thomas Scogland, Hariharan Devarajan, Bronis R. de Supinski, Gian Luca Delzanno, Michela Taufer |
eScience | 10 |
| 2025 | WisIO: Automated I/O Bottleneck Detection with Multi-Perspective Views for HPC WorkflowsabstractWhy I/O Bottlenecks Matter in HPC• Modern HPC workloads (AI, simulations) involve massive data transfers that are crucial for enabling scientific discoveries• The large volume of these data transfers often lead to workloads spending significant amount of time performing I/O• Recent studies show that it is between 25-40% of total runtime • As a result, tuning the performance of data transfers via I/O analysis has become a routine task for application developers 6/27/2025 High-Level Execution Flow of WisIO 6/27/2025 WisIO 7 • Transform raw trace data into multi-perspective views • File, process, timeline, or user-defined High-Level Execution Flow of WisIO 6/27/2025 WisIO 8 • Transform raw trace data into multi-perspective views • File, process, timeline, or user-defined High-Level Execution Flow of WisIO 6/27/2025 WisIO 9 • Severity-based classification using I/O metrics • Quantifies how "bad" an I/O behavior is via a relative severity angle High-Level Execution Flow of WisIO 6/27/2025 WisIO 10 • Severity-based classification using I/O metrics • Quantifies how "bad" an I/O behavior is via a relative severity angle High-Level Execution Flow of WisIO 6/27/2025 WisIO 11 • Explains bottlenecks via rule-based reasoning • Identifies one or more causes (e.g., small reads, metadata overhead) High-Level Execution Flow of WisIO 6/27/2025 WisIO 12 • Explains bottlenecks via rule-based reasoning • Identifies one or more causes (e.g., small reads, metadata overhead) Implementation & API 6/27/2025 WisIO 13 Implemented in Python, for versions 3.8 and above • Parallel and distributed via Dask Works out-of-the-box with trace data from common I/O monitoring tools • Darshan, DFTracer, Recorder Two user-facing interfaces: • CLI: Installable via pip, highly configurable • Python API: Allows interactive analysis Multiple output types: Izzet Yildirim, Hariharan Devarajan, Antonios Kougkas, Xian-He Sun, Kathryn Mohror |
ICS | 2 |
| 2025 | XIO: Toward eXplainable I/O for HPC Systems
Sarah Neuwirth, Hariharan Devarajan, Chen Wang 0004, Jay F. Lofstead |
SSDBM | 2 |
| 2025 | H5Intent: Autotuning HDF5 With User IntentabstractThe complexity of data management in HPC systems stems from the diversity in I/O behavior exhibited by new workloads, multistage workflows, and multitiered storage systems. The HDF5 library is a popular interface to interact with storage systems in HPC workloads. The library manages the complexity of diverse I/O behaviors by providing user-level configurations to optimize the I/O for HPC workloads. The HDF5 library exposes hundreds of configuration properties that can be set to alter how HDF5 manages I/O requests for better performance. However, determining which properties to set is quite challenging for users who lack expertise in HDF5 library internals. We propose a paradigm change through our H5Intent software, where users specify the intent of I/O operations and the software can set various HDF5 properties automatically to optimize the I/O behavior. This work demonstrates several use cases where mapping user-defined intents to HDF5 properties can be exploited to optimize I/O. In this study, we make three observations. First, I/O intents can accurately define HDF5 properties while managing conflicts between various properties and improving the I/O performance of microbenchmarks by up to 22×. Second, I/O intents can be efficiently passed to HDF5 with a small footprint of 6.74MB per node for thousands of intents per process. Third, an H5Intent VOL connector can dynamically map I/O intents to HDF5 properties for various I/O behaviors exhibited by our microbenchmark and improve I/O performance by up to 8.8×. Overall, H5Intent software improves the I/O performance of complex large-scale workloads we studied by up to 11×. Hariharan Devarajan, Gerd Heber, Kathryn Mohror |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | ML-based Modeling to Predict I/O Performance on Different Storage Sub-systemsabstractParallel applications can spend a significant amount of time performing I/O on large-scale supercomputers. Fast near-compute storage accelerators called burst buffers can reduce the time a processor spends performing I/O and mitigate I/O bottlenecks. However, determining if a given application could be accelerated using burst buffers is not straightforward even for storage experts. The relationship between an application's I/O characteristics (such as I/O volume, processes involved, etc.) and the best storage sub-system for it can be complicated. As a result, adapting parallel applications to use burst buffers efficiently is a trial-and-error process. In this work, we present a Python-based tool called PrismIO that enables programmatic analysis of I/O traces. Using PrismIO, we identify performance bottlenecks when using burst buffers and parallel file systems, and explain why certain I/O patterns perform poorly. Further, we use machine learning to model the relationship between I/O characteristics and file system selections. We use IOR, an I/O benchmark to gather performance data for training the machine learning model. Our model can predict the better performing storage system for unseen IOR scenarios with an accuracy of 94.47% and for four real applications with an accuracy of 95.86%. Yiheng Xu, Pranav Sivaraman, Hariharan Devarajan, Kathryn Mohror, Abhinav Bhatele |
HiPC | 3 |
| 2024 | DYAD: Locality-aware Data Management for accelerating Deep Learning TrainingabstractDeep Learning (DL) is increasingly applied across various fields to solve complex scientific challenges in modern high-performance computing (HPC) systems that are beyond the reach of traditional algorithms. Training DL models for scientific applications involves processing multi-terabyte datasets in each epoch. The data access behavior during DL training exposes optimization opportunities to cache these datasets in near-compute storage accelerators in HPC systems, enhancing I/O throughput. However, current middleware solutions employ near-compute storage accelerators primarily as exclusive caches, which limits the effectiveness of cache access locality. To address this problem, we introduce DYAD, a system designed to maximize sample locality in the cache, thereby significantly increasing I/O throughput in HPC systems.DYAD optimizes I/O for DL training based on three key features. First, DYAD boosts inter-node access speeds by using a novel streaming RPC with RDMA protocol, achieving a 1.25x performance gain over state-of-the-art solutions. Second, DYAD further enhances inter-node access by coordinating data movement, which mitigates network congestion and increases throughput for inter-node accesses by up to 8.78x. Last, DYAD uses smart metadata caching that outperforms traditional global metadata access methods by several orders of magnitude in terms of lookup throughput. We demonstrate how DYAD accelerates large-scale DL training on a high-end HPC cluster with 512 GPUs by up to 10.82x faster epochs compared to UnifyFS by performing locality-aware caching on near-compute storage accelerators. Hariharan Devarajan, Ian Lumsden, Chen Wang 0004, Konstantia Georgouli, Thomas Scogland, Jae-Seung Yeom, Michela Taufer |
SBAC-PAD | 1 |
| 2024 | TailorFS: An Adaptive File System to Support Dynamic I/O requirements of HPC WorkloadsabstractHigh-Performance Computing (HPC) systems typically provide a global storage system to support large-scale workload’s I/O requirements. However, these workloads have diversified from traditional simulation to include big data analytics and AI workloads. HPC systems include hardware accelerators and storage software with configurations to support this workload diversification. However, selecting the correct combination of these configurations for each workload is a complex task even for I/O experts. To address this problem, we designed TailorFS, a software abstraction using FSView that transparently selects the appropriate file system characteristics including software abstractions, hardware accelerators, and optimizations to accelerate I/O for a given workload. Given a workload’s behavior, TailorFS chooses an appropriate software existing in the system and a corresponding configuration to dynamically optimize the workload. With TailorFS, users can utilize any interface to perform I/O on the file system and get the correct software selected for them based on the described behavior. Finally, TailorFS can accelerate I/O for different workload types by up to 46 × by utilizing workload characteristics. In conclusion, TailorFS can accelerate complex HPC workflows by up to 6.5 × better I/O performance on the Lassen supercomputer by using dynamically created workload-aware FSViews. Hariharan Devarajan, Kathryn Mohror |
SBAC-PAD | 1 |
| 2024 | DFTracer: An Analysis-Friendly Data Flow Tracer for AI-Driven WorkflowsabstractModern HPC workflows involve intricate coupling of simulation, data analytics, and artificial intelligence (AI) applications to improve time to scientific insight. These workflows require a cohesive set of performance analysis tools to provide a comprehensive understanding of data exchange patterns in HPC systems. However, current tools are not designed to work with an AI-based I/O software stack that requires tracing at multiple levels of the application. To this end, we developed a data flow tracer called DFTracer to capture data-centric events from workflows and the I/O stack to build a detailed understanding of the data exchange within AI-driven workflows. DFTracer has the following three novel features, including a unified interface to capture trace data from different layers in the software stack, a trace format that is analysis-friendly and optimized to support efficiently loading multi-million events in a few seconds, and the capability to tag events with workflow-specific context to perform domain-centric data flow analysis for workflows. Additionally, we demonstrate that DFTracer has a $1.44 x$ smaller runtime overhead and 1.3-7.1x smaller trace size than state-of-the-art tracing tools such as Score-P, Recorder, and Darshan. Moreover, with AI-driven workflows, Score-P, Recorder, and Darshan cannot find I/O accesses from dynamically spawned processes, and their load performance of 100 M events is three orders of magnitude slower than DFTracer. In conclusion, we demonstrate that DFTracer can capture multi-level performance data, including contextual event tagging with a low overhead of 1-5% from AI-driven workflows such as MuMMI and Microsoft’s Megatron Deepspeed running on large-scale HPC systems. Hariharan Devarajan, Loïc Pottier, Kaushik Velusamy, Huihuo Zheng, Izzet Yildirim, Olga Kogiou, Weikuan Yu, Antonios Kougkas, Xian-He Sun, Jae-Seung Yeom, Kathryn Mohror |
SC | 1 |
| 2023 | Mimir: Extending I/O Interfaces to Express User Intent for Complex Workloads in HPCabstractThe complexity of data management in HPC systems stems from the diversity in I/O behavior exhibited by new workloads, multistage workflows, and the presence of multitiered storage systems. This complexity is managed by the storage systems, which provide user-level configurations to allow the tuning of workload I/O within the system. However, these configurations are difficult to set by users who lack expertise in I/O subsystems. We propose a paradigm change in which users specify the intent of I/O operations and storage systems automatically set various configurations based on the supplied intent. To this end, we developed the Mimir infrastructure to assist users in passing I/O intent to the underlying storage system. We demonstrate several use cases that map user-defined intents to storage configurations that lead to optimized I/O. In this study, we make three observations. First, I/O intents should be applied to each level of the I/O storage stack, from HDF5 to MPI-IO to POSIX, and integrated using lightweight adaptors in the existing stack. Second, the Mimir infrastructure supports up to 400M Ops/sec throughput of intents in the system, with a low memory overhead of 6.85KB per node. Third, intents assist in configuring a hierarchical cache to preload I/O, buffer in a node-local device, and store data in a global cache to optimize I/O workloads by 2.33×, 4×, and 2.1×, respectively. Our Mimir infrastructure optimizes complex large-scale workflows by up to 4× better I/O performance on the Lassen supercomputer by using automatically derived I/O intents. Hariharan Devarajan, Kathryn Mohror |
IPDPS | 1 |
| 2022 | Stimulus: Accelerate Data Management for Scientific AI applications in HPCabstractModern scientific workflows couple simulations with AI-powered analytics by frequently exchanging data to accelerate time-to-science to reduce the complexity of the simulation planes. However, this data exchange is limited in performance and portability due to a lack of support for scientific data formats in AI frameworks. We need a cohesive mechanism to effectively integrate at scale complex scientific data formats such as HDF5, PnetCDF, ADIOS2, GNCF, and Silo into popular AI frameworks such as TensorFlow, PyTorch, and Caffe. To this end, we designed Stimulus, a data management library for ingesting scientific data effectively into the popular AI frameworks. We utilize the StimOps functions along with StimPack abstraction to enable the integration of scientific data formats with any AI framework. The evaluations show that Stimulus outperforms several large-scale applications with different use-cases such as Cosmic Tagger (consuming HDF5 dataset in PyTorch), Distributed FFN (consuming HDF5 dataset in TensorFlow), and CosmoFlow (converting HDF5 into TFRecord and then consuming that in TensorFlow) by 5.3 x, 2.9 x, and 1.9 x respectively with ideal I/O scalability up to 768 GPUs on the Summit supercomputer. Through Stimulus, we can portably extend existing popular AI frameworks to cohesively support any complex scientific data format and efficiently scale the applications on large-scale supercomputers. Hariharan Devarajan, Antonios Kougkas, Huihuo Zheng, Venkatram Vishwanath, Xian-He Sun |
CCGRID | 1 |
| 2022 | Extracting and characterizing I/O behavior of HPC workloadsabstractSystem administrators set default storage-system configuration parameters with the goal of providing high per-formance for their system's I/O workloads. However, this gener-alized configuration can lead to suboptimal I/O performance for individual workloads. Users can provide parameter settings to the storage system to obtain better performance for individual applications, but it can be very challenging to determine which parameters to set and to what values. This problem is further ex-acerbated by the increased complexity of modern storage systems. In this work, we move towards solving this problem by providing a systematic categorization of workload-related information that users or middleware libraries can pass to the storage system to optimize I/O performance for specific workloads. We study applications and workflows from different scientific domains to cover a broad range of HPC use cases. Through our categorization, we find that a) workload features differ based on the hardware, software, and data components involved in the execution of workloads and b) multiple workload features together drive I/O optimizations. The methodology proposed in this work optimizes complex scientific workloads by 2.2 x−8 x, using workload-aware I/Ooptimizations. Using the proposed methodology, users can pragmatically characterize their workload, and this characterization can assist the storage system in configuring itself to optimize I/Operformance for individual workloads in HPC systems. Hariharan Devarajan, Kathryn Mohror |
CLUSTER | 1 |
| 2021 | DLIO: A Data-Centric Benchmark for Scientific Deep Learning ApplicationsabstractDeep learning has been shown as a successful method for various tasks, and its popularity results in numerous open-source deep learning software tools. Deep learning has been applied to a broad spectrum of scientific domains such as cosmology, particle physics, computer vision, fusion, and astrophysics. Scientists have performed a great deal of work to optimize the computational performance of deep learning frameworks. However, the same cannot be said for I/O performance. As deep learning algorithms rely on big-data volume and variety to effectively train neural networks accurately, I/O is a significant bottleneck on large-scale distributed deep learning training. This study aims to provide a detailed investigation of the I/O behavior of various scientific deep learning workloads running on the Theta supercomputer at Argonne Leadership Computing Facility. In this paper, we present DLIO, a novel representative benchmark suite built based on the I/O profiling of the selected workloads. DLIO can be utilized to accurately emulate the I/O behavior of modern scientific deep learning applications. Using DLIO, application developers and system software solution architects can identify potential I/O bottlenecks in their applications and guide optimizations to boost the I/O performance leading to lower training times by up to 6.7x. Hariharan Devarajan, Huihuo Zheng, Antonios Kougkas, Xian-He Sun, Venkatram Vishwanath |
CCGRID | 1 |
| 2021 | HFlow: A Dynamic and Elastic Multi-Layered I/O ForwarderabstractModern applications are highly data-intensive, leading to the well-known I/O bottleneck problem. Scientists have proposed the placement of fast intermediate storage resources which aim to mask the I/O penalties. To manage these resources, three core software abstractions are being used in leadership-class computing facilities: IO Forwarders, Burst Buffers, and Data Stagers. Yet, with the rise of multi-tenant deployment in HPC systems, these software abstractions are: managed and maintained in isolation, leading to inefficient interactions; allocated statically, leading to load imbalance; exclusively bifurcated between the intermediate storage, leading to under-utilization of resources, and, in many cases, do not support in-situ operations. To this end, we present HFlow, a new class of data forwarding system that leverages a real-time data movement paradigm. HFlow introduces a unified data movement abstraction (the ByteFlow) providing data-independent tasks that can be executed anywhere and thus, enabling dynamic resource provisioning. Moreover, the processing elements executing the ByteFlows are designed to be ephemeral and, hence, enable elastic management of intermediate storage resources. Our results show that applications running under HFlow display an increase in performance of 3x when compared with state-of-the-art software solutions. Jaime Cernuda, Hariharan Devarajan, Luke Logan, Keith Bateman, Neeraj Rajesh, Antonios Kougkas, Xian-He Sun |
CLUSTER | 2 |
| 2021 | Apollo: : An ML-assisted Real-Time Storage Resource ObserverabstractApplications and middleware services, such as data placement engines, I/O scheduling, and prefetching engines, require low-latency access to telemetry data in order to make optimal decisions. However, typical monitoring services store their telemetry data in a database in order to allow applications to query them, resulting in significant latency penalties. This work presents Apollo: a low-latency monitoring service that aims to provide applications and middleware libraries with direct access to relational telemetry data. Monitoring the system can create interference and overhead, slowing down raw performance of the resources for the job. However, having a current view of the system can aid middleware services in making more optimal decisions which can ultimately improve the overall performance. Apollo has been designed from the ground up to provide low latency, using Publish-Subscriber Pub-Sub semantics, and low overhead, using adaptive intervals in order to change the length of time between polling the resource for telemetry data and machine learning in order to predict changes to the telemetry data between actual resource polling. This work also provides some high level abstractions called I/O curators, which can further aid middleware libraries and applications to make optimal decisions. Evaluations showcase that Apollo can achieve sub-millisecond latency for acquiring complex insights with a memory overhead of ~57 MB and CPU overhead being only 7% more than existing state-of-the-art systems. Neeraj Rajesh, Hariharan Devarajan, Jaime Cernuda, Keith Bateman, Luke Logan, Antonios Kougkas, Xian-He Sun |
HPDC | 2 |
| 2020 | HReplica: A Dynamic Data Replication Engine with Adaptive Compression for Multi-Tiered StorageabstractAs the diversity of big data applications increases, their requirements diverge and often conflict with one other. Managing this diversity in any supercomputer or data center is a major challenge for system designers. Data replication is a popular approach to meet several of these requirements, such as low latency, read availability, durability, etc. This approach can be enhanced using new modern heterogeneous hardware and software techniques such as data compression. However, both these enhancements work in isolation to the detriment of both. In this work, we present HReplica: a dynamic data replication engine which harmoniously leverages data compression and hierarchical storage to increase the effectiveness of data replication. We have developed a novel dynamic selection algorithm that facilitates the optimal matching of replication schemes, compression libraries, and tiered storage. Our evaluation shows that HReplica can improve scientific and cloud application performance by 5.2x when compared to other state-of-the-art replication schemes. Hariharan Devarajan, Antonios Kougkas, Xian-He Sun |
IEEE BigData | 1 |
| 2020 | HCL: Distributing Parallel Data Structures in Extreme ScalesabstractMost parallel programs use irregular control flow and data structures, which are perfect for one-sided communication paradigms such as MPI or PGAS programming languages. However, these environments lack the presence of efficient function-based application libraries that can utilize popular communication fabrics such as TCP, Infinity Band (IB), and RDMA over Converged Ethernet (RoCE). Additionally, there is a lack of high-performance data structure interfaces. We present Hermes Container Library (HCL), a high-performance distributed data structures library that offers high-level abstractions including hash-maps, sets, and queues. HCL uses a RPC over RDMA technology that implements a novel procedural programming paradigm. In this paper, we argue a RPC over RDMA technology can serve as a high-performance, flexible, and co-ordination free backend for implementing complex data structures. Evaluation results from testing real workloads shows that HCL programs are 2x to 12x faster compared to BCL, a state-of-the-art distributed data structure library. Hariharan Devarajan, Antonios Kougkas, Keith Bateman, Xian-He Sun |
CLUSTER | 1 |
| 2020 | HCompress: Hierarchical Data Compression for Multi-Tiered Storage EnvironmentsabstractModern scientific applications read and write massive amounts of data through simulations, observations, and analysis. These applications spend the majority of their runtime in performing I/O. HPC storage solutions include fast node-local and shared storage resources to elevate applications from this bottleneck. Moreover, several middleware libraries (e.g., Hermes) are proposed to move data between these tiers transparently. Data reduction is another technique that reduces the amount of data produced and, hence, improve I/O performance. These two technologies, if used together, can benefit from each other. The effectiveness of data compression can be enhanced by selecting different compression algorithms according to the characteristics of the different tiers, and the multi-tiered hierarchy can benefit from extra capacity. In this paper, we design and implement HCompress, a hierarchical data compression library that can improve the application's performance by harmoniously leveraging both multi-tiered storage and data compression. We have developed a novel compression selection algorithm that facilitates the optimal matching of compression libraries to the tiered storage. Our evaluation shows that HCompress can improve scientific application's performance by 7x when compared to other state-of-the-art tiered storage solutions. Hariharan Devarajan, Antonios Kougkas, Luke Logan, Xian-He Sun |
IPDPS | 1 |
| 2020 | HFetch: Hierarchical Data Prefetching for Scientific Workflows in Multi-Tiered Storage EnvironmentsabstractIn the era of data-intensive computing, accessing data with a high-throughput and low-latency is more imperative than ever. Data prefetching is a well-known technique for hiding read latency. However, existing solutions do not consider the new deep memory and storage hierarchy and also suffer from under-utilization of prefetching resources and unnecessary evictions. Additionally, existing approaches implement a client-pull model where understanding the application's I/O behavior drives prefetching decisions. Moving towards exascale, where machines run multiple applications concurrently by accessing files in a workflow, a more data-centric approach can resolve challenges such as cache pollution and redundancy. In this study, we present HFetch, a truly hierarchical data prefetcher that adopts a server-push approach to data prefetching. We demonstrate the benefits of such an approach. Results show 10-35% performance gains over existing prefetchers and over 50% when compared to systems with no prefetching. Hariharan Devarajan, Antonios Kougkas, Xian-He Sun |
IPDPS | 1 |
| 2020 | I/O Acceleration via Multi-Tiered Data Buffering and Prefetching
Antonios Kougkas, Hariharan Devarajan, Xian-He Sun |
J. Comput. Sci. Technol. | 2 |
| 2020 | Bridging Storage Semantics Using Data Labels and Asynchronous I/OabstractIn the era of data-intensive computing, large-scale applications, in both scientific and the BigData communities, demonstrate unique I/O requirements leading to a proliferation of different storage devices and software stacks, many of which have conflicting requirements. Further, new hardware technologies and system designs create a hierarchical composition that may be ideal for computational storage operations. In this article, we investigate how to support a wide variety of conflicting I/O workloads under a single storage system. We introduce the idea of a Label , a new data representation, and, we present LABIOS: a new, distributed, Label- based I/O system. LABIOS boosts I/O performance by up to 17× via asynchronous I/O, supports heterogeneous storage resources, offers storage elasticity, and promotes in situ analytics and software defined storage support via data provisioning. LABIOS demonstrates the effectiveness of storage bridging to support the convergence of HPC and BigData workloads on a single platform. Antonios Kougkas, Hariharan Devarajan, Xian-He Sun |
ACM Trans. Storage | 2 |
| 2019 | NIOBE: An Intelligent I/O Bridging Engine for Complex and Distributed WorkflowsabstractIn the age of data-driven computing, integrating High Performance Computing(HPC) and Big Data(BD) environments may be the key to increasing productivity and to driving scientific discovery forward. Scientific workflows consist of diverse applications (i.e., HPC simulations and BD analysis) each with distinct representations of data that introduce a semantic barrier between the two environments. To solve scientific problems at scale, accessing semantically different data from different storage resources is the biggest unsolved challenge. In this work, we aim to address a critical question: ”How can we exploit the existing resources and efficiently provide transparent access to data from/to both environments”. We propose iNtelligent I/O Bridging Engine(NIOBE), a new data integration framework that enables integrated data access for scientific workflows with asynchronous I/O and data aggregation. NIOBE performs the data integration using available I/O resources, in contrast to existing optimizations that ignore the I/O nodes present on the data path. In NIOBE, data access is optimized to consider both the ongoing production and the consumption of the data in the future. Experimental results show that with NIOBE, an integrated scientific workflow can be accelerated by up to 10x when compared to a no-integration baseline and by up to 133% compared to other state-of-the-art integration solutions. Hariharan Devarajan, Antonios Kougkas, Xian-He Sun |
IEEE BigData | 2 |
| 2019 | An Intelligent, Adaptive, and Flexible Data Compression FrameworkabstractThe data explosion phenomenon in modern applications causes tremendous stress on storage systems. Developers use data compression, a size-reduction technique, to address this issue. However, each compression library exhibits different strengths and weaknesses when considering the input data type and format. We present Ares, an intelligent, adaptive, and flexible compression framework which can dynamically choose a compression library for a given input data based on the type of the workload and provides an appropriate infrastructure to users to fine-tune the chosen library. Ares is a modular framework which unifies several compression libraries while allowing the addition of more compression libraries by the user. Ares is a unified compression engine that abstracts the complexity of using different compression libraries for each workload. Evaluation results show that under real-world applications, from both scientific and Cloud domains, Ares performed 2-6x faster than competitive solutions with a low cost of additional data analysis (i.e., overheads around 10%) and up to 10x faster against a baseline of no compression at all. Hariharan Devarajan, Antonios Kougkas, Xian-He Sun |
CCGRID | 1 |
| 2019 | LABIOS: A Distributed Label-Based I/O SystemabstractIn the era of data-intensive computing, large-scale applications, in both scientific and the BigData communities, demonstrate unique I/O requirements leading to a proliferation of different storage devices and software stacks, many of which have conflicting requirements. In this paper, we investigate how to support a wide variety of conflicting I/O workloads under a single storage system. We introduce the idea of a Label, a new data representation, and, we present LABIOS: a new, distributed, Label- based I/O system. LABIOS boosts I/O performance by up to 17x via asynchronous I/O, supports heterogeneous storage resources, offers storage elasticity, and promotes in-situ analytics via data provisioning. LABIOS demonstrates the effectiveness of storage bridging to support the convergence of HPC and BigData workloads on a single platform. Antonios Kougkas, Hariharan Devarajan, Jay F. Lofstead, Xian-He Sun |
HPDC | 2 |
| 2018 | Harmonia: An Interference-Aware Dynamic I/O Scheduler for Shared Non-volatile Burst BuffersabstractModern HPC systems employ burst buffer installations to reduce the peak I/O requirements for external storage and deal with the burstiness of I/O in modern scientific applications. These I/O buffering resources are shared between multiple applications that run concurrently. This leads to severe performance degradation due to contention, a phenomenon called cross-application I/O interference. In this paper, we first explore the negative effects of interference at the burst buffer layer and we present two new metrics that can quantitatively describe the slowdown applications experience due to interference. We introduce Harmonia, a new dynamic I/O scheduler that is aware of interference, adapts to the underlying system, implements a new 2-way decision-making process and employs several scheduling policies to maximize the system efficiency and applications' performance. Our evaluation shows that Harmonia, through better I/O scheduling, can outperform by 3× existing state-of-the-art buffering management solutions and can lead to better resource utilization. Antonios Kougkas, Hariharan Devarajan, Xian-He Sun, Jay F. Lofstead |
CLUSTER | 2 |
| 2018 | Vidya: Performing Code-Block I/O Characterization for Data Access OptimizationabstractUnderstanding, characterizing and tuning scientific applications' I/O behavior is an increasingly complicated process in HPC systems. Existing tools use either offline profiling or online analysis to get insights into the applications' I/O patterns. However, there is lack of a clear formula to characterize applications' I/O. Moreover, these tools are application specific and do not account for multi-tenant systems. This paper presents Vidya, an I/O profiling framework which can predict application's I/O intensity using a new formula called Code-Block I/O Characterization (CIOC). Using CIOC, developers and system architects can tune an application's I/O behavior and better match the underlying storage system to maximize performance. Evaluation results show that Vidya can predict an application's I/O intensity with a variance of 0.05%. Vidya can profile applications with a high accuracy of 98% while reducing profiling time by 9x. We further show how Vidya can optimize an application's I/O time by 3.7x. Hariharan Devarajan, Antonios Kougkas, Prajwal Challa, Xian-He Sun |
HiPC | 1 |
| 2018 | Hermes: a heterogeneous-aware multi-tiered distributed I/O buffering systemabstractModern High-Performance Computing (HPC) systems are adding extra layers to the memory and storage hierarchy named deep memory and storage hierarchy (DMSH), to increase I/O performance. New hardware technologies, such as NVMe and SSD, have been introduced in burst buffer installations to reduce the pressure for external storage and boost the burstiness of modern I/O systems. The DMSH has demonstrated its strength and potential in practice. However, each layer of DMSH is an independent heterogeneous system and data movement among more layers is significantly more complex even without considering heterogeneity. How to efficiently utilize the DMSH is a subject of research facing the HPC community. In this paper, we present the design and implementation of Hermes: a new, heterogeneous-aware, multi-tiered, dynamic, and distributed I/O buffering system. Hermes enables, manages, supervises, and, in some sense, extends I/O buffering to fully integrate into the DMSH. We introduce three novel data placement policies to efficiently utilize all layers and we present three novel techniques to perform memory, metadata, and communication management in hierarchical buffering systems. Our evaluation shows that, in addition to automatic data movement through the hierarchy, Hermes can significantly accelerate I/O and outperforms by more than 2x state-of-the-art buffering platforms. Antonios Kougkas, Hariharan Devarajan, Xian-He Sun |
HPDC | 2 |
| 2018 | IRIS: I/O Redirection via Integrated StorageabstractThere is an ocean of available storage solutions in modern high-performance and distributed systems. These solutions consist of Parallel File Systems (PFS) for the more traditional high-performance computing (HPC) systems and of Object Stores for emerging cloud environments. More of ten than not, these storage solutions are tied to specific APIs and data models and thus, bind developers, applications, and entire computing facilities to using certain interfaces. Each storage system is designed and optimized for certain applications but does not perform well for others. Furthermore, modern applications have become more and more complex consisting of a collection of phases with different computation and I/O requirements. In this paper, we propose a unified storage access system, called IRIS (i.e., I/O Redirection via Integrated Storage). IRIS enables unified data access and seamlessly bridges the semantic gap between file systems and object stores. With IRIS, emerging High-Performance Data Analytics software has capable and diverse I/O support. IRIS can bring us closer to the convergence of HPC and Cloud environments by combining the best storage subsystems from both worlds. Experimental results show that IRIS can grant more than 7x improvement in performance than existing solutions. Antonios Kougkas, Hariharan Devarajan, Xian-He Sun |
ICS | 2 |