Pradeep Subedi

dblp:137/0625 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0003-4281-9674ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 5 first-author · 8 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2024 ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
abstract
Compute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coherent memory space between host and CXL storage devices such as such as Solid-state drive (SSD) for memory expansion, with a corresponding DRAM implemented as the device cache. However, this introduces challenges such as substantial cache miss penalties, sub-optimal caching due to data access granularity mismatch between the DRAM "cache" and SSD "memory", and inefficient hardware cache management. To address these issues, we propose a novel solution, named ICGMM, which optimizes caching and eviction directly on hardware, employing a Gaussian Mixture Model (GMM)-based approach. We prototype our solution on an FPGA board, which demonstrates a noteworthy improvement compared to the classic Least Recently Used (LRU) cache strategy. We observe a decrease in the cache miss rate ranging from 0.32% to 6.14%, leading to a substantial 16.23% to 39.14% reduction in the average SSD access latency. Furthermore, when compared to the state-of-the-art Long Short-Term Memory (LSTM)-based cache policies, our GMM algorithm on FPGA showcases an impressive latency reduction of over 10,000 times. Remarkably, this is achieved while demanding much fewer hardware resources.
Hanqiu Chen, Yitu Wang, Vitorio Cargnini, Mohammadreza Soltaniyeh, Gongjin Sun, Pradeep Subedi, Yiran Chen 0001, Cong Hao
DAC7
2024 Residual-INR: Communication Efficient On-Device Learning Using Implicit Neural Representation
abstract
Edge computing is a distributed computing paradigm that collects and processes data at or near the source of data generation. The on-device learning at edge relies on device-to-device wireless communication to facilitate real-time data sharing and collaborative decision-making among multiple devices. This significantly improves the adaptability of the edge computing system to the changing environments. However, as the scale of the edge computing system is getting larger, communication among devices is becoming the bottleneck because of the limited bandwidth of wireless communication leads to large data transfer latency. To reduce the amount of device-to-device data transmission and accelerate on-device learning, in this paper, we propose Residual-INR, a fog computing-based communication-efficient on-device learning framework by utilizing implicit neural representation (INR) to compress images/videos into neural network weights. Residual-INR enhances data transfer efficiency by collecting JPEG images from edge devices, compressing them into INR format at the fog node, and redistributing them for on-device learning. By using a smaller INR for full image encoding and a separate object INR for high-quality object region reconstruction through residual encoding, our technique can reduce the encoding redundancy while maintaining the object quality. Residual-INR is a promising solution for edge on-device learning because it reduces data transmission by up to 5.16 × across a network of 10 edge devices. It also facilitates CPU-free accelerated on-device learning, achieving up to 2.9 × speedup without sacrificing accuracy. Our code is available at: https://github.com/sharc-lab/Residual-INR.
Hanqiu Chen, Xuebin Yao, Pradeep Subedi, Cong Hao
ICCAD3
2023 Benesh: a Framework for Choreographic Coordination of In Situ Workflows
abstract
The growing scale of high-performance computing systems increasingly enables scientists to develop more complex applications as in situ workflows composed of coupled simulation and analysis codes. It is therefore important that workflow programming systems and runtime middleware support the composition and execution of these complex applications intuitively and efficiently. The scientific computing community has put significant effort into purpose-built coupled simulation codes that have been optimized for specialized use cases. However, the development effort involving the coupling of established codes has been largely ad hoc. The Benesh programming system was recently proposed to support the development of coupled simulation workflows from existing code bases. Benesh allows a shared data model to be defined across established codes, so that they can be interfaced in a flexible, coupled workflow. In this paper, we develop Benesh into a workflow development framework. Using Benesh, we develop workflow data model and data exchange definitions for coupled, in situ workflows. We evaluate the cost of development using Benesh in terms of development time and overhead, showing that Benesh offers development advantages without undue impact upon workflow performance.
Philip E. Davis, Jacob S. Merson, Pradeep Subedi, Lee F. Ricketson, Cameron W. Smith, Mark S. Shephard, Manish Parashar
HiPC3
2023 SRC: Mitigate I/O Throughput Degradation in Network Congestion Control of Disaggregated Storage Systems
abstract
The industry has adopted disaggregated storage systems to provide high-quality services for hyper-scale architectures. This infrastructure enables organizations to access storage resources that can be independently managed, configured, and scaled. It is supported by the recent advances of all-flash arrays and NVMe-over-Fabric protocol, enabling remote access to NVMe devices over different network fabrics. A surge of research has been proposed to mitigate network congestion in traditional remote direct memory access protocol (RDMA). However, NVMe-oF raises new challenges in congestion control for disaggregated storage systems.In this work, we investigate the performance degradation of the read throughput on storage nodes caused by traditional network congestion control mechanisms. We design a storage-side rate control (SRC) to relieve network congestion while avoiding performance degradation on storage nodes. First, we design an I/O throughput control mechanism in the NVMe driver layer to enable throughput control on storage nodes. Second, we construct a throughput prediction model to learn a mapping function between workload characteristics and I/O throughput. Third, we deploy SRC on storage nodes to cooperate with traditional network congestion control on an NVMe-over-RDMA architecture. Finally, we evaluate SRC with varying workloads, SSD configurations, and network topologies. The experimental results show that SRC achieves significant performance improvement.
Danlin Jia, Xuebin Yao, Mahsa Bayati, Pradeep Subedi, Bo Sheng, Ningfang Mi
IPDPS8
2023 Adaptive elasticity policies for staging-based in situ visualization
Zhe Wang 0059, Matthieu Dorier, Pradeep Subedi, Philip E. Davis, Manish Parashar
Future Gener. Comput. Syst.3
2022 Assembling Portable In-Situ Workflow from Heterogeneous Components using Data Reorganization
abstract
Heterogeneous computing is becoming common in the HPC world. The fast-changing hardware landscape is pushing programmers and developers to rely on performance-portable programming models to rewrite old and legacy applications and develop new ones. While this approach is suitable for individual applications, outstanding challenges still remain when multiple applications are combined into complex workflows. One critical difficulty is the exchange of data between communicating applications where performance constraints imposed by heterogeneous hardware advantage different data layouts. We attempt to solve this problem by exploring asynchronous data layout conversions for applications requiring different memory access patterns for shared data. We implement the proposed solution within the DataSpaces data staging service, extending it to support heterogeneous application workflows across a broad spectrum of programming models. In addition, we integrate heterogeneous DataSpaces with the Kokkos programming model and propose the Kokkos Staging Space as an extension of the Kokkos data abstraction. This new abstraction enables us to express data on a virtual shared space for multiple Kokkos applications, thus guaranteeing the portability of each application when assembling them into an efficient heterogeneous workflow. We present performance results for the Kokkos Staging Space using a synthetic workflow emulator and three different scenarios representing access frequency and use patterns in shared data. The results show that the Kokkos Staging Space is a superior solution in terms of time-to-solution and scalability compared to existing file-based Kokkos data abstractions for inter-application data exchange.
Bo Zhang 0120, Pradeep Subedi, Philip E. Davis, Francesco Rizzi, Keita Teranishi, Manish Parashar
CCGRID2
2021 RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows
abstract
While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. We experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.
Pradeep Subedi, Philip E. Davis, Manish Parashar
CLUSTER1
2021 Adaptive Placement of Data Analysis Tasks For Staging Based In-Situ Processing
abstract
In-situ processing addresses the gap between speeds of computing and I/O capabilities by processing data close to the data source, i.e., on the same system as the data source (e.g., a simulation). However, the effective implementation of in-situ processing workflows requires the optimization of several design parameters such as where on the system workflow data analysis/visualization (ana/vis) as placed and how execution as well as the interaction and data exchanges between ana/vis are coordinated. For example, in the case of hybrid in-situ processing, interacting ana/vis may be tightly or loosely coupled depending on their placement, and this can lead to very different performance and scalability. A key challenge is deciding the most appropriate ana/vis placement, which depends on dynamic applications, workflow, and system characteristics that might change at runtime. In this paper, we present a framework to support online adaptive data analysis placement during the execution of an in-situ workflow. Specifically, the paper presents a model and architecture, and explores several data analysis placement strategies. Evaluation results show that dynamically choosing appropriate data analysis placement strategies can balance the benefits and overhead of different data analysis placement patterns to reduce in-situ processing time.
Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar
HiPC2
2020 Staging Based Task Execution for Data-driven, In-Situ Scientific Workflows
abstract
As scientific workflows increasingly use extreme-scale resources, the imbalance between higher computational capabilities, generated data volumes, and available I/O bandwidth is limiting the ability to translate these scales into insights. In-situ workflows (and the in-situ approach) are leveraging storage levels close to the computation in novel ways in order to reduce the required I/O. However, to be effective, it is important that the mapping and execution of such in-situ workflows adopts a data-driven approach, enabling in-situ tasks to be executed flexibly based upon data content. This paper first explores the design space for data-driven in-situ workflows. Specifically, it presents a model that captures different factors that influence the mapping, execution, and performance of data-driven in-situ workflows and experimentally studies the impact of different mapping decisions and execution patterns. The paper then presents the design, implementation, and experimental evaluation of a data-driven in-situ workflow execution framework that leverages in-memory distributed data management and user-defined task-triggers to enable efficient and scalable in-situ workflow execution.
Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar
CLUSTER2
2019 Leveraging Machine Learning for Anticipatory Data Delivery in Extreme Scale In-situ Workflows
abstract
Extreme scale scientific workflows are composed of multiple applications that exchange data at runtime. Several data-related challenges are limiting the potential impact of such workflows. While data staging and in-situ models of execution have emerged as approaches to address data-related costs at extreme scales, increasing data volumes and complex data exchange patterns impact the effectiveness of such approaches. In this paper, we design and implement DESTINY, which is an autonomic data delivery mechanism for staging-based in-situ workflows. DESTINY dynamically learns the data access patterns of scientific workflow applications and leverages these patterns to decrease data access costs. Specifically, DESTINY uses machine learning techniques to anticipate future data accesses, proactively packages and delivers the data necessary to satisfy these requests as close to the consumer as possible and, when data staging processes and consumer processes are colocated, removes the need for inter-process communication by making these data available to the consumer as shared-memory objects. When consumer processes reside on nodes other than staging nodes, the data is packaged and stored in a format the client will likely access in future. This amortizes expensive data discovery and assembly operations typically associated with data staging. We experimentally evaluate the performance and scalability of DESTINY on leadership class platforms using synthetic applications and the S3D combustion workflow. We demonstrate that DESTINY is scalable and can achieve a reduction of up to 75% in read response time as compared to in-memory staging service for production scientific workflows.
Pradeep Subedi, Philip E. Davis, Manish Parashar
CLUSTER1
2019 Addressing data resiliency for staging based scientific workflows
abstract
As applications move towards extreme scales, data-related challenges are becoming significant concerns, and in-situ workflows based on data staging and in-situ/in-transit data processing have been proposed to address these challenges. Increasing scale is also expected to result in an increase in the rate of silent data corruption errors, which will impact both the correctness and performance of applications. Furthermore, this impact is amplified in the case of in-situ workflows due to the dataflow between the component applications of the workflow. While existing research has explored silent error detection at the application level, silent error detection for workflows remains an open challenge. This paper addresses silent error detection for extreme scale in-situ workflows. The presented approach leverages idle computation resource in data staging to enable timely detection and recovery from silent data corruption, effectively reducing the propagation of corrupted data and end-to-end workflow execution time in the presence of silent errors. As an illustration of this approach, we use a spatial outlier detection approach in staging to detect errors introduced in data transfer and storage. We also provide a CPU-GPU hybrid staging framework for error detection in order to achieve faster error identification. We have implemented our approach within the DataSpaces staging service, and evaluated it using both synthetic and real workflows on a Cray XK7 system (Titan) at different scales. We demonstrate that, in the presence of silent errors, enabling error detection on staged data alongside a checkpoint/restart scheme improves the total in-situ workflow execution time by up to 22% in comparison with using checkpoint/restart alone.
Shaohua Duan, Pradeep Subedi, Philip E. Davis, Manish Parashar
SC2
2018 Scalable Data Resilience for In-memory Data Staging
abstract
The dramatic increase in the scale of current and planned high-end HPC systems is leading new challenges, such as the growing costs of data movement and IO, and the reduced mean times between failures (MTBF) of system components. In-situ workflows, i.e., executing the entire application workflows on the HPC system, have emerged as an attractive approach to address data-related challenges by moving computations closer to the data, and staging-based frameworks have been effectively used to support in-situ workflows at scale. However, the resilience of these staging-based solutions has not been addressed and they remain susceptible to expensive data failures. Furthermore, naive use of data resilience techniques such as n-way replication and erasure codes can impact latency and/or result in significant storage overheads. In this paper, we present CoREC, a scalable resilient in-memory data staging runtime for large-scale in-situ workflows. CoREC uses a novel hybrid approach that combines dynamic replication with erasure coding based on data access patterns. The paper also presents optimizations for load balancing and conflict avoiding encoding, and a low overhead, lazy data recovery scheme. We have implemented the CoREC runtime and have deployed with the DataSpaces staging service on Titan at ORNL, and present an experimental evaluation in the paper. The experiments demonstrate that CoREC can tolerate in-memory data failures while maintaining low latency and sustaining high overall storage efficiency at large scales.
Shaohua Duan, Pradeep Subedi, Keita Teranishi, Philip E. Davis, Hemanth Kolla, Marc Gamell, Manish Parashar
IPDPS2
2018 Exploring Power Budget Scheduling Opportunities and Tradeoffs for AMR-Based Applications
abstract
Computational demand has brought major changes to Advanced Cyber-Infrastructure (ACI) architectures. It is now possible to run scientific simulations faster and obtain more accurate results. However, power and energy have become critical concerns. Also, the current roadmap toward the new generation of ACI includes power budget as one of the main constraints. Current research efforts have studied power and performance tradeoffs and how to balance these (e.g., using Dynamic Voltage and Frequency Scaling (DVFS) and power capping for meeting power constraints, which can impact performance). However, applications may not tolerate degradation in performance, and other tradeoffs need to be explored to meet power budgets (e.g., involving the application in making energy-performance-quality tradeoff decisions). This paper proposes using the properties of AMR-based algorithms (e.g., dynamically adjusting the resolution of a simulation in combination with power capping techniques) to schedule or re-distribute the power budget. It specifically explores the opportunities to realize such an approach using checkpointing as a proof-of-concept use case and provides a characterization of a representative set of applications that use Adaptive Mesh Refinement (AMR) methods, including a Low-Mach-Number Combustion (LMC) application. It also explores the potential of utilizing power capping to understand power-quality tradeoffs via simulation.
Yubo Qin, Ivan Rodero, Pradeep Subedi, Manish Parashar, Sandro Rigo
SBAC-PAD3
2018 Stacker: an autonomic data movement engine for extreme-scale data staging-based in-situ workflows
Pradeep Subedi, Philip E. Davis, Shaohua Duan, Scott Klasky, Hemanth Kolla, Manish Parashar
SC1
2016 CoARC: Co-operative, Aggressive Recovery and Caching for Failures in Erasure Coded Hadoop
abstract
Cloud file systems like Hadoop have become a norm for handling big data because of the easy scaling and distributed storage layout. However, these systems are susceptible to failures and data needs to be recovered when a failure is detected. During temporary failures, MapReduce jobs or file system clients perform degraded reads and satisfy the read request. We argue that lack of sharing of the recovered data during degraded reads and recovery of only the requested data block places a heavy strain on the system's network resources and increases the job execution time. To this end, we propose CoARC (Co-operative, Aggressive Recovery and Caching), which is a new data-recovery mechanism for unavailable data during degraded reads in distributed file systems. The main idea is to recover not only the data block that was requested but also other temporarily unavailable blocks in the same strip and cache them in a separate data node. We also propose an LRF (Least Recently Failed) cache replacement algorithm for such a kind of recovery caches. We also show that CoARC significantly reduces the network usage and job runtime in erasure coded Hadoop.
Pradeep Subedi, Ping Huang 0001, Tong Liu 0030, Joseph Moore, Stan Skelton, Xubin He
ICPP1
2015 PPM: A Partitioned and Parallel Matrix Algorithm to Accelerate Encoding/Decoding Process of Asymmetric Parity Erasure Codes
abstract
Erasure codes are widely deployed in storage systems and the encoding/decoding process is a common operation in erasure-coded systems. Parity-check matrix method is a general method employed in erasure codes to conduct encoding/decoding process. However, the process is serial and generates high computational cost in dealing with matrix operations, and hence, causes low encoding/decoding performance. Especially for some recently proposed erasure codes, including SD code, PMDS code, and LRC code, the disadvantages are more obvious. To address this issue, in this paper, we present an optimization algorithm, called Partitioned and Parallel Matrix (PPM) algorithm, to accelerate the encoding/decoding processes of these codes by partitioning the parity-check matrix, parallelizing the encoding/decoding operations, and optimizing the calculation sequence, so as to achieve the goal of fast encoding/decoding. Experimental results show that PPM can speed up the encoding/decoding process of these codes by up to 210.81%.
Qiang Cao 0001, Shenggang Wan, Wenhui Zhang 0005, Changsheng Xie 0001, Xubin He, Pradeep Subedi
ICPP7
2015 FINGER: A novel erasure coding scheme using fine granularity blocks to improve Hadoop write and update performance
abstract
With the explosive increase of the data by volume in various fields of science, engineering, information services, etc., data-intensive computing has gained significant interest in recent years. Various challenges ranging from efficient peta-scale data management to the adoption of highly scalable cloud computing have become a norm for data center administrators. Highly scalable architectures such as Hadoop, BlobSeer and MapR are used in large data centers for efficient data management, and employ 3-way replication for fault tolerance or data availability. One means of reducing storage overhead of replication in data-centers is erasure coding. However, HDFS-RAID (erasure-coded Hadoop) uses large block sizes and does not support update operations. Therefore, changing any file-block content requires recreating the whole file, which effectively reduces the overall write and update performance of the system. We propose FINe Grained ERasure coding scheme (FINGER) for the erasure-coded Hadoop FileSystem, which improves both write and update performance without sacrificing the read performance. The main idea is to chunk the large block size (64 or 128 MB) into smaller chunks; the chunk layout is designed to mitigate extra reads when performing erasure coding on a large block update and maintains the same metdata size as HDFS-RAID. We implement the update operation in Hadoop and conduct testbed experiments to demonstrate that FINGER improves the write and update performance by 38.20% and 8.6% w.r.t. 3-way replication and by 8.08% and up to 5.68×w.r.t HDFS-RAID respectively.
Pradeep Subedi, Ping Huang 0001, Xubin He
NAS1
2014 A hybrid erasure-coded ECC scheme to improve performance and reliability of solid state drives
abstract
The high performance and ever-increasing capacity of flash memory has led to the rapid adoption of Solid-State Disks (SSDs) in mass storage systems. In order to increase disk capacity, multi-level cells (MLC) are used in the design of SSDs, but the use of such SSDs in persistent storage systems raise concerns for users due to the low reliability of such disks. In this paper, we present a hybrid erasure-coded (EECC) architecture that incorporates ECC schemes and erasure codes to improve both performance and reliability. As weak error-correction codes have faster decoding speed than complex error correction codes (ECC), we propose the use of weak-ECC at the segment level rather than complex ECC. To compensate the reduced correction ability of weak-ECC, we use an erasure code that is striped across segments rather than pages or blocks. We use a small sized HDD to store parities so that we can leverage parallelism across multiple devices and remove the parity updates from the critical write path. We carry out simulation experiments based on Disksim to demonstrate that our proposed scheme is able reduce the SSD average read-latency by up to 31.23% and along with tolerance from double chip failures, it dramatically reduces the uncorrectable page error rate.
Pradeep Subedi, Ping Huang 0001, Xubin He, Ming Zhang 0026, Jizhong Han
IPCCC1
2014 FlexECC: Partially Relaxing ECC of MLC SSD for Better Cache Performance
Ping Huang 0001, Pradeep Subedi, Xubin He, Shuang He, Ke Zhou 0001
USENIX ATC2