EDBT 2026 Demo / reviewers in the wild / expert
Md Hasanur Rahman 0001
dblp:304/5686
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
0009-0002-5540-8751ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modeling Rate-Distortion for Endpoint-Aware Lossy Compression in Scientific Data TransferabstractHigh-performance computing (HPC) systems generate massive scientific datasets, often stored in remote data repositories. Limited bandwidth and resource-constrained endpoints pose challenges for efficient large data transfer. Error-bounded lossy compression addresses this by reducing data sizes (higher bit-rates) while controlling distortion. However, different compressors exhibit distinct rate-distortion behaviors even under the same error bounds. Thus, selecting an optimal compressor before transfer is essential to meet endpoint-specific requirements e.g., maximizing data reduction at a fixed distortion. Existing trial-and-error approaches require multiple costly full-scale compression runs to reach at target requirements, making them impractical for such online use. To address this, we propose OptRD, a compressor-agnostic framework that efficiently models rate-distortion trade-offs across multiple lossy compressors by analyzing spatial data traits at reduced resolutions. Evaluated using 3 state-of-the-art lossy compressors on 30 scientific datasets from 4 HPC applications, OPTRD incurs only$\sim 5 \%$average estimation error and achieves over$100 \times$runtime speedup compared to trial-and-error methods, significantly improving optimal compressor selection during such data transfer use cases. Md Hasanur Rahman 0001, Sheng Di, Guanpeng Li, Franck Cappello |
IPCCC | 1 |
| 2025 | Deploying Lightweight Input-Aware Selective Instruction Duplication in HPC ApplicationsabstractModern high-performance computing (HPC) applications are increasingly vulnerable to silent data corruptions (SDCs) caused by transient hardware faults. While selective instruction duplication (SID) offers an efficient software-level protection strategy, existing SID methods rely on SDC vulnerability profile derived from only the default reference input often found in application suites. However, they overlook the input-dependent nature of SDC propagation. This leads to significant SDC coverage loss when inputs vary. We present Protego, a novel input-aware SID protection framework that efficiently adapts protection to runtime inputs. Protego performs a one-time vulnerability-guided input exploration to identify a small number of input groups with distinct SID protection patterns. At runtime, Protego uses lightweight features derived from input arguments to select and deploy the appropriate SID protection. Our evaluation across 10 HPC applications demonstrates the effectiveness and efficiency of Protego in mitigating SDC coverage loss across diverse inputs, compared to existing SID techniques. Md Hasanur Rahman 0001, Guanpeng Li |
SC | 1 |
| 2024 | Significantly Improving Fixed-Ratio Compression Framework for Resource-limited ApplicationsabstractScientific simulations running on HPC facilities generate massive amount of data, putting significant pressure onto supercomputers’ storage capacity and network bandwidth. To alleviate this problem, there has been a rich body of work on reducing data volumes via error-controlled lossy compression. However, fixed-ratio compression is not very well-supported, not allowing users to appropriately allocate memory/storage space or know the data transfer time over the network in advance. To address this problem, recent ratio-controlled frameworks, such as FXRZ, have incorporated methods to predict required error bound settings to reach a user-specified compression ratio. However, these approaches fail to achieve fixed-ratio compression in an accurate, efficient and scalable fashion on diverse datasets and compression algorithms. Tri Nguyen 0002, Md Hasanur Rahman 0001, Sheng Di, Michela Becchi |
ICPP | 2 |
| 2024 | Druto: Upper-Bounding Silent Data Corruption Vulnerability in GPU ApplicationsabstractDue to the increasing scale of high-performance computing (HPC) systems, transient hardware faults have become a major reliability concern. Consequently, Silent Data Corruptions (SDCs) due to these faults have been a common insidious consequence in GPU applications. Developers often measure the application resilience with a set of program test inputs available in the benchmark suite, assuming the resilience would not fluctuate much among different inputs. However, we observe that this assumption often results in an over-optimistic evaluation for GPU applications. As a result, the subsequent SDC protection following the evaluation can hardly meet the expected reliability bar in the production environment, where applications would run with potentially arbitrary input values. To this end, we propose Druto – a compiler-based automated technique that searches for inputs to incrementally approach the upper bound of a GPU application’s SDC probability. We develop Druto based on the property that the resilience profiles of a small group of representative threads in a GPU kernel can approximately rank various inputs in terms of the overall SDC probability. Therefore, Druto strategically steers the search towards new program inputs that efficiently portray the overall SDC probability. Evaluation shows that the SDC probability derived from Druto’s input generation is as much as 74× higher than that from existing techniques. Moreover, existing techniques cannot find our generated inputs even given 5× more search time. Md Hasanur Rahman 0001, Sheng Di, Shengjian Guo, Xiaoyi Lu 0001, Guanpeng Li, Franck Cappello |
IPDPS | 1 |
| 2024 | Investigating the impact of transient hardware faults on deep learning neural network inferenceabstractSummary Safety‐critical applications, such as autonomous vehicles, healthcare, and space applications, have witnessed widespread deployment of deep neural networks (DNNs). Inherent algorithmic inaccuracies have consistently been a prevalent cause of misclassifications, even in modern DNNs. Simultaneously, with an ongoing effort to minimize the footprint of contemporary chip design, there is a continual rise in the likelihood of transient hardware faults in deployed DNN models. Consequently, researchers have wondered the extent to which these faults contribute to DNN misclassifications compared to algorithmic inaccuracies. This article delves into the impact of DNN misclassifications caused by transient hardware faults and intrinsic algorithmic inaccuracies in safety‐critical applications. Initially, we enhance a cutting‐edge fault injector,TensorFI, for TensorFlow applications to facilitate fault injections on modern DNN non‐sequential models in a scalable manner. Subsequently, we analyse the DNN‐inferred outcomes based on our defined safety‐critical metrics. Finally, we conduct extensive fault injection experiments and a comprehensive analysis to achieve the following objectives: (1) investigate the impact of different target class groupings on DNN failures and (2) pinpoint the most vulnerable bit locations within tensors, as well as DNN layers accountable for the majority of safety‐critical misclassifications. Our findings regarding different grouping formations reveal that failures induced by transient hardware faults can have a substantially greater impact (with a probability up to 4 higher) on safety‐critical applications compared to those resulting from algorithmic inaccuracies. Additionally, our investigation demonstrates that higher order bit positions in tensors, as well as initial and final layers of DNNs, necessitate prioritized protection compared to other regions. Md Hasanur Rahman 0001, Sabuj Laskar, Guanpeng Li |
Softw. Test. Verification Reliab. | 1 |
| 2023 | A Feature-Driven Fixed-Ratio Lossy Compression Framework for Real-World Scientific DatasetsabstractToday’s scientific applications and advanced instruments are producing extremely large volumes of data everyday, so that error-controlled lossy compression has become a critical technique to the scientific data storage and management. Existing lossy scientific data compressors, however, are designed mainly based on error-control driven mechanism, which cannot be efficiently applied in the fixed-ratio use-case, where a desired compression ratio needs to be reached because of the restricted data processing/management resources such as limited memory/storage capacity and network bandwidth. To address this gap, we propose a low-cost compressor-agnostic feature-driven fixed-ratio lossy compression framework (FXRZ). The key contributions are three-fold. (1) We perform an in-depth analysis of the correlation between diverse data features and compression ratios based on a wide range of application datasets, which is a fundamental work for our framework. (2) We propose a series of optimization strategies that can enable the framework to reach a fairly high accuracy in identifying the expected error configuration with very low computational cost. (3) We comprehensively evaluate our framework using 4 state-of-the-art error-controlled lossy compressors on 10 different snapshots and simulation configuration-based real-world scientific datasets from 4 different applications across different domains. Our experiment shows that FXRZ outperforms the state-of-the-art related work by 108×. The experiments with 4,096 cores on a supercomputer show a performance gain of 1.18∼8.71× than the related work in overall parallel data dumping. Md Hasanur Rahman 0001, Sheng Di, Kai Zhao 0008, Robert Underwood, Guanpeng Li, Franck Cappello |
ICDE | 1 |
| 2022 | Characterizing Deep Learning Neural Network Failures Between Algorithmic Inaccuracy and Transient Hardware FaultsabstractDeep Neural Networks (DNNs) have been widely deployed in safety-critical applications such as autonomous vehicles, healthcare, and space applications. Though DNN models have long suffered intrinsic algorithmic inaccuracies, the increasing number of hardware transient faults in computer systems has been raising safety and reliability concerns in safety-critical applications. This paper investigates the impact of DNN misclassifications that caused by hardware transient faults and intrinsic algorithmic inaccuracy in safety-critical applications. We first extend a state-of-the-art fault injector for TensorFlow application, TensorFI, to support fault injections on modern DNN models in a scalable way, then characterize the outcome classes of the models, analyzing them based on safety related metrics. Finally, we conduct a large-scale fault injection experiment to measure the failures according to the metrics and study their impact on safety. We observe that failures caused by hardware transient faults could have much more significant impact (up to 4 times higher probability) on safety-critical applications than that of the DNN algorithmic inaccuracies, advocating the potential needs to protect DNNs from hardware faults in safety-critical applications. Sabuj Laskar, Md Hasanur Rahman 0001, Guanpeng Li |
PRDC | 2 |
| 2021 | PEPPA-X: finding program test inputs to bound silent data corruption vulnerability in HPC applicationsabstractTransient hardware faults have become prevalent due to the shrinking size of transistors, leading to silent data corruptions (SDCs). Therefore, HPC applications need to be evaluated (e.g., via fault injections) and protected to meet the reliability target. In the evaluation, the target programs exercise with a set of given inputs which are usually from program benchmark suite. However, these inputs rarely manifest the SDC vulnerabilities, leading to over-optimistic assessment and unexpectedly higher failure rates in production. We propose Peppa-X, which efficiently identifies the test inputs that estimate the bound of program SDC resiliency. Our key insight is that the SDC sensitivity distribution in a program often remains stationary across input space. Thereby, we can guide the search of SDC-bound inputs by a sampled distribution. Our evaluation shows that Peppa-X can identify the SDC-bound input of a program that existing methods cannot find even with 5x more search time. Md Hasanur Rahman 0001, Aabid Shamji, Shengjian Guo, Guanpeng Li |
SC | 1 |