EDBT 2026 Demo / reviewers in the wild / expert
Guanpeng Li
dblp:151/4108
· DBLP profile ↗
55ranked-venue papers
7as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 5 first-author · 26 since 2021Security and privacy · 12 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAGMA: A Multi-Graph based Agentic Memory Architecture for AI AgentsabstractMemory-Augmented Generation (MAG) extends Large Language Models with external memory to support long-context reasoning, but existing approaches largely rely on semantic similarity over monolithic memory stores, entangling temporal, causal, and entity information.This design limits interpretability and alignment between query intent and retrieved evidence, leading to suboptimal reasoning accuracy.In this paper, we propose MAGMA, a multi-graph agentic memory architecture that represents each memory item across orthogonal semantic, temporal, causal, and entity graphs.MAGMA formulates retrieval as policy-guided traversal over these relational views, enabling query-adaptive selection and structured context construction.By decoupling memory representation from retrieval logic, MAGMA provides transparent reasoning paths and fine-grained control over retrieval.Experiments on LoCoMo and LongMemEval demonstrate that MAGMA consistently outperforms state-of-the-art agentic memory systems in long-horizon reasoning tasks. Dongming Jiang, Guanpeng Li, Bingzhe Li |
ACL (1) | 3 |
| 2026 | Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model InferenceabstractLarge language models (LLMs) are increasingly integrated into high-performance computing (HPC) workflows, accelerating scientific discovery through diverse perspectives such as code generation and domain-specific decision-making. Yet, how soft errors propagate and affect LLM inference remains largely unexplored. To bridge this gap, we present a comprehensive study on error propagation in LLM inference, enabled by our proposed LLMFI, a configurable and deterministic fault-injection framework. Using LLMFI, we systematically inject faults across three open-weighted LLMs and thirteen representative tasks, covering reasoning, multilingual, mathematical, and coding domains. In addition, we conduct fine-grained case studies that reveal critical vulnerability patterns. Overall, our study yields 17 takeaways that advance the understanding of error propagation in LLM inference and introduces four low-overhead directions to improve reliability through software-only modification, offering practical guidance for future error detection and mitigation. Yafan Huang, Sheng Di, Guanpeng Li |
ICS | 3 |
| 2026 | GPZ: GPU-Accelerated Lossy Compressor for Particle DataabstractParticle-based simulations and point-cloud applications generate massive, irregular datasets that challenge storage, I/O, and real-time analytics. Traditional compression techniques struggle with irregular particle distributions and GPU architectural constraints, often resulting in limited throughput and suboptimal compression ratios. In this paper, we present GPZ, a high-performance, error-bounded lossy compressor designed specifically for large-scale particle data on modern GPUs. GPZ employs a novel four-stage parallel pipeline that synergistically balances high compression efficiency with the architectural demands of massively parallel hardware. We introduce a suite of targeted optimizations for computation, memory access, and GPU occupancy that enable GPZ to achieve near-hardware-limit throughput. We conduct an extensive evaluation on three distinct GPU architectures (workstation, data center, and edge) using six large-scale, real-world scientific datasets from four distinct domains. The results demonstrate that GPZ consistently and significantly outperforms four state-of-the-art GPU compressors, delivering up to 8x higher end-to-end throughput while achieving superior compression ratios and data quality. Yafan Huang, Zhuoxun Yang, Sheng Di, Boyuan Zhang 0002, Jiajun Huang 0001, Jinyang Liu 0003, Jiannan Tian, Guanpeng Li, Fengguang Song, Hanqi Guo 0001, Franck Cappello, Kai Zhao 0008 |
ICS | 10 |
| 2026 | PACER: A Userspace Network Rate Controller in MPI with Adaptive Compression for Parallel Applications
Yuke Li 0003, Darren Ng, Arjun Kashyap, Sheng Di, Guanpeng Li, Xiaoyi Lu 0001 |
ICS | 5 |
| 2026 | Feed-Forward Controller-Based Recovery for Robotic Vehicles From Physical AttacksabstractRobotic Vehicles (RV) rely extensively on sensor inputs to operate autonomously. Physical attacks such as sensor tampering and spoofing can feed erroneous sensor measurements to deviate RVs from their course and result in mission failures. In this paper, we present a Feed-Forward Controller based framework for automatically recovering RVs from physical attacks. We use machine learning (ML) to design an attack resilient Feed-Forward Controller (FFC), which runs in tandem with the RV's primary controller and monitors it. Under attacks, the FFC takes over from the RV's primary controller to recover the RV, and allows the RV to complete its mission successfully. Our evaluation on 6 RV systems including 3 real RVs shows that our proposed framework prevents crashes and allows RVs to complete their missions successfully despite attacks in 86% of the cases. Further, we propose designs to streamline the implementation of the FFC-based recovery and its application in new RV systems. Pritam Dash, Guanpeng Li, Zitao Chen 0001, Mehdi Karimibiuki, Karthik Pattabiraman |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Understanding Error Sensitivity in Checkpointing for Linear System SolversabstractFault tolerance in large-scale iterative solvers is critical, yet traditional checkpointing methods often impose significant storage overhead. In this study, we evaluate the compression error introduced from lossy compression in the Conjugate Gradient method. We systematically investigate how various compressor configurations such as numerical error bounds, error modes, and prediction algorithms influence compressed checkpoint size, compression error, and extra iterations after recovery. Our analysis reveals key trade-offs between storage efficiency and recovery overhead. Yafan Huang, Guanpeng Li |
HPDC | 4 |
| 2025 | Pushing the Limits of GPU Lossy Compression: A Hierarchical Delta Approach
Boyuan Zhang 0002, Yafan Huang, Sheng Di, Fengguang Song, Guanpeng Li, Franck Cappello |
ICS | 5 |
| 2025 | FedDES: Discrete Event Based Performance Simulation for Federated Learning SystemsabstractFederated Learning (FL) is a scalable and privacy-preserving paradigm well-suited for edge computing. Real-world FL deployments face substantial systems challenges such as compute variability and communication delays, motivating researchers to leverage simulation before real deployment. Most existing FL simulators, however, struggle to scale efficiently and incur long runtimes even for small workloads. To address this, we present FedDES, a high-fidelity, framework-agnostic discrete-event simulation platform that accurately models the runtime behavior of FL systems, including client training, communication overhead, network dynamics, and aggregation strategies. FedDES supports flexible configurations and diverse aggregation approaches, achieving simulation error within 2% of real deployments and delivering over 1000× speedup compared to prior tools. Large-scale experiments with up to 131,072 clients further show that the aggregation strategy critically affects performance, especially under heterogeneous and variable network conditions typical of edge environments. Zhonghao Chen, Weicong Chen 0002, Kibaek Kim, Guanpeng Li, Sheng Di, Xiaoyi Lu 0001 |
SEC | 5 |
| 2025 | Modeling Rate-Distortion for Endpoint-Aware Lossy Compression in Scientific Data TransferabstractHigh-performance computing (HPC) systems generate massive scientific datasets, often stored in remote data repositories. Limited bandwidth and resource-constrained endpoints pose challenges for efficient large data transfer. Error-bounded lossy compression addresses this by reducing data sizes (higher bit-rates) while controlling distortion. However, different compressors exhibit distinct rate-distortion behaviors even under the same error bounds. Thus, selecting an optimal compressor before transfer is essential to meet endpoint-specific requirements e.g., maximizing data reduction at a fixed distortion. Existing trial-and-error approaches require multiple costly full-scale compression runs to reach at target requirements, making them impractical for such online use. To address this, we propose OptRD, a compressor-agnostic framework that efficiently models rate-distortion trade-offs across multiple lossy compressors by analyzing spatial data traits at reduced resolutions. Evaluated using 3 state-of-the-art lossy compressors on 30 scientific datasets from 4 HPC applications, OPTRD incurs only$\sim 5 \%$average estimation error and achieves over$100 \times$runtime speedup compared to trial-and-error methods, significantly improving optimal compressor selection during such data transfer use cases. Md Hasanur Rahman 0001, Sheng Di, Guanpeng Li, Franck Cappello |
IPCCC | 3 |
| 2025 | You Only Spectralize Once: Taking a Spectral Detour to Accelerate Graph Neural NetworkabstractTraining Graph Neural Networks (GNNs) often relies on repeated, irregular, and expensive message-passing operations over all nodes (e.g., $N$), leading to high computational overhead. To alleviate this inefficiency, we revisit the GNNs training from a spectral perspective. In many real-world graphs, node features and embeddings exhibit sparse representation in the Graph Fourier domain. This inherent spectral sparsity aligns well with the principles of Compressed Sensing, which posits that signals sparse in one transform domain can be accurately reconstructed from a significantly reduced number of measurements. This observation motivates the design of a more efficient GNNs that operates predominantly in compressed spectral subspace. Thus, we propose You Only Spectralize Once (YOSO), a GNN training scheme that performs single Graph Fourier Transformation to project features onto a learnable orthonormal Fourier basis, retaining only $M$ spectral coefficients ($M \ll N$). The entire GNN computation is then carried out in reduced spectral domain. Final full-graph embeddings are recovered only at output layer by solving a bounded $\ell_{2,1}$-regularized optimization problem. Theoretically, drawing upon Compressed Sensing theory, we prove stable recovery throughout training by showing that the projection onto our learnable Fourier basis can satisfy the Restricted Isometry Property when $M=\mathcal{O}(k \log N)$ for $k$-row-sparse spectra, acting as the measurement process. Empirically, YOSO achieves an average 74\% reduction in training time across five benchmark datasets compared to state-of-the-art methods, while maintaining competitive accuracy. Zhichun Guo, Guanpeng Li, Bingzhe Li |
NeurIPS | 3 |
| 2025 | Deploying Lightweight Input-Aware Selective Instruction Duplication in HPC ApplicationsabstractModern high-performance computing (HPC) applications are increasingly vulnerable to silent data corruptions (SDCs) caused by transient hardware faults. While selective instruction duplication (SID) offers an efficient software-level protection strategy, existing SID methods rely on SDC vulnerability profile derived from only the default reference input often found in application suites. However, they overlook the input-dependent nature of SDC propagation. This leads to significant SDC coverage loss when inputs vary. We present Protego, a novel input-aware SID protection framework that efficiently adapts protection to runtime inputs. Protego performs a one-time vulnerability-guided input exploration to identify a small number of input groups with distinct SID protection patterns. At runtime, Protego uses lightweight features derived from input arguments to select and deploy the appropriate SID protection. Our evaluation across 10 HPC applications demonstrates the effectiveness and efficiency of Protego in mitigating SDC coverage loss across diverse inputs, compared to existing SID techniques. Md Hasanur Rahman 0001, Guanpeng Li |
SC | 2 |
| 2025 | GPU Lossy Compression for HPC Can Be Versatile and Ultra-FastabstractThis work proposes VGC, a versatile and ultra-fast GPU lossy compression framework designed to address the growing data challenges in high-performance computing (HPC). VGC captures dimension information in scientific data and supports three compression algorithms, achieving high compression ratios across diverse HPC domains. Built with a highly optimized GPU kernel, VGC delivers state-of-the-art throughput with error control. In addition to compression ratio and speed, VGC supports two distinctive modes that enhance its versatility. Memory-efficient Compression uses a kernel fission design to compute compressed size, allocate only the required GPU memory, and compress data without waste, effectively reducing memory footprint. Selective Decompression introduces an early stopping mechanism that enables direct access to regions of interest without decompressing the entire dataset. Yafan Huang, Sheng Di, Guanpeng Li, Franck Cappello |
SC | 3 |
| 2025 | lsCOMP: Efficient Light Source CompressionabstractLight source facilities, which generate X-rays for probing microstructures and dynamic processes, produce intense data streams, reaching up to 250 GB/s and projected to exceed 1 TB/s by the end of this decade. Managing such massive data poses critical challenges due to limited local processing capacity and bandwidth constraints when offloading data to HPC systems. To address these challenges, we propose lsCOMP, a GPU compressor that operates within a single kernel. lsCOMP supports both lossless and configurable lossy compression, ensuring high compression ratios and preserved data quality across diverse light source applications. On a single NVIDIA A100 GPU, lsCOMP achieves compression throughputs of 380.89 to 509.21 GB/s in lossless mode, delivering up to 20 times higher performance than industry-leading GPU compressors while achieving superior compression ratios. In lossy modes, lsCOMP further improves throughput and ratios significantly. Additionally, lsCOMP demonstrates versatile performance across various integer datasets and supports TB/s-level random access throughput. Yafan Huang, Sheng Di, Robert Underwood, Peco Myint, Miaoqi Chu, Guanpeng Li, Nicholas Schwarz, Franck Cappello |
SC | 6 |
| 2025 | GEREM: Fast and Precise Error Resilience Assessment for GPU MicroarchitecturesabstractGPUs are widely used hardware acceleration platforms in many areas due to their great computational throughput. In the meanwhile, GPUs are vulnerable to transient hardware faults in the post-Moore era. Analyzing the error resilience of GPUs are critical for both hardware and software. Statistical fault injection approaches are commonly used for error resilience analysis, which are highly accurate but very time consuming. In this work, we propose GEREM, a first framework to speed up fault injection process so as to estimate the error resilience of GPU microarchitectures swiftly and precisely. We find early fault behaviors can be used to accurately predict the final outcomes of program execution. Based on this observation, we categorize the early behaviors of hardware faults into GPU Early Fault Manifestation models (EFMs). For data structures, EFMs are early propagation characteristics of faults, while for pipeline instructions, EFMs are heuristic properties of several instruction contexts. We further observe that EFMs are determined by static microarchitecture states, so we can capture them without actually simulating the program execution process under fault injections. Leveraging these observations, our GEREM framework first profiles the microarchitectural states related for EFMs at one time. It then injects faults into the profiled traces to immediately generate EFMs. For data storage structures, EFMs are directly used to predict final fault outcomes, while for pipeline instructions, machine learning is used for prediction. Evaluation results show GEREM precisely assesses the error resilience of GPU microarchitecture structures with$237\times$speedup on average comparing with traditional fault injections. Jingweijia Tan, An Zhong, Kaige Yan, Xiaohui Wei 0002, Guanpeng Li |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy CompressionabstractCross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. To bridge this gap, we proposeFedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show thatFedEFszimproves the fairness across different benchmarks by up to$60.88\%$and meanwhile reduces the communication traffic by up to$315\times$. Sheng Di, Benben Liu, Zhuoran Ji, Guanpeng Li, Xiaoyi Lu 0001, Amelie Chi Zhou, Khalid Ayedh Alharthi, Jiannong Cao 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Diagnosis-guided Attack Recovery for Securing Robotic Vehicles from Sensor Deception AttacksabstractSensors are crucial for perception and autonomous operation in robotic vehicles (RV). Unfortunately, RV sensors can be compromised by physical attacks such as sensor tampering or spoofing. In this paper, we present DeLorean, a unified framework for attack detection, attack diagnosis, and recovering RVs from sensor deception attacks (SDA). DeLorean can recover RVs even from strong SDAs in which the adversary targets multiple heterogeneous sensors simultaneously. We propose a novel attack diagnosis technique that inspects the attack-induced errors under SDAs, and identifies the targeted sensors using causal analysis. DeLorean then uses historic state information to selectively reconstruct physical states for compromised sensors, enabling targeted attack recovery under single or multi-sensor SDAs. We evaluate DeLorean on four real and two simulated RVs under SDAs targeting various sensors, and we find that it successfully recovers RVs from SDAs in 93% of the cases. Pritam Dash, Guanpeng Li, Mehdi Karimibiuki, Karthik Pattabiraman |
AsiaCCS | 2 |
| 2024 | A Fast Low-Level Error Detection Technique
Zhengyang He, Hui Xu 0009, Guanpeng Li |
DSN | 3 |
| 2024 | Druto: Upper-Bounding Silent Data Corruption Vulnerability in GPU ApplicationsabstractDue to the increasing scale of high-performance computing (HPC) systems, transient hardware faults have become a major reliability concern. Consequently, Silent Data Corruptions (SDCs) due to these faults have been a common insidious consequence in GPU applications. Developers often measure the application resilience with a set of program test inputs available in the benchmark suite, assuming the resilience would not fluctuate much among different inputs. However, we observe that this assumption often results in an over-optimistic evaluation for GPU applications. As a result, the subsequent SDC protection following the evaluation can hardly meet the expected reliability bar in the production environment, where applications would run with potentially arbitrary input values. To this end, we propose Druto – a compiler-based automated technique that searches for inputs to incrementally approach the upper bound of a GPU application’s SDC probability. We develop Druto based on the property that the resilience profiles of a small group of representative threads in a GPU kernel can approximately rank various inputs in terms of the overall SDC probability. Therefore, Druto strategically steers the search towards new program inputs that efficiently portray the overall SDC probability. Evaluation shows that the SDC probability derived from Druto’s input generation is as much as 74× higher than that from existing techniques. Moreover, existing techniques cannot find our generated inputs even given 5× more search time. Md Hasanur Rahman 0001, Sheng Di, Shengjian Guo, Xiaoyi Lu 0001, Guanpeng Li, Franck Cappello |
IPDPS | 5 |
| 2024 | NVMe-oPF: Designing Efficient Priority Schemes for NVMe-over-Fabrics with Multi-Tenancy SupportabstractResource disaggregation is prevalent in datacenters since it provides high resource utilization when compared to servers dedicated to either compute, memory, or storage. NVMe-over-Fabrics (NVMe-oF) is the standardized protocol for accessing disaggregated network storage. Currently, the NVMe-oF specification lacks semantics to prioritize I/O requests based on different application needs. Since applications have varying goals — latency-sensitive or throughput-critical I/O — we need to design efficient schemes to allow applications to specify the type of performance they wish to achieve. To this end, we propose a new NVMe-over-Priority-Fabrics (NVMe-oPF) protocol with multi-tenancy support that allows applications to specify whether to optimize for latency or throughput. NVMe-oPF proposes coalescing request completions, lock-free optimization, zero-copy queues, out-of-order request completion handling, and window size optimization for the specific I/O patterns, queue depths, and I/O sizes that yield the best performance. Our NVMe-oPF-10Gbps can achieve up to 2.94X improvement in throughput and reduces tail latency by up to 32.1% for highly concurrent multi-tenant read workloads when compared to the state-of-the-art userspace NVMe-oF runtime design in Intel Storage Performance Development Kit (SPDK). For write workloads with 100Gbps, NVMe-oPF achieves a 32.6% increase in throughput while maintaining low latency compared to SPDK. We also bring performance benefits to the application level with HDF5 by increasing write workload throughput by 25.2% in larger-scale experiments. Darren Ng, Andrew Lin, Arjun Kashyap, Guanpeng Li, Xiaoyi Lu 0001 |
IPDPS | 4 |
| 2024 | cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationabstractError-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. Compared to CPU-based compressors, GPU-based compressors exhibit substantially higher throughputs, fitting better for today’s HPC applications. However, the critical limitations of existing GPU-based compressors are their low compression ratios and qualities, severely restricting their applicability. To overcome these, we introduce a new GPU-based error-bounded scientific lossy compressor named CUSZ-i, with the following contributions: (1) A novel GPU-optimized interpolation-based prediction method significantly improves the compression ratio and decompression data quality. (2) The Huffman encoding module in CUSZ-i is optimized for better efficiency. (3) CUSZ-i is the first to integrate the NVIDIA Bitcomp-lossless as an additional compression-ratio-enhancing module. Evaluations show that CUSZ-i significantly outperforms other latest GPU-based lossy compressors in compression ratio under the same error bound (hence, the desired quality), showcasing a 476% advantage over the second-best. This leads to CUSZ-i’s optimized performance in several real-world use cases. Jinyang Liu 0003, Jiannan Tian, Shixun Wu, Sheng Di, Boyuan Zhang 0002, Robert Underwood, Yafan Huang, Jiajun Huang 0001, Kai Zhao 0008, Guanpeng Li, Dingwen Tao, Zizhong Chen, Franck Cappello |
SC | 10 |
| 2024 | cuSZp2: A GPU Lossy Compressor with Extreme Throughput and Optimized Compression RatioabstractExisting GPU lossy compressors suffer from expensive data movement overheads, inefficient memory access patterns, and high synchronization latency, resulting in limited throughput. This work proposes cuSZP2, a generic single-kernel error-bounded lossy compressor purely on GPUs designed for applications that require high speed, such as large-scale GPU simulation and large language model training. In particular, CUSZP2 proposes a novel lossless encoding method, optimizes memory access patterns, and hides synchronization latency, achieving extreme end-to-end throughput and optimized compression ratio. Experiments on NVIDIA A100 GPU with 9 real-world HPC datasets demonstrate that, even with higher compression ratios and data quality, CUSZP2 can deliver on average 332.42 and $513.04 \mathrm{~GB} / \mathrm{s}$ end-to-end throughput for compression and decompression, respectively, which is around $2 \times$ of existing pure-GPU compressors and $200 \times$ of CPU-GPU hybrid compressors. Yafan Huang, Sheng Di, Guanpeng Li, Franck Cappello |
SC | 3 |
| 2024 | Versatile Datapath Soft Error Detection on the Cheap for HPC ApplicationsabstractWith the ongoing reduction in technology sizes and voltage levels, modern microprocessors are increasingly susceptible to soft errors, corrupting datapath units during program execution. While these error types have received considerable attention recently, existing solutions either confine themselves to limited scopes or incur massive overheads in performance and power consumption, hindering practical usage. In this work, we propose CONDA, a novel error detection technique based on code transformation and static program analysis, achieving versatile datapath protection at low cost. At compile time, ConDa analyzes program characteristics and transforms the original program code without complicating its control-flow and memory access patterns. At runtime, ConDa detects datapath errors with low overhead and latency. The evaluation of 38 benchmarks and a parallel HPC simulation reveals that CONDA only incurs 57.79% runtime overhead, which is 41.84% faster than existing state-of-the-art, with the same level of error detection effectiveness and low detection latency. Yafan Huang, Sheng Di, Xiaoyi Lu 0001, Guanpeng Li |
SC | 5 |
| 2024 | Investigating the impact of transient hardware faults on deep learning neural network inferenceabstractSummary Safety‐critical applications, such as autonomous vehicles, healthcare, and space applications, have witnessed widespread deployment of deep neural networks (DNNs). Inherent algorithmic inaccuracies have consistently been a prevalent cause of misclassifications, even in modern DNNs. Simultaneously, with an ongoing effort to minimize the footprint of contemporary chip design, there is a continual rise in the likelihood of transient hardware faults in deployed DNN models. Consequently, researchers have wondered the extent to which these faults contribute to DNN misclassifications compared to algorithmic inaccuracies. This article delves into the impact of DNN misclassifications caused by transient hardware faults and intrinsic algorithmic inaccuracies in safety‐critical applications. Initially, we enhance a cutting‐edge fault injector,TensorFI, for TensorFlow applications to facilitate fault injections on modern DNN non‐sequential models in a scalable manner. Subsequently, we analyse the DNN‐inferred outcomes based on our defined safety‐critical metrics. Finally, we conduct extensive fault injection experiments and a comprehensive analysis to achieve the following objectives: (1) investigate the impact of different target class groupings on DNN failures and (2) pinpoint the most vulnerable bit locations within tensors, as well as DNN layers accountable for the majority of safety‐critical misclassifications. Our findings regarding different grouping formations reveal that failures induced by transient hardware faults can have a substantially greater impact (with a probability up to 4 higher) on safety‐critical applications compared to those resulting from algorithmic inaccuracies. Additionally, our investigation demonstrates that higher order bit positions in tensors, as well as initial and final layers of DNNs, necessitate prioritized protection compared to other regions. Md Hasanur Rahman 0001, Sabuj Laskar, Guanpeng Li |
Softw. Test. Verification Reliab. | 3 |
| 2023 | Towards Improving Reverse Time Migration Performance by High-speed Lossy CompressionabstractSeismic imaging is an exploration method for estimating the seismic characteristics of the earth's sub-surface for geologists and geophysicists. Reverse time migration (RTM) is a critical method in seismic imaging analysis. It can produce huge volumes of data that need to be stored for later use during its execution. The traditional solution transfers the vast amount of data to peripheral devices and loads them back to memory whenever needed, which may cause a substantial burden to I/O and storage space. As such, an efficient data compressor turns out to be a very critical solution. In order to get the best overall RTM analysis performance, we develop a novel hybrid lossy compression method (called HyZ), which is not only fairly fast in both compression and decompression but also has a good compression ratio with satisfactory reconstructed data quality for post hoc analysis. We evaluate several state-of-the-art error-controlled lossy compression algorithms (including HyZ, BR, SZx, SZ, SZ-Interp, ZFP, etc.) in a supercomputer. Experiments show that HyZ not only significantly improves the overall performance for RTM by 6.29∼6.60× but also obtains fairly good qualities for both RTM single snapshots and the final stacking image. Yafan Huang, Kai Zhao 0008, Sheng Di, Guanpeng Li, Maxim Dmitriev, Thierry-Laurent D. Tonellot, Franck Cappello |
CCGrid | 4 |
| 2023 | A Feature-Driven Fixed-Ratio Lossy Compression Framework for Real-World Scientific DatasetsabstractToday’s scientific applications and advanced instruments are producing extremely large volumes of data everyday, so that error-controlled lossy compression has become a critical technique to the scientific data storage and management. Existing lossy scientific data compressors, however, are designed mainly based on error-control driven mechanism, which cannot be efficiently applied in the fixed-ratio use-case, where a desired compression ratio needs to be reached because of the restricted data processing/management resources such as limited memory/storage capacity and network bandwidth. To address this gap, we propose a low-cost compressor-agnostic feature-driven fixed-ratio lossy compression framework (FXRZ). The key contributions are three-fold. (1) We perform an in-depth analysis of the correlation between diverse data features and compression ratios based on a wide range of application datasets, which is a fundamental work for our framework. (2) We propose a series of optimization strategies that can enable the framework to reach a fairly high accuracy in identifying the expected error configuration with very low computational cost. (3) We comprehensively evaluate our framework using 4 state-of-the-art error-controlled lossy compressors on 10 different snapshots and simulation configuration-based real-world scientific datasets from 4 different applications across different domains. Our experiment shows that FXRZ outperforms the state-of-the-art related work by 108×. The experiments with 4,096 cores on a supercomputer show a performance gain of 1.18∼8.71× than the related work in overall parallel data dumping. Md Hasanur Rahman 0001, Sheng Di, Kai Zhao 0008, Robert Underwood, Guanpeng Li, Franck Cappello |
ICDE | 5 |
| 2023 | Characterizing Runtime Performance Variation in Error Detection by Duplicating InstructionsabstractSoft error rate has been increasing due to the shrinking size of transistors, leading to an elevated risk of catastrophic failures in modern computer systems. Error detection by duplicating instructions (EDDI) is a software-based technique to mitigate soft errors with a low runtime performance overhead and has been widely adopted in many safety- and mission-critical real-time systems such as space applications. However, these systems are commonly sensitive to runtime performance overheads the protection techniques incur. Few studies have investigated the performance of EDDI across various system designs and operational parameters, hence lacking a complete understanding in the literature. In this paper, we conduct comprehensive experiments to study the variation of EDDI runtime performance overhead and characterize the root causes. We find that there exist significant variations in performance overheads of EDDI, due to a few architectural and program-level factors. Based on the findings, we propose two practical techniques FuzzyB and Celer: FuzzyB uses an input searching technique to bound EDDI runtime performance overhead across different inputs for a given program; while Celer reduces EDDI run-time performance overheads using compiler transformation (by 25.08% reduction). Yafan Huang, Zhengyang He, Lingda Li, Guanpeng Li |
ISSRE | 4 |
| 2023 | Demystifying and Mitigating Cross-Layer Deficiencies of Soft Error Protection in Instruction DuplicationabstractSoft errors are prevalent in modern High-Performance Computing (HPC) systems, resulting in silent data corruptions (SDCs), compromising system reliability. Instruction duplication is a widely used software-based protection technique against SDCs. Existing instruction duplication techniques are mostly implemented at LLVM level and may suffer from low SDC coverage at assembly level. In this paper, we evaluate instruction duplication at both LLVM and assembly levels. Our study shows that existing instruction duplication techniques have protection deficiency at assembly level and are usually over-optimistic in the protection. We investigate the root-causes of the protection deficiency and propose a mitigation technique, Flowery, to solve the problem. Our evaluation shows that Flowery can effectively protect programs from SDCs evaluated at assembly level. Zhengyang He, Yafan Huang, Hui Xu 0009, Dingwen Tao, Guanpeng Li |
SC | 5 |
| 2023 | cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceabstractModern scientific applications and supercomputing systems are generating large amounts of data in various fields, leading to critical challenges in data storage footprints and communication times. To address this issue, error-bounded GPU lossy compression has been widely adopted, since it can reduce the volume of data within a customized threshold on data distortion. In this work, we propose an ultra-fast error-bounded GPU lossy compressor cuSZp. Specifically, cuSZp computes the linear recurrences with hierarchical parallelism to fuse the massive computation into one kernel, drastically improving the end-to-end throughput. In addition, cuSZp adopts a block-wise design along with a lightweight fixed-length encoding and bit-shuffle inside each block such that it achieves high compression ratios and data quality. Our experiments on NVIDIA A100 GPU with 6 representative scientific datasets demonstrate that cuSZp can achieve an ultra-fast end-to-end throughput (95.53x compared with cuSZ) along with a high compression ratio and high reconstructed data quality. Yafan Huang, Sheng Di, Xiaodong Yu 0001, Guanpeng Li, Franck Cappello |
SC | 4 |
| 2023 | Fault Injection for TensorFlow ApplicationsabstractAs machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application's resilience. In this work, we present TensorFI 1 and TensorFI 2, high-level fault injection (FI) frameworks for TensorFlow-based applications. TensorFI 1 and 2 are able to inject both hardware and software faults in any general TensorFlow 1 and 2 program respectively. Both are configurable FI tools that are flexible, easy to use, and portable. They can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., bit-flips in particular operations or layers). We use TensorFI 1 and TensorFI 2 to evaluate the resilience of 11 and 10 ML programs respectively, all written in TensorFlow, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also measure the performance overheads of the two injectors, and present 4 case studies, two for each tool, to demonstrate their utility. Niranjhana Narayanan, Zitao Chen 0001, Bo Fang 0002, Guanpeng Li, Karthik Pattabiraman, Nathan DeBardeleben |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | Hardening selective protection across multiple program inputs for HPC applicationsabstractWith the ever-shrinking size of transistors and increasing scale of applications, silent data corruptions (SDCs) have become a common yet serious issue in HPC applications. Selective instruction duplication (SID) is a popular fault-tolerance technique that can obtain a high SDC coverage with low-performance overhead, as it selects the most vulnerable parts of a program for protection with priority. However, existing studies of SID are confined to single program input in the evaluation, assuming that the error resilience of the program remains similar across inputs, leading to a drastic loss of SDC coverage from SID when the protected program runs different inputs. Hence, we proposed Sentinel, an automated compiler-based framework to mitigate the loss of SDC coverage. Evaluation results show that Sentinel can effectively mitigate the loss of SDC coverage (up to 97.00%) across multiple inputs, which significantly hardens existing SID techniques. Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, Franck Cappello |
PPoPP | 4 |
| 2022 | Characterizing Deep Learning Neural Network Failures Between Algorithmic Inaccuracy and Transient Hardware FaultsabstractDeep Neural Networks (DNNs) have been widely deployed in safety-critical applications such as autonomous vehicles, healthcare, and space applications. Though DNN models have long suffered intrinsic algorithmic inaccuracies, the increasing number of hardware transient faults in computer systems has been raising safety and reliability concerns in safety-critical applications. This paper investigates the impact of DNN misclassifications that caused by hardware transient faults and intrinsic algorithmic inaccuracy in safety-critical applications. We first extend a state-of-the-art fault injector for TensorFlow application, TensorFI, to support fault injections on modern DNN models in a scalable way, then characterize the outcome classes of the models, analyzing them based on safety related metrics. Finally, we conduct a large-scale fault injection experiment to measure the failures according to the metrics and study their impact on safety. We observe that failures caused by hardware transient faults could have much more significant impact (up to 4 times higher probability) on safety-critical applications than that of the DNN algorithmic inaccuracies, advocating the potential needs to protect DNNs from hardware faults in safety-critical applications. Sabuj Laskar, Md Hasanur Rahman 0001, Guanpeng Li |
PRDC | 4 |
| 2022 | Salus: A Novel Data-Driven Monitor that Enables Real-Time Safety in Autonomous Driving SystemsabstractThis paper proposes Salus, a data-driven real-time safety monitor, that detects and mitigates safety violations of an autonomous vehicle (AV). The key insight is that traffic situations that lead to AV safety violations fall into patterns and can be identified by learning from the safety violations of the AV. Our approach is to use machine learning (ML) techniques to model the traffic behaviors that result in safety violations in the AV, characterize their early symptoms for training a preemptive model, hence deploy and detect real-time safety violations before the actual crashes happen to the AV. In order to train our ML model, we leverage a pipeline of fuzzing techniques to tailor AV-specific safety violation symptoms and generate the training data via data argumentation techniques. Our evaluation demonstrates our proposed technique is effective in reducing over 97.2% of safety violations in industry-level autonomous driving systems, such as Baidu Apollo, with no more than 0.018 false positive values. Yafan Huang, Guanpeng Li |
QRS | 3 |
| 2022 | Mitigating Silent Data Corruptions in HPC Applications across Multiple Program InputsabstractWith the ever-shrinking size of transistors, silent data corruptions (SDCs) are becoming a common yet serious issue in HPC. Selective instruction duplication (SID) is a widely used fault-tolerance technique that can obtain high SDC coverage with low performance overhead. However, existing SID methods are confined to single program input in its assessment, assuming that error resilience of a program remains similar across inputs. Nevertheless, we observe that the assumption cannot always hold, leading to a drastic loss in SDC coverage across different inputs, compromising HPC reliability. We notice that the SDC coverage loss correlates with a small set of instructions - we call them incubative instructions, which reveal elusive error propagation characteristics across multiple inputs. We propose Minpsid, an automated SID framework that automatically identifies and re-prioritizes incubative instructions in a given program to enhance SDC coverage. Evaluation shows Minpsid can effectively mitigate the loss of SDC coverage across multiple inputs. Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, Franck Cappello |
SC | 4 |
| 2022 | Improving the Accuracy of IR-Level Fault InjectionabstractFault injection (FI) is a commonly used experimental technique to evaluate the resilience of software techniques for tolerating hardware faults. Software-implemented FI can be performed at different levels of abstraction in the system stack; FI performed at the compiler’s intermediate representation (IR) level has the advantage that it is closer to the program being evaluated and is hence easier to derive insights from for the design of software fault-tolerance mechanisms. Unfortunately, it is not clear how accurate IR-level FI is vis-a-vis FI performed at the assembly code level, and prior work has presented contradictory findings. In this article, we perform a comprehensive evaluation of the accuracy of IR-level FI across a range of benchmark programs and compiler optimization levels. Our results show that IR-level FI is as accurate as assembly-level FI for silent data corruption (SDC) probability estimation across different benchmarks and optimization levels. Further, we present a machine-learning-based technique for improving the accuracy ofcrashprobability measurements made by IR-level FI, which takes advantage of an observed correlation between program crash probabilities and instructions that operate on memory address values. We find that the machine learning technique provides comparable accuracy for IR-level FI as assembly code level FI for program crashes. Lucas Palazzi, Guanpeng Li, Bo Fang 0002, Karthik Pattabiraman |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | A Low-cost Fault Corrector for Deep Neural Networks through Range RestrictionabstractDeep neural networks (DNNs) have seen growing adoption in safety-critical domains. Unfortunately, they are also subject to unexpected failures due to hardware transient faults (soft errors). Traditional fault tolerance techniques require significant implementation efforts and/or incur major performance overheads. This work introducesRanger, a low-cost fault corrector that can directly correct the faulty prediction output due to transient faults without re-computation. This research laid the foundations of improving the fault tolerance of DNN applications under hardware transient faults and it has influenced subsequent work in the area, both in academia and industry. Zitao Chen 0001, Guanpeng Li, Karthik Pattabiraman |
DSN | 2 |
| 2021 | PID-Piper: Recovering Robotic Vehicles from Physical AttacksabstractRobotic Vehicles (RV) rely extensively on sensor inputs to operate autonomously. Physical attacks such as sensor tampering and spoofing can feed erroneous sensor measurements to deviate RVs from their course and result in mission failures. In this paper, we present PID-Piper, a novel framework for automatically recovering RVs from physical attacks. We use machine learning (ML) to design an attack resilient Feed-Forward Controller (FFC), which runs in tandem with the RV's primary controller and monitors it. Under attacks, the FFC takes over from the RV's primary controller to recover the RV, and allows the RV to complete its mission successfully. Our evaluation on 6 RV systems including 3 real RVs shows that PID-Piper achieves high accuracy in emulating the RV's controller, in the absence of attacks, with no false positives. Further, PID-Piper allows RVs to complete their missions successfully despite attacks in 83% of the cases, while incurring low performance overheads. Pritam Dash, Guanpeng Li, Zitao Chen 0001, Mehdi Karimibiuki, Karthik Pattabiraman |
DSN | 2 |
| 2021 | A novel memory-efficient deep learning training framework via error-bounded lossy compressionabstractDNNs are becoming increasingly deeper, wider, and nonlinear due to the growing demands on prediction accuracy and analysis quality. When training a DNN model, the intermediate activation data must be saved in the memory during forward propagation and then restored for backward propagation. Traditional memory saving techniques such as data recomputation and migration either suffers from a high performance overhead or is constrained by specific interconnect technology and limited bandwidth. In this paper, we propose a novel memory-driven high performance CNN training framework that leverages error-bounded lossy compression to significantly reduce the memory requirement for training in order to allow training larger neural networks. Specifically, we provide theoretical analysis and then propose an improved lossy compressor and an adaptive scheme to dynamically configure the lossy compression error-bound and adjust the training batch size to further utilize the saved memory space for additional speedup. We evaluate our design against state-of-the-art solutions with four widely-adopted CNNs and the ImangeNet dataset. Results demonstrate that our proposed framework can significantly reduce the training memory consumption by up to 13.5× and 1.8× over the baseline training and state-of-the-art framework with compression, respectively, with little or no accuracy loss. The full paper can be referred to at https://arxiv.org/abs/2011.09017. Sian Jin, Guanpeng Li, Shuaiwen Song, Dingwen Tao |
PPoPP | 2 |
| 2021 | PEPPA-X: finding program test inputs to bound silent data corruption vulnerability in HPC applicationsabstractTransient hardware faults have become prevalent due to the shrinking size of transistors, leading to silent data corruptions (SDCs). Therefore, HPC applications need to be evaluated (e.g., via fault injections) and protected to meet the reliability target. In the evaluation, the target programs exercise with a set of given inputs which are usually from program benchmark suite. However, these inputs rarely manifest the SDC vulnerabilities, leading to over-optimistic assessment and unexpectedly higher failure rates in production. We propose Peppa-X, which efficiently identifies the test inputs that estimate the bound of program SDC resiliency. Our key insight is that the SDC sensitivity distribution in a program often remains stationary across input space. Thereby, we can guide the search of SDC-bound inputs by a sampled distribution. Our evaluation shows that Peppa-X can identify the SDC-bound input of a program that existing methods cannot find even with 5x more search time. Md Hasanur Rahman 0001, Aabid Shamji, Shengjian Guo, Guanpeng Li |
SC | 4 |
| 2021 | COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionabstractDeep neural networks (DNNs) are becoming increasingly deeper, wider, and non-linear due to the growing demands on prediction accuracy and analysis quality. Training wide and deep neural networks require large amounts of storage resources such as memory because the intermediate activation data must be saved in the memory during forward propagation and then restored for backward propagation. However, state-of-the-art accelerators such as GPUs are only equipped with very limited memory capacities due to hardware design constraints, which significantly limits the maximum batch size and hence performance speedup when training large-scale DNNs. Traditional memory saving techniques either suffer from performance overhead or are constrained by limited interconnect bandwidth or specific interconnect technology. In this paper, we propose a novel memory-efficient CNN training framework (called COMET) that leverages error-bounded lossy compression to significantly reduce the memory requirement for training in order to allow training larger models or to accelerate training. Our framework purposely adopts error-bounded lossy compression with a strict error-controlling mechanism. Specifically, we perform a theoretical analysis on the compression error propagation from the altered activation data to the gradients, and empirically investigate the impact of altered gradients over the training process. Based on these analyses, we optimize the error-bounded lossy compression and propose an adaptive error-bound control scheme for activation data compression. Experiments demonstrate that our proposed framework can significantly reduce the training memory consumption by up to 13.5X over the baseline training and 1.8X over another state-of-the-art compression-based framework, respectively, with little or no accuracy loss. Sian Jin, Chengming Zhang 0006, Yunhe Feng, Hui Guan 0001, Guanpeng Li, Shuaiwen Song, Dingwen Tao |
Proc. VLDB Endow. | 6 |
| 2020 | LCFI: A Fault Injection Tool for Studying Lossy Compression Error Propagation in HPC ProgramsabstractError-bounded lossy compression is becoming more and more important to today's extreme-scale HPC applications because of the ever-increasing volume of data generated because it has been widely used in in-situ visualization, data stream intensity reduction, storage reduction, I/O performance improvement, checkpoint/restart acceleration, memory footprint reduction, etc. Although many works have optimized ratio, quality, and performance for different error-bounded lossy compressors, there is none of the existing works attempting to systematically understand the impact of lossy compression errors on HPC application due to error propagation.In this paper, we propose and develop a lossy compression fault injection tool, called LCFI. To the best of our knowledge, this is the first fault injection tool that helps both lossy compressor developers and users to systematically and comprehensively understand the impact of lossy compression errors on HPC programs. The contributions of this work are threefold: (1) We propose an efficient approach to inject lossy compression errors according to a statistical analysis of compression errors for different state-of-the-art compressors. (2) We build a fault injector which is highly applicable, customizable, easy-to-use in generating top-down comprehensive results, and demonstrate the use of LCFI. (3) We evaluate LCFI on four representative HPC benchmarks with different abstracted fault models and make several observations about error propagation and their impacts on program outputs. Baodi Shan, Aabid Shamji, Jiannan Tian, Guanpeng Li, Dingwen Tao |
IEEE BigData | 4 |
| 2020 | Error Resilient Machine Learning for Safety-Critical Systems: Position PaperabstractMachine learning (ML) has increasingly been adopted in safety-critical systems such as autonomous vehicles (AVs) and industrial robotics. In these domains, reliability and safety are important considerations, and hence it is critical to ensure the resilience of ML systems to faults and errors. On the other hand, soft errors are becoming more frequent in commodity computer systems due to the effects of technology scaling and reduced supply voltages. Further, traditional solutions for masking hardware faults such as Triple-Modular Redundancy (TMR) are prohibitively expensive in terms of their energy and performance overheads. Therefore, there is a compelling need to ensure the resilience of ML applications to soft errors on commodity hardware platforms.We first experimentally assess the resilience of safety-critical ML applications to soft errors. We demonstrate through fault injection experiments that even a single bit flip due to a soft error can lead to misclassification in Deep Neural Network (DNN) applications deployed in AVs, leading to safety violations. However, not all the errors in an DNN will result in serve consequences such as safety violations, and hence it is sufficient to protect the DNN from the ones that do. Unfortunately, finding all possible errors that result in safety violations is a very compute intensive task. We propose BinFI, a fault injection approach that efficiently injects critical faults that are highly likely to result in safety violations, based on the unique properties of DNNs. Finally, we propose Ranger, an approach to protect DNNs from critical faults with minimal performance overheads and no accuracy loss. We will conclude by presenting some of our ongoing work, and the future challenges in this area. Karthik Pattabiraman, Guanpeng Li, Zitao Chen 0001 |
IOLTS | 2 |
| 2020 | TensorFI: A Flexible Fault Injection Framework for TensorFlow ApplicationsabstractAs machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application's resilience. In this work, we present TensorFI, a high-level fault injection (FI) framework for TensorFlow-based applications. TensorFI is able to inject both hardware and software faults in general TensorFlow programs. TensorFI is a configurable FI tool that is flexible, easy to use, and portable. It can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., faults in particular operators). We use TensorFI to evaluate the resilience of 12 ML programs, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also present two case studies to demonstrate the usefulness of the tool. TensorFI is publicly available at https://github.com/DependableSystemsLab/TensorFI. Zitao Chen 0001, Niranjhana Narayanan, Bo Fang 0002, Guanpeng Li, Karthik Pattabiraman, Nathan DeBardeleben |
ISSRE | 4 |
| 2020 | AV-FUZZER: Finding Safety Violations in Autonomous Driving SystemsabstractThis paper proposes AV-FUZZER, a testing framework, to find the safety violations of an autonomous vehicle (AV) in the presence of an evolving traffic environment. We perturb the driving maneuvers of traffic participants to create situations in which an AV can run into safety violations. To optimally search for the perturbations to be introduced, we leverage domain knowledge of vehicle dynamics and genetic algorithm to minimize the safety potential of an AV over its projected trajectory. The values of the perturbation determined by this process provide parameters that define participants' trajectories. To improve the efficiency of the search, we design a local fuzzer that increases the exploitation of local optima in the areas where highly likely safety-hazardous situations are observed. By repeating the optimization with significantly different starting points in the search space, AV-FUZZER determines several diverse AV safety violations. We demonstrate AV-FUZZER on an industrial-grade AV platform, Baidu Apollo, and find five distinct types of safety violations in a short period of time. In comparison, other existing techniques can find at most two. We analyze the safety violations found in Apollo and discuss their overarching causes. Guanpeng Li, Saurabh Jha, Timothy Tsai 0002, Michael B. Sullivan 0001, Siva Kumar Sastry Hari, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ISSRE | 1 |
| 2020 | GPU-trident: efficient modeling of error propagation in GPU programsabstractFault injection (FI) techniques are typically used to determine the reliability profiles of programs under soft errors. However, these techniques are highly resource- and time-intensive. Prior research developed a model, TRIDENT to analytically predict Silent Data Corruption (SDC, i.e., incorrect output without any indication) probabilities of single-threaded CPU applications without requiring FIs. Unfortunately, TRIDENT is incompatible with GPU programs, due to their high degree of parallelism and different memory architectures than CPU programs. The main challenge is that modeling error propagation across thousands of threads in a GPU kernel requires enormous amounts of data to be profiled and analyzed, posing a major scalability bottleneck for HPC applications. In this paper, we propose GPU-TRIDENT, an accurate and scalable technique for modeling error propagation in GPU programs. We find that GPU-TRIDENT is 2 orders of magnitude faster than FI-based approaches, and nearly as accurate in determining the SDC rate of GPU programs. Abdul Rehman Anwer, Guanpeng Li, Karthik Pattabiraman, Michael B. Sullivan 0001, Timothy Tsai 0002, Siva Kumar Sastry Hari |
SC | 2 |
| 2019 | A Tale of Two Injectors: End-to-End Comparison of IR-Level and Assembly-Level Fault InjectionabstractFault injection (FI) is a commonly used experimental technique to evaluate the resilience of software techniques for tolerating hardware faults. Software-implemented FI can be performed at different levels of abstraction in the system stack; FI performed at the compiler's intermediate representation (IR) level has the advantage that it is closer to the program being evaluated and is hence easier to derive insights from for the design of software fault-tolerance mechanisms. Unfortunately, it is not clear how accurate IR-level FI is vis-a-vis FI performed at the assembly code level, and prior work has presented contradictory findings. In this paper, we perform an analysis of said prior work, find an inconsistency in the FI methodology used in one study, and show that it results in a flawed comparison between IR-level and assembly-level FI. We further confirm this finding by performing a comprehensive evaluation of the accuracy of IR-level FI across a range of benchmark programs and compiler optimization levels. Our results show that IR-level FI is as accurate as assembly-level FI for silent data corruptions (SDCs) across different benchmarks and optimization levels. Lucas Palazzi, Guanpeng Li, Bo Fang 0002, Karthik Pattabiraman |
ISSRE | 2 |
| 2019 | BinFI: an efficient fault injector for safety-critical machine learning systemsabstractAs machine learning (ML) becomes pervasive in high performance computing, ML has found its way into safety-critical domains (e.g., autonomous vehicles). Thus the reliability of ML has grown in importance. Specifically, failures of ML systems can have catastrophic consequences, and can occur due to soft errors, which are increasing in frequency due to system scaling. Therefore, we need to evaluate ML systems in the presence of soft errors. Zitao Chen 0001, Guanpeng Li, Karthik Pattabiraman, Nathan DeBardeleben |
SC | 2 |
| 2018 | Modeling Input-Dependent Error Propagation in ProgramsabstractTransient hardware faults are increasing in computer systems due to shrinking feature sizes. Traditional methods to mitigate such faults are through hardware duplication, which incurs huge overhead in performance and energy consumption. Therefore, researchers have explored software solutions such as selective instruction duplication, which require fine-grained analysis of instruction vulnerabilities to Silent Data Corruptions (SDCs). These are typically evaluated via Fault Injection (FI), which is often highly time-consuming. Hence, most studies confine their evaluations to a single input for each program. However, there is often significant variation in the SDC probabilities of both the overall program and individual instructions across inputs, which compromises the correctness of results with a single input. In this work, we study the variation of SDC probabilities across different inputs of a program, and identify the reasons for the variations. Based on the observations, we propose a model, VTRIDENT, which predicts the variations in programs' SDC probabilities without any FIs, for a given set of inputs. We find that VTRIDENT is nearly as accurate as FI in identifying the variations in SDC probabilities across inputs. We demonstrate the use of VTRIDENT to bound overall SDC probability of a program under multiple inputs, while performing FI on only a single input. Guanpeng Li, Karthik Pattabiraman |
DSN | 1 |
| 2018 | Modeling Soft-Error Propagation in ProgramsabstractAs technology scales to lower feature sizes, devices become more susceptible to soft errors. Soft errors can lead to silent data corruptions (SDCs), seriously compromising the reliability of a system. Traditional hardware-only techniques to avoid SDCs are energy hungry, and hence not suitable for commodity systems. Researchers have proposed selective software-based protection techniques to tolerate hardware faults at lower costs. However, these techniques either use expensive fault injection or inaccurate analytical models to determine which parts of a program must be protected for preventing SDCs. In this work, we construct a three-level model, TRIDENT, that captures error propagation at the static data dependency, control-flow and memory levels, based on empirical observations of error propagations in programs. TRIDENT is implemented as a compiler module, and it can predict both the overall SDC probability of a given program and the SDC probabilities of individual instructions, without fault injection. We find that TRIDENT is nearly as accurate as fault injection and it is much faster and more scalable. We also demonstrate the use of TRIDENT to guide selective instruction duplication to efficiently mitigate SDCs under a given performance overhead bound. Guanpeng Li, Karthik Pattabiraman, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Timothy Tsai 0002 |
DSN | 1 |
| 2017 | Understanding error propagation in deep learning neural network (DNN) accelerators and applicationsabstractDeep learning neural networks (DNNs) have been successful in solving a wide range of machine learning problems. Specialized hardware accelerators have been proposed to accelerate the execution of DNN algorithms for high-performance and energy efficiency. Recently, they have been deployed in datacenters (potentially for business-critical or industrial applications) and safety-critical systems such as self-driving cars. Soft errors caused by high-energy particles have been increasing in hardware systems, and these can lead to catastrophic failures in DNN systems. Traditional methods for building resilient systems, e.g., Triple Modular Redundancy (TMR), are agnostic of the DNN algorithm and the DNN accelerator's architecture. Hence, these traditional resilience approaches incur high overheads, which makes them challenging to deploy. In this paper, we experimentally evaluate the resilience characteristics of DNN systems (i.e., DNN software running on specialized accelerators). We find that the error resilience of a DNN system depends on the data types, values, data reuses, and types of layers in the design. Based on our observations, we propose two efficient protection techniques for DNN systems. Guanpeng Li, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Timothy Tsai 0002, Karthik Pattabiraman, Joel S. Emer, Stephen W. Keckler |
SC | 1 |
| 2017 | Configurable Detection of SDC-causing Errors in ProgramsabstractSilent Data Corruption (SDC) is a serious reliability issue in many domains, including embedded systems. However, current protection techniques are brittle and do not allow programmers to trade off performance for SDC coverage. Further, many require tens of thousands of fault-injection experiments, which are highly time- and resource-intensive. In this article, we propose two empirical models, SDCTune and SDCAuto , to predict the SDC proneness of a program’s data. Both models are based on static and dynamic features of the program alone and do not require fault injections to be performed. The main difference between them is that SDCTune requires manual tuning while SDCAuto is completely automated, using machine-learning algorithms. We then develop an algorithm using both models to selectively protect the most SDC-prone data in the program subject to a given performance overhead bound. Our results show that both models are accurate at predicting the relative SDC rate of an application compared to fault injection, for a fraction of the time taken. Further, in terms of efficiency of detection (i.e., ratio of SDC coverage provided to performance overhead), our technique outperforms full duplication by a factor of 0.78x to 1.65x with the SDCTune model and 0.62x to 0.96x with SDCAuto model. Qining Lu, Guanpeng Li, Karthik Pattabiraman, Meeta Sharma Gupta, Jude A. Rivers |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Understanding error propagation in GPGPU applicationsabstractGPUs have emerged as general-purpose accelerators in high-performance computing (HPC) and scientific applications. However, the reliability characteristics of GPU applications have not been investigated in depth. While error propagation has been extensively investigated for non-GPU applications, GPU applications have a very different programming model which can have a significant effect on error propagation in them. We perform an empirical study to understand and characterize error propagation in GPU applications. We build a compilerbased fault-injection tool for GPU applications to track error propagation, and define metrics to characterize propagation in GPU applications. We find GPU applications exhibit significant error propagation for some kinds of errors, but not others, and the behaviour is highly application specific. We observe the GPUCPU interaction boundary naturally limits error propagation in these applications compared to traditional non-GPU applications. We also formulate various guidelines for the design of faulttolerance mechanisms in GPU applications based on our results. Guanpeng Li, Karthik Pattabiraman, Chen-Yong Cher, Pradip Bose |
SC | 1 |
| 2016 | Automatic fault localization for client-side JavaScriptabstractSummary JAVASCRIPTis a scripting language that plays a prominent role in web applications today. It is dynamic, loosely typed and asynchronous and is extensively used to interact with the Document Object Model (DOM) at runtime. All these characteristics makeJAVASCRIPTcode error‐prone; unfortunately,JAVASCRIPTfault localization remains a tedious and mainly manual task. Despite these challenges, the problem has received very limited research attention. This paper proposes an automated technique to localizeJAVASCRIPTfaults based on dynamic analysis, tracing and backward slicing ofJAVASCRIPTcode. This technique is capable of handling features ofJAVASCRIPTcode that have traditionally been difficult to analyse, includingeval, anonymous functions and minified code. The approach is implemented in an open source tool calledAUTOFLOX, and evaluation results indicate that it is capable of (1) automatically localizing DOM‐relatedJAVASCRIPTfaults with high accuracy (over 96%) and no false‐positives and (2) isolatingJAVASCRIPTfaults in production websites and actual bugs from real‐world web applications. Copyright © 2015 John Wiley & Sons, Ltd. Frolin S. Ocariza Jr., Guanpeng Li, Karthik Pattabiraman, Ali Mesbah 0001 |
Softw. Test. Verification Reliab. | 2 |
| 2015 | Fine-Grained Characterization of Faults Causing Long Latency Crashes in ProgramsabstractAs the rate of transient hardware faults increases, researchers have investigated software techniques to tolerate these faults. An important class of faults are those that cause long- latency crashes (LLCs), or faults that can persist for a long time in the program before causing it to crash. In this paper, we develop a technique to automatically find program locations where LLC causing faults originate so that the locations can be protected to bound the program's crash latency. We first identify program code patterns that are responsible for the majority of LLC causing faults through an empirical study. We then build CRASHFINDER, a tool that finds LLC locations by statically searching the program for the patterns, and then refining the static analysis results with a dynamic analysis and selective fault injection-based approach. We find that CRASHFINDER can achieve an average of 9.29 orders of magnitude time reduction to identify more than 90% of LLC causing locations in the program, compared to exhaustive fault injection techniques, and has no false-positives. Guanpeng Li, Qining Lu, Karthik Pattabiraman |
DSN | 1 |
| 2015 | Experience report: An application-specific checkpointing technique for minimizing checkpoint corruptionabstractCheckpointing is widely deployed in computer systems to recover from failures due to both hardware and software errors. However, as faults propagate, checkpoints may become corrupted by saving erroneous states and make errors unrecoverable, especially at aggressive checkpoint frequencies. In this paper, we proposed a technique that automatically analyzes a given program to guide checkpoint strategies in order to minimize checkpoint corruptions. To understand checkpoint corruptions, we first perform a large-scale fault injection study across ten benchmark applications. We then classify checkpoint corruptions, and comprehensively characterize the fault propagations leading to these corruptions. Leveraging these findings, we build ReCov, a compiler-based tool that automatically identifies the program locations that have lowest density of fault propagation for placing checkpoints, and combines it with low-overhead protection techniques. Our experimental results shows that ReCov can eliminate nearly 92% of the checkpoint corruptions with about 5% performance overhead. ReCov reduces the unavailability of the system by 8.25 times even at very aggressive checkpoint frequencies, showing that it is effective in practice. Guanpeng Li, Karthik Pattabiraman, Chen-Yong Cher, Pradip Bose |
ISSRE | 1 |
| 2014 | Quantifying the Accuracy of High-Level Fault Injection Techniques for Hardware FaultsabstractHardware errors are on the rise with reducing feature sizes, however tolerating them in hardware is expensive. Researchers have explored software-based techniques for building error resilient applications. Many of these techniques leverage application-specific resilience characteristics to keep overheads low. Understanding application-specific resilience characteristics requires software fault-injection mechanisms that are both accurate and capable of operating at a high-level of abstraction to allow developers to reason about error resilience. In this paper, we quantify the accuracy of high-level software fault injection mechanisms vis-à-vis those that operate at the assembly or machine code levels. To represent high-level injection mechanisms, we built a fault injector tool based on the LLVM compiler, called LLFI. LLFI performs fault injection at the LLVM intermediate code level of the application, which is close to the source code. We quantitatively evaluate the accuracy of LLFI with respect to assembly level fault injection, and understand the reasons for the differences. Jiesheng Wei, Anna Thomas, Guanpeng Li, Karthik Pattabiraman |
DSN | 3 |