Bo Fang 0002

dblp:86/388-2 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-9721-3982ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-author · 18 since 2021Security and privacy · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating AI Compression through Lightweight Lossless Encoding and Pipelined Workflows
Boyuan Zhang 0002, Luanzheng Guo, Jiannan Tian, Jinyang Liu 0003, Daoce Wang, Chengming Zhang 0006, Bo Fang 0002, Fengguang Song, Jan Strube 0001, Nathan R. Tallent, Dingwen Tao
IPDPS7
2025 FT2: First-Token-Inspired Online Fault Tolerance on Critical Layers for Generative Large Language Models
abstract
Generative Large Language Models (LLMs) are deployed on large-scale computing systems, where such tasks unavoidably suffer from soft errors, leading to quality degradation of content generated by LLMs. Enhancing LLM resilience is particularly challenging because of its complicated model architecture and tremendous size. State-of-the-art protections have limitations such as high overhead and incomplete coverage, and often require offline profiling.
Cherish Mulpuru, Roberto Gioiosa, Zhao Zhang 0007, Bo Fang 0002, Lishan Yang 0001
HPDC6
2025 UQ-VarQA: Benchmarking and Characterizing NISQ Computers Through Uncertainty Quantification of Variational Quantum Algorithms
abstract
Quantum computing offers speedups, but NISQ processors face hardware-induced errors that degrade fidelity and reproducibility. We introduce an uncertainty-aware benchmarking framework that combines uncertainty quantification with global sensitivity analysis to evaluate not only peak fidelity but also its reliability over time. Using Bayesian Optimization with SGLD refinement under fixed budgets, seeds, bounds, and trust-region rules, and repeating runs across days, we capture calibration drift and quantify efficiency, stability, landscape complexity, and maintenance cost. A noise-aware Gaussian process surrogate provides scalable sensitivity estimates without error mitigation. Applied to VQAMET and VQC on three IBMQ backends, the framework delivers actionable guidance for co-tuning and backend selection, complementing and often outperforming quantum volume style metrics.
Priyabrata Senapati, Shengye Zhu, Bo Peng 0024, Bo Fang 0002, Qiang Guan
ICCD4
2025 Can Large Language Models Understand Intermediate Representations in Compilers?
abstract
Intermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by *Large Language Models* (LLMs) remains underexplored. In this paper, we present an explorative empirical study evaluating the capabilities of six state-of-the-art LLMs—GPT-4, GPT-3, DeepSeek, Gemma 2, Llama 3, and Code Llama—in understanding IRs. Specifically, we assess model performance across four core tasks: *control flow graph reconstruction*, *decompilation*, *code summarization*, and *execution reasoning*. While LLMs exhibit competence in parsing IR syntax and identifying high-level structures, they consistently struggle with instruction-level reasoning, especially in control flow reasoning, loop handling, and dynamic execution. Common failure modes include misinterpreting branching instructions, omitting critical operations, and relying on heuristic reasoning rather than on precise instruction-level logic. Our findings highlight the need for IR-specific enhancements in LLM design. We recommend fine-tuning on structured IR datasets and integrating control-flow-sensitive architectures to improve the models’ effectiveness on IR-related tasks. All the experimental data and source code are publicly available at [https://github.com/hjiang13/LLM4IR](https://github.com/hjiang13/LLM4IR).
Hailong Jiang, Yao Wan 0001, Bo Fang 0002, Hongyu Zhang 0002, Ruoming Jin, Qiang Guan
ICML4
2025 BMQSim: Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
Boyuan Zhang 0002, Bo Fang 0002, Fanjiang Ye, Luanzheng Guo, Fengguang Song, Nathan R. Tallent, Dingwen Tao
ICS2
2025 Be Aware of Metadata Corruption in Parallel File System: It can be Silent and Catastrophic
abstract
High-Performance Computing (HPC) systems rely on parallel file systems (PFS) like Lustre to reliably manage large-scale data. Unfortunately, such data stored in PFS are subjected to corruption due to hardware failures, power outages, on-disk corruption, in-memory corruption, and software bugs. Previous work has shown the impacts of general data corruptions on scientific applications, but barely investigated the impacts of corruptions on metadata, which is a special type of data for many critical PFS functionalities. Since the metadata are accessed and updated frequently, the metadata corruption is non-negligible, and the impact on both HPC systems and applications is often complicated. In this study, we systematically studied the effects of PFS metadata corruption on representative scientific applications and workflows through fault injections. We observed various abnormal behaviors of applications against metadata corruptions, such as failing to finish execution or successfully finishing execution but generating wrong outputs silently. We also showed that neither the defensive programming practice from the application developers nor existing metadata checking mechanism deployed in PFS can resolve the metadata corruption issues or guild applications to react correctly. To address this issue, we further propose a metadata checksum mechanism as a system-level mitigation strategy to detect and handle metadata corruptions at runtime. We implement a prototype of the checksum mechanism on Lustre file system using FUSE interface and evaluate its functionality and performance overhead. Our results show the proposed checksum mechanism can effectively detect metadata corruptions during data accesses with extremely low overhead and stop the applications from continuous execution and generating wrong results.
Saisha Kamat, Mai Zheng, Bo Fang 0002, Dong Dai 0001
IPDPS3
2025 ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.
Yuhang Liang, Jie Ren 0015, Ang Li 0006, Bo Fang 0002, Jieyang Chen
PPoPP5
2025 Demystifying the Resilience of Large Language Model Inference: An End-to-End Perspective
abstract
Deep neural networks are known to be resilient to random bitwise faults in their parameters. However, this resilience has primarily been established through studies of classification models. The extent to which this claim holds for large-language models remains under-explored. In this work, we conduct an extensive measurement study on the impact of random bitwise faults in commercial-scale language model inference. We first expose that these language models are not truly resilient to random bit-flips. While aggregate metrics such as accuracy may suggest resilience, an in-depth inspection of the generated outputs shows significant degradation in text quality. Our analysis also shows that tasks requiring more complex reasoning suffer more from performance and quality degradation. Moreover, we extend our resilience analysis to models with augmented reasoning capabilities, such as Chain-of-Thought or Mixture of Experts architectures.
Zachary Coalson, Shiyang Chen 0004, Hang Liu 0001, Zhao Zhang 0007, Sanghyun Hong 0001, Bo Fang 0002, Lishan Yang 0001
SC7
2025 QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers
abstract
Protein structure prediction is a core challenge in computational biology, particularly for fragments within ligand-binding regions, where accurate modeling is still difficult. Quantum computing offers a novel first-principles modeling paradigm, but its application is currently limited by hardware constraints, high computational cost, and the lack of a standardized benchmarking dataset. In this work, we present QDockBank—the first large-scale protein fragment structure dataset generated entirely using utility-level quantum computers, specifically designed for protein–ligand docking tasks. QDockBank comprises 55 protein fragments extracted from ligand-binding pockets. The dataset was generated through tens of hours of execution on superconducting quantum processors, making it the first quantum-based protein structure dataset with a total computational cost exceeding one million USD. Experimental evaluations demonstrate that structures predicted by QDockBank outperform those predicted by AlphaFold2 and AlphaFold3 in terms of both RMSD and docking affinity scores. QDockBank serves as a new benchmark for evaluating quantum-based protein structure prediction.
Yuxin Yang 0001, Cheng-Chang Lu, Weiwen Jiang, Feixiong Cheng, Bo Fang 0002, Qiang Guan
SC6
2024 Red-QAOA: Efficient Variational Optimization through Circuit Reduction
abstract
The Quantum Approximate Optimization Algorithm (QAOA) addresses combinatorial optimization challenges by converting inputs to graphs. However, the optimal parameter searching process of QAOA is greatly affected by noise. Larger problems yield bigger graphs, requiring more qubits and making their outcomes highly noise-sensitive. This paper introduces Red-QAOA, leveraging energy landscape concentration via a simulated annealing-based graph reduction.
Meng Wang 0033, Bo Fang 0002, Ang Li 0006, Prashant J. Nair
ASPLOS (2)2
2024 FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
abstract
NVIDIA Tensor Cores and AMD Matrix Cores (together called Matrix Accelerators) are of growing interest in high-performance computing and machine learning owing to their high performance. Unfortunately, some of their crucial numerical attributes pertaining to departures from full IEEE floating-point compatibility are not documented. This makes it impossible to reliably port codes across these differing accelerators. This paper contributes a collection of Feature Targeted Tests for Numerical Properties that that help determine these features across five floating-point formats, four rounding modes and additional that highlight the rounding behaviors and preservation of extra precision bits. To show the practical relevance of FTTN, we design a simple matrix-multiplication test designed with insights gathered from our feature-tests. We executed this very simple test on five platforms, producing different answers: V100, A100, and MI250X produced 0, MI100 produced 255.875, and Hopper H100 produced 191.875. Our matrix multiplication tests employ patterns found in iterative refinement-based algorithms, highlighting the need to check for significant result variability when porting code across GPUs.
Ang Li 0006, Bo Fang 0002, Katarzyna Swirydowicz, Ignacio Laguna, Ganesh Gopalakrishnan
CCGrid3
2024 Discovery of Floating-Point Differences Between NVIDIA and AMD GPUs
abstract
NVIDIA and AMD GPUs are fundamental components in contemporary high-performance systems, boosting computational capabilities in the HPC and AI fields.However, a clear understanding of the nuances in floating-point operations between these GPU variants is crucial to avoid introducing errors during software development or porting, and such clarity is currently insufficient.The complexity of this issue is amplified when considering the variety of floating-point precision options (such as FP16, FP32, etc.), floating-point formats (like standard floats, bfloats, etc.), and the different execution units (elementary units, matrix/tensor cores, etc.).As it stands, much of this information is either not well-known or is difficult to obtain. Our work aims to shed light on these areas through a pioneering testing-guided methodology that seeks to unravel many of these uncertainties.We are in the process of developing a series of tests that uncover the numerical discrepancies in elementary computing units, the built-in math libraries, and the numerical properties of matrix accelerators present in both NVIDIA (tensor cores) and AMD GPUs (matrix cores).The significance of this testing approach extends beyond current GPU models; it is designed to be forward-compatible with upcoming GPU technologies. We have already identified discrepancies as significant as 7 ulps for trigonometric functions at FP32 precision and 3 ulps at FP64 precision between NVIDIA and AMD GPUs. Additionally, our comprehensive examination has documented the behaviors of matrix cores (NVIDIA) and tensor cores (AMD), including their rounding modes (such as truncation and round-to-nearest), the extent of extra internal bits maintained (specifically, whether an additional 3 bits are retained), the handling of subnormal numbers in inputs and outputs and the FMA features in these units. This analysis spans four distinct floating-point formats and multiple GPU models, including NVIDIA’s V100, A100, H100 and AMD’s MI100 and MI250X.We believe that the information now being disclosed will reduce the risk of porting errors when codes are adapted across these different hardware platforms.
Ang Li 0006, Bo Fang 0002, Katarzyna Swirydowicz, Ignacio Laguna, Ganesh Gopalakrishnan
CCGrid3
2024 Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience
abstract
Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on error resilience characteristics, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.
Bo Fang 0002, Harvey Dam, Cheng Tan 0002, Siva Kumar Sastry Hari, Timothy Tsai 0002, Ignacio Laguna, Dingwen Tao, Ganesh Gopalakrishnan, Prashant J. Nair, Kevin J. Barker, Ang Li 0006
CLUSTER1
2024 Privacy-Preserving Artificial Intelligence on Edge Devices: A Homomorphic Encryption Approach
abstract
Recent advancements in privacy-preserving artificial intelligence (AI) have paved the way for enhanced privacy in computational processes. A standing challenge, however, is the robust privacy preservation in AI algorithms, especially when integrated into edge devices and Internet-of-Thing (IoT) infrastructures. Most prevailing solutions have adopted traditional encryption methods which, though secure, often introduce significant overhead and potential dips in accuracy. In this study, we put forth an innovative approach, utilizing the CKKS encryption scheme, aiming to harmoniously balance computational efficiency with stringent data privacy. By harnessing the capabilities of Full Homomorphic Encryption (FHE) under the CKKS scheme, we ensure the preservation of privacy, successfully curbing the inherent noise traditionally linked with accuracy reductions in similar encryption-oriented solutions. Through comprehensive experiments, our approach showcased its potential as a strong contender for privacy preservation, demonstrating commendable performance across all tests, affirming that FHE is indeed viable for devices with constrained computational power and energy resources.
Muhammad Jahanzeb Khan, Bo Fang 0002, Gaetano Cimino, Stefano Cirillo, Lei Yang 0001, Dongfang Zhao 0001
ICWS2
2024 HAppA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded
abstract
High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method - HAppA-LSTM - achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of key words representing the source code. A comprehensive importance analysis of these key words further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.
Hailong Jiang, Bo Fang 0002, Kevin J. Barker, Ruoming Jin, Qiang Guan
SRDS3
2023 Design and Evaluation of GPU-FPX: A Low-Overhead tool for Floating-Point Exception Detection in NVIDIA GPUs
abstract
Floating-point exceptions occurring during numerical computations can be a serious threat to the validity of the computed results if they are not caught and diagnosed Unfortunately, on NVIDIA GPUs-today's most widely used types and which do not have hardware exception traps-this task must be carried out in software. Given the prevalence of closed-source kernels, efficient binary-level exception tracking is essential. It is also important to know how exceptions flow through the code, whether they alter the code behavior and additionally whether these exceptions can be detected at the program outputs or are killed inside program flow-paths.
Ignacio Laguna, Bo Fang 0002, Katarzyna Swirydowicz, Ang Li 0006, Ganesh Gopalakrishnan
HPDC3
2023 Visilience: An Interactive Visualization Framework for Resilience Analysis using Control-Flow Graph
abstract
Soft errors have become one of the main concerns for the resilience of HPC applications, as these errors can cause HPC applications to generate serious outcomes such as silent data corruption (SDC). Many approaches have been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to perform resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of detail, which can create obstacles to compare and explore the resilience of the HPC program execution. To this end, we present Visilience, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, Visilience leverages an effective visualization approach, Control Flow Graph (CFG) to present a function execution. Furthermore, three widely used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly integrated into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework Visilience.
Hailong Jiang, Shaolun Ruan, Bo Fang 0002, Yong Wang 0021, Qiang Guan
PRDC3
2023 AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications
abstract
As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences exponential growth. Adaptive Mesh Refinement (AMR) has emerged as an effective solution to address these two challenges. Concurrently, error-bounded lossy compression is recognized as one of the most efficient approaches to tackle the latter issue. Despite their respective advantages, few attempts have been made to investigate how AMR and error-bounded lossy compression can function together. To this end, this study presents a novel in-situ lossy compression framework that employs the HDF5 filter to improve both I/O costs and boost compression quality for AMR applications. We implement our solution into the AMReX framework and evaluate on two real-world AMR applications, Nyx and WarpX, on the Summit supercomputer. Experiments with 4096 CPU cores demonstrate that AMRIC improves the compression ratio by up to 81× and the I/O performance by up to 39× over AMReX's original compression solution.
Daoce Wang, Jesus Pulido, Pascal Grosset, Jiannan Tian, Sian Jin, Houjun Tang, Jean M. Sexton, Sheng Di, Kai Zhao 0008, Bo Fang 0002, Zarija Lukic, Franck Cappello, James P. Ahrens, Dingwen Tao
SC10
2023 Fault Injection for TensorFlow Applications
abstract
As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application's resilience. In this work, we present TensorFI 1 and TensorFI 2, high-level fault injection (FI) frameworks for TensorFlow-based applications. TensorFI 1 and 2 are able to inject both hardware and software faults in any general TensorFlow 1 and 2 program respectively. Both are configurable FI tools that are flexible, easy to use, and portable. They can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., bit-flips in particular operations or layers). We use TensorFI 1 and TensorFI 2 to evaluate the resilience of 11 and 10 ML programs respectively, all written in TensorFlow, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also measure the performance overheads of the two injectors, and present 4 case studies, two for each tool, to demonstrate their utility.
Niranjhana Narayanan, Zitao Chen 0001, Bo Fang 0002, Guanpeng Li, Karthik Pattabiraman, Nathan DeBardeleben
IEEE Trans. Dependable Secur. Comput.3
2022 Efficient Hierarchical State Vector Simulation of Quantum Circuits via Acyclic Graph Partitioning
abstract
Early but promising results in quantum computing have been enabled by the concurrent development of quan-tum algorithms, devices, and materials. Classical simulation of quantum programs has enabled the design and analysis of algorithms and implementation strategies targeting current and anticipated quantum device architectures. In this paper, we present a graph-based approach to achieving efficient quantum circuit simulation. Our approach involves partitioning the graph representation of a given quantum circuit into acyclic sub-graphs/circuits that exhibit better data locality. Simulation of each sub-circuit is organized hierarchically, with the iterative construction and simulation of smaller state vectors, improving overall performance. Also, this partitioning reduces the number of passes through data, improving the total computation time. We present three partitioning strategies and observe that acyclic graph partitioning typically results in the best time-to-solution. In contrast, other strategies reduce the partitioning time at the expense of potentially increased simulation times. Experimental evaluation demonstrates the effectiveness of our approach.
Bo Fang 0002, M. Yusuf Özkaya, Ang Li 0006, Ümit V. Çatalyürek, Sriram Krishnamoorthy
CLUSTER1
2022 ASAP: automatic synthesis of area-efficient and precision-aware CGRAs
abstract
Coarse-grained reconfigurable accelerators (CGRAs) are a promising accelerator design choice that strikes a balance between performance and adaptability to different computing patterns across various applications domains. Designing a CGRA for a specific application domain involves enormous software/hardware engineering effort. Recent research works explore loop transformations, functional unit types, network topology, and memory size to identify optimal CGRA designs given a set of kernels from a specific application domain. Unfortunately, the impact of functional units with different precision support has rarely been investigated. To address this gap, we propose ASAP - a hardware/software co-design framework that automatically identifies and synthesizes optimal precision-aware CGRA for a set of applications of interest. Our evaluation shows that ASAP generates specialized designs 3.2X, 4.21X, and 5.8X more efficient (in terms of performance per unit of energy or area) than non-specialized homogeneous CGRAs, for the scientific computing, embedded, and edge machine learning domains, respectively, with limited accuracy loss. Moreover, ASAP provides more efficient designs than other state-of-the-art synthesis frameworks for specialized CGRAs.
Cheng Tan 0002, Thierry Tambe, Jeff Zhang 0001, Bo Fang 0002, Tong Geng, Gu-Yeon Wei, David Brooks 0001, Antonino Tumeo, Ganesh Gopalakrishnan, Ang Li 0006
ICS4
2022 MARS: Malleable Actor-Critic Reinforcement Learning Scheduler
abstract
In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARSensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the predefined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARSupdates the Deep Neural Network (DNN) model based on the reward. MARSis designed to optimize the existing models through reinforcement mechanisms. MARSadapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARSwith different real-world workflow traces. MARS can achieve 5%–60% increased performance compared to the state-of-the-art approaches.
Betis Baheri, Jake Tronge, Bo Fang 0002, Ang Li 0006, Vipin Chaudhary, Qiang Guan
IPCCC3
2022 Improving the Accuracy of IR-Level Fault Injection
abstract
Fault injection (FI) is a commonly used experimental technique to evaluate the resilience of software techniques for tolerating hardware faults. Software-implemented FI can be performed at different levels of abstraction in the system stack; FI performed at the compiler’s intermediate representation (IR) level has the advantage that it is closer to the program being evaluated and is hence easier to derive insights from for the design of software fault-tolerance mechanisms. Unfortunately, it is not clear how accurate IR-level FI is vis-a-vis FI performed at the assembly code level, and prior work has presented contradictory findings. In this article, we perform a comprehensive evaluation of the accuracy of IR-level FI across a range of benchmark programs and compiler optimization levels. Our results show that IR-level FI is as accurate as assembly-level FI for silent data corruption (SDC) probability estimation across different benchmarks and optimization levels. Further, we present a machine-learning-based technique for improving the accuracy ofcrashprobability measurements made by IR-level FI, which takes advantage of an observed correlation between program crash probabilities and instructions that operate on memory address values. We find that the machine learning technique provides comparable accuracy for IR-level FI as assembly code level FI for program crashes.
Lucas Palazzi, Guanpeng Li, Bo Fang 0002, Karthik Pattabiraman
IEEE Trans. Dependable Secur. Comput.3
2021 Characterizing Impacts of Storage Faults on HPC Applications: A Methodology and Insights
abstract
In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to equip with more advanced storage infrastructures such as solid-state disks (SSDs). While SSDs offer high-performance I/O, the reliability challenges faced by the HPC applications under the SSD-related failures remains unclear, in particular for failures resulting in data corruptions. The goal of this paper is to understand the impact of SSD-related faults on the behaviors of complex HPC applications. To this end, we propose FFIS, a FUSE-based fault injection framework that systematically introduces storage faults into the application layer to model the errors originated from SSDs. FFIS is able to plant different I/O related faults into the data returned from underlying file systems, which enables the investigation on the error resilience characteristics of the scientific file format. We demonstrate the use of FFIS with three representative real HPC applications, showing how each application reacts to the data corruptions, and provide insights on the error resilience of the widely adopted HDF5 file format for the HPC applications.
Bo Fang 0002, Daoce Wang, Sian Jin, Quincey Koziol, Zhao Zhang 0007, Qiang Guan, Surendra Byna, Sriram Krishnamoorthy, Dingwen Tao
CLUSTER1
2021 A Hybrid System for Learning Classical Data in Quantum States
abstract
Deep neural network powered artificial intelligence has rapidly changed our daily life with various applications. However, as one of the essential steps of deep neural networks, training a heavily-weighted network requires a tremendous amount of computing resources. Especially in the post Moore’s Law era, the limit of semiconductor fabrication technology has restricted the development of learning algorithms to cope with the increasing high intensity training data. Meanwhile, quantum computing has demonstrated its significant potential in terms of speeding up the traditionally compute-intensive workloads. For example, Google illustrated quantum supremacy by completing a sampling calculation task in 200 seconds, which is otherwise impracticable on the world’s largest supercomputers. To this end, quantum-based learning has become an area of interest, with the potential of a quantum speedup. In this paper, we propose GenQu, a hybrid and general-purpose quantum framework for learning classical data through quantum states. We evaluate GenQu with real datasets and conduct experiments on both simulations and real quantum computer IBM-Q. Our evaluation demonstrates that, compared with classical solutions, the proposed models running on GenQu framework achieve similar accuracy with a much smaller number of qubits, while significantly reducing the parameter size by up to 95.86% and converging speedup by 33.33% faster.
Samuel A. Stein, Ryan L'Abbate, Wenrui Mu, Betis Baheri, Ying Mao 0001, Qiang Guan, Ang Li 0006, Bo Fang 0002
IPCCC9
2021 SV-sim: scalable PGAS-based state vector simulation of quantum circuits
abstract
High-performance quantum circuit simulation in a classic HPC is still imperative in the NISQ era. Observing that the major obstacle of scalable state-vector quantum simulation arises from the massively fine-grained irregular data-exchange with remote nodes, in this paper we present SV-Sim to apply the PGAS-based communication models (i.e., direct peer access for intra-node CPUs/GPUs and SHMEM for inter-node CPU/GPU clusters) for efficient generalpurpose quantum circuit simulation. Through an orchestrated design based on device functional pointer, SV-Sim is able to abstract various quantum gates across multiple heterogeneous backends, including IBM/Intel/AMD CPUs, NVIDIA/AMD GPUs, and Intel Xeon Phi, in a unified framework, but still asserting outstanding performance and tractable interface to higher-level quantum programming environments, such as IBM Qiskit, Microsoft Q# and Google Cirq. Circumventing the obstacle from the lack of polymorphism in GPUs and leveraging the device-initiated one-sided communication, SV-Sim can process circuit that are dynamically generated in Python using a single GPU/CPU kernel without the need of expensive JIT or runtime parsing, significantly simplifying the programming complexity and improving performance for QC simulation. This is especially appealing for the variational quantum algorithms given the circuits are synthesized online per iteration. Evaluations on the latest NVIDIA DGX-A100, V100-DGX-2, ALCF Theta, OLCF Spock, and OLCF Summit HPCs show that SV-Sim can deliver scalable performance on various state-of-the-art HPC platforms, offering a useful tool for quantum algorithm validation and verification. SV-Sim has been released at http://github.com/pnnl/sv-sim. A version specially tweaked for Q#/QDK is also provided.
Ang Li 0006, Bo Fang 0002, Christopher E. Granade, Guen Prawiroatmodjo, Bettina Heim, Martin Rötteler, Sriram Krishnamoorthy
SC2
2020 Chaser: An Enhanced Fault Injection Tool for Tracing Soft Errors in MPI Applications
abstract
Resilient computation has been an emerging topic in the field of high-performance computing (HPC). In particular, studies show that tolerating faults on leadership-class supercomputers (such as exascale supercomputers) is expected to be one of the main challenges. In this paper, we utilize dynamic binary instrumentation and virtual machine based fault injection to emulate soft errors and study the soft errors' impact on the behavior of applications. We propose Chaser, a fine-grained, accountable, flexible, and efficient fault injection framework built on top of QEMU. Chaser offers just-in-time fault injection, the ability to trace fault propagation, and flexible and programable interfaces. In the case study, we demonstrate the usage of Chaser on Matvec and a real DOE mini MPI application
Qiang Guan, Xunchao Hu, Terence Grove, Bo Fang 0002, Hailong Jiang, Heng Yin 0001, Nathan DeBardeleben
DSN4
2020 TensorFI: A Flexible Fault Injection Framework for TensorFlow Applications
abstract
As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application's resilience. In this work, we present TensorFI, a high-level fault injection (FI) framework for TensorFlow-based applications. TensorFI is able to inject both hardware and software faults in general TensorFlow programs. TensorFI is a configurable FI tool that is flexible, easy to use, and portable. It can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., faults in particular operators). We use TensorFI to evaluate the resilience of 12 ML programs, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also present two case studies to demonstrate the usefulness of the tool. TensorFI is publicly available at https://github.com/DependableSystemsLab/TensorFI.
Zitao Chen 0001, Niranjhana Narayanan, Bo Fang 0002, Guanpeng Li, Karthik Pattabiraman, Nathan DeBardeleben
ISSRE3
2019 BonVoision: leveraging spatial data smoothness for recovery from memory soft errors
abstract
Detectable but Uncorrectable Errors (DUEs) in the memory subsystem are becoming increasingly frequent. Today, upon encountering a DUE, applications crash, and the recovery methods used incur significant performance, storage, and energy overheads. To mitigate the impact of these errors, we start from two high-level observations that apply to some classes of HPC applications (e.g., stencil computations on regular grids or irregular meshes): first, these applications, display a property we dub spatial data smoothness: i.e., data items that are nearby in the application's logical space are relatively similar. Second, since these data items are generally used together, programmers go to great lengths to place them in nearby memory locations to improve application's performance by improving access locality. Based on these observations we explore the feasibility of a roll-forward recovery scheme that leverages spatial data smoothness to repair the memory location corrupted by a DUE and continues the application execution. We present BonVoision, a run-time system that intercepts DUE events, analyzes the application binary at runtime to identify the data elements in the neighborhood of the memory location that generates a DUE, and uses them to fix the corrupted data. Our evaluation demonstrates that BonVoision is: (i) efficient - it incurs negligible overhead, (ii) effective - it is frequently successful in continuing the application with benign outcomes, and (iii) user friendly - as it does not require programmer input to expose the data layout or access to source code. We demonstrate that using BonVoision can lead to significant savings in the context of a checkpointing/restart schemes by enabling longer checkpoint intervals.
Bo Fang 0002, Hassan Halawa, Karthik Pattabiraman, Matei Ripeanu, Sriram Krishnamoorthy
ICS1
2019 A Tale of Two Injectors: End-to-End Comparison of IR-Level and Assembly-Level Fault Injection
abstract
Fault injection (FI) is a commonly used experimental technique to evaluate the resilience of software techniques for tolerating hardware faults. Software-implemented FI can be performed at different levels of abstraction in the system stack; FI performed at the compiler's intermediate representation (IR) level has the advantage that it is closer to the program being evaluated and is hence easier to derive insights from for the design of software fault-tolerance mechanisms. Unfortunately, it is not clear how accurate IR-level FI is vis-a-vis FI performed at the assembly code level, and prior work has presented contradictory findings. In this paper, we perform an analysis of said prior work, find an inconsistency in the FI methodology used in one study, and show that it results in a flawed comparison between IR-level and assembly-level FI. We further confirm this finding by performing a comprehensive evaluation of the accuracy of IR-level FI across a range of benchmark programs and compiler optimization levels. Our results show that IR-level FI is as accurate as assembly-level FI for silent data corruptions (SDCs) across different benchmarks and optimization levels.
Lucas Palazzi, Guanpeng Li, Bo Fang 0002, Karthik Pattabiraman
ISSRE3
2017 LetGo: A Lightweight Continuous Framework for HPC Applications Under Failures
abstract
Requirements for reliability, low power consumption, and performance place complex and conflicting demands on the design of high-performance computing (HPC) systems. Fault-tolerance techniques such as checkpoint/restart (C/R) protect HPC applications against hardware faults. These techniques, however, have non negligible overheads particularly when the fault rate exposed by the hardware is high: it is estimated that in future HPC systems, up to 60% of the computational cycles/power will be used for fault tolerance.
Bo Fang 0002, Qiang Guan, Nathan DeBardeleben, Karthik Pattabiraman, Matei Ripeanu
HPDC1
2016 ePVF: An Enhanced Program Vulnerability Factor Methodology for Cross-Layer Resilience Analysis
abstract
The Program Vulnerability Factor (PVF) has been proposed as a metric to understand the impact of hardware faults on software. The PVF is calculated by identifying the program bits required for architecturally correct execution (ACE bits). PVF, however, is conservative as it assumes that all erroneous executions are a major concern, not just those that result in silent data corruptions, and it also does not account for errorsthat are detected at runtime, i.e., lead to program crashes. A more discriminating metric can inform the choice of the appropriate resilience techniques with acceptable performance and energy overheads. This paper proposes ePVF, an enhancement of the original PVF methodology, which filters out the crash-causing bits from the ACE bits identified by the traditional PVF analysis. The ePVF methodology consists of an error propagation model that reasons about error propagation in the program, and a crash model that encapsulates the platform-specific characteristics for handling hardware exceptions. ePVF reduces the vulnerable bits estimated by the original PVF analysis by between 45% and 67% depending on the benchmark, and has high accuracy (89% recall, 92% precision) in identifying the crash-causing bits. We demonstrate the utility of ePVF by using it to inform selectiveprotection of the most SDC-prone instructions in a program.
Bo Fang 0002, Qining Lu, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi
DSN1
2016 A Systematic Methodology for Evaluating the Error Resilience of GPGPU Applications
abstract
The wide adoption of graphics processing units (GPUs) as accelerators for general-purpose applications makes the end-to-end reliability implications of their use increasingly significant. Fault injection is a widely adopted method to evaluate the resilience of applications. However, building a fault injector for general-purpose GPU applications is challenging due to their massive parallelism, which makes it difficult to achieve representativeness while being time-efficient. This paper makes four key contributions. First, it presents a fault-injection methodology to evaluate the end-to-end reliability properties of application kernels running on GPUs. Second, it introduces GPU-Qin, a fault-injection tool that uses real GPU hardware and offers a tunable and efficient balance between the representativeness and the cost of a fault-injection campaign. Third, it characterizes the error resilience characteristics of seventeen application kernels. Finally, it provides preliminary insights on correlations between the algorithmic properties of applications and their error resilience.
Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi
IEEE Trans. Parallel Distributed Syst.1
2014 GPGPUs: How to combine high computational power with high reliability
abstract
GPGPUs are used increasingly in several domains, from gaming to different kinds of computationally intensive applications. In many applications GPGPU reliability is becoming a serious issue, and several research activities are focusing on its evaluation. This paper offers an overview of some major results in the area. First, it shows and analyzes the results of some experiments assessing GPGPU reliability in HPC datacenters. Second, it provides some recent results derived from radiation experiments about the reliability of GPGPUs. Third, it describes the characteristics of an advanced fault-injection environment, allowing effective evaluation of the resiliency of applications running on GPGPUs.
Leonardo Arturo Bautista-Gomez, Franck Cappello, Luigi Carro, Nathan DeBardeleben, Bo Fang 0002, Sudhanva Gurumurthi, Karthik Pattabiraman, Paolo Rech, Matteo Sonza Reorda
DATE5
2014 Evaluating the Error Resilience of Parallel Programs
abstract
As a consequence of increasing hardware fault rates, HPC systems face significant challenges in terms of reliability. Evaluating the error resilience of HPC applications is an essential step for building efficient fault-tolerant mechanisms for these applications. In this paper, we propose a methodology to characterize the resilience of OpenMP programs using fault-injection experiments. We find that the error resilience of OpenMP applications depends on the program structure and thread model, hence, these need to be taken into account while characterizing error resilience. We also report preliminary results about the correlation between the application's error resilience and the algorithm(s) used in the application.
Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi
DSN1
2014 GPU-Qin: A methodology for evaluating the error resilience of GPGPU applications
abstract
While graphics processing units (GPUs) have gained wide adoption as accelerators for general-purpose applications (GPGPU), the end-to-end reliability implications of their use have not been quantified. Fault injection is a widely used method for evaluating the reliability of applications. However, building a fault injector for GPGPU applications is challenging due to their massive parallelism, which makes it difficult to achieve representativeness while being time-efficient. This paper makes three key contributions. First, it presents the design of a fault-injection methodology to evaluate end-to-end reliability properties of application kernels running on GPUs. Second, it introduces a fault-injection tool that uses real GPU hardware and offers a good balance between the representativeness and the efficiency of the fault injection experiments. Third, this paper characterizes the error resilience characteristics of twelve GPGPU applications.
Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi
ISPASS1