EDBT 2026 Demo / reviewers in the wild / expert
Yi He 0010
dblp:65/425-10
· DBLP profile ↗
7ranked-venue papers
5as first author
5since 2021 · last 2023
0000-0001-7206-4845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Understanding Permanent Hardware Failures in Deep Learning Training Accelerator SystemsabstractHardware failures pose critical threats to deep neural network (DNN) training workloads, and the urgency of tackling this challenge (known as the Silent Data Corruption challenge in a broader context) has been raised widely by the industry. Based on industry reports, a large number of the failures observed in real systems are permanent hardware failures in logic. However, there is a very limited understanding of the effects that these failures can impose on DNN training workloads. In this paper, we present the first resilience study on this subject, focusing on deep learning (DL) training accelerator systems. We developed a fault injection framework to accurately simulate the effects of permanent faults, and conducted 100K fault injection experiments. Our results provide the fundamental understanding on how logic permanent hardware failures affect training workloads and eventually generate unexpected training outcomes. Based on this new knowledge, we developed efficient software-based detection and recovery techniques to mitigate logic permanent hardware failures that are likely to generate unexpected outcomes. Evaluation on Google Cloud TPUs shows that our techniques are effective and practical: they require 15−25 lines of code change, and introduce 0.004%−0.025% performance/energy overhead for various representative neural network models. Yi He 0010, Yanjing Li |
ETS | 1 |
| 2023 | Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsabstractDeep neural network (DNN) training workloads are increasingly susceptible to hardware failures in datacenters. For example, Google experienced "mysterious, difficult to identify problems" in their TPU training systems due to hardware failures [7]. Although these particular problems were subsequently corrected through significant efforts, they have raised the urgency of addressing the growing challenges emerging from hardware failures impacting many DNN training workloads. Yi He 0010, Mike Hutton, Robert De Gruijl, Rama Govindaraju, Nishant Patil, Yanjing Li |
ISCA | 1 |
| 2022 | Achieving Automotive Safety Requirements through Functional In-Field Self-Test for Deep Learning AcceleratorsabstractDeep learning (DL) accelerators are prominent in automotive systems, and it is essential to guarantee that these accelerators can meet the stringent automotive safety standard even in the presence of various hardware failures. In our previous work [1], we developed an efficient functional in-field self-test generation technique targeting DL accelerators, which achieves high (99.9%) stuck-at fault coverage. In this paper, we present an industry case study that extends our previous work to generate functional in-field self-tests with high transition test coverage, which is critical for screening timing degradation (e.g., caused by circuit aging). We will first present an overview of the general in-vehicle system architecture and discuss reliability/safety requirements and goals. Next, we will discuss the details of our functional in-field self-test generation technique for the transition fault model. Finally, through detailed evaluation on an industrial DL accelerator design, we will show that our approach is able to achieve extremely high transition test coverage (> 99.0%), thereby successfully achieving the reliability/safety requirements for DL accelerators in automotive applications. Moreover, the total in-field self-test time and test storage costs of our technique are low, within the required constraints. Takumi Uezono, Yi He 0010, Yanjing Li |
ITC | 2 |
| 2022 | Special Session: On the Reliability of Conventional and Quantum Neural Network HardwareabstractNeural Networks (NNs) are being extensively used in critical applications such as aerospace, healthcare, autonomous driving, and military, to name a few. Limited precision of the underlying hardware platforms, permanent and transient faults injected unintentionally as well as maliciously, and voltage/temperature fluctuations can potentially result in malfunctions in NNs with consequences ranging from substantial reduction in the network accuracy to jeopardizing the correct prediction of the network in worst cases. To alleviate such reliability concerns, this paper discusses the state-of-the-art reliability enhancement schemes that can be tailored for deep learning accelerators. We will discuss the errors associated with the hardware implementation of Deep-Learning (DL) algorithms along with their corresponding countermeasures. An in-field self-test methodology with a high test coverage is introduced, and an accurate high-level framework, so-called FIdelity, is proposed that enables the designers to evaluate DL accelerators in presence of such errors. Then, a state-of-the-art robustness-preserving training algorithm based on the Hessian Regularization is introduced. This algorithm alleviates the perturbations during inference time with negligible degradation in the accuracy of the network. Finally, Quantum Neural Networks (QNNs) and the methods to make them resilient against a variety of vulnerabilities such as fault injection, spatial and temporal variations in Qubits, and noise in QNNs are discussed. Mehdi Sadi, Yi He 0010, Yanjing Li, Mahabubul Alam, Satwik Kundu, Swaroop Ghosh, Javad Bahrami, Naghmeh Karimi |
VTS | 2 |
| 2021 | Efficient Functional In-Field Self-Test for Deep Learning AcceleratorsabstractWe present a technique that generates high-quality functional in-field self-tests specifically targeting deep learning (DL) accelerators. These functional tests can be applied in the field during normal operation of a DL accelerator, which is crucial to ensure that the safety and/or reliability requirements are met for any given application, including safety-critical applications such as self-driving cars, robotics, and more.Our technique takes advantage of special architectural characteristics and application properties to achieve high functional test coverage while incurring minimal system-level costs. Moreover, we devise different strategies for the compute units (which support computation operations) and the control units (which control data movement) because these two types of units exhibit different properties. For the compute units of a DL accelerator, we first use combinational ATPG to generate test patterns with high test coverage, which is possible because these units do not contain complex sequential logic. Next, we map the ATPG patterns to one or more equivalent deep neural networks (DNNs) that can be directly executed on the accelerator, which is possible given the well-defined dataflow/reuse algorithm of a DL accelerator. For the control units, we leverage the property that typically only one or a few fixed DNNs are deployed at a time in many application domains (e.g., self-driving cars). Thus, it is sufficient to target only the faults that can directly affect the correctness of the DNNs that are currently deployed. This is done by executing different layers of each target DNN using carefully-crafted input and weight values to maximize test coverage while minimizing test time.We apply our technique using Nvidia’s open-source accelerator as a case study to demonstrate its efficacy. Our results show that our technique achieves high test coverage. For the compute units, 99.9% single stuck-at functional test coverage is achieved. For the control units, we are able to prove that, given any target DNN, 100% coverage can be achieved for a large class of single and multiple fault models. The in-field functional self-test time is also very low, < 17 ms for various representative DNNs. These functional tests can be applied during boot-up, reset, and even concurrently with normal operation by executing DNN test programs directly on the accelerator, without requiring any test support in the hardware. Yi He 0010, Takumi Uezono, Yanjing Li |
ITC | 1 |
| 2020 | FIdelity: Efficient Resilience Analysis Framework for Deep Learning AcceleratorsabstractWe present a resilience analysis framework, called FIdelity, to accurately and quickly analyze the behavior of hardware errors in deep learning accelerators. Our framework enables resilience analysis starting from the very beginning of the design process to ensure that the reliability requirements are met, so that these accelerators can be safely deployed for a wide range of applications, including safety-critical applications such as self-driving cars.Existing resilience analysis techniques suffer from the following limitations: 1. general-purpose hardware techniques can achieve accurate results, but they require access to RTL to perform time-consuming RTL simulations, which is not feasible for early design exploration; 2. general-purpose software techniques can produce results quickly, but they are highly inaccurate; 3. techniques targeting deep learning accelerators only focus on memory errors.Our FIdelity framework overcomes these limitations. FIdelity only requires a minimal amount of high-level design information that can be obtained from architectural descriptions/block diagrams, or estimated and varied for sensitivity analysis. By leveraging unique architectural properties of deep learning accelerators, we are able to systematically model a major class of hardware errors – transient errors in logic components – in software with high fidelity. Therefore, FIdelity is both quick and accurate, and does not require access to RTL.We thoroughly validate our FIdelity framework using Nvidia’s open-source accelerator called NVDLA, which shows that the results are highly accurate – out of 60K fault injection experiments, the software fault models derived using FIdelity closely match the behaviors observed from RTL simulations. Using the validated FIdelity framework, we perform a large-scale resilience study on NVDLA, which consists of 46M fault injection experiments running various representative deep neural network applications. We report the key findings and architectural insights, which can be used to guide the design of future accelerators. Yi He 0010, Prasanna Balaprakash, Yanjing Li |
MICRO | 1 |
| 2019 | Time-Slicing Soft Error Resilience in Microprocessors for Reliable and Energy-Efficient ExecutionabstractResilience to soft errors is essential for ensuring the robustness of a computing system. In this paper, we present a new soft error resilience approach called TSSER (time-sliced soft error resilience), which enables resilience features for instructions that are most likely to cause errors only to minimize system-level energy costs while achieving high levels of resilience. Our TSSER idea (1) takes advantage of the observation that protecting a fraction of the instructions in an application already achieves most of the resilience benefits, (2) utilizes circuit-level features that allow resilience mode to be turned on/off, and (3) bridges the gap between application knowledge and circuit features by devising novel ISA and microarchitectural techniques to achieve optimized tradeoffs. Our results obtained from RTL implementation and detailed simulation show that, for various applications from the SPEC and PARSEC benchmark suites, TSSER achieves 65X reduction in SDC rate (a common metric to measure soft error resilience) while imposing 16.8% (11.3%) processor-level energy cost for in-order (out-of-order) processors. This is a significant improvement compared to existing techniques that impose 32% - 81% (18% -83 %) energy overhead. Our technique also enables flexible tradeoffs between SDC rate and system costs. Yi He 0010, Yanjing Li |
ITC | 1 |