Lishan Yang 0001

dblp:182/6868-1 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-0735-8617ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 7 since 2021Security and privacy · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quantifying Performance Variability in GPU Clusters
abstract
Modern supercomputers are equipped with massive amounts of Graphics Processing Units (GPUs) to meet the growing demands of scientific computing and machine learning. However, GPUs, even of the same type, exhibit performance variability, leading to resource under-utilization and prolonged execution times. In this work, we perform a characterization study on the performance variability of NVIDIA A100 GPUs and GH200 superchips using the GEMM and STREAM micro-benchmarks and seven real-world applications across two supercomputing systems: TACC Vista (GH200) and NERSC Perlmutter (A100). Our results reveal GEMM performance variability ranging from 0.1% to 8.8%, though outliers can push deviations significantly higher. For scientific applications, the variability ranges from 1.0% to 12.2% on Vista. The variability of single-GPU GPT 4.8B and Llama 3B training is 10.4% and 8.9%, respectively. We further extend this study by training a GPT 19B model using eight GH200s in the 3D parallel configuration, observing up to 3.6% throughput variability and 8.5% slowdown in the worst case. By comparing the variability across GPU types, compute paths, data precisions, and applications, we observe that applications that use Tensor Cores and FP64 on CUDA cores show higher variability than other applications. Supercomputer users can leverage the observed variability to diagnose performance issues and enhance execution efficiency. Computing centers can design variability-aware scheduling approaches to achieve higher machine utilization without sacrificing individual application performance.
Michael Mogilevsky, Hazem Zaky, Mingkai Zheng, Lishan Yang 0001, Zhao Zhang 0007
IEEE Trans. Parallel Distributed Syst.5
2025 Safety Interventions against Adversarial Patches in an Open-Source Driver Assistance System
abstract
Drivers are becoming increasingly reliant on advanced driver assistance systems (ADAS) as autonomous driving technology becomes more popular and developed with advanced safety features to enhance road safety. However, the increasing complexity of the ADAS makes autonomous vehicles (AVs) more exposed to attacks and accidental faults. In this paper, we evaluate the resilience of a widely used ADAS against safety-critical attacks that target perception inputs. Various safety mechanisms are simulated to assess their impact on mitigating attacks and enhancing ADAS resilience. Experimental results highlight the importance of timely intervention by human drivers and automated safety mechanisms in preventing accidents in both driving and lateral directions and the need to resolve conflicts among safety interventions to enhance system resilience and reliability. [Code Available at https://doi.org/10.6084/m9.figshare.28691090]
Grant Xiao, Daehyun Lee, Lishan Yang 0001, Evgenia Smirni, Homa Alemzadeh, Xugui Zhou
DSN4
2025 FT2: First-Token-Inspired Online Fault Tolerance on Critical Layers for Generative Large Language Models
abstract
Generative Large Language Models (LLMs) are deployed on large-scale computing systems, where such tasks unavoidably suffer from soft errors, leading to quality degradation of content generated by LLMs. Enhancing LLM resilience is particularly challenging because of its complicated model architecture and tremendous size. State-of-the-art protections have limitations such as high overhead and incomplete coverage, and often require offline profiling.
Cherish Mulpuru, Roberto Gioiosa, Zhao Zhang 0007, Bo Fang 0002, Lishan Yang 0001
HPDC7
2025 Demystifying the Resilience of Large Language Model Inference: An End-to-End Perspective
abstract
Deep neural networks are known to be resilient to random bitwise faults in their parameters. However, this resilience has primarily been established through studies of classification models. The extent to which this claim holds for large-language models remains under-explored. In this work, we conduct an extensive measurement study on the impact of random bitwise faults in commercial-scale language model inference. We first expose that these language models are not truly resilient to random bit-flips. While aggregate metrics such as accuracy may suggest resilience, an in-depth inspection of the generated outputs shows significant degradation in text quality. Our analysis also shows that tasks requiring more complex reasoning suffer more from performance and quality degradation. Moreover, we extend our resilience analysis to models with augmented reasoning capabilities, such as Chain-of-Thought or Mixture of Experts architectures.
Zachary Coalson, Shiyang Chen 0004, Hang Liu 0001, Zhao Zhang 0007, Sanghyun Hong 0001, Bo Fang 0002, Lishan Yang 0001
SC8
2024 GPU Reliability Assessment: Insights Across the Abstraction Layers
abstract
Graphics Processing Units (GPUs) are widely de-ployed and utilized across various computing domains including cloud and high-performance computing. Considering its extensive usage and increasing popularity, ensuring GPU reliability is cru-cial. Software-based reliability evaluation methodologies, though fast, often neglect the complex hardware details of modern GPU designs. This oversight could lead to misleading measurements and misguided decisions regarding protection strategies. This paper breaks new ground by conducting an in-depth examination of well-established vulnerability assessment methods for modern GPU architectures, from the microarchitecture all the way to the software layers. It highlights divergences between popular software-based vulnerability evaluation methods and the ground truth cross-layer evaluation, which persist even under strong protections like triple modular redundancy. Accurate evaluation requires considering fault distribution from hardware to software. Our comprehensive measurements offer valuable insights into the accurate assessment of GPU reliability.
Lishan Yang 0001, George Papadimitriou 0001, Dimitris Sartzetakis, Adwait Jog, Evgenia Smirni, Dimitris Gizopoulos
CLUSTER1
2024 Probing Weaknesses in GPU Reliability Assessment: A Cross-Layer Approach
abstract
Due to extensive deployment and heavy usage of GPUs, ensuring the reliability of such devices is crucial. Current software-based reliability evaluation methodologies, albeit fast, often neglect the intricate hardware complexities of modern GPU designs. This oversight could result in misleading measurements and misguided decisions regarding protection strategies. This work breaks new ground by examining well-established vulnerability assessment methods for modern GPU architectures, from the microarchitecture all the way to the software layers. It highlights divergences between popular software-based vulnerability evaluation methods and the ground truth cross-layer evaluation (which, as we show, holds even when strong protection like triple modular redundancy is employed); accurate evaluation requires considering fault distribution from hardware to software. Our comprehensive measurements offer valuable insights into accurately assessing GPU reliability.
Lishan Yang 0001, George Papadimitriou 0001, Dimitris Sartzetakis, Adwait Jog, Evgenia Smirni, Dimitris Gizopoulos
ISPASS1
2024 Aspis: Lightweight Neural Network Protection Against Soft Errors
abstract
Convolutional neural networks (CNN) are incorporated into many image-based tasks across a variety of domains. Some of these are safety critical tasks such as object classification/detection and lane detection for self-driving cars. These applications have strict safety requirements and must guarantee the reliable operation of the neural networks in the presence of soft errors (i.e., transient faults) in DRAM. Standard safety mechanisms (e.g., triplication of data/computation) provide high resilience, but introduce intolerable overhead. We perform detailed characterization and propose an efficient methodology for pinpointing critical weights by using an efficient proxy, the Taylor criterion. Using this characterization, we design Aspis, an efficient software protection scheme that does selective weight hardening and offers a performance/reliability tradeoff. Aspis provides higher resilience comparing to state-of-the-art methods and is integrated into PyTorch as a fully-automated library.
Anna Schmedding, Lishan Yang 0001, Adwait Jog, Evgenia Smirni
ISSRE2
2024 Strategic Resilience Evaluation of Neural Networks Within Autonomous Vehicle Software
Anna Schmedding, Philip Schowitz, Xugui Zhou, Yiyang Lu 0001, Lishan Yang 0001, Homa Alemzadeh, Evgenia Smirni
SAFECOMP5
2023 Epidemic Spread Modeling for COVID-19 Using Cross-Fertilization of Mobility Data
abstract
We present an individual-centric model for COVID-19 spread in an urban setting. We first analyze patient and route data of infected patients from January 20, 2020, to May 31, 2020, collected by the Korean Center for Disease Control & Prevention (KCDC) and discover how infection clusters develop as a function of time. This analysis offers a statistical characterization of mobility habits and patterns of individuals at the beginning of the pandemic. While the KCDC data offer a wealth of information, they are also by their nature limited. To compensate for their limitations, we use detailed mobility data from Berlin, Germany after observing that mobility of individuals is surprisingly similar in both Berlin and Seoul. Using information from the Berlin mobility data, we cross-fertilize the KCDC Seoul data set and use it to parameterize an agent-based simulation that models the spread of the disease in an urban environment. After validating the simulation predictions with ground truth infection spread in Seoul, we study the importance of each input parameter on the prediction accuracy, compare the performance of our model to state-of-the-art approaches, and show how to use the proposed model to evaluate differentwhat-ifcounter-measure scenarios.
Anna Schmedding, Riccardo Pinciroli, Lishan Yang 0001, Evgenia Smirni
IEEE Trans. Big Data3
2023 Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models
abstract
Data center downtime typically centers around IT equipment failure. Storage devices are the most frequently failing components in data centers. We present a comparative study of hard disk drives (HDDs) and solid state drives (SSDs) that constitute the typical storage in data centers. Using six-year field data of 100,000 HDDs of different models from the same manufacturer from the Backblaze dataset and six-year field data of 30,000 SSDs of three models from a Google data center, we characterize the workload conditions that lead to failures. We illustrate that their root failure causes differ from common expectations and that they remain difficult to discern. For the case of HDDs we observe that young and old drives do not present many differences in their failures. Instead, failures may be distinguished by discriminating drives based on the time spent for head positioning. For SSDs, we observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.
Riccardo Pinciroli, Lishan Yang 0001, Jacob Alter, Evgenia Smirni
IEEE Trans. Dependable Secur. Comput.2
2022 Strategic Safety-Critical Attacks Against an Advanced Driver Assistance System
abstract
A growing number of vehicles are being transformed into semi-autonomous vehicles (Level 2 autonomy) by relying on advanced driver assistance systems (ADAS) to improve the driving experience. However, the increasing complexity and connectivity of ADAS expose the vehicles to safety-critical faults and attacks. This paper investigates the resilience of a widely-used ADAS against safety-critical attacks that target the control system at opportune times during different driving scenarios and cause accidents. Experimental results show that our proposed Context-Aware attacks can achieve an 83.4% success rate in causing hazards, 99.7% of which occur without any warnings. These results highlight the intolerance of ADAS to safety-critical attacks and the importance of timely interventions by human drivers or automated recovery mechanisms to prevent accidents.
Xugui Zhou, Anna Schmedding, Haotian Ren, Lishan Yang 0001, Philip Schowitz, Evgenia Smirni, Homa Alemzadeh
DSN4
2021 Enabling Software Resilience in GPGPU Applications via Partial Thread Protection
abstract
Graphics Processing Units (GPUs) are widely used by various applications in a broad variety of fields to accelerate their computation but remain susceptible to transient hardware faults (soft errors) that can easily compromise application output. By taking advantage of a general purpose GPU application hierarchical organization in threads, warps, and cooperative thread arrays, we propose a methodology that identifies the resilience of threads and aims to map threads with the same resilience characteristics to the same warp. This allows to engage partial replication mechanisms for error detection/correction at the warp level. By exploring 12 benchmarks (17 kernels) from 4 benchmark suites, we illustrate that threads can be remapped into reliable or unreliable warps with only 1.63% introduced overhead (on average), and then enable selective protection via replication to those groups of threads that truly need it. Furthermore, we show that thread remapping to different warps does not sacrifice application performance. We show how this remapping facilitates warp replication for error detection and/or correction and achieves average reduction of 20.61% and 27.15% execution cycles, respectively comparing to standard duplication/triplication.
Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni
ICSE1
2021 Practical Resilience Analysis of GPGPU Applications in the Presence of Single- and Multi-Bit Faults
abstract
Graphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU (GPGPU) applications is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space in the order of billions, even for some simple applications and even when considering the occurrence of just a single-bit fault. We present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience becomes practical. The key insight behind our proposed methodology stems from the fact that while GPGPU applications spawn a lot of threads, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by careful analysis. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be primarily classified based on the number of the dynamic instructions they execute. We therefore achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, a subset of loop iterations within the representative threads, and a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications. We show the effectiveness of the proposed progressive pruning technique for a single-bit model and illustrate its application to even more challenging cases with three distinct multi-bit fault models.
Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni
IEEE Trans. Computers1
2020 Simulating COVID-19 containment measures using the South Korean patient data: poster abstract
abstract
As the COVID-19 outbreak evolves around the world, the World Health Organization (WHO) and its Member States have been heavily relying on staying at home and lock down measures to control the spread of the virus. In last months, various signs showed that the COVID-19 curve was flattening, but the premature lifting of some containment measures (e.g., school closures and telecommuting) are favouring a second wave of the disease. The accurate evaluation of possible countermeasures and their well-timed revocation are therefore crucial to avoid future waves or reduce their duration. In this paper, we analyze patient and route data collected by the Korea Centers for Disease Control & Prevention (KCDC). We extract information from real-world data sets and use them to parameterize simulations and evaluate different what-if scenarios.
Lishan Yang 0001, Anna Schmedding, Riccardo Pinciroli, Evgenia Smirni
SenSys1
2018 Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications
abstract
Graphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU applications is the purpose of this study. To this end, it is imperative to explore the range of application output by injecting faults at all the potential fault sites. This problem is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space – in the order of billions even for some simple applications. In this paper, we present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience can be practical. The key insight behind our proposed methodology stems from the fact that GPGPU applications spawn a lot of threads, however, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by a careful analysis of faults across threads and instructions. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be first classified based on the number of the dynamic instructions they execute. We achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying and analyzing: a) the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, b) a subset of loop iterations within the representative threads, and c) a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications.
Bin Nie, Lishan Yang 0001, Adwait Jog, Evgenia Smirni
MICRO2
2018 G-NET: Effective GPU Sharing in NFV Systems
Kai Zhang 0006, Bingsheng He, Zeke Wang, Bei Hua, Jiayi Meng, Lishan Yang 0001
NSDI7
2018 Evaluating Scalability and Performance of a Security Management Solution in Large Virtualized Environments
abstract
Virtualized infrastructure is a key capability of modern enterprise data centers and cloud computing, enabling a more agile and dynamic IT infrastructure with fast IT provisioning, simplified, automated management, and flexible resource allocation to handle a broad set of workloads. However, at the same time, virtualization introduces new challenges, since securing virtual servers is more difficult than physical machines. HyTrust Inc. has developed an innovative security solution, called HyTrust Cloud Control (HTCC), to mitigate risks associated with virtualization and cloud technologies. HTCC is a virtual appliance deployed as a transparent proxy in front of a VMware-based virtualized environment. Since HTCC serves as a gateway to a customer virtualized environment, it is important to carefully assess its performance and scalability as well as provide its accurate resource sizing. In this work, we introduce a novel approach for accomplishing this goal. First, we describe a special framework, based on a nested virtualization technique, which enables the creation and deployment of a large scale virtualized environment (with 30,000 VMs) using a limited number of physical servers (4 servers in our experiments). Second, we introduce a design and implementation of a novel, extensible benchmark, called HT-vmbench, that allows to mimic the session-based activities of different system administrators and users in virtualized environments. The benchmark is implemented using VMware Web Service SDK. By executing HT-vmbench in the emulated large-scale virtualized environments, we can support an efficient performance assessment of management and security solutions (such as HTCC), their overhead, and provide capacity planning rules and resource sizing recommendations.
Lishan Yang 0001, Ludmila Cherkasova, Rajeev Badgujar, Jack Blancaflor, Rahul Konde, Jason Mills, Evgenia Smirni
ICPE1