EDBT 2026 Demo / reviewers in the wild / expert
Hengshan Yue
dblp:237/1510
· DBLP profile ↗
35ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0003-2189-8385ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptGuard: Safeguarding large vision-language models via adversarial prompt tuning
Changbao Zhou, Hengshan Yue, Ming Yan 0007, Xiaohui Wei 0002 |
Knowl. Based Syst. | 2 |
| 2026 | Sift: Channel-Wise Historical Embedding for High Efficiency Distributed Graph Neural Network Training with Accuracy GuaranteeabstractDistributed Graph Neural Network (DGNN) is a powerful tool in large-scale graph representation learning. However, high data-transfer overhead among workers in a DGNN training job confines its scalability and thus the overall performance. Vertex-wise historical embedding methods have demonstrated high potential to alleviate the problems, but still suffer from severe accuracy loss and limited performance scalability, which has been attributed to the information loss of critical channels in historical vertices and redundant information in local channels. This article explores the optimization of channel level and construct a quantitative accuracy model for channel-wise historical embedding. We propose Sift, a novel DGNN training framework, supporting channel-wise partial historical embedding with accuracy guarantee. Sift has three components: a historical embedding evaluator with channel-wise quantitative accuracy model, a sawtooth-like matrix rearrangement for accelerating message passing, and a hybrid parallel framework for overlapping communication overhead. Comprehensive experimental results show that Sift achieves near-linear parallel convergence speedup, outperforming the state-of-the-art baselines by up to 72% in total training performance and up to 21% in convergence speed. Zhewen Xu, Hongliang Li 0003, Junze Han, Hengshan Yue, Hairui Zhao 0002, Dongyuan Tian, Zijian Li 0007, Xiaohui Wei 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | FR2eRAM: A Fault-Resilient Graph Processing Paradigm in Realistic ReRAMsabstractWith the explosive growth of modern graph data, ReRAM-based graph processing paradigms are emerging as promising solutions to the “memory wall” bottleneck. However, existing paradigms often lean on overly idealized ReRAM architectures, overlooking the critical influence of various hardware faults due to the analog nature and immature fabrication processes of realistic ReRAMs. These hardware faults can lead to convergence anomalies or unacceptable output deviations (i.e., Severe Errors), undermining the reliability of ReRAM-based graph processing. While some studies enhance the reliability of ReRAM-based computation through fault-aware remapping or robust algorithm design, the unique graph execution characteristics make these efforts challenging to migrate effectively. In this work, we first develop ReGFI, a microarchitecture-level Fault Injection framework for ReRAM-based Graph processing. Unlike traditional fault injection methods that only introduce random algorithm-level faults, ReGFI precisely maps hardware faults into the microarchitectural graph execution flow to effectively characterize their effects on the execution correctness. Based on ReGFI, we propose FReRAM, a Fault-Resilient graph processing paradigm in realistic ReRAMs. Firstly, observing the fault robustness of graph vertices compared to graph edges, we reverse-map the vertex values to the non-ideal crossbar while using the edge values as inputs, for proactive SE avoidance. Then, leveraging the cell idleness in ReRAM crossbars and bit-wise reliability discrepancies of graph data, we recycle idle crossbar cells and squeeze out approximable Least Significant Bits to robust-encode the fault-sensitive bits, for further SE elimination. Experimental results exhibit that FReRAM achieves 88.84% SE reduction while incurring negligible overhead for fault-resilient ReRAM-based graph processing. Furthermore, we evaluate the effectiveness and performance of FReRAM under different hardware configurations, graph datasets, and data formats. Hengshan Yue, Nan Jiang 0013, Zongdian Li, Jiaguo Deng, Yu Huang 0013, Meikang Qiu, Xiaohui Wei 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Achieving Efficient Temporal Graph Transformation on the GPU
Linchen Yu, Jin Zhao 0003, Longlong Lin, Hengshan Yue |
APPT | 5 |
| 2025 | GraphFI: An Efficient Fault Injection Framework for Graph Processing on GPGPUsabstractAs graph tasks become pervasive in real-time and safetycritical domains (e.g., financial fraud detection and electrical power systems), it is also essential to guarantee their reliable execution beyond pursuing extraordinary performance. However, due to the neglect of consideration for graph-specific execution paradigm, existing Fault Injection (FI) reliability analysis methods typically incur inaccurate system error resilience characterization, making it challenging to provide helpful guidance for efficient and reliable graph processing paradigm design. This paper proposes GraphFI, an efficient Graph Fault Injection framework on the universal parallel tasks acceleration platform (i.e., GPGPUs). Our key insight is progressively excavating the graph-specific error propagation and effect mechanisms, thereby avoiding blind FI trials. Firstly, observing that iterations with similar active vertex set exhibit similar error behavior, we propose iteration-driven GraphFI (ID-GraphFI) to solely select representative iterations for fast error resilience profile assessment. Secondly, by detecting resilience-similarity communities in graph topology, we propose topology-driven GraphFI (TD-GraphFI) that only selects representative vertices for community overall reliability evaluation. Thirdly, by exploring the graph-specific fault monotonic property, we propose the monotonicity-driven GraphFI (MD-GraphFI) to granularly draw system severe error boundaries for predictable/unnecessary fault injection avoidance. Merging them all, GraphFI can reduce system fault site space by up to two orders of magnitude, which achieves $2.1 \sim 15.2 \times$ speedup compared to SOTA methods while providing better reliability assessment accuracy. Nan Jiang 0013, Hengshan Yue, Jingweijia Tan, Mengting Zhou, Wenda Wei, Meikang Qiu, Xiaohui Wei 0002 |
DAC | 2 |
| 2025 | OHMiner: An Overlap-centric System for Efficient Hypergraph Pattern MiningabstractHypergraph Pattern Mining (HPM) aims to identify all the instances of user-interested subhypergraphs (patterns) in hypergraphs, which has been widely used in various applications. However, existing solutions either need significant enumeration overhead because they extend subhypergraphs at the granularity of vertices, or suffer from massive redundant computations because they often need to repeatedly fetch and process the same incident hyperedges for different vertices. This paper presents an overlap-centric system named OHMiner to efficiently support HPM. OHMiner proposes an overlap-centric execution model to determine the subhypergraphs isomorphism through computing and comparing overlaps among hyperedges using set operations. This model aims to efficiently handle the vertices that collectively share the same incident hyperedges. To automatically and precisely retrieve an arbitrary pattern's overlapping semantics without performing redundant set computations, OHMiner further proposes a redundancy-free compiler, which constructs an Overlap Intersection Graph (OIG) for the pattern, optimizes the OIG, and generates an overlap-centric execution plan to guide the procedure of HPM. Moreover, OHMiner designs an overlap-centric parallel execution engine, which adopts an incremental overlap-pruned approach to fast validate candidates for HPM. Additionally, it proposes a degree-aware data store to support efficient generation of candidates. Through evaluating OHMiner on a broad range of real-world hypergraphs with various patterns, our experimental results show that OHMiner outperforms the state-of-the-art HPM system by 5.4×-22.2×. Hao Qi 0004, Ligang He, Yu Zhang 0027, Minzhi Cai, Jingxin Dai, Bingsheng He, Hai Jin 0001, Zhan Zhang 0003, Jin Zhao 0003, Hengshan Yue, Xiaofei Liao |
EuroSys | 11 |
| 2025 | Accelerating Graph Sampling in GNN Systems Through Accuracy-Aware Data Reuse
Nan Jiang 0013, Hengshan Yue, Xiaohui Wei 0002 |
ICA3PP (3) | 4 |
| 2025 | A Mapping Strategy Optimization Framework for Systolic Array Accelerators
Hengshan Yue, Haixiao Xu, Xiaohui Wei 0002 |
ICA3PP (1) | 4 |
| 2025 | GraphFT: A Lightweight Fault-tolerant Framework for Iterative Graph Processing
Xiaohui Wei 0002, Mengting Zhou, Nan Jiang 0013, Xiang Li 0197, Hengshan Yue |
WASA (3) | 6 |
| 2025 | Harnessing dynamic graph differential operators for efficient data-driven wind prediction
Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Jieyun Hao, Hengshan Yue, Changzheng Liu |
GeoInformatica | 5 |
| 2025 | ResCheckpointer: Building Program Error Resilience-Aware Checkpointing Mechanism for HPC Systems
Xiaohui Wei 0002, Shiyu Tong, Zhongao Sun, Xiang Li 0197, Hengshan Yue |
J. Comput. Sci. Technol. | 5 |
| 2024 | PGSampler: Accelerating GPU-Based Graph Sampling in GNN Systems via Workload FusionabstractGraph Neural Networks (GNNs) have demonstrated remarkable performance across various domains. Sample-based training, a practical strategy for training on large-scale graphs, often faces time-consuming graph sampling challenges. To address this, GPU-based graph sampling has been introduced, while there is still room for further efficiency improvements. Though several prior works have been proposed to accelerate the computation or memory access for GPU-based graph sampling, we show that the performance bottlenecks induced by small workload cannot be ignored. In this paper, we propose PGSampler, an efficient system for accelerating GPU-based graph sampling. First, PGSampler leverages a barrier-free execution mode to fuse workload, significantly improving the resource utilization. By altering the sampling execution mode, PGSampler also reduces the preprocessing time before kernel execution, thus accelerating the whole sampling process. Next, based on the new sampling execution mode, considering the dynamically generated nature of sampling tasks, PGSampler adopts a persistent kernel design and uses the task queue to assign tasks, achieving dynamic load balancing. Evaluations with diverse parameter settings show that PGSampler can achieve up to 2.22 × performance speedup over the state-of-the-art GNN system DGL. Xiaohui Wei 0002, Weikai Tang, Hao Qi 0004, Hengshan Yue |
CLUSTER | 4 |
| 2024 | HAp-FT: A Hybrid Approximate Fault Tolerance Framework for DNN AcceleratorabstractNowadays, as Deep Neural Networks (DNNs) become ubiquitous in mission-critical domains (e.g., automatic driving systems), ensuring the reliable execution of domain-specific DNN accelerators in the presence of hardware faults is increasingly essential. However, recent fault tolerance attempts either suffer from expensive performance overhead or are limited in scalability by specific fault modes. This paper proposes HAp-FT, a Hybrid Approximate fault Tolerance framework that integrates both proactive error alleviation and efficient error detection design philosophies. First, considering the vertical accumulation execution paradigm of the systolic array, HAp-FT proactively transfers the reliability risk of vulnerable weight to neighboring Processing Elements (PEs) for advance fault effect alleviation. Then, leveraging the inter-filter similarity characteristics, HAp-FT remaps similar filters to adjacent columns of PEs to achieve real-time intragroup approximate error detection. Experimental results exhibit that HAp-FT can recover 98.38%95.74% accuracy degradation incurred by transient/permanent hardware faults while only introducing 0.18% performance overhead, 2.69% extra area, and 2.38% energy overhead. Moreover, to satisfy the reliability requirements of various application scenarios, HAp-FT is also portable to systolic arrays under all dataflow strategies and adaptable quantized models. Xiaohui Wei 0002, Zeyu Guan, Fengyi Li, Hengshan Yue |
ICCD | 5 |
| 2024 | EFNAS: Efficient Federated Neural Architecture Search Across AIoT DevicesabstractFederated neural architecture search tailors deep learning models to accommodate varied client data in Federated Learning (FL) scenarios. However, the simultaneous optimization of multiple subnetworks leads to substantial GPU memory overhead in differentiable Neural Architecture Search (NAS) methods. Additionally, after each client searches local architecture, the conventional weighted averaging approach may result in a loss of architectural feature information and failure to capture architectural diversity, limiting model performance and expressiveness. To address these challenges, we propose EFNAS, a novel computation-efficient and aggregation-effective Federated NAS framework. Specifically, we propose Single Path Local Search (SPLS) to automatically search for the optimal network architecture with minimal complexity. SPLS uses Gumbel Softmax to continually reparameterize the probability distribution to reduce complexity. Furthermore, we propose Client-Centric Architecture Aggregation (CCAA), considering the aggregated architecture as a graph with subgraphs representing client architectures. CCAA leverages probabilistic distributions extracted from subgraphs that frequently appear across multiple clients to derive a global architecture. Guided by this global architecture, each client compares the accuracy of its locally searched architecture with the global architecture, selecting the architecture best suited for its final model architecture. Comprehensive experiments on various datasets demonstrate that EFNAS achieves excellent performance while guaranteeing efficiency during searching compared to other methods. Xiaohui Wei 0002, Guanhua Chen 0003, Hairui Zhao 0002, Hengshan Yue |
IJCNN | 6 |
| 2024 | Improving INT Inference Resilience of CNNs by Vulnerability-Aware Mixed-Precision QuantizationabstractThe application of INT quantization allows CNN model size and computation to be reduced exponentially, making extreme quantization a hot research topic. However, firstly the decrease in bitwidth increases the probability of error in critical bits, which is unacceptable for safety-critical scenarios that require high resilience of CNNs. Secondly, extreme quantization often requires retraining with high overhead. Considering the higher tolerance of errors in high-bitwidth INTs and the local resilience differences of CNNs, we use mixed precision quantization (MPQ) to trade off resilience and compression effects. Existing MPQ methods mostly focus on compression as the first goal and lack indicators of resilience or vulnerability. Based on a large number of vulnerability analyses, we propose to use the maximum bit average offset and the approximation parameter average quantization error to measure layer vulnerability. And we measure the vulnerability of the entire CNN in terms of the accumulation of layer offsets. Accordingly, we design a complete vulnerability-aware MPQ (VulnA-MPQ) framework. From the perspective of minimizing vulnerability and compression ratio, VulnA-MPQ is performed on multiple CNNs using genetic algorithm (GA) with low overhead. The CNNs only need to be fine-tuned to restore accuracy after VulnA-MPQ. Experiments demonstrate that VulnA-MPQ not only maintains accuracy and compression ratio better but also improves CNNs fault tolerance. Within the 2% accuracy threshold, the random error data tolerated by VGG16 weights and activations are improved by 33%. The number of random errors tolerated by MobileNetV3-large weights is improved by 8.87 times, and that of activations is improved by 4.64 times. Xiaohui Wei 0002, Zeyu Guan, Hengshan Yue |
ISPA | 6 |
| 2024 | CraftRGP: A Comprehensive Reliability Analysis Framework Towards ReRAM-Based Graph ProcessingabstractThe burgeoning proliferation of graph data has rendered graph processing within conventional von Neumann computing architectures increasingly inefficient. Resistive Random Access Memory (ReRAM) emerges as a promising solution to circumvent the "memory wall" bottleneck encountered in graph processing. Nevertheless, the inherent non-idealities of ReRAM devices raise reliability concerns, the ramifications of which for graph processing necessitate a comprehensive exploration. In this work, we propose CraftRGP, a Comprehensive Reliability Analysis Framework Towards ReRAM-Based Graph Processing. CraftRGP simulates the operation of graph processing on ReRAM devices at the underlying hardware level, and employs software fault injection to simulate three primary fault models on ReRAM devices: permanent faults, transient faults, and sensing faults. Based on CraftRGP, we systematically evaluate the fault resilience of ReRAM-based graph processing along the dimensions of graph applications, fault models, and hardware architectures. Finally, utilizing the observed reliability characteristics with CraftRGP, we exemplify two use cases to demonstrate that CraftRGP can guide system designers to facilitate better reliability designs. Xiaohui Wei 0002, Jiaguo Deng, Zongdian Li, Nan Jiang 0013, Hengshan Yue |
ITC-Asia | 6 |
| 2024 | DUAL-C: Building a "soft error efficient" on-the-fly compression mechanism for raw video data at edge devices
Xiaohui Wei 0002, Hengshan Yue, Nan Jiang 0013, Jianpeng Zhao 0001, Meikang Qiu |
Future Gener. Comput. Syst. | 3 |
| 2024 | ALERT: A lightweight defense mechanism for enhancing DNN robustness against T-BFA
Xiaohui Wei 0002, Yumin Yan, Nan Jiang 0013, Hengshan Yue |
J. Syst. Archit. | 5 |
| 2024 | SAR: Sharpness-Aware minimization for enhancing DNNs' Robustness against bit-flip errors
Changbao Zhou, Jiawei Du 0002, Ming Yan 0007, Hengshan Yue, Xiaohui Wei 0002, Joey Tianyi Zhou |
J. Syst. Archit. | 4 |
| 2024 | ReIPE: Recycling Idle PEs in CNN Accelerator for Vulnerable Filters Soft-Error DetectionabstractTo satisfy prohibitively massive computational requirements of current deep Convolutional Neural Networks (CNNs), CNN-specific accelerators are widely deployed in large-scale systems. Caused by high-energy neutrons and α-particle strikes, soft error may lead to catastrophic failures when CNN is deployed on high integration density accelerators. As CNNs become ubiquitous in mission-critical domains, ensuring the reliable execution of CNN accelerators in the presence of soft errors is increasingly essential. In this article, we propose to Re cycle I dle P rocessing E lements (PEs) in the CNN accelerator for vulnerable filters soft error detection (ReIPE). Considering the error-sensitivity of filters, ReIPE first carries out a filter-level gradient analysis process to replace fault injection for fast filter-wise error resilience estimation. Then, to achieve maximal reliability benefits, combining the hardware-level systolic array idleness and software-level CNN filter-wise error resilience profile, ReIPE preferentially duplicated loads the most vulnerable filters onto systolic array to recycle idle-column PEs for opportunistically redundant execution (error detection). Exploiting the data reuse properties of accelerators, ReIPE incorporates the error detection process into the original computation flow of accelerators to perform real-time error detection. Once the error is detected, ReIPE will trigger a correction round to rectify the erroneous output. Experimental results performed on LeNet-5, Cifar-10-CNN, AlexNet, ResNet-20, VGG-16, and ResNet-50 exhibit that ReIPE can cover 96.40% of errors while reducing 75.06% performance degradation and 67.79% energy consumption of baseline dual modular redundancy on average. Moreover, to satisfy the reliability requirements of various application scenarios, ReIPE is also applicable for pruned, quantized, and Transformer-based models, as well as portable to other accelerator architectures. Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan, Zeyu Guan, Nan Jiang 0013, Xinyang Zheng, Jianpeng Zhao 0001, Meikang Qiu |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | ApproxDup: Developing an Approximate Instruction Duplication Mechanism for Efficient SDC Detection in GPGPUsabstractNowadays, selective instruction duplication (SelDup) is the typical approach to detect silent data corruption (SDC) in GPGPU. However, owing to the up-to-billions fault sites of parallel GPGPU kernel functions, it usually introduces tremendous overhead to perform fault injections (FIs) for obtaining the duplication-candidate instruction set (although can be conducted in parallel). Moreover, current SelDup typically considers all SDCs severe and tends to duplicate more instructions. The nontrivial duplication overhead seriously restricts the deployment of current SelDup on resource-constrained systems (e.g., embedded GPGPUs). To address the above challenges, this article proposes an approximate instruction duplication (ApproxDup) mechanism for efficient SDC detection in GPGPUs. First, to replace the expensive FI-based duplication-candidate instructions identified method, we drive out a machine learning (ML)-based model (SDC-predictor) for instructionwise SDC proneness and severity estimation. Our key insight is that instruction type/functionality and instruction dependency set can efficaciously characterize the instructionwise SDC proneness in GPGPUs. In contrast, the instruction’s original data magnitude, fault propagation range, and error detected features can distinguish its SDC severity. Second, incorporating the concept of approximate computing, we propose ApproxDup that preferentially duplicates severe-SDC-prone instructions while relaxing the detection of minor/detectable SDCs for traditional SelDup overhead reduction. Experimental results exhibit that ApproxDup can cover 92.51% of severe SDCs while merely increasing 38% of dynamic instructions, which achieves a better tradeoff between reliability and performance compared with the state-of-the-art SelDup. Furthermore, we discuss the effectiveness of the proposed method on different ML models/applications/GPGPU architectures. Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue, Jianpeng Zhao 0001, Guangli Li, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Exploiting Complex Network-Based Clustering for Personalization-Enhanced Hierarchical Federated Edge LearningabstractFederated Learning (FL) has been extensively applied in urban environmental prediction tasks of mobile edge computing by training a global machine learning model without data sharing. However, the training of FL faces the challenges such as the poor generalization capability of a single global model over heterogeneous data and hefty communication overhead caused by the frequent model exchange between massive edge servers and remote cloud servers. To address such issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework with Complex Network clustering. HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. Based on the clustering results of PFC, anedge-mediator-cloudhierarchical architecture is proposed to realize personalization at the cluster level by Effective Hierarchical Scheduling (EHS). Furthermore, to adapt to dynamic scenarios of new edge servers joining and streaming data generation, we further extend HPFL-CN to Adaptive personalized federated learning with dynamic grouping (Ada-HPFL-CN), which can flexibly re-group edge servers and adjust mixed model weights and the model aggregation frequency adaptively. Our extensive experiments on real-world datasets demonstrate the efficacy of our framework, which outperforms state-of-the-art FL methods regarding personalization and communication efficiency performance. Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Hengshan Yue, Zhewen Xu, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Detecting SDCs in GPGPUs Through Efficient Partial Thread Redundancy
Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue |
ICA3PP (7) | 4 |
| 2023 | GLAM-SERP: Building a Graph Learning-Assisted Model for Soft Error Resilience Prediction in GPGPUs
Xiaohui Wei 0002, Jianpeng Zhao 0001, Nan Jiang 0013, Hengshan Yue |
ICA3PP (4) | 4 |
| 2023 | CFPA: Cognitive Federated Partial Adaptation for Effective PersonalizationabstractFederated Learning (FL) is a promising machine learning paradigm to train a global model from multiple clients while ensuring privacy. A key challenge in FL is the data heterogeneity that impairs the generalization of the single global model on each client. As a prevalent solution, the existing personalized FL primarily concentrates on the client-level personalization process, which may pose an overfitting problem since the local data of a single device is usually limited. Additionally, these methods lack the cognitive ability for collaborative learning among similar clients, which also restricts the performance of personalized models. To address this, we propose CFPA, a novel communication-efficient, and computation-efficient Cognitive Federated Partial Adaptation framework to boost personalization performance via the Soft-grouping Weighted Aggregation (SWA) strategy. Specifically, CFPA decouples the model into the body and the head and requires clients collaboratively train a well-performing federated pre-personalized model. Then, CFPA leverages the distribution knowledge extracted from the output feature maps of the convolutional layers in the federated pre-personalized model to identify the similarity among clients efficiently. Guided by the similarity matrix, CFPA further performs weighted federated adaption on the head of each local model, ultimately generating a personalized local model for each client. Comprehensive experiments on three benchmark datasets with various heterogeneous settings demonstrate that CFPA outperforms other state-of-the-art FL approaches. Xiaohui Wei 0002, Didi Jiao, Shiyu Tong, Zijian Li 0007, Chenghao Ren, Hengshan Yue |
MSN | 6 |
| 2023 | FASS-pruner: customizing a fine-grained CNN accelerator-aware pruning framework via intra-filter splitting and inter-filter shuffling
Xiaohui Wei 0002, Xinyang Zheng, Guangli Li, Hengshan Yue |
CCF Trans. High Perform. Comput. | 5 |
| 2023 | TC-SEPM: Characterizing soft error resilience of CNNs on Tensor Cores from program and microarchitecture perspectives
Xiaohui Wei 0002, Changbao Zhou, Hengshan Yue, Joey Tianyi Zhou |
J. Syst. Archit. | 3 |
| 2022 | MSSA-FL: High-Performance Multi-stage Semi-asynchronous Federated Learning with Non-IID Data
Xiaohui Wei 0002, Mingkai Hou, Chenghao Ren, Xiang Li 0197, Hengshan Yue |
KSEM (2) | 5 |
| 2022 | Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning
Guangli Li, Xiu Ma, Xueying Wang 0003, Hengshan Yue, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002, Jingling Xue |
J. Syst. Archit. | 4 |
| 2022 | Eff-ECC: Protecting GPGPUs Register File With a Unified Energy-Efficient ECC MechanismabstractGraphics processing units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread-level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) error correction code (ECC) for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage. In this article, we propose to leverage the error sensitivity of instructions, the duplicate characteristics of the same-named registers, and the error sensitivity of data bits to build a unified energy-efficient ECC mechanism for a GPGPUs register file (Eff-ECC), which consists of instruction-aware ECC (IA-ECC), duplication-aware ECC (DA-ECC), and bit-aware ECC (BA-ECC). Considering the error sensitivity of instructions, IA-ECC merely implements ECCs for the write registers of critical instructions. Observing the same-named registers across threads usually keeps the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Leveraging the inherent error-tolerance features of the program, BA-ECC merely protects significant bits of registers to combat the crucial error. Experimental results demonstrate that Eff-ECC tremendously reduces 86.46% energy consumption of traditional SEC-DED ECC. Hengshan Yue, Xiaohui Wei 0002, Jingweijia Tan, Nan Jiang 0013, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Detecting SDCs in GPGPUs Through an Efficient Instruction Duplication Mechanism
Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue |
KSEM | 4 |
| 2021 | G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsabstractAs GPUs become ubiquitous in large-scale general purpose HPC systems (GPGPUs), ensuring the reliable execution of such systems in the presence of soft errors is increasingly essential. To provide insights into how resilient GPU programs are toward soft errors, researchers typically rely on random Fault Injection (FI) to evaluate the tolerance of programs. However, it is expensive to obtain a statistically significant resilience profile and not suitable to identify all the error-critical fault sites of GPU programs. Hengshan Yue, Xiaohui Wei 0002, Guangli Li, Jianpeng Zhao 0001, Nan Jiang 0013, Jingweijia Tan |
SC | 1 |
| 2020 | LAD-ECC: Energy-Efficient ECC Mechanism for GPGPUs Register FileabstractGraphics Processing Units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) ECC for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage.In this paper, we propose to Leverage Approximation and Duplication characteristics of register values to build an energy-efficient ECC mechanism (LAD-ECC) in GPGPUs, which consists of APproximation-aware ECC (AP-ECC) and Duplication-Aware ECC (DA-ECC). Leveraging the inherent error tolerance features, AP-ECC merely protects significant bits of registers to combat the critical error. Observing same-named registers across threads usually keep the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Experimental results demonstrate that our LAD-ECC tremendously reduces 69.72% energy consumption of traditional SEC-DED ECC. Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan |
DATE | 2 |
| 2020 | A Novel Clustering-Based Filter Pruning Method for Efficient Deep Neural Networks
Xiaohui Wei 0002, Xiaoxian Shen, Changbao Zhou, Hengshan Yue |
ICA3PP (2) | 4 |
| 2020 | G-SEAP: Analyzing and characterizing soft-error aware approximation in GPGPUs
Xiaohui Wei 0002, Hengshan Yue, Shang Gao 0005, Ruyu Zhang, Jingweijia Tan |
Future Gener. Comput. Syst. | 2 |