Jingweijia Tan

dblp:118/8945 · DBLP profile ↗
← Back
37ranked-venue papers
19as first author
18since 2021 · last 2026
0000-0001-8256-7585ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 19 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ACOPT: Adaptive continuity-aware address translation for performance optimization of MCM-GPU architectures
Jingweijia Tan, Zhanyuntian Li, Weiren Wang, Jiashuo Wang, Kaige Yan
Future Gener. Comput. Syst.1
2026 Voltage Channel: Exploiting GPU Voltage Noise for Covert and Side Channel Attacks
Zhanyuntian Li, Jingweijia Tan, Kaige Yan, Haixiao Xu, Xiaohui Wei 0002
IEEE Trans. Computers2
2025 GraphFI: An Efficient Fault Injection Framework for Graph Processing on GPGPUs
abstract
As graph tasks become pervasive in real-time and safetycritical domains (e.g., financial fraud detection and electrical power systems), it is also essential to guarantee their reliable execution beyond pursuing extraordinary performance. However, due to the neglect of consideration for graph-specific execution paradigm, existing Fault Injection (FI) reliability analysis methods typically incur inaccurate system error resilience characterization, making it challenging to provide helpful guidance for efficient and reliable graph processing paradigm design. This paper proposes GraphFI, an efficient Graph Fault Injection framework on the universal parallel tasks acceleration platform (i.e., GPGPUs). Our key insight is progressively excavating the graph-specific error propagation and effect mechanisms, thereby avoiding blind FI trials. Firstly, observing that iterations with similar active vertex set exhibit similar error behavior, we propose iteration-driven GraphFI (ID-GraphFI) to solely select representative iterations for fast error resilience profile assessment. Secondly, by detecting resilience-similarity communities in graph topology, we propose topology-driven GraphFI (TD-GraphFI) that only selects representative vertices for community overall reliability evaluation. Thirdly, by exploring the graph-specific fault monotonic property, we propose the monotonicity-driven GraphFI (MD-GraphFI) to granularly draw system severe error boundaries for predictable/unnecessary fault injection avoidance. Merging them all, GraphFI can reduce system fault site space by up to two orders of magnitude, which achieves $2.1 \sim 15.2 \times$ speedup compared to SOTA methods while providing better reliability assessment accuracy.
Nan Jiang 0013, Hengshan Yue, Jingweijia Tan, Mengting Zhou, Wenda Wei, Meikang Qiu, Xiaohui Wei 0002
DAC3
2025 A Novel Frequency-Spatial Domain Aware Network for Fast Thermal Prediction in 2.5D ICs
abstract
In the post-Moore era, 2.5D chiplet-based ICs present significant challenges in thermal management due to increased power density and thermal hotspots. Neural network-based thermal prediction models can perform real-time predictions for many unseen new designs. However, existing CNN-based and GCN-based methods cannot effectively capture the global thermal features, especially for high-frequency components, hindering pre-diction accuracy enhancement. In this paper, we propose a novel frequency-spatial dual domain aware prediction network (FSA-Heat) for fast and high-accuracy thermal prediction in 2.5D ICs. It integrates high-to-low frequency and spatial domain encoder (FSTE) module with frequency domain cross-scale interaction module (FCIFormer) to achieve high-to-low frequency and global-to-local thermal dissipation feature extraction. Additionally, a frequency-spatial hybrid loss (FSL) is designed to effectively attenuate high-frequency thermal gradient noise and spatial mis-alignments. The experimental results show that the performance enhancements offered by our proposed method are substantial, outperforming the newly-proposed 2.5D method, GCN+PNA, by considerable margins (over 99% RMSE reduction, 4.23X inference time speedup). Moreover, extensive experiments demonstrate that FSA-Heat also exhibits robust generalization capabilities.
Dekang Zhang, Dan Niu, Zhou Jin 0001, Yichao Dong, Jingweijia Tan, Changyin Sun 0001
DATE5
2025 Evaluating GPU's Instruction-Level Error Characteristics Under Low Supply Voltages
abstract
Supply voltage underscaling has been an effective approach to improve the energy-efficiency of modern high-performance processors, such as GPUs. However, energy efficiency and reliability are two sides of a trade-off. Undervolting will inevitably undermine reliability, since it reduces chip manufacturers’ voltage guardbands that is designed to ensure correct operations under worst-case scenarios. To achieve optimal energy efficiency while maintaining enough reliability, it is necessary to deeply understand the error characteristics caused by undervolting. Unlike previous works which focus mostly on program level, we perform the first comprehensive instruction-level voltage margin and error characteristics evaluation for GPU architectures. We systematically measure the error probability and patterns of GPU instructions during undervolting. Then, we also analyze the impact of locations (SMs, threads, and bits) and operand data values on the error characteristics. Based on our observations, we reduce the voltage to the minimum safe limit for different instructions which achieves 18.37% energy saving, and we further propose an error detection strategy which reduces the performance and energy overhead by 14.8% with negligible 0.01% degradation for error detection rate.
Jingweijia Tan, Jiashuo Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001
IEEE Trans. Computers1
2025 GEREM: Fast and Precise Error Resilience Assessment for GPU Microarchitectures
abstract
GPUs are widely used hardware acceleration platforms in many areas due to their great computational throughput. In the meanwhile, GPUs are vulnerable to transient hardware faults in the post-Moore era. Analyzing the error resilience of GPUs are critical for both hardware and software. Statistical fault injection approaches are commonly used for error resilience analysis, which are highly accurate but very time consuming. In this work, we propose GEREM, a first framework to speed up fault injection process so as to estimate the error resilience of GPU microarchitectures swiftly and precisely. We find early fault behaviors can be used to accurately predict the final outcomes of program execution. Based on this observation, we categorize the early behaviors of hardware faults into GPU Early Fault Manifestation models (EFMs). For data structures, EFMs are early propagation characteristics of faults, while for pipeline instructions, EFMs are heuristic properties of several instruction contexts. We further observe that EFMs are determined by static microarchitecture states, so we can capture them without actually simulating the program execution process under fault injections. Leveraging these observations, our GEREM framework first profiles the microarchitectural states related for EFMs at one time. It then injects faults into the profiled traces to immediately generate EFMs. For data storage structures, EFMs are directly used to predict final fault outcomes, while for pipeline instructions, machine learning is used for prediction. Evaluation results show GEREM precisely assesses the error resilience of GPU microarchitecture structures with$237\times$speedup on average comparing with traditional fault injections.
Jingweijia Tan, An Zhong, Kaige Yan, Xiaohui Wei 0002, Guanpeng Li
IEEE Trans. Parallel Distributed Syst.1
2024 HSAS: Efficient task scheduling for large scale heterogeneous systolic array accelerator cluster
Kaige Yan, Yanshuang Song, Jingweijia Tan, Xiaohui Wei 0002, Xin Fu 0001
Future Gener. Comput. Syst.4
2024 Efficient one-shot Neural Architecture Search with progressive choice freezing evolutionary search
Qiyu Wan, Yu Wen 0003, Mingsong Chen 0001, Jingweijia Tan, Kaige Yan, Xin Fu 0001
Neurocomputing6
2024 ReIPE: Recycling Idle PEs in CNN Accelerator for Vulnerable Filters Soft-Error Detection
abstract
To satisfy prohibitively massive computational requirements of current deep Convolutional Neural Networks (CNNs), CNN-specific accelerators are widely deployed in large-scale systems. Caused by high-energy neutrons and α-particle strikes, soft error may lead to catastrophic failures when CNN is deployed on high integration density accelerators. As CNNs become ubiquitous in mission-critical domains, ensuring the reliable execution of CNN accelerators in the presence of soft errors is increasingly essential. In this article, we propose to Re cycle I dle P rocessing E lements (PEs) in the CNN accelerator for vulnerable filters soft error detection (ReIPE). Considering the error-sensitivity of filters, ReIPE first carries out a filter-level gradient analysis process to replace fault injection for fast filter-wise error resilience estimation. Then, to achieve maximal reliability benefits, combining the hardware-level systolic array idleness and software-level CNN filter-wise error resilience profile, ReIPE preferentially duplicated loads the most vulnerable filters onto systolic array to recycle idle-column PEs for opportunistically redundant execution (error detection). Exploiting the data reuse properties of accelerators, ReIPE incorporates the error detection process into the original computation flow of accelerators to perform real-time error detection. Once the error is detected, ReIPE will trigger a correction round to rectify the erroneous output. Experimental results performed on LeNet-5, Cifar-10-CNN, AlexNet, ResNet-20, VGG-16, and ResNet-50 exhibit that ReIPE can cover 96.40% of errors while reducing 75.06% performance degradation and 67.79% energy consumption of baseline dual modular redundancy on average. Moreover, to satisfy the reliability requirements of various application scenarios, ReIPE is also applicable for pruned, quantized, and Transformer-based models, as well as portable to other accelerator architectures.
Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan, Zeyu Guan, Nan Jiang 0013, Xinyang Zheng, Jianpeng Zhao 0001, Meikang Qiu
ACM Trans. Archit. Code Optim.4
2023 Saca-FI: A microarchitecture-level fault injection framework for reliability analysis of systolic array based CNN accelerator
Jingweijia Tan, Qixiang Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001
Future Gener. Comput. Syst.1
2023 Saca-AVF: A Quantitative Approach to Analyze the Architectural Vulnerability Factors of CNN Accelerators
abstract
Convolutional neural network (CNN) accelerators are widely used in artificial intelligence applications, such as image recognitions, due to their superior computing performance. However, as manufacturing technology scales down, the shrinking of chip size and the increasing of integration density make CNN accelerators vulnerable to soft errors, which are key factors affecting the reliability of integrated circuits. For emerging CNN applications where reliability is critical (such as self-driving cars), visible faults in applications’ outputs caused by soft errors may lead to catastrophic consequences. Therefore, it is important to consider reliability into CNN accelerators’ architecture design. Recent modeling and fault injection approaches analyze the reliability features of CNN models, but none evaluate the soft error reliability of CNN accelerators from architecture level. The Architectural Vulnerability Factor (AVF) indicates the probability that a soft error results in faulty computation outcomes. The Architecturally Correct Execution (ACE) method analyzes AVF of structures by identifying ACE bits and calculating their residency time. Based on the traditional ACE-based AVF analysis model, we propose a quantitative approach saca-AVF to analyze the AVF of systolic array based CNN accelerators that perform homogeneous matrix arithmetic operations instead of various kinds of instructions. We explore the reliability features of CNN accelerators under different design choices at different levels leveraging saca-AVF. We also observe several reliability characteristics of CNN accelerator architecture. We believe our proposed saca-AVF and the observations we made are able to guide designers to discover the reliability hotspots of CNN accelerators and build highly reliable CNN accelerators.
Jingweijia Tan, Liqi Ping, Qixiang Wang, Kaige Yan
IEEE Trans. Computers1
2023 MCM-GPU Voltage Noise Characterization and Architecture-Level Mitigation
abstract
Due to manufacturing process and yield constraints, scaling GPU performance via increasing chip area becomes difficult. In the meanwhile, the demand for high computational throughput is increasing for high performance computing applications. As an alternative, multichip module GPU (MCM-GPU) achieves performance scalability via integrating multiple GPU chip modules (GPMs) on the same package. However, large MCM-GPU systems are susceptible to voltage noise effects, which cause voltage instability during program execution and result in energy inefficiency. In this work, we first model and analyze the voltage noise of MCM-GPUs at architecture level in detail. We characterize the voltage noise distributions of MCM-GPUs at different levels and under various design parameters. We further propose two architecture level voltage noise mitigation approaches, including GPM aware mitigation (GAM) and droop magnitude aware smoothing (DMAS), that leverage the voltage noise characteristics of MCM-GPUs. Evaluation shows both techniques are effective in reducing the voltage droop magnitudes and achieve good energy savings with negligible performance degradation.
Jingweijia Tan, Weiren Wang, Kaige Yan, Xiaohui Wei 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Improving the Performance of CNN Accelerator Architecture under the Impact of Process Variations
abstract
Convolutional neural network (CNN) accelerators are popular specialized platforms for efficient CNN processing. As semiconductor manufacturing technology scales down to nano scale, process variation dramatically affects the chip’s quality. Process variation causes delay variation within the chip due to transistor parameter differences. CNN accelerators adopt a large number of processing elements (PEs) for parallel computing, which are highly susceptible to process variation effects. Fast CNN processing desires consistent performance among PEs; otherwise the processing speed is limited by the slowest PE within the chip. In this work, we first quantitatively model and analyze the impact of process variation on CNN accelerators’ operating frequency. We further analyze the utilization of CNN accelerators and the characteristics of CNN models. We then leverage the PE underutilization to propose a sub-matrix reformation mechanism and leverage the pixel similarity of images to propose a weight transfer technique. Both techniques are able to tolerate the low-frequency PEs and achieve performance improvement at chip level. Furthermore, a novel resilience-aware mapping technique that exploits the diversity in the importance of weights is also proposed to improve the performance. Evaluation results show that our techniques are able to achieve significant processing speed improvement with negligible accuracy loss.
Jingweijia Tan, Weiren Wang, Maodi Ma, Xiaohui Wei 0002, Kaige Yan
ACM Trans. Design Autom. Electr. Syst.1
2022 MG-Voltage: Characterizing and Mitigating Voltage Noise in MCM-GPU Architectures
abstract
Multi-chip-module (MCM) based package-level integration, a technology of forming large logical processing systems with small chiplets, is promising in the post-Moore era. It achieves great performance scalability and yield improvement for GPUs, which are highly desired by emerging high performance computing applications. However, critical problems, such as energy overhead and voltage instability, introduced by MCM-GPU designs, must be carefully resolved. Modern processors adopt large voltage guardbands to tolerate worst-case variations, such as voltage noises, which is pessimistic for most of the execution time. Exploiting supply voltage guardbands are effective to improve the processor’s energy efficiency. Even though voltage noise is studied for traditional monolithic GPUs, its effect on MCM-GPUs is unexplored.In this work, we characterize voltage noise effects on MCM-GPU architectures in detail. We analyze the maximum voltage droops of different applications, the temporal and spatial voltage droop distributions, and location-aware power variations. Based on the observations of staggered voltage droops across GPU modules (GPMs), we propose a GPM-aware voltage noise mitigation technique. Evaluation shows our technique is able to reduce the worst-case voltage droop by 29.4% on average, and up to 38.1% with negligible performance degradation.
Jingweijia Tan, Kaige Yan
ICCD1
2022 Eff-ECC: Protecting GPGPUs Register File With a Unified Energy-Efficient ECC Mechanism
abstract
Graphics processing units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread-level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) error correction code (ECC) for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage. In this article, we propose to leverage the error sensitivity of instructions, the duplicate characteristics of the same-named registers, and the error sensitivity of data bits to build a unified energy-efficient ECC mechanism for a GPGPUs register file (Eff-ECC), which consists of instruction-aware ECC (IA-ECC), duplication-aware ECC (DA-ECC), and bit-aware ECC (BA-ECC). Considering the error sensitivity of instructions, IA-ECC merely implements ECCs for the write registers of critical instructions. Observing the same-named registers across threads usually keeps the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Leveraging the inherent error-tolerance features of the program, BA-ECC merely protects significant bits of registers to combat the crucial error. Experimental results demonstrate that Eff-ECC tremendously reduces 86.46% energy consumption of traditional SEC-DED ECC.
Hengshan Yue, Xiaohui Wei 0002, Jingweijia Tan, Nan Jiang 0013, Meikang Qiu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Towards Fine-Grained Online Adaptive Approximation Control for Dense SLAM on Embedded GPUs
abstract
Dense SLAM is an important application on an embedded environment. However, embedded platforms usually fail to provide enough computation resources for high-accuracy real-time dense SLAM, even with high-parallelism architecture such as GPUs. To tackle this problem, one solution is to design proper approximation techniques for dense SLAM on embedded GPUs. In this work, we propose two novel approximation techniques, critical data identification and redundant branch elimination. We also analyze the error characteristics of the other two techniques—loop skipping and thread approximation. Then, we propose SLaPP, an online adaptive approximation controller, which aims to control the error to be under an acceptable threshold. The evaluation shows SLaPP can achieve 2.0× performance speedup and 30% energy saving on average compared to the case without approximation.
Tiancong Bu, Kaige Yan, Jingweijia Tan
ACM Trans. Design Autom. Electr. Syst.3
2021 AERO: Towards Energy-Efficient Autonomous Flight in MAVs Using Approximate Execution
abstract
Micro aerial vehicles (MAVs) are popular in many intelligent robot areas nowadays. As an battery-powered vehicle, the energy-efficiency of MAVs becomes the bottleneck for its wide adoption. This issue exacerbates for autonomous flight MAVs, since inefficient action executions may results in flight mission failure. In this work, we address the energy-inefficiency of deep Q-network (DQN) based autonomous flight of MAVs via approximate execution. We first investigate data sensing during flight process and observe two consecutive steps tend to perceive consistent data. We further analyze the decision making characteristics of DQN algorithm, and find the margins between maximum and the second largest Q-value in different steps are usually symmetrically distributed between action changes. Leveraging these two characteristics, we propose to Approximately Execute the autonomous flight pROcessing of MAVs for energy-efficiency improvement (AERO). AERO is composed of two techniques of AERO-S and AERO-AP. AERO-S uses the sensed data from the previous step to approximately infer the current sensing data. AERO-AP uses the symmetry of margins between the maximum and the second largest Q-value to skip the decision makings of some steps. Our evaluations show in environments with no obstacle, AERO achieves relative improvement of time (RIT) of 24.28% and relative improvement of energy (RIE) of 18.75% with no reduction of success rate. For environments with obstacles, AERO achieves 19.11% RIT and 12.72% RIE with no reduction of success rate.
Jingweijia Tan, Kaige Yan
ASAP2
2021 G-SEPM: building an accurate and efficient soft error prediction model for GPGPUs
abstract
As GPUs become ubiquitous in large-scale general purpose HPC systems (GPGPUs), ensuring the reliable execution of such systems in the presence of soft errors is increasingly essential. To provide insights into how resilient GPU programs are toward soft errors, researchers typically rely on random Fault Injection (FI) to evaluate the tolerance of programs. However, it is expensive to obtain a statistically significant resilience profile and not suitable to identify all the error-critical fault sites of GPU programs.
Hengshan Yue, Xiaohui Wei 0002, Guangli Li, Jianpeng Zhao 0001, Nan Jiang 0013, Jingweijia Tan
SC6
2020 LAD-ECC: Energy-Efficient ECC Mechanism for GPGPUs Register File
abstract
Graphics Processing Units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) ECC for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage.In this paper, we propose to Leverage Approximation and Duplication characteristics of register values to build an energy-efficient ECC mechanism (LAD-ECC) in GPGPUs, which consists of APproximation-aware ECC (AP-ECC) and Duplication-Aware ECC (DA-ECC). Leveraging the inherent error tolerance features, AP-ECC merely protects significant bits of registers to combat the critical error. Observing same-named registers across threads usually keep the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Experimental results demonstrate that our LAD-ECC tremendously reduces 69.72% energy consumption of traditional SEC-DED ECC.
Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan
DATE3
2020 SERN: Modeling and Analyzing the Soft Error Reliability of Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) are popular in artificial intelligence areas due to their high accuracy. Meanwhile, as manufacturing process technology scales, the probability of soft errors occurrence in computer systems increases, which causes reliability challenges for CNNs. For emerging safety and reliability critical CNN applications, such as autonomous vehicles, soft errors may cause catastrophic consequences. Thus it is important to analyze the vulnerability characteristics of CNNs. A common approach to analyze CNNs' reliability is fault injection, which requires a lot of computation resources and is time consuming. Thus a more efficient method is desired. In this work, we propose an analytical model named SERN to analyze the soft error reliability of CNNs, which requires only a small number of parameters in CNN models. Validation on several common CNN models shows SERN can efficiently and accurately perform soft error reliability analysis for CNNs. We also observe that the reliability of CNNs depend on data types, values, the sign of data and types of layers. Take advantage of these observations, we propose to protect vulnerable bits through ECC and protect error-prone layers through redundancy.
Liqi Ping, Jingweijia Tan, Kaige Yan
ACM Great Lakes Symposium on VLSI2
2020 G-SEAP: Analyzing and characterizing soft-error aware approximation in GPGPUs
Xiaohui Wei 0002, Hengshan Yue, Shang Gao 0005, Ruyu Zhang, Jingweijia Tan
Future Gener. Comput. Syst.6
2020 Toward Customized Hybrid Fuel-Cell and Battery-powered Mobile Device for Individual Users
abstract
Rapidly evolving technologies and applications of mobile devices inevitably increase the power demands on the battery. However, the development of batteries can hardly keep pace with the fast-growing demands, leading to short battery life, which becomes the top complaints from customers. In this article, we investigate a novel energy supply technology, fuel cell (FC), and leverage its advantages of providing long-term energy storage to build a hybrid FC-battery power system. Therefore, mobile device operation time is dramatically extended, and users are no longer bothered by battery recharging. We examine real-world smartphone usage data and find that a naive hybrid power system cannot meet many users’ highly diversified power demands. We thus propose an OS-level power management policy that reduces the device power consumption for each power peak to solve this mismatch. This technique trades the quality-of-service (QoS) for a larger FC ratio in the system and thus much longer device operation time. We further observe that the user’s personality largely determines his/her satisfaction with the QoS degradation and the operation time extension. Thus, applying a hybrid system with fixed configuration (i.e., peak throttling level coupled with corresponding FC/battery ratio) fails to satisfy every user. We then explore customized hybrid system configuration based on each individual user’s personality to deliver the optimal satisfaction for him/her. The experimental results show that our personality-aware hybrid FC-battery solution can achieve 4× longer operation time and 25% higher satisfaction score compared to the common setting for state-of-the-art mobile devices.
Kaige Yan, Jingweijia Tan, Longjun Liu, Xingyao Zhang 0002, Stanko R. Brankovic, Jinghong Chen, Xin Fu 0001
ACM Trans. Embed. Comput. Syst.2
2020 Energy-Efficient GPU L2 Cache Design Using Instruction-Level Data Locality Similarity
abstract
This article presents a novel energy-efficient cache design for massively parallel, throughput-oriented architectures like GPUs. Unlike L1 data cache on modern GPUs, L2 cache shared by all of the streaming multiprocessors is not the primary performance bottleneck, but it does consume a large amount of chip energy. We observe that L2 cache is significantly underutilized by spending 95.6% of the time storing useless data. If such “dead time” on L2 is identified and reduced, L2’s energy efficiency can be drastically improved. Fortunately, we discover that the SIMT programming model of GPUs provides a unique feature among threads: instruction-level data locality similarity, which can be used to accurately predict the data re-reference counts at L2 cache block level. We propose a simple design that leverages this Lo cality S imilarity to build an energy-efficient GPU L2 Cache , named LoSCache . Specifically, LoSCache uses the data locality information from a small group of cooperative thread arrays to dynamically predict the L2-level data re-reference counts of the remaining cooperative thread arrays. After that, specific L2 cache lines can be powered off if they are predicted to be “dead” after certain accesses. Experimental results on a wide range of applications demonstrate that our proposed design can significantly reduce the L2 cache energy by an average of 64% with only 0.5% performance loss. In addition, LoSCache is cost effective, independent of the scheduling policies, and compatible with the state-of-the-art L1 cache designs for additional energy savings.
Jingweijia Tan, Kaige Yan, Shuaiwen Song, Xin Fu 0001
ACM Trans. Design Autom. Electr. Syst.1
2019 LoSCache: Leveraging Locality Similarity to Build Energy-Efficient GPU L2 Cache
abstract
This paper presents a novel energy-efficient cache design for massively parallel, throughput-oriented architectures like GPUs. Unlike L1 data cache on modern GPUs, L2 cache shared by all the streaming multiprocessors is not the primary performance bottleneck but it does consume a large amount of chip energy. We observe that L2 cache is significantly under-utilized by spending 95.6% of the time storing useless data. If such "dead time" on L2 is identified and reduced, L2's energy efficiency can be drastically improved. Fortunately, we discover that the SIMT programming model of GPUs provides a unique feature among threads: instruction-level data locality similarity, which can be used to accurately predict the data re-reference counts at L2 cache block level. We propose a simple design that leverages this Locality Similarity to build an energy-efficient GPU L2 Cache, named LoSCache. Specifically, LoSCache uses the data locality information from a small group of CTAs to dynamically predict the L2-level data re-reference counts of the remaining CTAs. After that, specific L2 cache lines can be powered off if they are predicted to be "dead" after certain accesses. Experimental results on a wide range of applications demonstrate that our proposed design can significantly reduce the L2 cache energy by an average of 64% with only 0.5% performance loss.
Jingweijia Tan, Kaige Yan, Shuaiwen Song, Xin Fu 0001
DATE1
2019 Process Variation Mitigation on Convolutional Neural Network Accelerator Architecture
abstract
Convolutional Neural Network (CNN) accelerators are popular specialized platforms for efficient CNN processing. As semiconductor manufacturing technology scales down to nano scale, process variation dramatically affects the chip's quality. Process variation causes delay variation within the chip due to transistor parameter differences. CNN accelerators adopt a large number of Processing Elements (PEs) for parallel computing, which are highly susceptible to process variation effects. Fast CNN processing desires consistent performance among PEs, otherwise the processing speed is limited by the slowest PE within the chip. In this work, we first quantitatively model and analyze the impact of process variation on CNN accelerator's operating frequency. We further analyze the utilization of CNN accelerator and the characteristics of CNN models. We then leverage the PE underutilization to propose a sub-matrix reformation mechanism and leverage the pixel similarity of images to propose a weight transfer technique. Both techniques are able to tolerate the low-frequency PEs, and achieve performance improvement at chip level. Evaluation results show our techniques are able to achieve significant processing speed improvement with negligible accuracy loss.
Maodi Ma, Jingweijia Tan, Xiaohui Wei 0002, Kaige Yan
ICCD2
2019 Improving energy efficiency of mobile devices by characterizing and exploring user behaviors
Kaige Yan, Jingweijia Tan, Xin Fu 0001
J. Syst. Archit.2
2019 Bridging mobile device configuration to the user experience under budget constraint
Kaige Yan, Jingweijia Tan, Xin Fu 0001
Pervasive Mob. Comput.2
2019 Efficiently Managing the Impact of Hardware Variability on GPUs' Streaming Processors
abstract
Graphics Processing Units (GPUs) are widely used in general-purpose high-performance computing fields due to their highly parallel architecture. In recent years, a new era with the nanometer scale integrated circuit manufacture process has come. As a consequence, GPUs’ computation capability gets even stronger. However, as process technology scales down, hardware variability, e.g., process variations (PVs) and negative bias temperature instability (NBTI), has a higher impact on the chip quality. The parallelism of GPU desires high consistency of hardware units on chip; otherwise, the worst unit will inevitably become the bottleneck. So the hardware variability becomes a pressing concern to further improve GPUs’ performance and lifetime, not only in integrated circuit fabrication, but more in GPU architecture design. Streaming Processors (SPs) are the key units in GPUs, which perform most of parallel computing operations. Therefore, in this work, we focus on mitigating the impact of hardware variability in GPU SPs. We first model and analyze SPs’ performance variations under hardware variability. Then, we observe that both PV and NBTI have a large impact on SPs’ performance. We further observe unbalanced SP utilization, e.g., some SPs are idle when others are active, during program execution. Leveraging this observation, we propose a Hardware Variability-aware SPs’ Management policy (HVSM), which dynamically dispatches computation in appropriate SPs to balance the utilizations. In addition, we find that a large portion of compute operations are duplicate. We also propose an Operation Compression (OC) technique to minimize the unnecessary computations to further mitigate the hardware variability effects. Our experimental results show the combined HVSM and OC technique effectively reduces the impact of hardware variability, which can translate to 37% performance improvement or 18.3% lifetime extension for a GPU chip.
Jingweijia Tan, Kaige Yan
ACM Trans. Design Autom. Electr. Syst.1
2018 HVSM: Hardware-variability aware streaming processors' management policy in GPUs
abstract
GPUs are widely used in general-purpose high performance computing field due to their highly parallel architecture. In recent years, a new era with nanometer scale integrated circuit manufacture process has come, as a consequence, GPUs' computation capability gets even stronger. However, as process technology scales down, hardware variability, e.g., process variations (PVs) and negative bias temperature instability (NBTI), has a higher impact on the chip quality. The parallelism of GPU desires high consistency of hardware units on chip, otherwise, the worst unit will inevitably become the bottleneck. So the hardware variability becomes a pressing concern to further improve GPUs' performance and lifetime, not only in integrated circuit fabrication, but more in GPU architecture design. Streaming Processors (SPs) are the key units in GPUs, which perform most of parallel computing operations. Therefore, in this work, we focus on mitigating the impact of hardware variability in GPU SPs. We first model and analyze SPs' performance variations under hardware variability. Then, we observe that both PV and NBTI have large impact on SP's performance. We further observe unbalanced SP utilization, e.g., some SPs are idle when others are active, during program execution. Leveraging both observations, we propose a Hardware Variability-aware SPs' Management policy (HVSM), which dynamically prioritizes the fast SPs, regroups SPs in a two-level granularity and dispatches computation in appropriate SPs. Our experimental results show HVSM effectively reduces the impact of hardware variability, which can translate to 28% performance improvement or 14.4% lifetime extension for a GPU chip.
Jingweijia Tan, Kaige Yan
DATE1
2016 Combating the Reliability Challenge of GPU Register File at Low Supply Voltage
abstract
Supply voltage reduction is an effective approach to significantly reduce GPU energy consumption. As the largest on-chip storage structure, the GPU register file becomes the reliability hotspot that prevents further supply voltage reduction below the safe limit ($V_{min}$) due to process variation effects. This work addresses the reliability challenge of the GPU register file at low supply voltages, which is an essential first step for aggressive supply voltage reduction of the entire GPU chip. To better understand the reliability issues posed by undervolting and its energy-saving potential, we first rigorously model and analyze the process variation impact on the GPU register file at different voltages. By further analyzing the GPU architecture, we make a key observation that the time GPU registers contain useless data (i.e., dead time) is long, providing a unique opportunity to enhance register reliability. We then propose GR-Guard, an architectural solution that leverages long register dead time to enable reliable operations from unreliable register file at low voltages. GR-Guard is both effective and low-cost, and does not affect normal (i.e., non-faulty) register accesses. Experimental results show that for a 28nm baseline GPU under aggressive voltage reduction, GR-Guard can maintain the register file reliability with less than 2\% overall performance degradation, while achieving an average of 31% energy reduction across various applications.
Jingweijia Tan, Shuaiwen Song, Kaige Yan, Xin Fu 0001, Andrés Márquez 0001, Darren J. Kerbyson
PACT1
2016 Redefining QoS and customizing the power management policy to satisfy individual mobile users
abstract
Delivering an excellent use experience to the customers is the top challenge faced by today's mobile device designers and producers. There have been multiple studies on achieving the good trade-offs between QoS and energy to enhance the user experience, however, they generally lack a comprehensive and accurate understanding of QoS, and ignore the fact that each individual user has his/her own preference between QoS and energy. In this study, we overcome these two drawbacks and propose a customized power management policy that dynamically configures the mobile platform to achieve the user-specific optimal QoS and energy trade-offs and hence, satisfy each individual mobile user. We first introduce a novel and comprehensive definition of QoS, and propose the accurate QoS measurement and management methodologies. We then observe that user's personality greatly determines his/her preferences between QoS and energy, and propose an online personality-guided user satisfaction prediction model based on the QoS and energy, guided by the user personality inferred from his/her device usage history. Our validation proves our model can achieve very high prediction accuracy. Finally, we propose our customized power management policy based on the prediction model for individual users. The experiment results show that our technique can improve the user experience by around 36% compared with the state-of-the-art power management policies.
Kaige Yan, Xingyao Zhang 0002, Jingweijia Tan, Xin Fu 0001
MICRO3
2016 Exploring Soft-Error Robust and Energy-Efficient Register File in GPGPUs using Resistive Memory
abstract
The increasing adoption of graphics processing units (GPUs) for high-performance computing raises the reliability challenge, which is generally ignored in traditional GPUs. GPUs usually support thousands of parallel threads and require a sizable register file. Such large register file is highly susceptible to soft errors and power-hungry. Although ECC has been adopted to register file in modern GPUs, it causes considerable power overhead, which further increases the power stress. Thus, an energy-efficient soft-error protection mechanism is more desirable. Besides its extremely low leakage power consumption, resistive memory (e.g., spin-transfer torque RAM) is also immune to the radiation induced soft errors due to its magnetic field based storage. In this article, we propose to LEverage reSistive memory to enhance the Soft-error robustness and reduce the power consumption (LESS) of registers in the General-Purpose computing on GPUs (GPGPUs). Since resistive memory experiences longer write latency compared to SRAM, we explore the unique characteristics of GPGPU applications to obtain the win-win gains: achieving the near-full soft-error protection for the register file, and meanwhile substantially reducing the energy consumption with negligible performance degradation. Our experimental results show that LESS is able to mitigate the registers soft-error vulnerability by 86% and achieve 61% energy savings with negligible (e.g., 1%) performance degradation.
Jingweijia Tan, Zhi Li 0016, Mingsong Chen 0001, Xin Fu 0001
ACM Trans. Design Autom. Electr. Syst.1
2016 Mitigating the Impact of Hardware Variability for GPGPUs Register File
abstract
As technology keeps scaling down, hardware variability, such as process variations (PV) and negative bias temperature instability (NBTI), emerges as a growing challenge in the modern GPGPUs (general-purpose computing on graphics processing units). PV induces significant delay variations statically, while NBTI dynamically slows down the GPGPUs. Each computing core (i.e., streaming multiprocessor) in GPGPUs supports thousands of simultaneously active threads, and requires a large register file. Such a sizable register file is very sensitive to the hardware variability, and becomes one of the major units in determining the core frequency. In this study, we propose a set of techniques that mitigate both the PV and NBTI impacts on GPGPUs register file. In order to mitigate the susceptibility to PV, we first develop a novel mechanism that classifies registers into fast and slow categories in the highly-banked register architecture to maximize the frequency improvement. We then leverage the unique features in GPGPU applications to effectively tolerate the extra access delay to the slow registers. Moreover, we propose to dynamically balance the utilization across registers to further tolerate the NBTI degradation. Our experimental results show that our proposed techniques optimize GPGPUs performance by 22 percent on average under both PV and NBTI effects.
Jingweijia Tan, Mingsong Chen 0001, Yang Yi 0002, Xin Fu 0001
IEEE Trans. Parallel Distributed Syst.1
2015 Soft-error reliability and power co-optimization for GPGPUS register file using resistive memory
Jingweijia Tan, Zhi Li 0016, Xin Fu 0001
DATE1
2015 Mitigating the Susceptibility of GPGPUs Register File to Process Variations
abstract
As technology keeps scaling down at nano-scale, the increasing process variations (PV) induce significant delay variations and limit the maximum clock frequency in GPGPUs (general-purpose computing on graphics processing units). Each computing core (i.e. streaming multiprocessor) in GPGPUs supports thousands of simultaneously active threads, and requires a large register file. Such a sizeable register file is very sensitive to process variations, and becomes one of the major units in determining the core frequency. In this study, we first develop a novel mechanism that classifies registers into fast and slow categories in the highly-banked register architecture to maximize the frequency improvement. We then leverage the unique features in GPGPU applications to effectively tolerate the extra access delay to the slow registers. Our experimental results show that our proposed techniques are able to significantly optimize GPGPUs performance under process variations.
Jingweijia Tan, Xin Fu 0001
IPDPS1
2013 Modeling and characterizing GPGPU reliability in the presence of soft errors
Jingweijia Tan, Yang Yi 0002, Fangyang Shen, Xin Fu 0001
Parallel Comput.1
2012 RISE: improving the streaming processors reliability against soft errors in gpgpus
abstract
With hundreds of cores integrated into a single chip, the general-purpose computing on graphic processing units (GPGPUs) provide high computing power to accelerate parallel applications. However, they are prone to manifest high soft-error vulnerability due to the lack of fault detection and tolerance. Especially, streaming processors become the reliability hot-spot in GPGPUs. This paper explores two opportunistic soft-error detection techniques to cost-effectively improve the streaming processors reliability. Observing that the streaming processors are not fully utilized during the branch divergence and pipeline stalls caused by the long latency operations, we propose to Recycle the streaming processors Idle time for Soft-Error detection (RISE) and obtain the good fault coverage with negligible performance degradation. RISE is composed of full-RISE and partial-RISE. Full-RISE selectively triggers the redundancy for a set of warps so that leverages the fully idled streaming processors during the pipeline stall time for the error detection. Partial-RISE performs the redundancy for a number of threads in certain warps using the partially idled streaming processors during the branch divergence. Our experimental results show that RISE shows strong capability in improving the SPs soft-error reliability by 43% with negligible (e.g. 4%) performance loss.
Jingweijia Tan, Xin Fu 0001
PACT1