EDBT 2026 Demo / reviewers in the wild / expert
Lan Gao 0004
dblp:55/5352-4
· DBLP profile ↗
12ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0001-5637-9417ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Bit Level Acceleration of Mixed Precision Neural NetworkabstractWith the growth of the convolutional neural network (CNN) parameters, the hardware resources become limited when deploying CNN models. Single bit-width quantization may lead to degradation of accuracy, while mixed-precision quantization models maintain higher accuracy. However, current mixed-precision quantization accelerators only consider the algorithm level quantization and fixed-bit-width processing elements (PEs) without fully utilizing the resources. Therefore, we propose a neural network acceleration architecture based on bit-level computational units to improve resource utilization and throughput of mixed-precision quantization accelerators. The 2-bit and 3-bit low-bit computational units are designed to implement the high-bit quantization. We also propose spatio-temporal fusion to satisfy the unique bit widths in each layer with mixed precision quantization. In particular, we use the 2-bit and 3-bit computational units to achieve dynamic layer-level quantization. We also discuss various combinations of different bit widths, which can be dynamically implemented according to the requirements of accuracy, execution time, etc. The proposed acceleration architecture is implemented in Verilog, verified using three network models: ResNet18, VGG7 and LeNet-5, and tested against three accelerators: Eyeriss, Stripes and Bit Fusion. The experimental results show that our accelerators provide more accuracy and the grouping operation reduces the area overhead. It provides a 3.08 to 3.54 times acceleration ratio and 2.4 to 2.57 times reduction in energy consumption on the three networks. Dehui Qiu, Jing Wang 0055, Weigong Zhang, Lan Gao 0004 |
ICPADS | 5 |
| 2023 | Accelerating Look-Up Table based Matrix Multiplication on GPUsabstractMultiplying matrices is among the most fundamental and compute-intensive operations in machine learning. Approximated Matrix Multiplication (AMM) based on table look-ups can significantly reduce the pressure on computing units and memory bandwidth, and has great potential in large-scale machine learning applications. In this work, we speed up table look-ups on GPUs to improve the performance of matrix multiplication. To avoid random memory accesses in table look-ups, we propose a novel warp-wide data sharing execution model. With this execution model, we develop a GPU AMM library to speed up MADDNESS (the state-of-the-art AMM), named GPU-MADDNESS. The experimental results show that GPU-MADDNESS improves the performance by 103X on average, and outperforms the tiling implementation by up to 42%. Lan Gao 0004, Weigong Zhang, Jing Wang 0055, Dehui Qiu |
ICPADS | 2 |
| 2022 | Adaptive Contention Management for Fine-Grained Synchronization on Commodity GPUsabstractAs more emerging applications are moving to GPUs, fine-grained synchronization has become imperative. However, their performance can be severely impaired in case of frequent synchronization failures caused by high data contention. Differently from CPUs, GPUs own thousands of hardware threads and adopt single instruction multiple threads paradigm, making it impractical to deploy the CPU contention management mechanisms directly on GPUs. In this article, we design a Software Warp Controlling Framework (SWCF), which employs producer-consumer execution model and leverages GPU hardware barriers to dynamically control the execution of warps at runtime. On the basis of SWCF, we propose a contention management strategy to decrease frequent synchronization failures while avoiding the over-reducing of parallelism. We evaluate SWCF and the proposed strategy on commodity GPUs using a set of applications with fine-grained synchronization. The results show that on V100 GPU our contention management achieves a 4.7X speedup and outperforms the conventional GPU software backoff solution by 42% on average. Lan Gao 0004, Jing Wang 0055, Weigong Zhang |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Multi-dimensional optimization for approximate near-threshold computingabstractThe demise of Dennard’s scaling has created both power and utilization wall challenges for computer systems. As transistors operating in the near-threshold region are able to obtain flexible trade-offs between power and performance, it is regarded as an alternative solution to the scaling challenge. A reduction in supply voltage will nevertheless generate significant reliability challenges, while maintaining an error-free system that generates high costs in both performance and energy consumption. The main purpose of research on computer architecture has therefore shifted from performance improvement to complex multi-objective optimization. In this paper, we propose a three-dimensional optimization approach which can effectively identify the best system configuration to establish a balance among performance, energy, and reliability. We use a dynamic programming algorithm to determine the proper voltage and approximate level based on three predictors: system performance, energy consumption, and output quality. We propose an output quality predictor which uses a hardware/software co-design fault injection platform to evaluate the impact of the error on output quality under near-threshold computing (NTC). Evaluation results demonstrate that our approach can lead to a 28% improvement in output quality with a 10% drop in overall energy efficiency; this translates to an approximately 20% average improvement in accuracy, power, and performance. Jing Wang 0055, Wei-wei Liang, Yuehua Niu, Lan Gao 0004, Weigong Zhang |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2020 | Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage ScalingabstractWith the platforms of running deep neural networks (DNNs) move from large-scale data centers to handheld devices, power emerge as one of the most significant obstacles. Voltage scaling is a promising technique that enables power saving. Nevertheless, it raises reliability and performance concerns that may undesirably deteriorate NNs accuracy and performance. Consequently, an energy-efficient and reliable scheme is required for NNs to balance the above three aspects with satisfied user experience. To this end, we propose a neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation characteristics in NNs at both inter- and intra-network layers to precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Furthermore, we combine a voltage clustering method and the multi-objective optimization to identify the optimal voltage islands and apply the same voltage to neurons with similar fault tolerance capability. We perform three case studies to demonstrate the efficacy of the proposed techniques. Jing Wang 0055, Xin Fu 0001, Lan Gao 0004, Weigong Zhang |
IEEE Trans. Computers | 5 |
| 2020 | Thread-Level Locking for SIMT ArchitecturesabstractAs more emerging applications are moving to GPUs, thread-level synchronization has become a requirement. However, GPUs only provide warp-level and thread-block-level rather than thread-level synchronization. Moreover, it is highly possible to cause live-locks by using CPU synchronization mechanisms to implement thread-level synchronization for GPUs. In this article, we first propose a software-based thread-level synchronization mechanism called lock stealing for GPUs to avoid live-locks. We then describe how to implement our lock stealing algorithm in mutual exclusive locks and readers-writer locks with high performance. Finally, by putting it all together, we develop a thread-level locking library (TLLL) for commercial GPUs. To evaluate TLLL and show its general applicability, we use it to implement six widely used programs. We compare TLLL against the state-of-the-art ad-hoc GPU synchronization, GPU software transactional memory (STM), and CPU hardware transactional memory (HTM), respectively. The results show that, compared with the ad-hoc GPU synchronization for Delaunay mesh refinement (DMR), TLLL improves the performance by 22 percent on average on a GTX970 GPU, and shows up to 11 percent of performance improvement on a Volta V100 GPU. Moreover, it significantly reduces the required memory size. Such low memory consumption enables DMR to successfully run on the GTX970 GPU with the 10-million mesh size, and the V100 GPU with the 40-million mesh size, with which the ad-hoc synchronization can not run successfully. In addition, TLLL outperforms the GPU STM by 65 percent, and the CPU HTM (running on a Xeon E5-2620 v4 CPU with 16 hardware threads) by 43 percent on average. Lan Gao 0004, Rui Wang 0014, Zhongzhi Luan, Zhibin Yu 0001, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage ScalingabstractAs the application scope of deep neural networks (DNNs) moves from large-scale data centers to small-scale mobile devices, power wall has become one of the most important obstacles. Voltage scaling is a typical technique enables power saving, but it causes reliability and performance challenges. Therefore, an energy-efficient and reliable scheme for NNs is required to balance above three aspects according to users' requirements for excellent user experience. In this paper, we innovatively propose neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation in NNs and precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Multi-objective optimization and clustering method are combined to find the optimal voltage islands. Finally, we conduct experiment to demonstrate the efficacy of the proposed technique. Jing Wang 0055, Xin Fu 0001, Xingyao Zhang 0002, Lan Gao 0004, Weigong Zhang, Tao Li 0006 |
ICPADS | 6 |
| 2019 | Multiple Algorithms Against Multiple Hardware Architectures: Data-Driven Exploration on Deep Convolution Neural Network
Chongyang Xu, Zhongzhi Luan, Lan Gao 0004, Rui Wang 0014, Lianyi Zhang, Yi Liu 0013, Depei Qian 0001 |
NPC | 3 |
| 2019 | Accelerating in-memory transaction processing using general purpose graphics processing units
Lan Gao 0004, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
Future Gener. Comput. Syst. | 1 |
| 2018 | SRAM- and STT-RAM-based hybrid, shared last-level cache for on-chip CPU-GPU heterogeneous architectures
Lan Gao 0004, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001, Jihong Cai |
J. Supercomput. | 1 |
| 2016 | Scheduling Tasks with Mixed Timing Constraints in GPU-Powered Real-Time SystemsabstractDue to the cost-effective, massive computational power of graphics processing units (GPUs), there is a growing interest of utilizing GPUs in real-time systems. For example GPUs have been applied to automotive systems to enable new advanced and intelligent driver assistance technologies, accelerating the path to self-driving cars. In such systems, GPUs are shared among tasks with mixed timing constraints: real-time (RT) tasks that have to be accomplished before specified deadlines, and non-real-time, best-effort (BE) tasks. In this paper, (1) we propose resource-aware non-uniform slack distribution to enhance the schedulability of RT tasks (the total amount of work of RT tasks whose deadlines can be satisfied on a given amount of resources) in GPU-enabled systems; (2) we propose deadline-aware dynamic GPU partitioning to allow RT and BE tasks to run on a GPU simultaneously, such that BE tasks are not blocked for a long time. Rui Wang 0014, Tao Li 0006, Mingcong Song, Lan Gao 0004, Zhongzhi Luan, Depei Qian 0001 |
ICS | 5 |
| 2014 | Software Transactional Memory for GPU Architectures
Rui Wang 0014, Nilanjan Goswami, Tao Li 0006, Lan Gao 0004, Depei Qian 0001 |
CGO | 5 |