EDBT 2026 Demo / reviewers in the wild / expert
Lei Liu 0030
dblp:21/2715-30
· DBLP profile ↗
31ranked-venue papers
8as first author
6since 2021 · last 2022
0000-0001-5063-2864ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A systematic study on benchmarking AI inference accelerators
Zihan Jiang 0006, Jiansong Li, Fangxin Liu, Wanling Gao, Lei Wang 0004, Chuanxin Lan, Fei Tang 0003, Lei Liu 0030, Tao Li 0022 |
CCF Trans. High Perform. Comput. | 8 |
| 2022 | Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning
Guangli Li, Xiu Ma, Xueying Wang 0003, Hengshan Yue, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002, Jingling Xue |
J. Syst. Archit. | 6 |
| 2022 | An Application-oblivious Memory Scheduling System for DNN AcceleratorsabstractDeep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes. Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 10 |
| 2021 | Unleashing the Low-Precision Computation Potential of Tensor Cores on GPUsabstractTensor-specialized hardware for supporting low-precision arithmetic has become an inevitable trend due to the ever-increasing demand on computational capability and energy efficiency in intelligent applications. The main challenge faced when accelerating a tensor program on tensor-specialized hardware is how to achieve the best performance possible in reduced precision by fully utilizing its computational resources while keeping the precision loss in a controlled manner. In this paper, we address this challenge by proposing QUANTENSOR, a new approach for accelerating general-purpose tensor programs by replacing its tensor computations with low-precision quantized tensor computations on NVIDIA Tensor Cores. The key novelty is a new residual-based precision refinement technique for controlling the quantization errors, allowing tradeoffs between performance and precision to be made. Evaluation with GEMM, deep neural networks, and linear algebra applications shows that QUANTENSOR can achieve remarkable performance improvements while reducing the precision loss incurred significantly at acceptable overheads. Guangli Li, Jingling Xue, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Jiansong Li, Xiaobing Feng 0002 |
CGO | 3 |
| 2021 | NRHI: A Concurrent Non-Rehashing Hash Index for Persistent MemoryabstractPersistent memory (PM) featured with data persistence, byte-addressability, and DRAM-like performance has been commercially available with the advent of Intel®Optane™ DC persistent memory. The DRAM-like performance and disk-like persistence invite shifting hashing-based index schemes, which are important building blocks of today’s internet service infrastructures to provide fast queries, from DRAM onto persistent memory. Numerous hash indexes for persistent memory have been proposed to optimize writes and crash consistency, but with poor scalability under resizing. Generally, resizing consists of allocating a new hash table and rehashing items from the old table into the new one. We argue that resizing with rehashing performed in either blocking or non-blocking way can degrade the overall performance and limit the scalability.In order to mitigate the limitation of resizing, this paper proposes a Non-Rehashing Hash Index (NRHI) scheme to perform resizing with no necessity of rehashing items. NRHI leverages a layered structure to link hash tables without moving key-value pairs across layers, thus reducing the time spent on rehashing in blocking way and alleviating slots contention occurred in non-blocking way. Furthermore, the compare-and-swap primitive is utilized to support concurrent lock-free hashing operations. Experimental results on real PM hardware show that NRHI outperforms the state-of-the-art PM hash indexes by 1.7× to 3.59×, and scales linearly with the number of threads. Huimin Cui, Lei Liu 0030 |
ICCD | 3 |
| 2021 | Pinpointing the Memory Behaviors of DNN TrainingabstractThe training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns. Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ISPASS | 11 |
| 2020 | Accelerating Deep Learning Inference with Cross-Layer Data Reuse on GPUs
Xueying Wang 0003, Guangli Li, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002 |
Euro-Par | 5 |
| 2020 | Lance: efficient low-precision quantized winograd convolution for neural networks based on graphics processing unitsabstractAccelerating deep convolutional neural networks has become an active topic and sparked an interest in academia and industry. In this paper, we propose an efficient low-precision quan-tized Winograd convolution algorithm, called LANCE, which combines the advantages of fast convolution and quantization techniques. By embedding linear quantization operations into the Winograd-domain, the fast convolution can be performed efficiently under low-precision computation on graphics processing units. We test neural network models with LANCE on representative image classification datasets, including SVHN, CIFAR, and ImageNet. The experimental results show that our 8-bit quantized Winograd convolution improves the performance by up to 2.40× over the full-precision convolution with trivial accuracy loss. Guangli Li, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Xiaobing Feng 0002 |
ICASSP | 2 |
| 2020 | Compiler-Assisted Operator Template Library for DNN Accelerators
Jiansong Li, Wei Cao 0010, Guangli Li, Xueying Wang 0003, Lei Liu 0030, Xiaobing Feng 0002 |
NPC | 6 |
| 2020 | Fusion-Catalyzed Pruning for Optimizing Deep Learning on Intelligent Edge DevicesabstractThe increasing computational cost of deep neural network models limits the applicability of intelligent applications on resource-constrained edge devices. While a number of neural network pruning methods have been proposed to compress the models, prevailing approaches focus only on parametric operators (e.g., convolution), which may miss optimization opportunities. In this article, we present a novel fusion-catalyzed pruning approach, called FuPruner, which simultaneously optimizes the parametric and nonparametric operators for accelerating neural networks. We introduce an aggressive fusion method to equivalently transform a model, which extends the optimization space of pruning and enables nonparametric operators to be pruned in a similar manner as parametric operators, and a dynamic filter pruning method is applied to decrease the computational cost of models while retaining the accuracy requirement. Moreover, FuPruner provides configurable optimization options for controlling fusion and pruning, allowing much more flexible performance-accuracy tradeoffs to be made. Evaluation with state-of-the-art residual neural networks on five representative intelligent edge platforms, Jetson TX2, Jetson Nano, Edge tensor processing unit, neural compute stick, and neural compute stick 2, demonstrates the effectiveness of our approach, which can accelerate the inference of models on CIFAR-10 and ImageNet datasets. Guangli Li, Xiu Ma, Xueying Wang 0003, Lei Liu 0030, Jingling Xue, Xiaobing Feng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Acorns: A Framework for Accelerating Deep Neural Networks with Input SparsityabstractDeep neural networks have been employed in a broad range of applications, including face detection, natural language processing, and autonomous driving. Yet, the neural networks with the capability to tackle real-world problems are intrinsically expensive in computation, hindering the usage of these models. Sparsity in the input data of neural networks provides an optimizing opportunity. However, harnessing the potential performance improvement on modern CPU faces challenges raised by sparse computations of the neural network, such as cache-unfriendly memory accesses and efficient sparse kernel implementation. In this paper, we propose Acorns, a framework to accelerate deep neural networks with input sparsity. In Acorns, sparse input data is organized into our designed sparse data layout, which allows memory-friendly access for kernels in neural networks and opens the door for many performance-critical optimizations. Upon that, Acorns generates efficient sparse kernels for operators in neural networks from kernel templates, which combine directions that express specific optimizing transformations to be performed, and straightforward code that describes the computation. Comprehensive evaluations demonstrate Acorns can outperform state-of-the-art baselines by significant speedups. On the real-world detection task in autonomous driving, Acorns demonstrates 1.8-22.6× performance improvement over baselines. Specifically, the generated programs achieve 1.8-2.4× speedups over Intel MKL-DNN, 3.0-8.8× speedups over TensorFlow, and 11.1-13.2× speedups over Intel MKL-Sparse. Lei Liu 0030, Peng Zhao 0008, Guangli Li, Jiansong Li, Xueying Wang 0003, Xiaobing Feng 0002 |
PACT | 2 |
| 2019 | Accelerating GPU Computing at Runtime with Binary OptimizationabstractNowadays, many applications use GPUs (Graphics Processing Units) to achieve high performance. When we use GPU servers, the idle CPU resource of the servers is often ignored. In this paper, we explore the idea: using the idle CPU resource to speed up GPU programs. We design a dynamic binary optimization framework for accelerating GPU computing at runtime. A template-based binary optimization method is proposed to optimize kernels, which can avoid the high cost of kernel compilation. This method replaces determined variables with constant values and generates an optimized binary kernel. Based on the analysis results of optimization opportunities, we replace the original kernels with optimized kernels during program execution. The experimental results show that it is feasible to accelerate GPU programs via binary optimization. After applying binary optimization to five convolution layers of deep neural networks, the average performance improvement can reach 20%. Guangli Li, Lei Liu 0030, Xiaobing Feng 0002 |
CGO | 2 |
| 2019 | Exploiting the input sparsity to accelerate deep neural networks: posterabstractEfficient inference of deep learning models are challenging and of great value in both academic and industrial community. In this paper, we focus on exploiting the sparsity in input data to improve the performance of deep learning models. We propose an end-to-end optimization pipeline to generate programs for the inference with sparse input. The optimization pipeline contains both domain-specific and general optimization techniques and is capable of generating efficient code without relying on the off-the-shelf libraries. Evaluations show that we achieve significant speedups over the state-of-the-art frameworks and libraries on a real-world application, e.g., 9.8× over TensorFlow and 3.6× over Intel MKL on the detection in autonomous driving. Lei Liu 0030, Guangli Li, Jiansong Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Feng 0002 |
PPoPP | 2 |
| 2019 | Cacheap: Portable and Collaborative I/O Optimization for Graph Processing
Peng Zhao 0008, Chen Ding 0001, Lei Liu 0030, Jiping Yu, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 3 |
| 2018 | Fast CNN Pruning via Redundancy-Aware Training
Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 2 |
| 2018 | Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
Guangli Li, Lei Liu 0030, Xueying Wang 0003, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 2 |
| 2018 | Background Subtraction on Depth Videos with Convolutional Neural NetworksabstractBackground subtraction is a significant component of computer vision systems. It is widely used in video surveillance, object tracking, anomaly detection, etc. A new data source for background subtraction appeared as the emergence of low-cost depth sensors like Microsof t Kinect, Asus Xtion PRO, etc. In this paper, we propose a background subtraction approach on depth videos, which is based on convolutional neural networks (CNNs), called BGSNet-D (BackGround Subtraction neural Networks for Depth videos). The method can be used in color unavailable scenarios like poor lighting situations, and can also be applied to combine with existing RGB background subtraction methods. A preprocessing strategy is designed to reduce the influences incurred by noise from depth sensors. The experimental results on the SBM-RGBD dataset show that the proposed method outperforms existing methods on depth data, and even reaches the performance of the methods that use RGB-D data. Xueying Wang 0003, Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
IJCNN | 2 |
| 2017 | SysMon: Monitoring Memory Behaviors via OS Approach
Mengyao Xie, Lei Liu 0030, Chenggang Wu 0002, Hongna Geng |
APPT | 2 |
| 2017 | Redundancy checking algorithms based on parallel novel extension ruleabstractRedundancy checking (RC) is a key knowledge reduction technology. Extension rule (ER) is a new reasoning method, first presented in 2003 and well received by experts at home and abroad. Novel extension rule (NER) is an improved ER-based reasoning method, presented in 2009. In this paper, we first analyse the characteristics of the extension rule, and then present a simple algorithm for redundancy checking based on extension rule (RCER). In addition, we introduce MIMF, a type of heuristic strategy. Using the aforementioned rule and strategy, we design and implement RCHER algorithm, which relies on MIMF. Next we design and implement an RCNER (redundancy checking based on NER) algorithm based on NER. Parallel computing greatly accelerates the NER algorithm, which has weak dependence among tasks when executed. Considering this, we present PNER (parallel NER) and apply it to redundancy checking and necessity checking. Furthermore, we design and implement the RCPNER (redundancy checking based on PNER) and NCPPNER (necessary clause partition based on PNER) algorithms as well. The experimental results show that MIMF significantly influences the acceleration of algorithm RCER in formulae on a large scale and high redundancy. Comparing PNER with NER and RCPNER with RCNER, the average speedup can reach up to the number of task decompositions when executed. Comparing NCPNER with the RCNER-based algorithm on separating redundant formulae, speedup increases steadily as the scale of the formulae is incrementing. Finally, we describe the challenges that the extension rule will be faced with and suggest possible solutions. Lei Liu 0030, Guangli Li, Shuai Lü 0001 |
J. Exp. Theor. Artif. Intell. | 1 |
| 2016 | Memos: A full hierarchy hybrid memory management frameworkabstractIn this paper, we introduce memos, which integrates suitable memory management policies and schedules resources over the entire memory hierarchy in hybrid memory system. Powered by an OS kernel level monitoring tool, memos captures memory patterns online, and then leverages them to guide the memory page placement and data mapping. Experimental results show, on average, memos can benefit memory utilization, contributing to system throughput and QoS by 19.1% and 23.6%. Moreover, memos can reduce the NVM side memory latency by 3∼83.3%, energy consumption by 25.1∼99%, and benefit the NVM lifetime significantly (40× improvement on average). Lei Liu 0030, Mengyao Xie, Chenggang Wu 0002 |
ICCD | 1 |
| 2016 | Pragma Directed Shared Memory Centric Optimizations on GPUs
Lei Liu 0030, Xiang-Hua Liu, Xiaobing Feng 0002, Chengyong Wu |
J. Comput. Sci. Technol. | 2 |
| 2016 | Rethinking Memory Management in Modern Operating System: Horizontal, Vertical or Random?abstractOn modern multicore machines, the memory management typically combines address interleaving in hardware and random allocation in the operating system (OS) to improve performance of both memory and cache. The conventional solutions, however, are increasingly strained as a wide variety of workloads run on complicated memory hierarchy and cause contention at multiple levels. We describe a new framework (named HVR) in OS memory management to support a flexible policy space for tackling diverse application needs, integrating vertical partitioning across layers, horizontal partitioning and random-interleaved allocation at a single layer. We exhaustively study the performance of these policies for over 2,000 workloads and correlate performance with application characteristics. Based on this correlation we derive several practical rules of memory allocation that we integrate into the unified HVR framework to guide resource partitioning and sharing for dynamic and diverse workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Experimental results show that our framework consistently outperforms the unmodified Linux kernel, with up to 21 percent performance gains, and outperforms prior solutions at individual levels of the memory hierarchy. Lei Liu 0030, Chen Ding 0001, Chengyong Wu |
IEEE Trans. Computers | 1 |
| 2015 | WiseThrottling: a new asynchronous task scheduler for mitigating I/O bottleneck in large-scale datacenter servers
Lei Liu 0030, Huimin Cui, Lei Wang 0004, Ying Liu 0055, Xiaobing Feng 0002, Pen-Chung Yew |
J. Supercomput. | 2 |
| 2014 | Going vertical in memory management: Handling multiplicity by multi-policyabstractMany emerging applications from various domains often exhibit heterogeneous memory characteristics. When running in combination on parallel platforms, these applications present a daunting variety of workload behaviors that challenge the effectiveness of any memory allocation strategy. Prior partitioning-based or random memory allocation schemes typically manage only one level of the memory hierarchy and often target specific workloads. To handle diverse and dynamically changing memory and cache allocation needs, we augment existing “horizontal” cache/DRAM bank partitioning with vertical partitioning and explore the resulting multi-policy space. We study the performance of these policies for over 2000 workloads and correlate the results with application characteristics via a data mining approach. Based on this correlation we derive several practical memory allocation rules that we integrate into a unified multi-policy framework to guide resources partitioning and coalescing for dynamic and diverse multi-programmed/threaded workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Extensive experiments show that, in practice, our framework can select proper memory allocation policy and consistently outperforms the unmodified Linux kernel, achieving up to 11% performance gains compared to prior techniques. Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ISCA | 1 |
| 2014 | Dynamic I/O-Aware Scheduling for Batch-Mode Applications on Chip Multiprocessor Systems of Cluster Platforms
Huimin Cui, Lei Wang 0004, Lei Liu 0030, Chenggang Wu 0002, Xiaobing Feng 0002, Pen-Chung Yew |
J. Comput. Sci. Technol. | 4 |
| 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systemsabstractThe main memory system is a shared resource in modern multicore machines that can result in serious interference leading to reduced throughput and unfairness. Many new memory scheduling mechanisms have been proposed to address the interference problem. However, these mechanisms usually employ relative complex scheduling logic and need modifications to Memory Controllers (MCs), which incur expensive hardware design and manufacturing overheads. This article presents a practical software approach to effectively eliminate the interference without any hardware modifications. The key idea is to modify the OS memory management system and adopt a page-coloring-based Bank-level Partitioning Mechanism (BPM) that allocates dedicated DRAM banks to each core (or thread). By using BPM, memory requests from distinct programs are segregated across multiple memory banks to promote locality/fairness and reduce interference. We further extend BPM to BPM+ by incorporating channel-level partitioning, on which we demonstrate additional gain over BPM in many cases. To achieve benefits in the presence of diverse application memory needs and avoid performance degradation due to resource underutilization, we propose a dynamic mechanism upon BPM/BPM+ that assigns appropriate bank/channel resources based on application memory/bandwidth demands monitored through PMU (performance-monitoring unit) and a low-overhead OS page table scanning process. We implement BPM/BPM+ in Linux 2.6.32.15 kernel and evaluate the technique on four-core and eight-core real machines by running a large amount of randomly generated multiprogrammed and multithreaded workloads. Experimental results show that BPM/BPM+ can improve the overall system throughput by 4.7%/5.9%, on average, (up to 8.6%/9.5%) and reduce the unfairness by an average of 4.2%/6.1% (up to 15.8%/13.9%). Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Access Annotation for Safe Program Parallelization
Chen Ding 0001, Lei Liu 0030 |
NPC | 2 |
| 2012 | A software memory partition approach for eliminating bank-level interference in multicore systemsabstractMain memory system is a shared resource in modern multicore machines, resulting in serious interference, which causes performance degradation in terms of throughput slowdown and unfairness. Numerous new memory scheduling algorithms have been proposed to address the interference problem. However, these algorithms usually employ complex scheduling logic and need hardware modification to memory controllers, as a result, industrial venders seem to have some hesitation in adopting them. Lei Liu 0030, Zehan Cui, Mingjie Xing, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
PACT | 1 |
| 2011 | Safe parallel programming using dynamic dependence hintsabstractSpeculative parallelization divides a sequential program into possibly parallel tasks and permits these tasks to run in parallel if and only if they show no dependences with each other. The parallelization is safe in that a speculative execution always produces the same output as the sequential execution. Chuanle Ke, Lei Liu 0030, Tongxin Bai, Bryan Jacobs, Chen Ding 0001 |
OOPSLA | 2 |
| 2008 | Global Tiling for Communication Minimal Parallelization on Distributed Memory Systems
Lei Liu 0030, Chengyong Wu, Xiaobing Feng 0002 |
Euro-Par | 1 |
| 2008 | Automatic Implementation of Multi-partitioning Using Global TilingabstractStrategies for partitioning an application’s data and computation play fundamental role in determining the efficiency of parallelization. This paper describes a sophisticated strategy for partitioning data and computation known as multi-partitioning, which can support the best parallelization for some applications such as the line sweep computations. However, the implementation of multi-partitioning is very difficult and, as we know, there is none automatic parallelizing compiler supports such partitioning strategy. Though the dHPF compiler implemented multi-partitioning as a special extension for block style HPF partitioning, it still needs the programmer’s participation to analyze the application and decide the data distribution scheme. In this paper, we present a global tiling transformation algorithm and a tile-to-processors mapping strategy called hyper-diagonal modular mapping, to implement the multi-partitioning strategy. The experimentation with NPB2.3-serial SP shows that the code generated by the compiler achieves scalable performance. Lei Liu 0030, Dingfei Zhang, Hengjie Li |
ICPADS | 1 |