VLDB 2026 Research / reviewers in the wild / expert
Peng Zhao 0008
dblp:93/4324-8
· DBLP profile ↗
8ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0003-4668-1852ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 64% Hardware accelerators and domain-specific architectures · 36% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 69% Runtime systems and virtual machines · 31% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › memory management
DNN training memory management |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › memory controller
memory scheduling |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › virtual memory management
memory swapping |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Compilers and program optimization
domain-specific compilation |
0.4 | 1 | 2019 | Exploiting the input sparsity to accelerate deep neural networks: poster · PPoPP 2019 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | Exploiting the input sparsity to accelerate deep neural networks: poster · PPoPP 2019 |
Runtime systems and virtual machines
runtime memory management |
0.2 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Methods — techniques the papers use, named apart from their topics
memory behavior tracing · 1.1application-oblivious scheduling · 1.1program optimization · 0.8code generation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | An Application-oblivious Memory Scheduling System for DNN AcceleratorsabstractDeep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes. Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 6 |
| 2021 | Pinpointing the Memory Behaviors of DNN TrainingabstractThe training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns. Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ISPASS | 4 |
| 2019 | Acorns: A Framework for Accelerating Deep Neural Networks with Input SparsityabstractDeep neural networks have been employed in a broad range of applications, including face detection, natural language processing, and autonomous driving. Yet, the neural networks with the capability to tackle real-world problems are intrinsically expensive in computation, hindering the usage of these models. Sparsity in the input data of neural networks provides an optimizing opportunity. However, harnessing the potential performance improvement on modern CPU faces challenges raised by sparse computations of the neural network, such as cache-unfriendly memory accesses and efficient sparse kernel implementation. In this paper, we propose Acorns, a framework to accelerate deep neural networks with input sparsity. In Acorns, sparse input data is organized into our designed sparse data layout, which allows memory-friendly access for kernels in neural networks and opens the door for many performance-critical optimizations. Upon that, Acorns generates efficient sparse kernels for operators in neural networks from kernel templates, which combine directions that express specific optimizing transformations to be performed, and straightforward code that describes the computation. Comprehensive evaluations demonstrate Acorns can outperform state-of-the-art baselines by significant speedups. On the real-world detection task in autonomous driving, Acorns demonstrates 1.8-22.6× performance improvement over baselines. Specifically, the generated programs achieve 1.8-2.4× speedups over Intel MKL-DNN, 3.0-8.8× speedups over TensorFlow, and 11.1-13.2× speedups over Intel MKL-Sparse. Lei Liu 0030, Peng Zhao 0008, Guangli Li, Jiansong Li, Xueying Wang 0003, Xiaobing Feng 0002 |
PACT | 3 |
| 2019 | Exploiting the input sparsity to accelerate deep neural networks: posterabstractEfficient inference of deep learning models are challenging and of great value in both academic and industrial community. In this paper, we focus on exploiting the sparsity in input data to improve the performance of deep learning models. We propose an end-to-end optimization pipeline to generate programs for the inference with sparse input. The optimization pipeline contains both domain-specific and general optimization techniques and is capable of generating efficient code without relying on the off-the-shelf libraries. Evaluations show that we achieve significant speedups over the state-of-the-art frameworks and libraries on a real-world application, e.g., 9.8× over TensorFlow and 3.6× over Intel MKL on the detection in autonomous driving. Lei Liu 0030, Guangli Li, Jiansong Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Feng 0002 |
PPoPP | 5 |
| 2019 | Cacheap: Portable and Collaborative I/O Optimization for Graph Processing
Peng Zhao 0008, Chen Ding 0001, Lei Liu 0030, Jiping Yu, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 1 |
| 2018 | Fast CNN Pruning via Redundancy-Aware Training
Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 4 |
| 2018 | Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
Guangli Li, Lei Liu 0030, Xueying Wang 0003, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 5 |
| 2018 | Background Subtraction on Depth Videos with Convolutional Neural NetworksabstractBackground subtraction is a significant component of computer vision systems. It is widely used in video surveillance, object tracking, anomaly detection, etc. A new data source for background subtraction appeared as the emergence of low-cost depth sensors like Microsof t Kinect, Asus Xtion PRO, etc. In this paper, we propose a background subtraction approach on depth videos, which is based on convolutional neural networks (CNNs), called BGSNet-D (BackGround Subtraction neural Networks for Depth videos). The method can be used in color unavailable scenarios like poor lighting situations, and can also be applied to combine with existing RGB background subtraction methods. A preprocessing strategy is designed to reduce the influences incurred by noise from depth sensors. The experimental results on the SBM-RGBD dataset show that the proposed method outperforms existing methods on depth data, and even reaches the performance of the methods that use RGB-D data. Xueying Wang 0003, Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
IJCNN | 5 |