VLDB 2026 Research / reviewers in the wild / expert
Pengcheng Dai
dblp:116/6968
· DBLP profile ↗
15ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 48% Memory systems · 34% Emerging computing paradigms · 7% | |
| Artificial intelligence
5 papers |
Efficient and distributed learning · 55% Reinforcement learning · 32% Deep learning architectures and training · 10% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 2 | 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022 SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.0 | 1 | 2026 | Global Convergence for Multi-agent Reinforcement Learning in Unreliable Communication Networks · INFOCOM 2026 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 2 | 2020 | Accelerating CNN Training by Pruning Activation Gradients · ECCV (25) 2020 SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.9 | 2 | 2020 | Accelerating CNN Training by Pruning Activation Gradients · ECCV (25) 2020 SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.6 | 1 | 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022 |
Hardware accelerators and domain-specific architectures
systolic array |
0.6 | 1 | 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022 |
Memory systems
cache |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN training accelerator |
0.4 | 1 | 2020 | SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems
non-volatile memory |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Energy-efficient computing
power management |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Emerging computing paradigms › approximate and stochastic computing
stochastic computing |
0.4 | 1 | 2020 | SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › non-volatile memory › magnetic random access memory › STT-MRAM
STT-MRAM cache |
0.4 | 1 | 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal Consideration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Distributed systems
fault tolerance |
0.3 | 1 | 2026 | Global Convergence for Multi-agent Reinforcement Learning in Unreliable Communication Networks · INFOCOM 2026 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
sparse convolutional networks |
0.2 | 1 | 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.1 | 1 | 2020 | SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network training |
0.1 | 1 | 2020 | Accelerating CNN Training by Pruning Activation Gradients · ECCV (25) 2020 |
Methods — techniques the papers use, named apart from their topics
multi-agent reinforcement learning · 2.0dynamic data selection · 1.1compressed dataflow · 1.1stochastic pruning · 0.9stochastic computing · 0.9magnetic tunnel junction · 0.9bayesian inference · 0.91-dimensional convolution dataflow · 0.9thermal-aware data migration · 0.4dynamic region partitioning · 0.4activation gradient pruning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Global Convergence for Multi-agent Reinforcement Learning in Unreliable Communication Networks
Pengcheng Dai, Lingjie Duan |
INFOCOM | 1 |
| 2024 | Applications in Traffic Signal Control: A Distributed Policy Gradient Decomposition AlgorithmabstractThis article explores the application of the multiagent reinforcement learning (MARL) algorithm in addressing the large-scale traffic signal control (TSC) problem. To address the TSC problem in complex urban traffic networks, most existing algorithms focus on optimizing local traffic flow at each intersection through decentralized training based on either the local observations or messages from its neighboring intersections, which, however, lacks the concept of cooperative learning. To conquer such limitations, a novel distributed critic with decentralized actor (DCDA) framework is proposed, which allows the communication messages and temporal difference (TD) losses to be exchanged among neighboring intersections. Specially, by considering the traffic network as a communication network of agents (more precisely, intersections and lanes are considered as agents and edges, respectively), a distributed global average TD loss estimation algorithm is designed in the distributed critic step to estimate the global average TD loss estimation and enhance collaboration among agents. Moreover, in the decentralized actor step, the policy gradient decomposition method is adopted for each agents to learn its local policy solely based on its local action-value function. By adhering to the DCDA framework, a novel distributed policy gradient decomposition (DPGD) algorithm is further proposed to address the TSC problem. Empirical experiments demonstrate that the efficiency, robustness, and stability of the DPGD algorithm outperform the state-of-the-art MARL algorithms in both the environments of cooperative adaptive cruise control and adaptive traffic signal control. Pengcheng Dai, Wenwu Yu, He Wang 0006 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Distributed Neural Learning Algorithms for Multiagent Reinforcement LearningabstractIn this article, the fully distributed neural learning algorithms by neural network approximation for networked multiagent reinforcement learning (NMARL) are studied. To tackle the convergence analysis of methods in NMARL with tremendous state-action space, most of the existing distributed algorithms are designed by linear function approximation, which however would fall into a situation of poor expression. To conquer such limitation, the distributed neural learning algorithms are developed by using a novel neural network approximation that bridges the theory and practice of deep NMARL (DMARL). Specifically, inspired by the overparametrization method for minimizing mean-squared projected bellman error (MSPBE), the distributed neural learning algorithms with population semigradients and stochastic semigradients are respectively, proposed to solve the NMARL problem. Furthermore, the convergence of the proposed algorithms are strictly given by employing the overparametrization method to establish the approximate stationary point of MSPBE to characterize the algorithms toward the global optimum. Finally, some numerical simulations demonstrate the effectiveness of the distributed neural learning algorithms. Pengcheng Dai, Hongzhe Liu 0002, Wenwu Yu, He Wang 0006 |
IEEE Internet Things J. | 1 |
| 2023 | Distributed Actor-Critic Algorithms for Multiagent Reinforcement Learning Over Directed GraphsabstractActor-critic (AC) cooperative multiagent reinforcement learning (MARL) over directed graphs is studied in this article. The goal of the agents in MARL is to maximize the globally averaged return in a distributed way, i.e., each agent can only exchange information with its neighboring agents. AC methods proposed in the literature require the communication graphs to be undirected and the weight matrices to be doubly stochastic (more precisely, the weight matrices are row stochastic and their expectation are column stochastic). Differently from these methods, we propose a distributed AC algorithm for MARL over directed graph with fixed topology that only requires the weight matrix to be row stochastic. Then, we also study the MARL over directed graphs (possibly not connected) with changing topologies, proposing a different distributed AC algorithm based on the push-sum protocol that only requires the weight matrices to be column stochastic. Convergence of the proposed algorithms is proven for linear function approximation of the action value function. Simulations are presented to demonstrate the effectiveness of the proposed algorithms. Pengcheng Dai, Wenwu Yu, He Wang 0006, Simone Baldi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) have achieved great success in performing cognitive tasks. However, execution of CNNs requires a large amount of computing resources and generates heavy memory traffic, which impose a severe challenge on computing system design. Through optimizing parallel executions and data reuse in convolution, systolic architecture demonstrates great advantages in accelerating CNN computations. However, regular internal data transmission path in traditional systolic architecture prevents the systolic architecture from completely leveraging the benefits introduced by neural network sparsity.Deployment of fine-grained sparsity on the existing systolic architectures is greatly hindered by the incurred computational overheads.In this work, we propose S2Engine a novel systolic architecture that can fully exploit the sparsity in CNNs with maximized data reuse. S2Engine transmits compressed data internally and allows each processing element to dynamically select an aligned data from the compressed dataflow in convolution. Compared to the naive systolic array, S2Engine achieves about 3.2 and about 3.0 improvements on speed and power efficiency, respectively. Jianlei Yang 0001, Wenzhi Fu, Xingzhou Cheng, Xucheng Ye, Pengcheng Dai, Weisheng Zhao 0001 |
IEEE Trans. Computers | 5 |
| 2022 | Distributed Q-Learning Algorithm for Dynamic Resource Allocation With Unknown Objective Functions and Application to MicrogridabstractDynamic resource allocation problem (DRAP) with unknown cost functions and unknown resource transition functions is studied in this article. The goal of the agents is to minimize the sum of cost functions over given time periods in a distributed way, that is, by only exchanging information with their neighboring agents. First, we propose a distributed Q -learning algorithm for DRAP with unknown cost functions and unknown resource transition functions under discrete local feasibility constraints (DLFCs). It is theoretically proved that the joint policy of agents produced by the distributed Q -learning algorithm can always provide a feasible allocation (FA), that is, satisfying the constraints at each time period. Then, we also study the DRAP with unknown cost functions and unknown resource transition functions under continuous local feasibility constraints (CLFCs), where a novel distributed Q -learning algorithm is proposed based on function approximation and distributed optimization. It should be noted that the update rule of the local policy of each agent can also ensure that the joint policy of agents is an FA at each time period. Such property is of vital importance to execute the ε -greedy policy during the whole training process. Finally, simulations are presented to demonstrate the effectiveness of the proposed algorithms. Pengcheng Dai, Wenwu Yu, Duxin Chen |
IEEE Trans. Cybern. | 1 |
| 2021 | Brief Industry Paper: optimizing Memory Efficiency of Graph Neural Networks on Edge Computing PlatformsabstractGraph neural networks (GNN) have achieved state-of-the-art performance on various industrial tasks. However, the poor efficiency of GNN inference and frequent Out-of-Memory (OOM) problem limit the successful application of GNN on edge computing platforms. To tackle these problems, a feature decomposition approach is proposed for memory efficiency optimization of GNN inference. The proposed approach could achieve outstanding optimization on various GNN models, covering a wide range of datasets, which speeds up the inference by up to 3×. Furthermore, the proposed feature decomposition could significantly reduce the peak memory usage (up to 5× in memory efficiency improvement) and mitigate OOM problems during GNN inference. Jianlei Yang 0001, Yeqi Gao, Yingjie Qi, Yunli Chen, Pengcheng Dai, Weisheng Zhao 0001, Chunming Hu |
RTAS | 8 |
| 2020 | SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingabstractTraining Convolutional Neural Networks (CNNs) usually requires a large number of computational resources. In this paper, SparseTrain is proposed to accelerate CNN training by fully exploiting the sparsity. It mainly involves three levels of innovations: activation gradients pruning algorithm, sparse training dataflow, and accelerator architecture. By applying a stochastic pruning algorithm on each layer, the sparsity of back-propagation gradients can be increased dramatically without degrading training accuracy and convergence rate. Moreover, to utilize both natural sparsity (resulted from ReLU or Pooling layers) and artificial sparsity (brought by pruning algorithm), a sparse-aware architecture is proposed for training acceleration. This architecture supports forward and back-propagation of CNN by adopting 1-Dimensional convolution dataflow. We have built a cycle-accurate architecture simulator to evaluate the performance and efficiency based on the synthesized design with 14nm FinFET technologies. Evaluation results on AlexNet/ResNet show that SparseTrain could achieve about 2.7× speedup and 2.2× energy efficiency improvement on average compared with the original training process. Pengcheng Dai, Jianlei Yang 0001, Xucheng Ye, Xingzhou Cheng, Junyu Luo 0002, Linghao Song, Yiran Chen 0001, Weisheng Zhao 0001 |
DAC | 1 |
| 2020 | Accelerating CNN Training by Pruning Activation Gradients
Xucheng Ye, Pengcheng Dai, Junyu Luo 0002, Xin Guo 0008, Yingjie Qi, Jianlei Yang 0001, Yiran Chen 0001 |
ECCV (25) | 2 |
| 2020 | SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic ComputingabstractBayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-intensive task whose efficiency is physically limited by the bottlenecks of conventional computing platforms. In this paper, a spintronics-based stochastic computing (SC) approach is proposed for efficient Bayesian inference. The inherent stochastic switching behaviors of spintronic devices are exploited to build a stochastic bitstream generator (SBG) for SC with hybrid CMOS/magnetic tunnel junction (MTJ) circuits design. Aiming to improve the inference efficiency, an SBG sharing strategy is leveraged to reduce the required SBG array scale by integrating a switch network between SBG array and SC logic. A device-to-architecture level framework is proposed to evaluate the performance of spintronics-based Bayesian inference system (SPINBIS). Experimental results on data fusion applications have shown that SPINBIS could improve the energy efficiency about 12× than MTJ-based approach with 45% design area overhead and about 26× than FPGA-based approach. Xiaotao Jia, Jianlei Yang 0001, Pengcheng Dai, Runze Liu 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Distributed Reinforcement Learning Algorithm for Dynamic Economic Dispatch With Unknown Generation Cost FunctionsabstractIn this article, the dynamic economic dispatch (DED) problem for smart grid is solved under the assumption that no knowledge of the mathematical formulation of the actual generation cost functions is available. The objective of the DED problem is to find the optimal power output of each unit at each time so as to minimize the total generation cost. To address the lack of a priori knowledge, a new distributed reinforcement learning optimization algorithm is proposed. The algorithm combines the state-action-value function approximation with a distributed optimization based on multiplier splitting. Theoretical analysis of the proposed algorithm is provided to prove the feasibility of the algorithm, and several case studies are presented to demonstrate its effectiveness. Pengcheng Dai, Wenwu Yu, Guanghui Wen, Simone Baldi |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Supporting Japanese Language Learners with an Onomatopoeia Learning SiteabstractThis paper describes our study where we aim to support onomatopoeia learning for Japanese language learners using our developed website called SLL Ono (To See, to Listen, to Learn onomatopoeia). The lack of English proficiency of Japanese people has been pointed out for a long time. Therefore it is a must for foreigners to have a Japanese language skill to survive in Japan. The focus in our study is onomatopoeia learning since there are more than 4,500 onomatopoeia words in Japanese. The result of our pilot evaluation shows that there was no statistically significant difference between our system and the blogger site. However the mean score increased more when they learned with SLL Ono and 15 out of 16 participants preferred our system. It was found that our automatically generated quiz function did not work effectively enough and that the refinement of the learning contents was necessary. It is among our future works to improve our quiz system and add more contents in order to improve quiz efficiency and to enrich its contents. Noriko Uosaki, Pengcheng Dai, Hye Rin Kong, Jacky Chun Kit Lam, Mehrasa Alizadeh |
ICCE | 2 |
| 2018 | A Scalable Pipelined Dataflow Accelerator for Object Region Proposals on FPGA PlatformabstractRegion proposal is critical for object detection while it usually poses a bottleneck in improving the computation efficiency on traditional control-flow architectures. We have observed region proposal tasks are potentially suitable for performing pipelined parallelism by exploiting dataflow driven acceleration. In this paper, a scalable pipelined dataflow accelerator is proposed for efficient region proposals on FPGA platform. The accelerator processes image data by a streaming manner with three sequential stages: resizing, kernel computing and sorting. First, Ping-Pong cache strategy is adopted for rotation loading in resize module to guarantee continuous output streaming. Then, a multiple pipelines architecture with tiered memory is utilized in kernel computing module to complete the main computation tasks. Finally, a bubble-pushing heap sort method is exploited in sorting module to find the top-k largest candidates efficiently. Our design is implemented with high level synthesis on FPGA platforms, and experimental results on VOC2007 datasets show that it could achieve about 3.67X speedups than traditional desktop CPU platform and >250X energy efficiency improvement than embedded ARM platform. Wenzhi Fu, Jianlei Yang 0001, Pengcheng Dai, Yiran Chen 0001, Weisheng Zhao 0001 |
FPT | 3 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 3 |