VLDB 2026 Research / reviewers in the wild / expert
Guanrui Wang
dblp:238/0286
· DBLP profile ↗
5ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-9109-7535ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 43% Hardware accelerators and domain-specific architectures · 34% Memory systems · 12% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 87% Efficient and distributed learning · 13% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
1.2 | 2 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on Crossbars · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Parallel and multicore computing
parallel programming models and runtimes |
0.8 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
Parallel and multicore computing
task scheduling |
0.8 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
Memory systems
in-memory computing |
0.4 | 1 | 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on Crossbars · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 1 | 2019 | Convolution with even-sized kernels and symmetric padding · NeurIPS 2019 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore processors |
0.2 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
Electronic design automation
design optimization |
0.1 | 1 | 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on Crossbars · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Methods — techniques the papers use, named apart from their topics
prototype chip design · 0.8mapping strategy · 0.8semi-folded convolution mapping · 0.4cycle-accurate simulation · 0.4symmetric padding · 0.4depthwise convolution · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN TasksabstractThe rapid development of deep learning has propelled many real-world artificial intelligence applications. Many of these applications integrate multiple neural networks (multi-NN) to cater to various functionalities. There are two challenges of multi-NN acceleration: (1) competition for shared resources becomes a bottleneck, and (2) heterogeneous workloads exhibit remarkably different computing-memory characteristics and various synchronization requirements. Therefore, resource isolation and fine-grained resource allocation for each task are two fundamental requirements for multi-NN computing systems. Although a number of multi-NN acceleration technologies have been explored, few can completely fulfill both of these requirements, especially for mobile scenarios. This paper reports a Hierarchical Asynchronous Parallel Model (HASP) to enhance multi-NN performance to meet both requirements. HASP can be implemented on a multicore processor that adopts Multiple Instruction Multiple Data (MIMD) or Single Instruction Multiple Thread (SIMT) architectures, with minor adaptive modification needed. Further, a prototype chip is developed to validate the hardware effectiveness of this design. A corresponding mapping strategy is also developed, allowing the proposed architecture to simultaneously promote resource utilization and throughput. With the same workload, the prototype chip demonstrates 3.62$\boldsymbol{\times}$, and 3.51$\boldsymbol{\times}$higher throughput over Planaria and 8.68$\boldsymbol{\times}$, 2.61$\boldsymbol{\times}$over Jetson AGX Orin for MobileNet-V1 and ResNet50, respectively. Songchen Ma, Taoyi Wang, Guanrui Wang, Chenhang Song, Huanyu Qu, Junfeng Lin, Jing Pei |
IEEE Trans. Computers | 5 |
| 2023 | Policy Gradient-Based Core Placement Optimization for Multichip Many-Core SystemsabstractAs many deep neural network models become deeper and more complex, processing devices with stronger computing performance and communication capability are required. Following this trend, the dependence on multichip many-core systems that have high parallelism and reasonable transmission costs is on the rise. In this work, in order to improve routing performance of the system, such as routing runtime and power consumption, we propose a reinforcement learning (RL)- based core placement optimization approach, considering application constraints, such as deadlock caused by multicast paths. We leverage the capability of deep RL from indirect supervision as a direct nonlinear optimizer, and the parameters of the policy network are updated by proximal policy optimization. We treat the routing topology as a network graph, so we utilize a graph convolutional network to embed the features into the policy network. One step size environment is designed, so all cores are placed simultaneously. To handle large dimensional action space, we use continuous values matching with the number of cores as the output of the policy network and discretize them again for obtaining the new placement. For multichip system mapping, we developed a community detection algorithm. We use several datasets of multilayer perceptron and convolutional neural networks to evaluate our agent. We compare the optimal results obtained by our agent with other baselines under different multicast conditions. Our approach achieves a significant reduction of routing runtime, communication cost, and average traffic load, along with deadlock-free performance for inner chip data transmission. The traffic of interchip routing is also significantly reduced after integrating the community detection algorithm to our agent. Wooshik Myung, Donghyun Lee 0002, Chenhang Song, Guanrui Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | A deadlock-free physical mapping method on the many-core neural network chip
Guoqi Li 0002, Lei Deng 0003, Guanrui Wang |
Neurocomputing | 5 |
| 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on CrossbarsabstractCrossbar architecture has been widely used in neural network (NN) accelerators, involving conventional and emerging devices. It performs well on the fully connected layer through efficient vector-matrix multiplication. Whereas, the advantages degrade on the convolutional layer with huge data reuse, since the execution speed and resource overhead are imbalanced when using existing fully unfolded or fully folded mapping strategy. To address this issue, we propose a novel semi-folded mapping (SemiMap) framework for implementing the convolution on crossbars. It simultaneously folds the physical resources along the row dimension of feature maps (FMs) and unfolds them along the column dimension. The former reduces the resource overhead, and the latter maintains the parallelism. An FM slicing scheme is further proposed to enable the processing of large-size image. Via our mapping framework, a row-by-row streaming pipeline for intraimage dataflow and periodical pipeline for interimage dataflow are easy to be obtained. To validate the idea, we build a many-crossbar architecture with several designs to guarantee the overall functionality and performance. Based on the measurement data of a fabricated chip, a mapping compiler and a cycle-accurate simulator are developed for the hardware simulation of large-scale networks. We evaluate the proposed SemiMap on various convolutional NNs across different network scale. ${>} 35 {\times }$ resource saving and several hundred times cycle reduction are demonstrated compared to the existing fully unfolded and fully folded strategies, respectively. This paper jumps out of the current extreme mapping schemes, and provides a balanced solution on how to efficiently deploy the computational graphs with data reuse on many-crossbar architecture. Lei Deng 0003, Yuan Xie 0001, Ling Liang 0003, Guanrui Wang, Liang Chang 0002, Xing Hu 0001, Liu Liu 0017, Jing Pei, Guoqi Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Convolution with even-sized kernels and symmetric paddingabstractCompact convolutional neural networks gain efficiency mainly through depthwise convolutions, expanded channels and complex topologies, which contrarily aggravate the training process. Besides, 3x3 kernels dominate the spatial representation in these models, whereas even-sized kernels (2x2, 4x4) are rarely adopted. In this work, we quantify the shift problem occurs in even-sized kernel convolutions by an information erosion hypothesis, and eliminate it by proposing symmetric padding on four sides of the feature maps (C2sp, C4sp). Symmetric padding releases the generalization capabilities of even-sized kernels at little computational cost, making them outperform 3x3 kernels in image classification and generation tasks. Moreover, C2sp obtains comparable accuracy to emerging compact models with much less memory and time consumption during training. Symmetric padding coupled with even-sized convolutions can be neatly implemented into existing frameworks, providing effective elements for architecture designs, especially on online and continual learning occasions where training efforts are emphasized. Guanrui Wang, Pei Tang, Luping Shi |
NeurIPS | 2 |