VLDB 2026 Research / reviewers in the wild / expert
Meng Wu 0006
dblp:19/5921-6
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-9841-5962ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A GCN Accelerator with Unified Architecture
Meng Wu 0006, Mingyu Yan, Lei Deng 0003, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan |
ICA3PP (1) | 1 |
| 2025 | Revisiting Edge Perturbation for Graph Neural Network in Graph Data Augmentation and AttackabstractEdge perturbation is a basic method to modify graph structures. It can be categorized into two veins based on their effects on the performance of graph neural networks (GNNs), i.e., graph data augmentation and attack. Surprisingly, both veins of edge perturbation methods employ the same operations, yet yield opposite effects on GNNs' accuracy. A distinct boundary between these methods in using edge perturbation has never been clearly defined. Consequently, inappropriate perturbations may lead to undesirable outcomes, necessitating precise adjustments to achieve desired effects. Therefore, questions of “why edge perturbation has a two-faced effect?” and “what makes edge perturbation flexible and effective?” still remain unanswered. In this paper, we will answer these questions by proposing a unified formulation and establishing a quantizable boundary between two categories of edge perturbation methods. Specifically, we conduct experiments to elucidate the differences and similarities between these methods and theoretically unify the workflow of these methods by casting it to one optimization problem. Then, we devise Edge Priority Detector (EPD) to generate a novel priority metric, bridging these methods up in the workflow. Experiments show that EPD can make augmentation or attack flexibly and achieve comparable or superior performance to other counterparts with less time overhead. Xin Liu 0073, Yuxiang Zhang 0011, Meng Wu 0006, Mingyu Yan, Wei Yan 0005, Shirui Pan, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | DFU-E: A Dataflow Architecture for Edge DSP and AI ApplicationsabstractEdge computing aims to enable swift, real-time data processing, analysis, and storage close to the data source. However, edge computing platforms are often constrained by limited processing power and efficiency. This paper presents DFU-E, a dataflow-based accelerator specifically designed to meet the demands of edge digital signal processing (DSP) and artificial intelligence (AI) applications. Our design addresses real-world requirements with three main innovations. First, to accommodate the diverse algorithms utilized at the edge, we propose a multi-layer dataflow mechanism capable of exploiting task-level, instruction block-level, instruction-level, and data-level parallelism. Second, we develop an edge dataflow architecture that includes a customized processing element (PE) array, memory, and on-chip network microarchitecture optimized for the multi-layer dataflow mechanism. Third, we design an edge dataflow software stack that enables automatic optimizations through operator fusion, dataflow graph mapping, and task scheduling. We utilize representative real-world DSP and AI applications for evaluation. Comparing with Nvidia's state-of-the-art edge computing processor, DFU-E achieves up to 1.42× geometric mean performance improvement and 1.27× energy efficiency improvement. Zhihua Fan, Tianyu Liu 0007, Zhen Wang 0045, Meng Wu 0006, Kunming Zhang, Yanhuan Liu, Ninghui Sun, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Survey on Characterizing and Understanding GNNs From a Computer Architecture PerspectiveabstractCharacterizing and understanding graph neural networks (GNNs) is essential for identifying performance bottlenecks and facilitating their deployment in parallel and distributed systems. Despite substantial work in this area, a comprehensive survey on characterizing and understanding GNNs from a computer architecture perspective is lacking. This article presents a comprehensive survey, proposing a triple-level classification method to categorize, summarize, and compare existing efforts, particularly focusing on their implications for parallel architectures and distributed systems. We identify promising future directions for GNN characterization that align with the challenges of optimizing hardware and software in parallel and distributed systems. Our survey aims to help scholars systematically understand GNN performance bottlenecks and execution patterns from a computer architecture perspective, thereby contributing to the development of more efficient GNN implementations across diverse parallel architectures and distributed systems. Meng Wu 0006, Mingyu Yan, Xiaochun Ye, Dongrui Fan, Yuan Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | GDL-GNN: Applying GPU Dataloading of Large Datasets for Graph Neural Network Inference
Haoran Dang, Meng Wu 0006, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 2 |
| 2024 | ADE-HGNN: Accelerating HGNNs Through Attention Disparity Exploitation
Dengke Han, Meng Wu 0006, Runzhen Xue, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 2 |
| 2024 | Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
Yuxiang Zhang 0011, Xin Liu 0073, Meng Wu 0006, Wei Yan 0005, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 3 |
| 2024 | Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
Meng Wu 0006, Jingkai Qiu, Mingyu Yan, Yang Zhang 0163, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan |
ICA3PP (3) | 1 |
| 2023 | Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output ActivationabstractDeep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2021 | An efficient scheduling algorithm for dataflow architecture using loop-pipelining
Yi Li 0043, Meng Wu 0006, Xiaochun Ye, Hao Zhang 0009, Dongrui Fan |
Inf. Sci. | 2 |
| 2020 | An Efficient Multicast Router using Shared-Buffer with Packet Merging for Dataflow ArchitectureabstractDataflow architecture has native advantages in achieving high instruction parallelism and power efficiency for today's emerging applications such as high performance computing and deep neural network. For the dataflow computing, the execution of instructions is driven by data, so the data transfer efficiency of the network on chip (NoC) is a key factor affecting performance. In the NoC, the latest router uses the multicast routing scheme and output buffer structure to improve network transfer efficiency. However, the effective utilization rate of the router's buffer is low due to the multicast transfer characteristics and unbalanced network load. This observation motivates us to design MRSB, a router architecture that effectively improves buffer utilization by allowing to share data and buffer resources among input ports. As the multicast packet is continuously split during transferring, the effective bandwidth utilization of the packet decreases. Packets with small size waste more buffer cell space, so we expanded packet merging based on MRSB according to the bandwidth occupied by different types of packets. For our experimental workloads, experimental results show that MRSB is 221.48% higher effective buffer utilization and 32.98% less latency than a state-of-the-art router with 31.39% smaller area and 29.14% lower power. The performance of the dataflow accelerator using MRSB is improved by 25.61%, and the average energy of experimental workloads is reduced by 24.27%. Yi Li 0043, Meng Wu 0006, Dongrui Fan, Yuqing Ji, Xiaochun Ye |
NOCS | 2 |
| 2020 | An efficient dataflow accelerator for scientific applications
Xiaochun Ye, Xu Tan 0001, Meng Wu 0006, Yujing Feng, Hao Zhang 0009, Songwen Pei, Dongrui Fan |
Future Gener. Comput. Syst. | 3 |
| 2019 | Applying CNN on a scientific application accelerator based on dataflow architecture
Xiaochun Ye, Taoran Xiang, Xu Tan 0001, Yujing Feng, Meng Wu 0006, Dongrui Fan |
CCF Trans. High Perform. Comput. | 6 |