VLDB 2026 Research / reviewers in the wild / expert
Mingyu Yan
dblp:151/5654
· DBLP profile ↗
37ranked-venue papers
3as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FZKP: Alleviating Dataflow Complexity to Exploit Fine-Grained Parallelism for ZKP AccelerationabstractZero-knowledge proof (ZKP) is a promising cryptographic protocol, but its practical deployment is hindered by the time-consuming proof generation. The proof generation inherently exhibits high-degree parallelism, yet challenges persist in exploiting fine-grained parallelism due to the dataflow complexity, impeding previous work to achieve optimal acceleration. In this work, we propose FZKP, a ZKP accelerator that utilizes two novel fine-grained dataflows coupled with two forward-flow microarchitectures to alleviate dataflow complexity, efficiently exploiting fine-grained parallelism. The proposed dataflows simplify the dataflow pattern for parallel execution, disclosing fine-grained parallelism at a low cost. The microarchitectures employ a base design to handle large bit-width intermediate results for timely consumption. They then replicate and combine the base design following the proposed dataflow to facilitate parallel execution. When evaluated in 12 nm, FZKP achieves an average speedup of 10.3× and 2.2× over the state-of-the-art GPU-based solution and ZKP accelerator on real-world workloads, respectively. Ziheng Xiao, Mingyu Yan, Mingyu Gao 0001, Runzhen Xue, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space ExplorationabstractCross-workload design space exploration (DSE) is crucial in CPU architecture design. Existing DSE methods typically employ the transfer learning technique to leverage knowledge from source workloads, aiming to minimize the requirement of target workload simulation. However, these methods struggle with overfitting, data ambiguity, and workload dissimilarity. To address these challenges, we reframe the cross-workload CPU DSE task as a few-shot meta-learning problem and further introduce MetaDSE. By leveraging model agnostic meta-learning, MetaDSE swiftly adapts to new target workloads, greatly enhancing the efficiency of cross-workload CPU DSE. Additionally, MetaDSE introduces a novel knowledge transfer method called the workload-adaptive architectural mask algorithm, which uncovers the inherent properties of the architecture. Experiments on SPEC CPU 2017 demonstrate that MetaDSE significantly reduces prediction error by 44.3% compared to the state-of-theart. MetaDSE is open-sourced and available at this anonymous GitHub. Runzhen Xue, Hao Wu 0070, Mingyu Yan, Ziheng Xiao, Xiaochun Ye, Dongrui Fan |
DAC | 3 |
| 2025 | LiGNN: Accelerating GNN Training Through Locality-Aware DropoutabstractGraph Neural Networks (GNNs) have demonstrated significant success in graph learning and are widely adopted across various critical domains. However, the irregular connectivity between vertices leads to inefficient neighbor aggregation, resulting in substantial irregular and coarse-grained DRAM accesses. This lack of data locality presents significant challenges for execution platforms, ultimately degrading performance. While previous accelerator designs have leveraged on-chip memory and data access scheduling strategies to address this issue, they still inevitably access features at irregular addresses from DRAM. In this work, we propose LiGNN, a hardware-based solution that enhances locality and applies dropout to aggregation to accelerate GNN training. Unlike algorithmic dropout approaches that primarily focus on improving accuracy and neglects hardware costs, LiGNN is specifically designed to drop nodes' features with data locality awareness, directly targeting the reduction of irregular DRAM accesses, meanwhile maintaining accuracy. LiGNN introduces locality-aware ordering and a DRAM row integrity policy, enabling configurable burst and row-granularity dropout at the DRAM level. This approach improves data locality and ensures more efficient DRAM access. Compared to state-of-the-art methods, under classic 0.5 droprate, LiGNN achieves a 1.62~2.2× speedup, reduces DRAM accesses by 44~50% and DRAM row activation by 41~82%, all without losing accuracy. Gongjian Sun, Mingyu Yan, Dengke Han, Runzhen Xue, Xiaochun Ye, Dongrui Fan |
DATE | 2 |
| 2025 | A GCN Accelerator with Unified Architecture
Meng Wu 0006, Mingyu Yan, Lei Deng 0003, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan |
ICA3PP (1) | 2 |
| 2025 | TLV-HGNN: Thinking Like a Vertex for Memory-Efficient HGNN InferenceabstractHeterogeneous graph neural networks (HGNNs) excel at processing heterogeneous graph data and are widely applied in critical domains. In HGNN inference, the neighbor aggregation stage is the primary performance determinant, yet it suffers from two major sources of memory inefficiency. First, the commonly adopted per-semantic execution paradigm stores intermediate aggregation results for each semantic prior to semantic fusion, causing substantial memory expansion. Second, the aggregation process incurs extensive redundant memory accesses, including repeated loading of target vertex features across semantics and repeated accesses to shared neighbors due to crosssemantic neighborhood overlap. These inefficiencies severely limit scalability and reduce HGNN inference performance. In this work, we first propose a semantics-complete execution paradigm from a vertex perspective that eliminates per-semantic intermediate storage and redundant target vertex accesses. Building on this paradigm, we design TVL-HGNN, a reconfigurable hardware accelerator optimized for efficient aggregation. In addition, we introduce a vertex grouping technique based on crosssemantic neighborhood overlap, with hardware implementation, to reduce redundant accesses to shared neighbors. Experimental results demonstrate that TVL-HGNN achieves average speedups of 7.85× and 1.41× over the NVIDIA A100 GPU and the state-of-the-art HGNN accelerator HiHGNN, respectively, while reducing energy consumption by 98.79 % and 32.61 %. Dengke Han, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ICCD | 3 |
| 2025 | Leveraging Large Language Models for Effective Label-free Node Classification in Text-Attributed GraphsabstractGraph neural networks (GNNs) have become the preferred models for node classification in graph data due to their robust capabilities in integrating graph structures and attributes. However, these models heavily depend on a substantial amount of high-quality labeled data for training, which is often costly to obtain. With the rise of large language models (LLMs), a promising approach is to utilize their exceptional zero-shot capabilities and extensive knowledge for node labeling. Despite encouraging results, this approach either requires numerous queries to LLMs or suffers from reduced performance due to noisy labels generated by LLMs. To address these challenges, we introduce Locle, an active self-training framework that does Label-free nOde Classification with LLMs cost-Effectively. Locle iteratively identifies small sets of ''critical'' samples using GNNs and extracts informative pseudo-labels for them with both LLMs and GNNs, serving as additional supervision signals to enhance model training. Specifically, Locle comprises three key components: (i) an effective active node selection strategy for initial annotations; (ii) a careful sample selection scheme to identify ''critical'' nodes based on label disharmonicity and entropy; and (iii) a label refinement module that combines LLMs and GNNs with a rewired topology. Extensive experiments on five benchmark text-attributed graph datasets demonstrate that Locle significantly outperforms state-of-the-art methods under the same query budget to LLMs in terms of label-free node classification. Notably, on the DBLP dataset with 14.3k nodes, Locle achieves an 8.08% improvement in accuracy over the state-of-the-art at a cost of less than one cent. Our code is available at https://github.com/HKBU-LAGAS/Locle. Taiyan Zhang, Renchi Yang, Yurui Lai, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
SIGIR | 4 |
| 2025 | DropNaE: Alleviating irregularity for large-scale graph representation learning
Xin Liu 0073, Xunbin Xiong, Mingyu Yan, Runzhen Xue, Shirui Pan, Songwen Pei, Lei Deng 0003, Xiaochun Ye, Dongrui Fan |
Neural Networks | 3 |
| 2025 | Characterizing and Understanding HGNN Training on GPUsabstractOwing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world domains such as recommendation systems and medical analysis. Prior to their practical application, identifying the optimal HGNN model parameters tailored to specific tasks through extensive training is a time-consuming and costly process. To enhance the efficiency of HGNN training, it is essential to characterize and analyze the execution semantics and patterns within the training process to identify performance bottlenecks. In this study, we conduct a comprehensive quantification and in-depth analysis of two mainstream HGNN training scenarios, including single-GPU and multi-GPU distributed training. Based on the characterization results, we reveal the performance bottlenecks and their underlying causes in different HGNN training scenarios and propose optimization guidelines from both software and hardware perspectives. Dengke Han, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN AccelerationabstractHeterogeneous graph neural networks (HGNNs) have expanded graph representation learning to heterogeneous graph fields. Recent studies have demonstrated their superior performance across various applications, including circuit representation, chip design automation, and placement optimization, often surpassing existing methods. However, GPUs often experience inefficiencies when executing HGNNs due to their unique and complex execution patterns. Compared to traditional graph neural networks (GNNs), these patterns further exacerbate irregularities in memory access. To tackle these challenges, recent studies have focused on developing domain-specific accelerators for HGNNs. Nonetheless, most of these efforts have concentrated on optimizing the datapath or scheduling data accesses, while largely overlooking the potential benefits that could be gained from leveraging the inherent properties of the semantic graph, such as its topology, layout, and generation. In this work, we focus on leveraging the properties of semantic graphs to enhance HGNN performance. First, we analyze the semantic graph build (SGB) stage and identify significant opportunities for data reuse during semantic graph generation. Next, we uncover the phenomenon of buffer thrashing during the graph feature processing (GFP) stage, revealing potential optimization opportunities in semantic graph layout. Furthermore, we propose a lightweight hardware accelerator frontend for HGNNs, called SiHGNN. This accelerator frontend incorporates a tree-based SGB for efficient semantic graph generation and features a novel Graph Restructurer for optimizing semantic graph layouts. Experimental results show that SiHGNN enables the state-of-the-art HGNN accelerator to achieve an average performance improvement of$2.95\times $. Runzhen Xue, Mingyu Yan, Dengke Han, Ziheng Xiao, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Revisiting Edge Perturbation for Graph Neural Network in Graph Data Augmentation and AttackabstractEdge perturbation is a basic method to modify graph structures. It can be categorized into two veins based on their effects on the performance of graph neural networks (GNNs), i.e., graph data augmentation and attack. Surprisingly, both veins of edge perturbation methods employ the same operations, yet yield opposite effects on GNNs' accuracy. A distinct boundary between these methods in using edge perturbation has never been clearly defined. Consequently, inappropriate perturbations may lead to undesirable outcomes, necessitating precise adjustments to achieve desired effects. Therefore, questions of “why edge perturbation has a two-faced effect?” and “what makes edge perturbation flexible and effective?” still remain unanswered. In this paper, we will answer these questions by proposing a unified formulation and establishing a quantizable boundary between two categories of edge perturbation methods. Specifically, we conduct experiments to elucidate the differences and similarities between these methods and theoretically unify the workflow of these methods by casting it to one optimization problem. Then, we devise Edge Priority Detector (EPD) to generate a novel priority metric, bridging these methods up in the workflow. Experiments show that EPD can make augmentation or attack flexibly and achieve comparable or superior performance to other counterparts with less time overhead. Xin Liu 0073, Yuxiang Zhang 0011, Meng Wu 0006, Mingyu Yan, Wei Yan 0005, Shirui Pan, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Survey on Characterizing and Understanding GNNs From a Computer Architecture PerspectiveabstractCharacterizing and understanding graph neural networks (GNNs) is essential for identifying performance bottlenecks and facilitating their deployment in parallel and distributed systems. Despite substantial work in this area, a comprehensive survey on characterizing and understanding GNNs from a computer architecture perspective is lacking. This article presents a comprehensive survey, proposing a triple-level classification method to categorize, summarize, and compare existing efforts, particularly focusing on their implications for parallel architectures and distributed systems. We identify promising future directions for GNN characterization that align with the challenges of optimizing hardware and software in parallel and distributed systems. Our survey aims to help scholars systematically understand GNN performance bottlenecks and execution patterns from a computer architecture perspective, thereby contributing to the development of more efficient GNN implementations across diverse parallel architectures and distributed systems. Meng Wu 0006, Mingyu Yan, Xiaochun Ye, Dongrui Fan, Yuan Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and RecouplingabstractHeterogeneous Graph Neural Networks (HGNNs) have broadened the applicability of graph representation learning to heterogeneous graphs. However, the irregular memory access pattern of HGNNs leads to the buffer thrashing issue in HGNN accelerators. Runzhen Xue, Mingyu Yan, Dengke Han, Yihan Teng, Xiaochun Ye, Dongrui Fan |
DAC | 2 |
| 2024 | GDL-GNN: Applying GPU Dataloading of Large Datasets for Graph Neural Network Inference
Haoran Dang, Meng Wu 0006, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 3 |
| 2024 | ADE-HGNN: Accelerating HGNNs Through Attention Disparity Exploitation
Dengke Han, Meng Wu 0006, Runzhen Xue, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 4 |
| 2024 | Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
Yuxiang Zhang 0011, Xin Liu 0073, Meng Wu 0006, Wei Yan 0005, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 5 |
| 2024 | Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
Meng Wu 0006, Jingkai Qiu, Mingyu Yan, Yang Zhang 0163, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan |
ICA3PP (3) | 3 |
| 2024 | MoDSE: A High-Accurate Multiobjective Design Space Exploration Framework for CPU MicroarchitecturesabstractTo accelerate time-consuming multi-objective design space exploration of CPU microarchitecture, previous work trains prediction models using a set of performance metrics derived from a few simulations, then predicts the rest. Unfortunately, the low accuracy of models limits the exploration effect, and how to achieve a good trade-off between multiple objectives while reducing exploration time is challenging. In this paper, we investigate various prediction models and find out the most accurate basic model. We enhance the model by ensemble learning and generate Pareto-rank-based sample weights to improve prediction accuracy. A hypervolume-improvement-based optimization method to trade off between multiple objectives is proposed together with a uniformity-aware selection algorithm to jump out of the local optimum. Furthermore, the exploration time is reduced owing to a proposed Pareto-aware filter algorithm. Experiments demonstrate that our open-source framework can reduce the distance to the Pareto optimal set by 39% compared with the state-of-the-art framework. Mingyu Yan, Yihan Teng, Dengke Han, Xin Liu 0073, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | HiHGNN: Accelerating HGNNs Through Parallelism and Data Reusability ExploitationabstractHeterogeneous graph neural networks (HGNNs) have emerged as powerful algorithms for processing heterogeneous graphs (HetGs), widely used in many critical fields. To capture both structural and semantic information in HetGs, HGNNs first aggregate the neighboring feature vectors for each vertex in each semantic graph and then fuse the aggregated results across all semantic graphs for each vertex. Unfortunately, existing graph neural network accelerators are ill-suited to accelerate HGNNs. This is because they fail to efficiently tackle the specific execution patterns and exploit the high-degree parallelism as well as data reusability inside and across the processing of semantic graphs in HGNNs. In this work, we first quantitatively characterize a set of representative HGNN models on GPU to disclose the execution bound of each stage, inter-semantic-graph parallelism, and inter-semantic-graph data reusability in HGNNs. Guided by our findings, we propose a high-performance HGNN accelerator, HiHGNN, to alleviate the execution bound and exploit the newfound parallelism and data reusability in HGNNs. Specifically, we first propose a bound-aware stage-fusion methodology that tailors to HGNN acceleration, to fuse and pipeline the execution stages being aware of their execution bounds. Second, we design an independency-aware parallel execution design to exploit the inter-semantic-graph parallelism. Finally, we present a similarity-aware execution scheduling to exploit the inter-semantic-graph data reusability. Compared to the state-of-the-art software framework running on NVIDIA GPU T4 and GPU A100, HiHGNN respectively achieves an average 40.0× and 8.3× speedup as well as 99.59% and 99.74% energy reduction with quintile the memory bandwidth of GPU A100. Runzhen Xue, Dengke Han, Mingyu Yan, Mo Zou, Xiaocheng Yang, John Kim 0001, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Simple and Efficient Heterogeneous Graph Neural NetworkabstractHeterogeneous graph neural networks (HGNNs) have the powerful capability to embed rich structural and semantic information of a heterogeneous graph into node representations. Existing HGNNs inherit many mechanisms from graph neural networks (GNNs) designed for homogeneous graphs, especially the attention mechanism and the multi-layer structure. These mechanisms bring excessive complexity, but seldom work studies whether they are really effective on heterogeneous graphs. In this paper, we conduct an in-depth and detailed study of these mechanisms and propose the Simple and Efficient Heterogeneous Graph Neural Network (SeHGNN). To easily capture structural information, SeHGNN pre-computes the neighbor aggregation using a light-weight mean aggregator, which reduces complexity by removing overused neighbor attention and avoiding repeated neighbor aggregation in every training epoch. To better utilize semantic information, SeHGNN adopts the single-layer structure with long metapaths to extend the receptive field, as well as a transformer-based semantic fusion module to fuse features from different metapaths. As a result, SeHGNN exhibits the characteristics of a simple network structure, high prediction accuracy, and fast training speed. Extensive experiments on five real-world heterogeneous graphs demonstrate the superiority of SeHGNN over the state-of-the-arts on both accuracy and training speed. Xiaocheng Yang, Mingyu Yan, Shirui Pan, Xiaochun Ye, Dongrui Fan |
AAAI | 2 |
| 2023 | A High-accurate Multi-objective Exploration Framework for Design Space of CPUabstractTo accelerate time-consuming multi-objective design space exploration of CPU, previous work trains prediction models using a set of performance metrics derived from few simulations, then predicts the rest. Unfortunately, the low accuracy of models limits the exploration effect, and how to achieve a good trade-off between multiple objectives is challenging.In this paper, we investigate various prediction models and find out the most accurate basic model. We enhance the model by ensemble learning to improve prediction accuracy. A hypervolume-improvement-based optimization method to trade off between multiple objectives is proposed together with a uniformity-aware selection algorithm to jump out of the local optimum. Experiments demonstrate that our open-source framework can reduce the distance to the Pareto optimal set by 76% and prediction error by 97% compared with the state-of-the-art work. Mingyu Yan, Xin Liu 0073, Mo Zou, Tianyu Liu 0007, Xiaochun Ye, Dongrui Fan |
DAC | 2 |
| 2023 | A High-accurate Multi-objective Ensemble Exploration Framework for Design Space of CPU MicroarchitectureabstractTo accelerate the time-consuming multi-objective design space exploration of CPU, previous work trains prediction models using a set of cycle per instruction and power performance metrics derived from a few simulations for sampled design points, then exploits the predicted metrics of the rest design points to perform exploration. Unfortunately, the low accuracy of models limits the exploration effect, and how to balance exploitation and exploration while reducing time is challenging. In this paper, we design an open-source high-accurate multi-objective exploration framework. A bagging ensemble prediction model is designed for high-accurate prediction. An upper confidence bound hypervolume improvement optimization method is proposed to approach the Pareto optimal set and balance exploitation and exploration. A Pareto-aware filter algorithm is proposed to reduce the exploration time. Experiments demonstrate that our framework can reduce the distance to the Pareto optimal set by 17.2%, prediction error by 64.8%, and exploration time by 75.1% compared with the state-of-the-art work. Mingyu Yan, Yihan Teng, Dengke Han, Xiaochun Ye, Dongrui Fan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | A Transfer Learning Framework for High-Accurate Cross-Workload Design Space Exploration of CPUabstractTo perform cross-workload design space exploration of CPU, previous works implicitly transfer knowledge from several existing source workloads and try to make predictions on the target one. However, they do not fully explore the transferability across workloads and their single basic prediction models limit the prediction accuracy. In this paper, an open-source Transfer learning Ensemble Design Space Exploration framework (TrEnDSE) is proposed to perform cross-workload performance predictions. The black-box transferability between workloads is quantitatively dissected and explicitly utilized as sample weights for training. Moreover, an ensemble bagging learning model and an uncertainty-driven iterative optimization method are proposed to perform accurate and robust prediction, with these sample weights leveraged. Experiments on SPEC CPU 2017 demonstrate that TrEnDSE can reduce cycle per instruction prediction error by 54% and power prediction error by 34% compared with the state-of-the-art work. Mingyu Yan, Yihan Teng, Dengke Han, Haoran Dang, Xiaochun Ye, Dongrui Fan |
ICCAD | 2 |
| 2023 | A Comprehensive Survey on Distributed Training of Graph Neural NetworksabstractGraph neural networks (GNNs) have been demonstrated to be a powerful algorithmic model in broad application fields for their effectiveness in learning over graphs. To scale GNN training up for large-scale and ever-growing graphs, the most promising solution is distributed training that distributes the workload of training across multiple computing nodes. At present, the volume of related research on distributed GNN training is exceptionally vast, accompanied by an extraordinarily rapid pace of publication. Moreover, the approaches reported in these studies exhibit significant divergence. This situation poses a considerable challenge for newcomers, hindering their ability to grasp a comprehensive understanding of the workflows, computational patterns, communication strategies, and optimization techniques employed in distributed GNN training. As a result, there is a pressing need for a survey to provide correct recognition, analysis, and comparisons in this field. In this article, we provide a comprehensive survey of distributed GNN training by investigating various optimization techniques used in distributed GNN training. First, distributed GNN training is classified into several categories according to their workflows. In addition, their computational patterns and communication patterns, as well as the optimization techniques proposed by recent work, are introduced. Second, the software frameworks and hardware platforms of distributed GNN training are also introduced for a deeper understanding. Third, distributed GNN training is compared with distributed training of deep neural networks (DNNs), emphasizing the uniqueness of distributed GNN training. Finally, interesting issues and opportunities in this field are discussed. Haiyang Lin, Mingyu Yan, Xiaochun Ye, Dongrui Fan, Shirui Pan, Yuan Xie 0001 |
Proc. IEEE | 2 |
| 2022 | Alleviating datapath conflicts and design centralization in graph analytics accelerationabstractPrevious graph analytics accelerators have achieved great improvement on throughput by alleviating irregular off-chip memory accesses. However, on-chip side datapath conflicts and design centralization have become the critical issues hindering further throughput improvement. In this paper, a general solution, Multiple-stage Decentralized Propagation network (MDP-network), is proposed to address these issues, inspired by the key idea of trading latency for throughput. Besides, a novel High throughput Graph analytics accelerator, HiGraph, is proposed by deploying MDP-network to address each issue in practice. The experiment shows that compared with state-of-the-art accelerator, HiGraph achieves up to 2.2× speedup (1.5× on average) as well as better scalability. Haiyang Lin, Mingyu Yan, Mo Zou, Fengbin Tu, Xiaochun Ye, Dongrui Fan, Yuan Xie 0001 |
DAC | 2 |
| 2022 | HetGraph: A High Performance CPU-CGRA Architecture for Matrix-based Graph AnalyticsabstractIn this paper, we explore graph analytics on a heterogeneous platform named HetGraph integrating with CPU and a flexible CGRA accelerator called RFU for matrix-based paradigm in this paper. RFU utilizes the lightweight pipeline without data hazards to support various generalized Sparse Matrix-Vector multiplications (SpMVs) of matrix-based graph analytics effectively. HetGraph utilizes the degree-aware workload distribution with vector-scanning sparsity removing scheme to alleviate the impact of highly sparse graph. Furthermore, we propose a heterogeneous work-stealing strategy to balance the workloads between CPU and RFU for HetGraph. To the best of our knowledge, HetGraph is the first heterogeneous CPU-CGRA architecture for matrix-based graph analytics. Overall, HetGraph achieves 9.42x, 2.45x speedup, and 9.80x, 7.70x energy savings on average compared to state-of-the-art (SOTA) CPU-based and GPGPU-based solutions respectively. Compared to the SOTA graph analytics accelerator, HetGraph also achieves 1.42x speedup and 1.06x less energy. Long Tan, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | MatGraph: An Energy-Efficient and Flexible CGRA Engine for Matrix-Based Graph Analytics
Long Tan, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ICA3PP | 2 |
| 2022 | GEM: Execution-Aware Cache Management for Graph Analytics
Mo Zou, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ICA3PP | 2 |
| 2022 | Survey on Graph Neural Network Acceleration: An Algorithmic PerspectiveabstractGraph neural networks (GNNs) have been a hot spot of recent research and are widely utilized in diverse applications. However, with the use of huger data and deeper models, an urgent demand is unsurprisingly made to accelerate GNNs for more efficient execution. In this paper, we provide a comprehensive survey on acceleration methods for GNNs from an algorithmic perspective. We first present a new taxonomy to classify existing acceleration methods into five categories. Based on the classification, we systematically discuss these methods and highlight their correlations. Next, we provide comparisons from aspects of the efficiency and characteristics of these methods. Finally, we suggest some promising prospects for future research. Xin Liu 0073, Mingyu Yan, Lei Deng 0003, Guoqi Li 0002, Xiaochun Ye, Dongrui Fan, Shirui Pan, Yuan Xie 0001 |
IJCAI | 2 |
| 2022 | GNNSampler: Bridging the Gap Between Sampling Algorithms of GNN and Hardware
Xin Liu 0073, Mingyu Yan, Shuhan Song, Zhengyang Lv, Guangyu Sun 0003, Xiaochun Ye, Dongrui Fan |
ECML/PKDD (5) | 2 |
| 2022 | Multi-Node Acceleration for Large-Scale GCNsabstractLimited by the memory capacity and compute power, singe-node graph convolutional neural network (GCN) accelerators cannot complete the execution of GCNs within a reasonable amount of time, due to the explosive size of graphs nowadays. Thus, large-scale GCNs call for a multi-node acceleration system (MultiAccSys) like TPU-Pod for large-scale neural networks. In this work, we aim to scale up single-node GCN accelerators to accelerate GCNs on large-scale graphs. We first identify the communication pattern and challenges of multi-node acceleration for GCNs on large-scale graphs. We observe that (1) coarse-grained communication patterns exist in the execution of GCNs in MultiAccSys, which introduces massive amount of redundant network transmissions and off-chip memory accesses; (2) overall, the acceleration of GCNs in MultiAccSys is bandwidth-bound and latency-tolerant. Guided by these two observations, we then propose MultiGCN, the first MultiAccSys for large-scale GCNs that trades network latency for network bandwidth. Specifically, by leveraging the network latency tolerance, wefirstpropose a topology-aware multicast mechanism with a oneputpermulticastmessage-passing model to reduce transmissions and alleviate network bandwidth requirements.Second, we introduce a scatter-based round execution mechanism which cooperates with the multicast mechanism and reduces redundant off-chip memory accesses. Compared to the baseline MultiAccSys, MultiGCN achieves 4$\sim 12\times$speedup using only 28%$\sim$68% energy, while reducing 32% transmissions and 73% off-chip memory accesses on average. It not only achieves 2.5$\sim 8\times$speedup over the state-of-the-art multi-GPU solution, but also scales to large-scale graphs as opposed to single-node GCN accelerators. Gongjian Sun, Mingyu Yan, Han Li 0011, Xiaochun Ye, Dongrui Fan, Yuan Xie 0001 |
IEEE Trans. Computers | 2 |
| 2022 | Rubik: A Hierarchical Architecture for Efficient Graph Neural Network TrainingabstractThe graph convolutional network (GCN) emerges as a promising direction to learn the inductive representation in graph data commonly used in widespread applications, such as E-commerce, social networks, and knowledge graphs. However, learning from graphs is nontrivial because of its mixed computation model involving both graph analytics and neural network computing. To this end, we decompose the GCN learning into two hierarchical paradigms: 1) graph-level and 2) node-level computing. Such a hierarchical paradigm facilitates the software and hardware accelerations for GCN learning. We propose a lightweight graph reordering methodology, incorporated with a GCN accelerator architecture that equips a customized cache design to fully utilize the graph-level data reuse. We also propose a mapping methodology aware of data reuse and task-level parallelism to handle various graphs inputs effectively. The results show that Rubik accelerator design improves energy efficiency by$26.3\times $–$1375.2\times $than GPU platforms across different datasets and GCN models. Xiaobing Chen, Xinfeng Xie, Xing Hu 0001, Abanti Basak, Ling Liang 0003, Mingyu Yan, Lei Deng 0003, Yufei Ding 0001, Zidong Du, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | Triangle Counting by Adaptively Resampling over Evolving Graph StreamsabstractTriangle counting is a fundamental graph mining problem, widely used in many real-world application scenarios.Due to the large scale of graph streams and limited memory space, it is appropriate to achieve the estimation of global and local triangles by sampling.Existing streaming algorithms for triangle counting can be generalized into two categories.One is Reservoir-based methods employing a fixed memory budget, whose size is difficult to set for accurate estimation without any prior knowledge about graph streams.The other is Bernoullibased methods, which sample edges by a given probability with uncontrollable memory budget.In this work, we propose a novel and bounded-sampling-ratio method, called BSR-Sample, by adaptively resizing memory budget upwards over evolving graph streams.BSR-Sample can keep the sampling ratio always greater than or equal to a specified threshold with available memory space.Then, we design BSR-TC, a single-pass streaming algorithm for both global and local triangle counting, based on BSR-Sample.Experimental results show that BSR-TC achieves accuracy of at least 99.8% for global triangles, when the ratio of initial memory budget to whole graph streams ≥ 0.002% and given threshold = 20%.And our proposed BSR-TC can gain more advantage than the state-of-the-art algorithms over the continuous growth of graph streams. Huawei Cao, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
SEKE | 3 |
| 2021 | BSR-TC: Adaptively Sampling for Accurate Triangle Counting over Evolving Graph StreamsabstractTriangle counting is a fundamental graph mining problem, widely employed in various real-world application scenarios. Given the large scale of graph streams and limited memory space, it is feasible to achieve the estimation of global and local triangles by sampling. Existing streaming algorithms for triangle counting can be generalized into two categories: Reservoir-based methods and Bernoulli-based methods. The former use a fixed memory budget, whose size is difficult to set for accurate estimation without any prior knowledge about graph streams. The latter sample edges by a specified probability, but memory budget is uncontrollable for following a binomial distribution. In this work, we propose a novel and bounded-sampling-ratio algorithm for both global and local triangle counting, called BSR-TC, by adaptively resizing memory budget upwards over evolving graph streams. Specifically, our proposed single-pass BSR-TC can gain more advantage than the state-of-the-art algorithms over the continuous growth of graph streams. Experimental results show that BSR-TC achieves accuracy of at least 99.8% for global triangles, when the ratio of initial memory budget against whole graph streams [Formula: see text] and given [Formula: see text], respectively. Huawei Cao, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2020 | HyGCN: A GCN Accelerator with Hybrid ArchitectureabstractInspired by the great success of neural networks, graph convolutional neural networks (GCNs) are proposed to analyze graph data. GCNs mainly include two phases with distinct execution patterns. The Aggregation phase, behaves as graph processing, showing a dynamic and irregular execution pattern. The Combination phase, acts more like the neural networks, presenting a static and regular execution pattern. The hybrid execution patterns of GCNs require a design that alleviates irregularity and exploits regularity. Moreover, to achieve higher performance and energy efficiency, the design needs to leverage the high intra-vertex parallelism in Aggregation phase, the highly reusable inter-vertex data in Combination phase, and the opportunity to fuse phase-by-phase execution introduced by the new features of GCNs. However, existing architectures fail to address these demands. In this work, we first characterize the hybrid execution patterns of GCNs on Intel Xeon CPU. Guided by the characterization, we design a GCN accelerator, HyGCN, using a hybrid architecture to efficiently perform GCNs. Specifically, first, we build a new programming model to exploit the fine-grained parallelism for our hardware design. Second, we propose a hardware design with two efficient processing engines to alleviate the irregularity of Aggregation phase and leverage the regularity of Combination phase. Besides, these engines can exploit various parallelism and reuse highly reusable data efficiently. Third, we optimize the overall system via inter-engine pipeline for inter-phase fusion and priority-based off-chip memory access coordination to improve off-chip bandwidth utilization. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA V100 GPU, our work achieves on average 1509× speedup with 2500× energy reduction and average 6.5× speedup with 10× energy reduction, respectively. Mingyu Yan, Lei Deng 0003, Xing Hu 0001, Ling Liang 0003, Yujing Feng, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001 |
HPCA | 1 |
| 2020 | fuseGNN: Accelerating Graph Convolutional Neural Network Training on GPGPUabstractGraph convolutional neural networks (GNN) have achieved state-of-the-art performance on tasks like node classification. It has become a new workload family member in data-centers. GNN works on irregular graph-structured data with three distinct phases: Combination, Graph Processing, and Aggregation. While Combination phase has been well supported by sgemm kernels in cuBLAS, the other two phases are still inefficient on GPGPU due to the lack of optimized CUDA kernels. In particular, Aggregation phase introduces large volume of DRAM storage footprint and data movement, and both Aggregation and Graph Processing phases suffer from high kernel launching time. These inefficiencies not only decrease training throughput but also limit users from training GNNs on larger graphs on GPGPU. Although these problems have been partially alleviated by recent studies, their optimizations are still not sufficient. In this paper, we propose fuseGNN, an extension of PyTorch that provides highly optimized APIs and CUDA kernels for GNN. First, two different programming abstractions for Aggregation phase are utilized to handle graphs with different average degrees. Second, dedicated GPGPU kernels are developed for Aggregation and Graph Processing in both forward and backward passes, in which kernel-fusion along with other optimization strategies are applied to reduce kernel launching time and latency as well as exploit data reuse opportunities. Evaluation on multiple benchmarks shows that fuseGNN achieves up to 5.3× end-to-end speedup over state-of-the-art frameworks, and the DRAM storage footprint is reduced by several orders of magnitude on large datasets. Zhaodong Chen 0001, Mingyu Yan, Maohua Zhu, Lei Deng 0003, Guoqi Li 0002, Shuangchen Li, Yuan Xie 0001 |
ICCAD | 2 |
| 2019 | Balancing Memory Accesses for Energy-Efficient Graph Analytics AcceleratorsabstractDomain-specific accelerators for graph analytics leverage a large on-chip memory in order to tackle the intensive random memory accesses, offering higher performance and energy efficiency than conventional architectures. However, limited by the inefficient usage of on-chip memory, current accelerators suffer from energy and performance bottlenecks due to the large amount of off-chip memory accesses. In this work, we introduce an online preprocessing step for the vertex-centric programming model based on our observation of imbalanced memory bandwidth utilization between two execution phases. Our scheme improves energy efficiency and performance by significantly reducing off-chip accesses in two ways. First, we sequence random off-chip memory accesses to balance memory bandwidth demands and improve the utilization of on-chip memory. Second, we prune active leaf vertices to avoid redundant memory accesses. We evaluate our method on a state-of-the-art graph analytics accelerator and achieve 1.6× speedup while reducing energy consumption by 42% on average. Mingyu Yan, Xing Hu 0001, Shuangchen Li, Itir Akgun, Han Li 0011, Lei Deng 0003, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001 |
ISLPED | 1 |
| 2019 | Alleviating Irregularity in Graph Analytics Acceleration: a Hardware/Software Co-Design ApproachabstractGraph analytics is an emerging application which extracts insights by processing large volumes of highly connected data, namely graphs. The parallel processing of graphs has been exploited at the algorithm level, which in turn incurs three irregularities onto computing and memory patterns that significantly hinder an efficient architecture design. Certain irregularities can be partially tackled by the prior domain-specific accelerator designs with well-designed scheduling of data access, while others remain unsolved. Mingyu Yan, Xing Hu 0001, Shuangchen Li, Abanti Basak, Han Li 0011, Itir Akgun, Yujing Feng, Peng Gu 0008, Lei Deng 0003, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001 |
MICRO | 1 |