EDBT 2026 Demo / reviewers in the wild / expert
Dengke Han
dblp:348/4669
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-0641-5779ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LiGNN: Accelerating GNN Training Through Locality-Aware DropoutabstractGraph Neural Networks (GNNs) have demonstrated significant success in graph learning and are widely adopted across various critical domains. However, the irregular connectivity between vertices leads to inefficient neighbor aggregation, resulting in substantial irregular and coarse-grained DRAM accesses. This lack of data locality presents significant challenges for execution platforms, ultimately degrading performance. While previous accelerator designs have leveraged on-chip memory and data access scheduling strategies to address this issue, they still inevitably access features at irregular addresses from DRAM. In this work, we propose LiGNN, a hardware-based solution that enhances locality and applies dropout to aggregation to accelerate GNN training. Unlike algorithmic dropout approaches that primarily focus on improving accuracy and neglects hardware costs, LiGNN is specifically designed to drop nodes' features with data locality awareness, directly targeting the reduction of irregular DRAM accesses, meanwhile maintaining accuracy. LiGNN introduces locality-aware ordering and a DRAM row integrity policy, enabling configurable burst and row-granularity dropout at the DRAM level. This approach improves data locality and ensures more efficient DRAM access. Compared to state-of-the-art methods, under classic 0.5 droprate, LiGNN achieves a 1.62~2.2× speedup, reduces DRAM accesses by 44~50% and DRAM row activation by 41~82%, all without losing accuracy. Gongjian Sun, Mingyu Yan, Dengke Han, Runzhen Xue, Xiaochun Ye, Dongrui Fan |
DATE | 3 |
| 2025 | TLV-HGNN: Thinking Like a Vertex for Memory-Efficient HGNN InferenceabstractHeterogeneous graph neural networks (HGNNs) excel at processing heterogeneous graph data and are widely applied in critical domains. In HGNN inference, the neighbor aggregation stage is the primary performance determinant, yet it suffers from two major sources of memory inefficiency. First, the commonly adopted per-semantic execution paradigm stores intermediate aggregation results for each semantic prior to semantic fusion, causing substantial memory expansion. Second, the aggregation process incurs extensive redundant memory accesses, including repeated loading of target vertex features across semantics and repeated accesses to shared neighbors due to crosssemantic neighborhood overlap. These inefficiencies severely limit scalability and reduce HGNN inference performance. In this work, we first propose a semantics-complete execution paradigm from a vertex perspective that eliminates per-semantic intermediate storage and redundant target vertex accesses. Building on this paradigm, we design TVL-HGNN, a reconfigurable hardware accelerator optimized for efficient aggregation. In addition, we introduce a vertex grouping technique based on crosssemantic neighborhood overlap, with hardware implementation, to reduce redundant accesses to shared neighbors. Experimental results demonstrate that TVL-HGNN achieves average speedups of 7.85× and 1.41× over the NVIDIA A100 GPU and the state-of-the-art HGNN accelerator HiHGNN, respectively, while reducing energy consumption by 98.79 % and 32.61 %. Dengke Han, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ICCD | 1 |
| 2025 | Characterizing and Understanding HGNN Training on GPUsabstractOwing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world domains such as recommendation systems and medical analysis. Prior to their practical application, identifying the optimal HGNN model parameters tailored to specific tasks through extensive training is a time-consuming and costly process. To enhance the efficiency of HGNN training, it is essential to characterize and analyze the execution semantics and patterns within the training process to identify performance bottlenecks. In this study, we conduct a comprehensive quantification and in-depth analysis of two mainstream HGNN training scenarios, including single-GPU and multi-GPU distributed training. Based on the characterization results, we reveal the performance bottlenecks and their underlying causes in different HGNN training scenarios and propose optimization guidelines from both software and hardware perspectives. Dengke Han, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN AccelerationabstractHeterogeneous graph neural networks (HGNNs) have expanded graph representation learning to heterogeneous graph fields. Recent studies have demonstrated their superior performance across various applications, including circuit representation, chip design automation, and placement optimization, often surpassing existing methods. However, GPUs often experience inefficiencies when executing HGNNs due to their unique and complex execution patterns. Compared to traditional graph neural networks (GNNs), these patterns further exacerbate irregularities in memory access. To tackle these challenges, recent studies have focused on developing domain-specific accelerators for HGNNs. Nonetheless, most of these efforts have concentrated on optimizing the datapath or scheduling data accesses, while largely overlooking the potential benefits that could be gained from leveraging the inherent properties of the semantic graph, such as its topology, layout, and generation. In this work, we focus on leveraging the properties of semantic graphs to enhance HGNN performance. First, we analyze the semantic graph build (SGB) stage and identify significant opportunities for data reuse during semantic graph generation. Next, we uncover the phenomenon of buffer thrashing during the graph feature processing (GFP) stage, revealing potential optimization opportunities in semantic graph layout. Furthermore, we propose a lightweight hardware accelerator frontend for HGNNs, called SiHGNN. This accelerator frontend incorporates a tree-based SGB for efficient semantic graph generation and features a novel Graph Restructurer for optimizing semantic graph layouts. Experimental results show that SiHGNN enables the state-of-the-art HGNN accelerator to achieve an average performance improvement of$2.95\times $. Runzhen Xue, Mingyu Yan, Dengke Han, Ziheng Xiao, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and RecouplingabstractHeterogeneous Graph Neural Networks (HGNNs) have broadened the applicability of graph representation learning to heterogeneous graphs. However, the irregular memory access pattern of HGNNs leads to the buffer thrashing issue in HGNN accelerators. Runzhen Xue, Mingyu Yan, Dengke Han, Yihan Teng, Xiaochun Ye, Dongrui Fan |
DAC | 3 |
| 2024 | ADE-HGNN: Accelerating HGNNs Through Attention Disparity Exploitation
Dengke Han, Meng Wu 0006, Runzhen Xue, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
Euro-Par (2) | 1 |
| 2024 | MoDSE: A High-Accurate Multiobjective Design Space Exploration Framework for CPU MicroarchitecturesabstractTo accelerate time-consuming multi-objective design space exploration of CPU microarchitecture, previous work trains prediction models using a set of performance metrics derived from a few simulations, then predicts the rest. Unfortunately, the low accuracy of models limits the exploration effect, and how to achieve a good trade-off between multiple objectives while reducing exploration time is challenging. In this paper, we investigate various prediction models and find out the most accurate basic model. We enhance the model by ensemble learning and generate Pareto-rank-based sample weights to improve prediction accuracy. A hypervolume-improvement-based optimization method to trade off between multiple objectives is proposed together with a uniformity-aware selection algorithm to jump out of the local optimum. Furthermore, the exploration time is reduced owing to a proposed Pareto-aware filter algorithm. Experiments demonstrate that our open-source framework can reduce the distance to the Pareto optimal set by 39% compared with the state-of-the-art framework. Mingyu Yan, Yihan Teng, Dengke Han, Xin Liu 0073, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | HiHGNN: Accelerating HGNNs Through Parallelism and Data Reusability ExploitationabstractHeterogeneous graph neural networks (HGNNs) have emerged as powerful algorithms for processing heterogeneous graphs (HetGs), widely used in many critical fields. To capture both structural and semantic information in HetGs, HGNNs first aggregate the neighboring feature vectors for each vertex in each semantic graph and then fuse the aggregated results across all semantic graphs for each vertex. Unfortunately, existing graph neural network accelerators are ill-suited to accelerate HGNNs. This is because they fail to efficiently tackle the specific execution patterns and exploit the high-degree parallelism as well as data reusability inside and across the processing of semantic graphs in HGNNs. In this work, we first quantitatively characterize a set of representative HGNN models on GPU to disclose the execution bound of each stage, inter-semantic-graph parallelism, and inter-semantic-graph data reusability in HGNNs. Guided by our findings, we propose a high-performance HGNN accelerator, HiHGNN, to alleviate the execution bound and exploit the newfound parallelism and data reusability in HGNNs. Specifically, we first propose a bound-aware stage-fusion methodology that tailors to HGNN acceleration, to fuse and pipeline the execution stages being aware of their execution bounds. Second, we design an independency-aware parallel execution design to exploit the inter-semantic-graph parallelism. Finally, we present a similarity-aware execution scheduling to exploit the inter-semantic-graph data reusability. Compared to the state-of-the-art software framework running on NVIDIA GPU T4 and GPU A100, HiHGNN respectively achieves an average 40.0× and 8.3× speedup as well as 99.59% and 99.74% energy reduction with quintile the memory bandwidth of GPU A100. Runzhen Xue, Dengke Han, Mingyu Yan, Mo Zou, Xiaocheng Yang, John Kim 0001, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | A High-accurate Multi-objective Ensemble Exploration Framework for Design Space of CPU MicroarchitectureabstractTo accelerate the time-consuming multi-objective design space exploration of CPU, previous work trains prediction models using a set of cycle per instruction and power performance metrics derived from a few simulations for sampled design points, then exploits the predicted metrics of the rest design points to perform exploration. Unfortunately, the low accuracy of models limits the exploration effect, and how to balance exploitation and exploration while reducing time is challenging. In this paper, we design an open-source high-accurate multi-objective exploration framework. A bagging ensemble prediction model is designed for high-accurate prediction. An upper confidence bound hypervolume improvement optimization method is proposed to approach the Pareto optimal set and balance exploitation and exploration. A Pareto-aware filter algorithm is proposed to reduce the exploration time. Experiments demonstrate that our framework can reduce the distance to the Pareto optimal set by 17.2%, prediction error by 64.8%, and exploration time by 75.1% compared with the state-of-the-art work. Mingyu Yan, Yihan Teng, Dengke Han, Xiaochun Ye, Dongrui Fan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | A Transfer Learning Framework for High-Accurate Cross-Workload Design Space Exploration of CPUabstractTo perform cross-workload design space exploration of CPU, previous works implicitly transfer knowledge from several existing source workloads and try to make predictions on the target one. However, they do not fully explore the transferability across workloads and their single basic prediction models limit the prediction accuracy. In this paper, an open-source Transfer learning Ensemble Design Space Exploration framework (TrEnDSE) is proposed to perform cross-workload performance predictions. The black-box transferability between workloads is quantitatively dissected and explicitly utilized as sample weights for training. Moreover, an ensemble bagging learning model and an uncertainty-driven iterative optimization method are proposed to perform accurate and robust prediction, with these sample weights leveraged. Experiments on SPEC CPU 2017 demonstrate that TrEnDSE can reduce cycle per instruction prediction error by 54% and power prediction error by 34% compared with the state-of-the-art work. Mingyu Yan, Yihan Teng, Dengke Han, Haoran Dang, Xiaochun Ye, Dongrui Fan |
ICCAD | 4 |