Hengliang Guo

dblp:261/8385 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-5796-233XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Optimizing sparse-dense matrix-matrix multiplication for DCUs
Hengliang Guo, Haolei Wang, Shengguang Zhu, Chuanqiang Li
CCF Trans. High Perform. Comput.1
2026 A KAN-enhanced graphSAGE model for ethereum account classification on heterophilic graphs
abstract
Ethereum account classification is essential for identifying individuals engaged in illicit transactions and analyzing behavioral patterns across various account types. This process serves as a critical mechanism for monitoring and regulating unlawful activities within transactional markets. However, the Ethereum network exhibits the characteristics of a complex heterophilic graph which poses significant challenges to the effectiveness and performance of conventional graph neural networks (GNNs). To address this challenge, the present study proposes FSGCN(Fourier-Sage GCN), a novel architecture for heterophilic graph neural networks (GNNs) that integrates Kolmogorov–Arnold Networks (KANs) with GraphSAGE. FSGCN is specifically designed to adapt efficiently to the structural complexity of heterophilic graphs. By leveraging KANs to extract high-order neighborhood information and employing GraphSAGE to capture low-order neighborhood patterns, FSGCN effectively aggregates both homophilic and heterophilic features, thereby improving classification performance. Furthermore, to improve training efficiency and generalization, we propose the MLPInit weight initialization scheme and the DropEdge graph augmentation technique. Experiments on a large-scale Ethereum transaction dataset show that FSGCN achieves an F1-score of 91.8% and a classification accuracy of 91.6%, significantly outperforming traditional homophilic and heterophilic GNN baselines. Additionally, FSGCN demonstrates high training efficiency, completing each epoch in just 2.302 s per epoch and improving overall training speed by 130.4% compared to conventional GraphSAGE.
Hengliang Guo, Yizhe Sui, Jiaru Li, Fuchang Gao
Peer Peer Netw. Appl.1
2025 Best-Fit Document: Enhancing Compositional Generalization in Multi-label Text Classification
Hengliang Guo, Shengguang Zhu, Jiaru Li, Ruikai Ma
NLPCC (3)1
2025 Optimizing 2D convolution for DCUs
Wenlong Fan, Haobo Hua, Jiandong Shang, Zhuxin Wen, Hengliang Guo, Litao Zhang
CCF Trans. High Perform. Comput.5
2025 GeoProspect: A domain-specific geological large language model with enhanced continual learning
Kunyan Zhang, Shengguang Zhu, Mengzhe Fan, Hengliang Guo, Gubin Zhang, Dujuan Zhang, Haitao Wei
Neurocomputing8
2025 CEGT: Smart contract vulnerability detection via Connectivity-Enhanced GCN-Transformer
Jiandong Shang, Jiaru Li, Yizhe Sui, Hengliang Guo, Dujuan Zhang
J. Syst. Softw.4
2025 VBATS: an adaptive strategy for grouped GEMM on GPUs
Jiandong Shang, Zhuxin Wen, Haobo Hua, Hengliang Guo, Wenlong Fan, Guangsheng Qin
J. Supercomput.4
2024 Optimizing sparse general matrix-matrix multiplication for DCUs
abstract
Abstract Sparse general matrix–matrix multiplication (SpGEMM) is a crucial and complex computational task in many practical applications. Improving the performance of SpGEMM on SIMT processors like modern GPUs is challenging due to the unpredictable sparsity of sparse matrices. Although existing GPU solutions have made progress in improving performance through advanced algorithm design, they ignore some optimizations related to specific processor architectures. This can result in a partially inefficient implementation of their algorithms. This paper focuses on optimizing four inefficient parts of the NSparse algorithm on DCU (a GPU-like accelerator). The optimizations include: 1) setting parameters to improve the load balance of the second matrix by extracting maximum row information at runtime; 2) reducing overhead of binning operations by making full use of registers and shared memory effectively; 3) improving numerical SpGEMM performance by adjusting its calculation mode; and 4) enhancing global load balance through finer-grained grouping and kernel configurations. Experiment results demonstrate that when compared to five state-of-the-art SpGEMM algorithms (bhSparse, KokkosKernels, NSparse, rocSparse, and spECK), our optimized method achieves an average of 7.99x (up to 18.2x), 8.01x (up to 20.83x), 2.37x (up to 6.16x), 1.82x (up to 4.20x), and 1.63x (up to 5.01x) speedups on 29 sparse matrices with different sparse structures, respectively.
Hengliang Guo, Haolei Wang, Wanting Chen, Congxiang Zhang, Shengguang Zhu, Dujuan Zhang, Jiandong Shang
J. Supercomput.1
2024 OpenMP offloading data transfer optimization for DCUs
abstract
Abstract OpenMP supports the use of target offloading compile guidance instructions to invoke heterogeneous-platform accelerators to compute core code segments; however, unreasonable use of target offloading instructions can make the data transfer process time-consuming. The problem of unused array transfer and unused data segment transfer arises when the amount of data transferred from the host side to the device side exceeds the amount of data required for the core computation on the device side. For the transmission of unused arrays, the use of the transmitted arrays is guided by adding a filter to eliminate the transmission of redundant data; for the transmission of unused data segments, the use of arrays is quickly determined on the basis of the filter, and valid data are transmitted by optimizing Clang’s code generation strategy after obtaining the lengths of the data segments in core computation. Experiments are performed using the Polybench benchmark; the optimized speedup for unused array transfer reaches 7%, and the optimized speedup for unused data segment transfer reaches 10%. The experimental results show that data transfer optimization for target offloading characteristics can help improve program performance.
Hengliang Guo, Xiaoyue Xu, Kuangsheng Cai, Shuxin Yang, Lingbo Kong
J. Supercomput.1
2021 Robust Deep Neural Networks for Road Extraction From Remote Sensing Images
abstract
The application of deep neural networks (DNNs) for road extraction from remote sensing images has gained broad interest because of the competence concerning complex nonlinear relations; however, the presence of noisy labels in the training data sets adversely affects the performance of DNNs. The existing methods of improving the robustness of DNNs focus on modeling the noise distribution. However, these approaches are not satisfactory because of the inaccurate high-level image features obtained by the DNNs. To address this issue, we develop a noise probabilistic model for learning the label noise based on the relationship between the input images, noisy labels, and true labels. The key idea of the probabilistic model is to directly explore the information from the input images and apply it to model the label noise. Then, a robust deep neural network (RDNN) is proposed to instantiate the noise probabilistic model, which consists of two important modules: the true label predictor (TLP) and the noise label estimator (NLE). Especially, the TLP is made of a DNN with softmax, which is used to learn the true label distribution. The NLE is applied to model the label noise distribution, which aims to absorb the label noise in the training process. Moreover, to tackle the challenges in the optimization, we deduce a loss function with the novel regularization, which allows the RDNN to conduct effective training on the noise data set. The effectiveness of the proposed method is validated by experiments on three road data sets that contain various resolutions and imaging conditions. The results demonstrate its superiority over state-of-the-art methods in visual performance and classification accuracy.
Panle Li, Xiaohui He 0001, Mengjia Qiao, Xijie Cheng, Haotian Luo, Dingjun Song, Daidong Li, Shaokai Hu, Runchuan Li, Pu Han, Fangbing Qiu, Hengliang Guo, Jiandong Shang, Zhihui Tian
IEEE Trans. Geosci. Remote. Sens.13