Yuedan Chen

dblp:201/1805 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0001-5665-268XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Mitigating Channel Redundancy for Multivariate Time Series Forecasting
abstract
Transformer-based methods have been widely used in multivariate time series forecasting (MTSF), typically following either channel-independent (CI) or channel-dependent (CD) mod-eling approaches. However, CI methods overlook inter-channel information, while CD methods struggle with the complexity and redundancy of channel relationships, leading to suboptimal performance and high computational costs. In this paper, we propose a novel Channel Aggregation Network, dubbed CANet, to efficiently model both intra and inter-channel dependencies. Specifically, CANet embeds input sequences into patch tokens and uses probabilistic masking in the Channel Aggregator to effectively filter out redundant and noisy information. The refined inter-channel features are then injected into the temporal dimension, allowing CANet to focus on temporal attention while efficiently leveraging cross-channel information. After that, the Temporal Sampler selects key temporal tokens along the time axis, enhancing long-range dependency modeling and accelerating global representation learning. Extensive experiments on multiple real-world datasets demonstrate that CANet achieves state-of-the-art performance with superior accuracy and significantly reduced computational complexity.
Guoqing Xiao 0001, Boxiang Qin, Yuedan Chen, Wangdong Yang
IJCNN3
2025 G-SPAC: a more granular greedy graph partition algorithm with spatial locality and judgment-aware edge folding
Yuedan Chen, Guoqing Xiao 0001, Peixin Xu, Kenli Li 0001
Sci. China Inf. Sci.1
2025 DCGG: A Dynamically Adaptive and Hardware-Software Coordinated Runtime System for GNN Acceleration on GPUs
abstract
Graph neural networks (GNNs) are a prominent trend in graph-based deep learning, known for their capacity to produce high-quality node embeddings. However, the existing GNN framework design is only implemented from the algorithm level, and the hardware architecture of the GPU is not fully utilized.To this end, we propose DCGG, a dynamic runtime adaptive framework, which can accelerate various GNN workloads on GPU platforms. DCGG has carried out deeper optimization work mainly in terms of load balancing and software and hardware matching. Accordingly, three optimization strategies are proposed. First, we propose dynamic 2D workload management methods and perform customized optimization based on it, effectively reducing additional memory operations. Second, a new slicing strategy is adopted, combined with hardware features, to effectively improve the efficiency of data reuse. Third, DCGG uses the Quantitative Dimension Parallel Strategy to optimize dimensions and parallel methods, greatly improving load balance and data locality. Extensive experiments demonstrate that DCGG outperforms the state-of-the-art GNN computing frameworks, such as Deep Graph Library (up to 3.10× faster) and GNNAdvisor (up to 2.80× faster), on mainstream GNN architectures across various datasets.
Guoqing Xiao 0001, Yuedan Chen, Hongyang Chen 0001, Wangdong Yang
IEEE Trans. Computers3
2024 PUSHGNN: A Low-communication Runtime System for GNN Acceleration on Multi-GPUs
abstract
The need for multi-GPU platforms in graph neural networks (GNNs) has been driven by the growing size of input graphs. However, although the existing multi-gpu GNN framework has been optimized from the perspective of optimizing computing and communication operations, communication competition still exists. To this end, we introduce PUSHGNN, a runtime system designed to reduce communication overhead across GPUs, boosting GNN performance. Therefore, we designed a push-based pipeline communication model and made custom tuning to significantly reduce pipeline contention. Comparative assessments demonstrate that PUSHGNN consistently outperforms leading full-graph GNN systems on average 1.97× and 7.57× faster than MGG and MGG-UVM, respectively.
Guoqing Xiao 0001, Yuedan Chen, Wangdong Yang
IEEE Big Data3
2024 Accurate and Scalable Graph Convolutional Networks for Recommendation Based on Subgraph Propagation
abstract
In recommendation systems, Graph Convolutional Networks (GCNs) often suffer from significant computational and memory cost when propagating features across the entire user-item graph. While various sampling strategies have been introduced to reduce the cost, the challenge of neighbor explosion persists, primarily due to the iterative nature of neighbor aggregation. This work focuses on exploring subgraph propagation for scalable recommendation by addressing two primary challenges:efficient and effective subgraph constructionandsubgraph sparsity. To address these challenges, we propose a novelGCNmodel for recommendation based onSubgraph propagation, called SubGCN. One key component of SubGCN is BiPPR, a technique that fuses both source- and target-based Personalized PageRank (PPR) approximations, to overcome the challenge ofefficient and effective subgraph construction. Furthermore, we propose a source-target contrastive learning scheme to mitigate the impact ofsubgraph sparsityfor SubGCN. We conduct extensive experiments on two large and two medium-sized datasets to evaluate the scalability, efficiency, and effectiveness of SubGCN. On medium-sized datasets, compared to full-graph GCNs, SubGCN achieves competitive accuracy while using only 23.79% training time on Gowalla and 16.3% on Yelp2018. On large datasets, where full-graph GCNs ran out of the GPU memory, our proposed SubGCN outperforms widely used sampling strategies in terms of training efficiency and recommendation accuracy.
Xueqi Li 0002, Guoqing Xiao 0001, Yuedan Chen, Kenli Li 0001, Gao Cong
IEEE Trans. Knowl. Data Eng.3
2024 Efficient Utilization of Multi-Threading Parallelism on Heterogeneous Systems for Sparse Tensor Contraction
abstract
Many fields of scientific simulation, such as chemistry and condensed matter physics, are increasingly eschewing dense tensor contraction in favor of sparse tensor contraction. In this work, we center around binary sparse tensor contraction (SpTC) which has the challenges of index matching and accumulation. To address these difficulties, we present GSpTC, an efficient element-wise SpTC framework on CPU-GPU heterogeneous systems. GSpTC first introduces a fine-grained partitioning strategy based on element-wise tensor contraction. By analyzing and selecting appropriate dimension partitioning strategies, we can efficiently utilize the multi-threading parallelism on GPUs and optimize the overall performance of GSpTC. In particular, GSpTC leverages multi-threading parallelism on GPUs for the contraction phase and merging phase, which greatly accelerates the computation phase in sparse tensor contraction computations. Furthermore, GSpTC employs parallel pipeline technology to hide the data transmission time between the host and the device, further enhancing its performance. As a result, GSpTC achieves an average performance improvement of 267% compared to the previous state-of-the-art framework Sparta.
Guoqing Xiao 0001, Chuanghui Yin, Yuedan Chen, Mingxing Duan, Kenli Li 0001
IEEE Trans. Parallel Distributed Syst.3
2024 APPQ-CNN: An Adaptive CNNs Inference Accelerator for Synergistically Exploiting Pruning and Quantization Based on FPGA
abstract
Convolutional neural networks (CNNs) are widely utilized in intelligent edge computing applications such as computational vision and image processing. However, as the number of layers of the CNN model increases, the number of parameters and computations gets larger, making it increasingly challenging to accelerate in edge computing applications. To effectively adapt to the tradeoff between the speed and accuracy of CNNs inference for smart applications. This paper proposes an FPGA-based adaptive CNNs inference accelerator synergistically utilizing filter pruning, fixed-point parameter quantization, and multi-computing unit parallelism called APPQ-CNN. First, the article devises a hybrid pruning algorithm based on the L1- norm and APoZ to measure the filter impact degree and a configurable parameter quantization fixed-point computing architecture instead of floating-point architecture. Then, design a cascade of the CNN pipelined kernel architecture and configurable multiple computation units. Finally, conduct extensive performance exploration and comparison experiments on various real and synthetic datasets. With negligible accuracy loss, the speed performance of our accelerator APPQ-CNN compares with current state-of-the-art FPGA-based accelerators PipeCNN and OctCNN by 2.15x and 1.91x, respectively. Furthermore, APPQCNN provides settable fixed-point quantization bit-width parameters, filter pruning rate, and multiple computation unit counts to cope with practical application performance requirements in edge computing.
Guoqing Xiao 0001, Mingxing Duan, Yuedan Chen, Kenli Li 0001
IEEE Trans. Sustain. Comput.4
2023 PH-CF: A Phased Hybrid Algorithm for Accelerating Subgraph Matching Based on CPU-FPGA Heterogeneous Platform
abstract
Nowadays, more data are represented and stored by a graph structure, and subgraph matching is a fundamental problem in a variety of scientific machine learning and industrial applications, such as remote sensing image registration, industrial inspection, etc. Due to the nondeterministic polynominal hard (NP-hard) problem of subgraph matching, the explosive growth of graph data, the disadvantages of high energy consumption, and the high overhead of CPU and graphics processing unit (GPU) platforms, computing subgraph matching is becoming more and more challenging. To alleviate this problem, we propose a phased hybrid algorithm to accelerate the enumeration task of subgraph matching, calledPH-CF, based on the CPU-field programmable gate arrays (FPGA) heterogeneous platform. This approach can make full use of the pipeline and data flow mechanism, low power consumption, and configurable characteristics of FPGA. First, the matching order of query vertices automatically selects GraphQL (GQL) or RI methods according to the sparsity of the data (query) graph. Second, a candidate vertex auxiliary data structure set partitioning method is designed to effectively realize the load balance of multiple computing units at the FPGA and CPU host sides. Third, FPGA's pipeline and data flow mechanism is used to accelerate the enumeration phase of subgraph matching. Experimental results on real-world and synthetic datasets show that the performance of thePH-CFoutperforms the state-of-the-arts.PH-CFcan obtain the average performance improvement of up to$16.07\times$,$38.61\times$, and$11.46\times$over CFL, CECI, and DP-iso, respectively. Moreover, our approach has good stability and robustness on various datasets.
Guoqing Xiao 0001, Mingxing Duan, Yuedan Chen, Kenli Li 0001
IEEE Trans. Ind. Informatics4
2023 An Explicitly Weighted GCN Aggregator based on Temporal and Popularity Features for Recommendation
abstract
Graph convolutional network (GCN) has been extensively applied to recommender systems (RS) and achieved significant performance improvements through iteratively aggregating high-order neighbors to model the relevance between users and items as well as their characteristics. In the aggregation process, GCN models usually give neighbors the same or trainable weights based on implicit features, ignoring explicit ones. In this work, we take the features with explicit meanings or extracted with specific purpose as explicit ones (e.g., temporal features) and the others contained in user-item network as implicit ones (e.g., user preferences). However, some explicit features and knowledge play an essential role in improving the model representation ability and explainability in recommendation systems. To deal with the limitation, we propose a GCN based framework to embed the explicit features or those extracted with explicit intentions in this work. We also provide specific implementations based on two commonly researched features, temporal evolution and popularity bias. Specifically, we first experimentally analyze the popularity bias of the representation learning in RS based on two commonly used GCN models. Secondly, we propose a general framework to weigh neighbors based on explicit features or intentions. Thirdly, we implement a Temporal and Popularity weighted Aggregator (TPA) for GCN. The Interest-Forgetting Curve is utilized to capture temporal evolution as temporal weights and the data-driven Beta distribution is employed to tune the weights based on the node popularity flexibly. At last, we conduct extensive experiments on three real-world datasets to demonstrate the effectiveness of TPA in improving recommendation accuracy and alleviating the popularity bias.
Xueqi Li 0002, Guoqing Xiao 0001, Yuedan Chen, Zhuo Tang, Kenli Li 0001
Trans. Recomm. Syst.3
2022 Exploiting Hierarchical Parallelism and Reusability in Tensor Kernel Processing on Heterogeneous HPC Systems
abstract
Canonical Polyadic Decomposition (CPD) of sparse tensors is an effective tool in various machine learning and data analytics applications, in which sparse Matricized Tensor Times Khatri-Rao Product (MTTKRP) is the major performance bottleneck. To overcome this bottleneck and support efficient applications, this paper presents HPSpTM, an efficient sparse MTTKRP framework, to exploit the multi-level parallelism and reusability on heterogeneous HPC systems. HPSpTM incorporates: (1) a multi-level matrix-driven tiling engine that leverages the process- and thread-level parallelism of the underlying platform and data reusability based on the derived factor matrix-driven MTTKRP algorithm; (2) a tensor-driven parallel execution that enables buffering-aware scheduling and pipeline scheduling to optimize the performance in the tile granularity; (3) a partition-aware light weight data storage that exploits better data locality based on the proposed hierarchical and fine-grained execution; and (4) a performance auto-tuning technique that offers large flexibility for tile size auto-adjusting across various input datasets based on a designed runtime model. Our experiments show that HPSpTM on a Nvidia Tesla P100 obtains the average performance improvement of up to 76.46% over the state-of-the-arts, and HPSpTM achieves the speedup of up to 15.39× when scaling from 8 to 128 core groups, corresponding to processes, on the Sunway TaihuLight supercomputer.
Yuedan Chen, Guoqing Xiao 0001, M. Tamer Özsu, Zhuo Tang, Albert Y. Zomaya, Kenli Li 0001
ICDE1
2021 CASpMV: A Customized and Accelerative SpMV Framework for the Sunway TaihuLight
abstract
The Sunway TaihuLight, equipped with 10 million cores, is currently the world's third fastest supercomputer. SpMV is one of core algorithms in many high-performance computing applications. This paper implements a fine-grained design for generic parallel SpMV based on the special Sunway architecture and finds three main performance limitations, i.e., storage limitation, load imbalance, and huge overhead of irregular memory accesses. To address these problems, this paper introduces a customized and accelerative framework for SpMV (CASpMV) on the Sunway. The CASpMV customizes an auto-tuning four-way partition scheme for SpMV based on the proposed statistical model, which describes the sparse matrix structure characteristics, to make it better fit in with the computing architecture and memory hierarchy of the Sunway. Moreover, the CASpMV provides an accelerative method and customized optimizations to avoid irregular memory accesses and further improve its performance on the Sunway. Our CASpMV achieves a performance improvement that ranges from 588.05 to 2118.62 percent over the generic parallel SpMV on a CG (which corresponds to an MPI process) of the Sunway on average and has good scalability on multiple CGs. The performance comparisons of the CASpMV with state-of-the-art methods on the Sunway indicate that the sparsity and irregularity of data structures have less impact on CASpMV.
Guoqing Xiao 0001, Kenli Li 0001, Yuedan Chen, Wangquan He, Albert Y. Zomaya, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.3
2020 ahSpMV: An Autotuning Hybrid Computing Scheme for SpMV on the Sunway Architecture
abstract
The prevalence of the Internet of Things (IoT) and the explosion of available information on the Web have led to an enormous amount of widely available IoT data sets with sparsity. Sparse matrix-vector multiplication (SpMV) is one of the most essential algorithms in various kinds of IoT applications. This article designs an autotuning hybrid computing scheme for SpMV, named ahSpMV, on the powerful and unique architecture of Sunway TaihuLight supercomputer, to combine the advantages of the heterogeneous parallel Sunway architecture and the Hybrid (HYB) sparse matrix format and optimize the SpMV's performance. First, we propose a heterogeneous parallelization design for ahSpMV based on the heterogeneous manycore architecture of the SW26010 of Sunway TaihuLight and the hybrid feature of the HYB format. Second, we propose several optimization techniques for computation and communication of ahSpMV, to fully utilize the computing power of Sunway. Third, we analyze the execution time of ahSpMV on Sunway. Fourth, based on the performance analysis, we propose an autotuning scheme for ahSpMV to set the proper parameter for the HYB format. We evaluate ahSpMV's performance on the Sunway architecture. The result analysis indicates that ahSpMV has obvious performance improvement over parallel SpMV based on other related sparse matrix formats. The optimization techniques and the autotuning scheme for ahSpMV also yield expected optimization effects. Moreover, the experimental results illustrate that ahSpMV has good scalability on the Sunway architecture.
Guoqing Xiao 0001, Yuedan Chen, Chubo Liu, Xu Zhou 0001
IEEE Internet Things J.2
2020 tpSpMV: A two-phase large-scale sparse matrix-vector multiplication kernel for manycore architectures
Yuedan Chen, Guoqing Xiao 0001, Fan Wu 0016, Zhuo Tang, Keqin Li 0001
Inf. Sci.1
2020 MalFCS: An effective malware classification framework with automated feature extraction based on deep convolutional neural networks
Guoqing Xiao 0001, Jingning Li, Yuedan Chen, Kenli Li 0001
J. Parallel Distributed Comput.3
2020 Optimizing partitioned CSR-based SpGEMM on the Sunway TaihuLight
Yuedan Chen, Guoqing Xiao 0001, Wangdong Yang
Neural Comput. Appl.1
2020 aeSpTV: An Adaptive and Efficient Framework for Sparse Tensor-Vector Product Kernel on a High-Performance Computing Platform
abstract
Multi-dimensional, large-scale, and sparse data, which can be neatly represented by sparse tensors, are increasingly used in various applications such as data analysis and machine learning. A high-performance sparse tensor-vector product (SpTV), one of the most fundamental operations of processing sparse tensors, is necessary for improving efficiency of related applications. In this article, we propose aeSpTV, an adaptive and efficient SpTV framework on Sunway TaihuLight supercomputer, to solve several challenges of optimizing SpTVon high-performance computing platforms. First, to map SpTV to Sunway architecture and tame expensive memory access latency and parallel writing conflict due to the intrinsic irregularity of SpTV, we introduce an adaptive SpTV parallelization. Second, to co-execute with the parallelization design while still ensuring high efficiency, we design a sparse tensor data structure named CSSoCR. Third, based on the adaptive SpTV parallelization with the novel tensor data structure, we present an autotuner that chooses the most befitting tensor partitioning method for aeSpTV using the variance analysis theory of mathematical statistics to achieve load balance. Fourth, to further leverage the computing power of Sunway, we propose customized optimizations for aeSpTV. Experimental results show that aeSpTV yields good sacalability on both thread-level and process-level parallelism of Sunway. It achieves a maximum GFLOPS of 195.69 on 128 processes. Additionally, it is proved that optimization effects of the partitioning autotuner and optimization techniques are remarkable.
Yuedan Chen, Guoqing Xiao 0001, M. Tamer Özsu, Chubo Liu, Albert Y. Zomaya, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2019 Implementation and optimization of a data protecting model on the Sunway TaihuLight supercomputer with heterogeneous many-core processors
abstract
Summary With the rapid development of information technology, the security of massive amounts of digital data has attracted huge attention in recent years. The Advanced Encryption Standard (AES) algorithm and the Security Hash Algorithm 3 (SHA3) are extensively used as cryptographic algorithms for protecting the security of information. The Sunway TaihuLight, with massive heterogeneous many‐core SW26010 processors, has the peak performance of over 100 PFlops. To achieve high efficiency of data encryption/decryption and guarantee the data integrity for large‐scale applications, this paper proposes a fast and secure data protecting model using the parallel AES algorithm and the SHA3 on the Sunway TaihuLight. According to the particular computing architecture and memory hierarchy of the Sunway TaihuLight, we propose a fine‐grained software design for the data protecting model to fully exploit the parallelism and properly arranges the data on the Sunway. Furthermore, we propose optimization strategies for the parallel AES algorithm of the data protecting model. It is proved that our data protecting model has high security, good scalability, and excellent encryption/decryption efficiency on the Sunway. Our data protecting model achieves a high throughput of 269.95 Gbits/s, and the optimized parallel AES algorithm achieves 511.28 Gbits/s on the Sunway.
Yuedan Chen, Kenli Li 0001, Xiongwei Fei, Zhe Quan, Keqin Li 0001
Concurr. Comput. Pract. Exp.1
2019 Performance-Aware Model for Sparse Matrix-Matrix Multiplication on the Sunway TaihuLight Supercomputer
abstract
General sparse matrix-sparse matrix multiplication (SpGEMM) is one of the fundamental linear operations in a wide variety of scientific applications. To implement efficient SpGEMM for many large-scale applications, this paper proposes scalable and optimized SpGEMM kernels based on COO, CSR, ELL, and CSC formats on the Sunway TaihuLight supercomputer. First, a multi-level parallelism design for SpGEMM is proposed to exploit the parallelism of over 10 millions cores and better control memory based on the special Sunway architecture. Optimization strategies, such as load balance, coalesced DMA transmission, data reuse, vectorized computation, and parallel pipeline processing, are applied to further optimize performance of SpGEMM kernels. Second, we thoroughly analyze the performance of the proposed kernels. Third, a performance-aware model for SpGEMM is proposed to select the most appropriate compressed storage formats for the sparse matrices that can achieve the optimal performance of SpGEMM on the Sunway. The experimental results show the SpGEMM kernels have good scalability and meet the challenge of the high-speed computing of large-scale data sets on the Sunway. In addition, the performance-aware model for SpGEMM achieves an absolute value of relative error rate of 8.31 percent on average when the kernels are executed in one single process and achieves 8.59 percent on average when the kernels are executed in multiple processes. It is proved that the proposed performance-aware model can perform at high accuracy and satisfies the precision of selecting the best formats for SpGEMM on the Sunway TaihuLight supercomputer.
Yuedan Chen, Kenli Li 0001, Wangdong Yang, Guoqing Xiao 0001, Xianghui Xie 0001, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2016 Implementation and Optimization of AES Algorithm on the Sunway TaihuLight
abstract
With the rapid development of information technology, the security of massive amounts of digital data has attracted huge attention in recent years. In this paper, we provide an efficient parallel implementation of the Advanced Encryption Standard (AES) algorithm, a widely used symmetrical block encryption algorithm, based on the Sunway TaihuLight. The Sunway TaihuLight is a China's independently developed heterogeneous supercomputer with peak performance over 100 PFlops. We also optimize the parallel implementation of the AES algorithm based on the Sunway TaihuLight to achieve more optimized performance. The optimization of the parallel AES algorithm in a single SW26010 node is provided. Specifically, we expand the scale to 1024 nodes and achieve the throughput of about 63.91 GB/s (511.28 Gbits/s). Our parallel implementation of the AES algorithm has great parallel scalability and the speedup ratio can be very high with the number of nodes increasing.
Yuedan Chen, Kenli Li 0001, Xiongwei Fei, Zhe Quan, Keqin Li 0001
PDCAT1