Xiangdong Pei

dblp:147/0447 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0005-5454-9482ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Optimize Winograd Convolution for a Novel MIMD Many-core Architecture PEZY-SC3s
abstract
Optimizing convolution operations is critical for enhancing the performance of convolutional neural networks (CNNs). The Winograd convolution algorithm, renowned for its significant reduction in computational complexity, has been widely adopted in convolution acceleration. While the Winograd convolution algorithm has been highly optimized for traditional computing platforms such as SIMD-based CPUs and SIMT-based GPUs, research on alternative architectures better suited for convolution operations remains ongoing. MIMD architectures, characterized by their high parallelism and thread divergence mitigation capabilities, present promising potential; nevertheless, their applicability and performance in convolution operations remain underexplored. Additionally, GPU-based convolution acceleration, despite delivering exceptional performance, faces escalating energy consumption challenges, necessitating a balanced optimization of hardware performance and energy efficiency. To address these, we implement and optimize the Winograd convolution algorithm on a low-power MIMD many-core processor PEZY-SC3s, aiming to investigate its viability for convolution workloads. Experimental evaluations using convolutional layer parameters from widely adopted CNNs demonstrate: 1) 78.37% average bandwidth utilization in data-intensive stages; 2) 98% single-core and 92.52% system-wide computational efficiency in compute-intensive phases. The optimized Winograd algorithm achieves significantly superior computational efficiency compared to cuDNN-based implementation on Nvidia A100 GPU and oneDNN-based implementation on Intel Xeon Silver 4314 CPU, with energy efficiency ratios (GFLOPS/W) of $2.58 \times$ and $27.52 \times$, respectively.
Zhiyan Liu, Bingwei Wang, Feiming Liu, Xiangdong Pei
PACT7
2025 Optimizing Incomplete Cholesky Factorization on MIMD Many-core Architecture
abstract
Incomplete Cholesky (IC) factorization is widely used to precondition large sparse positive definite symmetric linear equations, since it can effectively reduce the number of iterations and enhance the solution efficiency compared to the unconditioned Conjugate Gradient (CG) iterative algorithm. In this paper, we have carried out optimization research on the IC factorization parallel algorithms based on the MIMD many-core architecture PEZY-SC3s processor. We propose a parallel algorithm for IC factorization under the MIMD architecture. This algorithm is based on the level-scheduling and graph coloring reordering algorithm. We optimize algorithm performance using various measures, including memory access optimization based on vector units and on-chip local memory, load balancing optimization based on dynamic thread scheduling, and task allocation optimization based on graph partitioning. Experimental results demonstrate that our IC parallel factorization achieves average speedups of 17.5x and 8.4x over the two cuSPARSE implementations on an NVIDIA A30 GPU, and 230.6x and 83.0x over Ginkgo’s serial and OpenMP implementations on a dual-socket Intel Xeon 4314 platform. Our study demonstrates that the PEZY architecture with effective algorithmic design still achieves good performance gains while maintaining low power consumption.
Yongzhen Shi, Jie Liu 0002, Zhiyan Liu, Bingwei Wang, Feiming Liu, Xiangdong Pei
ICPP8
2024 High performance dilated convolutions on multi-core DSPs
Xiangdong Pei, Songzhu Mei, Rongchun Li, Jie Liu 0002
CCF Trans. High Perform. Comput.3
2023 Optimizing Pointwise Convolutions on Multi-core DSPs
Xiangdong Pei, Songzhu Mei, Jie Liu 0002
ICA3PP (7)3
2023 Predicting gene regulatory links from single-cell RNA-seq data using graph neural networks
abstract
Single-cell RNA-sequencing (scRNA-seq) has emerged as a powerful technique for studying gene expression patterns at the single-cell level. Inferring gene regulatory networks (GRNs) from scRNA-seq data provides insight into cellular phenotypes from the genomic level. However, the high sparsity, noise and dropout events inherent in scRNA-seq data present challenges for GRN inference. In recent years, the dramatic increase in data on experimentally validated transcription factors binding to DNA has made it possible to infer GRNs by supervised methods. In this study, we address the problem of GRN inference by framing it as a graph link prediction task. In this paper, we propose a novel framework called GNNLink, which leverages known GRNs to deduce the potential regulatory interdependencies between genes. First, we preprocess the raw scRNA-seq data. Then, we introduce a graph convolutional network-based interaction graph encoder to effectively refine gene features by capturing interdependencies between nodes in the network. Finally, the inference of GRN is obtained by performing matrix completion operation on node features. The features obtained from model training can be applied to downstream tasks such as measuring similarity and inferring causality between gene pairs. To evaluate the performance of GNNLink, we compare it with six existing GRN reconstruction methods using seven scRNA-seq datasets. These datasets encompass diverse ground truth networks, including functional interaction networks, Loss of Function/Gain of Function data, non-specific ChIP-seq data and cell-type-specific ChIP-seq data. Our experimental results demonstrate that GNNLink achieves comparable or superior performance across these datasets, showcasing its robustness and accuracy. Furthermore, we observe consistent performance across datasets of varying scales. For reproducibility, we provide the data and source code of GNNLink on our GitHub repository: https://github.com/sdesignates/GNNLink.
Guo Mao, Zhengbin Pang, Ke Zuo, Xiangdong Pei, Xinhai Chen 0001, Jie Liu 0002
Briefings Bioinform.5