Kangkang Chen

dblp:50/8685 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 AMXStencil: Boosting the Performance of Stencil Computations on AMX-Powered CPUs via Fusion
Zitong An, Kangkang Chen, Huayou Su, Jinwei Xu, Xi Yang 0020
APPT2
2026 Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUs
abstract
Graph Neural Networks (GNNs) are becoming increasingly popular in graph data processing due to their excellent performance in feature extraction on graph datasets. Compared to GPUs, CPUs are more widely accessible and serve as a practical platform for GNN inference. However, achieving efficient GNN execution on CPUs remains a challenge. We first comprehensively evaluate and quantitatively analyze the performance of GNN inference on multi-core CPUs using the state-of-the-art frameworks, identifying four key performance bottlenecks: inefficient sparse computation, poor data locality, workload imbalance, and inefficient General Matrix Multiplication (GEMM). To tackle these issues, we introduce a set of joint optimizations. Specifically, for the aggregation phase, we propose three optimizations: a register padding and tiling Graph Sparse-dense Matrix Multiplication (GSpMM) algorithm that leverages the computation capability of long vector processing units on modern multi-core CPUs, a destination node-oriented indexes reorganization to enhance data locality, and a boundary buffer-based method to balance the workloads. Additionally, for the update phase, we develop an efficient bias fusion GEMM algorithm, tailored for the irregular matrices. We evaluate the proposed optimizations extensively with three popular GNN models on three typical multi-core CPU platforms. Experimental results on Intel, AMD, and ARM platforms show that our optimizations outperform the state-of-the-art GNN framework DGL by an average factor of 2.41×, 1.58×, and 2.04× (up to 4.75×, 2.70×, and 3.55×), respectively. Compared to PyG, our implementations achieve an average speedup of 1.70×, 1.86×, and 2.44×, respectively.
Kangkang Chen, Huayou Su, Xi Yang 0020, Zitong An, Yong Dou, Dongsheng Li 0001
ACM Trans. Archit. Code Optim.1
2023 Optimizing GNN Inference Processing on Very Long Vector Processor
Kangkang Chen, Huayou Su, Chaorun Liu
ICA3PP (6)1
2022 An Efficient Transformer Inference Engine on DSP
Kangkang Chen, Huayou Su, Chaorun Liu, Xiaoli Gong
ICA3PP1
2022 Deep Reinforcement Learning with Parametric Episodic Memory
abstract
Deep Reinforcement Learning methods are widely acknowledged to be sample inefficient, while incorporating episodic memory significantly improves it through rapidly latching onto successful experiences to guide the action of agents. Previous episodic methods, utilizing discrete memory, cannot well accommodate the continuous control tasks and have limited generalization ability to aggregate the experience across trajectories. We propose an improved episodic memory-based RL algorithm, combining the one-step method in off-policy algorithm with Parametric Episodic Memory (PEM), which leverages the discrete memory by neural networks, and thereby enhances both sample efficiency and generalization ability. Moreover, an adaptive k-nearest-neighbors is used in determining the volume of retrieved memory, further improving its efficiency. Our algorithm, evaluated on various MuJoCo continuous control tasks, outperforms the model-free baseline methods and latest episodic memory-based RL algorithms.
Kangkang Chen, Zhongxue Gan 0001, Siyang Leng, Chun Guan
IJCNN1
2010 An improved section-wise exploiting modification direction method
Yiting Sun, Huan Xu 0002, Kangkang Chen, Hyoung Joong Kim, Sang-Hyun Joo
Signal Process.4