Chen Zhuang

dblp:207/2573 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FRUGAL: Pushing GPU Applications beyond Memory Limits
abstract
GPUs power modern scientific and AI applications, but their limited memory capacity restricts scalability. Buying GPUs with larger HBM is prohibitively expensive and still bounded by market limits. Existing solutions either exploit application-specific knowledge through out-of-core techniques, which lack generality, or rely on system-level page faulting, which is transparent but inefficient. We propose FRUGAL, an application-agnostic framework and methodology that reduces GPU memory footprint while sustaining high performance. FRUGAL formulates memory management as an optimization over an application’s execution graph, encompassing prefetching, kernel execution, and offloading. Using static analysis and profiling, FRUGAL applies a two-phase scheduling and migration strategy, solving an otherwise intractable optimization efficiently. Evaluations on Tiled Cholesky Decomposition, Tiled LU Decomposition, Tiny-CUDA-NN, and QuEST show that FRUGAL significantly reduces maximum GPU memory usage by 80.21%, 80.20%, 64.75% and 60.86% with only a geometric mean of 28.31% slowdown. FRUGAL allows applications to exceed hardware-imposed limits, and maintains strong performance scalability beyond existing GPU memory constraints, without additional hardware cost.
Lingqi Zhang 0001, Jiajun Huang 0001, Chen Zhuang, Ivan R. Ivanov, Peng Chen 0035, Toshio Endo, Mohamed Wahib
CGO4
2026 SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
abstract
Distributed Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in high-performance computing and deep learning applications. The major performance bottleneck in distributed SpMM lies in substantial communication overhead, which limits both performance and scalability. In this paper, we identify two key sources of communication inefficiency in distributed SpMM: redundant data transfer due to sparsity unawareness, and suboptimal utilization of hierarchical network topology. To address these, we propose (1) a fine-grained, sparsity-aware communication strategy that reduces communication overhead by exploiting the sparsity pattern of the sparse matrix, and (2) a hierarchical communication strategy that maps the sparsity-aware strategy onto two-tier GPU network architectures, minimizing redundant data movement across slower inter-node links. We implement these optimizations in a comprehensive distributed SpMM framework, SHIRO. Extensive evaluations on real-world datasets show that SHIRO demonstrates strong scalability up to 128 GPUs, achieving geometric mean speedups of 221.5 ×, 56.0 ×, 23.4 ×, and 8.8 × in SpMM over four state-of-the-art baselines (CAGNET, SPA, BCL, and CoLa, respectively) at this scale.
Chen Zhuang, Lingqi Zhang 0001, Benjamin Brock, Du Wu, Peng Chen 0035, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib
ICS1
2025 Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
abstract
Graph Convolutional Networks (GCNs), particularly for largescale graphs, are crucial across numerous domains.However, training distributed full-batch GCNs on large-scale graphs suffers from inefficient memory access patterns and high communication overhead.To address these challenges, we introduce SuperGCN, an efficient and scalable distributed GCN
Chen Zhuang, Lingqi Zhang 0001, Du Wu, Peng Chen 0035, Jiajun Huang 0001, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib
ICS1
2025 A General and Scalable GCN Training Framework on CPU Supercomputers
abstract
Graph Convolutional Networks (GCNs) are widely used in various domains. However, training distributed full-batch GCNs on large-scale graphs poses challenges due to inefficient memory access patterns and high communication overhead. This paper presents a general and efficient GCN training framework on CPU supercomputers. It comprises a general aggregation kernel designed to optimize irregular memory access and a quantization method with label propagation to reduce communication overhead. Experimental results show that our method achieves a speedup of up to 4.1× compared with the SoTA implementations.
Chen Zhuang, Peng Chen 0035, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Lingqi Zhang 0001, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib
PPoPP1
2025 Outlier-Resistant Cooperative Positioning Method Using Robust Factor Graph Optimization
abstract
Cooperative positioning (CP) is able to improve the vehicular positioning performance by introducing the data of multiple vehicles into the position estimation. However, CP methods are vulnerable to measurement outliers in dense urban areas. The existing outlier-resistant CP methods are easy to trap in local optimum and may wrongly reject the outliers when the ratio of outliers to inliers is relatively high. To deal with this problem, a factor graph optimization (FGO) based CP method using Graduated Non-Convexity (GNC) Welsch cost is proposed in this paper. The state-of-art FGO algorithm is used to integrate multi-node and multi-epoch measurements including the Global Navigation Satellite Systems (GNSS) pseudoranges, inter-epoch baselines estimated by GNSS time-differenced carrier phase (TDCP), and inter-vehicle ranging measurements in a centralized framework. The least-square cost in traditional FGO is replaced with the GNC-based Welsch cost so as to enhance the robustness of the proposed method to any kind of outliers in our CP system. The use of GNC can reduce the risk of local optimum by gradually increasing the non-convexity of the Welsch cost. The proposed method can de-weight the outliers correctly even if a large number of outliers exist. The experimental results show the superiority of the proposed method over the existing CP methods in resisting multiple outliers.
Jianrong Wang, Chen Zhuang, Hongbo Zhao 0001, Rongke Liu
IEEE Trans. Intell. Transp. Syst.2
2024 A Symbol-Level Optimization Method for Block Precoding in Wireless Communication
abstract
Compared to block precoding (BP), interference exploitation symbol-level precoding (IESLP) can effectively reduce transmit power in multiantenna wireless communication systems. However, because most existing IESLP methods are symbol-level optimizations of zero-forcing (ZF) BP which is usually not the best BP, the performance and application range of IESLP are restricted. In this article, we investigate the symbol-level optimization of any given BP method to reduce transmit power. We first define generalized constructive interference (GCI) as a new constructive interference metric for any received BP signals, where the interference is constructive as long as it will not impair the symbol error rate (SER). Then, convex GCI regions (GCIRs) of received BP signals for both phase shift keying (PSK) and quadrature amplitude modulation (QAM) constellations are designed. With the aid of GCIR, we propose BP-based symbol-level precoding (BPSLP) methods to reduce transmit power. Moreover, due to the difficulty of obtaining perfect channel state information (CSI), we also propose robust BPSLP (RBPSLP) methods to reduce transmit power for PSK and QAM constellations, where outage probability (OP) is constrained for robustness. Numerical results show that compared to existing precoding methods, our proposed precoding methods effectively reduce transmit power. It is worth noting that due to the bounded or even one-point constructive interference region (CIR), there are usually no feasible solutions for existing robust IESLP methods under QAM constellations, but our proposed RBPSLP method for QAM constellations is feasible as long as a feasible robust BP matrix exists.
Qi Wang 0078, Chen Zhuang, Wuyang Zhou
IEEE Internet Things J.2
2024 Plane Constraints Aided Multi-Vehicle Cooperative Positioning Using Factor Graph Optimization
abstract
The development of vehicle-to-vehicle (V2V) communication facilitates the study of cooperative positioning (CP) techniques for vehicular applications. The CP methods can improve the positioning availability and accuracy by inter-vehicle ranging and data exchange between vehicles. However, the inter-vehicle ranging can be easily interrupted due to many factors such as obstacles in-between two cars. Without inter-vehicle ranging, the other cooperative data such as vehicle positions will be wasted, leading to performance degradation of range-based CP methods. To fully utilize the cooperative data and mitigate the impact of inter-vehicle ranging loss, a novel cooperative positioning method aided by plane constraints is proposed in this paper. The positioning results received from cooperative vehicles are used to construct the road plane for each vehicle. The plane parameters are then introduced into CP scheme to impose constraints on positioning solutions. The state-of-art factor graph optimization (FGO) algorithm is employed to integrate the plane constraints with raw data of Global Navigation Satellite Systems (GNSS) as well as inter-vehicle ranging measurements. The proposed CP method has the ability to resist the interruptions of inter-vehicle ranging since the plane constraints are computed by just using position-related data. A vehicle can still benefit from the position data of cooperative vehicles even if the inter-vehicle ranging is unavailable. The experimental results indicate the superiority of the proposed CP method in positioning performance over the existing methods, especially when the inter-ranging interruptions occur.
Chen Zhuang, Hongbo Zhao 0001, Jianrong Wang
IEEE Trans. Intell. Transp. Syst.1
2022 Multi-criteria Selection of Rehearsal Samples for Continual Learning
Chen Zhuang, Shaoli Huang, Gong Cheng 0003, Jifeng Ning
Pattern Recognit.1
2022 Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning
abstract
We presentFastConv, a template-based code auto-generation open-source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming layers of convolutional neural networks. ARM CPUs cover a wide range of designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. Results show 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, ARM NN, and FeatherCNN on Kunpeng 920. Furthermore, performance portability experiments with various convolution shapes show that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920. CPU performance portability evaluation on VGG–16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.
Jintao Meng 0001, Chen Zhuang, Peng Chen 0035, Mohamed Wahib, Bertil Schmidt, Xiao Wang 0004, Haidong Lan, Dou Wu, Minwen Deng, Yanjie Wei, Shengzhong Feng
IEEE Trans. Parallel Distributed Syst.2
2020 A Big Data Driven Model for Screening Electricity Customers
Xingping Liu, Zhihan Xie, Chenmin Zhang, Chenhui Zhou, Chen Zhuang
ICIC (1)5