Guangnan Feng

dblp:322/6692 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-1382-280XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations
abstract
Particle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model.
Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen 0001, Yutong Lu
EuroSys5
2026 POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
abstract
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a \(99.1\%\) overlap ratio. In cross-platform comparisons, POLAR-PIC achieves \(13.2\%\) of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches \(9.6\%\) on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains \(67.5\%\) weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems.
Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Languang Gao, Zhiguang Chen 0001, Yutong Lu
HPDC5
2026 TADS: Trend-Aware Dynamic Load Balancing for Large-Scale SNN Simulations with Delay-Sharded Graph Infrastructure
abstract
Large-scale simulation of Spiking Neural Networks (SNNs) on supercomputers is pivotal for unraveling the mechanisms of brain function and advancing brain-inspired intelligence. However, efficiently mapping billions of neurons onto distributed nodes presents a significant challenge due to the heterogeneity of neuronal activities and complex, irregular network connectivity. While static partitioning strategies perform well in stable states, they often fail under metastable neurodynamics where theoretical models cannot accurately predict neuronal firing rates, leading to severe load imbalance. To address this, we propose TADS (Trend-Aware Dynamic load balancing with Delay-Sharded graph infrastructure), a framework tailored for large-scale SNN simulations. First, we identify the specific failure modes of static partitioning under metastable dynamics, establishing the necessity for runtime intervention. Second, we introduce a trend-aware dynamic load balancing strategy. By analyzing the temporal evolution of loads, this approach effectively distinguishes persistent imbalance from transient fluctuations, thereby avoiding unnecessary migrations caused by momentary jitter. Third, we design a delay-sharded graph infrastructure that leverages synaptic delays to parallelize graph modifications, significantly reducing the overhead associated with dynamic load balancing. Experimental results on the Tianhe-Xingyi supercomputer demonstrate that TADS effectively handles metastable scenarios, achieving up to 2.13 × speedup over state-of-the-art static partitioning methods, while sustaining 1.58 × performance improvement at the largest evaluated scale of 192 nodes.
Shangzhi Pang, Yangle Zeng, Guangnan Feng, Zhiguang Chen 0001, Yutong Lu
ICS4
2025 HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLA
abstract
Stencil computations are fundamental to various HPC and intelligent computing applications, often consuming significant execution time. The emergence of specialized matrix units presents new opportunities to accelerate stencil computations. While scalable matrix compute units provide substantial computing horsepower, prior efforts fail to fully utilize the computing capabilities for stencils due to suboptimal matrix-unit utilization, limited instruction-level parallelism, and low cache hit rates. This paper introduces HStencil, a novel stencil computing framework utilizing matrix and vector units. HStencil addresses these challenges through three contributions: 1) microkernels that jointly leverage matrix and vector units to enhance hardware utilization; 2) fine-grained instruction scheduling with interleaved execution to enhance instruction-level parallelism; and 3) spatial prefetch to sustain high performance when working sets exceed cache capacity. Evaluations on representative benchmarks demonstrate that HStencil achieves maximum speedups of 1.81x – 5.76x over auto-vectorization across different CPU platforms, delivers 31% - 91% higher performance versus state-of-the-art methods.
Jiabin Xie, Guangnan Feng, Xianwei Zhang 0001, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu
SC3
2025 Critique of "Productivity, Portability, Performance Data-Centric Python" by SCC Team From Sun Yat-sen University
abstract
In SC21, Ziogas et al. proposed Data-Centric (DaCe) Python. It attains high performance and portability, and further extends the original productivity of Python. This paper analyzes the reproducibility of the DaCe paper as part of the SC22 Student Cluster Competition (SCC). The reproduction experiments are conducted on the Azure CycleCloud. Different from the DaCe paper, we use AMD EPYC 7V73X processors for CPU-based experiments. We successfully reproduce most of the results of the DaCe paper. The remaining results are also explainable.
Tengyang Zheng, Tianxing Yang, Siran Liu, Shengyou Lu, Guangnan Feng, Zhiguang Chen 0001, Dan Huang 0001
IEEE Trans. Parallel Distributed Syst.8
2024 ATM: Area-based Partition and Topology-aware Mapping for Large-scale SNN Simulation
abstract
Spiking Neural Network (SNN) is an effective tool for the simulation of neuronal dynamics as well as the understanding of brain structure and functions. However, scaling up SNN for large-scale simulations poses significant computational demands that necessitate the supercomputers. The advent of distributed simulation introduces the requirement of SNN partition and process mapping, which becomes a critical challenge in the context of large-scale distributed SNN simulations. In this paper, we propose an Area-based partition and Topology-aware process Mapping (ATM) strategy to balance the computation workload while coping with the heterogeneity of communication interconnect. We first model the computation workload and communication volume of the SNN simulation according to its biological features. Based on this model, we design an area-based SNN partition strategy to balance the computation workload. Subsequently, we introduce a topology-aware strategy for process mapping, Bottleneck Fulfilling (BF), tailored specifically for collective communication paradigms. Experiments are conducted on an HPC cluster with a multi-area model of the marmoset brain. The results demonstrate that the proposed approach achieves up to 2.2x speedup compared with the baseline on 290 compute nodes.
Yangle Zeng, Guangnan Feng, Zhiguang Chen 0001, Yutong Lu, Nong Xiao 0001
ISPA2
2024 Extreme-scale Direct Numerical Simulation of Incompressible Turbulence on the Heterogeneous Many-core System
abstract
Direct numerical simulation (DNS) is a technique that directly solves the fluid Navier-Stokes equations with high spatial and temporal resolutions, which has driven much research regarding the nature of turbulence. For high-Reynolds number (Re) incompressible turbulence of particular interest, where the nondimensional Re characterizes the flow regime, the application of DNS is hindered by the fact that the numerical grid size (i.e., the memory requirement) scales with Re3, while the overall computational cost scales with Re4. Recent studies have shown that developing efficient parallel methods for heterogeneous many-core systems is promising to solve this computational challenge.
Jiabin Xie, Guangnan Feng, Junxuan Feng, Zhiguang Chen 0001, Yutong Lu
PPoPP2
2024 UNR: Unified Notifiable RMA Library for HPC
abstract
Remote Memory Access (RMA) enables direct access to remote memory to achieve high performance for HPC applications. However, most modern parallel programming models lack schemes for the remote process to detect the completion of RMA operations. Many previous works have proposed programming models and extensions to notify the communication peer, but they did not solve the multi-NIC aggregation, portability, hardware-software co-design, and usability problems. In this work, we proposed a Unified Notifiable RMA (UNR) library for HPC to address these challenges. In addition, we demonstrate the best practice of utilizing UNR within a real-world scientific application, PowerLLEL. We deployed UNR across four HPC systems, each with a different interconnect. The results show that PowerLLEL powered by UNR achieves up to a 36% acceleration on 1728 nodes of the Tianhe-Xingyi supercomputing system.
Guangnan Feng, Jiabin Xie, Dezun Dong, Yutong Lu
SC1
2023 GRAP: Group-level Resource Allocation Policy for Reconfigurable Dragonfly Network in HPC
abstract
Dragonfly is a highly scalable, low-diameter, and cost-efficient network topology, which has been adopted in new exascale High Performance Computing (HPC) systems. However, Dragonfly topology suffers from the limited direct links between groups. The reconfigurable network can solve this problem by reconfiguring topology to adjust the number of direct links between groups. While the performance improvement of a single job on reconfigurable HPC network has been evaluated in previous works, the performance of HPC workloads has not been studied because of the lack of an appropriate resource allocation policy.
Guangnan Feng, Dezun Dong, Shizhen Zhao, Yutong Lu
ICS1
2022 Optimized MPI collective algorithms for dragonfly topology
abstract
The Message Passing Interface (MPI) is the most prominent and dominant programming model for scientific computing in super-computing systems today. Although many general and efficient algorithms have been proposed for MPI collective operations, there is still room for topology-aware optimization. Dragonfly is a high-scalability, low-diameter, and cost-efficient network topology adopted in more and more supercomputing networks. However, Dragonfly topology limits the performance of some MPI collective operations. In this paper, our analysis shows that the bottlenecks of collective algorithms in Dragonfly topology are intra-job interference, inter-job interference, and topology mismatch. We propose 5 different optimizations, i.e., Pseudo-random Pairwise, Tree-based Shuffle, Reversed Recursive Doubling, Reordered Bruck, and Matched Rabenseifner, for MPI collective operations including All-Gather, All-to-All, All-Reduce, and Reduce-Scatter. We evaluate each optimization through CODES network simulation framework with minimal, non-minimal, and adaptive routing. The simulation results demonstrate that the performance of All-to-All, All-Gather, All-Reduce, and Reduce-Scatter can be improved by 4.7X, 3.4X, 12.7%, and 4.1X, respectively, for 32768-node jobs with adaptive routing.
Guangnan Feng, Dezun Dong, Yutong Lu
ICS1