Weicheng Xue

dblp:225/4769 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0003-0816-2453ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain II
abstract
Mix-precision computation is crucial for artificial intelligence and scientific computing applications. However, as novel chips with innovative architectures emerge, harnessing their computational capabilities presents significant challenges. While existing algorithms for the HPL-MxP LU factorization excel on homogeneous systems, they often encounter difficulties on specialized heterogeneous architectures. This deficiency arises from inadequate optimization for computation, memory access, and communication, hindering effective mixed-precision acceleration. This work introduces an algorithm-hardware co-optimization approach for LU factorization on specialized NPUs and CPUs, leveraging their unique architectures. A novel multi-iteration fusion method for general matrix multiplication is proposed, strategically designed to maximize on-chip L1 buffer utilization, effectively overcoming the notorious “memory wall”. Additionally, a multi-stage, multi-level heterogeneous pipeline for LU factorization in an accelerator-CPU cloud environment is presented, where compute-intensive matrix multiplications are offloaded to NPUs while CPUs handle the remaining tasks. The co-optimization approach fosters deep collaboration between CPUs and accelerators, thereby unlocking enhanced performance.
Weicheng Xue, Kai Yang 0051, Yongxiang Liu, Dengdong Fan, Pengxiang Xu, Yonghong Tian 0001
SC1
2024 CPU-GPU heterogeneous code acceleration of a finite volume Computational Fluid Dynamics solver
Weicheng Xue, Christopher J. Roy
Future Gener. Comput. Syst.1
2021 Multi-GPU performance optimization of a computational fluid dynamics code using OpenACC
abstract
Summary This article investigates the multi‐GPU performance of a 3D buoyancy driven cavity solver using MPI and OpenACC directives on multiple platforms. The article shows that decomposing the total problem in different dimensions affects the strong scaling performance significantly for the GPU. Without proper performance optimizations, it is shown that 1D domain decomposition scales poorly on multiple GPUs due to the noncontiguous memory access. The performance using whatever decompositions can be benefited from a series of performance optimizations in the article. Since the buoyancy driven cavity code is communication‐bounded on the clusters examined, a series of optimizations both agnostic and tailored to the platforms are designed to reduce the communication cost and improve memory throughput between hosts and devices efficiently. First, the parallel message packing/unpacking strategy developed for noncontiguous data movement between hosts and devices improves the overall performance by about a factor of 2. Second, transferring different data based on the stencil sizes for different variables further reduces the communication overhead. These two optimizations are general enough to be beneficial to stencil computations having ghost exchanges. Third, GPUDirect is used to improve the communication on clusters which have the hardware and software support for direct communication between GPUs without staging the memory of CPU. Finally, overlapping the communication and computations is shown to be not efficient on multi‐GPUs if only using MPI or MPI+OpenACC. Although we believe our implementation has revealed enough communication and computation overlap, the actual running does not utilize the overlap well due to a lack of enough asynchronous progression.
Weicheng Xue, Christopher J. Roy
Concurr. Comput. Pract. Exp.1
2021 An improved framework of GPU computing for CFD applications on structured grids using OpenACC
Weicheng Xue, Charles W. Jackson, Christoper J. Roy
J. Parallel Distributed Comput.1