Wu Yuan 0002

dblp:230/9614 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6528-9181ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 WindStencil: Unleashing GPU Potential for High-Order Stencil Computation in High-Performance Inviscid CFD Simulations
Xiazhen Liu, Runfeng Jin, Jian Zhang 0070, Wu Yuan 0002, Shan Liang 0005, Zhonghua Lu
ICS7
2026 From Optimal Solutions to Reusable Rules: A Knowledge-Driven Stencil Scheduling Framework
Jian Zhang 0070, Xiazhen Liu, Wu Yuan 0002, Shan Liang 0005
KSEM (6)5
2025 An Integrated Topology-Aware Mapping Scheme for Large-Scale CFD Applications on the ORISE Supercomputer
Lin Gao 0004, Wu Yuan 0002, Xiazhen Liu, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICA3PP (6)2
2025 FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
abstract
Efficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %.
Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICCD8
2024 MIST: Efficient Mixed-Precision Preconditioning Through Iterative Sparse- Triangular Solver Design
abstract
Exact sparse-triangular solvers are highly sequential and difficult to implement efficiently on GPUs with ILU preconditioning. Lower precision is crucial for reducing data movement and storage demands in memory-bound problems. However, current mixed-precision systems struggle to achieve performance gains for ILU preconditioning on multi-GPU platforms due to challenges in (1) utilizing two levels of parallelism and (2) minimizing off-chip memory bandwidth while maintaining accuracy. Additionally, these systems focus on scalar operations and lack support for point-block matrices, which arise naturally in multiphysics problems and require tailored algorithm designs. To address these challenges, we propose MIST, a novel Mixed-precision Iterative Sparse-Triangular solver optimized for GPUs to accelerate preconditioning in Krylov methods. We (1) implement an efficient mixed-precision Jacobi iterative local solver to harness single-GPU parallelism and scale it to multi- GPU via domain decomposition, and (2) design a BSpMVA kernel to reduce bandwidth while achieving high double-precision accuracy. Integrated into a widely-used numerical library, MIST offers end-to-end support for solving sparse linear systems, balancing efficiency and convergence. Experimental results show that MIST provides a 3.38× average speedup over cuSPARSE�s exact sparse-triangular solver, with an additional 1.37× speedup when using low-precision, while maintaining robustness
Yidong Chen 0014, Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu
ICCD4
2024 Importance-Guided Sequential Training for Physics-Informed Neural Networks
abstract
In recent years, the emergence of deep learning has brought Physics-Informed Neural Networks (PINNs) into the spotlight as a promising method for solving partial differential equations (PDEs). Despite the attention received by PINNs, their accuracy and convergence face significant challenges, particularly when dealing with complex equations containing multiple components. In PDEs, due to their inherent rigidity, distinct components may converge at varying rates, resulting in an unbalanced convergence process. This imbalance during optimization has the potential to yield suboptimal solutions. We introduce an importance-guided sequential training method to regulate the competition of different components in PDEs. The importance can be quantitatively defined through correlation analysis, enabling the formulation of an algorithm that systematically instructs neural networks to acquire insights from individual terms in accordance with their respective importance levels. To evaluate the performance of our proposed approach, we applied it to solve the Burgers equation and the Klein-Gordon equation. The results clearly demonstrate that our proposed approaches outperform the original model, showcasing the effectiveness of our internal weighting method in improving the accuracy and convergence of PINNs. Furthermore, our research serves as a stepping stone for exploring and harnessing the power of correlation analysis in the realm of PINNs.
Nanxi Chen, Jiyan Qiu, Xuesong Wu 0005, Wu Yuan 0002
IJCNN6
2024 Mixed-precision block incomplete sparse approximate preconditioner on Tensor core
Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu
CCF Trans. High Perform. Comput.3
2023 udPINNs: An Enhanced PDE Solving Algorithm Incorporating Domain of Dependence Knowledge
Nanxi Chen, Jiyan Qiu, Wu Yuan 0002, Jian Zhang 0070
KSEM (4)4
2022 Sparse Reconstruction Method for Flow Fields Based on Mode Decomposition Autoencoder
Jiyan Qiu, Wu Yuan 0002, Jian Zhang 0070, Xuebin Chi
PRICAI (1)2