EDBT 2026 Demo / reviewers in the wild / expert
Yaqian Gao
dblp:157/4734
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0003-8689-3750ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUsabstractEfficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %. Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu |
ICCD | 2 |
| 2024 | Large-scale Phase-Field Simulations for Solid-Solid Phase Transformations involving Elastic EnergyabstractPhase-field models have been used extensively in studying microstructure evolution in alloys and have the superiority of comprehending, predicting, and optimizing microstructure-sensitive macroscopic material properties. The elastic strain energy is a vital factor in modeling crystal structure formation in solid-solid phase transformations. Conventionally, it is computed in the reciprocal space according to the famous Khachaturyan-Shatalov theory. In large-scale simulations, the full-space Fourier transform becomes extremely time-consuming. Yaqian Gao, Jian Zhang 0070, Huang Ye, Xuebin Chi |
ICPP | 1 |
| 2024 | High-Performance 3D convolution on the Latest Generation Sunway ProcessorabstractThe emergence of High-Performance Computing (HPC) and Artificial Intelligence (AI) has significantly expanded the applications of three-dimensional convolutional neural networks (3D CNNs). At the same time, the next-generation Sunway supercomputer has evidenced its superior computational capabilities in the HPC+AI domain. However, complex 3D convolution remains a primary performance limitation in many applications. The optimization of tensor-like operators on the Sunway processor is usually implemented via a multi-level blocking approach, adapting to its architecture. Although it can effectively mitigate the differences in memory access latency among different memory hierarchies, the performance of 3D convolutions is still frequently limited by the transfer bandwidth. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
ICPP | 3 |
| 2024 | A Feature Extraction Framework for 3D Scientific Voxel Object Using SVD and Neural NetworkabstractLatest advances in computational methods and high-performance computing have enabled large-scale scientific simulations to become feasible. However, efficiently analyzing the resulting large datasets remains challenging. Typically, 3D voxel data occupies huge storage space which causes serious storage pressure. Performing real-time feature extraction during simulations could mitigate this demand. In this paper, We propose a three-decker structure framework for locating and extracting voxel object features such as pose, class, and size. The framework first utilizes a 3D CNN to localize the object’s center and size along each axis, enabling adaptive cropping of the target from the original voxel space. Then the singular value decomposition (SVD) will applied to the extracted object to preliminarily extract and refine its posture. Finally, a neural network will refine the SVD-extracted pose and predict the category and size. Finally, a neural network is utilized to infer residual pose corrections as well as category and size information for the object after preliminary pose alignment via SVD decomposition. Our method achieves accurate feature extraction for voxel objects with complex poses. More importantly, we achieve pose estimation without an initial position. By storing extracted features instead of full data, we significantly reduce storage needs. We validate our approach using the phase field simulations dataset and ModelNet40 dataset. At last, based on the proposed framework we develop an in-situ feature extraction library for running with large-scale scientific computing programs and performing real-time feature extraction on the large amount of computation data it generates. As a concrete running example, we selected the microstructure evolution program governed by the phase-field method for a common run and achieved high-accuracy feature extraction. Zhichen Feng, Yaqian Gao, Huang Ye, Jian Zhang 0070 |
IJCNN | 2 |
| 2024 | POSTER: Enabling Extreme-Scale Phase Field Simulation with In-situ Feature ExtractionabstractIn this paper, we present an integrated framework composed of a highly efficient phase field simulator and an in-situ feature extraction library. This novel framework enables us to conduct extreme-scale micro-structure evolution simulations while the characteristic features of each individual grain are extracted on the fly. After systematic design and optimization on the new generation Sunway supercomputer, the code scales up to 39 million cores and achieves 582 PFlops in double precision and 637 POps in mixed precision. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
PPoPP | 3 |
| 2016 | Enabling efficient approximate nearest neighbor search for outsourced database in cloud computing
Jianfeng Wang 0001, Meixia Miao, Yaqian Gao, Xiaofeng Chen 0001 |
Soft Comput. | 3 |