VLDB 2026 Research / reviewers in the wild / expert
Meiyue Shao
dblp:40/8946
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-4914-7666ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021Theory of computation · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Shadow tomography of quantum states with prediction
Jiyu Jiang, Zongqi Wan, Tongyang Li, Meiyue Shao |
Frontiers Comput. Sci. | 4 |
| 2025 | A Mixed Precision Jacobi SVD AlgorithmabstractWe propose a mixed precision Jacobi algorithm for computing the singular value decomposition (SVD) of a dense matrix. After appropriate preconditioning, the proposed algorithm computes the SVD in a lower precision as an initial guess and then performs one-sided Jacobi rotations in the working precision as iterative refinement. By carefully transforming a lower precision solution to a higher precision one, our algorithm achieves about \(2\times\) speedup on the x86-64 architecture compared to the usual one-sided Jacobi SVD algorithm in LAPACK, without sacrificing the accuracy. Weiguo Gao, Yuxin Ma 0002, Meiyue Shao |
ACM Trans. Math. Softw. | 3 |
| 2024 | An improved mixed-precision FEAST algorithm for solving symmetric eigenvalue problemsabstractSolving symmetric eigenvalue problems is vital in many areas of scientific computing. FEAST is a well-known package designed for large-scale eigenvalue problems, incorporating mixed-precision techniques to accelerate linear equation solving. Unlike FEAST’s approach, this work introduces a new mixed-precision method that approximates the original eigenvalue problem at a lower precision to quickly provide a good initial guess. These results are then used to accelerate the convergence in working precision. Extensive experiments on large sparse matrices from real applications and randomly generated banded matrices demonstrate the effectiveness of this approach. In appropriate circumstances, our improved mixed-precision FEAST algorithm achieves an average speedup of 1.58× compared to the double-precision FEAST algorithm, with a maximum speedup of 1.79×. Additionally, compared to the original mixed-precision approach in the latest FEAST library, our method offers up to a 1.43× performance improvement. Shengguo Li, Meiyue Shao, Ruixuan Ren |
HPCC | 4 |
| 2024 | Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million AtomsabstractRaman spectroscopy offers invaluable insights into the chemical composition and structural characteristics of various materials, making it a powerful tool for structural analysis. However, accurate quantum mechanical simulations of Raman spectra for large systems, such as biological materials, have been limited due to immense computational costs and technical challenges. In this study, we developed efficient algorithms and optimized implementations on heterogeneous computing architectures to enable fast and highly scalable ab initio simulations of Raman spectra for large-scale biological systems with up to 100 million atoms. Our simulations have achieved nearly linear strong and weak scaling on two cutting-edge high-performance computing systems, with peak FP64 performances reaching 400 PFLOPS on 96,000 nodes of new Sunway supercomputer and 85 PFLOPS on 6,000 node of ORISE supercomputer. These advances provide promising prospects for extending quantum mechanical simulations to biological systems. Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen, Jinfeng Liu 0004, Meiyue Shao, Yingzhou Li, Bowen Kan, Huimin Cui, Xiaobing Feng 0002, Yunquan Zhang, Donald G. Truhlar, Hong An, Xiao He 0004, Jinlong Yang 0003 |
SC | 6 |
| 2023 | Fault-Tolerant LOBPCG for Nuclear CI CalculationsabstractExascale computing platforms with millions of compute units and with thousands of nodes are predicted to experience frequent faults which interrupt applications’ execution. In this context resilience against faults becomes important. We examine user and software level fault mitigation strategies in a distributed LOBPCG algorithm targeting nuclear CI calculations. In particular, we present and evaluate one strategy that keeps the total number of fault-tolerant LOBPCG iterations close to that of the standard LOBPCG algorithm ran on a fault-free machine. Meiyue Shao, Dossay Oryspayev, Chao Yang 0001, Pieter Maris, Brandon Cook 0001 |
HPC Asia | 1 |
| 2017 | A High Performance Block Eigensolver for Nuclear Configuration Interaction CalculationsabstractAs on-node parallelism increases and the performance gap between the processor and the memory system widens, achieving high performance in large-scale scientific applications requires an architecture-aware design of algorithms and solvers. We focus on the eigenvalue problem arising in nuclear Configuration Interaction (CI) calculations, where a few extreme eigenpairs of a sparse symmetric matrix are needed. We consider a block iterative eigensolver whose main computational kernels are the multiplication of a sparse matrix with multiple vectors (SpMM), and tall-skinny matrix operations. We present techniques to significantly improve the SpMM and the transpose operation SpMMT by using the compressed sparse blocks (CSB) format. We achieve 3-4× speedup on the requisite operations over good implementations with the commonly used compressed sparse row (CSR) format. We develop a performance model that allows us to correctly estimate the performance of our SpMM kernel implementations, and we identify cache bandwidth as a potential performance bottleneck beyond DRAM. We also analyze and optimize the performance of LOBPCG kernels (inner product and linear combinations on multiple vectors) and show up to 15× speedup over using high performance BLAS libraries for these operations. The resulting high performance LOBPCG solver achieves 1.4× to 1.8× speedup over the existing Lanczos solver on a series of CI computations on high-end multicore architectures (Intel Xeons). We also analyze the performance of our techniques on an Intel Xeon Phi Knights Corner (KNC) processor. Hasan Metin Aktulga, Md. Afibuzzaman, Samuel Williams 0001, Aydin Buluç, Meiyue Shao, Chao Yang 0001, Esmond G. Ng, Pieter Maris, James P. Vary |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | Algorithm 953: Parallel Library Software for the Multishift QR Algorithm with Aggressive Early DeflationabstractLibrary software implementing a parallel small-bulge multishift QR algorithm with Aggressive Early Deflation (AED) targeting distributed memory high-performance computing systems is presented. Starting from recent developments of the parallel multishift QR algorithm [Granat et al., SIAM J. Sci. Comput. 32(4), 2010], we describe a number of algorithmic and implementation improvements. These include communication avoiding algorithms via data redistribution and a refined strategy for balancing between multishift QR sweeps and AED. Guidelines concerning several important tunable algorithmic parameters are provided. As a result of these improvements, a computational bottleneck within AED has been removed in the parallel multishift QR algorithm. A performance model is established to explain the scalability behavior of the new parallel multishift QR algorithm. Numerous computational experiments confirm that our new implementation significantly outperforms previous parallel implementations of the QR algorithm. Robert A. Granat, Bo Kågström, Daniel Kressner, Meiyue Shao |
ACM Trans. Math. Softw. | 4 |
| 2011 | A Supernodal Approach to Incomplete LU Factorization with Partial PivotingabstractWe present a new supernode-based incomplete LU factorization method to construct a preconditioner for solving sparse linear systems with iterative methods. The new algorithm is primarily based on the ILUTP approach by Saad, and we incorporate a number of techniques to improve the robustness and performance of the traditional ILUTP method. These include new dropping strategies that accommodate the use of supernodal structures in the factored matrix and an area-based fill control heuristic for the secondary dropping strategy. We present numerical experiments to demonstrate that our new method is competitive with the other ILU approaches and is well suited for modern architectures with memory hierarchy. Xiaoye S. Li, Meiyue Shao |
ACM Trans. Math. Softw. | 2 |