Min Tian 0005

dblp:91/5889-5 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
0009-0000-9930-4802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2025 RKOP: A Parallel Randomized Kaczmarz Algorithom Based on Oblique Projection for Large-Scale Overdetermined Equations
abstract
The randomized Kaczmarz algorithm is a simple iterative method for solving overdetermined linear systems. However, the classical randomized Kaczmarz algorithm relies on orthogonal projections, and its convergence rate deteriorates significantly when the system exhibits a high linear dependence. This paper describes a novel oblique projection solution, RKOP, a randomized Kaczmarz algorithm with oblique projections based on maximum cosine similarity. Firstly, we select hyperplanes based on maximum cosine similarity and construct oblique projection directions using two hyperplanes to accelerate convergence. Secondly, a dynamic Monte Carlo error estimation method is employed to reduce the computational overhead of error evaluation effectively. Finally, we implement a multi-level parallel framework that achieves effective load balancing through optimized data distribution and uses a delayed update strategy to reduce computational overhead and improve overall efficiency significantly. Experimental results demonstrate that the serial version of RKOP achieves a$64.25 \times$speedup over the traditional randomized Kaczmarz algorithm. When scaled to 32 cores, RKOP achieves a$57.91 \times$speedup compared to its single-core version.
Min Tian 0005, Yunhui Zeng, Jidong Huo
HPCC2
2024 swPTS: an efficient parallel Thomas split algorithm for tridiagonal systems on Sunway manycore processors
Min Tian 0005, Qi Liu 0065, Jingshan Pan, Ying Gou, Zanjun Zhang
J. Supercomput.1
2023 SW-LeNet: Implementation and Optimization of LeNet-1 Algorithm on Sunway Bluelight II Supercomputer
Zenghui Ren, Tao Liu 0029, Zhaoyuan Liu, Min Tian 0005, Ying Guo 0028, Jingshan Pan
ICA3PP (5)4
2023 AH-TDMA: An Adaptive Heterogeneous Tridiagonal Matrix Algorithm on the New Sunway Supercomputer
abstract
The solution of tridiagonal linear systems is used in in various fields and plays a crucial role in numerical simulations. However, there is few efficient solver for tridiagonal linear systems on the new Sunway supercomputer. Based on a three-dimensional heat conduction problem, we propose an adaptive heterogeneous tridiagonal matrix algorithm (AH-TDMA). Our major innovations include: (1) To address computational hotspots within AH-TDMA, a multi-level parallel approach involving MPI+Athread has been adopted. (2) An adaptive data partitioning scheme has been set up to achieve load balance. (3) Employing direct memory access and establishing shared space between registers and main memory, instead of employing discrete memory access, is done to enhance memory access efficiency. (4) The optimization of the loop structure has been made in adjusting the sequencing of the dual-layered loops to reduce communication overhead. The experimental results show that, with one core group, the AH-TDMA achieves a speedup of 99.6 times for hotspot and total time speedup up to 59.4 times compared to the Parallel and Scalable Library for Tridiagonal Matrix Algorithm (PaScal TDMA). The AH-TDMA is scalable up to 2048 core groups, with a parallel efficiency of 69.2%.
Min Tian 0005, Qi Liu 0065, Zenghui Ren, Yue Liu 0027
ICPADS1
2023 hcaPCG: A Heterogeneous and communication-avoid PCG with Jacobi preconditioner on SW26010-Pro architecture
abstract
Due to its efficiency and versatility, the preconditioned conjugate gradient algorithm has long been a staple in the realm of iterative linear system solvers. In this paper, we proposed an optimized preconditioned conjugate gradient algorithm tailored for the SW26010-Pro manycore processor, The main work includes: Optimizing data block sizes based on the processor’s storage structure; combining thread-level and data-level parallelism, utilizing manual SIMD to improve efficiency; optimizing the memory access pattern by employing direct memory access and overlapping of computation and communication; employing a shared-memory approach to store long vectors across all cores. Furthermore, we design an accelerated algorithm for reduction operations, to avoid data communication. Experimental results show that the hcaPCG yields up to 28.1× speedups on average compared to the original implementation.
Min Tian 0005, Yue Liu 0027, Qi Liu 0065, Jingshan Pan
ICPADS1
2023 SPM-GCN: An adaptive reordering algorithm for sparse LU factorization via GCN
abstract
Sparse LU factorization is a critical kernel in scientific computing and engineering applications. A better nonzero pattern of sparse matrixes can accelerate LU factorization by reordering. Traditionally it’s difficult to predict which non-zero pattern is optimal for a sparse matrix. In this paper, we proposed a graph convolutional neural network (GCN) for adaptively selecting the optimal reordering algorithm from five candidate reordering algorithms in the Matrix preprocessing step, referred to as SPM-GCN. After using our SPM-GCN, the average numerical factorization time outperformed the other five algorithms, and the average numerical factorization time was reduced by 17.3% compared with the default method of swSuperLU.
Min Tian 0005, Huazeng Liu, Qi Liu 0065, Zanjun Zhang
ICPADS1
2023 PIOD: An Efficient Parallel Iterative Algorithm for Solving Over-determined Equations
abstract
The solution of over-determined equations plays a very important role in fields such as data fitting, signal processing, and machine learning. It is of great significance in predicting natural phenomena, optimizing engineering design, and other fields. However, there is currently no efficient method to solve over-determined equations, either being not accurate enough or consuming a lot of time. In this article, we propose a parallel iterative method for solving over-determined equations, called the PIOD algorithm. By using a sub-convergence condition to terminate the iterative calculation, we have developed a task partitioning strategy for the algorithm and implemented parallelization of the solution of over-determined equations on a distributed memory system. Our proposed algorithm achieves an average speedup of 152 ×times compared to the open-source Eigen solver. Additionally, it also achieves a parallel efficiency of over 30%.
Min Tian 0005, Zhenguo Wei, Lei Xiao 0002, Chaoshuai Xu
ICPADS1
2023 Parallel optimization of method of characteristics based on Sunway Bluelight II supercomputer
Renjiang Chen, Tao Liu 0029, Zhaoyuan Liu, Min Tian 0005, Ying Guo 0028, Jingshan Pan, Meihong Yang
J. Supercomput.5
2023 swParaFEM: a highly efficient parallel finite element solver on Sunway many-core architecture
Jingshan Pan, Lei Xiao 0002, Min Tian 0005, Tao Liu 0029, Yinglong Wang 0001
J. Supercomput.3
2022 swSuperLU: A highly scalable sparse direct solver on Sunway manycore architecture
Min Tian 0005, Zanjun Zhang, Jingshan Pan, Tao Liu 0029
J. Supercomput.1