Yunhui Zeng

dblp:143/5266 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing SPAD HDR Imaging via Cross-Modality Frequency-Spatial Fusion
Yuhe Chen, Yunhui Zeng
ISCAS2
2026 S-Lop: An accuracy-improved performance model of MPI communication based on network contention
Chen Yuan 0016, Yunhui Zeng
Future Gener. Comput. Syst.3
2025 RKOP: A Parallel Randomized Kaczmarz Algorithom Based on Oblique Projection for Large-Scale Overdetermined Equations
abstract
The randomized Kaczmarz algorithm is a simple iterative method for solving overdetermined linear systems. However, the classical randomized Kaczmarz algorithm relies on orthogonal projections, and its convergence rate deteriorates significantly when the system exhibits a high linear dependence. This paper describes a novel oblique projection solution, RKOP, a randomized Kaczmarz algorithm with oblique projections based on maximum cosine similarity. Firstly, we select hyperplanes based on maximum cosine similarity and construct oblique projection directions using two hyperplanes to accelerate convergence. Secondly, a dynamic Monte Carlo error estimation method is employed to reduce the computational overhead of error evaluation effectively. Finally, we implement a multi-level parallel framework that achieves effective load balancing through optimized data distribution and uses a delayed update strategy to reduce computational overhead and improve overall efficiency significantly. Experimental results demonstrate that the serial version of RKOP achieves a$64.25 \times$speedup over the traditional randomized Kaczmarz algorithm. When scaled to 32 cores, RKOP achieves a$57.91 \times$speedup compared to its single-core version.
Min Tian 0005, Yunhui Zeng, Jidong Huo
HPCC4
2025 An Improved Distributed SpMV with Cache-Aware Optimization and Load Balancing
abstract
Sparse Matrix-Vector Multiplication (SpMV) is a critical operation in scientific and high-performance computing, playing a key role in large-scale numerical simulations. Handling extremely large sparse matrices on a single node is challenging, making distributed computing an effective solution. Existing distributed SpMV optimizations often overlook data locality and suffer from communication bottlenecks. This paper proposes two optimization strategies for distributed SpMV based on the Compressed Sparse Row (CSR) format to improve multi-node efficiency. First, combining cache-aware memory access with binary search, we partition thread tasks across nodes to enable fine-grained non-zero element distribution, reducing memory conflicts, load imbalance, and communication overhead. Second, we design a distributed SpMV kernel incorporating vectorization and load balancing to enhance performance. These strategies improve coordination across nodes and resource utilization. Experiments show average speedups of 81.82× and 7.99× over classical and graph-partitioned distributed SpMV, with peak speedups of 462.34× and 45.85×. Compared to recent distributed SpMV algorithms that balance computation and communication, our method achieves average and peak speedups of 1.23× and 4.27×. The algorithm scales well on large irregular sparse matrices, demonstrating the effectiveness of multi-node collaborative optimization.
Yunkun Shang, Yunhui Zeng
SMC3
2025 A barotropic solver capable of reducing global synchronization latency in parallel ocean program
Xiaoyong Fan, Yunhui Zeng
CCF Trans. High Perform. Comput.3
2024 A transmission optimization method for MPI communications
Jubin Wang, Yunhui Zeng
J. Supercomput.3
2020 Redistributing and Optimizing High-Resolution Ocean Model POP2 to Million Sunway Cores
Yunhui Zeng
ICA3PP (1)1
2017 A Security Analysis Method for Supercomputing Users' Behavior
abstract
Supercomputers are widely applied in various domains, which have advantage of high processing capability and mass storage. With growing supercomputing users, the system security receives comprehensive attentions, and becomes more and more important. In this paper, according to the characteristics of supercomputing environment, we perform an in-depth analysis of existing security problems in the process of using resources. To solve these problems, we propose a security analysis method and a prototype system for supercomputing users' behavior. The basic idea is to restore the complete users' behavior paths and operation records based on the supercomputing business process and track the use of resources. Finally, the method is evaluated and the results show that the security analysis method of users' behavior can help administrators detect security incidents in time and respond quickly. The final purpose is to optimize and improve the security level of the whole system.
Yunhui Zeng
CSCloud2