Zhixiang Zhao

dblp:246/5222 · also Zhi-Xiang Zhao · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Post-quantum Internet Key Exchange via Authenticated Forward-Secure KEM
Yunlei Zhao, Biming Zhou, Zhixiang Zhao, Haodong Jiang
CRYPTO (10)3
2025 Dual-Band Sensing for Passive Target Surveillance in ISAC Systems
abstract
A ubiquitous and high-performance sensing service is required in the next generation of wireless communication systems. The legacy Sub-6GHz bands can provide a larger sensing range, but the available bandwidth limits the resolution. In contrast, the millimeter-wave (mmWave) frequencies enhance the sensing performance with wider bandwidth. However, high-directional beamforming poses new challenges in passive target surveillance. The conventional beam sweeping in mm Wave is costly due to the significant overhead, which conflicts with the goal of communication on efficient data transmission. To reduce the overhead caused by target surveillance, we propose a dual-band sensing approach. By combining the advantages of both Sub-6GHz and mmWave frequencies, our method enables fine target detection in a wide area. Compared to beam sweeping, the pro-posed method significantly reduces the communication overhead and time consumption, while also lightening the computational load of radar signal processing.
Zhixiang Zhao, Carsten Smeenk, Sebastian Semper, Christian Schneider 0003, Reiner S. Thomä
WCNC1
2025 Recursive Hybrid Compression for Sparse Matrix-Vector Multiplication on GPU
abstract
ABSTRACT Sparse Matrix‐Vector Multiplication (SpMV) is a fundamental operation in scientific computing, machine learning, and data analysis. The performance of SpMV on GPUs is crucial for accelerating various applications. However, the efficiency of SpMV on GPUs is significantly affected by irregular memory access patterns, high memory bandwidth requirements, and insufficient exploitation of parallelism. In this paper, we propose a Recursive Hybrid Compression (RHC) method to address these challenges. RHC begins by splitting the initial matrix into two portions: an Ellpack (ELL) portion and a Coordinate (COO) portion. This partitioning is followed by further recursive division of the COO portion into additional ELL and COO portions, continuing this process until predefined termination criteria, based on a percentage threshold of the number of nonzero elements, are met. Additionally, we introduce a dynamic partitioning method to determine the optimal threshold for partitioning the matrix into ELL and COO portions based on the distribution of nonzero elements and the memory footprint. We develop the RHC algorithm to fully exploit the advantages of the ELL kernel on GPUs and achieve high thread‐level parallelism. We evaluated our proposed method on two different NVIDIA GPUs: the GeForce RTX 2080 Ti and the A100, using a set of sparse matrices from the SuiteSparse Matrix Collection. We compare RHC with NVIDIA's cuSPARSE library and three state‐of‐the‐art methods: SELLP, MergeBase, and BalanceCSR. RHC achieves average speedups of 2.13, 1.13, 1.87, and 1.27 over cuSPARSE, SELLP, MergeBase, and BalanceCSR, respectively.
Zhixiang Zhao, Yanxia Wu 0001, Guoyin Zhang, Yiqing Yang, Ruize Hong
Concurr. Comput. Pract. Exp.1
2025 The Time and Frequency Distribution Characteristics of Interference Signals Based on Artificial Intelligence Technology
abstract
In this paper, we propose a method with artificial intelligence to optimize manual astronomical observation works. For the large amount of data generated by radio astronomy monitoring, we compile 4 algorithms including VTD, WSV, MAD, and MAS in the procedure of data analysis. Then the platform can recognize the radio interference signals from radio astronomy monitoring data and analyze the spatiotemporal distribution characteristics, and generate reports automatically. Through this method, the amount of work would greatly improve work efficiency and accuracy. The distribution patterns and changes of radio frequency interference signals in the area can be grasped and analyzed efficiently and quickly by astronomy researchers.
Shengyang Li, Zhixiang Zhao, Junwen Tang, Ning Fu
Int. J. Pattern Recognit. Artif. Intell.2
2025 FSD-GAN: Generative Adversarial Training for Face Swap Detection via the Latent Noise Fingerprint
Jiawei Ge 0002, Jiuxin Cao, Zhixiang Zhao, Bo Liu 0004
J. Comput. Sci. Technol.3
2024 Optimizing Radio Resources for Radar Services in ISAC Systems by Deep Reinforcement Learning
abstract
In integrated sensing and communication (ISAC) systems, the radar and communication functionality share the same infrastructure and radio resources. The flexible access scheme in mobile communication systems allows for variable and efficient radio resource allocation for multiple users and services. In this paper, we present a resource allocation strategy for Orthogonal Frequency Division Multiple (OFDM) based radar sensing. Furthermore, we propose a Deep Reinforcement Learning (DRL) method combined with classic OFDM radar signal processing techniques to optimize the radio resources for radar sensing. Two agents are trained with the DRL method on simulated data to predict the radio resources for the subsequent signal, aiming to achieve the required radar performance while improving the resource efficiency. One agent optimizes the transmission power, while the other optimizes the signal bandwidth/duration. The investigated scenario is a highway with traffic. We evaluate the performance of the agents based on detection loss and radio resource efficiency the metrics. Further, we compare the results against random and maximum resource-selecting agents.
Carsten Smeenk, Zhixiang Zhao, Christian Schneider 0003, Jörg Robert, Giovanni Del Galdo
PIMRC2
2024 Split-bucket partition (SBP): a novel execution model for top-K and selection algorithms on GPUs
abstract
Abstract Top-K and selection operations are critical in data processing and analysis, and their efficient implementation on GPUs is increasingly important due to the growing demands of data analysis. Existing methods, primarily relying on the bucket partition execution model, encounter challenges such as uneven bucket distribution and latency in merging processes. To address these issues, we introduce a novel Split-Bucket Partition (SBP) execution model that specifically addresses these challenges. Additionally, we propose task and control flow optimizations targeted at top-K and selection algorithms, which further contribute to performance improvements. Our optimized algorithms significantly outperform existing approaches, delivering performance gains of up to $$2.3$$ 2.3 times and $$1.6$$ 1.6 times for different bucket partitioning rules. Our algorithms show robust performance improvements in non-uniform data scenarios, with gains ranging from $$1.9$$ 1.9 times to $$15.5$$ 15.5 times. However, it should be noted that the SBP model has limitations related to shared memory and register utilization, potentially impacting performance. Tests on TU102 and A100 GPU architectures validate the effectiveness of our approach, achieving a maximum speedup of $$2.9$$ 2.9 times. The study suggests that while the SBP model is effective for top-K and selection algorithms, it also holds promise for other computational tasks, setting the stage for future research.
Yiqing Yang, Guoyin Zhang, Yanxia Wu 0001, Zhixiang Zhao
J. Supercomput.4
2024 Block-wise dynamic mixed-precision for sparse matrix-vector multiplication on GPUs
abstract
Abstract Sparse matrix-vector multiplication (SpMV) plays a critical role in a wide range of linear algebra computations, particularly in scientific and engineering disciplines. However, the irregular memory access patterns, extensive memory usage, high bandwidth requirements, and underutilization of parallelism hinder the computational efficiency of SpMV on GPUs. In this paper, we propose a novel approach called block-wise dynamic mixed-precision (BDMP) to address these challenges. Our methodology involves partitioning the original matrix into uniformly sized blocks, with each block’s size determined by considering architectural characteristics and accuracy requirements. Additionally, we dynamically assign precision to each block using a precision selection method that takes into account the value distribution of the original sparse matrix. We develop two distinct SpMV computation algorithms for BDMP: BDMP-PBP (Precision-based partitioning) and BDMP-TCKI (Tailored compression and kernel implementation). BDMP-PBP partitions the matrix into two independent matrices for separate computations based on block precision, offering flexibility for integration with other optimization techniques. Meanwhile, BDMP-TCKI focuses on achieving significant thread-level parallelism and memory utilization by tailoring an appropriate compressed storage format and kernel implementation for each block. We compare BDMP with NVIDIA’s cuSPARSE library and three state-of-the-art SpMV methods, including SELLP, MergeBase, and BalanceCSR, using matrices from the University of Florida’s SuiteSparse dataset collection. BDMP-PBP and BDMP-TCKI show average speedups up to 2.64 $$\times $$ × and 2.91 $$\times $$ × on Turing RTX 2080Ti, and up to 2.99 $$\times $$ × and 3.22 $$\times $$ × on Ampere A100. The results demonstrate that BDMP enables the optimization of computation speed without compromising the precision necessary for reliable results.
Zhixiang Zhao, Guoyin Zhang, Yanxia Wu 0001, Ruize Hong, Yiqing Yang
J. Supercomput.1
2020 Number Theoretic Transform: Generalization, Optimization, Concrete Analysis and Applications
Zhichuang Liang, Shiyu Shen 0001, Yuantao Shi, Dongni Sun, Chongxuan Zhang, Guoyun Zhang, Yunlei Zhao, Zhixiang Zhao
Inscrypt8