Zhuo Tian

dblp:223/6542 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Pattern-Aware Finite Element Matrix Assembly Method on GPUs
abstract
The Finite Element Method (FEM) is a fundamental technique for solving large-scale and complex engineering problems. During the construction of the system equations, the efficiency of finite element matrix assembly plays a crucial role in the overall performance. However, existing approaches often overlook the sensitivity of assembly algorithm performance to mesh characteristics, making it difficult to achieve optimal performance across diverse problems. In this work, we propose a novel pattern-aware FEM matrix assembly method on GPUs. To this end, we thoroughly analyze the key factors affecting performance and extract a set of potentially influential mesh features and density representations. Based on this, we construct a Deep learning-based prediction model that fully captures the input mesh characteristics to predict the performance-optimal assembly strategy. Experimental results on mesh datasets with a wide range of feature variations demonstrate that our method achieves remarkable prediction accuracy and delivers up to$7.34 \times$speedup in execution time compared to state-of-the-art approaches. To the best of our knowledge, this is the first work that introduces auto-tuning for the FEM matrix assembly process.
Changyou Zhang, Zhuo Tian, Guangzhao Li, Chen Ju
CLUSTER4
2025 RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property Protection
abstract
Reversible Adversarial Examples (RAE) are designed to protect the intellectual property of datasets. Such examples can function as imperceptible adversarial examples to erode the model performance of unauthorized users while allowing authorized users to remove the adversarial perturbations and recover the original samples for normal model training. With the rise of Self-Supervised Learning (SSL), an increasing number of unlabeled datasets and pre-trained encoders are available in the community. However, existing RAE methods not only rely on well-labeled datasets for training Supervised Learning (SL) models but also exhibit poor adversarial transferability when attacking SSL pre-trained encoders. To address these challenges, we propose RAEncoder, the first framework for RAEs without the need for labeled samples. RAEncoder aims to generate universal adversarial perturbations by targeting SSL pretrained encoders. Unlike traditional RAE approaches, the pre-trained encoder outputs the feature distribution of the protected dataset rather than classification labels, enhancing both the attack success rate and transferability of RAEs. Extensive experiments are conducted on six pre-trained encoders and four SL models, covering aspects such as imperceptibility and transferability. Our results demonstrate that RAEncoder effectively protects unlabeled datasets from malicious infringements. Additional robustness experiments further confirm the security of RAEncoder in practical application scenarios.
Fan Xing, Zhuo Tian, Xuefeng Fan, Xiaoyi Zhou
CVPR2
2024 RAEDiff: Diffusion Models Enable Self-Generation and Self-Recovery of Reversible Adversarial Examples
Fan Xing, Xiaoyi Zhou, Hongli Peng, Zhuo Tian, Xuefeng Fan
ICONIP (6)4
2024 TRAE: Reversible Adversarial Example with Traceability
Zhuo Tian, Xiaoyi Zhou, Fan Xing, Wentao Hao, Ruiyang Zhao
PRCV (1)1
2024 Towards the Transferable Reversible Adversarial Example via Distribution-Relevant Attack
Zhuo Tian, Xiaoyi Zhou, Fan Xing, Ruiyang Zhao
PRCV (11)1
2024 A Photovoltaic Hot-Spot Fault Detection Network for Aerial Images Based on Progressive Transfer Learning and Multiscale Feature Fusion
abstract
The number of samples is one of the key factors affecting the performance of deep learning-based detection networks. Aiming at the problem that the detection network is difficult to accurately detect the hot-spot fault targets under the condition of small samples, a photovoltaic hot-spot fault detection network based on progressive transfer learning and multiscale feature fusion is proposed. First, a large number of artificial hot-spot samples are generated through the artificial model, and the mixed dataset containing real and artificial samples is constructed to improve the data diversity. On this basis, a pre-trained model based on artificial samples is established to learn the shallow features of hot-spot faults. Then, to fuse the multiscale features and improve feature aggregation ability of detection network, a novel feature pyramid structure based on reparameterized generalized and multiscale feature fusion (RepG-MSFF) is designed. Moreover, to balance the detection accuracy and speed, the spatial and channel reconstruction convolution (SCConv) is utilized to replace conventional convolution in the backbone network. Finally, to further accurately locate hot-spot targets, an adaptive threshold focal loss (TFL) function is introduced. The experimental results indicate that, in three different scenarios datasets, the detection accuracy can reach 87.9%, 88.6%, and 87.7%, respectively, which is higher than that of other nine detection algorithms.
Shuai Hao 0003, Siya Sun, Zhuo Tian, Yifeng Hou
IEEE Trans. Geosci. Remote. Sens.5
2023 Accelerating Sparse General Matrix-Matrix Multiplication for NVIDIA Volta GPU and Hygon DCU
abstract
Sparse general matrix-matrix multiplication (SpGEMM) is challenging especially on graphic accelerators. Existing solutions do not fully utilize the shared memory of the graphics accelerator. Our proposal could effectively utilize the graphics accelerator's on-chip shared memory and dynamically assign the device resources by grouping the rows based on a hybrid strategy for load balancing. Experiments show that our proposal achieves speedups of up to x7.43 in double precision compared to existing SpGEMM libraries. Our implementation is fully general and our optimization strategy adaptively processes the SpGEMM workload row-wise to substantially improve performance by decreasing the work complexity and utilizing the memory hierarchy more effectively.
Zhuo Tian, Changyou Zhang
HPDC1
2022 An Asynchronous Parallel Algorithm to Improve the Scalability of Finite Element Solvers
abstract
Large-scale finite element equations are usually solved by the Preconditioned Conjugate Gradient (PCG) iterative method, and the computational hotspots are sparse matrix-vector multiplication and inner product, which requires local and global communication. But, the latency of the high-performance cluster is too long for the PCG algorithm. This paper proposes a new asynchronous algorithm that could reduce the number of communication among cluster nodes to improve the scalability of finite element solvers. The performance could be improved up to 6.31 x.
Zhuo Tian, Changyou Zhang
CLUSTER1
2019 An Asynchronous Algorithm to Reduce the Number of Data Exchanges
Zhuo Tian
ICA3PP (2)1