Ruonan Zhao

dblp:211/6290 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-7479-585XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Based on Tensor Core Sparse Kernels Accelerating Deep Neural Networks
abstract
Large language models in deep learning have numerous parameters, requiring significant storage space and computational resources. Compression techniques are highly effective in addressing these challenges. With the development of hardware like Graphics Processing Unit (GPU), Tensor Core can accelerate low-precision matrix multiplication but achieve acceleration for sparse matrices is challenging. Due to its sparsity, the utilization of Tensor Cores is relatively low. To address this, we propose the based onTensorCoreCompressedSparseRow format (TC-CSR), which facilitates data loading on GPUs and matrix operations on Tensor Cores. Based on this format, we designed block Sparse Matrix-Matrix Multiplication (SpMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM) kernels, which are common operations in deep learning. Utilizing these designs, we achieved a$\mathbf {1.41\times }$speedup on Sputnik in scenarios of moderate sparsity and a$\mathbf {1.38\times }$speedup with large-scale highly sparse matrices. Benefit from our design, we achieved a$\mathbf {1.75\times }$speedup in end-to-end inference with sparse Transformers and save memory.
Shijie Lv, Debin Liu, Laurence T. Yang, Xiaosong Peng, Ruonan Zhao, Zecan Yang, Jun Feng 0007
IEEE Trans. Parallel Distributed Syst.5
2025 Improving Ethereum Mixing Address Linking With Tensor Computation, Neighbor Data Utilization, and Asymmetric Information Modeling
abstract
Due to the strong untraceability of mixing services, numerous criminals exploit these services to engage in illicit activities, posing a significant threat to the blockchain ecosystem. This paper addresses the challenge of linking transaction addresses in Tornado Cash, a popular mixing service on Ethereum. While existing state-of-the-art solutions like MixBroker attempt to address this problem, two fundamental limitations persist: insufficient utilization of neighbor information and neglect of address information asymmetry. To address these gaps, a novel framework termed “MixLinker” is proposed, which enhances neighbor information utilization and models information asymmetry. Specifically, a Normalized Adjusted Personal PageRank (NAPPR) module is designed to prioritize significant neighbor nodes while mitigating interference from super and irrelevant addresses. Additionally, tensors are employed to model transactions, capturing rich interaction features related to transaction attributes. Based on historical transaction sequences, Tensor Long Short-Term Memory (TLSTM) is used to obtain high-quality initial input features for the Graph Neural Network (GNN) module, enabling effective learning of nonlinear dynamics. To ensure symmetric output results and model asymmetric information, a temporal-aware symmetry classifier is constructed that leverages asymmetric information through permutation operations and an order-aware classifier. Extensive experiments demonstrate that MixLinker outperforms other methods, validating the effectiveness of the proposed approach and confirming the two underlying motivations.
Shuilong Wang, Laurence T. Yang, Debin Liu, Ruonan Zhao, Xianjun Deng, Cannian Zou, Xiaoxuan Fan
IEEE Trans. Inf. Forensics Secur.4
2025 Lightweight Tensor-Enabled GRU for Trustworthy and Communication Efficient Federated Learning in Industrial IoT
abstract
Deep learning provides an intelligent analytical approach for Big Data analysis and feature extraction in Industrial Internet of Things (IIoT). However, due to concerns about data security and privacy disclosure, conventional data-centralized deep learning often faces difficulties about data famine and data islands. Federated learning (FL) as a novel privacy-preserving deep learning paradigm breaks the data islands among different smart factories by sharing their model parameters instead of raw data, which essentially solves the data famine problem for training a high-quality deep learning model. Nevertheless, exchanging numerous model parameters not only generates considerable communication overhead but also poses the risk of privacy information disclosure hidden in model parameters due to inference attacks launched by external attackers orhonest-but-curiousservers. The purpose of this article is to build a high-quality trustworthy FL architecture dubbed TrustFedGRU for IIoT while alleviating the communication overhead. First, a multikey decryption assisted privacy-preserving homomorphic encryption scheme is proposed in FL to meet the distinct privacy-preserving requirements of different data owners without impairing model performance. Furthermore, tensor decomposition is leveraged to convert the weight tensor of the gated recurrent unit (GRU) into a low-rank approximation, so as to reduce the communication bandwidth overhead and decrease the storage requirements. Meanwhile, a novel dynamic update-based FL approach is investigated to improve the model performance. The experimental results show that the proposed TrustFedGRU greatly reduces the communication overhead while guaranteeing the model performance and security.
Ruonan Zhao, Laurence T. Yang, Debin Liu, Wanli Lu, Xiangli Yang
IEEE Trans. Ind. Informatics1
2024 TD3D: Tensor-based Discrete Diffusion Process for 3D Shape Generation
abstract
Based on the recent popularity of diffusion models, we have proposed a tensor-based diffusion model for 3D shape generation (TD3D). This generator is capable of tasks such as unconditional shape generation, shape completion and cross-modal shape generation. TD3D utilizes the Vector Quantized Variational Autoencoder (VQ-VAE) for encoding, compressing 3D shapes into compact latent representations, and then learns the discrete diffusion model based on it. To preserve high-dimensional feature information, we propose a tensor-based noise injection process. To capture 3D features in space faster and use them, a ResNet3D module is introduced during the denoising process. To fuse shape features obtained from ResNet3D and Self-Attention mechanisms, we employ a tensor-based Self-Attention mechanism (T-SA) fusion method. Lastly, a ResNet3D-assisted Multi-Frequency Fusion Module (R-MFM) is designed to aggregate high and low frequency features. Based on the aforementioned design, TD3D provides high fidelity, diverse generated samples, and have the ability to generate 3D shapes across modalities. Extensive experiments have demonstrated its superior performance in various 3D shape generation tasks.
Jinglin Zhao, Debin Liu, Laurence T. Yang, Ruonan Zhao, Zhe Li 0038
ICME4
2024 Incentive Mechanism Against Bounded Rationality for Federated Learning-Enabled Internet of UAVs: A Prospect Theory-Based Approach
abstract
Unmanned aerial vehicles (UAVs) equipped with high definition (HD) cameras, intelligent sensors, computing, and communication modules can be deployed to execute crowdsensing tasks by leveraging federated learning (FL), e.g., air quality perception and ground target detection. FL can reduce transmission stress and protect data privacy when training models, which is suitable for resource constrained Internet of UAVs. Nevertheless, the incentive issues about information asymmetry and bounded rationality impede the applications of FL-enabled Internet of UAVs. The existing FL incentive approaches focus on the risk-free condition, where task publishers are capable of making decisions with complete rationality by utilizing expected utility theory. In fact, task publishers under risk conditions are often bounded rational, whose risk-awareness makes the utility models more sophisticated. To overcome the above problems, we present a prospect theory (PT)-based incentive mechanism for FL-enabled Internet of UAVs. We first leverage PT to model the task publisher’s risk-awareness behavior and construct the subjective utility model. Thereafter, we utilize the framing effect of PT to design the optimal contract to maximize the subjective utility. Simulation results demonstrate that, compared with the baseline method, the proposed incentive mechanism has better performance.
Fang Fu, Yan Wang 0002, Laurence T. Yang, Ruonan Zhao, Yueyue Dai, Zhaohui Yang 0001, Zhicai Zhang
IEEE Internet Things J.5
2024 Dual-Grained Lightweight Strategy
abstract
Removing redundant parameters and computations before the model training has attracted a great interest as it can effectively reduce the storage space of the model, speed up the training and inference of the model, and save energy consumption during the running of the model. In addition, the simplification of deep neural network models can enable high-performance network models to be deployed to resource-constrained edge devices, thus promoting the development of the intelligent world. However, current pruning at initialization methods exhibit poor performance at extreme sparsity. In order to improve the performance of the model under extreme sparsity, this paper proposes a dual-grained lightweight strategy-TEDEPR. This is the first time that TEDEPR has used tensor theory in the pruning at initialization method to optimize the structure of a sparse sub-network model and improve its performance. Specifically, first, at the coarse-grained level, we represent the weight matrix or weight tensor of the model as a low-rank tensor decomposition form and use multi-step chain operations to enhance the feature extraction capability of the base module to construct a low-rank compact network model. Second, unimportant weights are pruned at a fine-grained level based on the trainability of the weights in the low-rank model before the training of the model, resulting in the final compressed model. To evaluate the superiority of TEDEPR, we conducted extensive experiments on MNIST, UCF11, CIFAR-10, CIFAR-100, Tiny-ImageNet and ImageNet datasets with LeNet, LSTM, VGGNet, ResNet and Transformer architectures, and compared with state-of-the-art methods. The experimental results show that TEDEPR has higher accuracy, faster training and inference, and less storage space than other pruning at initialization methods under extreme sparsity.
Debin Liu, Xiang Bai, Ruonan Zhao, Xianjun Deng, Laurence T. Yang
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Multi-Tree Compact Hierarchical Tensor Recurrent Neural Networks for Intelligent Transportation System Edge Devices
abstract
Recurrent neural networks (RNNs) and their variants can efficiently capture the features of time-series characteristic data and are widely used for intelligent transportation tasks. Internet of Vehicles (IoV) edge devices deploying RNN models are an important impetus for the development of intelligent transportation system (ITS) and provide convenient services for users and managers. However, the input data of some transportation tasks have high dimensional characteristics, resulting in the number of training parameters and computational complexity of RNN models being too large, making it difficult to deploy high-performance RNN models on resource-constrained IoV edge devices. To overcome this problem, we compress the training parameters of the RNN model using the proposed multi-tree compact hierarchical tensor representation-Dtensor Block Decomposition (DBD), which reduces the computational complexity of the model and speeds up the training process of the model, thus making the network model lightweight. We evaluate the performance of Dtensor Block-Long Short-Term Memory (DB-LSTM) and Improved Dtensor Block-LSTM (IDB-LSTM) models on multiple real datasets and compare them with the current state-of-the-art LSTM compression models. Experimental results demonstrate that our proposed method can massively compress the number of training parameters of the models on different datasets and shorten the training time of the models without degrading the testing accuracy of the models. In addition, our proposed DB-LSTM and IDB-LSTM models have better comprehensive performance compared with other models and are more suitable for deployment on resource-constrained IoV edge devices.
Debin Liu, Laurence T. Yang, Ruonan Zhao, Xianjun Deng, Chenlu Zhu, Yiheng Ruan
IEEE Trans. Intell. Transp. Syst.3
2024 A Multi-Modal Tensor Ring Decomposition for Communication-Efficient and Trustworthy Federated Learning for ITS in COVID-19 Scenario
abstract
Traffic and the movement of people are inextricably associated with the potential spread of COVID-19. In Intelligent Transportation System (ITS), Deep Learning (DL) traffic detection approaches driven by transportation big data have significant application values in monitoring, counting and classifying traffic vehicle information during the COVID-19 epidemic blockade, while DL COVID-19 medical diagnostic technology is also very important. However, due to concerns about data privacy and security, traditional data-centralized DL techniques that require uploading training data from multiple cameras or hospitals are no longer suitable. Federated Learning (FL) as a novel collaborative privacy-preserving DL paradigm could address this issue well. Nevertheless, in FL, most existing works train learning models with full-precision weights and communicate them over multiple iterations, which may incur massive additional communication costs and disclose the privacy implied in the trained local models. To tackle these issues, we first propose a novel multi-modal tensor ring decomposition TR-TSVD that not only achieves efficient data reduction but also keeps the correlations among multi-modes. Afterward, applying TR-TSVD to the training process of a convolutional neural network under the FL framework to achieve the goal of reducing communication overhead while ensuring model performance. Additionally, since the weight parameters are transmitted with the TR-TSVD format, attackers cannot infer the data privacy without knowing the specific restoration method. Besides, the additively homomorphic encryption is leveraged to further preserve model security. Extensive experimental results on MNIST, BIT-Vehicle and COVID-CT datasets show that the proposed approach could achieve a better performance.
Ruonan Zhao, Laurence T. Yang, Debin Liu, Xiaokang Zhou, Xianjun Deng, Xueming Tang
IEEE Trans. Intell. Transp. Syst.1
2024 Tensor-Empowered LSTM for Communication-Efficient and Privacy-Enhanced Cognitive Federated Learning in Intelligent Transportation Systems
abstract
Multimedia cognitive computing as a revolutionary emerging concept of artificial intelligence emulating the reasoning process like human brains can facilitate the evolution of intelligent transportation systems (ITS) to be smarter, safer, and more efficient. Massive multimedia traffic big data is an important prerequisite for the success of cognitive computing in ITS. However, traditional data-centralized artificial intelligence approaches often face the problems of data islands and data famine due to concerns about data privacy and security. To this end, we propose the concept of cognitive federated learning leveraging federated learning as the learning paradigm for cognitive computing, which solves the preceding concerns by sharing updated models rather than raw data. Nevertheless, the exchange of numerous model parameters not only generates significant communication overhead but also suffers from the risk of privacy leakage due to inference attacks. This article aims to design a novel lightweight and privacy-enhanced cognitive federated learning architecture to facilitate the development of ITS. First, a privacy-enhanced model protection scheme with homomorphic encryption as the underlying technology is proposed to simultaneously defend against the inference attacks launched by external malicious attackers, honest-but-curious cognitive platforms, and internal participants. Furthermore, a novel tensor ring-block decomposition and its corresponding deep computation model converting the weight tensor into a set of matrices and third-order core tensors are proposed, which could reduce the communication overhead and storage requirements without compromising model performance. Experimental results on real-world datasets show that the proposed approach performs well.
Ruonan Zhao, Laurence T. Yang, Debin Liu, Wanli Lu, Chenlu Zhu, Yiheng Ruan
ACM Trans. Multim. Comput. Commun. Appl.1
2023 CP-Decomposition Based Federated Learning with Shapley Value Aggregation
abstract
Federated learning enables multiple data providers to collaborate on training models without exposing personal data. During the training process, frequent communication is required between the data provider and the central server, which puts great pressure on federated learning. To reduce the communication pressure of federated learning, we use the CP-decomposition processing model to reduce the size of data that needs to be transmitted during the communication process. In addition, we aggregate the global model based on the Shapley value, and eliminate nodes that are not beneficial to federated learning as soon as possible, which reduces the communication pressure and can stimulate the participating nodes and enhance the enthusiasm of participants in federated learning, thus improving the training results of the global model. We named the system CPSV, which stands for Federated learning of CP-decomposition models based on Shapley value aggregation. Numerous experiments on CPSV have shown that CPSV can motivate and supervise participating nodes to aggregate better global models while reducing the stress of federal learning communication.
Chengqian Wu, Xuemei Fu, Xiangli Yang, Ruonan Zhao, Qidong Wu, Tinghua Zhang
ICPADS4
2023 Tensor-Enabled Communication-Efficient and Trustworthy Federated Learning for Heterogeneous Intelligent Space-Air-Ground-Integrated IoT
abstract
Federated learning (FL) could provide a promising privacy-preserving intelligent learning paradigm for space–air–ground-integrated Internet of Things (SAGI-IoT) by breaking down data islands and solving the dilemma between data privacy and data sharing. Currently, adaptivity, communication efficiency and model security are the three main challenges faced by FL, and they are rarely considered by existing works simultaneously. Concretely, most existing FL works assume that local models share the same architecture with the global model, which is less adaptive and cannot meet the heterogeneous requirements of SAGI-IoT. Exchanging numerous model parameters not only generates massive communication overhead but also poses the risk of privacy leakage. The security of FL based on homomorphic encryption with a single private key is weak as well. Given this, this article proposes a tensor-empowered communication-efficient and trustworthy heterogeneous FL, where various participants could choose suitable heterogeneous local models according to their actual computing and communication environment, so that clients with different capabilities could do what they are good at. Additionally, tensor train decomposition is leveraged to reduce communication parameters while maintaining model performance. The storage requirements and communication overhead for heterogeneous clients are reduced further. Finally, the homomorphic encryption with double trapdoor property is utilized to provide a robust and trustworthy environment, which can defend against the inference attacks from malicious external attackers,honest-but-curiousserver and internal participating clients. Extensive experimental results show that the proposed approach is more adaptive and can improve communication efficiency as well as protect model security compared with the state-of-the-art.
Ruonan Zhao, Laurence T. Yang, Debin Liu, Wanli Lu
IEEE Internet Things J.1
2022 Lightweight Tensor Deep Computation Model With Its Application in Intelligent Transportation Systems
abstract
Deep computation models (DCMs) are widely used in intelligent transportation systems (ITS), like driving behavior detection, intelligent parking navigation and real-time road condition detection. Due to the multi-source heterogeneous nature of big data of the ITS, it is difficult for traditional DCMs to learn effective multi-modal data features. Although, the DCMs in tensor space can efficiently represent multi-modal data, it further worsens the problem of model learning parameter explosion. In this paper, we propose a lightweight tensor DCM. The model compresses the redundant learning parameters of the model and reduces the consumption of computational resources while maintaining the learning characterization capability of the DCM in tensor space, thus making the network model more general and lightweight for deploying the DCM to smart cars and edge devices. The proposed lightweight tensor DCM is evaluated on several real datasets. The experimental results show that the number of learning parameters is massively compressed while keeping the performance of the network model almost constant, while also reducing the computational complexity and training time of the model.
Debin Liu, Laurence T. Yang, Ruonan Zhao
IEEE Trans. Intell. Transp. Syst.3
2022 A Tensor-Based Truthful Incentive Mechanism for Blockchain-Enabled Space-Air-Ground Integrated Vehicular Crowdsensing
abstract
Space-Air-Ground Integrated Network (SAGIN) as an efficient newly integration network could provide more comprehensive network services to meet the multifarious quality of service requirements in different Intelligent Transportation Systems (ITS). By taking advantage of SAGIN, Space-Air-Ground Integrated Vehicular Crowdsensing (SAGI-VCS) would have great potential and the services regarding ITS could be facilitated. However, centralized SAGI-VCS is usually vulnerable to malicious attacks and the trust issues are one of the main reasons that hinder its further development. Blockchain as a distributed hyperledger shows a vital potential to solve the trust problem of multiple participants who do not trust each other and tackle the security issues in SAGI-VCS. Additionally, selfishness is another factor that prevents vehicles from participating in SAGI-VCS. The vast majority of existing incentives for vehicular crowdsensing only focus on the terrestrial networks which cannot be directly used in SAGI-VCS. Meanwhile, the redundant winner phenomenon and the multi-attributes of participants are less considered by them. Toward this end, we first illustrate a blockchain-enabled service architecture for SAGI-VCS and then construct a unified representation model. Afterwards, a tensor computing based truthful incentive mechanism TensorBC for blockchain-enabled SAGI-VCS is proposed to motivate vehicles to participate in completing tasks, ensure the security of the whole process and maximize the social welfare. TensorBC not only can eliminate the redundant winner phenomenon, but also can guarantee the economic properties such as truthfulness, individual rationality and profitability. Finally, both the rigorous theoretical analysis and extensive experimental results show that TensorBC could achieve a better performance.
Ruonan Zhao, Laurence T. Yang, Debin Liu, Xianjun Deng, Yijun Mo
IEEE Trans. Intell. Transp. Syst.1
2022 TT-TSVD: A Multi-modal Tensor Train Decomposition with Its Application in Convolutional Neural Networks for Smart Healthcare
abstract
Smart healthcare systems are generating a large scale of heterogenous high-dimensional data with complex relationships. It is hard for current methods to analyze such high-dimensional healthcare data. Specifically, the traditional data reduction methods can not keep the correlation among different modalities of data objects, while the latest methods based on tensor singular value decomposition are not effective for data reduction, although they can keep the correlation. This article presents a tensor train-tensor singular value decomposition (TT-TSVD) algorithm for data reduction. Particularly, the presented algorithm balances the correlation-preservation ability of modalities and data reduction ability by combining the advantages of the train structure of the tensor train decomposition and the association relationship between the tensor singular value decomposition retention mode. Extensive experiments are conducted on the convolutional neural network and the results clearly show that the presented algorithm performs effectively for data reduction with a low-loss classification accuracy; what is more, classification accuracy on medical image dataset has been improved a little.
Debin Liu, Laurence T. Yang, Puming Wang, Ruonan Zhao, Qingchen Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2018 An Efficient Energy-Aware Probabilistic Routing Approach for Mobile Opportunistic Networks
Ruonan Zhao, Lichen Zhang 0001, Xiaoming Wang 0001, Chunyu Ai, Fei Hao 0001, Yaguang Lin
WASA1
2018 An on-demand coverage based self-deployment algorithm for big data perception in mobile sensing networks
Yaguang Lin, Xiaoming Wang 0001, Fei Hao 0001, Liang Wang 0014, Lichen Zhang 0001, Ruonan Zhao
Future Gener. Comput. Syst.6