VLDB 2026 Research / reviewers in the wild / expert
Chong Yu 0002
dblp:128/4478-2
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0002-6244-3486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 3 first-author · 11 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multidiffusion Information Centrality for the Identification of Influential Spreaders in Temporal Social NetworksabstractIdentifying influential spreaders in a temporal social network, which has potential applications including network immunization, epidemic control, and viral marketing, is a fundamental class of problems. In this context, various centrality algorithms have been introduced to quantify influential spreaders, focusing on three categories: topology-based methods, dynamics-based methods, and machine learning-based methods. However, topology-based methods tend to consider single temporal features, while the consideration of multi-temporal features is subject to the same challenges of high temporal complexity as dynamics-based methods, and machine learning-based methods face challenges related to dependency on the training dataset. In this paper, we propose a novel centrality algorithm based on multiple diffusion information (RPT: Multi-diffusion information centrality based on R-path trees) to identify the influential node in a temporal social network. This algorithm considers three different temporal features and has lower temporal complexity using a newly proposed representation structure known as an R-path tree (a distinctive inverted tree that encompasses the earliest arrival paths from other nodes to the root node). Through experiments carried out on 12 empirical social networks, the results show that the effectiveness of RPT in identifying influential spreaders generally exceeds that of other baseline measures. Xuelong Yu, Shukun Yang, Hai Zhao 0002, Kuan Zhang 0001, Chong Yu 0002 |
IEEE Internet Things J. | 6 |
| 2025 | Beyond Access Pattern: Efficient Volume-Hiding Multi-Range Queries Over Outsourced Data ServicesabstractMulti-range query (MRQ) is a typical multi-attribute data query widely used in various practical applications. It is capable of searching all data objects contained in a query request. Many privacy-preserving MRQ schemes have been proposed to realize MRQ on encrypted data. However, existing MRQ schemes only consider the security threat caused by access pattern leakage, not the harm of volume pattern leakage. Moreover, most existing schemes cannot achieve efficient queries and updates while preserving the access pattern. In this paper, we propose an efficient MRQ scheme for hiding volume and access patterns. We first design a joint data index using Order-Revealing Encryption (ORE) and Pseudo-random functions (PRFs) to realize volume-hiding range queries. Then, we combine the private set intersection (PSI) and hardware Software Guard Extensions (SGX) to compute each attribute’s intersection of query results. In addition, we preserve access patterns during queries by designing a batch refresh algorithm and an update protocol. Finally, rigorous security analysis and extensive experiments demonstrate the security and performance of our scheme in real-world scenarios. Haoyang Wang 0005, Kai Fan 0001, Chong Yu 0002, Kuan Zhang 0001, Fenghua Li 0001, Haojin Zhu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Hide Yourself: Multi-Dimensional Range Queries for Responses-Hiding Over Outsourced DataabstractMulti-dimensional range query (MRQ) over outsourced data has been extensively applied in various domains. However, security and efficiency are still two aspects that cannot be easily balanced in private MRQs, as improving security inevitably incurs high computation, storage, and communication costs. Several schemes perform encrypted data retrieval in the trusted execution environment (TEE), which balances security and performance. Unfortunately, they focused on keywords or single-dimensional range queries, failing to address private MRQs. With the TEE (i.e., Intel SGX), we propose a response-hiding MRQ scheme over encrypted data (SGX-MRQ) in this paper. We first design an index structure called SDic, which can achieve efficient range queries while hiding the responses to each query from the server. Moreover, based on the security properties of SGX, we construct the encrypted polynomials of each dimension on the enclave and implement the intersection computation of multi-attribute queries by the server, which greatly improves the system efficiency. We present the formal definition of SGX-MRQ and perform a rigorous proof. We implement a prototype of SGX-MRQ and conduct extensive experiments on real datasets. The evaluation results validate the feasibility of our scheme in practical applications. Haoyang Wang 0005, Kai Fan 0001, Chong Yu 0002, Kuan Zhang 0001, Fenghua Li 0001, Haojin Zhu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Joint Communication and Control Optimization of a Multi-Vehicle Platooning SystemabstractIn the context of vehicle-road-cloud integration, multi-vehicle platooning systems have become an important approach for improving road traffic efficiency, driver comfort, driving safety, energy consumption, and mitigating traffic congestion. However, under high-speed mobility, Vehicle-to-Vehicle (V2V) communication within multi-vehicle platoons is susceptible to delays caused by interference and the inherent uncertainties of wireless communication channels. These delays present considerable challenges to achieving effective multi-vehicle cooperative control. To overcome the limitations of existing research, this paper proposes a joint communication and control optimization strategy for multi-vehicle platooning systems. A novel spacing error metric is introduced, which uses the real-time velocity of each vehicle to improve the platooning system responsiveness. Furthermore, we derive the Signal-to-Interference-plus-Noise Ratio (SINR) threshold to ensure the stability and reliability of the platoon. This ensures safe distances and synchronized speeds among all vehicles, even when communication delays occur. Finally, the proposed joint optimization strategy is validated through performance comparisons, demonstrating its effectiveness and superior performance. Xuelong Yu, Fa Zhu, Xingchi Chen, Kuan Zhang 0001, Chong Yu 0002, Hai Zhao 0002, Athanasios V. Vasilakos |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Secure and Efficient Federated Learning Against Model Poisoning Attacks in Horizontal and Vertical Data PartitioningabstractIn distributed systems, data may partially overlap in sample and feature spaces, that is, horizontal and vertical data partitioning. By combining horizontal and vertical federated learning (FL), hybrid FL emerges as a promising solution to simultaneously deal with data overlapping in both sample and feature spaces. Due to its decentralized nature, hybrid FL is vulnerable to model poisoning attacks, where malicious devices corrupt the global model by sending crafted model updates to the server. Existing work usually analyzes the statistical characteristics of all updates to resist model poisoning attacks. However, training local models in hybrid FL requires additional communication and computation steps, increasing the detection cost. In addition, due to data diversity in hybrid FL, solutions based on the assumption that malicious models are distinct from honest models may incorrectly classify honest ones as malicious, resulting in low accuracy. To this end, we propose a secure and efficient hybrid FL against model poisoning attacks. Specifically, we first identify two attacks to define how attackers manipulate local models in a harmful yet covert way. Then, we analyze the execution time and energy consumption in hybrid FL. Based on the analysis, we formulate an optimization problem to minimize training costs while guaranteeing accuracy considering the effect of attacks. To solve the formulated problem, we transform it into a Markov decision process and model it as a multiagent reinforcement learning (MARL) problem. Then, we propose a malicious device detection (MDD) method based on MARL to select honest devices to participate in training and improve efficiency. In addition, we propose an alternative poisoned model detection (PMD) method considering model change consistency. This method aims to prevent poisoned models from being used in the model aggregation. Experimental results validate that under the random local model poisoning attack, the proposed MDD method can save over 50% training costs while guaranteeing accuracy. When facing the advanced adaptive local model poisoning (ALMP) attack, utilizing both the proposed MDD and PMD methods achieves the desired accuracy while reducing execution time and energy consumption. Chong Yu 0002, Zhenyu Meng, Wenmiao Zhang, Lei Lei 0004, Jianbing Ni, Kuan Zhang 0001, Hai Zhao 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Communication-Efficient Hybrid Federated Learning for E-Health With Horizontal and Vertical Data PartitioningabstractElectronic healthcare (e-health) allows smart devices and medical institutions to collaboratively collect patients' data, which is trained by artificial intelligence (AI) technologies to help doctors make diagnosis. By allowing multiple devices to train models collaboratively, federated learning is a promising solution to address the communication and privacy issues in e-health. However, applying federated learning in e-health faces many challenges. First, medical data are both horizontally and vertically partitioned. Since single horizontal federated learning (HFL) or vertical federated learning (VFL) techniques cannot deal with both types of data partitioning, directly applying them may consume excessive communication cost due to transmitting a part of raw data when requiring high modeling accuracy. Second, a naive combination of HFL and VFL has limitations including low training efficiency, unsound convergence analysis, and lack of parameter tuning strategies. In this article, we provide a thorough study on an effective integration of HFL and VFL, to achieve communication efficiency and overcome the above limitations when data are both horizontally and vertically partitioned. Specifically, we propose a hybrid federated learning framework with one intermediate result exchange and two aggregation phases. Based on this framework, we develop a hybrid stochastic gradient descent (HSGD) algorithm to train models. Then, we theoretically analyze the convergence upper bound of the proposed algorithm. Using the convergence results, we design adaptive strategies to adjust the training parameters and shrink the size of transmitted data. The experimental results validate that the proposed HSGD algorithm can achieve the desired accuracy while reducing communication cost, and they also verify the effectiveness of the adaptive strategies. Chong Yu 0002, Shuaiqi Shen, Shiqiang Wang 0001, Kuan Zhang 0001, Hai Zhao 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | ESR-MHFL: Edge Server Reallocation for Multi-Hierarchical Federated LearningabstractFederated Learning (FL) enables efficient and privacy-preserving Edge Intelligence (EI) in Mobile Edge Computing (MEC). However, implementing FL-enabled EI services faces critical challenges, including data and device heterogeneity, limited network resources, uneven distribution of network infrastructure, etc., which may intensify with increasing system scale. These challenges are particularly acute in multi-provider environments where edge servers are suboptimally allocated across federations, leading to degraded convergence and increased training costs. In this paper, we present a novel Multiple Hierarchical Federated Learning (MHFL) architecture for large-scale FL and design an Edge Server Reallocation scheme (ESR-MHFL) to enhance training efficiency by optimally redistributing edge servers among federations based on their contribution to model convergence. We first develop a closed-form analysis model for MHFL to quantify training time, computation, and communication costs. To improve training efficiency, we analyze the impacts of edge server allocation on convergence and formulate server reallocation as a multi-item auction problem with theoretical guarantees. We then propose ESR-MHFL, which leverages Coalition Structure Generation (CSG) and greedy matching methods to simplify the reallocation problem and enhance efficiency. Extensive numerical simulations demonstrate that ESR-MHFL not only improves model accuracy while reducing training cost but also exhibits strong compatibility with existing client selection methods, achieving improved training efficiency. The total economic expenditure combining all components Tianao Xiang, Yuanguo Bi, Lin Cai 0001, Chong Yu 0002, Mingjian Zhi, Rongfei Zeng, Tom H. Luan |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | Secure Interaction-Based Feature Selection for Vertical Federated LearningabstractFederated learning enables decentralized data own-ers to collaborate and train models in a distributed manner. A special type is Vertical Federated Learning (VFL), where each of the participated data owners only has a portion of the data features. To maintain a high accuracy and reasonable computational cost, selecting a set of features among the entire dataset is essential. Although some existing work selects features by calculating their individual contributions to the learning outcomes, knowing the joint contribution from multiple features becomes necessary but challenging. Meanwhile, security concerns are raised when calculating the joint contribution of a set of features where the feature data are stored by different owners. Using homomorphic encryption or secure computing over en-crypted data is possible, but it may cost too much when complex calculations are involved and repeated. To this end, this paper proposes a privacy-preserving feature selection protocol that considers the interactions between features stored across different data owners. Specifically, we first propose an interaction-based feature selection algorithm for vertically distributed datasets. This algorithm estimates the features' joint contributions to the model training outcomes. Then, we propose a privacy-preservation protocol to prevent the semi-honest cloud server from obtaining or inferring the raw data when aggregating the knowledge and calculating the complex interaction measure for feature selection. We create a new approximation method for interaction measures to address the high computational cost when securely calculating the interaction measure while maintaining the training accuracy. The security discussions show that the proposed protocol preserves data owner's privacy. The extensive simulations validate the achieved training accuracy and efficiency. Zhenyu Meng, Wenmiao Zhang, Shuaiqi Shen, Chong Yu 0002, Kuan Zhang 0001 |
ICC | 4 |
| 2024 | Explore Patterns to Detect Sybil Attack during Federated Learning in Mobile Digital Twin NetworkabstractDigital twins represent users in the cyber world and interact between users and network controllers to better manage the mobile network. Due to communications and other resource constraints, transmitting raw data for a traditional, centralized machine learning in the mobile network has been replaced by federated learning. Federated learning allows participants to train a complex model in a distributed manner, through a group of participants' local training and a global aggregation with model updates as feedback. Although federated learning can save communications costs, address data heterogeneity and protect privacy by stopping the raw data transmission, it faces various se-curity challenges. For example, poisoning attacks may inject false models or modify existing model parameters to bias the gradient descent of federated learning. Some literature attempted to detect poisoning attacks, but the attackers can still strengthen their power by creating many identities to build their group advantage, which overturns the existing detection. In this paper, we propose a digital-twin-based Sybil detection by creating new community detection among participants in federated learning. Specifically, we first identify Sybil attackers on several levels according to their attacking strength and strategies. Then, we integrate digital twins as a side channel to distinguish Sybil identities which in fact belong to the same attacker. This could leak the attacker's correlated behavior patterns which are automatically recorded in digital twins. Under this observation, we build a DT-graph that tightly connects Sybil-controlled identities belonging to the same attacker. We propose a graph-based community detection algorithm to further partition the DT-graph and distinguish Sybil attacks. Extensive simulations validate our proposed method compared with existing work. Wenmiao Zhang, Chong Yu 0002, Zhenyu Meng, Shuaiqi Shen, Kuan Zhang 0001 |
ICC | 2 |
| 2024 | Privacy-Preserving Anomaly Detection of Encrypted Smart Contract for Blockchain-Based Data TradingabstractIn a blockchain-based data trading platform, data users can purchase data sets and computing power through encrypted smart contracts. The security of smart contracts is important as it relates to that of the data platform. However, due to the inability to apply to detection rules with complex structures and the inefficiency of detection, existing malicious code detection methods are not suitable for the encrypted smart contracts in blockchain-based data trading platforms with high transaction rate requirements. In this paper, a practical and privacy-preserving malicious code detection method is proposed for encrypted smart contract in blockchain-based data trading platform. Specifically, we design two kinds of miners to act as the malicious rule processor and the detector respectively for inspecting the encrypted smart contract. The rule processor generates an obfuscated map with the original open-source malicious rule set. The detector performs a malicious inspection algorithm by inputting the obfuscated map and the randomized tokens, where the latter is generated from smart contract. Then, we theoretically analyze the security syntax of the proposed method. The analysis results demonstrate the proposed scheme can achieve$\mathcal {L}$-secure against adaptive attacks. Extensive experiments are carried out through the open-source real rule sets, which show that the proposed scheme can reduce communication time and communication overhead. Dajiang Chen, Zeyu Liao, Rui-dong Chen, Hao Wang 0229, Chong Yu 0002, Kuan Zhang 0001, Ning Zhang 0007, Xuemin Shen |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | LSPSS: Constructing Lightweight and Secure Scheme for Private Data Storage and Sharing in Aerial ComputingabstractAerial computing is gradually playing an essential role in edge and fog computing paradigms by virtue of mobility, availability, scalability, flexibility, and simultaneity, where the Low-altitude Computing (LAC) platform, as the end close to the data sources, is mainly responsible for data collection and storage. However, because of the long physical distance of data transmission and the vulnerability of the transmission link to various attacks, how to efficiently share the stored data while ensuring data privacy is a critical issue for LAC at present. In this paper, we propose a lightweight and secure private data storage and sharing scheme to support range queries over encrypted multi-dimensional data. Specifically, we first propose two data conversion methods for transforming location features and collected log files with multi-dimensional attributes in Unmanned Aerial Vehicles (UAVs). Based on the ideas of asymmetric scalar-product-preserving encryption (ASPE) and inner product comparison (IPC), we design a privacy-preserving storage and sharing technique for the converted data. In addition, to achieve secure and efficient data querying and result verification, we design a secure data index and build a data authentication structure (DAS) with G-tree. Finally, we rigorously analyze the security of our proposed scheme and conduct extensive experiments on a real-world database to prove that our proposed scheme is secure and easy to use in practical application scenarios. Haoyang Wang 0005, Kai Fan 0001, Chong Yu 0002, Kuan Zhang 0001, Fenghua Li 0001, Hui Li 0006, Yintang Yang, Haojin Zhu |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Privacy-Preserving Task Allocation and Decentralized Dispute Protocol in Mobile CrowdsourcingabstractMobile crowdsourcing is an emerging network architecture that can outsource tasks to a group of people or devices. Most recent literature studied the location privacy of users during task allocation in mobile crowdsourcing. But in some cases, a dispute may occur when users have an argument about the payment of the task. To adjudicate the dispute, users need to provide their private information such as credentials to the third party. Furthermore, finding a trusted third party for mobile crowdsourcing is hard in real life, raising the opportunity for decentralized disputes. However, existing work cannot avoid information leakage during decentralized dispute adjudication, and these schemes can also lower the efficiency of arbitrating. In this paper, we propose a privacy-preserving and decentralized dispute arbitration protocol, which allows the dispute adjudication to proceed without a trusted third party but prevents users' private information from disclosing. Specifically, we first propose a secure task allocation protocol for mobile crowdsourcing, which preserves users' location privacy while enabling efficient task release and allocation. Then, in order to avoid employing a trusted third party, we design a privacy-preserving and decentralized dispute arbitration protocol, which does not reveal any private information during the dispute adjudication. Security and privacy discussions show our protocol can resist forgery attacks but preserve privacy during the decentralized dispute adjudication. In addition, performance evaluation validates the efficiency of our protocol. Zhenyu Meng, Chong Yu 0002, Yi Qian 0001 |
ICC | 2 |
| 2022 | Collaborative Edge Caching with Personalized Modeling of Content Popularity over Indoor Mobile Social NetworksabstractMobile social networks allow users to acquire multimedia contents to their mobile devices via wireless communications. To alleviate the network traffic and latency for transmitting contents, user preferences can be predicted and popular contents can be cached at the edge of network. However, for edge caching over the indoor mobile social networks raises challenging issues. The user preference for mobile data is location-dependent in different areas of indoor environment, such that various edge nodes need to maintain its distinctive prediction model instead of using the universal one. The limited computing power for edge nodes over indoor mobile social networks also hinders effective model training on the edge. In this paper, we propose a collaborative edge caching framework that enables personalized modeling for content popularity prediction. Specifically, a non-additive measure based feature selection scheme is proposed to realize efficient yet accurate modeling on resource-constrained edge nodes. A collaborative learning algorithm is designed to reduce the computational overheads over mobile social networks by extracting global knowledge on simplifying model training through feature selection. Extensive simulation validates the effectiveness and efficiency of our proposed framework. Shuaiqi Shen, Chong Yu 0002, Kuan Zhang 0001, Song Ci |
ICC | 2 |
| 2022 | Efficient Multi-Layer Stochastic Gradient Descent Algorithm for Federated Learning in E-healthabstractE-health systems consist of intelligent devices, medical institutions, edge nodes, and cloud servers to improve healthcare service quality and efficiency. In e-health systems, patients’ data are cooperatively collected by their wearable devices and the hospital they have visited, i.e., vertically distributed data. The data on wearable devices share the same feature set but are different in sample spaces, i.e., horizontally partitioned data. Meanwhile, hospitals target various user groups resulting in high data diversity, i.e., non-identically distributed data. These three characteristics cause that existing federated learning frameworks cannot efficiently train models on medical data. Furthermore, model training in e-health is time-sensitive because some diseases mutate very quickly and spread easily, which requires fast convergence of machine learning algorithms. In this paper, we address the problem of how to efficiently and rapidly train global models on e-health data. Specifically, we propose a multilayer federated learning framework to cope with data that are vertically, horizontally, and non-identically distributed. Moreover, we develop a Multi-Layer Stochastic Gradient Descent (MLSGD) algorithm towards the proposed framework to learn the optimal global model. To improve training efficiency, partial models learned by devices are aggregated on edge nodes before exchanging intermediate results with hospitals. The weight of local models is proportional to local data size when performing global aggregation to balance the impact of local models on the global model. We also prove the convergence of the MLSGD algorithm from a theoretical perspective. The experimental results from the real-world dataset MIMIC-III validate that the proposed algorithm converges fast and achieves desired accuracy. Chong Yu 0002, Shuaiqi Shen, Shiqiang Wang 0001, Kuan Zhang 0001, Hai Zhao 0002 |
ICC | 1 |
| 2022 | Energy-Aware Device Scheduling for Joint Federated Learning in Edge-assisted Internet of Agriculture ThingsabstractEdge-assisted Internet of Agriculture Things (Edge-IoAT) connects massive smart devices managed by edge nodes to collect crop data for distributed computing, such as federated learning, to guide agricultural production. In Edge-IoAT, data are cooperatively collected by edge nodes and the server, i.e., vertically partitioned. In addition, sample size and distribution are different for edge nodes, i.e., horizontally partitioned. Existing federated learning frameworks are not applicable for Edge-IoAT because they do not consider both types of data partitioning simultaneously. Moreover, the excessive energy consumption may cause premature interruption of model training, and spectrum scarcity prevents a portion of edge nodes from communicating with the server. Given limited energy and communication resources, training accuracy relies on how to schedule devices. In this paper, we first propose a joint federated learning framework for Edge-IoAT to cope with both vertically and horizontally partitioned data. After that, we formulate an energy-aware device scheduling problem to assign communication resources to the optimal edge node subset for minimizing the global loss function. Then, we develop a greedy algorithm to find the optimal solution. Experiments in a Nebraska farm show that the proposed framework with energy-aware device scheduling achieves a fast convergence rate, low communication cost, and high modeling accuracy under resource constraints. Chong Yu 0002, Shuaiqi Shen, Kuan Zhang 0001, Hai Zhao 0002, Yeyin Shi |
WCNC | 1 |
| 2022 | Leveraging Energy, Latency, and Robustness for Routing Path Selection in Internet of Battlefield ThingsabstractInternet of Battlefield Things (IoBT) connects massive tactical devices to collect battlefield situations and share perceived information. The IoBT can enhance the intelligent battlefield command, collaborative attack, and other applications, such as landmine trigger and post-war clearance. Existing routing path selection methods designed for wireless sensor networks (WSNs) are effective but still face challenges in IoBT scenarios. First, tactical devices follow nonuniform distributions with high density on boundaries in IoBT to prevent the location of devices from being speculated and protect strategic positions, which results in unbalanced energy consumption. Second, increasing latency in IoBT is caused by various data generation probabilities of tactical devices. Third, the military task features, such as landmine explosion, disconnection, and failure of tactical devices, may put forward special requirements on network robustness. To this end, we propose a routing path selection method with joint optimization in IoBT based on nonuniform node distributions and location-related data generation probabilities. Specifically, we first investigate and formulate the distribution and data generation probability of tactical devices. Based on the special features, energy consumption, latency, and network robustness are analyzed during multihop communications in IoBT. Then, a joint optimization problem is formulated to minimize energy consumption and latency, while maximizing the network robustness simultaneously. Furthermore, two path assignment algorithms are developed to solve this optimization problem. Finally, our simulation results show that the proposed routing path selection method can reduce energy consumption and latency with the guaranteed robustness of IoBT. Chong Yu 0002, Shuaiqi Shen, Haojun Yang, Kuan Zhang 0001, Hai Zhao 0002 |
IEEE Internet Things J. | 1 |
| 2021 | Exploiting Ensemble Learning for Edge-assisted Anomaly Detection Scheme in e-healthcare SystemabstractWith the thriving of wearable devices and the widespread use of smartphones, the e-healthcare system emerges to cope with the high demand of health services. However, this integrated smart health system is vulnerable to various attacks, including intrusion attacks. Traditional detection schemes generally lack the classifier diversity to identify attacks in complex scenarios that contain a small amount of training data. Moreover, the use of cloud-based attack detection may result in higher detection latency. In this paper, we propose an Edge-assisted Anomaly Detection (EAD) scheme to detect malicious attacks. Specifically, we first identify four types of attackers according to their attacking capabilities. To distinguish attacks from normal behaviors, we then propose a wrapper feature selection method. This selection method eliminates the impact of irrelevant and redundant features so that the detection accuracy can be improved. Moreover, we investigate the diversity of classifiers and exploit ensemble learning to improve the detection rate. To reduce high detection latency in the cloud, edge nodes are used to concurrently implement the proposed lightweight scheme. We evaluate the EAD performance based on two real-world datasets, i.e., NSL-KDD and UNSW-NB15 datasets. The simulation results show that the EAD outperforms other state-of-the-art methods in terms of accuracy, detection rate, and computational complexity. The analysis of detection time validates the fast detection of the proposed EAD compared with cloud-assisted schemes. Wei Yao 0016, Kuan Zhang 0001, Chong Yu 0002, Hai Zhao 0002 |
GLOBECOM | 3 |
| 2021 | Exploiting Feature Interactions for Malicious Website Detection with Overhead-accuracy TradeoffabstractMalicious websites attempt to install malware on user’s devices without permission, which can disrupt device operation, steal personal information, and even acquire access to the device for future attacks. Accurate detection of malicious website behaviors is crucial for network security but still faces challenges. Firstly, various types and semantics of website features are required to identity the wide range of malicious characteristics, leading to massive training data and computational overhead. Secondly, to reduce model dimensionality, a proper selection of website features is essential but difficult due to the complex relations among features that can affect each other’s contribution to detection outcomes. In this paper, we propose a lightweight feature-based detection scheme against malicious websites considering the interaction measures among features and the overhead-accuracy tradeoff. Specifically, we systematically characterize the interactions among website features in a non-additive manner to indicate the aggregated impacts of feature subsets. Then we propose a quantification method to measure the feature interactions based on multivariate regression. With this method, important features are selected to substantially reduce the model dimension and computational complexity while maintaining desirable accuracy. Meanwhile, the proposed scheme provides an interpretable model that preserves the physical meanings of original features. It allows users to balance the overhead-accuracy tradeoff for detection model training through feature subset selection to fit the requirements and constraints of real applications. Shuaiqi Shen, Chong Yu 0002, Kuan Zhang 0001, Song Ci |
ICC | 2 |
| 2021 | Communication-Efficient Federated Learning for Connected Vehicles with Constrained ResourcesabstractWith the upcoming next generation wireless network, vehicles are expected to be empowered by artificial intelligence (AI). By connecting vehicles and cloud server via wireless communication, federated learning (FL) allows vehicles to collaboratively train deep learning models to support intelligent services, such as autonomous driving. However, the large number of vehicles and increasing size of model parameters bring challenges to FL-empowered connected vehicles. Since communication bandwidth is insufficient to upload full-precision local models from numerous vehicles, model compression is usually conducted to reduce transmitted data size. Nevertheless, conventional model compression methods may not be practical for resource-constrained vehicles due to the increasing computational overhead for FL training. The overhead for downloading global model can also be omitted by existing methods since they are originally designed for centralized learning instead of FL. In this paper, we propose a ternary quantization based model compression method on communication-efficient FL for resource-constrained connected vehicles. Specifically, we firstly propose a ternary quantization based local model training algorithm that optimizes quantization factors and parameters simultaneously. Then, we design a communication-efficient FL approach that reduces overhead for both upstream and downstream communications. Finally, simulation results validate that the proposed method demands the lowest communication and computational overheads for FL training, while maintaining desired model accuracy compared to existing model compression methods. Shuaiqi Shen, Chong Yu 0002, Kuan Zhang 0001, Song Ci |
IWCMC | 2 |
| 2021 | Adaptive Artificial Intelligence for Resource-Constrained Connected Vehicles in Cybertwin-Driven 6G NetworkabstractThe emerging technology of cybertwin is expected to bring revolutionary benefits to the sixth-generation (6G) network in respect of communication, resources allocation, and digital asset management. Empowered by ubiquitous artificial intelligence (AI), cybertwin is capable of adjusting the requests for computing resources to support network services by analyzing user’s demands for quality of experience and resource scarcity in the market. For resource-constrained applications, such as connected vehicles in the 6G network, cybertwin can intelligently determine the time-varying requests of computing resources for various vehicles at different times. However, the current service architecture executes AI algorithms with universal configurations for all vehicles. This causes the difficulty of customizing the complexity of AI algorithms to maintain adaptive to cybertwin’s decisions on dynamic resources allocation. In this article, we propose an adaptive AI framework based on efficient feature selection to cooperate with cybertwin’s resource allocation. This proposed framework can adaptively customizing AI model complexity with available computing resources. Specifically, we systematically characterize the aggregated impacts of all feature combinations on the modeling outcomes of AI algorithms. By utilizing nonadditive measures, the interactions among features can be quantified to indicate their contributions to the modeling process. Then, we propose an efficient algorithm to obtain accurate interaction measures for adaptive feature selection to balance the tradeoff between modeling accuracy and computational overhead. Finally, extensive simulations are conducted to validate that our proposed framework substantially reduces the overhead of AI algorithms while guaranteeing desired modeling accuracy for cybertwin-driven connected vehicles in 6G. Shuaiqi Shen, Chong Yu 0002, Kuan Zhang 0001, Song Ci |
IEEE Internet Things J. | 2 |