Wenchao Xu 0001

dblp:96/8862-1 · DBLP profile ↗
← Back
121ranked-venue papers
12as first author
99since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 63 · 9 first-author · 47 since 2021Artificial intelligence and machine learning · 31 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Backdooring Rationalization
abstract
Rationalization model has recently garnered significant attention for enhancing the interpretability of natural language processing by first using a generator to select the most relevant pieces from the text with respect to the label, before passing the text input to the predictor. However, the robustness of the rationalization models is not sufficiently investigated. Specifically, this paper explores the robustness of rationalization models against backdoor attacks, which has been ignored by previous studies. Surprisingly, we find that conventional backdoor attack techniques fail to inject triggers into the rationalization model because its generator can filter out bad triggers. Considering this, we further propose a novel backdoor attack method named as BadRNL designed specially for the rationalization models. The core idea of BadRNL is first to search for the personalized trigger for each specific dataset and then manipulate the rationales and labels to conduct attacks. Besides, BadRNL controls the order of sample learning through poison-priority sampling strategies. Experimental results show that our method can successfully craft the predictions of samples containing triggers while maintaining the performance of the model on clean data.
Lingxiao Kong, Wenchao Xu 0001
AAAI3
2026 Causality-inspired Federated Learning for Dynamic Spatio-Temporal Graphs
abstract
Federated Graph Learning (FGL) has emerged as a powerful paradigm for decentralized training of graph neural networks while preserving data privacy. However, existing FGL methods are predominantly designed for static graphs and rely on parameter averaging or distribution alignment, which implicitly assume that all features are equally transferable across clients, overlooking both the spatial and temporal heterogeneity and the presence of client-specific knowledge in real-world graphs. In this work, we identify that such assumptions create a vicious cycle of spurious representation entanglement, client-specific interference, and negative transfer, degrading generalization performance in Federated Learning over Dynamic Spatio-Temporal Graphs (FSTG). To address this issue, we propose a novel causality-inspired framework named SC-FSGL, which explicitly decouples transferable causal knowledge from client-specific noise through representation-level interventions. Specifically, we introduce a Conditional Separation Module that simulates soft interventions through client conditioned masks, enabling the disentanglement of invariant spatio-temporal causal factors from spurious signals and mitigating representation entanglement caused by client heterogeneity. In addition, we propose a Causal Codebook that clusters causal prototypes and aligns local representations via contrastive learning, promoting cross-client consistency and facilitating knowledge sharing across diverse spatio-temporal patterns. Experiments on five diverse heterogeneity Spatio-Temporal Graph (STG) datasets show that SC-FSGL outperforms state-of-the-art methods.
Yuxuan Liu 0017, Wenchao Xu 0001, Haozhao Wang, Zhiming He, Zhaofeng Shi, Chongyang Xu, Peichao Wang
AAAI2
2026 Distilling Cross-Modal Knowledge via Feature Disentanglement
abstract
Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and state-of-the-art cross-modal KD approaches.
Wenchao Xu 0001, Renyu Yang
AAAI4
2026 HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
abstract
The deployment of large language models' (LLMs) inference at the edge can facilitate prompt service responsiveness while protecting user privacy. However, it is critically challenged by the resource constraints of a single edge node. Distributed inference has emerged to aggregate and leverage computational resources across multiple devices. Yet, existing methods typically require strict synchronization, which is often infeasible due to the unreliable network conditions. In this paper, we propose HALO , a novel framework that can boost the distributed LLM inference in lossy edge network. The core idea is to enable a relaxed yet effective synchronization by strategically allocating less critical neuron groups to unstable devices, thus avoiding the excessive waiting time incurred by delayed packets. HALO introduces three key mechanisms: (1) a semantic-aware predictor to assess the significance of neuron groups prior to activation. (2) a parallel execution scheme of neuron group loading during the model inference. (3) a load-balancing scheduler that efficiently orchestrates multiple devices with heterogeneous resources. Experimental results from a Raspberry Pi cluster demonstrate that HALO achieves a 3.41x end-to-end speedup for LLaMA-series LLMs under unreliable network conditions. It maintains performance comparable to optimal conditions and significantly outperforms the state-of-the-art in various scenarios.
Peirong Zheng, Wenchao Xu 0001, Haozhao Wang, Xuemin Shen
INFOCOM2
2026 Black-Box RF Fingerprint Spoofing via Surrogate-Guided Generative Perturbations
abstract
We study the feasibility of black-box radio frequency fingerprint (RFF) spoofing, where an adversary lacks access to the target receiver’s data or model. We present a surrogate-guided generative perturbation framework that jointly trains a generator with feedback from multiple surrogate receivers to synthesize low-power, fingerprint-level perturbations. Using real RF fingerprint datasets, we evaluate the spoofing effectiveness of forged perturbations on unseen receivers. Our results show that, even without any feedback from the target, an attacker can successfully impersonate a selected transmitter at an unseen receiver. However, the success remains limited and varies across devices, reflecting both the feasibility and the boundary of black-box RF fingerprint spoofing. These findings provide the first empirical evidence that multi-surrogate training can partially narrow the black-box gap in RF fingerprint spoofing.
Zhaoyi Lu 0001, Wenchao Xu 0001, Cunqing Hua
WISEC2
2026 pFedDKS: Detached Knowledge Sharing for Personalized Federated Learning
abstract
By allowing each client to refer to the knowledge from other clients while retaining their specific characteristics, partial knowledge sharing has become one of the main approaches to realizing personalized federated learning (pFL). Representative techniques of partial knowledge sharing propose sharing the feature extractor while customizing the classifier head of the neural network. Although such methods achieve great success, the underlying principle behind them remains yet to be comprehensively understood. A fundamental problem is whether it is really appropriate to fully share the feature extractor. Based on the theory of neural collapse, in this paper, we demonstrate both theoretically and empirically that the feature extractor should be partially shared rather than fully shared. More specifically, we identify a substantial inconsistency between the fused global feature representations and expected local feature representations, and thus it is necessary to preserve partially customized layers of the feature extractor for enhancing personalized representations. Based on this discovery, we further propose a novel method called pFedDKS which detaches the shared global knowledge and customized local knowledge by providing detached feature prototypes. Extensive experiments on various datasets and models show that pFedDKS outperforms state-of-the-arts.
Haozhao Wang, Wenchao Xu 0001, Jingzhi Wang, Yunfeng Fan, Xiaoquan Yi, Rui Zhang 0003
WWW2
2026 Self-Speculative Decoding for On-device MoE Acceleration
abstract
The sparse mixture-of-experts (MoE) architecture is a promising backbone of foundation models for a wide range of applications in edge. However, deploying them locally presents a significant challenge to memory-constrained GPUs. Previous techniques utilize CPUs for expert offloading, which suffer from inaccurate expert prefetching and on-demand loading latency. To address these challenges, we propose self-speculative MoE (SS-MoE), an algorithm-system co-design framework that facilitates inference under limited GPU memory. Our insight is that only a subset of routed experts, i.e., draft model, can still tackle easy tasks and generate draft tokens. Second, we deem GPU memory as the experts cache, and on-demand update it to mitigate IO overhead. Draft tokens from fewer routed experts are generated quickly, and these experts are then routed for verification. Additionally, we design a confidence-based policy to adaptively accept or verify draft tokens, which selectively decreases or increases the number of verification tokens of speculative decoding and achieves acceleration. Notably, under conservative verification, our approach preserves model accuracy and surpasses the decoding speed of the 4-bit quantized counterpart model. Under adaptive verification, our method significantly enhances decoding speed by 3.72x over state-of-the-art methods while maintaining nearly lossless accuracy.
Peirong Zheng, Wenchao Xu 0001, Haozhao Wang
WWW2
2026 FedCHG: Graph autoencoder enhanced federated learning for cross-Domain heterogeneous graph
Jiyuan He, Yichen Li 0006, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Hongwei Lu, Ruixuan Li 0001
Expert Syst. Appl.4
2026 FedGSE: Gradient-Based Submodel Extraction for Resource-Constrained Federated Learning
abstract
Federated Learning (FL) has emerged as a pivotal paradigm for multi-client collaborative learning, primarily due to its inherent capability to safeguard privacy. Nonetheless, the heterogeneity among FL clients, characterized by their disparate resource capabilities and not independent and identical (NonIID) local datasets, presents a significant challenge. Specifically, low-resource clients, such as edge devices, grapple with the inadequacy to accommodate the entire model parameter set for training purposes. To mitigate this issue, preceding research has ventured into devising methodologies that entail extracting sub-models from the overarching global model, tailored to the specific communication, computational, and memory constraints of individual clients. Despite these advancements, prevailing sub-model extraction techniques, which predominantly hinge on pre-established rules, overlook a crucial factor: the impact of NonIID local data on the trajectory of neuron update dynamics. This oversight can amplify discrepancies between the practical local updates inferred by the sub-model and those anticipated via the entire model, thereby undermining overall performance. In this paper, we introduceFedGSE, an innovative Gradient-based Neuron Selection methodology designed explicitly for FL environments. This methodology aims to curate sub-models that significantly reduce discrepancies in local updates, enhancing alignment with the global model's learning trajectory. Central to theFedGSEapproach is a sophisticated algorithm that handpicks critical neurons for sub-model construction. These neurons are identified through their pronounced gradient magnitudes, resulting from the training of the global model on a dataset mirroring the client's data distribution. Consequently, the sub-model's induced local gradient updates closely emulate those derived from directly training the client's data on the full global model, fostering enhanced alignment and performance. Extensive experiments over diverse datasets and tasks demonstrate the superiority ofFedGSEover existing baselines.
Genlang Chen, Yabo Jia, Haozhao Wang, Chaoyi Pang, Wenchao Xu 0001
IEEE Trans. Mob. Comput.5
2026 A Channel-Triggered Backdoor Attack on Wireless Semantic Image Reconstruction
abstract
This paper investigates backdoor attacks in image oriented semantic communications. The threat of backdoor at tacks on symbol reconstruction in semantic communication (Sem Com) systems has received limited attention. Existing research on backdoor attacks targeting SemCom symbol reconstruction primarily focuses on input-level triggers, which are impractical in scenarios with strict input constraints. In this paper, we propose a novel channel-triggered backdoor attack (CT-BA) framework that exploits inherent wireless channel characteristics as activation triggers. Our key innovation involves utilizing fundamental channel statistics parameters, specifically channel gain with different fading distributions or channel noise with different power, as potential triggers. This approach enhances stealth by eliminating explicit input manipulation, provides flexibility through trigger selection from diverse channel conditions, and enables automatic activation via natural channel variations without adversary intervention. We extensively evaluate CT-BA across four joint source-channel coding (JSCC) communication system architectures and three benchmark datasets. Simulation results demonstrate that our attack achieves near-perfect attack success rate (ASR) while maintaining effective stealth. Finally, we discuss potential defense mechanisms against such attacks.
Jialin Wan, Jinglong Shen, Nan Cheng 0001, Zhisheng Yin, Yiliang Liu, Wenchao Xu 0001, Xuemin Shen
IEEE Trans. Mob. Comput.6
2026 Graph Neural Network-Based Multicast Routing for On-Demand Streaming Services in 6G Networks
abstract
The increase of bandwidth-intensive applications in sixth-generation (6 G) wireless networks, such as real-time volumetric streaming, and multi-sensory extended reality, demands intelligent multicast routing solutions capable of delivering differentiated quality-of-service (QoS) at scale. Traditional shortest-path and multicast routing algorithms are either computationally prohibitive or structurally rigid, and they often fail to support heterogeneous user demands, leading to suboptimal resource utilization. Neural network-based approaches, while offering improved inference speed, typically lack topological generalization and scalability. To address these limitations, this paper presents a graph neural network (GNN)-based multicast routing framework that jointly minimizes total transmission cost and supports user-specific video quality requirements. The routing problem is formulated as a constrained minimum-flow optimization task, and a reinforcement learning algorithm is developed to sequentially construct efficient multicast trees by reusing paths and adapting to network dynamics. A graph attention network (GAT) is employed as the encoder to extract context-aware node embeddings, while a long short-term memory (LSTM) module models the sequential dependencies in routing decisions. Extensive simulations demonstrate that the proposed method closely approximates optimal dynamic programming-based solutions while significantly reducing computational complexity. The results also confirm strong generalization to large-scale and dynamic network topologies, highlighting the method's potential for real-time deployment in 6 G multimedia delivery scenarios. Code is available athttps://github.com/UNIC-Lab/GNN-Routing.
Xiucheng Wang, Zien Wang, Nan Cheng 0001, Wenchao Xu 0001, Wei Quan 0001, Xuemin Shen
IEEE Trans. Mob. Comput.4
2026 Spatial-Temporal Attention Model for Traffic State Estimation With Sparse Internet of Vehicles Data
abstract
The rapid growth of connected vehicles creates new opportunities to exploit internet of vehicles (IoV) data for traffic state estimation (TSE), which is a key enabler of intelligent transportation systems (ITS). In this paper, we propose a cost-effective TSE framework that leverages sparse IoV data, which significantly reducing the data collection overhead associated with large-scale IoV datasets. We further analyze the impact of data sparsification and show that the induced estimation errors can be well approximated by Gaussian noise, thereby reformulating sparse IoV-based TSE as a denoising problem. To enhance estimation accuracy, we develop a spatial-temporal attention model, termed the convolutional retentive network (CRNet), which integrates convolutional neural networks (CNNs) for spatial correlation learning with a retentive network (RetNet) for temporal dependency modeling. Extensive experiments conducted on a large-scale real-world IoV dataset validate the feasibility of TSE under sparse IoV sensing conditions. Notably, even when only 5% of the data is available, CRNet achieves a mean absolute error (MAE) below 5 km/h, demonstrating both the high accuracy of the proposed approach and its practical applicability in real-world scenarios.
Jianzhe Xue, Dongcheng Yuan, Yu Sun 0032, Wenchao Xu 0001, Xuemin Shen
IEEE Trans. Mob. Comput.5
2026 ERA: A QoE-Aware Collaborative Inference Algorithm for NOMA-Based Edge Intelligence
abstract
Although AI has been extensively adopted and has profoundly transformed our lives, it is not feasible to directly deploy large AI models on edge devices with limited resources. To enhance the performance of Edge Intelligence (EI), model split inference has been proposed. In this approach, an AI model is segmented into sub-models, with the most resource-intensive parts offloaded wirelessly to the edge server. This reduces the resource demands and inference latency on the device. However, previous studies have primarily focused on enhancing and optimizing system Quality of Service (QoS), often overlooking Quality of Experience (QoE), which is another crucial aspect for users. Even though QoE has been extensively studied in Edge Computing (EC), the distinct differences between task offloading in EC and split inference in EI, along with specific QoE issues that remain unaddressed in both fields, render these algorithms ineffective for edge split inference scenarios. Therefore, this paper introduces an effective resource allocation algorithm, dubbed ERA, which aims to: 1) expedite split inference in EI, and 2) balance inference delay, QoE, and resource consumption. ERA incorporates resource consumption, QoE, and inference latency to determine the most optimal model split and resource allocation strategies. Given that it is impossible to simultaneously minimize inference delay and resource consumption while maximizing QoE, we employ a gradient descent-based algorithm to find the best possible compromise. Furthermore, to address the complexity arising from parameter discretization in the gradient descent algorithm, we have developed a pipeline gradient descent approach, known as PipGD. We have also examined the properties of the proposed algorithms, including their convergence, complexity, and approximation error. The experimental results clearly show that ERA outperforms previous studies significantly in terms of performance.
Xin Yuan 0003, Ning Li 0003, Quan Chen 0003, Wenchao Xu 0001, Song Guo 0001
IEEE Trans. Mob. Comput.4
2025 Boosting Fine-Grained Visual Anomaly Detection with Coarse-Knowledge-Aware Adversarial Learning
abstract
Many unsupervised visual anomaly detection methods train an auto-encoder to reconstruct normal samples and then leverage the reconstruction error map to detect and localize the anomalies. However, due to the powerful modeling and generalization ability of neural networks, some anomalies can also be well reconstructed, resulting in unsatisfactory detection and localization accuracy. In this paper, a small coarsely-labeled anomaly dataset is first collected. Then, a coarse-knowledge-aware adversarial learning method is developed to align the distribution of reconstructed features with that of normal features. The alignment can effectively suppress the auto-encoder's reconstruction ability on anomalies and thus improve the detection accuracy. Considering that anomalies often only occupy very small areas in anomalous images, a patch-level adversarial learning strategy is further developed. Although no patch-level anomalous information is available, we rigorously prove that by simply viewing any patch features from anomalous images as anomalies, the proposed knowledge-aware method can also align the distribution of reconstructed patch features with the normal ones. Experimental results on four medical datasets and two industrial datasets demonstrate the effectiveness of our method in improving the detection and localization performance.
Qingqing Fang, Qinliang Su, Wenxi Lv, Wenchao Xu 0001, Jianxing Yu
AAAI4
2025 Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
abstract
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). To alleviate this issue, some methods, known as contrastive decoding, induce hallucinations by manually disturbing the raw vision or instruction inputs and then mitigate them by contrasting the outputs of the original and disturbed LVLMs. However, these holistic input disturbances sometimes induce potential noise and also double the inference cost. To tackle these issues, we propose a simple yet effective method named $\textit{Self-Introspective Decoding}$ (SID). Our empirical investigations reveal that pre-trained LVLMs can introspectively assess the importance of vision tokens based on preceding vision and text (both instruction and generated) tokens. Leveraging this insight, we develop the Context and Text-aware Token Selection (CT$^2$S) strategy, which preserves only the least important vision tokens after the early decoder layers, thereby adaptively amplify vision-and-text association hallucinations during auto-regressive decoding. This strategy ensures that multimodal knowledge absorbed in the early decoder layers induces multimodal contextual rather than aimless hallucinations, and significantly reduces computation burdens. Subsequently, the original token logits subtract the amplified fine-grained hallucinations, effectively alleviating hallucinations without compromising the LVLMs' general ability. Extensive experiments illustrate SID generates less-hallucination and higher-quality texts across various metrics, without much additional computation cost.
Fushuo Huo, Wenchao Xu 0001, Zhong Zhang 0014, Haozhao Wang, Peilin Zhao
ICLR2
2025 One-for-All Few-Shot Anomaly Detection via Instance-Induced Prompt Learning
abstract
Anomaly detection methods under the 'one-for-all' paradigm aim to develop a unified model capable of detecting anomalies across multiple classes. However, these approaches typically require a large number of normal samples for model training, which may not always be feasible in practice. Few-shot anomaly detection methods can address scenarios with limited data but often require a tailored model for each class, struggling within the 'one-for-one' paradigm. In this paper, we first proposed the one-for-all few-shot anomaly detection method with the assistance of vision-language model. Different from previous CLIP-based methods learning fix prompts for each class, our method learn a class-shared prompt generator to adaptively generate suitable prompt for each instance. The prompt generator is trained by aligning the prompts with the visual space and utilizing guidance from general textual descriptions of normality and abnormality. Furthermore, we address the mismatch problem of the memory bank within one-for-all paradigm. Extensive experimental results on MVTec and VisA demonstrate the superiority of our method in few-shot anomaly detection task under the one-for-all paradigm.
Wenxi Lv, Qinliang Su, Wenchao Xu 0001
ICLR3
2025 FedRF: Input-side Client Drift Mitigation for Federated Learning via Reusing Features
abstract
A major challenge in Federated Learning (FL) is that the data is heterogeneous in terms of features among clients, which significantly degrades the performance of the global model due to its incurred client drift. To solve this problem, some works propose leveraging the global knowledge, e.g., class prototypes, to regularize the output of the local model to mitigate the local model drift. Despite their effectiveness, the input data of the local model are still heterogeneous, causing the parameters to not fully align. Therefore, to fill this gap, this paper proposes a novel method named FedRF which aligns local models on the input side by reusing the intermediate features to approach the global ones. Specifically, we use the intermediate features output by the global model on local data as input to assist in training the local model. Experimental results demonstrate that the proposed method significantly improves the performance of the global model as compared to baselines.
Lingxiao Kong, Wenchao Xu 0001, Haozhao Wang, Ruixuan Li 0001
ICME3
2025 BSemiFL: Semi-supervised Federated Learning via a Bayesian Approach
abstract
Semi-supervised Federated Learning (SSFL) is a promising approach that allows clients to collaboratively train a global model in the absence of their local data labels. The key step of SSFL is the re-labeling where each client adopts two types of available models, namely global and local models, to re-label the local data. While various technologies such as using the global model or the average of two models have been proposed to conduct the re-labeling step, little literature delves deeply into the performance dominance and limitations of the two models. In this paper, we first theoretically and empirically demonstrate that the local model achieves higher re-labeling accuracy over local data while the global model can progressively improve the re-labeling performance by introducing the extra data knowledge of other clients. Based on these findings, we propose BSemiFL which re-labels the local data through the collaboration between the local and global model in a Bayesian approach. Specifically, to re-label any given local sample, BSemiFL first uses Bayesian inference to assess the closeness of the local/global model to the sample. Then, it applies a weighted combination of their pseudo labels, using the closeness as the weights. Theoretical analysis shows that the labeling error of our method is smaller than that of simply using the global model, the local model, or their simple average. Experimental results show that BSemiFL improves the performance by up to $9.8\%$ as compared to state-of-the-art methods.
Haozhao Wang, Shengyu Wang, Hao Ren 0001, Xingshuo Han, Wenchao Xu 0001, Shangwei Guo, Tianwei Zhang 0004, Ruixuan Li 0001
ICML6
2025 Poster: Coda: Context-aware Acceleration for Distributed LLM Decoding in Edge
abstract
The inference of large language models (LLMs) on distributed edge devices is crucial for privacy-preserving applications. However, its performance is severely degraded in a practically lossy edge network due to frequent synchronization. In this paper, we propose Coda, a novel context-aware distributed acceleration framework tailored for packet loss scenarios. We observe the consistency of sparse patterns of neuron group activations in the same context. By identifying sparsity patterns during the prefill phase and updating neuron groups mapping strategically, we enable a relaxed synchronization for decoding, without slowing down, while preserving accuracy. Our approach offers up to 5.62x speedup, significantly outperforming the state-of-the-art in various scenarios.
Peirong Zheng, Wenchao Xu 0001, Haozhao Wang
MobiCom2
2025 Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
abstract
Model-Heterogeneous Federated Learning (Hetero-FL) has attracted growing attention for its ability to aggregate knowledge from heterogeneous models while keeping private data locally. To better aggregate knowledge from clients, ensemble distillation, as a widely used and effective technique, is often employed after global aggregation to enhance the performance of the global model. However, simply combining Hetero-FL and ensemble distillation does not always yield promising results and can make the training process unstable. The reason is that existing methods primarily focus on logit distillation, which, while being model-agnostic with softmax predictions, fails to compensate for the knowledge bias arising from heterogeneous models. To tackle this challenge, we propose a stable and efficient Feature Distillation for model-heterogeneous Federated learning, dubbed FedFD, that can incorporate aligned feature information via orthogonal projection to integrate knowledge from heterogeneous models better. Specifically, a new feature-based ensemble federated knowledge distillation paradigm is proposed. The global model on the server needs to maintain a projection layer for each client-side model architecture to align the features separately. Orthogonal techniques are employed to re-parameterize the projection layer to mitigate knowledge bias from heterogeneous models and thus maximize the distilled knowledge. Extensive experiments show that FedFD achieves superior performance compared to state-of-the-art methods.
Yichen Li 0006, Xiuying Wang 0015, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Jiahua Dong 0001, Ruixuan Li 0001
NeurIPS3
2025 LLM at Network Edge: A Layer-wise Efficient Federated Fine-tuning Approach
abstract
Fine-tuning large language models (LLMs) poses significant computational burdens, especially in federated learning (FL) settings. We introduce Layer-wise Efficient Federated Fine-tuning (LEFF), a novel method designed to enhance the efficiency of FL fine-tuning while preserving model performance and minimizing client-side computational overhead. LEFF strategically selects layers for fine-tuning based on client computational capacity, thereby mitigating the straggler effect prevalent in heterogeneous environments. Furthermore, LEFF incorporates an importance-driven layer sampling mechanism, prioritizing layers with greater influence on model performance. Theoretical analysis demonstrates that LEFF achieves a convergence rate of $\mathcal{O}(1/\sqrt{T})$. Extensive experiments on diverse datasets demonstrate that LEFF attains superior computational efficiency and model performance compared to existing federated fine-tuning methods, particularly under heterogeneous conditions.
Jinglong Shen, Nan Cheng 0001, Wenchao Xu 0001, Haozhao Wang, Yifan guo
NeurIPS3
2025 ChatbotID: Identifying Chatbots with Granger Causality Test
abstract
With the increasing sophistication of Large Language Models (LLMs), it is crucial to develop reliable methods to accurately identify whether an interlocutor in real-time dialogue is human or chatbot. However, existing detection methods are primarily designed for analyzing full documents, not the unique dynamics and characteristics of dialogue. These approaches frequently overlook the nuances of interaction that are essential in conversational contexts. This work identifies two key patterns in dialogues: (1) Human-Human (H-H) interactions exhibit significant bidirectional sentiment influence, while (2) Human-Chatbot (H-C) interactions display a clear asymmetric pattern. We propose an innovative approach named ChatbotID, which applies the Granger Causality Test (GCT) to extract a novel set of interactional features that capture the evolving, predictive relationships between conversational attributes. By synergistically fusing these GCT-based interactional features with contextual embeddings, and optimizing the model through a meticulous loss function. Experimental results across multiple datasets and detection models demonstrate the effectiveness of our framework, with significant improvements in accuracy for distinguishing between H-H and H-C dialogues.
Xiaoquan Yi, Haozhao Wang, Yining Qi, Wenchao Xu 0001, Rui Zhang 0003, Yuhua Li 0003, Ruixuan Li 0001
NeurIPS4
2025 Privacy-Friendly Cross-Domain Recommendation via Distilling User-irrelevant Information
abstract
Privacy-preserving Cross-Domain Recommendation (CDR) has been extensively studied to address the cold-start problem using auxiliary source domains while simultaneously protecting sensitive information. However, existing privacy-preserving CDR methods rely heavily on transferring sensitive user embeddings or behaviour logs, which leads to adopt privacy methods to distort the data patterns before transferring it to the target domain. The distorted information can compromise overall performance during the knowledge transfer process. To overcome these challenges, our approach differs from existing privacy-preserving methods that focus on safeguarding user-sensitive information. Instead, we concentrate on distilling transferable knowledge from insensitive item embeddings, which we refer to as prototypes. Specifically, we propose a conditional model inversion mechanism to accurately distill prototypes for individual users. We have designed a new data format and corresponding learning paradigm for distilling transferable prototypes from traditional recommendation models using model inversion. These prototypes facilitate bridging the domain shift between distinct source and target domains in a privacy-friendly manner. Additionally, they enable the identification of top-k users in the target domain to substitute for cold-start users prediction. We conduct extensive experiments across large real-world datasets, and the results substantiate the effectiveness of PFCDR https://github.com/walcheng/PFCDR.
Cheng Wang 0025, Wenchao Xu 0001, Haozhao Wang, Wei Liu 0144, Ruixuan Li 0001
WWW2
2025 A Lightweight Multi-BS Cooperation AKA Mechanism for Fully Decoupled RAN
abstract
The fully decoupled radio access network (FD-RAN) divides the base station (BS) into three components: 1) the uplink data BS (UBS); 2) the downlink data BS (DBS); and 3) the control BS (CBS) to achieve an ultra flexible network. FD-RAN can easily realize dynamic multi-BS cooperation to boost throughput and guarantee users’ quality of experience. However, as the number of cooperating BSs and concurrently accessed Internet of Things (IoT) devices increases, significant authentication and key agreement (AKA) overheads will be incurred, and the system becomes vulnerable to various security threats. This article proposes a novel lightweight AKA protocol that employs secret value sharing for secure key negotiation among BSs, user equipment and the core network (CN). Furthermore, multidevice of IoT access is optimized through aggregated message authentication codes with detecting functionality (AMAD) to efficiently manage and aggregate access requests. By utilizing interpolation polynomials and multiparty key agreement techniques, the proposed solution improves key negotiation efficiency while mitigating risks from man-in-the-middle (MitM) and Distributed Denial of Service (DDoS) attacks. Security analyses are conducted to validate the protocol while showcasing its superiority on computation, communication and transmission efficiency.
Ning Wang 0053, Wenchao Xu 0001, Liquan Chen
IEEE Internet Things J.3
2025 A Hybrid Method for Source Direction Finding With Radio Frequency Interference and Gaussian White Noise
abstract
This paper presents a hybrid data-driven method, termed moving average-Hankel-dynamic mode decomposition (MAHankDMD), for joint direction of arrival (DOA) and frequency estimation in environments affected by both radio frequency interference (RFI) and Gaussian white noise. The proposed approach integrates two key components: (1) a moving average-DMD filter that effectively mitigates Gaussian white noise and separates RFI from the source signal, and (2) a Hankel-DMD method that accurately estimates the DOA of the filtered signal and associates it with the corresponding frequency. The moving average-DMD stage first enhances the signal-to-noise ratio and improves the robustness of the estimation process through noise and inference mitigation, while the subsequent Hankel-DMD stage enables reliable parameter extraction even for overlapping sginals or strong interference conditions. Numerical simulations demonstrate the robustness of MAHankDMD, showing its ability to precisely estimate both DOA and frequency under challenging conditions involving RFI and Gaussian white noise interference. The proposed algorithm thus provides an effective solution for channel parameter estimation in complex noisy environments.
Wenchao Xu 0001, Antonios Argyriou, A-Long Jin, Tianquan Tang, Peifeng Ma, Lijun Jiang
IEEE Internet Things J.2
2025 A Tensor-Based Data-Driven Approach for Multidimensional Harmonic Retrieval and Its Application for MIMO Channel Sounding
abstract
In wireless channel sounding, accurately estimating multiple parameters within a multipath signal, such as azimuth, elevation, Doppler shift, and delay, necessitates addressing the challenges posed by the multidimensional harmonic retrieval (MHR) problem. To overcome these complexities, we propose a framework based on high-order dynamic mode decomposition (HODMD) that designed for robustly estimating frequencies of interest from high-dimensional sinusoidal signals, particularly in additive white Gaussian noise conditions. The HODMD approach, a hybrid algorithm amalgamating high-order singular value decomposition (HOSVD) and dynamic mode decomposition (DMD), operates by initially decomposing observed tensorial data into a core tensor and R mode matrices through HOSVD. Subsequently, DMD is applied to analyze each mode matrix individually, decomposing it into dynamic modes and DMD eigenvalues. The imaginary component of the DMD eigenvalues yields frequencies along the rth dimension. By uniformly applying this analysis to all mode matrices, multiple frequencies of interest are efficiently obtained. Furthermore, the integration of HOSVD, DMD, and moving average techniques in the proposed method is designed to mitigate noise interference during the MHR process. We conduct several numerical experiments and present a real-life example, i.e., the double-direction multiple-input and multiple-output (MIMO) channel sounding, to validate the effectiveness of the proposed HODMD approach. Results demonstrate that HODMD outperforms comparable approaches, particularly in scenarios characterized by high-signal-to-noise ratios. Notably, the proposed method exhibits the capability to estimate the number of tones in undamped cases during the decomposition process. Hence, our work contributes a practical and effective tensor-based solution to the MHR problem, particularly in the context of channel parameter estimation for MIMO systems.
Wenchao Xu 0001, A-Long Jin, Min Li 0032, Ping Yuan, Lijun Jiang
IEEE Internet Things J.2
2025 Enhanced Multidimensional Harmonic Retrieval in MIMO Wireless Channel Sounding
abstract
This article introduces a recursive parallel dynamic mode decomposition (RPDMD) scheme tailored for multidimensional harmonic retrieval (MHR), specifically applied to MIMO wireless channel sounding. The RPDMD algorithm is devised to address the complexities inherent in multidimensional scenarios, leveraging the dynamic mode decomposition (DMD) framework within a recursive parallel structure. Initially, the observed tensorial multidimensional harmonic data is transformed into a 2-D matrix format along the rth dimension. Subsequently, DMD dissects this matrix data into eigenvalues and their associated modes. The real and imaginary components of the DMD eigenvalues yield damping factors and frequencies in the rth dimension, respectively. Furthermore, recursive DMD is employed to scrutinize each mode independently for parameter retrieval across the remaining dimensions, enabling parallel analysis. Ultimately, this high-dimensional correlated decomposition scheme delivers paired damping factors and frequencies for all tones. Notably, the proposed approach can ascertain the number of tones in undamped sinusoidal signals, making it particularly suitable for MHR even without prior knowledge of the source count. Numerical experiments demonstrate the accuracy and robustness of the RPDMD scheme, with comparative analysis indicating that RPDMD outperforms similar methods, achieving optimal results with minimal mean square error in high signal-to-noise ratio scenarios. This work presents an effective data-driven solution for the MHR problem in MIMO wireless channel sounding.
Wenchao Xu 0001, A-Long Jin, Tianquan Tang, Min Li 0032, Peifeng Ma, Lijun Jiang
IEEE Internet Things J.2
2025 FeCoGraph: Label-Aware Federated Graph Contrastive Learning for Few-Shot Network Intrusion Detection
abstract
With increasing cyber attacks over the Internet, network intrusion detection systems (NIDS) have been an indispensable barrier to protecting network security. Taking advantage of automatically capturing topology connections, recent deep graph learning approaches have achieved remarkable performance in distinguishing different types of malicious flows. However, there remain some critical challenges. 1) previous supervised learning methods rely heavily on abundant and high-quality annotated samples, while label annotation requires abundant time and expert knowledge. 2) Centralized methods require all data to be uploaded to a server for learning behavior patterns, which results in high detection latency and critical privacy leakage. 3) Diverse attack scenarios exhibit highly imbalanced distribution, making it hard to characterize abnormal behaviors. To address these issues, we proposed FeCoGraph, a label-aware federated graph contrastive learning framework for intrusion detection in few-shot scenarios. The line graph is introduced to directly process flow embeddings, which are compatible with diverse GNNs. Furthermore, We formulate a graph contrastive learning task to effectively leverage label information, allowing intra-class embeddings more compact than inter-class embeddings. To improve the scalability of NIDS, we utilize federated learning to cover more attack scenarios while protecting data privacy. Experiment results show that FeCoGraph surpass E-graphSAGE with an average 8.36% accuracy on binary classification and 6.77% accuracy on multiclass classification, demonstrating the efficiency of our approach.
Qinghua Mao, Xi Lin 0003, Wenchao Xu 0001, Yuxin Qi 0001, Xiu Su, Gaolei Li, Jianhua Li 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Average AoI Minimization With Directional Charging for Wireless-Powered Network Edge
abstract
Age of Information (AoI) has been proposed as a new performance metric to capture the freshness of data. At wireless-powered network edge, the source nodes first need to be charged ready for update transmissions, which means the system AoI is not only decided by the scheduling of update transmissions but also by the designing of charging plan. However, the existing works either only focused on the point of scheduling update transmissions or have a rigid assumption that only one source node can be charged per time. Aiming at making the work more practical and general, we investigate the average AoI optimization problem at wireless-powered network edge with directional charging. Firstly, the theoretical bound of the weighted sum of average AoI of the entire network with a directional charger is analyzed, which is proved to be related to nodes' maximum transmitting interval and the charging strategy. An optimal charging time computation algorithm is proposed to obtain the maximum transmitting interval of each source node by considering the overlapped areas of different charging orientations. After then, an AoI-aware periodical charging scheduling algorithm is proposed to compute a periodical charging schedule while the average AoI is bounded, including a charging period$T$and the charging orientations assigned to each time slot within$T$. The proposed algorithm is proved to have an approximation ratio of up to 1.5625. Furthermore, several approximate algorithms are also proposed for average AoI optimization with multiple chargers and network bandwidth constraint. Finally, the extensive simulations demonstrate the high performance of the proposed algorithm in terms of AoI.
Quan Chen 0003, Song Guo 0001, Wenchao Xu 0001, Jing Li 0093, Hong Gao 0001, Zhipeng Cai 0001
IEEE Trans. Mob. Comput.3
2025 Fast Multimodal Edge Inference via Selective Feature Distillation
abstract
Inferring user status at the edge is essential for delivering personalized services, such as detecting emotional states. However, deploying large-scale models directly on user devices is impractical due to substantial computational overhead and the scarcity of labeled data. Conversely, uploading raw data to the cloud for processing raises significant privacy concerns and incurs prohibitive communication costs. To address this challenge, we propose a privacy-preserving multimodal inference framework that leverages large-scale public data while safeguarding sensitive information and optimizing computational efficiency. Specifically, we first train a teacher model in the cloud using publicly available data. Through a feature distillation process, the knowledge from this teacher model is transferred to a lightweight encoder deployed at the user end. This transfer is tailored to the user's data, ensuring that only relevant knowledge is distilled. To accommodate varying communication constraints, we introduce a feature compression mechanism that significantly reduces communication overhead without compromising inference accuracy. Extensive experiments on emotion recognition tasks demonstrate that the proposed framework effectively balances privacy preservation, resource efficiency, and inference accuracy, facilitating seamless collaboration between cloud and edge devices.
Wenchao Xu 0001, Yunfeng Fan, Haozhao Wang, Quan Chen 0003, Jing Li 0093
IEEE Trans. Mob. Comput.2
2025 Mobility and Cost Aware Inference Accelerating Algorithm for Edge Intelligence
abstract
The edge intelligence (EI) has been widely applied recently. Splitting the model between device, edge server, and cloud can significantly improve the performance of EI. The model segmentation without user mobility has been investigated in detail in previous studies. However, in most EI use cases, the end devices are mobile. Few studies have been conducted on this topic. These works still have many issues, such as ignoring the energy consumption of mobile device, inappropriate network assumption, and low effectiveness on adapting user mobility, etc. Therefore, to address the disadvantages of model segmentation and resource allocation in previous studies, we propose mobility and cost aware model segmentation and resource allocation algorithm for accelerating the inference at edge (MCSA). Specifically, in the scenario without user mobility, the loop iteration gradient descent (Li-GD) algorithm is provided. When the mobile user has a large model inference task that needs to be calculated, it will take the energy consumption of mobile user, the communication and computing resource renting cost, and the inference delay into account to find the optimal model segmentation and resource allocation strategy. In the scenario with user mobility, the mobility aware Li-GD (MLi-GD) algorithm is proposed to calculate the optimal strategy. Then, the properties of the proposed algorithms are investigated, including convergence, complexity, and approximation ratio. The experimental results demonstrate the effectiveness of the proposed algorithms.
Xin Yuan 0003, Ning Li 0003, Kang Wei 0004, Wenchao Xu 0001, Quan Chen 0003, Hao Chen 0045, Song Guo 0001
IEEE Trans. Mob. Comput.4
2025 Model Decomposition and Reassembly for Purified Knowledge Transfer in Personalized Federated Learning
abstract
Personalized federated learning (pFL) is to collaboratively train non-identical machine learning models for different clients to adapt to their heterogeneously distributed datasets. State-of-the-art pFL approaches pay much attention on exploiting clients’ inter-similarities to facilitate the collaborative learning process, meanwhile, can barely escape from the irrelevant knowledge pooling that is inevitable during the aggregation phase, and thus hindering the optimization convergence and degrading the personalization performance. To tackle such conflicts between facilitating collaboration and promoting personalization, we propose a novel pFL framework, dubbed pFedC, to first decompose the global aggregated knowledge into several compositional branches, and then selectively reassemble the relevant branches for supporting conflicts-aware collaboration among contradictory clients. Specifically, by reconstructing each local model into a shared feature extractor and multiple decomposed task-specific classifiers, the training on each client transforms into a mutually reinforced and relatively independent multi-task learning process, which provides a new perspective for pFL. Besides, we conduct a purified knowledge aggregation mechanism via quantifying the combination weights for each client to capture clients’ common prior, as well as mitigate potential conflicts from the divergent knowledge caused by the heterogeneous data. Extensive experiments over various models and datasets demonstrate the effectiveness and superior performance of the proposed algorithm.
Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Wenchao Xu 0001, Qihua Zhou, Jingcai Guo, Zicong Hong, Jun Shan
IEEE Trans. Mob. Comput.4
2025 Knowledge-Aware Parameter Coaching for Communication-Efficient Personalized Federated Learning in Mobile Edge Computing
abstract
Personalized Federated Learning (pFL) can improve the accuracy of local models and provide enhanced edge intelligence without exposing the raw data in Mobile Edge Computing (MEC). However, in the MEC environment with constrained communication resources, transmitting the entire model between the server and the clients in traditional pFL methods imposes substantial communication overhead, which can lead to inaccurate personalization and degraded performance of mobile clients. In response, we propose a Communication-Efficient pFL architecture to enhance the performance of personalized models while minimizing communication overhead in MEC. First, a Knowledge-Aware Parameter Coaching method (KAPC) is presented to produce a more accurate personalized model by utilizing the layer-wise parameters of other clients with adaptive aggregation weights. Then, convergence analysis of the proposed KAPC is developed in both the convex and non-convex settings. Second, a Bidirectional Layer Selection algorithm (BLS) based on self-relationship and generalization error is proposed to select the most informative layers for transmission, which reduces communication costs. Extensive experiments are conducted, and the results demonstrate that the proposed KAPC achieves superior accuracy compared to the state-of-the-art baselines, while the proposed BLS substantially improves resource utilization without sacrificing performance.
Mingjian Zhi, Yuanguo Bi, Lin Cai 0001, Wenchao Xu 0001, Haozhao Wang, Tianao Xiang, Qiang He 0002
IEEE Trans. Mob. Comput.4
2025 Harmonizing Global and Local Class Imbalance for Federated Learning
abstract
Federated Learning (FL) is to collaboratively train a global model among distributed clients by iteratively aggregating their local updates without sharing their raw data, whereby the global modal can approximately converge to the centralized training way over a global dataset that composed of all local datasets (i.e., union of all users’ local data). However, in real-world scenarios, the distributions of the data classes are often imbalanced not only locally, but also in the global dataset, which severely deteriorate the FL performance due to the conflicting knowledge aggregation. Existing solutions for FL class imbalance either focus on the local data to regulate the training process or purely aim at the global datasets, which often fail to alleviate the class imbalance problem if there is mismatch between the local and global imbalance. Considering these limitations, this paper proposes a Global-Local Joint Learning method, namely GLJL, which simultaneously harmonizes the global and local class imbalance issue by jointly embedding the local and the global factors into each client’s loss function. Through extensive experiments over popular datasets with various class imbalance settings, we show that the proposed method can significantly improve the model accuracy over minority classes without sacrificing the accuracy of other classes.
Wenchao Xu 0001, Haozhao Wang, Zhiming He, Yuxuan Liu 0017, Shuang Wang 0002
IEEE Trans. Mob. Comput.3
2025 Multimodal Dual-Embedding Networks for Malware Open-Set Recognition
abstract
Malware open-set recognition (MOSR) is an emerging research domain that aims at jointly classifying malware samples from known families and detecting the ones from novel unknown families, respectively. Existing works mostly rely on a well-trained classifier considering the predicted probabilities of each known family with a threshold-based detection to achieve the MOSR. However, our observation reveals that the feature distributions of malware samples are extremely similar to each other even between known and unknown families. Thus, the obtained classifier may produce overly high probabilities of testing unknown samples toward known families and degrade the model performance. In this article, we propose the multi\modal dual-embedding networks, dubbed MDENet, to take advantage of comprehensive malware features from different modalities to enhance the diversity of malware feature space, which is more representative and discriminative for down-stream recognition. Concretely, we first generate a malware image for each observed sample based on their numeric features using our proposed numeric encoder with a re- designed multiscale CNN structure, which can better explore their statistical and spatial correlations. Besides, we propose to organize tokenized malware features into a sentence for each sample considering its behaviors and dynamics, and utilize language models as the textual encoder to transform it into a representable and computable textual vector. Such parallel multimodal encoders can fuse the above two components to enhance the feature diversity. Last, to further guarantee the open-set recognition (OSR), we dually embed the fused multimodal representation into one primary space and an associated sub-space, i.e., discriminative and exclusive spaces, with contrastive sampling and -bounded enclosing sphere regularizations, which resort to classification and detection, respectively. Moreover, we also enrich our previously proposed large-scaled malware dataset MAL-100 with multimodal characteristics and contribute an improved version dubbed MAL-100+. Experimental results on the widely used malware dataset Mailing and the proposed MAL-100+ demonstrate the effectiveness of our method.
Jingcai Guo, Yuanyuan Xu 0004, Wenchao Xu 0001, Yufeng Zhan, Yuxia Sun, Song Guo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 ProCC: Progressive Cross-Primitive Compatibility for Open-World Compositional Zero-Shot Learning
abstract
Open-World Compositional Zero-shot Learning (OW-CZSL) aims to recognize novel compositions of state and object primitives in images with no priors on the compositional space, which induces a tremendously large output space containing all possible state-object compositions. Existing works either learn the joint compositional state-object embedding or predict simple primitives with separate classifiers. However, the former method heavily relies on external word embedding methods, and the latter ignores the interactions of interdependent primitives, respectively. In this paper, we revisit the primitive prediction approach and propose a novel method, termed Progressive Cross-primitive Compatibility (ProCC), to mimic the human learning process for OW-CZSL tasks. Specifically, the cross-primitive compatibility module explicitly learns to model the interactions of state and object features with the trainable memory units, which efficiently acquires cross-primitive visual attention to reason high-feasibility compositions, without the aid of external knowledge. Moreover, to alleviate the invalid cross-primitive interactions, especially for partial-supervision conditions (pCZSL), we design a progressive training paradigm to optimize the primitive classifiers conditioned on pre-trained features in an easy-to-hard manner. Extensive experiments on three widely used benchmark datasets demonstrate that our method outperforms other representative methods on both OW-CZSL and pCZSL settings by large margins.
Fushuo Huo, Wenchao Xu 0001, Song Guo 0001, Jingcai Guo, Haozhao Wang, Xiaocheng Lu
AAAI2
2024 Non-exemplar Online Class-Incremental Continual Learning via Dual-Prototype Self-Augment and Refinement
abstract
This paper investigates a new, practical, but challenging problem named Non-exemplar Online Class-incremental continual Learning (NO-CL), which aims to preserve the discernibility of base classes without buffering data examples and efficiently learn novel classes continuously in a single-pass (i.e., online) data stream. The challenges of this task are mainly two-fold: (1) Both base and novel classes suffer from severe catastrophic forgetting as no previous samples are available for replay. (2) As the online data can only be observed once, there is no way to fully re-train the whole model, e.g., re-calibrate the decision boundaries via prototype alignment or feature distillation. In this paper, we propose a novel Dual-prototype Self-augment and Refinement method (DSR) for NO-CL problem, which consists of two strategies: 1) Dual class prototypes: vanilla and high-dimensional prototypes are exploited to utilize the pre-trained information and obtain robust quasi-orthogonal representations rather than example buffers for both privacy preservation and memory reduction. 2) Self-augment and refinement: Instead of updating the whole network, we optimize high-dimensional prototypes alternatively with the extra projection module based on self-augment vanilla prototypes, through a bi-level optimization problem. Extensive experiments demonstrate the effectiveness and superiority of the proposed DSR in NO-CL.
Fushuo Huo, Wenchao Xu 0001, Jingcai Guo, Haozhao Wang, Yunfeng Fan
AAAI2
2024 Knowledge-Aware Parameter Coaching for Personalized Federated Learning
abstract
Personalized Federated Learning (pFL) can effectively exploit the non-IID data from distributed clients by customizing personalized models. Existing pFL methods either simply take the local model as a whole for aggregation or require significant training overhead to induce the inter-client personalized weights, and thus clients cannot efficiently exploit the mutually relevant knowledge from each other. In this paper, we propose a knowledge-aware parameter coaching scheme where each client can swiftly and granularly refer to parameters of other clients to guide the local training, whereby accurate personalized client models can be efficiently produced without contradictory knowledge. Specifically, a novel regularizer is designed to conduct layer-wise parameters coaching via a relation cube, which is constructed based on the knowledge represented by the layered parameters among all clients. Then, we develop an optimization method to update the relation cube and the parameters of each client. It is theoretically demonstrated that the convergence of the proposed method can be guaranteed under both convex and non-convex settings. Extensive experiments are conducted over various datasets, which show that the proposed method can achieve better performance compared with the state-of-the-art baselines in terms of accuracy and convergence speed.
Mingjian Zhi, Yuanguo Bi, Wenchao Xu 0001, Haozhao Wang, Tianao Xiang
AAAI3
2024 C2KD: Bridging the Modality Gap for Cross-Modal Knowledge Distillation
abstract
Existing Knowledge Distillation (KD) methods typically focus on transferring knowledge from a large-capacity teacher to a low-capacity student model, achieving sub-stantial success in unimodal knowledge transfer. However, existing methods can hardly be extended to Cross-Modal Knowledge Distillation (CMKD), where the knowledge is transferred from a teacher modality to a different student modality, with inference only on the distilled student modality. We empirically reveal that the modality gap, i.e., modality imbalance and soft label misalignment, incurs the in-effectiveness of traditional KD in CMKD. As a solution, we propose a novel f_ustomized Crossmodal Knowledge Distillation (C2KD). Specifically, to alleviate the modality gap, the pre-trained teacher performs bidirectional distillation with the student to provide customized knowledge. The On-the-Fly Selection Distillationi OFSD) strategy is applied to selectively filter out the samples with misaligned soft labels, where we distill cross-modal knowledge from non-target classes to avoid the modality imbalance issue. To further provide receptive cross-modal knowledge, proxy student and teacher, inheriting unimodal and cross-modal knowledge, is formulated to progressively transfer cross-modal knowledge through bidirectional distillation. Experimental results on audio-visual, image-text, and RGB-depth datasets demonstrate that our method can effectively transfer knowledge across modalities, achieving superior performance against traditional KD by a large margin.
Fushuo Huo, Wenchao Xu 0001, Jingcai Guo, Haozhao Wang, Song Guo 0001
CVPR2
2024 Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection
Yunfeng Fan, Wenchao Xu 0001, Haozhao Wang, Fushuo Huo, Song Guo 0001
ECCV (81)2
2024 Personalized Federated Domain-Incremental Learning Based on Adaptive Knowledge Matching
Yichen Li 0006, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Jingcai Guo, Ruixuan Li 0001
ECCV (46)2
2024 ALWNN: Automatic Modulation Classification via Adaptive Lightweight Wavelet Neural Network
abstract
Automatic Modulation Classification (AMC) plays a crucial role in non-cooperative communication systems and is an essential component of blind signal processing. The application of deep learning methods in modulation classification has shown tremendous potential, surpassing the performance of traditional methods by a large margin. However, the high storage and computational requirements of existing deep learning methods limit their practical applications. In this paper, we propose an AMC technique using an Adaptive Lightweight Wavelet Neural Network (ALWNN) that features a streamlined design and lower computational demands. This innovative model introduces an adaptive wavelet-based feature extraction method that effectively captures information at different frequencies in the input data, ensuring classification accuracy. Additionally, the model incorpo-rates depthwise separable convolution techniques, transforming traditional convolutions into depthwise convolutions and point-wise convolutions, Substantially diminishing the count of the model’s parameters and the complexity of its computations. The proposed ALWNN model strikes a balance between efficiency and accuracy. Simulation results demonstrate that with only 9899 and 9700 parameters, it achieves accuracies of 62.14% and 63.93% on the datasets known as RML2016.10a and RML2016.10b, respectively. Furthermore, we evaluate the model in terms of Floating Point Operations Per Second (FLOPS) and Normalized Multiply-Accumulate Complexity (NMACC) to provide a more comprehensive measure of computational complexity. Compared to existing methods, ALWNN reduces FLOPS by 1.25 to 1.91 orders of magnitude and NMACC by 0.81 to 1.6 orders of magnitude.
Yunhao Quan, Nan Cheng 0001, Xiucheng Wang, Zhisheng Yin, Wenchao Xu 0001
GLOBECOM5
2024 AWSSS: Adaptive Weighted Statistical Space Smoothing for Regression with Imbalance Data
abstract
Deep learning models, usually trained on large datasets, have witnessed great achievements for both classification and regression tasks during past years. However, in practical settings, the datasets are commonly imbalanced where the number of data samples differs among labels (i.e.,categories or target values), leading to serious performance degradation of the trained models. Although many prior works have been proposed to solve the data imbalance problem for the classification task, there are still limited works considering the regression task. In this work, we identify the distinct challenges in regression problems compared to traditional classification problems, such as managing the continuous nature of the target variable and navigating fuzzy decision-making boundaries. To address these challenges, we propose an Adaptive Weighted Statistical Space Smoothing method (AWSSS), which alleviates the impact of imbalanced data by learning from continuously valued, imbalanced data and extending local similarity across the entire target range. AWSSS extracts statistical values from the original dataset and smoothes statistical features. Subsequently, it aligns features across different regions and applies adaptive weights to facilitate knowledge transfer between adjacent areas. Finally, AWSSS calculates the loss values, which are used to refine the model’s parameters through iterative updates. Extensive experiments conducted on various datasets demonstrate the effectiveness of the proposed method as compared to state-of-the-art methods.
Xiaoquan Yi, Haozhao Wang, Zhenlong Zhu, Wei Liu 0144, Wenchao Xu 0001, Ruixuan Li 0001
HPCC5
2024 Disentangle Estimation of Causal Effects from Cross-Silo Data
abstract
Estimating causal effects among different events is of great importance to critical fields such as drug development. Nevertheless, the data features associated with events may be distributed across various silos and remain private within respective parties, impeding direct information exchange between them. This, in turn, can result in biased estimations of local causal effects, which rely on the characteristics of only a subset of the covariates. To tackle this challenge, we introduce an innovative disentangle architecture designed to facilitate the seamless cross-silo transmission of model parameters, enriched with causal mechanisms, through a combination of shared and private branches. Besides, we introduce global constraints into the equation to effectively mitigate bias within the various missing domains, thereby elevating the accuracy of our causal effect estimation. Extensive experiments conducted on new semi-synthetic datasets show that our method outperforms state-of-the-art baselines.
Yuxuan Liu 0017, Haozhao Wang, Shuang Wang 0002, Zhiming He, Wenchao Xu 0001
ICASSP5
2024 Physical Layer Overshadowing Attack on Semantic Communication System
abstract
Semantic communication systems (SCS) have gained extensive attention with the advancement of Artificial Intelligence (AI), which transmits the data feature instead of the raw bits, whereby the communication efficiency can be substantially enhanced, e.g., via a neural network or encoder to convert from the massive user data to corresponding light-weight feature map. However, SCS can be vulnerable to adversarial noise when transmitting the feature data, which may mislead the downstream tasks at the receiver side, e.g., leading to misclassification due to the disturbed receiving information. In this paper, we investigate the overshadowing-based attacks by perturbing the physical signal with artificial adversarial noise during the semantic feature transmission. Specifically, we directly attack the waveform after the modulation of the feature bits, and conduct both the white-box and black-box attacks to evaluate the vulnerability. In our attack methods, we use the local transfer model to acquire the gradient details and provide the gradient-based strategy for generating the perturbation. The experiment results demonstrate that both white-box and black-box attacks can be a critical threat for SCS and significantly degrade the performance of downstream tasks.
Zhaoyi Lu 0001, Wenchao Xu 0001, Haozhao Wang, Cunqing Hua
ICC2
2024 Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language Models
abstract
Fine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from invariant features, and employ a uniform alignment process when adapting to unseen target domains. This can impair the cross-modal feature alignment when the testing data significantly deviate from the distribution of the training data, resulting in a poor out-of-distribution (OOD) generalization performance. In this paper, we reveal that the prompt tuning failure in such OOD scenarios can be attribute to the undesired alignment between the textual and the spurious feature. As a solution, we propose **CoOPood**, a fine-grained prompt tuning method that can discern the causal features and deliberately align the text modality with the invariant feature. Specifically, we design two independent contrastive phases using two lightweight projection layers during the alignment, each with different objectives: 1) pulling the text embedding closer to invariant image embedding and 2) pushing text embedding away from spurious image embedding. We have illustrated that **CoOPood** can serve as a general framework for VLMs and can be seamlessly integrated with existing prompt tuning methods. Extensive experiments on various OOD datasets demonstrate the performance superiority over state-of-the-art methods.
Jie Zhang 0076, Xiaosong Ma, Song Guo 0001, Peng Li 0017, Wenchao Xu 0001, Xueyang Tang, Zicong Hong
ICML5
2024 FedBAT: Communication-Efficient Federated Learning via Learnable Binarization
abstract
Federated learning is a promising distributed machine learning paradigm that can effectively exploit large-scale data without exposing users’ privacy. However, it may incur significant communication overhead, thereby potentially impairing the training efficiency. To address this challenge, numerous studies suggest binarizing the model updates. Nonetheless, traditional methods usually binarize model updates in a post-training manner, resulting in significant approximation errors and consequent degradation in model accuracy. To this end, we propose Federated Binarization-Aware Training (FedBAT), a novel framework that directly learns binary model updates during the local training process, thus inherently reducing the approximation errors. FedBAT incorporates an innovative binarization operator, along with meticulously designed derivatives to facilitate efficient learning. In addition, we establish theoretical guarantees regarding the convergence of FedBAT. Extensive experiments are conducted on four popular datasets. The results show that FedBAT significantly accelerates the convergence and exceeds the accuracy of baselines by up to 9%, even surpassing that of FedAvg in some cases.
Shiwei Li 0002, Wenchao Xu 0001, Haozhao Wang, Xing Tang 0007, Yining Qi, Weihong Luo, Yuhua Li 0003, Xiuqiang He 0001, Ruixuan Li 0001
ICML2
2024 Contamination-Resilient Anomaly Detection via Adversarial Learning on Partially-Observed Normal and Anomalous Data
abstract
Many existing anomaly detection methods assume the availability of a large-scale normal dataset. But for many applications, limited by resources, removing all anomalous samples from a large un-labeled dataset is unrealistic, resulting in contaminated datasets. To detect anomalies accurately under such scenarios, from the probabilistic perspective, the key question becomes how to learn the normal-data distribution from a contaminated dataset. To this end, we propose to collect two additional small datasets that are comprised of partially-observed normal and anomaly samples, and then use them to help learn the distribution under an adversarial learning scheme. We prove that under some mild conditions, the proposed method is able to learn the correct normal-data distribution. Then, we consider the overfitting issue caused by the small size of the two additional datasets, and a correctness-guaranteed flipping mechanism is further developed to alleviate it. Theoretical results under incomplete observed anomaly types are also presented. Extensive experimental results demonstrate that our method outperforms representative baselines when detecting anomalies under contaminated datasets.
Wenxi Lv, Qinliang Su, Hai Wan, Hongteng Xu, Wenchao Xu 0001
ICML5
2024 OTAS: An Elastic Transformer Serving System via Token Adaptation
abstract
Transformer model empowered architectures have become a pillar of cloud services that keeps reshaping our society. However, the dynamic query loads and heterogeneous user requirements severely challenge current transformer serving systems, which rely on pre-training multiple variants of a foundation model, i.e., with different sizes, to accommodate varying service demands. Unfortunately, such a mechanism is unsuitable for large transformer models due to the additional training costs and excessive I/O delay. In this paper, we introduce OTAS, the first elastic serving system specially tailored for transformer models by exploring lightweight token management. We develop a novel idea called token adaptation that adds prompting tokens to improve accuracy and removes redundant tokens to accelerate inference. To cope with fluctuating query loads and diverse user requests, we enhance OTAS with application-aware selective batching and online token adaptation. OTAS first batches incoming queries with similar service-level objectives to improve the ingress throughput. Then, to strike a tradeoff between the overhead of token increment and the potentials for accuracy improvement, OTAS adaptively adjusts the token execution strategy by solving an optimization problem. We implement and evaluate a prototype of OTAS with multiple datasets, which show that OTAS improves the system utility by at least 18.2%.
Wenchao Xu 0001, Zicong Hong, Song Guo 0001, Haozhao Wang, Jie Zhang 0076, Deze Zeng
INFOCOM2
2024 FedNLR: Federated Learning with Neuron-wise Learning Rates
abstract
Federated Learning (FL) suffers from severe performance degradation due to the data heterogeneity among clients. Some existing work suggests that the fundamental reason is that data heterogeneity can cause local model drift, and therefore proposes to calibrate the direction of local updates to solve this problem. Though effective, existing methods generally take the model as a whole, which lacks a deep understanding of how the neurons within deep classification models evolve during local training to form model drift. In this paper, we bridge this gap by performing an intuitive and theoretical analysis of the activation changes of each neuron during local training. Our analysis shows that the high activation of some neurons on the samples of a certain class will be reduced during local training when these samples are not included in the client, which we call neuron drift, thus leading to the performance reduction of this class. Motivated by this, we propose a novel and simple algorithm called FedNLR, which utilizes Neuron-wise Learning Rates during the FL local training process. The principle behind this is to enhance the learning of neurons bound to local classes on local data knowledge while reducing the decay of non-local classes knowledge stored in neurons. Experimental results demonstrate that FedNLR achieves state-of-the-art performance on federated learning with popular deep neural networks.
Haozhao Wang, Peirong Zheng, Xingshuo Han, Wenchao Xu 0001, Ruixuan Li 0001, Tianwei Zhang 0004
KDD4
2024 Detached and Interactive Multimodal Learning
abstract
Recently, Multimodal Learning (MML) has gained significant interest as it compensates for single-modality limitations through comprehensive complementary information within multimodal data. However, traditional MML methods generally use the joint learning framework with a uniform learning objective that can lead to the modality competition issue, where feedback predominantly comes from certain modalities, limiting the full potential of others. In response to this challenge, this paper introduces DI-MML, a novel detached MML framework designed to learn complementary information across modalities under the premise of avoiding modality competition. Specifically, DI-MML addresses competition by separately training each modality encoder with isolated learning objectives. It further encourages cross-modal interaction via a shared classifier that defines a common feature space and employing a dimension-decoupled unidirectional contrastive (DUC) loss to facilitate modality-level knowledge transfer. Additionally, to account for varying reliability in sample pairs, we devise a certainty-aware logit weighting strategy to effectively leverage complementary information at the instance level during inference. Extensive experiments conducted on audio-visual, flow-image, and front-rear view datasets show the superior performance of our proposed method. The code is released at https://github.com/fanyunfeng-bit/DI-MML.
Yunfeng Fan, Wenchao Xu 0001, Haozhao Wang, Song Guo 0001
ACM Multimedia2
2024 Cross-modal Representation Flattening for Multi-modal Domain Generalization
abstract
Multi-modal domain generalization (MMDG) requires that models trained on multi-modal source domains can generalize to unseen target distributions with the same modality set. Sharpness-aware minimization (SAM) is an effective technique for traditional uni-modal domain generalization (DG), however, with limited improvement in MMDG. In this paper, we identify that modality competition and discrepant uni-modal flatness are two main factors that restrict multi-modal generalization. To overcome these challenges, we propose to construct consistent flat loss regions and enhance knowledge exploitation for each modality via cross-modal knowledge transfer. Firstly, we turn to the optimization on representation-space loss landscapes instead of traditional parameter space, which allows us to build connections between modalities directly. Then, we introduce a novel method to flatten the high-loss region between minima from different modalities by interpolating mixed multi-modal representations. We implement this method by distilling and optimizing generalizable interpolated representations and assigning distinct weights for each modality considering their divergent generalization capabilities. Extensive experiments are performed on two benchmark datasets, EPIC-Kitchens and Human-Animal-Cartoon (HAC), with various modality combinations, demonstrating the effectiveness of our method under multi-source and single-source settings. Our code is open-sourced.
Yunfeng Fan, Wenchao Xu 0001, Haozhao Wang, Song Guo 0001
NeurIPS2
2024 A General 3-D Geometry-Based Stochastic Channel Model for B5G mmWave IIoT
abstract
The Industrial Internet of Things (IIoT) is one of the typical application scenarios in the beyond fifth generation (B5G) wireless communication systems. Due to numerous metal obstacles and machines, the industrial channel, especially at the millimeter-wave (mmWave) bands, exhibits complex characteristics that have not been considered in existing literature. This article proposes an innovative 3-D nonstationary geometry-based stochastic model (GBSM) for IIoT scenarios at mmWave bands. In the proposed model, device reflections (DRs) caused by massive metal machines are modeled based on geometrical optics. Furthermore, the generalized extreme value (GEV) distribution and generalized Pareto (GP) distribution are used to parameterize the number of clusters and rays within a cluster, respectively. Further, the Doppler shift is modeled and analyzed using the Gaussian distribution. Some channel statistical characteristics are captured by the proposed model, such as the power delay profile, root-mean-square delay spread, root-mean-square angle spread, intercluster delay, and space–time–frequency correlation function. Then, these channel statistical characteristics are well fitted to the ray-tracing simulations and the channel measurements. The excellent fitting results demonstrate the high accuracy of the proposed model, which is crucial for future IIoT communication system design. What is more, this article shows the antenna height and propagation scenarios can significantly affect the DR ratio, which should adapt to various IIoT communication scenarios.
Wen Gu, Yang Liu 0065, Cheng-Xiang Wang 0001, Wenchao Xu 0001, Yu Yu 0002, Wen-Jun Lu, Hongbo Zhu 0002
IEEE Internet Things J.4
2024 GNN-Empowered Effective Partial Observation MARL Method for AoI Management in Multi-UAV Network
abstract
Unmanned aerial vehicles (UAVs), due to their low cost and high flexibility, have been widely used in various scenarios to enhance network performance. However, the optimization of UAV trajectories in unknown areas or areas without sufficient prior information still faces challenges related to poor planning performance and low distributed execution. These challenges arise when UAVs rely solely on their own observation information and the information from other UAVs within their communicable range, without access to global information. To address these challenges, this article proposes the Qedgix framework, which combines graph neural networks (GNNs) and the QMIX algorithm to achieve distributed optimization of the Age of Information (AoI) for users in unknown scenarios. The framework utilizes GNNs to extract information from UAVs, users within the observable range, and other UAVs within the communicable range, thereby enabling effective UAV trajectory planning. Due to the discretization and temporal features of AoI indicators, the Qedgix framework employs QMIX to optimize decentralized partially observable Markov decision processes (Dec-POMDP) based on centralized training and distributed execution (CTDE) with respect to mean AoI values of users. By modeling the UAV network optimization problem in terms of AoI and applying the Kolmogorov-Arnold representation theorem, the Qedgix framework achieves efficient neural network training through parameter sharing based on permutation invariance. Simulation results demonstrate that the proposed algorithm significantly improves convergence speed while reducing the mean AoI values of users. The code is available athttps://github.com/UNIC-Lab/Qedgix.
Yuhao Pan, Xiucheng Wang, Zhiyao Xu, Nan Cheng 0001, Wenchao Xu 0001, Jun-Jie Zhang 0007
IEEE Internet Things J.5
2024 UTDNet: A unified triplet decoder network for multimodal salient object detection
Fushuo Huo, Jingcai Guo, Wenchao Xu 0001, Song Guo 0001
Neural Networks4
2024 Distributed and latency-aware beaconing for asynchronous duty-cycled IoT networks
Qinglin Xie, Peng Long, Yuhang Wu 0008, Quan Chen 0003, Fanlong Zhang, Wenchao Xu 0001
Peer Peer Netw. Appl.7
2024 Mobility-Aware Utility Maximization in Digital Twin-Enabled Serverless Edge Computing
abstract
Driven by data and models, the digital twin technique presents a new concept of optimizing system design, process monitoring, decision-making and more, through performing comprehensive virtual-reality interaction and continuous mapping. By introducing serverless computing to Mobile Edge Computing (MEC) environments, the emerging serverless edge computing paradigm facilitates the communication-efficient digital twin services and promises agile, fine-grained and cost-efficient provisioning of limited edge resources, where serverless functions are implemented by containers in cloudlets (edge servers). However, the nonnegligible cold start delay of containers deteriorates the responsiveness of digital twin services dramatically and the perceived user service experience. In this paper, we investigate delay-sensitive query service provisioning in digital twin-empowered serverless edge computing by considering user mobility. With digital twins of users deployed in the remote cloud, referred to as primary digital twins, we deploy their digital twin replicas based on serverless functions in cloudlets to mitigate the query service delay while enhancing user service satisfaction that is expressed as a utility function. We study two optimization problems with the aim of maximizing the accumulative utility gain: the digital twin replica placement problem per time slot, and the dynamic digital twin replica placement problem over a finite time horizon. We first formulate an Integer Linear Program (ILP) solution for the digital twin replica placement problem when the problem size is small; otherwise, we propose an approximation algorithm for the problem with a provable approximation ratio. We then design an online algorithm for the dynamic digital twin replica placement problem, and a performance-guaranteed online algorithm for a special case of the problem by assuming each user issues a query at each time slot. Finally, we evaluate the performance of the proposed algorithms for placing digital twin replicas in MEC networks through simulations. The results demonstrate the proposed algorithms are promising, outperforming their counterparts.
Jing Li 0093, Song Guo 0001, Weifa Liang, Jianping Wang 0001, Quan Chen 0003, Wenchao Xu 0001, Kang Wei 0004, Xiaohua Jia
IEEE Trans. Computers6
2024 Towards Real-Time Inference Offloading With Distributed Edge Computing: The Framework and Algorithms
abstract
By combining edge computing and parallel computing, distributed edge computing has emerged as a new paradigm to exploit the booming IoT devices at the edge. To accelerate computation at the edge,i.e., the inference tasks for DNN-driven applications, the parallelism of both computation and communication needs to be considered for distributed edge computing, and thus, the problem of Minimum Latency joint Communication and Computation Scheduling (MLCCS) is proposed. However, existing works have rigid assumptions that the communication time of each device is fixed and the workload can be split arbitrarily small. Aiming at making the work more practical and general, the MLCCS problem without the above assumptions is studied in this paper. Firstly, the MLCCS problem under a general model is formulated and proved to be NP-hard. Secondly, a pyramid-based computing model is proposed to consider the parallelism of communication and computation jointly, which has an approximation ratio of$1+\delta$, where$\delta$is related to devices' communication rates. An interesting property under such a computing model is identified and proved,i.e., the optimal latency can be obtained under arbitrary scheduling order when all the devices share the same communication rate. When the workload cannot be split arbitrarily, an approximation algorithm with a ratio of at most$2\cdot (1+\delta )$is proposed. Additionally, for handling the dynamically changing network scenarios, several algorithms are also proposed accordingly. Finally, the theoretical analysis and simulation results verify that the proposed algorithm has high performance in terms of latency. Two testbed experiments are also conducted, which show that the proposed method outperforms the existing methods, reducing the latency by up to 29.2% for inference tasks at the edge.
Quan Chen 0003, Song Guo 0001, Kaijia Wang, Wenchao Xu 0001, Jing Li 0093, Zhipeng Cai 0001, Hong Gao 0001, Albert Y. Zomaya
IEEE Trans. Mob. Comput.4
2024 Caching User-Generated Content in Distributed Autonomous Networks via Contextual Bandit
abstract
The escalating proliferation of user generated contents such as videos and images are dominating the network traffic. The optimal strategy for mitigating backbone congestion and minimizing user request latency lies in prudent caching at edge stations within distributed autonomous networks, obviating the necessity to transmit data to the cloud. However, accurately caching content based on distributed autonomous networks requires elaborative collaboration between edge servers, which remains a great challenge, especially when content is highly dynamic and the storage resources of edge stations are limited. To tackle this challenge, this paper proposes a contextual bandit-based online caching algorithm for evaluating the optimal content hit rate reward, which can adapt to the constantly changing stream of emerging content. We build the content space, BS space, and a fine-grained space searching method to cache contents and corresponding edge stations. Furthermore, to perform collaborative caching and sharing between edges, we propose a federated autonomous multi-layer caching framework, whereby each server can locally learn the model for accurate caching and a synchronous mechanism is set up for global updating, further improving the hit rates. Finally, we perform theoretical proofs and simulations, demonstrating that our regret is sublinear and our caching algorithm outperforms several state-of-the-art algorithms.
Duyu Chen, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Ruixuan Li 0001, Pan Zhou 0001, Song Guo 0001
IEEE Trans. Mob. Comput.2
2024 PromptFL: Let Federated Participants Cooperatively Learn Prompts Instead of Models - Federated Learning in Age of Foundation Model
abstract
Quick global aggregation of effective distributed parameters is crucial to federated learning (FL), which requires adequate bandwidth for parameters communication and sufficient user data for local training. Otherwise, FL may cost excessive training time for convergence and produce inaccurate models. In this paper, we propose a brand-new FL framework, PromptFL, that replaces the federated model training with the federated prompt training, i.e., let federated participants train prompts instead of a shared model, to simultaneously achieve the efficient global aggregation and local training on insufficient data by exploiting the power of foundation models (FM) in a distributed way. PromptFL ships an off-the-shelf FM, i.e., CLIP, to distributed clients who would cooperatively train shared soft prompts based on very few local data. Since PromptFL only needs to update the prompts instead of the whole model, both the local training and the global aggregation can be significantly accelerated. And FM trained over large scale data can provide strong adaptation capability to distributed users tasks with the trained soft prompts. We empirically analyze the PromptFL via extensive experiments, and show its superiority in terms of system feasibility, user privacy, and performance.
Tao Guo 0004, Song Guo 0001, Xueyang Tang, Wenchao Xu 0001
IEEE Trans. Mob. Comput.5
2024 Tree Learning: Towards Promoting Coordination in Scalable Multi-Client Training Acceleration
abstract
Iteration based collaborative learning (CL) paradigms, such as federated learning (FL) and split learning (SL), faces challenges in training neural models over the rapidly growing yet resource-constrained edge devices. Such devices have difficulty in accommodating a full-size large model for FL or affording an excessive waiting time for the mandatory synchronization step in SL. To deal with such challenge, we propose a novel CL framework which adopts an tree-aggregation structure with an adaptive partition and ensemble strategy to achieve optimal synchronization and fast convergence at scale. To find the optimal split point for heterogeneous clients, we also design a novel partitioning algorithm by minimizing the idleness during communication and achieving the optimal synchronization between clients. In addition, a parallelism paradigm is proposed to unleash the potential of optimum synchronization between the clients and server to boost the distributed training process without losing model accuracy for edge devices. Furthermore, we theoretically prove that our framework can achieve better convergence rate than state-of-the-art CL paradigms. We conduct extensive experiments and show that our framework is 4.6× in training speed as compared with the traditional methods, without compromising training accuracy.
Tao Guo 0004, Song Guo 0001, Feijie Wu, Wenchao Xu 0001, Jiewei Zhang, Qihua Zhou, Quan Chen 0003, Weihua Zhuang
IEEE Trans. Mob. Comput.4
2024 RingSFL: An Adaptive Split Federated Learning Towards Taming Client Heterogeneity
abstract
Federated learning (FL) has gained increasing attention due to its ability to collaboratively train while protecting client data privacy. However, vanilla FL cannot adapt to client heterogeneity, leading to a degradation in training efficiency due to stragglers, and is still vulnerable to privacy leakage. To address these issues, this paper proposes RingSFL, a novel distributed learning scheme that integrates FL with a model split mechanism to adapt to client heterogeneity while maintaining data privacy. In RingSFL, all clients form a ring topology. For each client, instead of training the model locally, the model is split and trained among all clients along the ring through a pre-defined direction. By properly setting the propagation lengths of heterogeneous clients, the straggler effect is mitigated, and the training efficiency of the system is significantly enhanced. Additionally, since the local models are blended, it is less likely for an eavesdropper to obtain the complete model and recover the raw data, thus improving data privacy. The experimental results on both simulation and prototype systems show that RingSFL can achieve better convergence performance than benchmark methods on independently identically distributed (IID) and non-IID datasets, while effectively preventing eavesdroppers from recovering training data.
Jinglong Shen, Nan Cheng 0001, Xiucheng Wang, Feng Lyu 0001, Wenchao Xu 0001, Zhi Liu 0002, Khalid Aldubaikhy, Xuemin Shen
IEEE Trans. Mob. Comput.5
2024 Rethinking Personalized Client Collaboration in Federated Learning
abstract
Federated Learning (FL) has gained considerable attention recently, as it allows clients to cooperatively train a global machine learning model without sharing raw data. However, its performance can be compromised due to the high heterogeneity in clients' local data distributions, commonly known as Non-IID (non-independent and identically distributed). Moreover, collaboration among highly dissimilar clients exacerbates this performance degradation. Personalized FL seeks to mitigate this by enabling clients to collaborate primarily with others who have similar data characteristics, thereby producing personalized models. We noticed that existing methods for assessing model similarity often do not capture the genuine relevance of client domains. In response, our paper enhances personalized client collaboration in FL by introducing a metric for domain relevance between clients. Specifically, to facilitate optimal coalition formation, we measure the marginal contributions of client models using coalition game theory, providing a more accurate representation of potential client domain relevance within the FL privacy-preserving framework. Based on this metric, we then adjust each client's coalition membership and implement a personalized FL aggregation algorithm that is robust to Non-IID data domain. We provide a theoretical analysis of the algorithm's convergence and generalization capabilities. Our extensive evaluations on multiple datasets, including MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100, and under varying Non-IID data distributions (Pathological and Dirichlet), demonstrate that our personalized collaboration approach consistently outperforms contemporary benchmarks in terms of accuracy for individual clients.
Leijie Wu, Song Guo 0001, Yaohong Ding, Wenchao Xu 0001, Yufeng Zhan, Anne-Marie Kermarrec
IEEE Trans. Mob. Comput.5
2024 Long-Term Adaptive VCG Auction Mechanism for Sustainable Federated Learning With Periodical Client Shifting
abstract
Federated Learning (FL) system needs to incentivize clients since they may be reluctant to participate in the resource consuming process. Existing incentive mechanisms fail to construct a sustainable environment for the long-term development of FL system: 1) They seldom focus on system economic properties (e.g., social welfare, individual rationality, and incentive compatibility) to guarantee client attraction. 2) Current online auction modeling methods divide the whole continual process into multiple independent rounds and solve them one-by-one, which breaks the correlation between each round. Besides, the inherent characteristics of FL system (model-agnostic and privacy-sensitive) also prevent it from the optimal strategy by precise mathematical analysis. 3) Current system modelings ignore the practical problem of periodical client shifting, which cannot adaptively update its strategy to handle system dynamics. To overcome the above challenges, this paper proposes a long-term adaptive Vickrey–Clarke–Groves (VCG) auction mechanism for FL system, which incorporate a multi-branch deep reinforcement learning (DRL) algorithm. First, VCG auction is the only one that can simultaneously guarantee all crucial economic properties. Second, we extend the economic properties to long-term forms and apply the experience-driven DRL algorithm to directly obtain long-term optimal strategy, without any prior system knowledge. Third, we reconstruct a multi-branch DRL network to accommodate periodical client shifting by adaptive decision head switching for different time periods. Finally, we theoretically prove he extended economic properties (i.e., IC) and conduct extensive experiments on several real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL system increases by 36% with a 37% reduction in payment. Besides, the multi-branch network can adaptively handle periodical client shifting on the timeline.
Leijie Wu, Song Guo 0001, Zicong Hong, Yi Liu 0057, Wenchao Xu 0001, Yufeng Zhan
IEEE Trans. Mob. Comput.5
2024 AirCon: Over-the-Air Consensus for Wireless Blockchain Networks
abstract
Blockchain has been deemed as a promising solution for providing security and privacy protection in the next-generation wireless networks. Large-scale concurrent access for massive wireless devices to accomplish the consensus procedure may consume prohibitive communication and computing resources, and thus may limit the application of blockchain in wireless conditions. As most existing consensus protocols are designed for wired networks, directly apply them for wireless users equipment (UEs) may exhaust their scarce spectrum and computing resources. In this paper, we propose AirCon, a byzantine fault-tolerant (BFT) consensus protocol for wireless UEs via the over-the-air computation. The novelty of AirCon is to take advantage of the intrinsic characteristic of the wireless channel and automatically achieve the consensus in the physical layer while receiving from the UEs, which greatly reduces the communication and computational cost that would be caused by traditional consensus protocols. We implement the AirCon protocol integrated into an LTE system and provide solutions to the critical issues for over-the-air consensus implementation. Experimental results are provided to show the feasibility of the proposed protocol, and simulation results to show the performance of the AirCon protocol under different wireless conditions.
Cunqing Hua, Jianan Hong, Pengwenlong Gu, Wenchao Xu 0001
IEEE Trans. Mob. Comput.5
2024 Mobile Collaborative Learning Over Opportunistic Internet of Vehicles
abstract
Machine learning models are widely applied for vehicular applications, which are essential to future intelligent transportation system (ITS). Traditional model training methods commonly employ a client-server architecture to perform local training and global iterative aggregations, which can consume significant bandwidth resources that are often absent in vehicular networks, especially in high vehicle density scenarios. Modern vehicle users naturally can collaboratively train machine learning models as they are the data owner and have strong local computing power from the onboard units (OBU). In this paper, we propose a novel collaborative learning scheme for mobile vehicles that can utilize the opportunistic vehicle-to-roadside (V2R) communication to exploit the common priors of vehicular data without interaction with a centralized coordinator. Specifically, vehicles perform local training during the driving journey, and simply upload its local model to roadside unit (RSU) encountered on the way. RSU's model will be updated accordingly and sent back to the vehicle via the V2R communication. We have theoretically shown that RSUs' models can eventually converge without a backhaul connection. Extensive experiments upon various road configurations demonstrate that the proposed scheme can efficiently train models among vehicles without dedicated Internet access and scale well with both the road range and vehicle density.
Wenchao Xu 0001, Haozhao Wang, Zhaoyi Lu 0001, Cunqing Hua, Nan Cheng 0001, Song Guo 0001
IEEE Trans. Mob. Comput.1
2024 Fast Packet Loss Inferring via Personalized Simulation-Reality Distillation
abstract
Packet loss inferring can enable a transceiver to distinguish between channel impairment and collision for transmission failures, and thus can improve the network performance by exclusively performing rate adaptation or adjusting the medium access parameter. Machine learning methods from literature have shown great potential in producing models that can detect the loss causes over various network trace, however haven't considered accurate data-driven loss inferring on resource-constrained devices that cannot accommodate deep models. In this paper, we propose a novel packet loss inferring framework that can train lightweight models to distinguish between channel losses and collisions by learning the data trace from both simulation and real devices. Specifically, we first train a sophisticated teacher model based on extensive simulation datasets, whose knowledge is then transferred to a small student model that can be deployed on tiny device. The simulation-reality distillation is conducted via personalized trace from each client correspondingly, whose performance bound is analytically guaranteed. We have implemented our method on real testbed and show that the network access performance can be significantly improved, especially for sudden network variations.
Wenchao Xu 0001, Haodong Wan, Haozhao Wang, Nan Cheng 0001, Quan Chen 0003, Song Guo 0001
IEEE Trans. Mob. Comput.1
2024 SR-FDIL: Synergistic Replay for Federated Domain-Incremental Learning
abstract
Federated Learning (FL) is to allow multiple clients to collaboratively train a model while keeping their data locally. However, existing FL approaches typically assume that the data in each client is static and fixed, which cannot account for incremental data with domain shift, leading to catastrophic forgetting on previous domains, particularly when clients are common edge devices that may lack enough storage to retain full samples of each domain. To tackle this challenge, we proposeFederatedDomain-IncrementalLearning viaSynergisticReplay (SR-FDIL), which alleviates catastrophic forgetting by coordinating all clients to cache samples and replay them. More specifically, when new data arrives, each client selects the cached samples based not only on their importance in the local dataset but also on their correlation with the global dataset. Moreover, to achieve a balance between learning new data and memorizing old data, we propose a novel client selection mechanism by jointly considering the importance of both old and new data. We conducted extensive experiments on several datasets of which the results demonstrate that SR-FDIL outperforms state-of-the-art methods by up to 4.05% in terms of average accuracy of all domains.
Yichen Li 0006, Wenchao Xu 0001, Yining Qi, Haozhao Wang, Ruixuan Li 0001, Song Guo 0001
IEEE Trans. Parallel Distributed Syst.2
2024 Peak AoI Minimization With Directional Charging for Data Collection at Wireless-Powered Network Edge
abstract
Age of Information (AoI) has emerged as a new metric to measure data freshness from the destination's perspective. To optimize the system AoI, most existing works focused on the point of scheduling of update transmissions. While at wireless-powered network edge, the source nodes can only transmit their updates after being charged ready, which means the system AoI is not only determined by the update transmission strategies, but also the charging strategies. Thus, in this paper, we investigate the first work to optimize the weighted peak AoI from the point of charging at wireless-powered network edge. Firstly, the problem of optimizing the weighted sum of average peak AoI with a directional charger is formulated, and then transformed to a charging time optimization problem with respect to the charging orientations and peak AoI, and an approximate algorithm is proposed to obtain the required charging time for each source node. Secondly, an age-based scheduling algorithm is proposed to compute the charging decisions and transmission decisions simultaneously, which can not only optimize the weighted sum of average peak AoI, but also guarantee the maximum peak AoI of each source node is bounded. The proposed algorithm is proved to have an approximation ratio of up to (1+$\varphi$), where$\varphi$is a small value related to the weight of each source node. When there exist multiple chargers, an approximate algorithm is also proposed to minimize the weighted sum of average peak AoI by scheduling the orientations of these chargers cooperatively. Finally, the extensive simulations demonstrate the high performance of the proposed algorithms in terms of peak AoI.
Quan Chen 0003, Song Guo 0001, Wenchao Xu 0001, Jing Li 0093, Kang Wei 0004, Zhipeng Cai 0001, Hong Gao 0001
IEEE Trans. Serv. Comput.3
2024 AI-Generated Content-Based Edge Learning for Fast and Efficient Few-Shot Defect Detection in IIoT
abstract
Generative AI has garnered substantial attention due to the limited defect samples in the industrial Internet of Things (IIoT). However, addressing the challenge of few-shot defect detection in industrial edge networks remains a key issue. In this paper, we propose ABEL, a novel AI-generated content (AIGC)-based edge learning framework for fast and efficient few-shot defect detection. This framework facilitates fast few-shot defect detection by harnessing the capabilities of realistic sample synthesis and edge-based AIGC task execution. Specifically, we propose an energy-based model (EBM)-guided Langevin Markov chain Monte Carlo (L-MCMC) image generation algorithm, synthesizing high-resolution industrial defect samples for efficient few-shot defect detection. Then, we formulate a large-scale mixed cooperative-competitive AIGC computation offloading problem and propose an attention and memory-based multi-agent reinforcement learning (AMMARL) algorithm to ensure fast edge execution of heterogeneous defect samples generative tasks. Particularly, the challenges of partial observability and high-dimensional state space are addressed by introducing multi-head attention mechanisms and long-term memory modules. Comprehensive synthesis experiments are conducted utilizing real-world industrial datasets NEU-CLS and DeepPCB. The experimental results demonstrate the effectiveness of our framework and algorithm's effectiveness in efficiently synthesizing realistic industrial defect images and optimizing edge-based AIGC task execution.
Siyuan Li 0005, Xi Lin 0003, Wenchao Xu 0001, Jianhua Li 0001
IEEE Trans. Serv. Comput.3
2024 Optimization-Driven DRL-Based Joint Beamformer Design for IRS-Aided ITSN Against Smart Jamming Attacks
abstract
This paper investigates an intelligent reflecting surfaces (IRS) aided anti-jamming communication strategy in the integrated terrestrial-satellite network (ITSN), where the IRS is exploited to mitigate jamming interference and enhance the integrated system communication performance. In such a network, the terrestrial network and satellite network are co-existing with a spectrum-sharing scheme in the presence of a multi-antenna jammer. We aim at maximizing the weighted sum rate (WSR) of all users by jointly optimizing the terrestrial beamformers and IRS phase shifts while considering the signal-to-interference-plus-noise ratio (SINR) requirements of legitimate users. Different from the non-convex optimization techniques utilized in the IRS-related problem, a novel optimization-driven deep reinforcement learning (DRL) algorithm is proposed, which leverages both the robustness of model-free learning approaches and the efficiency of model-based optimization methods. In the optimization module of the proposed algorithm, we analyze the smart jammer under the unknown jamming model and derive a lower bound of the anti-jamming uncertainty, such that the IRS-aided anti-jamming problem can be solved by alteration method with second-order cone programming (SOCP) algorithm and semidefinite relaxation (SDR) technique. Simulation results demonstrate that the IRS can enhance the anti-jamming performance efficiently, and the proposed optimization-driven DRL algorithm can improve both the learning rate and the system performance compared with existing solutions.
Cunqing Hua, Lingya Liu, Wenchao Xu 0001, Song Guo 0001
IEEE Trans. Wirel. Commun.4
2024 Optimal Power Control and CSI Acquisition for Over-the-Air Computation in OFDM System
abstract
Over-the-air computation (AirComp) is a novel technology that utilizes the superposition characteristic of the wireless multiple-access channel to accomplish communication and computation tasks simultaneously, which can be used to achieve efficient data fusion in wireless networks. However, the performance of AirComp can be compromised due to non-ideal conditions in practical systems, such as limited transmitting power budget, receiving noise, etc. In this paper, we first propose a joint transmitting-receiving power control scheme for over-the-air computation in the OFDM-based multicarrier wireless system, which can minimize the mean square error (MSE) of the received signal by taking into account of limited transmitting power budget and receiving noise. Based on the special structure of the problem, which depends on the set of users that either use up their power budget or not, we decompose the problem into two sub-problems, one deals with the power allocation at the transmitters, the other deals with the power scaling at the receiver. The optimal results are obtained by searching the set of users with used up power budget and solving these two sub-problems accordingly. We then propose an efficient channel state information (CSI) acquisition and feedback scheme for the AirComp power control scheme, and the effect of imperfect CSI is also considered accordingly. We provide extensive simulation results to demonstrate the performance of the proposed scheme under different network conditions.
Cunqing Hua, Jianan Hong, Wenchao Xu 0001
IEEE Trans. Wirel. Commun.4
2023 PMR: Prototypical Modal Rebalance for Multimodal Learning
abstract
Multimodal learning (MML) aims to jointly exploit the common priors of different modalities to compensate for their inherent limitations. However, existing MML methods often optimize a uniform objective for different modalities, leading to the notorious “modality imbalance” problem and counterproductive MML performance. To address the problem, some existing methods modulate the learning pace based on the fused modality, which is dominated by the better modality and eventually results in a limited improvement on the worse modal. To better exploit the features of multimodal, we propose Prototypical Modality Rebalance (PMR) to perform stimulation on the particular slow-learning modality without interference from other modalities. Specifically, we introduce the prototypes that represent general features for each class, to build the non-parametric classifiers for uni-modal performance evaluation. Then, we try to accelerate the slow-learning modality by enhancing its clustering toward prototypes. Furthermore, to alleviate the suppression from the dominant modality, we introduce a prototype-based entropy regularization term during the early training stage to prevent premature convergence. Besides, our method only relies on the representations of each modality and without restrictions from model structures and fusion methods, making it with great application potential for various scenarios. The source code is available here11https://github.com/fanyunfeng-bit/Modal-Imbalance-PMR.
Yunfeng Fan, Wenchao Xu 0001, Haozhao Wang, Song Guo 0001
CVPR2
2023 DaFKD: Domain-aware Federated Knowledge Distillation
abstract
Federated Distillation (FD) has recently attracted increasing attention for its efficiency in aggregating multiple diverse local models trained from statistically heterogeneous data of distributed clients. Existing FD methods generally treat these models equally by merely computing the average of their output soft predictions for some given input distillation sample, which does not take the diversity across all local models into account, thus leading to degraded performance of the aggregated model, especially when some local models learn little knowledge about the sample. In this paper, we propose a new perspective that treats the local data in each client as a specific domain and design a novel domain knowledge aware federated distillation method, dubbed DaFKD, that can discern the importance of each model to the distillation sample, and thus is able to optimize the ensemble of soft predictions from diverse models. Specifically, we employ a domain discriminator for each client, which is trained to identify the correlation factor between the sample and the corresponding domain. Then, to facilitate the training of the domain discriminator while saving communication costs, we propose sharing its partial parameters with the classification model. Extensive experiments on various datasets and settings show that the proposed method can improve the model accuracy by up to 6.02% compared to state-of-the-art baselines.
Haozhao Wang, Yichen Li 0006, Wenchao Xu 0001, Ruixuan Li 0001, Yufeng Zhan, Zhigang Zeng
CVPR3
2023 Non-Inducible RF Fingerprint Hiding via Feature Perturbation
abstract
Machine learning mechanisms are applied to detect the unique characteristics of the wireless interface or signaler that can distinguish one device's signal pattern from the others, which has been widely researched as the fingerprint for user identification. However, such fingerprinting can also be used for malicious purposes, i.e., identification tracking, undesired positioning, etc., as the unique features of the radio signal from a device is determined at the manufacturing stage and often cannot be easily removed afterward. To prevent privacy leakage from such radio frequency (RF) fingerprinting, in this paper, we propose an adversarial mechanism to hide the fingerprint whereby the device's identification cannot be induced by machine learning models from the preamble. Specially, we apply the adversarial attack method to attack the fingerprinting model by adding optimized adversarial perturbation to the preamble that can mislead the model classification results. To alleviate the adversarial sample's impact on communications and ensure the execution of packet detection at receivers, we improve the identification protection strategy with sparse perturbed features. In order to prevent further fingerprinting of re-training over the perturbed RF feature, we extend our method with the time-varying perturbations to further hide the device's identity. Extensive experiments are conducted, and we show that the proposed method can effectively hide the device identification from both the dedicated fingerprint model and the re-trained one from perturbed signals without disturbing the preamble functionality, which provides a gratifying confirmation of the proposed method.
Zhaoyi Lu 0001, Jiazhong Bao, Wenchao Xu 0001, Cunqing Hua
ICC4
2023 Towards Unbiased Training in Federated Open-world Semi-supervised Learning
abstract
Federated Semi-supervised Learning (FedSSL) has emerged as a new paradigm for allowing distributed clients to collaboratively train a machine learning model over scarce labeled data and abundant unlabeled data. However, existing works for FedSSL rely on a closed-world assumption that all local training data and global testing data are from seen classes observed in the labeled dataset. It is crucial to go one step further: adapting FL models to an open-world setting, where unseen classes exist in the unlabeled data. In this paper, we propose a novel Federatedopen-world Semi-Supervised Learning (FedoSSL) framework, which can solve the key challenge in distributed and open-world settings, i.e., the biased training process for heterogeneously distributed unseen classes. Specifically, since the advent of a certain unseen class depends on a client basis, the locally unseen classes (exist in multiple clients) are likely to receive differentiated superior aggregation effects than the globally unseen classes (exist only in one client). We adopt an uncertainty-aware suppressed loss to alleviate the biased training between locally unseen and globally unseen classes. Besides, we enable a calibration module supplementary to the global aggregation to avoid potential conflicting knowledge transfer caused by inconsistent data distribution among different clients. The proposed FedoSSL can be easily adapted to state-of-the-art FL methods, which is also validated via extensive experiments on benchmarks and real-world datasets (CIFAR-10, CIFAR-100 and CINIC-10).
Jie Zhang 0076, Xiaosong Ma, Song Guo 0001, Wenchao Xu 0001
ICML4
2023 AOCC-FL: Federated Learning with Aligned Overlapping via Calibrated Compensation
abstract
Federated Learning enables collaboratively model training among a number of distributed devices with the coordination of a centralized server, where each device alternatively performs local gradient computation and communication to the server. FL suffers from significant performance degradation due to the excessive communication delay between the server and devices, especially when the network bandwidth of these devices is limited, which is common in edge environments. Existing methods overlap the gradient computation and communication to hide the communication latency to accelerate the FL training. However, the overlapping can also lead to an inevitable gap between the local model in each device and the global model in the server that seriously restricts the convergence rate of learning process. To address this problem, we propose a new overlapping method for FL, AOCC-FL, which aligns the local model with the global model via calibrated compensation such that the communication delay can be hidden without deteriorating the convergence performance. Theoretically, we prove that AOCC-FL admits the same convergence rate as the non-overlapping method. On both simulated and testbed experiments, we show that AOCC-FL achieves a comparable convergence rate relative to the non-overlapping method while outperforming the state-of-the-art overlapping methods.
Haozhao Wang, Wenchao Xu 0001, Yunfeng Fan, Ruixuan Li 0001, Pan Zhou 0001
INFOCOM2
2023 Optimizing Average AoI with Directional Charging for Wireless-Powered Network Edge
abstract
Age of Information, which emerged as a new performance metric to quantify the data freshness, has drawn increasing interests recently. At wireless-powered network edge, the source nodes can only transmit their updates after being charged ready, which means the system AoI is not only determined by the update transmission strategies, but also the charging strategies. However, the existing works either only focused on the point of scheduling of update transmissions or have a rigid assumption that only one source node can be charged per time. Aiming at making the work more practical and general, we investigate the average$A$oI optimization problem at wireless-powered network edge without such limitations in this paper. Firstly, the lower bound of the weighted sum of average$A$oI of the whole network with a directional charger is analyzed, which is proved to be related to nodes' maximum transmitting interval and the charging strategy. An optimal charging time allocation algorithm is proposed to obtain the maximum transmitting interval of each source node by considering the overlapped areas of different charging orientations. After then, an AoI-aware periodical charging scheduling algorithm is proposed, which can obtain a periodical charging schedule including a charging period$T$and the charging orientation assigned to each time slot within$T$, while the average AoI is bounded. The proposed algorithm is proved to have an approximation ratio of up to 1.5625. Finally, the extensive simulations demonstrate the high performance of the proposed algorithm in terms of AoI.
Quan Chen 0003, Song Guo 0001, Wenchao Xu 0001, Jing Li 0093, Zhipeng Cai 0001, Hong Gao 0001
IWQoS3
2023 SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models
abstract
Test-time adaptation (TTA) is a special and practical setting in unsupervised domain adaptation, which allows a pre-trained model in a source domain to adapt to unlabeled test data in another target domain. To avoid the computation-intensive backbone fine-tuning process, the zero-shot generalization potentials of the emerging pre-trained vision-language models (e.g., CLIP, CoOp) are leveraged to only tune the run-time prompt for unseen test domains. However, existing solutions have yet to fully exploit the representation capabilities of pre-trained models as they only focus on the entropy-based optimization and the performance is far below the supervised prompt adaptation methods, e.g., CoOp. In this paper, we propose SwapPrompt, a novel framework that can effectively leverage the self-supervised contrastive learning to facilitate the test-time prompt adaptation. SwapPrompt employs a dual prompts paradigm, i.e., an online prompt and a target prompt that averaged from the online prompt to retain historical information. In addition, SwapPrompt applies a swapped prediction mechanism, which takes advantage of the representation capabilities of pre-trained models to enhance the online prompt via contrastive learning. Specifically, we use the online prompt together with an augmented view of the input image to predict the class assignment generated by the target prompt together with an alternative augmented view of the same image. The proposed SwapPrompt can be easily deployed on vision-language models without additional requirement, and experimental results show that it achieves state-of-the-art test-time adaptation performance on ImageNet and nine other datasets. It is also shown that SwapPrompt can even achieve comparable performance with supervised prompt adaptation methods.
Xiaosong Ma, Jie Zhang 0076, Song Guo 0001, Wenchao Xu 0001
NeurIPS4
2023 Service-Oriented Resource Allocation in SDN Enabled LEO Satellite Networks
abstract
As an integral component of space-air-ground integrated networks (SAGINs), the low Earth orbit (LEO) satellite networks have displayed immense potential in providing ubiquitous connectivity and broadband mobile communication. However, the intrinsic dynamics of LEO satellites poses unprecedented challenges in network management, multi-dimensional resource scheduling, and service delivery. In this paper, we study the service function chain (SFC) orchestration in dynamic LEO satellite networks, with the aim of achieving flexible and efficient service provision. Considering the service requirements and the load fairness of LEO satellite networks, we formulate the SFC deployment problem as an integer nonlinear programming (INLP) problem. We then introduce a load-aware SFC orchestration algorithm to improve serving capacity and load fairness. Additionally, we address the issue of SFC migration in dynamic LEO satellite networks to ensure service continuity. To minimize the service interruption and network resource wastes, a Tabu search (TS)-based approach is presented to optimize the virtual network function (VNF) migration. Simulation results demonstrate that our proposed approaches outperform the benchmark by a substantial margin in terms of load fairness, without compromising service acceptance.
Jingchao He, Nan Cheng 0001, Zhisheng Yin, Wenchao Xu 0001, Haixia Peng, Conghao Zhou, Ruqian Zhang
PIMRC5
2023 A Practical Fast Model Inference System Over Tiny Wireless Device
abstract
The utilization of machine learning models has become prevalent in various wireless devices to deduce the network status from various tracing data, e.g., link capacity, channel fading, etc. To cope with the increasing complexity of the network environment, deep neural models are leveraged to mine the high-dimensional network tracing data for a variety of intelligent applications. However, due to the limited resource that allocated to the network stack process, it is infeasible to train deep neural models due to the constrained computing power and absence of large-scale labeled data. Besides, the network device can barely support quick inference of large model, thus cannot support promptly response to the network conditions. In this paper, we propose a practical fast model inference system that can run high accuracy model over tiny wireless devices that are constrained in both memory and CPU power. Specifically, we design a knowledge-distillation based training method for a light-weight model that deployed at device side that can migrate the knowledge from a well-trained deep model. It is shown that our system can support fast model inference over tiny devices, which can greatly improve the network throughput in a multi-user access system by inferring the transmission collision from channel error, and thus can improve the accuracy of the link adaptation. We have conducted practical experiments to verify our system and discuss the possible extensions.
Wenchao Xu 0001, Haodong Wan, Nan Cheng 0001, Meng Qin 0001
PIMRC1
2023 Decompose, Then Reconstruct: A Framework of Network Structures for Click-Through Rate Prediction
Lang Lang, Zhenlong Zhu, Haozhao Wang, Ruixuan Li 0001, Wenchao Xu 0001
ECML/PKDD (1)6
2023 Joint Beamformer Design and User Scheduling for Integrated Terrestrial-Satellite Networks
abstract
The integrated terrestrial satellite networks (ITSNs) have been deemed as a promising solution to ubiquitous Internet access anytime and anywhere. In this paper, we investigate the spectrum sharing problem in ITSN, in particular focusing on the downlink transmission of the satellite network, which adopts the framing structure for the satellite users (SUs). We model the interference from both terrestrial downlink and uplink transmissions to SUs according to the beamforming techniques. For the terrestrial downlink transmission, we assume that terrestrial users (TUs) are served cooperatively by multiple small base stations (SBSs) via joint beamforming, while the virtual multiple access channel (VMAC) scheme is adopted for the terrestrial uplink transmission. We propose the optimization framework by jointly considering the terrestrial beamformer design and satellite user scheduling to maximize the sum rate of all users. The optimization problems are decomposed into three sub-problems: satellite user scheduling, terrestrial beamformer design, and time slot allocation, which are solved by deep clustering, second-order cone programming (SOCP) (or fractional programming (FP)), and linear programming, respectively. Then, an alternating iterative algorithm is designed to obtain the optimal solution. Simulation results are provided to demonstrate the effectiveness of the proposed algorithm in multiple cases.
Cunqing Hua, Lingya Liu, Wenchao Xu 0001, Song Guo 0001, Rahim Tafazolli
IEEE Trans. Wirel. Commun.4
2023 Intelligent Reflecting Surface-Aided Integrated Terrestrial-Satellite Networks
abstract
Intelligent reflecting surface (IRS) is a novel technology to manipulate wireless propagation channels via smart and controllable signal reflection. In this paper, we investigate an IRS-aided integrated terrestrial-satellite network (ITSN) system, where the IRS is deployed to assist the co-existing transmissions of the terrestrial small base stations (SBSs) and the satellite. Because of the spectrum sharing in the ITSN, the interference between the two systems should be carefully mitigated. Our objective is to maximize the weighted sum rate (WSR) of all users by jointly optimizing the frame-based coordinated transmit beamforming vectors at the SBSs, the phase shift matrix at the IRS, and the frame user scheduling, subject to SBSs’ individual power constraints and unit modulus constraints of phase shifters. To this end, we first adopt the agglomerative hierarchical clustering (AHC) method to schedule the satellite users to different frames. Then the block coordinate descent (BCD) algorithm is proposed, which alternately optimizes the transmit beamforming vectors and the reflective phase shift matrix. In particular, the optimal transmit beamforming vectors are obtained via the fractional programming (FP) technique. Meanwhile, two efficient algorithms, i.e., the Riemannian manifold (RM) and the successive convex approximation (SCA), are proposed for the phase shift optimization. Finally, simulation results are provided to demonstrate the performance gain of our schemes over other benchmark schemes.
Cunqing Hua, Lingya Liu, Wenchao Xu 0001, Rahim Tafazolli
IEEE Trans. Wirel. Commun.4
2022 Layer-wised Model Aggregation for Personalized Federated Learning
abstract
Personalized Federated Learning (pFL) not only can capture the common priors from broad range of distributed data, but also support customized models for heterogeneous clients. Researches over the past few years have applied the weighted aggregation manner to produce personalized models, where the weights are determined by calibrating the distance of the entire model parameters or loss values, and have yet to consider the layer-level impacts to the aggregation process, leading to lagged model convergence and inadequate personalization over non-IID datasets. In this paper, we propose a novel pFL training framework dubbed Layer-wised Personalized Federated learning (pFedLA) that can discern the importance of each layer from different clients, and thus is able to optimize the personalized model aggregation for clients with heterogeneous data. Specifically, we employ a dedicated hyper-network per client on the server side, which is trained to identify the mutual contribution factors at layer granularity. Meanwhile, a parameterized mechanism is introduced to update the layer-wised aggregation weights to progressively exploit the inter-user similarity and realize accurate model personalization. Extensive experiments are conducted over different models and learning tasks, and we show that the proposed methods achieve significantly higher performance than state-of-the-art pFL methods.
Xiaosong Ma, Jie Zhang 0076, Song Guo 0001, Wenchao Xu 0001
CVPR4
2022 AoI Minimization Charging at Wireless-Powered Network Edge
abstract
Age of Information (AoI) has emerged as a new metric to measure data freshness from the destination’s perspective. The problem of optimizing AoI has been attracting extensive interests recently. However, existing works mainly focused on scheduling data transmission for AoI optimization. While at wireless-powered network edge, the charging plan of source nodes also requires to be computed in advance, which means the system AoI is determined by not only the data transmission decision but also the charging plan. Thus, in this paper, we investigate the first work to optimize the weighted peak AoI from the point of charging at wireless-powered network edge with a directional charger. Firstly, to minimize the weighted sum of average peak AoI, the AoI minimization problem is transformed to a charging time optimization problem with respect to the overlapped charging areas and average peak AoI, and an approximate algorithm is proposed to obtain the required charging time for each source node. Then, an age-based scheduling algorithm is proposed to compute the charging and data transmission decisions for each source node simultaneously, which can not only optimize the weighted sum of average peak AoI but also guarantee the maximum peak AoI for each source node. The proposed algorithm is proved to have an approximation ratio of up to (1+φ), where φ is a much smaller value related to the weight of each source node. Finally, the simulation results verify the high performance of proposed algorithms in terms of average and maximum peak AoI.
Quan Chen 0003, Song Guo 0001, Wenchao Xu 0001, Zhipeng Cai 0001, Lianglun Cheng, Hong Gao 0001
ICDCS3
2022 Sustainable Federated Learning with Long-term Online VCG Auction Mechanism
abstract
Federated learning (FL) clients may be reluctant to participate in the energy-consuming FL unless they are incentivized. Existing incentive mechanisms seldom consider the economic properties, e.g., social welfare, individual rationality and incentive compatibility, which significantly limits the sustainability of FL to attract more clients. The Vickrey–Clarke–Groves (VCG) auction is an ideal mechanism for simultaneously guaranteeing all crucial economic properties to maximize social welfare. However, VCG auction cannot be applied directly to FL scenarios due to the following challenges: 1) It requires precise analytical derivation of the optimal strategy, which is unavailable due to the inherent model-unknown and privacy-sensitive characteristics of FL. 2) Current auction modeling decomposes the entire process into multiple independent rounds and solves them one-by-one, which breaks the successive correlation between rounds in the long-term training process of FL. To overcome these challenges, this paper presents a long-term online VCG auction mechanism for FL that employs an experience-driven deep reinforcement learning algorithm to obtain the optimal strategy. Besides, we extend long-term forms of the crucial economic properties for the successive FL process. Furthermore, knowledge transfer is applied to reduce the excessive training overhead arising from the VCG payment rules. By exploiting the environmental similarity among sub-auctions, we develop the strategy sharing to significantly cut the training time by half. Finally, we theoretically prove the extended economic properties and conduct extensive experiments on multiple real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL increases by 36% with a 37% reduction in payment.
Leijie Wu, Song Guo 0001, Yi Liu 0057, Zicong Hong, Yufeng Zhan, Wenchao Xu 0001
ICDCS6
2022 Fusing Onboard Modalities with V2V Information for Autonomous Driving
abstract
Status quo autonomous driving mechanisms rely on fusing the multimodal sensing data to integrate the information from onboard units of a vehicle, e.g., lidar, camera, etc., and have yet to consider the information obtained via the inter-vehicle communication, such as the status of neighboring peers. In this paper, we consider to integrate not only the local onboard sensing data, but also the neighboring vehicle information from the vehicle-to-vehicle (V2V) data pipe, which is demonstrated to improve the autonomous driving performance significantly. Specifically, the opportunistic V2V messages are input to a transformer based fusing framework to improve the driving accuracy in both short and long routes in CARLA environment. Unlike previous rule-based mechanisms of dealing with the V2V messages, to the best of our knowledge, the proposed method is the first to integrate the V2V data to the neural network which implicitly induce the waypoints for accurate end-to-end autonomous driving. We conduct extensive experiments, whose results well demonstrate the utility of the V2V information, and can provide useful inspirations for future driving system design.
Haodong Wan, Wenchao Xu 0001, Nan Cheng 0001, Zhisheng Yin
VTC Spring2
2022 Design of Low-latency Overlay Protocol for Blockchain Delivery Networks
abstract
A major concern of blockchain systems is to scale up their throughput. Many improvements and novel consensus protocols have been proposed to address this issue, but they are intrinsically limited by the message synchronization latency of the underlying peer-to-peer (P2P) network. Most existing implementations of blockchain systems are based on the unstructured random overlay disseminating networks, which often results in a heavy-tailed delivery latency distribution, impairing the decentralization property of the blockchain system. To overcome these constraints, this research proposes Urocissa, a structured overlay protocol to reduce the delivery latency and to improve steadiness for blockchain systems. By exploiting the unique characteristics of blockchain traffic and network heterogeneity, the protocol maintains multiple minimum latency broadcasting trees in a distributed way, whereby each node communicates with its neighbors and makes decisions sovereignly to balance relaying tasks among participants. Experiments show that the proposed protocol significantly reduces block delivery and confirmation latency compared to the conventional blockchain delivery network protocols.
Yiqing Zhu, Cunqing Hua, Dingjie Zhong, Wenchao Xu 0001
WCNC4
2022 Gradient Scheduling With Global Momentum for Asynchronous Federated Learning in Edge Environment
abstract
Federated Learning has attracted widespread attention in recent years because it allows massive edge nodes to collaboratively train machine learning models without sharing their private data sets. However, these edge nodes are usually heterogeneous in computational capability and statistically different in data distribution, i.e., non-independent and identically distributed (IID), leading to significant performance degradation. Although status quo asynchronous training methods can solve the heterogeneity issue, they cannot prevent the non-IID problem from reducing the convergence rate. In this article, we propose a novel paradigm that schedules the gradient with partially averaged gradients and applies the global momentum (GSGM) for asynchronous training over non-IID data sets in an edge environment. Our key idea is to apply global momentum and partial average on the biased gradients calculated on edge nodes after scheduling, to make the training process stable. Empirical results demonstrate that GSGM can well adapt to different degrees of non-IID data and bring 20% performance gains in terms of training stability for popular optimization algorithms with enhanced accuracy over Fashion-Mnist and CIFAR-10 data sets.
Haozhao Wang, Ruixuan Li 0001, Pan Zhou 0001, Yuhua Li 0003, Wenchao Xu 0001, Song Guo 0001
IEEE Internet Things J.6
2022 A Comprehensive Survey on Training Acceleration for Large Machine Learning Models in IoT
abstract
The ever-growing artificial intelligence (AI) applications have greatly reshaped our world in many areas, e.g., smart home, computer vision, natural language processing, etc. Behind these applications are usually machine learning (ML) models with extremely large size, which require huge data sets for accurate training to mine the value contained in the big data. Large ML models, however, can consume tremendous computing resources to achieve decent performance and thus, it is difficult to train them in resource-constrained Internet of Things (IoT) environments, which would prevent further development and application of AI techniques in the future. To deal with such challenges, there are many efforts on accelerating the training process for large ML models in IoT. In this article, we provide a comprehensive review on the recent advances toward reducing the computing cost during the training stage while maintaining comparable model accuracy. Specifically, the optimization algorithms that aim to improve the convergence rate are emphasized over various distributed learning architectures that exploit ubiquitous computing resources. Then, the article elaborates the computation hardware acceleration and communication optimization for collaborative training among multiple learning entities. Finally, the remaining challenges, future opportunities, and possible directions are discussed.
Haozhao Wang, Zhihao Qu, Qihua Zhou, Haobo Zhang 0002, Boyuan Luo, Wenchao Xu 0001, Song Guo 0001, Ruixuan Li 0001
IEEE Internet Things J.6
2022 PS+: A Simple yet Effective Framework for Fast Training on Parameter Server
abstract
In distributed training, workers collaboratively refine the global model parameters by pushing their updates to the Parameter Server and pulling fresher parameters for the next iteration. This introduces high communication costs for training at scale, and incurs unproductive waiting time for workers. To minimize the waiting time, existing approachesoverlap communication and computationfor deep neural networks. Yet, these techniques not only require the layer-by-layer model structures, but also need significant efforts in runtime profiling and hyperparameter tuning. To make the overlapping optimizationsimpleandgeneric, in this article, we propose a new Parameter Server framework. Our solutiondecouplesthe dependency between push and pull operations, and allows workers toeagerlypull the global parameters. This way, both push and pull operations can be easily overlapped with computations. Besides, the overlapping manner offers a different way to address the straggler problem, where the stale updates greatly retard the training process. In the new framework, with adequate information available to workers, they can explicitly modulate the learning rates for their updates. Thus, the global parameters can be less compromised by stale updates. We implement a prototype system in PyTorch and demonstrate its effectiveness on both CPU/GPU clusters. Experimental results show that our prototype saves up to 54% less time for each iteration and up to 37% fewer iterations for model convergence, achieving up to 2.86× speedup over widely-used synchronization schemes.
A-Long Jin, Wenchao Xu 0001, Song Guo 0001, Bing Hu 0002, Kwan Lawrence Yeung
IEEE Trans. Parallel Distributed Syst.2
2021 Weighted Sum-Rate Maximization for Multi-IRS Aided Integrated Terrestrial-Satellite Networks
abstract
This paper investigates a multiple intelligent reflecting surfaces (IRSs) aided integrated terrestrial-satellite network (ITSN), where the IRSs are deployed to cooperatively assist the low channel gain users in the co-existing transmission system. In such a network, the coordinated beamforming and frame based transmission scheme are considered for the terrestrial network and the satellite network, respectively. We aim at maximizing the weighted sum rate (WSR) of all users by jointly designing the frame based beamforming at the small base stations (SBSs) and the phase shifts at the IRSs, subject to the individual maximum SBS's transmit power constraints and the IRSs' reflection constraints. This non-convex problem is firstly decomposed via fractional programming (FP) technique in the objective function, then transmit beamforming vectors and reflective phase shifts matrix are optimized alternatingly. A block coordinate descent (BCD) method is proposed to obtain the stationary solution. Simulation results verify the effectiveness of the proposed algorithm compared with different benchmark schemes.
Cunqing Hua, Lingya Liu, Wenchao Xu 0001, Rahim Tafazolli
GLOBECOM4
2021 Towards Integrated Terrestrial-Satellite Network via Intelligent Reflecting Surface
abstract
This paper investigates an intelligent reflecting surface (IRS)-aided integrated terrestrial-satellite network (ITSN) system for the low channel gain users, where an IRS is deployed to assist the co-existing transmissions of the terrestrial small base stations (SBSs) and the satellite. Because of the spectrum sharing in the ITSN, the interference between two systems should be carefully mitigated. We aim for maximizing the weighted sum rate (WSR) of all users through jointly optimizing the frame based coordinated transmit beamforming vectors at the SBSs and the phase shift matrix at the IRS, and the frame user scheduling subject to each SBS's power and unit modulus. To this end, we propose efficient algorithms based on alternating optimization, in which the transmit beamforming vectors and reflective phase shifts matrix are optimized in an alternating manner. In particular, we develop the second-order-cone programming (SOCP) for optimizing the coordinated transmit beamforming and propose the Riemannian conjugate gradient (RCG) for updating the reflecting shifts. For frame user scheduling, we propose the chordal distance measure method to improve the intra-fame correlation. Simulation results verify the effectiveness of the proposed algorithm compared with different benchmark schemes.
Cunqing Hua, Lingya Liu, Wenchao Xu 0001
ICC4
2021 Parameterized Knowledge Transfer for Personalized Federated Learning
abstract
In recent years, personalized federated learning (pFL) has attracted increasing attention for its potential in dealing with statistical heterogeneity among clients. However, the state-of-the-art pFL methods rely on model parameters aggregation at the server side, which require all models to have the same structure and size, and thus limits the application for more heterogeneous scenarios. To deal with such model constraints, we exploit the potentials of heterogeneous model settings and propose a novel training framework to employ personalized models for different clients. Specifically, we formulate the aggregation procedure in original pFL into a personalized group knowledge transfer training algorithm, namely, KT-pFL, which enables each client to maintain a personalized soft prediction at the server side to guide the others' local training. KT-pFL updates the personalized soft prediction of each client by a linear combination of all local soft predictions using a knowledge coefficient matrix, which can adaptively reinforce the collaboration among clients who own similar data distribution. Furthermore, to quantify the contributions of each client to others' personalized training, the knowledge coefficient matrix is parameterized so that it can be trained simultaneously with the models. The knowledge coefficient matrix and the model parameters are alternatively updated in each round following the gradient descent way. Extensive experiments on various datasets (EMNIST, Fashion_MNIST, CIFAR-10) are conducted under different settings (heterogeneous models and data distributions). It is demonstrated that the proposed framework is the first federated learning paradigm that realizes personalized model training via parameterized group knowledge transfer while achieving significant performance gain comparing with state-of-the-art algorithms.
Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Haozhao Wang, Wenchao Xu 0001, Feijie Wu
NeurIPS5
2021 Joint Resource Allocation and User Scheduling Scheme for Federated Learning
abstract
This paper investigates the impact of communication factors on the convergence performance of federated learning (FL) in wireless networks. Considering the limited communication resources in wireless networks, it is difficult to schedule all users to participate in a comprehensive training and the convergence performance of training model relies much on the user scheduling scheme. To minimize the maximum update delay of user training, we propose a joint resource allocation and user scheduling scheme in this paper. Particularly, the user communication delay and user training results are jointly considered to dynamically schedule users and allocate communication resources. Simulation results show that the convergence time can be reduced by 41.6% compared with the random scheduling allocation scheme.
Jinglong Shen, Nan Cheng 0001, Zhisheng Yin, Wenchao Xu 0001
VTC Fall4
2021 Service-Oriented Energy-Latency Tradeoff for IoT Task Partial Offloading in MEC-Enhanced Multi-RAT Networks
abstract
The development of the 5G network is envisioned to offer various types of services like virtual reality/augmented reality and autonomous vehicles applications with low-latency requirements in Internet-of-Things (IoT) networks. Mobile-edge computing (MEC) has become a promising solution for enhancing the computation capacity of mobile devices at the edge of the network in a 5G wireless network. Additionally, multiple radio access technologies (multi-RATs) have been verified with the potential in lowering the transmission latency and energy consumption, while improving the Quality of Services (QoS). Benefiting from the cooperation of multi-RATs, large latency-sensitive computing service tasks (L2SC) can be offloaded by different RATs simultaneously, which has great practical significance for data partitioned oriented applications with large task sizes. In this article, to enhance the L2SC offloading services for satisfying low-latency requirements with low energy consumption, we investigate the energy-latency tradeoff problem for partial task offloading in the MEC-enhanced multi-RAT network, considering the limitation of energy and computing in capability-constrained end devices in IoT networks. Specifically, we formulated the L2SC task computation offloading problem to minimize the weighted sum of the latency cost and the energy consumption by jointly optimizing the local computing frequency, task splitting, and transmit power, while guaranteeing the stringent latency requirement and the residual energy constraint. Due to the nonsmoothness and nonconvexity of the formulated problem with high complexity, we convert the tradeoff problem into a smooth biconvex problem and propose an alternate convex search-based algorithm, which can greatly reduce the computational complexity. Numerical simulation results show the effectiveness of the proposed algorithm with various performance parameters.
Meng Qin 0001, Nan Cheng 0001, Zewei Jing, Tingting Yang 0001, Wenchao Xu 0001, Qinghai Yang, Ramesh R. Rao
IEEE Internet Things J.5
2021 Space-air-ground integrated networks for future IoT: Architecture, management, service and performance
Feng Lyu 0001, Wenchao Xu 0001, Quan Yuan 0004, Katsuya Suto
Peer-to-Peer Netw. Appl.2
2021 Towards Rear-End Collision Avoidance: Adaptive Beaconing for Connected Vehicles
abstract
Connected vehicles have been considered as an effective solution to enhance driving safety as they can be well aware of nearby environments by exchanging safety beacons periodically. However, under dynamic traffic conditions, especially for dense-vehicle scenarios, the naive beaconing scheme where vehicles broadcast beacons at a fixed rate with a fixed transmission power can cause severe channel congestion and thus degrade the beaconing reliability. In this paper, by considering the kinematic status and beaconing rate together, we study the rear-end collision risk and define a danger coefficient ρ to capture the danger threat of each vehicle being in the rear-end collision. In specific, we propose a fully distributed adaptive beacon control scheme, called ABC, which makes each vehicle actively adopt a minimal but sufficient beaconing rate to avoid the rear-end collision in dense scenarios based on individually estimated ρ. With ABC, vehicles can broadcast at the maximum beaconing rate when the channel medium resource is enough and meanwhile keep identifying whether the channel is congested. Once a congestion event is detected, an NP-hard distributed beacon rate adaptation (DBRA) problem is solved with a greedy heuristic algorithm, in which a vehicle with a higher ρ is assigned with a higher beaconing rate while keeping the total required beaconing demand lower than the channel capacity. We prove the heuristic algorithm's close proximity to the optimal result and thoroughly analyze the communication overhead of ABC scheme. By using Simulation of Urban MObility (SUMO)-generated vehicular traces, we conduct extensive simulations to demonstrate the efficacy of our proposed ABC scheme. Simulation results show that vehicles can adapt beaconing rates according to the driving safety demand, and the beaconing reliability can be guaranteed even under high-dense vehicle scenarios.
Feng Lyu 0001, Nan Cheng 0001, Hongzi Zhu, Wenchao Xu 0001, Minglu Li 0001, Xuemin Shen
IEEE Trans. Intell. Transp. Syst.5
2020 Proactive Link Adaptation for Marine Internet of Things in TV White Space
abstract
By connecting the maritime users to Internet, e.g., boats, ships, etc., it is possible to operate maritime sensing and informatics across seas and oceans. Such marine Internet of things (MIoT) is urging intelligent maritime applications, e.g., real-time vessel tracking, navigation safety, autonomous shipping, etc. Due to the bandwidth limitation of conventional marine channels, broadband communication is desired for these emerging applications. In this paper, we consider operating the TV white space (TVWS) spectrum in 700MHz to support the near-sea surface communication for MIoT terminals. To better utilize the TV channel capacity, we propose a proactive and efficient link adaptation (LA) scheme based on nonlinear autoregressive neural network (NARNN) time series prediction. Specifically, the historical signal samplings are used to predict the near-sea-surface channel link status for the next transmission slot, which is then used to select a proper modulation and coding scheme (MCS) for the next egress frame. We have conducted extensive simulations, and show that the average channel utility can achieve almost 85% of the optimal capacity. The proposed LA scheme can provide useful inspirations for applying data analytics to efficient and adaptive LA schemes for mobile Internet of things.
Wenchao Xu 0001, Tingting Yang 0001, Huaqing Wu, Song Guo 0001
ICC1
2020 Autonomous Rate Control for Mobile Internet of Things: A Deep Reinforcement Learning Approach
abstract
With the ubiquitous deployment of mobile sensors and smart devices, the scope of Internet of things (IoT) has extended to the space of mobile networks, where IoT terminals are moving around instead of being fixed in buildings, ground infrastructures, etc. In this paper, we consider such mobile Internet of things (MIoT), and propose an autonomous rate control (RC) scheme for the uplink transmission from MIoT terminals to access stations. A deep reinforcement learning (DRL) based approach is designed to capture the channel variations of the link and to improve the effectiveness of the rate selection for each egress frame. Extensive simulations are conducted for MIoT terminals including vehicles and UAVs and show significant throughput performance improvement comparing with traditional methods, as well as the robustness and scalability of the DRL-RC algorithm. The proposed DRL-RC can provide inspirations for efficient and scalable link adaptation schemes for MIoT terminals.
Wenchao Xu 0001, Nan Cheng 0001, Ning Lu 0001, Lijuan Xu 0002, Meng Qin 0001, Song Guo 0001
VTC Fall1
2020 Optimal power allocation for non-orthogonal multiple access in wireless backhaul networks
abstract
In this study, the authors adopt the non‐orthogonal multiple access (NOMA) technique to improve the spectrum efficiency in the wireless backhaul networks, whereby the downlink and uplink NOMA techniques are applied for the backhaul and access links, respectively. Due to the coupling between the backhaul and access transmission stages, the transmission power should be carefully allocated in both stages so that the overall throughput can be maximised. They start with the single user equipment (UE) case and consider different scenarios and analyse the tradeoff between the access and backhaul links, and the optimal power allocation solutions are obtained accordingly. They then extend the analysis to the multi‐UE case and formulate the optimal power allocation problem, which is solved using the Lagrangian dual decomposition algorithm. Simulation results demonstrate that the proposed schemes are effective in improving the throughput and outperforms the conventional orthogonal multiple access technique under different network settings.
Xiaoqi Yang 0004, Cunqing Hua, Wenchao Xu 0001, Pengwenlong Gu
IET Commun.3
2020 Control Channel Anti-Jamming in Vehicular Networks via Cooperative Relay Beamforming
abstract
In vehicular networks, radio-frequency (RF) jamming attacks are considered a major threat to the availability of control channel (CCH). In particular, vehicles may not be able to receive control messages from roadside units (RSUs) due to persistent interference in the CCH, which may claim human lives and result in significant economic losses. In this article, a cooperative anti-jamming beamforming scheme is proposed to address the CCH jamming problems in vehicular networks. This scheme utilizes spatial diversity provided by the multiantenna RSU and relay vehicles to improve the transmission reliability of downlink control messages. In addition, to address the additive effects of the jamming signals and the intergroup interference, the relay selection problem and the beamformer design problem are jointly considered, which is modeled as a mixed-integer nonlinear programming (MINLP) problem. Then, we address this challenging problem by relaxing it into a series of convex subproblems via the semi-definite relaxation (SDR) and convex-concave process (CCP) methods, and then propose to solve these convex subproblems iteratively. The simulation results show that our proposed method convergences rapidly, and compared to the benchmark schemes, significant performance gains can be observed.
Pengwenlong Gu, Cunqing Hua, Wenchao Xu 0001, Rida Khatoun, Yue Wu 0010, Ahmed Serhrouchni
IEEE Internet Things J.3
2020 Enabling Security-Aware D2D Spectrum Resource Sharing for Connected Autonomous Vehicles
abstract
With the emergence of automated driving technology, wireless demand for secure information exchange among automated vehicles has increased dramatically. To this end, we design a security-aware dynamic device-to-device (D2D) spectrum resource sharing mechanism to enhance the security of vehicular D2D communications with the improved spectrum efficiency. Considering both resource block (RB) sharing and power control, we model this joint D2D spectrum resource sharing process as a weighted bipartite graph matching problem whose weights are obtained through deriving the closed-form solutions of power control in an algebraic method. Then, the global optimal solutions of the matching problem are obtained by using the Hungarian algorithm. Furthermore, a security-aware RB and power allocation (SA-RBPA) mechanism compatible with existing cellular networks is proposed for the small cell base station applications. Extensive simulation results have shown that, compared with existing approaches that consider RB reusing strategy or power control scheme optimization alone, the SA-RBPA scheme is able to realize a better spectrum efficiency and security performance.
Xuesen Peng, Bo Qian 0001, Kai Yu 0010, Feng Lyu 0001, Wenchao Xu 0001
IEEE Internet Things J.6
2020 Vehicular Networking-Enabled Vehicle State Prediction via Two-Level Quantized Adaptive Kalman Filtering
abstract
The accurate prediction of vehicle state based on the data acquired by the vehicular networking system plays an important role in improving traffic safety in the transportation section. However, it is difficult to accurately predict the vehicle state due to the highly dynamic road environment and various drivers' behaviors. To this end, in this article, we propose a two-level quantized adaptive Kalman filter (KF) algorithm based on the autoregressive moving average (MA) model to predict the vehicle state (including the moving direction, driving lane, vehicle speed, and acceleration). First, we propose a vehicular networking system to acquire the vehicle data by exchanging traffic data between the onboard unit and the roadside unit (RSU). Then, we predict the vehicle state at the edge cloud server (ECS) equipped at the RSU. Specifically, we utilize the autoregressive MA model to predict vehicle acceleration at the next moment. Then, the predicted vehicle acceleration is used as an input variable of the adaptive KF model to predict the vehicle location and speed at the next moment, in which we quantify the predicted vehicle location to the moving direction and the driving lane. Finally, the ECS broadcasts the predicted state to other RSUs. Through the communication with the road unit, all vehicles moving at the intersection can share vehicles states each other. In this doing, we can efficiently improve traffic safety in the intersection. We provide numerical simulations to validate the effectiveness of the autoregressive MA model used for predicting acceleration. Then, we evaluate the efficiency of the proposed two-level quantized adaptive KF algorithm. Compared with five conventional prediction algorithms, our proposed algorithm can improve the speed prediction accuracy by 90.62%, 89.81%, 88.91%, 82.76%, and 70.77%, respectively, which implies that our algorithm is a promising scheme for predicting the vehicle state in vehicular networks.
Li Ping Qian 0001, Anqi Feng, Ningning Yu, Wenchao Xu 0001, Yuan Wu 0001
IEEE Internet Things J.4
2020 Augmenting Drive-Thru Internet via Reinforcement Learning-Based Rate Adaptation
abstract
Drive-thru Internet has been considered as an effective Internet access method for Internet of Vehicles (IoV). Through the opportunistic vehicle-to-roadside WiFi connection, it can provide high throughput performance with low communication cost for IoV applications, such as intelligent transportation system, automotive infotainment, etc. However, its usability is highly affected by a fundamental issue called rate adaptation (RA), which is to adjust the modulation and coding rate to adapt to the dynamic wireless channel between the vehicle and the roadside access point (AP). Conventional WiFi RA schemes are designed for indoor or quasistatic scenarios and do not account for the channel variations in drive-thru Internet. In this article, we study the limitation of applying existing RA schemes in drive-thru Internet and propose a reinforcement learning (RL)-based RA scheme to capture the potential channel variation patterns and efficiently select the rate for every vehicle's egress frame. Simulation results demonstrate that the proposed RA scheme outperforms the existing schemes in network throughput and that the efficiency of the learning model can be generalized under various conditions. The proposed RA method can provide useful inspirations for designing robust and scalable link adaptation protocols in IoV.
Wenchao Xu 0001, Song Guo 0001, Shiheng Ma, Weihua Zhuang
IEEE Internet Things J.1
2020 Evolutionary V2X Technologies Toward the Internet of Vehicles: Challenges and Opportunities
abstract
To enable large-scale and ubiquitous automotive network access, traditional vehicle-to-everything (V2X) technologies are evolving to the Internet of Vehicles (IoV) for increasing demands on emerging advanced vehicular applications, such as intelligent transportation systems (ITS) and autonomous vehicles. In recent years, IoV technologies have been developed and achieved significant progress. However, it is still unclear what is the evolution path and what are the challenges and opportunities brought by IoV. For the aforementioned considerations, this article provides a thorough survey on the historical process and status quo of V2X technologies, as well as demonstration of emerging technology developing directions toward IoV. We first review the early stage when the dedicated short-range communications (DSRC) was issued as an important initial beginning and compared the cellular V2X with IEEE 802.11 V2X communications in terms of both the pros and cons. In addition, considering the advent of big data and cloud-edge regime, we highlight the key technical challenges and pinpoint the opportunities toward the big data-driven IoV and cloud-based IoV, respectively. We believe our comprehensive survey on evolutionary V2X technologies toward IoV can provide beneficial insights and inspirations for both academia and the IoV industry.
Wenchao Xu 0001, Wei Wang 0100
Proc. IEEE2
2020 Characterizing Urban Vehicle-to-Vehicle Communications for Reliable Safety Applications
abstract
The IEEE 802.11p-based dedicated short range communication (DSRC) is essential to enhance driving safety and improve road efficiency by enabling rapid cooperative message exchanging. However, there is a lack of good understanding on the DSRC performance in urban environments for vehicle-to-vehicle (V2V) communications, which impedes its reliable and efficient application. In this paper, we first conduct intensive data analytics on V2V performance, based on a large amount of real-world DSRC communications trace collected in Shanghai city, and obtain several key insights as follows. First, among many context factors, the non-line-of-sight (NLoS) link condition is the major factor degrading V2V performance. Second, the durations of line-of-sight (LoS) and NLoS transmission conditions follow power law distributions, which indicate that the probability of experiencing long LoS/NLoS conditions both could be high. Third, the packet inter-reception (PIR) time distribution follows an exponential distribution in the LoS conditions but a power law in the NLoS conditions, which means that the consecutive packet reception failures rarely appear in the LoS conditions but can constantly appear in the NLoS conditions. Based on these findings, we propose a context-aware reliable beaconing scheme, called CoBe, to enhance the broadcast reliability for safety applications. The CoBe is a fully distributed scheme, in which a vehicle first detects the link condition with each of its neighbors by machine learning algorithms, then exchanges such link condition information with its neighbors, and finally selects the minimal number of helper vehicles to rebroadcast its beacons to those neighbors in bad link condition. To analyze and evaluate the CoBe performance, a two-state Markov chain is devised to model beaconing behaviors. The extensive trace-driven simulations are conducted to demonstrate the efficacy of CoBe.
Feng Lyu 0001, Hongzi Zhu, Nan Cheng 0001, Wenchao Xu 0001, Minglu Li 0001, Xuemin Shen
IEEE Trans. Intell. Transp. Syst.5
2020 Delay-Minimized Edge Caching in Heterogeneous Vehicular Networks: A Matching-Based Approach
abstract
To enable ever-increasing vehicular applications, heterogeneous vehicular networks (HetVNets) are recently emerged to provide enhanced and cost-effective wireless network access. Meanwhile, edge caching is imperative to future vehicular content delivery to reduce the delivery delay and alleviate the unprecedented backhaul pressure. This work investigates content caching in HetVNets where Wi-Fi roadside units (RSUs), TV white space (TVWS) stations, and cellular base stations are considered to cache contents and provide content delivery. Particularly, to characterize the intermittent network connection provided by Wi-Fi RSUs and TVWS stations, we establish an on-off model with service interruptions to describe the content delivery process. Content coding then is leveraged to resist the impact of unstable network connections with optimized coding parameters. By jointly considering file characteristics and network conditions, we minimize the average delivery delay by optimizing the content placement, which is formulated as an integer linear programming (ILP) problem. Adopting the idea of student admission model, the ILP problem is then transformed into a many-to-one matching problem and solved by our proposed stable-matching-based caching scheme. Simulation results demonstrate that the proposed scheme can achieve near-optimal performances in terms of delivery delay and offloading ratio with low complexity.
Huaqing Wu, Wenchao Xu 0001, Nan Cheng 0001, Weisen Shi, Li Wang 0039, Xuemin Shen
IEEE Trans. Wirel. Commun.3
2019 A Queueing Analysis of the Opportunistic Vehicle-to-Vehicle Communication
abstract
The Vehicle-to-Vehicle (V2V) communication is critical in vehicular networks, besides the fact that it can enable the message sharing among neighboring vehicles, it can also help to relay neighbors' data content to remote destinations during the opportunistic V2V contacting period. The efficiency of V2V data pipe is affected drastically by the intermittent V2V connection, which requires a buffer queue to cache the data contents for random disruption and resuming of the transmission. In this paper, we propose a queueing theoretical framework to analysis the effectiveness of the opportunistic V2V communication, which captures the connection interruption due to the vehicle mobility and trades off between the V2V transmission effectiveness and average data service delay. An M/G/1/K queue model with service interruption is setup to analysis the relationship between how many data tasks can be fulfilled via V2V data pipe and the average time to accomplish the data service. We also consider two different resuming types, namely, resuming from beginning and resuming from break point. The proposed queueing framework is validated via simulation based on PTV VISSIM trace, which can provide guideline in design and development of future V2V communication protocols.
Wenchao Xu 0001, Song Guo 0001, Shiheng Ma
GLOBECOM1
2019 Efficient Rate Adaptation for 802.11af TVWS Vehicular Access via Deep Learning
abstract
Enabling connected vehicles is becoming an essential demand for modern mobility world. To envision the Internet of Vehicles (IoVs) and better support the transmission of the explosive vehicular data, the Ultra High Frequency(UHF) spectrum resource in TV White Space (TVWS) band is re-utilized to provide cognitively wireless access for vehicle users. The TVWS access can provide high throughput and wide coverage due to its wide bandwidth and penetrability. However, there remains a critical issue for vehicles to dynamically change the link rate for the egress frames to adapt to the channel variance. In this paper, we investigate deep learning based rate adaptation (RA) scheme for the vehicle users accessing to the TVWS band. We utilize a series of the recent Signal-to-Noise (SNR) records, and modeled the RA as a Time Series Classification (TSC) problem, which is solved by Deep Learning (DL) models to classifies the collected SNR series to the optimal rate selection for the frame to be transmitted. We compared three different DL models and show that the performance of the DL based RAs outperform conventional RAs. Such results could provide insightful guidance for applying machine learning in the RA problem for wireless vehicular access.
Wenchao Xu 0001, Song Guo 0001
GLOBECOM1
2018 Reinforcement Learning Policy for Adaptive Edge Caching in Heterogeneous Vehicular Network
abstract
The flourishing vehicular applications require vehicles to download huge amount of Internet data, which consumes significant backhaul bandwidth and considerable time for data content delivery. Caching the popular data at network edge station can alleviate the congestion at backhaul network and reduce the data delivery delay. In this paper, we propose a dynamic edge caching policy for Heterogeneous Vehicular Network via Reinforcement Learning on adaptive traffic intensity and hot content popularity. We aim to enhance the download rate of vehicles adapting to dynamic vehicle velocity and hot file pool by caching the popular content on heterogeneous network edge stations. The proposed policy makes use of real time information and jointly considers download rate for each file and utility for edge stations to improve overall download performance. Simulation results show that our proposed policy can achieve better download rate than random and fixed caching policies.
Wenchao Xu 0001, Nan Cheng 0001, Huaqing Wu, Shan Zhang 0001, Xuemin Shen
GLOBECOM2
2018 Matching-Based Content Caching in Heterogeneous Vehicular Networks
abstract
To cope with the explosive vehicular content demands from various applications and services, it is imperative to cache the content files close to end users to reduce both the core network traffic load and the delivery delay. The heterogeneous vehicular networks provide multiple data pipes to access to the Internet for vehicles, and thus can be leveraged to cache the required content files on different access network edges and further improve the caching effectiveness. This paper focuses on the content delivery in heterogeneous vehicular networks that allow content caching in roadside WiFi APs, TV White Space Stations, and Cellular Base Stations. With the objective of minimizing the content delivery delay, a Student Admission matching- based content caching scheme is proposed by considering the file popularity, vehicle mobility and cache capacity of the edge infrastructure in heterogeneous vehicular networks, whereby a stable result is further obtained via the Gale-Shapley algorithm. Our simulation results demonstrate the advantages of the proposed caching scheme.
Huaqing Wu, Wenchao Xu 0001, Li Wang 0039, Xuemin Shen
GLOBECOM2
2018 ViFi: Vehicle-to-Vehicle Assisted Traffic Offloading via Roadside WiFi Networks
abstract
Offloading vehicular data traffic from cellular networks to roadside WiFi networks is a very interesting issue since it can not only alleviate the traffic congestion for cellular networks, but also reduce the communication cost for vehicle users. In this paper, we study the vehicle-to-vehicle (V2V) assisted WiFi offloading, where nearby vehicles that associate to different access points (APs) can use their idle WiFi resource to offload part of peer's data traffic. We also consider the Internet access delay introduced by the network detection, user authentication and network address assignment between the vehicle and the AP prior to actual data transmission, which has impact on the WiFi cell sojourn duration of the vehicle. The offloading efficiency, which is the traffic offloaded from cellular network, is analyzed by modeling an M/G/1/K queueing process under various conditions. The accuracy of our analysis is validated through the conducted simulation.
Wenchao Xu 0001, Huaqing Wu, Weisen Shi, Nan Cheng 0001, Xuemin Shen
GLOBECOM1
2018 ABC: Adaptive Beacon Control for Rear-End Collision Avoidance in VANETs
abstract
Vehicular ad hoc network (VANET) has been widely recognized as a promising solution to enhance driving safety, by keeping vehicles well aware of the nearby environment through frequent beacon message exchanging. Due to the dynamic of transportation traffic, especially for those scenarios where the density of vehicles is high, the naive beaconing scheme where vehicles send beacon messages at a fixed rate with a fixed transmission power can cause severe channel congestion. In this paper, we investigate the risk of rear-end collision model and define a danger coefficient ρ to characterize the danger threat of each vehicle being in a rear-end collision. We then propose a fully-distributed beacon congestion control scheme, referred to as ABC, which guarantees each vehicle to actively adapt a minimal but sufficient beacon rate to avoid a rear-end collision based on individual estimates of ρ. In essence, ABC adopts a TDMA-based MAC protocol and solves a NP-hard optimal distributed beacon rate adapting (DBRA) problem with a greedy heuristic algorithm, in which a vehicle with a higher ρ will be assigned with a higher beacon rate while keeping the total required beacon demand lower than the channel capacity. We conduct extensive simulations to demonstrate the efficiency of ABC design in different traffic density and a large variety of underlying road topologies.
Feng Lyu 0001, Hongzi Zhu, Nan Cheng 0001, Yanmin Zhu 0006, Wenchao Xu 0001, Guangtao Xue, Minglu Li 0001
SECON6
2018 DBCC: Leveraging Link Perception for Distributed Beacon Congestion Control in VANETs
abstract
Under the IEEE 802.11p-based dedicated short range communication modules, vehicular safety applications rely on periodical broadcasts of safety beacons by each vehicle. However, the channel can be easily congested by high-frequency periodic beacons when the vehicle density becomes heavy. In this paper, through real-trace-based empirical study on vehicle-to-vehicle communication, we find that nonline-of-sight (NLoS) condition is the key factor on link performance degradation and blindly sending more packets in harsh NLoS conditions can hardly succeed but increase interferences to neighboring vehicles. Inspired by this, we propose a distributed beacon congestion control (DBCC) scheme to control beacon activities with considering link conditions, i.e., vehicles with more neighbors and better conditions of links with its neighbors, will be assigned with higher beacon rates. In DBCC, we first utilize two machine learning methods, i.e., naive Bayes and support vector machines, to train the features and output a classifier model which conducts online NLoS link condition prediction. With link status information, we then formulate a link-weighted safety benefit maximization (L-SBM) problem of the rate-adaptation under a TDMA broadcast MAC, which is proved to be NP-hard. A greedy heuristic algorithm for L-SBM is then proposed and the performance of the algorithm is evaluated. Extensive trace-driven simulations demonstrate the efficiency of DBCC design; particularly, the rate of beacon transmissions can be effectively controlled without exceeding the resource limit and the rate of transmission/reception collisions are greatly reduced.
Feng Lyu 0001, Nan Cheng 0001, Wenchao Xu 0001, Weisen Shi, Minglu Li 0001
IEEE Internet Things J.4
2017 Throughput Analysis of In-Vehicle Internet Access via On-Road WiFi Access Points
abstract
WiFi has been considered as the most promising radio technology to carry the rapidly growing in-vehicle Internet traffic, such as video streaming, user generated content sharing, etc. By offloading traffic from cellular networks to the roadside WiFi Access Points (APs), the traffic throughput can be improved with reduced network cost. Prior to Internet services, the vehicle has to accomplish the access procedure, which involves the transmission of the management frames such as probe request/response frames, authentication frames, etc. The access procedure can affect the throughput of the vehicular Internet connection, since the vehicle has to wait until all frames are successfully transmitted before accessing Internet services. To study the impact of such access procedure on the throughput performance of the in-vehicle Internet access via on-road WiFi APs, in this paper, we propose a two dimensional Markov chain model to study the progress of the access procedure when the vehicle drives through the consecutive zones within the AP coverage area. We evaluate the dependency of the throughput over different conditions, such as packet error rate, velocity of the vehicle, average packet delay, etc. The results of the paper will provide useful insights for future design and deployment of the roadside WiFi networks.
Wenchao Xu 0001, Weisen Shi, Feng Lyu 0001, Xuemin Shen
VTC Fall1
2017 Service-Oriented Dynamic Connection Management for Software-Defined Internet of Vehicles
abstract
Internet of vehicles (IoV) is an emerging paradigm for accommodating the requirements of future intelligent transportation systems (ITSs) with the overwhelming trend of equipping vehicles with versatile sensors and communications modules, and facilitating drivers and passengers with a variety of innovative ITS applications. However, the implementation of IoV still faces many challenges, such as flexible and efficient connections, quality of service guarantee, and multiple concurrent support requests. To this end, in this paper we introduce the software-defined IoV (SD-IoV), which is able to tackle the above-mentioned issues by adopting the software-defined networking framework. We first present the architecture of SD-IoV and develop a centralized vehicular connection management approach. Then, we aim to allocate dedicated communications resources and underlying vehicular nodes to satisfy each service. We formulate the dynamic vehicular connection as an overlay vehicular network creation (OVNC) problem. A comprehensive utility function is also designed to serve as the optimization objective of OVNC. Finally, we solve the OVNC problem by developing a graph-based genetic algorithm and a heuristic algorithm, respectively. Extensive simulation results are provided to demonstrate the effectiveness of our proposed solution of dynamic vehicular connection management.
Ning Zhang 0007, Wenchao Xu 0001, Lin Gui 0001, Xuemin Shen
IEEE Trans. Intell. Transp. Syst.4
2016 An Efficient PMIPv6-Based Handoff Scheme for Urban Vehicular Networks
abstract
In urban vehicular networks, traveling users can enjoy Internet multimedia services through various mobile devices, such as smart phones and laptops. To maintain seamless and ubiquitous Internet connectivity, an efficient handoff scheme has to be employed when mobile users travel across different access networks. However, in the urban vehicular environment, the high velocity of vehicles and the random mobility of users impose great challenges to the design of an effective handoff scheme. In this paper, we propose an Efficient Proxy Mobile IPv6 (E-PMIPv6)-based handoff scheme that guarantees session continuity for urban mobile users. In the registration process, E-PMIPv6 enables mobile users to obtain seamless Internet connectivity either from fixed roadside units or mobile routers and improves cache utilization at the local mobility anchor by merging the binding cache entries of the mobile users. In the handoff process, E-PMIPv6 comprehensively considers various handoff scenarios in the urban vehicular environment and provides transparent network-based mobility support to individual mobile users or a group of users in the same mobile network without disrupting ongoing sessions. In addition, E-PMIPv6 eliminates packet loss by either packet buffering or packet tunneling to improve handoff performance in each handoff scenario. Finally, a detailed analytical model is developed to study the performance of E-PMIPv6 in terms of handoff latency, signaling overhead, buffering cost, and tunneling cost. Analysis and simulation results demonstrate that the proposed E-PMIPv6 successfully extends the scalability of user mobility and greatly improves handoff efficiency in urban vehicular networks.
Yuanguo Bi, Wenchao Xu 0001, Xuemin Shen, Hai Zhao 0002
IEEE Trans. Intell. Transp. Syst.3
2011 Channel Assignment and User Association Game in Dense 802.11 Wireless Networks
abstract
In densely deployed IEEE 802.11 wireless networks, the transmission delay experienced by a user depends not only on the traffic load of the associated AP, but also the contention level of other APs operating on the same channel. However, due to the random distribution of users and inappropriate allocation of AP channels, the traffic loads of different APs are often uneven, leading to unfair delay experience to different users. In this paper, we consider the problem of channel assignment and user association for balancing the traffic load of APs operating on different channels, which is modeled as a non-cooperative game. We prove the existence of Nash equilibrium (NE) for this game, and derive the price of anarchy and the fairness index at NE. Simulation results are provided to compare the performance of the proposed algorithm with the theoretical bounds.
Wenchao Xu 0001, Cunqing Hua, Aiping Huang
ICC1
2010 A Game Theoretical Approach for Load Balancing User Association in 802.11 Wireless Networks
abstract
In the multi-cell IEEE 802.11 wireless networks, the traffic loads of access points(APs) are often uneven, which leads to inefficient use of network resources and unfair service to users. To alleviate such imbalance, different user association schemes have been proposed that use different metrics for measuring the congestion level of APs. In this paper, we propose a game theoretical model for the user association problem using the airtime cost as the congestion metric. The centralized and localized algorithms are designed for achieving the airtime- balancing Nash equilibrium. Simulation results show that the proposed algorithms outperform the existing scheme in terms of fairness and load balance.
Wenchao Xu 0001, Cunqing Hua, Aiping Huang
GLOBECOM1