VLDB 2026 Research / reviewers in the wild / expert
Yongheng Deng
dblp:251/2735
· DBLP profile ↗
30ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0003-3010-3812ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 8 first-author · 19 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSVAR: Enhancing Visual Privacy in Federated Learning via Adaptive Shuffling Against Overfitting
Zhenya Ma, Yan Zhang 0073, Donghua Cai, Qiushi Li 0002, Yongheng Deng, Ye Zhang 0033, Ju Ren 0001, Xuemin Shen |
ICC | 6 |
| 2026 | EarAuth: Towards Practical Cardiac Vibration Authentication on COTS Wireless Earbuds
Yongjian Fu 0004, Wenpeng Zhu, Yingjun Wu, Hao Pan 0003, Guanbo Wang, Yongheng Deng, Yaoxue Zhang, Ju Ren 0001 |
INFOCOM | 9 |
| 2026 | Revisiting Redundancy in Diffusion Transformers: A Temporal-Spatial Joint Caching Strategy for Efficient SamplingabstractDiffusion Transformers (DiTs) achieve impressive generative performance but suffer from significant inference latency. Feature caching–based acceleration methods reduce total computation by reusing results from earlier timesteps, but they largely ignore that temporal redundancy is dynamic and inconsistent across timesteps. Our analysis reveals this variability. More crucially, we identify a previously underexplored form of efficiency, namely spatial redundancy, characterized by high similarity between adjacent transformer blocks within the same timestep. Motivated by this dual-dimensional redundancy, we propose Temporal-Spatial Joint Cache, a training-free inference acceleration strategy that dynamically determines optimal reuse operations across temporal and spatial dimensions. Our approach features a redundancy-guided operation selector that estimates local feature stability using second-order divided differences, enabling fine-grained decisions between full computation, temporal cache, and spatial cache. Furthermore, we use interpolation-based feature prediction to capture local feature evolution for more accurate reuse. In addition, we propose a bounded cache distance control mechanism to mitigate error accumulation from excessive reuse. Together, these components allow our method to deliver substantial inference speedups without retraining or compromising generation fidelity, offering a new perspective on efficiency in diffusion transformer inference. Chenxi Du, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang |
KDD (1) | 2 |
| 2026 | A Fine-Tuning Data Recovery Attack on Generative Language Models via BackdooringabstractGenerative language models (GLMs) are increasingly integrated into modern intelligent applications to power intelligent functionalities. Developers often fine-tune open-source GLMs on proprietary data and deploy them in real-world applications. In this paper, we reveal a novel model supply chain attack that exploits this workflow: by injecting backdoors into the source code of an open-source GLM, an adversary can induce the model to memorize fine-tuning data and later regenerate it via crafted prompts. We propose LURE, a new backdoor-based data recovery attack that exploits memorization capabilities of fine-tuned models. During fine-tuning, LURE stealthily injects unique and attacker-enumerable hash prompts, and incorporates a Position-Decay Weighted Aligned Cross-Entropy Loss into the original fine-tuning loss, strengthening the association between injected prompts and corresponding data samples for effective data recovery. To achieve stealthy and transparent attack injection, LURE employs a stealthy backdoor within the model’s source code, enabling automatic injection of hash prompts during fine-tuning and thus maintaining the user’s original fine-tuning workflow. LURE also proposes several optimizations to maintain minimal impact on the performance of the original task and external training state. Extensive evaluations demonstrate the remarkable efficacy of LURE, achieving a 45%-68% data recovery rate while maintaining the attack’s transparency, stealthiness, and showcasing its ability to evade existing defenses. Zhenya Ma, Yongheng Deng, Ziqing Qiao, Quan Zhang 0003, Chijin Zhou, Fan Wu 0014, Yaoxue Zhang, Ju Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Toward Communication-Efficient and Data-Free Collaborative Fine-Tuning Between Small and Large Language ModelsabstractWhile large language models (LLMs) exhibit impressive general capabilities, their performance on domainspecific tasks often requires fine-tuning with private data that cannot be shared due to privacy constraints. Directly deploying LLMs on resource-constrained clients for local fine-tuning is impractical due to their significant computation and communication costs. In addition, pre-trained LLMs are valuable intellectual property, and model owners are reluctant to distribute full model weights. To address these challenges, we proposeCoT-LM, a communication-efficient, computation-light, and data-free framework for collaborative fine-tuning between small (SLMs) and large language models (LLMs). InCoT-LM, clients fine-tune lightweight SLMs locally without uploading models or private data. These SLMs provide task-specific feedback to guide server-side LLM enhancement via an efficient communication protocol that exchanges only lightweight synthetic data and feedback. The framework supports both synchronous and asynchronous collaboration and enables mutual enhancement: the LLM improves its task-specific capabilities, while clients benefit from refined synthetic data or distilled knowledge. Extensive experiments demonstrate thatCoT-LMsignificantly boosts natural language understanding (NLU, up to 18.3% for LLMs and 8.0% for SLMs) and natural language generation (NLG, up to 31.7% for LLMs) performance across diverse tasks while preserving data privacy, model intellectual property, and generalization capabilities, achieving significant reductions in computation and communication overhead. Zhenya Ma, Yongheng Deng, Ziqing Qiao, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | ConCISE: Confidence-guided Compression in Step-by-step Efficient ReasoningabstractZiqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziqing Qiao, Yongheng Deng, Jiali Zeng, Guanbo Wang, Fandong Meng, Jie Zhou 0016, Ju Ren 0001, Yaoxue Zhang |
EMNLP | 2 |
| 2025 | Efficient LOS/NLOS Classification in 5G NR Using Field Image Transformation and Modified Mean Teacher Frameworkabstract5G NR wireless positioning demonstrates exceptional potential in indoor and GNSS-denied environments due to its low latency, high signal quality, and precise measurement capabilities. However, non-line-of-sight (NLOS) propagation significantly degrades performance, necessitating robust LOS/NLOS signal identification. This paper presents an innovative framework combining Gramian Angular Fields (GAFs), Recurrence Plots (RPs), Kolmogorov-Arnold Networks (KAN), and a Mean Teacher semi-supervised approach to address these challenges.Our method converts raw channel impulse response (CIR) sequences into two-dimensional Field Images through GAF-GADF encoding, enabling rich spatial-temporal feature extraction. A hybrid feature learning strategy is employed: (1) a lightweight CNN extracts visual features from Field Images, and (2) a KAN-based module processes manually extracted CIR features through spline-activated layers to capture nonlinear relationships. A consistency factor is introduced in the fusion module to enhance model updates, while the Mean Teacher framework leverages unlabeled data through exponential moving average weight smoothing and consistency loss regularization.Extensive experiments on a real-world dataset validate the proposed approach: supervised classification achieves 99.1% accuracy, while semi-supervised learning attains 89.8% accuracy, both surpassing state-of-the-art methods. These results demonstrate the framework’s effectiveness in enhancing 5G NR positioning reliability in obstructive environments. Renhao Liu, Enwen Hu, Yongheng Deng |
IPIN | 5 |
| 2025 | Malva: A Jitter-Aware Online Pruning Framework for DNN Inference TasksabstractIn fields like autonomous driving, strict constraints are imposed on the computing latency of deep neural network (DNN) inference tasks on edge servers. However, it is typical for edge servers to execute multiple tasks in parallel to serve multiple users, causing severe latency jitter due to resource competition, which seriously affects timeliness. Existing works ignore the computing jitter and regard computing latency as a deterministic value, failing to meet the timeliness requirement. To address this issue, we propose Malva, a framework for finegrained online pruning for DNN tasks, allowing flexible pruning at runtime based on jitter conditions. Specifically, Malva first partitions the DNN model into blocks and applies early exiting and pruning methods to create block variants. Then, the Malva scheduler flexibly selects the variant to be executed or exits early according to urgency and jitter conditions. Moreover, we propose a novel urgency-aware prediction strategy to estimate the accuracy impact of variants with incomplete pathway information during scheduling. Stress testing shows Malva can strictly maintain a zero deadline miss rate and significantly increase the stress required to cause the first deadline miss while still outperforming state-of-the-art methods in accuracy. Ziyan Fu 0001, Yongheng Deng, Yingjun Wu, Zhibo Wang 0001, Su Yao, Yaoxue Zhang, Ju Ren 0001 |
IWQoS | 2 |
| 2025 | FedAF: Alignment-Augmented Fusion for Federated Multimodal Learning with Small LabelsabstractFederated multimodal learning is an emerging advancement in artificial intelligence, enabling the integration of data from diverse modalities while preserving data privacy. However, limited labeled data and modality heterogeneity on the clients pose significant challenges for effective federated multimodal model training. To address these challenges, this paper introduces FedAF, a novel alignment-augmented fusion framework tailored for federated multimodal learning. FedAF extracts unbiased and complementary information from multiple modalities with small data, enabling effective modality fusion and feature alignment for improving system performance. The framework introduces a three-stage strategy. First, FedAF utilizes labeled data to create unbiased anchor points, addressing disparities in client feature distributions. Second, FedAF employs a weighted enhancement contrast fusion scheme to improve feature clustering and reduce feature overlap. Finally, a multimodal semisupervised algorithm mitigates data heterogeneity and overfitting. Extensive experiments demonstrate that FedAF significantly outperforms baseline methods, showcasing its effectiveness in federated multimodal learning scenarios. Guanbo Wang, Yongheng Deng, Yingjun Wu, Xinyi Li 0005, Tuowei Wang, Yaoxue Zhang, Ju Ren 0001 |
IWQoS | 2 |
| 2025 | CrossLM: A Data-Free Collaborative Fine-Tuning Framework for Large and Small Language ModelsabstractWhile large language models (LLMs) are endowed with broad knowledge, their task-specific performance is often suboptimal. Fine-tuning LLMs with task-specific data from diverse nodes is necessary, but this data is typically safeguarded and not shared publicly due to privacy concerns. A common solution involves downstream nodes downloading the LLM locally and fine-tuning it with their proprietary data. However, owners often regard pre-trained LLMs as valuable assets and are reluctant to share them. Additionally, the significant computational resources required by LLMs make local fine-tuning impractical for many nodes. To mitigate these problems, this paper proposes CrossLM, a data-free collaborative fine-tuning framework for large and small language models. CrossLM enables resource-constrained nodes to train smaller language models (SLMs) using their private task-specific data. These SLMs are subsequently leveraged to promote the task-specific natural language generation and understanding capabilities of the LLMs. Simultaneously, the SLMs of nodes also benefit from enhancement by the fine-tuned LLMs. In this way, CrossLM avoids sharing private data and proprietary LLMs, and also reduces the resource requirements of nodes. Through extensive experiments across a range of benchmark tasks and popular language models, we demonstrate that CrossLM significantly boosts the task-specific performance of both LLMs and SLMs while preserving the generalization capabilities of LLMs. Yongheng Deng, Ziqing Qiao, Ye Zhang 0033, Zhenya Ma, Yang Liu 0165, Ju Ren 0001 |
MobiSys | 1 |
| 2025 | Gains: Fine-grained Federated Domain Adaptation in Open SetabstractConventional federated learning (FL) assumes a closed world with a fixed total number of clients. In contrast, new clients continuously join the FL process in real-world scenarios, introducing new knowledge. This raises two critical demands: detecting new knowledge, i.e., knowledge discovery, and integrating it into the global model, i.e., knowledge adaptation. Existing research focuses on coarse-grained knowledge discovery, and often sacrifices source domain performance and adaptation efficiency. To this end, we propose a fine-grained federated domain adaptation approach in open set (Gains). Gains splits the model into an encoder and a classifier, empirically revealing features extracted by the encoder are sensitive to domain shifts while classifier parameters are sensitive to class increments. Based on this, we develop fine-grained knowledge discovery and contribution-driven aggregation techniques to identify and incorporate new knowledge. Additionally, an anti-forgetting mechanism is designed to preserve source domain performance, ensuring balanced adaptation. Experimental results on multi-domain datasets across three typical data-shift scenarios demonstrate that Gains significantly outperforms other baselines in performance for both source-domain and target-domain clients. Code is available at: https://github.com/Zhong-Zhengyi/Gains. Zhengyi Zhong, Wenzheng Jiang, Weidong Bao 0001, Ji Wang 0002, Cheems Wang, Guanbo Wang, Yongheng Deng, Ju Ren 0001 |
NeurIPS | 7 |
| 2025 | MASA: Multimodal Federated Learning Through Modality-Aware and Secure AggregationabstractAs a promising paradigm, federated learning has been applied to multimodal sensing tasks due to its deployment convenience. However, the recent advances in multimodal federated learning emphasize learning a high-quality multimodal model but overlook the model usage requirements of massive unimodal clients. Moreover, the privacy risk in model sharing and client data heterogeneity impact the efficacy of federated learning. In this paper, we propose a novel multimodal federated learning system named MASA. As a departure from existing approaches, MASA simultaneously enhances the model learning efficiency of both multimodal and unimodal clients while ensuring their data privacy. First, we employ a gated cross-modal distillation scheme to achieve performance-aware knowledge transfer across modality-heterogeneous clients. To enhance the system security, MASA integrates a lightweight split-shuffle mechanism to realize the anonymization and encryption of model aggregation. Moreover, to reach personalized collaboration while protecting privacy, MASA features an attention-based spontaneous client clustering mechanism to form client cluster structures securely and distributedly. We evaluate our MASA on four public multimodal datasets for human activity recognition. The results show that our MASA outperforms leading multimodal federated learning methods on the model performance of both multimodal and unimodal clients. Jialin Guo, Yongjian Fu 0004, Zhiwei Zhai, Xinyi Li 0005, Yongheng Deng, Sheng Yue 0001, Hao Pan 0003, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | DualRec: A Collaborative Training Framework for Device and Cloud Recommendation ModelsabstractRecommendation systems (RS) play a vital role in various domains. However, under recent data regulations like General Data Protection Regulation (GDPR), traditional RS that rely on collecting user's interaction data centrally face significant challenges. Federated learning (FL) enables collaborative model training among users while keeping their private data locally. Yet, the constrained resources of devices often limit the size of the learned model, resulting in suboptimal recommendation performance. To overcome the dilemma of data accessibility and model size, we propose DualRec, a novel collaborative training framework for device and cloud recommendation models. In DualRec, users train lightweight models on devices to harness their local private data, while a larger model is simultaneously trained on the cloud server to exploit its substantial resources. Devices and the cloud server collaboratively train their models, compensating for individual limitations of model size and data availability, enabling mutual empowerment and benefits. Specifically, we introduce an efficient aggregation mechanism for recommendation models to boost the collaborative training performance of device models. With the learned device models, we propose to generate pseudo user interaction data to train the server model. To enhance the training performance of the server model, we design an automated denoising mechanism to mitigate the negative impact of noisy samples in the generated pseudo dataset. Finally, the learned knowledge of the server model is distilled to device models for enhanced on-device recommendation performance. Extensive experiments demonstrate the superior performance of DualRec compared to state-of-the-art baselines. Ye Zhang 0033, Yongheng Deng, Sheng Yue 0001, Qiushi Li 0002, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Momentum-Based Contextual Federated Reinforcement LearningabstractFederated Reinforcement Learning (FRL) is an attractive edge learning paradigm for decision-making applications, which has garnered significant interest recently. However, owing to the inherent spatio-temporal non-stationarity of local state-action distributions, current FRL approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL method, which incorporates momentum, importance sampling, and server-side adjustments, capable of controlling the gradient shifts induced by the non-stationary data. We prove that by proper selection of momentum parameters and interaction frequency, it can achieve$\tilde {\mathcal {O}}(H N^{-1}\epsilon ^{-3/2})$and$\tilde {\mathcal {O}}(\epsilon ^{-1})$interaction and communication complexities (N represents the agent number), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns with the best achievable among existing first-order FL algorithms. Further, we leverage attention-based contextual representation extraction to enable the learning policy to adapt to heterogeneous tasks and environments. Extensive experiments demonstrate that our proposed method significantly outperforms existing baselines on a range of complex, high-dimensional single-task and multi-task benchmarks. Sheng Yue 0001, Xingyuan Hua, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Netw. | 3 |
| 2025 | AugFL: Augmenting Federated Learning With Pretrained ModelsabstractFederated Learning (FL) has garnered widespread interest in recent years. However, owing to strict privacy policies or limited storage capacities of training participants such as IoT devices, its effective deployment is often impeded by the scarcity of training data in practical decentralized learning environments. In this paper, we study enhancing FL with the aid of (large) pre-trained models (PMs), that encapsulate wealthy general/domain-agnostic knowledge, to alleviate the data requirement in conducting FL from scratch. Specifically, we consider a networked FL system formed by a central server and distributed clients. First, we formulate the PM-aided personalized FL as a regularization-based federated meta-learning problem, where clients join forces to learn a meta-model with knowledge transferred from a private PM stored at the server. Then, we develop an inexact-ADMM-based algorithm, AugFL, to optimize the problem with no need to expose the PM or incur additional computational costs to local clients. Further, we establish theoretical guarantees for AugFL in terms of communication complexity, adaptation performance, and the benefit of knowledge transfer in general non-convex cases. Extensive experiments corroborate the efficacy and superiority of AugFL over existing baselines. Sheng Yue 0001, Zerui Qin, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang, Junshan Zhang |
IEEE Trans. Netw. | 3 |
| 2024 | Federated Offline Policy Optimization with Dual RegularizationabstractFederated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the environment during local updating, which can be prohibitively expensive or even infeasible in many real-world domains. To overcome this challenge, this paper proposes a novel offline federated policy optimization algorithm, named DRPO, which enables distributed agents to collaboratively learn a decision policy only from private and static data without further environmental interactions. DRPO leverages dual regularization, incorporating both the local behavioral policy and the global aggregated policy, to judiciously cope with the intrinsic two-tier distributional shifts in offline FRL. Theoretical analysis characterizes the impact of the dual regularization on performance, demonstrating that by achieving the right balance thereof, DRPO can effectively counteract distributional shifts and ensure strict policy improvement in each federative learning round. Extensive experiments validate the significant performance gains of DRPO over baseline methods. Sheng Yue 0001, Zerui Qin, Xingyuan Hua, Yongheng Deng, Ju Ren 0001 |
INFOCOM | 4 |
| 2024 | RelayRec: Empowering Privacy-Preserving CTR Prediction via Cloud-Device Relay LearningabstractClick-through rate (CTR) prediction holds paramount importance across numerous applications, profoundly impacting user experience and business profitability. The freshness of a CTR prediction model significantly influences its performance, since users’ needs and interests may be changing over time, thereby requiring the model to be updated frequently. However, stringent data protection regulations have constrained the collection of users’ personal data, posing challenges to traditional model refreshing strategies that rely on centralized data collection. On-device learning techniques, such as federated learning (FL), offer a viable solution by enabling model training on devices without compromising user privacy. Nevertheless, the scarcity of training data with diverse distributions among devices presents considerable obstacles to on-device learning effectiveness. To address these challenges, we introduce RelayRec, a cloud-device relay learning framework designed for privacy-preserving CTR prediction. To establish competent initial models for devices, RelayRec categorizes pre-regulation cloud data into user preference groups, training preference-specific models for devices. Furthermore, a cloud-based automated model selector is developed to identify suitable initial models for devices. To elevate the relay learning performance of these initial models, we incorporate a personalized collaborative learning mechanism that aggregates device models based on user preferences. Extensive experimental evaluations underscore RelayRec’s superior performance compared to state-of-the-art benchmarks, affirming its efficacy in privacy-preserving CTR prediction. Yongheng Deng, Guanbo Wang, Sheng Yue 0001, Wei Rao 0003, Qin Zu, Ju Ren 0001, Yaoxue Zhang |
IPSN | 1 |
| 2024 | A Communication-Efficient Hierarchical Federated Learning Framework via Shaping Data Distribution at EdgeabstractFederated learning (FL) enables collaborative model training over distributed computing nodes without sharing their privacy-sensitive raw data. However, in FL, iterative exchanges of model updates between distributed nodes and the cloud server can result in significant communication cost, especially when the data distributions at distributed nodes are imbalanced with requiring more rounds of iterations. In this paper, with our in-depth empirical studies, we disclose that extensive cloud aggregations can be avoided without compromising the learning accuracy if frequent aggregations can be enabled at edge network. To this end, we shed light on the hierarchical federated learning (HFL) framework, where a subset of distributed nodes can play as edge aggregators to support edge aggregations. Under the HFL framework, we formulate a communication cost minimization (CCM) problem to minimize the total communication cost required for model learning with a target accuracy by making decisions on edge aggragator selection and node-edge associations. Inspired by our data-driven insights that the potential of HFL lies in the data distribution at edge aggregators, we propose ShapeFL, i.e., SHaping dAta distRibution at Edge, to transform and solve the CCM problem. In ShapeFL, we divide the original problem into two sub-problems to minimize the per-round communication cost and maximize the data distribution diversity of edge aggregator data, respectively, and devise two light-weight algorithms to solve them accordingly. Extensive experiments are carried out based on several opened datasets and real-world network topologies, and the results demonstrate the efficacy of ShapeFL in terms of both learning accuracy and communication efficiency. Yongheng Deng, Feng Lyu 0001, Tengxi Xia, Yue-Zhi Zhou, Yaoxue Zhang, Ju Ren 0001, Yuanyuan Yang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | LEARN: Selecting Samples Without Training Verification for Communication-Efficient Vertical Federated LearningabstractIn the classical vertical federated learning (VFL) framework, feature maps and corresponding gradient information of all samples are transferred between the server and clients, which causes a significant communication burden. Therefore, to enable efficient VFL in resource-constrained wireless networks, we propose to select a part of the samples from the large training set to train models with minimal accuracy degradation. To this end, we propose LEARN, i.e., seLecting Efficient sAmples without tRaining verificatioN, to select efficient training samples for VFL. Particularly, LEARN integrates two major components named label distribution smoothing and feature center-based vertical sample filtering. The number of samples selected for each class is determined by the label distribution smoothing mechanism. Then the feature center-based vertical sample filtering component calculates the features centers and performs sample selection based on the distance between the samples and their corresponding feature center. Extensive experiments under various settings are carried out to corroborate the efficacy and robustness of LEARN. Tong Liu 0035, Feng Lyu 0001, Yongheng Deng, Qilong Tan, Yaoxue Zhang |
GLOBECOM | 4 |
| 2023 | A Prototype-Based Knowledge Distillation Framework for Heterogeneous Federated LearningabstractFederated learning (FL) is an emerging distributed machine learning paradigm, which has shown great potential in collaborative learning with privacy preservation. However, FL clients usually have disparate system resource capabilities (e.g., data, computation, and communication) for model training and aggregation, which can cause a series of system heterogeneity issues with performance degradation. To this end, we propose FedPKD, a Prototype-based Knowledge Distillation framework for FL. FedPKD integrates knowledge distillation and prototype learning with FL, which enables heterogeneous clients and the server to learn collaboratively, with different model architectures and resource capability adaptations. Specifically, FedPKD proposes to transfer dual knowledge of clients including the model output logits and prototypes to the server, and a prototype-based ensemble distillation mechanism is proposed to aggregate the logits and prototypes from clients, which can be used to train the server model with an unlabeled public dataset. The server model knowledge is then transferred back to clients to improve the performance of client models. Moreover, to improve learning performance and reduce communication overhead, we propose a prototype-based data filter mechanism to filter out the samples with low-quality knowledge. Extensive experiments under various settings demonstrate the superiority of FedPKD in learning performance and communication efficiency when compared to state-of-the-art benchmarks. Feng Lyu 0001, Yongheng Deng, Tong Liu 0035, Yongmin Zhang, Yaoxue Zhang |
ICDCS | 3 |
| 2023 | A Hierarchical Knowledge Transfer Framework for Heterogeneous Federated LearningabstractFederated learning (FL) enables distributed clients to collaboratively learn a shared model while keeping their raw data private. To mitigate the system heterogeneity issues of FL and overcome the resource constraints of clients, we investigate a novel paradigm in which heterogeneous clients learn uniquely designed models with different architectures, and transfer knowledge to the server to train a larger server model that in turn helps to enhance client models. For efficient knowledge transfer between client models and server model, we propose FedHKT, a Hierarchical Knowledge Transfer framework for FL. The main idea of FedHKT is to allow clients with similar data distributions to collaboratively learn to specialize in certain classes, then the specialized knowledge of clients is aggregated to a super knowledge covering all specialties to train the server model, and finally the server model knowledge is distilled to client models. Specifically, we tailor a hybrid knowledge transfer mechanism for FedHKT, where the model parameters based and knowledge distillation (KD) based methods are respectively used for client-edge and edge-cloud knowledge transfer, which can harness the pros and evade the cons of these two approaches in learning performance and resource efficiency. Besides, to efficiently aggregate knowledge for conducive server model training, we propose a weighted ensemble distillation scheme with server-assisted knowledge selection, which aggregates knowledge by its prediction confidence, selects qualified knowledge during server model training, and uses selected knowledge to help improve client models. Extensive experiments demonstrate the superior performance of FedHKT compared to state-of-the-art baselines. Yongheng Deng, Ju Ren 0001, Feng Lyu 0001, Yang Liu 0165, Yaoxue Zhang |
INFOCOM | 1 |
| 2023 | FedINC: An Exemplar-Free Continual Federated Learning Framework with Small Labeled DataabstractFederated learning (FL) has shown great promise for privacy-preserving learning by enabling collaborative training on decentralized clients. However, in realistic FL scenarios, clients often collect new data continuously, join or exit learning dynamically. As a result, the global model tends to forget old knowledge while learning new knowledge. Meanwhile, labeling the continuously arriving data in real-time is usually challenging. Therefore, the catastrophic forgetting problem intertwined with the label deficiency issue poses significant challenges for both learning new knowledge and consolidating old knowledge. To address these challenges, we develop a novel exemplar-free continual federated learning framework named FedINC, to learn a global incremental model with limited labeled data. We begin by excavating the cause of catastrophic forgetting via in-depth empirical studies. Based on that, we introduce targeted mechanisms for FedINC, including a hybrid contrastive learning mechanism to efficiently learn new knowledge with limited labeled data, a plastic feature regularization mechanism to preserve old task's representation space, a prototype-guided regularization mechanism to mitigate feature overlap between old and new classes while aligning the features of non-iid clients, and a prototype evolution mechanism for flexible and efficient incremental classification. Extensive experiments demonstrate the superior performance of FedINC in terms of both convergence speed and accuracy of the global model. Yongheng Deng, Sheng Yue 0001, Tuowei Wang, Guanbo Wang, Ju Ren 0001, Yaoxue Zhang |
SenSys | 1 |
| 2022 | HSFL: An Efficient Split Federated Learning Framework via Hierarchical OrganizationabstractFederated learning (FL) has emerged as a popular paradigm for distributed machine learning among vast clients. Unfortunately, resource-constrained clients often fail to participate in FL because they cannot pay for the memory resources required for model training due to their limited memory or bandwidth. Split federated learning (SFL) is a novel FL framework in which clients commit intermediate results of model training to a cloud server for client-server collaborative training of models, making resource-constrained clients also eligible for FL. However, existing SFL frameworks mostly require frequent communication with the cloud server to exchange intermediate results and model parameters, which results in significant communication overhead and elongated training time. In particular, this can be exacerbated by the imbalanced data distributions of clients. To tackle this issue, we propose HSFL, a hierarchical split federated learning framework that efficiently trains SFL model through hierarchical organization participants. Under the HSFL framework, we formulate a Cloud Aggregation Time Minimization (CATM) problem to minimize the global training time and design a light-weight client assignment algorithm based on dynamic programming to solve it. Moreover, we develop a self-adaption approach to cope with the dynamic computational resources of clients. Finally, we implement and evaluate HSFL on various real-world training tasks, elaborating on its effectiveness and superiority in terms of efficiency and accuracy compared to baselines. Tengxi Xia, Yongheng Deng, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
CNSM | 2 |
| 2022 | TailorFL: Dual-Personalized Federated Learning under System and Data HeterogeneityabstractFederated learning (FL) enables distributed mobile devices to collaboratively learn a shared model without exposing their raw data. However, heterogeneous devices usually have limited and different available resources, i.e., system heterogeneity, for model training and communicating, while the diverse data distribution among devices, i.e., data heterogeneity, may result in significant performance degradation. In this paper, we propose TailorFL, a dual-personalized FL framework, which tailors a submodel for each device with personalized structure for training and personalized parameters for local inference. To achieve this, we first excavate the personalization principle for data heterogeneous FL via in-depth empirical studies, and based on which, we propose a resource-aware and data-directed pruning strategy that makes each device's submodel structure match its resource capability and correlate with its local data distribution. To aggregate the submodels while preserving their dual personalization properties, we design a scaling-based aggregation strategy that scales parameters with the pruning rate of submodels and aggregates the overlapped parameters. Moreover, to further promote beneficial and restrain detrimental collaborations among devices, we propose a server-assisted model-tuning mechanism, which dynamically tunes device's submodel structure at the server side with the global view of device's data distribution similarities. Extensive experiments demonstrate that compared to the status quo approaches, TailorFL achieves an average of 22% increase in inference accuracy, and reduces the memory, computation, and communication costs for model training simultaneously. Yongheng Deng, Weining Chen, Ju Ren 0001, Feng Lyu 0001, Yang Liu 0165, Yunxin Liu 0001, Yaoxue Zhang |
SenSys | 1 |
| 2022 | Making resource adaptive to federated learning with COTS mobile devices
Yongheng Deng, Chengbo Jiao, Xing Bao, Feng Lyu 0001 |
Peer-to-Peer Netw. Appl. | 1 |
| 2022 | Improving Federated Learning With Quality-Aware User Incentive and Auto-Weighted Model AggregationabstractFederated learning enables distributed model training over various computing nodes, e.g., mobile devices, where instead of sharing raw user data, computing nodes can solely commit model updates without compromising data privacy. The quality of federated learning relies on the model updates contributed by computing nodes training with their local data. However, with various factors (e.g., training data size, mislabeled data samples, skewed data distributions), the model update qualities of computing nodes can vary dramatically, while inclusively aggregating low-quality model updates can deteriorate the global model quality. To achieve efficient federated learning, in this paper, we propose a novel framework namedFAIR, i.e.,Federated leArning with qualIty awaReness. Particularly,FAIRintegrates three major components: 1) learning quality estimation: we adopt the model aggregation weight (learned in the third component) to reversely quantify the individual learning quality of nodes in a privacy-preserving manner, and leverage the historical learning records to infer the next-round learning quality; 2) quality-aware incentive mechanism: within the recruiting budget, we model a reverse auction problem to stimulate the participation of high-quality and low-cost computing nodes, and the method is proved to be truthful, individually rational, and computationally efficient; and 3) auto-weighted model aggregation: based on the gradient descent method, we devise an auto-weighted model aggregation algorithm to automatically learn the optimal aggregation weights to further enhance the global model quality. Based on real-world datasets and learning tasks, extensive experiments are conducted to demonstrate the efficacy ofFAIR. Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yi-Chao Chen 0001, Peng Yang 0004, Yue-Zhi Zhou, Yaoxue Zhang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | AUCTION: Automated and Quality-Aware Client Selection Framework for Efficient Federated LearningabstractThe emergency of federated learning (FL) enables distributed data owners to collaboratively build a global model without sharing their raw data, which creates a new business chance for building data market. However, in practical FL scenarios, the hardware conditions and data resources of the participant clients can vary significantly, leading to different positive/negative effects on the FL performance, where the client selection problem becomes crucial. To this end, we proposeAUCTION, anAutomated and qUality-awareClient selecTIONframework for efficient FL, which can evaluate the learning quality of clients and select them automatically with quality-awareness for a given FL task within a limited budget. To designAUCTION, multiple factors such as data size, data quality, and learning budget that can affect the learning performance should be properly balanced. It is nontrivial since their impacts on the FL model are intricate and unquantifiable. Therefore,AUCTIONis designed to encode the client selection policy into a neural network and employ reinforcement learning to automatically learn client selection policies based on the observed client status and feedback rewards quantified by the federated learning performance. In particular, the policy network is built upon an encoder-decoder deep neural network with an attention mechanism, which can adapt to dynamic changes of the number of candidate clients and make sequential client selection actions to reduce the learning space significantly. Extensive experiments are carried out based on real-world datasets and well-known learning models to demonstrate the efficiency, robustness, and scalability ofAUCTION. Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Huaqing Wu, Yue-Zhi Zhou, Yaoxue Zhang, Xuemin Shen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | SHARE: Shaping Data Distribution at Edge for Communication-Efficient Hierarchical Federated LearningabstractFederated learning (FL) can enable distributed model training over mobile nodes without sharing privacy-sensitive raw data. However, to achieve efficient FL, one significant challenge is the prohibitive communication overhead to commit model updates since frequent cloud model aggregations are usually required to reach a target accuracy, especially when the data distributions at mobile nodes are imbalanced. With pilot experiments, it is verified that frequent cloud model aggregations can be avoided without performance degradation if model aggregations can be conducted at edge. To this end, we shed light on the hierarchical federated learning (HFL) framework, where a subset of distributed nodes are selected as edge aggregators to conduct edge aggregations. Particularly, under the HFL framework, we formulate a communication cost minimization (CCM) problem to minimize the communication cost raised by edge/cloud aggregations with making decisions on edge aggregator selection and distributed node association. Inspired by the insight that the potential of HFL lies in the data distribution at edge aggregators, we propose SHARE, i.e., SHaping dAta distRibution at Edge, to transform and solve the CCM problem. In SHARE, we divide the original problem into two sub-problems to minimize the per-round communication cost and mean Kullback-Leibler divergence of edge aggregator data, and devise two light-weight algorithms to solve them, respectively. Extensive experiments under various settings are carried out to corroborate the efficacy of SHARE. Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yongmin Zhang, Yue-Zhi Zhou, Yaoxue Zhang, Yuanyuan Yang 0001 |
ICDCS | 1 |
| 2021 | FAIR: Quality-Aware Federated Learning with Precise User Incentive and Model AggregationabstractFederated learning enables distributed learning in a privacy-protected manner, but two challenging reasons can affect learning performance significantly. First, mobile users are not willing to participate in learning due to computation and energy consumption. Second, with various factors (e.g., training data size/quality), the model update quality of mobile devices can vary dramatically, inclusively aggregating low-quality model updates can deteriorate the global model quality. In this paper, we propose a novel system named FAIR, i.e., Federated leArning with qualIty awaReness. FAIR integrates three major components: 1) learning quality estimation: we leverage historical learning records to estimate the user learning quality, where the record freshness is considered and the exponential forgetting function is utilized for weight assignment; 2) quality-aware incentive mechanism: within the recruiting budget, we model a reverse auction problem to encourage the participation of high-quality learning users, and the method is proved to be truthful, individually rational, and computationally efficient; and 3) model aggregation: we devise an aggregation algorithm that integrates the model quality into aggregation and filters out non-ideal model updates, to further optimize the global learning model. Based on real-world datasets and practical learning tasks, extensive experiments are carried out to demonstrate the efficacy of FAIR. Yongheng Deng, Feng Lyu 0001, Ju Ren 0001, Yi-Chao Chen 0001, Peng Yang 0004, Yue-Zhi Zhou, Yaoxue Zhang |
INFOCOM | 1 |
| 2019 | Toward Fast and Distributed Computation Migration System for Edge Computing in IoTabstractCode offload has been a key technique that provides new opportunities to achieve high performance with edge devices of weak computation capability. However, the implementation of code offload system on edge devices today is largely depending on backbone routers or cloud servers, imposes a significant burden on network traffic while ignores the potentials of utilizing Internet-of-Thing (IoT) devices physically nearby. In this article, we seek to obviate such limitations in improper code offload on edge devices, preferably focusing on minimizing interdomain data transfer and cross-device I/O cost while maximizing the CPU computation capability. First, we conduct the most in-depth and efficient analysis of both performance optimization and extra system burden. Second, we identify a new research problem of flexibly selecting the computation-intensive codes to offload at OS runtime across heterogeneous hardwares. Third, our measurement findings lead us to design and implement a fast and distributed code offload principle for edge devices called codeSpec, that shifts the destined devices from interdomain servers to IoT devices nearby, and only offloads binary code of user-specified regions across different instruction set architectures, thus to not only improve the system performance (by up to 80%) but also reduce burdens on data transfer (by at least 61%) and I/O latency (by on average 84%) than existing. codeSpec also makes user to flexibly specify the codes for offloading by plugging-in features, and is friendly to the commercial-off-the-shelf devices. Chao Wu 0002, Yaoxue Zhang, Yongheng Deng |
IEEE Internet Things J. | 3 |