EDBT 2026 Demo / reviewers in the wild / expert
Sheng Yue 0001
dblp:236/3241-1
· DBLP profile ↗
29ranked-venue papers
10as first author
29since 2021 · last 2026
0009-0001-3416-8181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 6 first-author · 22 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FORLER: Federated Offline Reinforcement Learning with Q-Ensemble and Actor Rectification
Nan Qiao, Sheng Yue 0001 |
ICC | 2 |
| 2026 | Two-Dimensional Stackelberg Game-Based Incentive Mechanism for Differential Private Federated Learning With Non-IID DataabstractIncentive mechanisms are essential for boosting client engagement in differential private federated learning (DP-FL). However, existing Stakelberg games-based incentive mechanisms typically assume that client decisions are one-dimension and that data is independent and identically distributed (IID) across clients. In reality, data distributions are often non-IID and clients have two-dimensional resources decisions, including data quantity and privacy. Therefore, in this paper, we present a novel two-dimensional Stackelberg game-based incentive mechanism for DP-FL with non-IID data, aiming to maximize the total utility of clients and server by seeking a balance between the clients' two-dimensional decisions and the server's payment. Specifically, we first formulate the utility functions of both server and clients under two-dimensional decisions and then model the interactions between server and clients as a single-leader-multiple-followers Stackelberg game. To derive the optimal decisions that maximize their utilities, we theoretically prove the existence of a Stackelberg equilibrium between server and clients. Due to the difficulty to directly calculate the Stackelberg equilibrium, we propose a bi-level multi-agent reinforment learning algorithm to learn the optimal decisions for both server and clients by trial and error. Extensive simulation results demonstrate that our proposed method outperforms the baselines in terms of total utility. Dan Wang 0031, Xiaoyi Pang, Jiahui Hu 0001, Sheng Yue 0001, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | HyStream: A Hybrid System for Application Streaming via Predictive Delivery and Sequence-Linearized CachingabstractTraditional application delivery requires full local installation, introducing persistent security risks from outdated software and imposing significant download delays. While advances in network bandwidth and latency have made remote content delivery more viable, existing dynamic loading mechanisms, such as network filesystems, often remain constrained by performance bottlenecks. Worse still, these solutions degrade sharply under variable or weak connectivity, where untimely code delivery can stall execution altogether. We propose HyStream, a hybrid application streaming system that combines predictive remote delivery with local sequence-linearized caching to sustain responsive and robust execution without requiring installation. HyStream addresses three key challenges: (1) maintaining microsecond-level latency comparable to local storage; (2) bridging the semantic gap between stateless remote storage and stateful execution; and (3) mitigating the limitations of purely network-based solutions under degraded connectivity. To achieve this, HyStream integrates three core components: a dual-mode transmission mechanism that decouples synchronous demand-driven requests from asynchronous speculative prefetching; a thread-aware Markov-chain model that captures fine-grained, concurrent access patterns for accurate prediction; and a sequence-linearized cache that persists streamed blocks in predicted execution order to support deterministic fallback behavior. Together, these components transform irregular, latency-sensitive I/O into efficient structured access that masks network variability. Evaluation shows HyStream delivers near-native performance across diverse networks. On mobile devices, it achieves 16–29% better per-page access latency than local UFS3.1, even over variable Wi-Fi connectivity. On desktops, it typically sustains startup overheads below 30% relative to local NVMe. Under variable and degraded network conditions, the sequence-linearized cache increasingly serves execution-critical accesses, rendering application performance largely insensitive to network latency and jitter within intra-city and inter-city deployments. Sheng Yue 0001, Xiang Liu 0017, Yongjian Fu 0004, Jialin Li 0001 |
IEEE Trans. Netw. | 2 |
| 2026 | Toward Communication-Efficient and Data-Free Collaborative Fine-Tuning Between Small and Large Language ModelsabstractWhile large language models (LLMs) exhibit impressive general capabilities, their performance on domainspecific tasks often requires fine-tuning with private data that cannot be shared due to privacy constraints. Directly deploying LLMs on resource-constrained clients for local fine-tuning is impractical due to their significant computation and communication costs. In addition, pre-trained LLMs are valuable intellectual property, and model owners are reluctant to distribute full model weights. To address these challenges, we proposeCoT-LM, a communication-efficient, computation-light, and data-free framework for collaborative fine-tuning between small (SLMs) and large language models (LLMs). InCoT-LM, clients fine-tune lightweight SLMs locally without uploading models or private data. These SLMs provide task-specific feedback to guide server-side LLM enhancement via an efficient communication protocol that exchanges only lightweight synthetic data and feedback. The framework supports both synchronous and asynchronous collaboration and enables mutual enhancement: the LLM improves its task-specific capabilities, while clients benefit from refined synthetic data or distilled knowledge. Extensive experiments demonstrate thatCoT-LMsignificantly boosts natural language understanding (NLU, up to 18.3% for LLMs and 8.0% for SLMs) and natural language generation (NLG, up to 31.7% for LLMs) performance across diverse tasks while preserving data privacy, model intellectual property, and generalization capabilities, achieving significant reductions in computation and communication overhead. Zhenya Ma, Yongheng Deng, Ziqing Qiao, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | FOVA: Offline Federated Reinforcement Learning With Mixed-Quality DataabstractOffline Federated Reinforcement Learning (FRL), a marriage of federated learning and offline reinforcement learning, has attracted increasing interest recently. Albeit with some advancement, we find that the performance of most existing offline FRL methods drops dramatically when provided with mixed-quality data, that is, the logging behaviors (offline data) are collected by policies with varying qualities across clients. To overcome this limitation, this paper introduces a new vote-based offline FRL framework, named FOVA. It exploits avote mechanismto identify high-return actions during local policy evaluation, alleviating the negative effect of low-quality behaviors from diverse local learning policies. Besides, building on advantage-weighted regression (AWR), we construct consistent local and global training objectives, significantly enhancing the efficiency and stability of FOVA. Further, we conduct an extensive theoretical analysis and rigorously show that the policy learned by FOVA enjoys strict policy improvement over the behavioral policy. Extensive experiments corroborate the significant performance gains of our proposed algorithm over existing baselines on widely used benchmarks. Nan Qiao 0008, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Netw. | 2 |
| 2025 | FocusX: All-in-Focus Image Synthesis for Dynamic Scenes on Mobile DevicesabstractWe propose FocusX, the first mobile-deployable system achieving artifact-free all-in-focus synthesis in dynamic scenes. Our approach introduces three key innovations: 1) For focal stack acquisition, our depth prior-based dynamic focusing method that adaptively selects focus distances using real-time scene depth distribution analysis and depth-of-field constrained spatial clustering, reducing redundant captures while ensuring full depth coverage; 2) To reduce pixel misalignment caused by lens breathing, we adopt a one-time offline calibration to map the relationship between field-of-view and focus distance, aligning the images by cropping accordingly; 3) We design the Diff-MotionAIFNet, a conditional diffusion-based model that decouples moving-static components for artifact-free AIF reconstruction in dynamic scene while preserving scene fidelity. We further contribute DynaAIFSet, containing 5,500 dynamic scenes (120K images) for training and evaluation. Experiments show FocusX achieves state-of-the-art performance, outperforming baselines up by 59.6% in SSIM and 49.1% in PSNR, respectively. The deployment latency of FocusX is 4.8s on Honor Magic7 Pro. This work bridges computational photography theory with mobile implementation constraints, delivering practical AIF enhancement for user-generated content. Pengkai Li, Fengzu Li, Wei Gao 0006, Sheng Yue 0001, Yaoxue Zhang, Ju Ren 0001 |
MobiCom | 6 |
| 2025 | Towards Distance-Adaptive Wireless ChargingabstractWireless charging holds significant promise for IoT devices and transportation networks by facilitating convenient and autonomous power supply. Traditional wireless charging technologies have typically adhered to a singular approach, choosing between near-field coupling or far-field radiation. However, our investigations uncover that each method outperforms the other at specific distances. This insight leads us to integrating the advantages of both to enable rapid wireless charging across any distance within the charging range. For this vision, we poses an intriguing question: "Can we develop a system that supports both near-field and far-field charging simultaneously?" Shuning Wang, Linghui Zhong, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 6 |
| 2025 | StreamSys: A Lightweight Executable Delivery System for Edge ComputingabstractEdge computing brings several challenges when it comes to data movement. First, moving large data from edge devices to the server is likely to waste bandwidth. Second, complex data patterns (e.g., traffic cameras) on devices require flexible handling. An ideal approach is to move code to data instead. However, since only a small portion of code is required, moving the executable as well as their libraries to the devices can be an overkill. While loading code on demand from remote such as NFS can be a stopgap, but on the other hand leads to low efficiency for irregular access patterns. This article presentsStreamSys, a lightweight executable delivery system that loads code on demand by redirecting the local disk IO to the server through optimized network IO. We employ a Markov-based prefetch mechanism on the server side. It learns the access pattern of code and predicts the block sequence for the client to reduce the network round trip. Meanwhile, server-sideStreamSysasynchronously prereads the block sequence from the disk to conceal disk IO latency beforehand. Evaluation shows that the latency ofStreamSysis up to 71.4% lower than the native Linux file system based on SD card and up to 62% lower than NFS in wired environments. Zhenya Ma, Yinggang Gao, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Cloud Comput. | 4 |
| 2025 | MAML-RAL: Learning Domain-Invariant HOI Rules for Real-Time Video MattingabstractReal-time video matting is essential for applications like online video conferencing but faces challenges in human-object interaction (HOI) scenarios, known as the HOI-matting problem. This problem is challenging due to its open-recognition nature, where no dataset can cover the wide range of potential HOI cases, making it difficult for feature-learning-based methods to generalize effectively. To address this issue, we present an HOI-matting dataset and introduce a Model-Agnostic Meta-Learning-based rule-aware learning approach (MAML-RAL). MAML-RAL combines transfer learning and meta-learning to capture domain-invariant HOI rules, complemented by a fast local adaptation strategy to counter domain shifts and background interference. Our method achieves a mean intersection-over-union (mIoU) of 92.3%, outperforming current algorithms, with local adaptation further boosting performance to a remarkable mIoU of 95.84%. Jiang Xin, Sheng Yue 0001, Ju Ren 0001, Feng Qian 0001, Yaoxue Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MASA: Multimodal Federated Learning Through Modality-Aware and Secure AggregationabstractAs a promising paradigm, federated learning has been applied to multimodal sensing tasks due to its deployment convenience. However, the recent advances in multimodal federated learning emphasize learning a high-quality multimodal model but overlook the model usage requirements of massive unimodal clients. Moreover, the privacy risk in model sharing and client data heterogeneity impact the efficacy of federated learning. In this paper, we propose a novel multimodal federated learning system named MASA. As a departure from existing approaches, MASA simultaneously enhances the model learning efficiency of both multimodal and unimodal clients while ensuring their data privacy. First, we employ a gated cross-modal distillation scheme to achieve performance-aware knowledge transfer across modality-heterogeneous clients. To enhance the system security, MASA integrates a lightweight split-shuffle mechanism to realize the anonymization and encryption of model aggregation. Moreover, to reach personalized collaboration while protecting privacy, MASA features an attention-based spontaneous client clustering mechanism to form client cluster structures securely and distributedly. We evaluate our MASA on four public multimodal datasets for human activity recognition. The results show that our MASA outperforms leading multimodal federated learning methods on the model performance of both multimodal and unimodal clients. Jialin Guo, Yongjian Fu 0004, Zhiwei Zhai, Xinyi Li 0005, Yongheng Deng, Sheng Yue 0001, Hao Pan 0003, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | DualRec: A Collaborative Training Framework for Device and Cloud Recommendation ModelsabstractRecommendation systems (RS) play a vital role in various domains. However, under recent data regulations like General Data Protection Regulation (GDPR), traditional RS that rely on collecting user's interaction data centrally face significant challenges. Federated learning (FL) enables collaborative model training among users while keeping their private data locally. Yet, the constrained resources of devices often limit the size of the learned model, resulting in suboptimal recommendation performance. To overcome the dilemma of data accessibility and model size, we propose DualRec, a novel collaborative training framework for device and cloud recommendation models. In DualRec, users train lightweight models on devices to harness their local private data, while a larger model is simultaneously trained on the cloud server to exploit its substantial resources. Devices and the cloud server collaboratively train their models, compensating for individual limitations of model size and data availability, enabling mutual empowerment and benefits. Specifically, we introduce an efficient aggregation mechanism for recommendation models to boost the collaborative training performance of device models. With the learned device models, we propose to generate pseudo user interaction data to train the server model. To enhance the training performance of the server model, we design an automated denoising mechanism to mitigate the negative impact of noisy samples in the generated pseudo dataset. Finally, the learned knowledge of the server model is distilled to device models for enhanced on-device recommendation performance. Extensive experiments demonstrate the superior performance of DualRec compared to state-of-the-art baselines. Ye Zhang 0033, Yongheng Deng, Sheng Yue 0001, Qiushi Li 0002, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Momentum-Based Contextual Federated Reinforcement LearningabstractFederated Reinforcement Learning (FRL) is an attractive edge learning paradigm for decision-making applications, which has garnered significant interest recently. However, owing to the inherent spatio-temporal non-stationarity of local state-action distributions, current FRL approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL method, which incorporates momentum, importance sampling, and server-side adjustments, capable of controlling the gradient shifts induced by the non-stationary data. We prove that by proper selection of momentum parameters and interaction frequency, it can achieve$\tilde {\mathcal {O}}(H N^{-1}\epsilon ^{-3/2})$and$\tilde {\mathcal {O}}(\epsilon ^{-1})$interaction and communication complexities (N represents the agent number), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns with the best achievable among existing first-order FL algorithms. Further, we leverage attention-based contextual representation extraction to enable the learning policy to adapt to heterogeneous tasks and environments. Extensive experiments demonstrate that our proposed method significantly outperforms existing baselines on a range of complex, high-dimensional single-task and multi-task benchmarks. Sheng Yue 0001, Xingyuan Hua, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Netw. | 1 |
| 2025 | AugFL: Augmenting Federated Learning With Pretrained ModelsabstractFederated Learning (FL) has garnered widespread interest in recent years. However, owing to strict privacy policies or limited storage capacities of training participants such as IoT devices, its effective deployment is often impeded by the scarcity of training data in practical decentralized learning environments. In this paper, we study enhancing FL with the aid of (large) pre-trained models (PMs), that encapsulate wealthy general/domain-agnostic knowledge, to alleviate the data requirement in conducting FL from scratch. Specifically, we consider a networked FL system formed by a central server and distributed clients. First, we formulate the PM-aided personalized FL as a regularization-based federated meta-learning problem, where clients join forces to learn a meta-model with knowledge transferred from a private PM stored at the server. Then, we develop an inexact-ADMM-based algorithm, AugFL, to optimize the problem with no need to expose the PM or incur additional computational costs to local clients. Further, we establish theoretical guarantees for AugFL in terms of communication complexity, adaptation performance, and the benefit of knowledge transfer in general non-convex cases. Extensive experiments corroborate the efficacy and superiority of AugFL over existing baselines. Sheng Yue 0001, Zerui Qin, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang, Junshan Zhang |
IEEE Trans. Netw. | 1 |
| 2024 | How to Leverage Diverse Demonstrations in Offline Imitation LearningabstractOffline Imitation Learning (IL) with imperfect demonstrations has garnered increasing attention owing to the scarcity of expert data in many real-world domains. A fundamental problem in this scenario is *how to extract positive behaviors from noisy data*. In general, current approaches to the problem select data building on state-action similarity to given expert demonstrations, neglecting precious information in (potentially abundant) *diverse* state-actions that deviate from expert ones. In this paper, we introduce a simple yet effective data selection method that identifies positive behaviors based on their *resultant states* - a more informative criterion enabling explicit utilization of dynamics information and effective extraction of both expert and beneficial diverse behaviors. Further, we devise a lightweight behavior cloning algorithm capable of leveraging the expert and selected data correctly. In the experiments, we evaluate our method on a suite of complex and high-dimensional offline IL benchmarks, including continuous-control and vision-based tasks. The results demonstrate that our method achieves state-of-the-art performance, outperforming existing methods on **20/21** benchmarks, typically by **2-5x**, while maintaining a comparable runtime to Behavior Cloning (BC). Sheng Yue 0001, Jiani Liu 0005, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang |
ICML | 1 |
| 2024 | OLLIE: Imitation Learning from Offline Pretraining to Online FinetuningabstractIn this paper, we study offline-to-online Imitation Learning (IL) that pretrains an imitation policy from static demonstration data, followed by fast finetuning with minimal environmental interaction. We find the naive combination of existing offline IL and online IL methods tends to behave poorly in this context, because the initial discriminator (often used in online IL) operates randomly and discordantly against the policy initialization, leading to misguided policy optimization and *unlearning* of pretraining knowledge. To overcome this challenge, we propose a principled offline-to-online IL method, named OLLIE, that simultaneously learns a near-expert policy initialization along with an *aligned discriminator initialization*, which can be seamlessly integrated into online IL, achieving smooth and fast finetuning. Empirically, OLLIE consistently and significantly outperforms the baseline methods in **20** challenging tasks, from continuous control to vision-based domains, in terms of performance, demonstration efficiency, and convergence speed. This work may serve as a foundation for further exploration of pretraining and finetuning in the context of IL. Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang |
ICML | 1 |
| 2024 | BR-DeFedRL: Byzantine-Robust Decentralized Federated Reinforcement Learning with Fast Convergence and Communication EfficiencyabstractIn this paper, we propose Byzantine-Robust Decentralized Federated Reinforcement Learning (BR-DeFedRL), an innovative framework that effectively combats the harmful influence of Byzantine agents by adaptively adjusting communication weights, thereby significantly enhancing the robustness of the learning system. By leveraging decentralized learning, our approach eliminates the dependence on a central server. Striking a harmonious balance between communication round count and sample complexity, BR-DeFedRL achieves efficient convergence with a rate of $\mathcal{O}\left( {\frac{1}{{TN}}} \right)$, where T denotes the communication rounds and N represents the local steps related to variance reduction. Notably, each agent attains an ϵ-approximation with a state-of-the-art sample complexity of $\mathcal{O}\left( {\frac{1}{{\varepsilon N}} + \frac{1}{\varepsilon }} \right)$. Extensive experimental validations further affirm the efficacy of BR-DeFedRL, making it a promising and practical solution for Byzantine-robust decentralized federated reinforcement learning. Jing Qiao, Zuyuan Zhang, Sheng Yue 0001, Yuan Yuan 0014, Zhipeng Cai 0001, Xiao Zhang 0015, Ju Ren 0001, Dongxiao Yu |
INFOCOM | 3 |
| 2024 | Momentum-Based Federated Reinforcement Learning with Interaction and Communication EfficiencyabstractFederated Reinforcement Learning (FRL) has garnered increasing attention recently. However, due to the intrinsic spatio-temporal non-stationarity of data distributions, the current approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL algorithm, named MFPO, that utilizes momentum, importance sampling, and additional server-side adjustment to control the shift of stochastic policy gradients and enhance the efficiency of data utilization. We prove that by proper selection of momentum parameters and interaction frequency, MFPO can achieve $\widetilde {\mathcal{O}}\left({H{N^{ - 1}}{\varepsilon ^{ - 3/2}}}\right)$ and $\widetilde {\mathcal{O}}\left({{\varepsilon ^{ - 1}}}\right)$ interaction and communication complexities (N represents the number of agents), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns the best achievable of existing first-order FL algorithms. Extensive experiments corroborate the substantial performance gains of MFPO over existing methods on a suite of complex and high-dimensional benchmarks. Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001 |
INFOCOM | 1 |
| 2024 | Federated Offline Policy Optimization with Dual RegularizationabstractFederated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the environment during local updating, which can be prohibitively expensive or even infeasible in many real-world domains. To overcome this challenge, this paper proposes a novel offline federated policy optimization algorithm, named DRPO, which enables distributed agents to collaboratively learn a decision policy only from private and static data without further environmental interactions. DRPO leverages dual regularization, incorporating both the local behavioral policy and the global aggregated policy, to judiciously cope with the intrinsic two-tier distributional shifts in offline FRL. Theoretical analysis characterizes the impact of the dual regularization on performance, demonstrating that by achieving the right balance thereof, DRPO can effectively counteract distributional shifts and ensure strict policy improvement in each federative learning round. Extensive experiments validate the significant performance gains of DRPO over baseline methods. Sheng Yue 0001, Zerui Qin, Xingyuan Hua, Yongheng Deng, Ju Ren 0001 |
INFOCOM | 1 |
| 2024 | RelayRec: Empowering Privacy-Preserving CTR Prediction via Cloud-Device Relay LearningabstractClick-through rate (CTR) prediction holds paramount importance across numerous applications, profoundly impacting user experience and business profitability. The freshness of a CTR prediction model significantly influences its performance, since users’ needs and interests may be changing over time, thereby requiring the model to be updated frequently. However, stringent data protection regulations have constrained the collection of users’ personal data, posing challenges to traditional model refreshing strategies that rely on centralized data collection. On-device learning techniques, such as federated learning (FL), offer a viable solution by enabling model training on devices without compromising user privacy. Nevertheless, the scarcity of training data with diverse distributions among devices presents considerable obstacles to on-device learning effectiveness. To address these challenges, we introduce RelayRec, a cloud-device relay learning framework designed for privacy-preserving CTR prediction. To establish competent initial models for devices, RelayRec categorizes pre-regulation cloud data into user preference groups, training preference-specific models for devices. Furthermore, a cloud-based automated model selector is developed to identify suitable initial models for devices. To elevate the relay learning performance of these initial models, we incorporate a personalized collaborative learning mechanism that aggregates device models based on user preferences. Extensive experimental evaluations underscore RelayRec’s superior performance compared to state-of-the-art benchmarks, affirming its efficacy in privacy-preserving CTR prediction. Yongheng Deng, Guanbo Wang, Sheng Yue 0001, Wei Rao 0003, Qin Zu, Ju Ren 0001, Yaoxue Zhang |
IPSN | 3 |
| 2024 | Dependent Task Offloading in Edge Computing Using GNN and Deep Reinforcement LearningabstractTask offloading is a widely used technology in Edge Computing (EC), which declines the makespan of user task with the aid of resourceful edge servers. How to solve the competition for computation and communication resources among tasks is a fundamental issue in task offloading. Besides, real-life user tasks often comprise multiple interdependent subtasks. Dependencies among subtasks significantly raises the complexity of task offloading, and makes it difficult to propose generalized approaches for scenarios of different size. In this paper, we study the Dependent Task Offloading (DTO) problem within both single-user single-edge and multi-user multi-edge scenario. First, we use Directed Acyclic Graph (DAG) to model dependent task, where nodes and directed edges represent the subtasks and their interdependencies respectively. Then, we propose a task scheduling method based on Graph Attention Network (GAT) and Deep Reinforcement Learning (DRL) to minimize the makespan of user tasks. More specifically, our method introduces a multi-discrete action DRL scheduler that simultaneously determines which subtask to consider and whether it should be offloaded at each step, and employs GAT to encode the graph-based state representation. To stabilize and speed up DRL scheduler training, we pretrain GAT encoder with unsupervised learning. Extensive experiments demonstrate that our proposed approach can be applied to various environments and outperforms prior methods. Zequn Cao, Xiaoheng Deng, Sheng Yue 0001, Ping Jiang 0001, Ju Ren 0001, Jinsong Gui |
IEEE Internet Things J. | 3 |
| 2024 | Towards Resource-Efficient Edge AI: From Federated Learning to Semi-Supervised Model PersonalizationabstractA central question in edge intelligence is “how can an edge device learn its local model with limited data and constrained computing capacity?” In this study, we explore the approach where a global model initialization is first obtained by running federated learning (FL) across multiple edge devices, based on which a semi-supervised algorithm is devised for a single edge device to carry out quick adaptation with its local data. Specifically, to account for device heterogeneity and resource constraints, a global model is first trained via FL, where each device conducts multiple local updates only for its customized subnet. A subset of devices can be selected to upload updates for aggregation during each training round. Further, device scheduling is optimized to minimize the training loss of FL, subject to resource constraints, based on the carefully crafted reward function defined as the one-round progress of FL each device can provide. We examine the convergence behavior of FL for the general non-convex case. For semi-supervised model personalization, we use the FL-based model initialization as a teacher network to impute soft labels on unlabeled data, thereby addressing the insufficiency of labeled data. Experiments are conducted to evaluate the performance of the proposed algorithms. Sheng Yue 0001, Junshan Zhang |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | SESAME: A Resource Expansion and Sharing Scheme for Multiple Edge Services ProvidersabstractAs a potential computing solution for fast-growing mobile and IoT applications, edge computing has been developed rapidly. However, due to the relatively limited resources of each edge node, it is difficult for edge nodes to provide quality-guaranteed services to dynamic and massive computation tasks individually. To address this challenge, this paper proposes a two-stage resource expansion and sharing scheme, named SESAME, to enable resource sharing among the edge nodes within/across multiple edge service providers (ESPs), to improve the overall efficiency of the edge computing system. To facilitate the operation and reduce complexity, the resource management scheme has both long-term and short-term decision periods. During the long-term period, an optimal conservative estimation-based resource expansion and pricing strategy has been designed to ensure the system stability and the interests of ESPs. During the short-term period, a resource-sharing strategy considering the internal and external behaviors of ESPs has been proposed to reduce resource-sharing costs while fully utilizing internal resources. In such a way, the resources from different ESPs can collaborate efficiently. Extensive experiments on real datasets show that our algorithm can effectively reduce ESP costs and improve system stability. Jiani Liu 0005, Ju Ren 0001, Yongmin Zhang, Sheng Yue 0001, Yaoxue Zhang |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | PoPeC: PAoI-Centric Task Offloading With Priority Over Unreliable ChannelsabstractFreshness-aware computation offloading has garnered increasing attention recently in the realm of edge computing, driven by the need to promptly obtain up-to-date information and mitigate the transmission of outdated data. However, most of the existing works assume that channels are reliable, neglecting the intrinsic fluctuations and uncertainty in wireless communication. More importantly, offloading tasks typically have diverse freshness requirements. Accommodation of various task priorities in the context of freshness-aware task scheduling and resource allocation remains an open and unresolved problem. To overcome these limitations, we cast the freshness-aware task offloading problem as a multi-priority optimization problem, considering the unreliability of wireless channels, prioritized users, and the heterogeneity of edge servers. Building upon the nonlinear fractional programming and the ADMM-Consensus method, we introduce a joint resource allocation and task offloading algorithm to solve the original problem iteratively. In addition, we devise a distributed asynchronous variant for the proposed algorithm to further enhance its communication efficiency. We rigorously analyze the performance and convergence of our approaches and conduct extensive simulations to corroborate their efficacy and superiority over the existing baselines. Nan Qiao 0008, Sheng Yue 0001, Yongmin Zhang, Ju Ren 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement Learning
Sheng Yue 0001, Guanbo Wang, Wei Shao 0006, Sen Lin 0001, Ju Ren 0001, Junshan Zhang |
ICLR | 1 |
| 2023 | FedINC: An Exemplar-Free Continual Federated Learning Framework with Small Labeled DataabstractFederated learning (FL) has shown great promise for privacy-preserving learning by enabling collaborative training on decentralized clients. However, in realistic FL scenarios, clients often collect new data continuously, join or exit learning dynamically. As a result, the global model tends to forget old knowledge while learning new knowledge. Meanwhile, labeling the continuously arriving data in real-time is usually challenging. Therefore, the catastrophic forgetting problem intertwined with the label deficiency issue poses significant challenges for both learning new knowledge and consolidating old knowledge. To address these challenges, we develop a novel exemplar-free continual federated learning framework named FedINC, to learn a global incremental model with limited labeled data. We begin by excavating the cause of catastrophic forgetting via in-depth empirical studies. Based on that, we introduce targeted mechanisms for FedINC, including a hybrid contrastive learning mechanism to efficiently learn new knowledge with limited labeled data, a plastic feature regularization mechanism to preserve old task's representation space, a prototype-guided regularization mechanism to mitigate feature overlap between old and new classes while aligning the features of non-iid clients, and a prototype evolution mechanism for flexible and efficient incremental classification. Extensive experiments demonstrate the superior performance of FedINC in terms of both convergence speed and accuracy of the global model. Yongheng Deng, Sheng Yue 0001, Tuowei Wang, Guanbo Wang, Ju Ren 0001, Yaoxue Zhang |
SenSys | 2 |
| 2022 | HSFL: An Efficient Split Federated Learning Framework via Hierarchical OrganizationabstractFederated learning (FL) has emerged as a popular paradigm for distributed machine learning among vast clients. Unfortunately, resource-constrained clients often fail to participate in FL because they cannot pay for the memory resources required for model training due to their limited memory or bandwidth. Split federated learning (SFL) is a novel FL framework in which clients commit intermediate results of model training to a cloud server for client-server collaborative training of models, making resource-constrained clients also eligible for FL. However, existing SFL frameworks mostly require frequent communication with the cloud server to exchange intermediate results and model parameters, which results in significant communication overhead and elongated training time. In particular, this can be exacerbated by the imbalanced data distributions of clients. To tackle this issue, we propose HSFL, a hierarchical split federated learning framework that efficiently trains SFL model through hierarchical organization participants. Under the HSFL framework, we formulate a Cloud Aggregation Time Minimization (CATM) problem to minimize the global training time and design a light-weight client assignment algorithm based on dynamic programming to solve it. Moreover, we develop a self-adaption approach to cope with the dynamic computational resources of clients. Finally, we implement and evaluate HSFL on various real-world training tasks, elaborating on its effectiveness and superiority in terms of efficiency and accuracy compared to baselines. Tengxi Xia, Yongheng Deng, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
CNSM | 3 |
| 2022 | Efficient Federated Meta-Learning Over Multi-Access Wireless NetworksabstractFederated meta-learning (FML) has emerged as a promising paradigm to cope with the data limitation and heterogeneity challenges in today’s edge learning arena. However, its performance is often limited by slow convergence and corresponding low communication efficiency. In addition, since the available radio spectrum and IoT devices’ energy capacity are usually insufficient, it is crucial to control the resource allocation and energy consumption when deploying FML in practical wireless networks. To overcome the challenges, in this paper, we rigorously analyze the contribution of each device to the global loss reduction in each round and develop an FML algorithm (called NUFM) with a non-uniform device selection scheme to accelerate the convergence. After that, we formulate a resource allocation problem integrating NUFM in multi-access wireless systems to jointly improve the convergence rate and minimize the wall-clock time along with energy cost. By deconstructing the original problem step by step, we devise a joint device selection and resource allocation strategy to solve the problem with theoretical guarantees. Further, we show that the computational complexity of NUFM can be reduced from$O(d^{2})$to$O(d)$(with the model dimension$d$) via combining two first-order approximation techniques. Extensive simulation results demonstrate the effectiveness and superiority of the proposed methods in comparison with existing baselines. Sheng Yue 0001, Ju Ren 0001, Jiang Xin, Yaoxue Zhang, Weihua Zhuang |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | TODG: Distributed Task Offloading With Delay Guarantees for Edge ComputingabstractEdge computing has been an efficient way to provide prompt and near-data computing services for resource-and-delay sensitive IoT applications via computation offloading. Effective computation offloading strategies need to comprehensively cope with several major issues, including 1) the allocation of dynamic communication and computational resources, 2) delay constraints of heterogeneous tasks, and 3) requirements for computationally inexpensive and distributed algorithms. However, most of the existing works mainly focus on part of these issues, which would not suffice to achieve expected performance in complex and practical scenarios. To tackle this challenge, in this paper, we systematically study a distributed computation offloading problem with delay constraints, where heterogeneous computational tasks require continually offloading to a set of edge servers via a limiting number of stochastic communication channels. The task offloading problem is formulated as a delay-constrained long-term stochastic optimization problem under unknown prior statistical knowledge. To solve this problem, we first provide a technical path to transform and decompose it into several slot-level sub-problems. Then, we devise a distributed online algorithm, namely TODG, to efficiently allocate resources and schedule offloading tasks. Further, we present a comprehensive analysis for TODG in terms of the optimality gap, the worst-case delay, and the impact of system parameters. Extensive simulation results demonstrate the effectiveness and efficiency of TODG. Sheng Yue 0001, Ju Ren 0001, Nan Qiao 0008, Yongmin Zhang, Hongbo Jiang 0001, Yaoxue Zhang, Yuanyuan Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Inexact-ADMM Based Federated Meta-Learning for Fast and Continual Edge LearningabstractIn order to meet the requirements for performance, safety, and latency in many IoT applications, intelligent decisions must be made right here right now at the network edge. However, the constrained resources and limited local data amount pose significant challenges to the development of edge AI. To overcome these challenges, we explore continual edge learning capable of leveraging the knowledge transfer from previous tasks. Aiming to achieve fast and continual edge learning, we propose a platform-aided federated meta-learning architecture where edge nodes collaboratively learn a meta-model, aided by the knowledge transfer from prior tasks. The edge learning problem is cast as a regularized optimization problem, where the valuable knowledge learned from previous tasks is extracted as regularization. Then, we devise an ADMM based federated meta-learning algorithm, namely ADMM-FedMeta, where ADMM offers a natural mechanism to decompose the original problem into many subproblems which can be solved in parallel across edge nodes and the platform. Further, a variant of inexact-ADMM method is employed where the subproblems are 'solved' via linear approximation as well as Hessian estimation to reduce the computational cost per round to O(n). We provide a comprehensive analysis of ADMM-FedMeta, in terms of the convergence properties, the rapid adaptation performance, and the forgetting effect of prior knowledge transfer, for the general non-convex case. Extensive experimental studies demonstrate the effectiveness and efficiency of ADMM-FedMeta, and showcase that it substantially outperforms the existing baselines. Sheng Yue 0001, Ju Ren 0001, Jiang Xin, Sen Lin 0001, Junshan Zhang |
MobiHoc | 1 |