Jie Zhang 0076

dblp:84/6889-76 · DBLP profile ↗
← Back
58ranked-venue papers
13as first author
44since 2021 · last 2026
0000-0002-8073-2118ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 3 first-author · 25 since 2021Computer networks · 16 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning Models
abstract
Recent Large Reasoning Models (LRMs) excel at complex reasoning tasks but often suffer from overthinking, generating overly long and redundant reasoning trajectories.To explore its essence, our empirical analysis reveals that LRMs are primarily limited to recognizing task properties (i.e., difficulty levels) like humans before solving the problem, leading to a one-size-fits-all reasoning strategy.This observation motivates a fundamental question: Can we explicitly bootstrap such ability to alleviate overthinking in LRMs?To this end, we propose Think-How-to-Think (TH2T), a novel two-stage fine-tuning strategy that progressively inspires LRMs' difficulty cognition and redundancy cognition of LRMs.Specifically, we first inject Difficulty Dypnosis into output prefixes as cues for global, prospective reasoning strategy selection, stimulating the model's sharper sensitivity to task complexity and adaptive control of reasoning depth.Then, we incorporate Redundancy Hypnosis into inprogress reasoning steps, which serve as local, retrospective signals for behavior correction by identifying and eliminating superfluous reasoning detours.Experiments across 7B/14B/32B models demonstrate that TH2T significantly reduces inference costs by over 70% on easy tasks and 40% on complex ones without compromising performance.The resultant models exhibit a nascent ability for difficulty-aware reasoning, effectively mitigating behaviors like excessive reflection and looping, thereby paving the way for more cognitively efficient LRMs.
Yongjiang Liu, Haoxi Li, Xiaosong Ma, Jie Zhang 0076, Song Guo 0001
ACL (1)4
2026 LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
abstract
Large Language Models (LLMs) exhibit enhanced capabilities by Chain-of-Thought reasoning.However, the extended reasoning sequences introduce significant GPU memory overhead due to increased key-value (KV) cache.Existing KV cache compression methods mitigate memory bottlenecks but struggle in long reasoning tasks.In this paper, we analyze attention patterns in reasoning tasks and reveal a Token Importance Recurrence phenomenon: a large proportion of tokens regain high attention after multiple decoding steps, which is failed to capture by existing works and may lead to unpredictable eviction on such periodically critical tokens.To address this, we propose LazyEviction, an observation windowbased lagged eviction framework retaining latent recurring tokens by an eviction policy informed by token recurrence patterns.Extensive experiments demonstrate that LazyEviction reduces KV cache by 50%~70% while maintaining comparable accuracy, outperforming existing KV cache compression baselines.Our implementation code can be found at https: //github.com/Halo-949/LazyEviction.
Xiaosong Ma, Jie Zhang 0076, Song Guo 0001
ACL (1)4
2026 Optimizing Prompts With MLLM for Multimodal Product Style Recognition in Industry 5.0
abstract
The transition to Industry 5.0, supported by the high-bandwidth and low-latency capabilities of 6G networks, accelerates the demand for personalized products, making efficient and accurate product style recognition a critical task. However, traditional recognition models, which rely heavily on large-labeled datasets, struggle with generalization and efficiency in such data-rich environments. In contrast, multimodal large language models (MLLMs), such as Qwen-VL, offer new possibilities by integrating visual and textual inputs, requiring less annotated data and demonstrating strong generalization. Nevertheless, most existing MLLMs are general-purpose and lack optimized prompts tailored for domain-specific tasks, limiting their recognition performance. To address this, we propose a prompt optimization framework that combines genetic algorithms (GA) and chain of thought (CoT) reasoning. GA is used to iteratively refine the instruction component of prompts, while CoT encourages step-by-step reasoning, improving the model’s interpretability and recognition accuracy. Validated on a custom-built apparel dataset and the public FashionStyle14 dataset, our method enables Qwen-VL to achieve significant accuracy improvements of 8.7% and 19.6%, respectively, over the baseline. This study makes three contributions: a framework to adapt general MLLMs for fine-grained style recognition; a demonstration of 6G-enabled real-time artificial intelligence reasoning loops; and a high-quality, multidimensional style dataset for future research. This demonstrates the effectiveness and stability of the proposed framework in handling fine-grained style recognition tasks, contributing to the development of intelligent design systems for Industry 5.0.
Mengxue Li, Hujiang Huang, Yan Hong 0002, Ziqiang Cao, Jie Zhang 0076, Song Guo 0001
IEEE Trans. Ind. Informatics5
2026 AFedLF: Adaptive Layer Freezing of Foundation Models in Heterogeneous Federated Learning
abstract
The rise of pre-trained foundation models (FMs) has popularized the trend of fine-tuning FMs to fit downstream tasks, while Federated Learning (FL) has become the de-facto approach for training distributed data with privacy-preservation. However, fine-tuning FMs in FL faces overwhelming overheads due to its bulky nature. While freezing parameters in FM have the potential to accelerate FL training, existing freezing strategies statically freeze parameters on specified or already converged layers, incur severe accuracy degradation, and resource-inefficiency in heterogeneous environments. In this paper, we propose AFedLF, an adaptive freezing framework for FM in FL, to accelerate its wall-clock time for convergence without losing its final accuracy. However, this poses great challenges, as different freezing strategies lead to different accuracy gains and time overheads, while unfreezing more layers may bring marginal accuracy gains but significant time overheads. To address this challenge, AFedLF mathematically establishes a correlation between the freezing strategy and the accuracy gain and time overhead, and allocates adaptive freezing strategies to clients, based on our insight that unfreezing more layers on devices with strong computation and communication capabilities helps improve resource efficiency. Besides, AFedLF incorporates our well-designed intermediate result caching scheme with constant approximation ratios utilizing the limited storage capacity on mobile devices to cache intermediate results to skip forward propagation, further saving wall-clock time. Finally, we implemented AFedLF using an open-source FL benchmark, and extensive trace-driven experimental results showed that AFedLF accelerates wall-clock time by up to 6.1× compared to state-of-the-art solutions, without sacrificing accuracy.
Yue Zeng 0002, Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Zicong Hong, Bin Tang 0002, Junlong Zhou, Jiaying Yu
IEEE Trans. Mob. Comput.2
2025 VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather Forecasting
abstract
This paper presents Variables-Adaptive Mixture of Experts (VA-MoE), a novel framework for incremental weather forecasting that dynamically adapts to evolving spatiotemporal patterns in real-time data. Traditional weather prediction models often struggle with exorbitant computational expenditure and the need to continuously update forecasts as new observations arrive. VA-MoE addresses these challenges by leveraging a hybrid architecture of experts, where each expert specializes in capturing distinct sub-patterns of atmospheric variables (e.g., temperature, humidity, wind speed). Moreover, the proposed method employs a variable-adaptive gating mechanism to dynamically select and combine relevant experts based on the input context, enabling efficient knowledge distillation and parameter sharing. This design significantly reduces computational overhead while maintaining high forecast accuracy. Experiments on real-world ERA5 dataset demonstrate that VA-MoE performs comparable against state-of-the-art models in both short-term (e.g., 1–3 days) and long-term (e.g., 5 days) forecasting tasks, with only about 25\% of trainable parameters and 50\% of the initial training data.
Hao Chen 0045, Tao Han 0002, Song Guo 0001, Jie Zhang 0076, Yonghan Dong, Yue Yu 0001, Lei Bai 0001
ICCV4
2025 Causally Motivated Sycophancy Mitigation for Large Language Models
abstract
Incorporating user preferences into large language models (LLMs) can enhance the personalization and reliability of model outputs and facilitate the application of LLMs to real-world scenarios. However, leveraging user preferences can be a double-edged sword. Recent studies have found that improper utilization can incur sycophancy, where LLMs prioritize alignment with user preferences over the correctness of their outputs. To address sycophancy in LLMs, we analyze and model the problem through the lens of structured causal models (SCMs). We attribute sycophancy to LLMs' reliance on spurious correlations between user preferences and model outputs in this paper. Based on the proposed SCMs, we develop a novel framework, termed **CAUSM**, to mitigate sycophancy in LLMs by exploiting a significant causal signature. Specifically, we eliminate the spurious correlations embedded in the intermediate layers of LLMs through causally motivated head reweighting, and then calibrate the intra-head knowledge along the causal representation direction. Extensive experiments are conducted across diverse language tasks to demonstrate the superiority of our method over state-of-the-art competitors in mitigating sycophancy in LLMs.
Haoxi Li, Xueyang Tang, Jie Zhang 0076, Song Guo 0001, Sikai Bai, Peiran Dong, Yue Yu 0001
ICLR3
2025 Exploring Prosocial Irrationality for LLM Agents: A Social Cognition View
abstract
Large language models (LLMs) have been shown to face hallucination issues due to the data they trained on often containing human bias; whether this is reflected in the decision-making process of LLM agents remains under-explored. As LLM Agents are increasingly employed in intricate social environments, a pressing and natural question emerges: Can we utilize LLM Agents' systematic hallucinations to mirror human cognitive biases, thus exhibiting irrational social intelligence? In this paper, we probe the irrational behavior among contemporary LLM agents by melding practical social science experiments with theoretical insights. Specifically, we propose CogMir, an open-ended Multi-LLM Agents framework that utilizes hallucination properties to assess and enhance LLM Agents’ social intelligence through cognitive biases. Experimental results on CogMir subsets show that LLM Agents and humans exhibit high consistency in irrational and prosocial decision-making under uncertain conditions, underscoring the prosociality of LLM Agents as social entities and highlighting the significance of hallucination properties. Additionally, CogMir framework demonstrates its potential as a valuable platform for encouraging more research into the social intelligence of LLM Agents.
Xuan Liu 0001, Jie Zhang 0076, Haoyang Shang, Song Guo 0001, Chengxu Yang, Quanyan Zhu
ICLR2
2025 Uni-IL: Unified Incremental Learning of Vision-Language Models via Mixture of Attribute-Guided Experts
abstract
With the advent of parameter-efficient fine-tuning techniques for pre-trained vision-language models, interest in adapting them for various incremental learning scenarios has grown, i.e., sequential increments on task, class, and domain. However, no high-performance incremental learning framework has integrated these three incremental scenarios to achieve Unified Incremental Learning (Uni-IL) in complex settings. In this work, we propose an incremental learning framework called Mixture of Attribute-Guided Experts (MAGE) to alleviate the long-term forgetting in vision-language model incremental learning. Our approach involves acquiring image attribute knowledge via LLMs to form an attribute pool. We match the most relevant attributes as inputs to the Mixture of Experts (MoE) to fine-tune the pre-trained CLIP. Then the expert routers learn to select specific expert combinations based on the data and attribute features, alleviating catastrophic forgetting. The attribute pool incorporates both domain and class knowledge, enabling our approach to adapt to the three types of incremental learning scenarios and thus facilitating unified incremental learning. Through extensive experiments on our newly proposed benchmark and existing incremental learning scenarios, the results demonstrate that our proposed method not only performs well on the new Uni-IL tasks but also consistently outperforms previous state-of-the-art methods. Source code is available at https://github.com/ElectricField/Uni-IL.
Yufeng Zhan, Jie Zhang 0076, Yuanqing Xia
MMAsia3
2025 DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
abstract
Despite the significant breakthrough of Mixture-of-Experts (MoE), the increasing scale of these MoE models presents huge memory and storage challenges. Existing MoE pruning methods, which involve reducing parameter size with a uniform sparsity across all layers, often lead to suboptimal outcomes and performance degradation due to varying expert redundancy in different MoE layers. To address this, we propose a non-uniform pruning strategy, dubbed Differentiable Expert Pruning (DiEP), which adaptively adjusts pruning rates at the layer level while jointly learning inter-layer importance, effectively capturing the varying redundancy across different MoE layers. By transforming the global discrete search space into a continuous one, our method handles exponentially growing non-uniform expert combinations, enabling adaptive gradient-based pruning. Extensive experiments on five advanced MoE models demonstrate the efficacy of our method across various NLP tasks. Notably, \textbf{DiEP} retains around 92\% of original performance on Mixtral 8$\times$7B with only half the experts, outperforming other pruning methods by up to 7.1% on the challenging MMLU dataset.
Sikai Bai, Haoxi Li, Jie Zhang 0076, Zicong Hong, Song Guo 0001
NeurIPS3
2025 Metaverse-Oriented User Preference Recommendation Systems Based on DSD-Transformer
abstract
The Metaverse, with its promise of immersive experiences and transformative user interactions, represents a new paradigm for IoT development. By integrating IoT with the Metaverse, real-world data can seamlessly enrich virtual environments, offering diverse choices to users. However, the sheer volume of products and user groups in the Metaverse poses challenges in effectively matching users with suitable products. Recommendation systems, particularly Collaborative Filtering (CF), emerge as a solution to this issue, leveraging user preferences and social dynamics. However, traditional CF algorithms face efficiency challenges in the multi-dimensional data landscape of the Metaverse. To address this, a clustering-based CF algorithm is proposed, enhancing recommendation efficiency by leveraging social connections. Additionally, the recommendation system is enhanced with a DSD-Transformer framework, optimizing recommendation accuracy. The experiments indicate that our proposed method may considerably enhance the Metaverse experience when compare to various sophisticated methods and can be utilized to build a range of product recommendation systems.
Yan Hong 0002, Ru Rao, Xinping Li, Jie Zhang 0076, Xiaoqun Dai, Meng Zhang 0011, Song Guo 0001
IEEE Internet Things J.4
2025 Advanced Product Personalization in Blockchain-Enabled Metaverse: A Diffusion Model for Automatic Style Generation
abstract
The Metaverse is a user-generated virtual world, aiming to provide highly personalized experiences for users. A product personalization design platform is a critical direction for the Metaverse’s future development, enhancing user experience by offering personalized services. Blockchain technology ensures the security and privacy of user data, and enables personalized services through smart contracts, offering opportunities for personalization platforms. However, blockchain’s decentralization can lead to excessive product data, resulting in ineffective data management and optimization, subsequently confusing personalized product design styles, which diminishes user experience. To address these issues, this study proposes the product style automatic generation system (PSAGS), centered on an image generation unit. The system outputs images with style information based on the input product text, achieving precise quantization and visualization of product styles, thereby enhancing user engagement and loyalty to the Metaverse. The image generation unit, with a standardization module can standardize the product style, namely, the relationship between product design elements and user emotions, addressing the problem of managing vast style data due to blockchain’s decentralization. The generation module utilizes a diffusion model enhanced with contrastive language-image pretraining (CLIP) to generate style images, deepening the Metaverse experience. Optimizations include dilation convolution in the UNet architecture to enhance image quality and fine-grained CLIP transformations for improved image and text alignment. Results demonstrate the system’s effectiveness in streamlining design processes and improving image quality in personalized product design, with wide applications in the Metaverse.
Mengsi Li, Jie Zhang 0076, Yan Hong 0002, Xiangpeng Xie 0001, Meng Zhang 0011, Song Guo 0001
IEEE Internet Things J.2
2025 Collaborative Neural Architecture Search for Personalized Federated Learning
abstract
Personalized federated learning (pFL) is a promising approach to train customized models for multiple clients over heterogeneous data distributions. However, existing works on pFL often rely on the optimization of model parameters and ignore the personalization demand on neural network architecture, which can greatly affect the model performance in practice. Therefore, generating personalized models with different neural architectures for different clients is a key issue in implementing pFL in a heterogeneous environment. Motivated by Neural Architecture Search (NAS), a model architecture searching methodology, this paper aims to automate the model design in a collaborative manner while achieving good training performance for each client. Specifically, we reconstruct the centralized searching of NAS into the distributed scheme called Personalized Architecture Search (PAS), where differentiable architecture fine-tuning is achieved via gradient-descent optimization, thus making each client obtain the most appropriate model. Furthermore, to aggregate knowledge from heterogeneous neural architectures, a knowledge distillation-based training framework is proposed to achieve a good trade-off between generalization and personalization in federated learning. Extensive experiments demonstrate that our architecture-level personalization method achieves higher accuracy under the non-iid settings, while not aggravating model complexity over state-of-the-art benchmarks.
Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Zicong Hong, Yufeng Zhan, Qihua Zhou
IEEE Trans. Computers3
2025 Model Decomposition and Reassembly for Purified Knowledge Transfer in Personalized Federated Learning
abstract
Personalized federated learning (pFL) is to collaboratively train non-identical machine learning models for different clients to adapt to their heterogeneously distributed datasets. State-of-the-art pFL approaches pay much attention on exploiting clients’ inter-similarities to facilitate the collaborative learning process, meanwhile, can barely escape from the irrelevant knowledge pooling that is inevitable during the aggregation phase, and thus hindering the optimization convergence and degrading the personalization performance. To tackle such conflicts between facilitating collaboration and promoting personalization, we propose a novel pFL framework, dubbed pFedC, to first decompose the global aggregated knowledge into several compositional branches, and then selectively reassemble the relevant branches for supporting conflicts-aware collaboration among contradictory clients. Specifically, by reconstructing each local model into a shared feature extractor and multiple decomposed task-specific classifiers, the training on each client transforms into a mutually reinforced and relatively independent multi-task learning process, which provides a new perspective for pFL. Besides, we conduct a purified knowledge aggregation mechanism via quantifying the combination weights for each client to capture clients’ common prior, as well as mitigate potential conflicts from the divergent knowledge caused by the heterogeneous data. Extensive experiments over various models and datasets demonstrate the effectiveness and superior performance of the proposed algorithm.
Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Wenchao Xu 0001, Qihua Zhou, Jingcai Guo, Zicong Hong, Jun Shan
IEEE Trans. Mob. Comput.1
2025 Feature Correlation-Guided Knowledge Transfer for Federated Self-Supervised Learning
abstract
Extensive attention has been paid to the application of self-supervised learning (SSL) approaches on federated learning (FL) to tackle the label scarcity problem. Previous works on federated SSL (FedSSL) generally fall into two categories: parameter-based model aggregation or data-based feature sharing to achieve knowledge transfer among multiple unlabeled clients. Despite the progress, they inevitably rely on some assumptions, such as homogeneous models or the existence of an additional public dataset, which hinder the universality of the training frameworks for more general scenarios (e.g., unlabeled clients with heterogeneous models). Therefore, in this article, we propose a novel and general method named federated self-supervised learning with feature-correlation-based aggregation (FedFoA) to tackle the above limitations. By exchanging feature correlation instead of model parameters or feature mappings, our approach reduces the discrepancies of local representations learning processes, thus promoting collaboration between heterogeneous clients. A factorization-based method is designed to extract the cross-feature relation matrix from local representations, which serves as a knowledge medium for the aggregation phase. We demonstrate that FedFoA is a heterogeneity-supportive and privacy-preserving training framework and can be easily compatible with state-of-the-art FedSSL methods. Extensive empirical experiments demonstrate our proposed approach outperforms the state-of-the-art methods by a significant margin.
Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Yufeng Zhan, Qihua Zhou, Yingchun Wang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2025 Data Quality-Aware Mixed-Precision Quantization via Hybrid Reinforcement Learning
abstract
Mixed-precision quantization mostly predetermines the model bit-width settings before actual training due to the non-differential bit-width sampling process, obtaining suboptimal performance. Worse still, the conventional static quality-consistent training setting, i.e., all data is assumed to be of the same quality across training and inference, overlooks data quality changes in real-world applications which may lead to poor robustness of the quantized models. In this article, we propose a novel data quality-aware mixed-precision quantization framework, dubbed DQMQ, to dynamically adapt quantization bit-widths to different data qualities. The adaption is based on a bit-width decision policy that can be learned jointly with the quantization training. Concretely, DQMQ is modeled as a hybrid reinforcement learning (RL) task that combines model-based policy optimization with supervised quantization training. By relaxing the discrete bit-width sampling to a continuous probability distribution that is encoded with few learnable parameters, DQMQ is differentiable and can be directly optimized end-to-end with a hybrid optimization target considering both task performance and quantization benefits. Trained on mixed-quality image datasets, DQMQ can implicitly select the most proper bit-width for each layer when facing uneven input qualities. Extensive experiments on various benchmark datasets and networks demonstrate the superiority of DQMQ against existing fixed/mixed-precision quantization methods.
Yingchun Wang 0002, Song Guo 0001, Jingcai Guo, Yuanhong Zhang, Weizhan Zhang, Jie Zhang 0076
IEEE Trans. Neural Networks Learn. Syst.7
2024 Combating Data Imbalances in Federated Semi-supervised Learning with Dual Regulators
abstract
Federated learning has become a popular method to learn from decentralized heterogeneous data. Federated semi-supervised learning (FSSL) emerges to train models from a small fraction of labeled data due to label scarcity on decentralized clients. Existing FSSL methods assume independent and identically distributed (IID) labeled data across clients and consistent class distribution between labeled and unlabeled data within a client. This work studies a more practical and challenging scenario of FSSL, where data distribution is different not only across clients but also within a client between labeled and unlabeled data. To address this challenge, we propose a novel FSSL framework with dual regulators, FedDure. FedDure lifts the previous assumption with a coarse-grained regulator (C-reg) and a fine-grained regulator (F-reg): C-reg regularizes the updating of the local model by tracking the learning effect on labeled data distribution; F-reg learns an adaptive weighting scheme tailored for unlabeled instances in each client. We further formulate the client model training as bi-level optimization that adaptively optimizes the model in the client with two regulators. Theoretically, we show the convergence guarantee of the dual regulators. Empirically, we demonstrate that FedDure is superior to the existing methods across a wide range of settings, notably by more than 11% on CIFAR-10 and CINIC-10 datasets.
Sikai Bai, Weiming Zhuang, Jie Zhang 0076, Shuai Yi, Junyu Gao 0001
AAAI4
2024 On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks
abstract
The explosive growth of video traffic on today's Internet promotes the rise of Neural-enhanced Video Streaming (NeVS), which effectively improves the rate-distortion trade-off by employing a cheap neural super-resolution model for quality enhancement on the receiver side. Missing by existing work, we reveal that the NeVS pipeline may suffer from a practical threat, where the crucial codec component (i.e., encoder for compression and decoder for restoration) can trigger adversarial attacks in a man-in-the-middle manner to significantly destroy video recovery performance and finally incurs the malfunction of downstream video perception tasks. In this paper, we are the first attempt to inspect the vulnerability of NeVS and discover a novel adversarial attack, called codec hijacking, where the injected invisible perturbation conspires with the malicious encoding matrix by reorganizing the spatial-temporal bit allocation within the bitstream size budget. Such a zero-day vulnerability makes our attack hard to defend because there is no visual distortion on the recovered videos until the attack happens. More seriously, this attack can be extended to diverse enhancement models, thus exposing a wide range of video perception tasks under threat. Evaluation based on state-of-the-art video codec benchmark illustrates that our attack significantly degrades the recovery performance of NeVS over previous attack methods. The damaged video quality finally leads to obvious malfunction of downstream tasks with over 75% success rate. We hope to arouse public attention on codec hijacking and its defence.
Qihua Zhou, Jingcai Guo, Song Guo 0001, Ruibin Li, Jie Zhang 0076, Zhenda Xu
AAAI5
2024 DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning
abstract
Federated learning (FL) has emerged as a powerful paradigm for learning from decentralized data, and federated domain generalization further considers the test dataset (target domain) is absent from the decentralized training data (source domains). However, most existing FL methods assume that domain labels are provided during training, and their evaluation imposes explicit constraints on the number of domains, which must strictly match the number of clients. Because of the underutilization of numerous edge devices and additional cross-client domain annotations in the real world, such restrictions may be impractical and involve potential privacy leaks. In this paper, we propose an efficient and novel approach, called Disentangled Prompt Tuning (DiPrompT), a method that tackles the above restrictions by learning adaptive prompts for domain generalization in a distributed manner. Specifically, we first design two types of prompts, i.e., global prompt to capture general knowledge across all clients and domain prompts to capture domain-specific knowledge. They eliminate the restriction on the one-to-one mapping between source domains and local clients. Furthermore, a dynamic query metric is introduced to automatically search the suitable domain label for each sample, which includes two-substep text-image alignments based on prompt tuning without labor-intensive annotation. Extensive experiments on multiple datasets demonstrate that our DiPrompT achieves superior domain generalization performance over state-of-the-art FL methods when domain labels are not provided, and even outperforms many centralized learning methods using domain labels.
Sikai Bai, Jie Zhang 0076, Song Guo 0001, Jingcai Guo, Tao Han 0002, Xiaocheng Lu
CVPR2
2024 Learning Personalized Causally Invariant Representations for Heterogeneous Federated Clients
abstract
Personalized federated learning (PFL) has gained great success in tackling the scenarios where target datasets are heterogeneous across the local clients. However, the application of the existing PFL methods to real-world setting is hindered by the common assumption that the test data on each client is in-distribution (IND) with respect to its training data. Due to the bias of training dataset, the modern machine learning model prefers to rely on shortcut which can perform well on the training data but fail to generalize to the unseen test data that is out-of-distribution (OOD). This pervasive phenomenon is called shortcut learning and has attracted plentiful efforts in centralized situations. In PFL, the limited data diversity on federated clients makes mitigating shortcut and meanwhile preserving personalization knowledge rather difficult. In this paper, we analyse this challenging problem by formulating the structural causal models (SCMs) for heterogeneous federated clients. From the proposed SCMs, we derive two significant causal signatures which inspire a provable shortcut discovery and removal method under federated learning, namely FedSDR. Specifically, FedSDR is divided into two steps: 1) utilizing the available training data distributed among local clients to discover all the shortcut features in a collaborative manner. 2) developing the optimal personalized causally invariant predictor for each client by eliminating the discovered shortcut features. We provide theoretical analysis to prove that our method can draw complete shortcut features and produce the optimal personalized invariant predictor that can generalize to unseen OOD data on each client. The experimental results on diverse datasets validate the superiority of FedSDR over the state-of-the-art PFL methods on OOD generalization performance.
Xueyang Tang, Song Guo 0001, Jie Zhang 0076, Jingcai Guo
ICLR3
2024 Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language Models
abstract
Fine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from invariant features, and employ a uniform alignment process when adapting to unseen target domains. This can impair the cross-modal feature alignment when the testing data significantly deviate from the distribution of the training data, resulting in a poor out-of-distribution (OOD) generalization performance. In this paper, we reveal that the prompt tuning failure in such OOD scenarios can be attribute to the undesired alignment between the textual and the spurious feature. As a solution, we propose **CoOPood**, a fine-grained prompt tuning method that can discern the causal features and deliberately align the text modality with the invariant feature. Specifically, we design two independent contrastive phases using two lightweight projection layers during the alignment, each with different objectives: 1) pulling the text embedding closer to invariant image embedding and 2) pushing text embedding away from spurious image embedding. We have illustrated that **CoOPood** can serve as a general framework for VLMs and can be seamlessly integrated with existing prompt tuning methods. Extensive experiments on various OOD datasets demonstrate the performance superiority over state-of-the-art methods.
Jie Zhang 0076, Xiaosong Ma, Song Guo 0001, Peng Li 0017, Wenchao Xu 0001, Xueyang Tang, Zicong Hong
ICML1
2024 Causally Motivated Personalized Federated Invariant Learning with Shortcut-Averse Information-Theoretic Regularization
abstract
Exploiting invariant relations and mitigating spurious correlation (a.k.a., shortcut) between representation and target across varied data distributions can tackle the challenging out-of-distribution (OOD) generalization problem. In personalized federated learning (PFL), heterogeneous data distribution across local clients offers the inherent prerequisites to extract the invariant features that maintain invariant relation with target. Nevertheless, personalized features are closely entangled with spurious features in PFL since they exhibit similar variability across different clients, which makes preserving personalization knowledge and eliminating shortcuts two conflicting objectives in PFL. To address the above challenge, we analyse the heterogeneous data generation on local clients through the lens of structured causal model and propose a crucial causal signature which can distinguish personalized features from spurious features with global invariant features as the anchor. Then the causal signature is quantified as an information-theoretic constraint that facilitates the shortcut-averse personalized invariant learning on each client. Theoretical analysis demonstrates our method, FedPIN, can yield a tighter bound on generalization error than the prevalent PFL approaches when train-test distribution shift exists on clients. Moreover, we provide a theoretical guarantee on the convergence rate of FedPIN in this paper. The results of extensive experiments show that our method can achieve superior OOD generalization performance compared with the state-of-the-art competitors.
Xueyang Tang, Song Guo 0001, Jingcai Guo, Jie Zhang 0076, Yue Yu 0001
ICML4
2024 Easing Concept Bleeding in Diffusion via Entity Localization and Anchoring
abstract
Recent diffusion models have manifested extraordinary capabilities in generating high-quality, diverse, and innovative images guided by textual prompts. Nevertheless, these state-of-the-art models may encounter the challenge of concept bleeding when generating images with multiple entities or attributes in the prompt, leading to the unanticipated merging or overlapping of distinct objects in the synthesized result. The current work exploits auxiliary networks to produce mask-constrained regions for entities, necessitating the training of an object detection network. In this paper, we investigate the bleeding reason and find that the cross-attention map associated with a specific entity or attribute tends to extend beyond its intended focus, encompassing the background or other unrelated objects and thereby acting as the primary source of concept bleeding. Motivated by this, we propose Entity Localization and Anchoring (ELA) to drive the entity to concentrate on the expected region accurately during inference, eliminating the necessity for training. Specifically, we initially identify the region corresponding to each entity and subsequently employ a tailored loss function to anchor entities within their designated positioning areas. Extensive experiments demonstrate its superior capability in precisely generating multiple objects as specified in the textual prompts.
Jiewei Zhang, Song Guo 0001, Peiran Dong, Jie Zhang 0076, Yue Yu 0001
ICML4
2024 ParsNets: A Parsimonious Composition of Orthogonal and Low-Rank Linear Networks for Zero-Shot Learning
Jingcai Guo, Qihua Zhou, Xiaocheng Lu, Ruibin Li, Jie Zhang 0076, Junyang Chen 0001, Xin Xie 0001, Song Guo 0001
IJCAI6
2024 Dual Expert Distillation Network for Generalized Zero-Shot Learning
Zhijie Rao, Jingcai Guo, Xiaocheng Lu, Jingming Liang, Jie Zhang 0076, Haozhao Wang, Kang Wei 0004, Xiaofeng Cao 0002
IJCAI5
2024 OTAS: An Elastic Transformer Serving System via Token Adaptation
abstract
Transformer model empowered architectures have become a pillar of cloud services that keeps reshaping our society. However, the dynamic query loads and heterogeneous user requirements severely challenge current transformer serving systems, which rely on pre-training multiple variants of a foundation model, i.e., with different sizes, to accommodate varying service demands. Unfortunately, such a mechanism is unsuitable for large transformer models due to the additional training costs and excessive I/O delay. In this paper, we introduce OTAS, the first elastic serving system specially tailored for transformer models by exploring lightweight token management. We develop a novel idea called token adaptation that adds prompting tokens to improve accuracy and removes redundant tokens to accelerate inference. To cope with fluctuating query loads and diverse user requests, we enhance OTAS with application-aware selective batching and online token adaptation. OTAS first batches incoming queries with similar service-level objectives to improve the ingress throughput. Then, to strike a tradeoff between the overhead of token increment and the potentials for accuracy improvement, OTAS adaptively adjusts the token execution strategy by solving an optimization problem. We implement and evaluate a prototype of OTAS with multiple datasets, which show that OTAS improves the system utility by at least 18.2%.
Wenchao Xu 0001, Zicong Hong, Song Guo 0001, Haozhao Wang, Jie Zhang 0076, Deze Zeng
INFOCOM6
2024 SFP: Spurious Feature-Targeted Pruning for Out-of-Distribution Generalization
abstract
Recent studies reveal that even highly biased dense networks can contain an invariant substructure with superior out-of-distribution (OOD) generalization. While existing works commonly seek these substructures using global sparsity constraints, the uniform imposition of sparse penalties across samples with diverse levels of spurious contents renders such methods suboptimal. The precise adaptation of model sparsity, specifically tailored for spurious features, remains a significant challenge. Motivated by the insight that in-distribution (ID) data containing spurious features may exhibit lower experiential risk, we propose a novel Spurious Feature-targeted Pruning framework, dubbed SFP, to induce the authentic invariant substructures without referring to the above concerns. Specifically, SFP distinguishes spurious features within ID instances during training by a theoretically validated threshold. It then penalizes the corresponding feature projections onto the model space, steering the optimization towards subspaces spanned by those invariant factors. Moreover, we also conduct detailed theoretical analysis to provide a rationality guarantee and a proof framework for OOD structures based on model sparsity. Experiments on various OOD datasets show that SFP can significantly outperform both structure-based and non-structure-based OOD generalization state-of-the-art (SOTA) methods by large margins.
Yingchun Wang 0001, Jingcai Guo, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Weizhan Zhang
ACM Multimedia5
2024 Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation Editing
abstract
Recent advancements in vision-language-to-image (VL2I) diffusion generation have made significant progress. While generating images from broad vision-language inputs holds promise, it also raises concerns about potential misuse, such as copying artistic styles without permission, which could have legal and social consequences. Therefore, it's crucial to establish governance frameworks to ensure ethical and copyright integrity, especially with widely used diffusion models. To address these issues, researchers have explored various approaches, such as dataset filtering, adversarial perturbations, machine unlearning, and inference-time refusals. However, these methods often lack either scalability or effectiveness. In response, we propose a new framework called causal representation editing (CRE), which extends representation editing from large language models (LLMs) to diffusion-based models. CRE enhances the efficiency and flexibility of safe content generation by intervening at diffusion timesteps causally linked to unsafe concepts. This allows for precise removal of harmful content while preserving acceptable content quality, demonstrating superior effectiveness, precision and scalability compared to existing methods. CRE can handle complex scenarios, including incomplete or blurred representations of unsafe concepts, offering a promising solution to challenges in managing harmful content generation in diffusion-based models.
Peiran Dong, Song Guo 0001, Jie Zhang 0076, Zicong Hong
NeurIPS5
2024 Towards performance-maximizing neural network pruning via global channel attention
Yingchun Wang 0001, Song Guo 0001, Jingcai Guo, Jie Zhang 0076, Weizhan Zhang, Caixia Yan, Yuanhong Zhang
Neural Networks4
2024 SafeDRL: Dynamic Microservice Provisioning With Reliability and Latency Guarantees in Edge Environments
abstract
As a key technology of 5G, network function virtualization enables each monolithic service to be divided into microservices, facilitating their deployment and management in edge environments. One of the most critical issues in 5G is how to support dynamically arriving mission-critical services with low-latency and high-reliability requirements in distributed edge environments. However, most existing works focus on how to provide reliable services without considering latency, and their heuristics struggle to cope with high-dimensional constraints and complex environments with heterogeneous infrastructure and services. In this paper, we propose a SafeDRL algorithm to resource-efficiently support these dynamically arriving services while meeting their reliability and latency requirements. Specifically, we first formulate the problem as an integer nonlinear programming and prove its NP-hardness. To tackle this problem, our SafeDRL algorithm captures delayed rewards in dynamic environments by reinforcement learning, and corrects constraint violations with high-quality feasible solutions based on expert intervention, and prunes unnecessary backup instances for optimality. The algorithm is proved to have a bounded approximation ratio in general cases. Extensive trace-driven simulations show that, compared with the state-of-the-art solution, SafeDRL can save resource costs by up to 49.32% and improve the service acceptance ratio by up to 55% with acceptable execution time.
Yue Zeng 0002, Zhihao Qu, Song Guo 0001, Jie Zhang 0076, Jing Li 0093, Bin Tang 0002
IEEE Trans. Computers5
2024 A Sustainable Supply Chain Design for Personalized Customization in Industry 5.0 Era
abstract
Purpose: This research aims to present a distributed localized manufacturing (DLM)-based personalized customization Supply Chain (PCSC) model with facility siting, with the objective of enhancing the sustainability level of PCSCs within the context of Industry 5.0. Design/methodology/approach: To accomplish the objectives, a DLM-based PCSC model is constructed, and the supply chain is optimized using a P-Median model with genetic algorithm. Furthermore, a hybrid simulation model, combining agent-based modeling and discrete event simulation, is employed to analyze and gather data on sustainable metrics. Findings: The implementation of the DLM-based PCSC model yields improvements in supply chain efficiency through cost reduction, risk mitigation, and enhanced responsiveness. By leveraging the decentralized manufacturing approach of DLM, organizations can minimize transportation distances, optimize resource utilization, and strengthen coordination among geographically dispersed facilities. Consequently, sustainability and overall performance within the supply chain are enhanced. Practical implications: This study offers valuable recommendations for stakeholders and managers, including the adoption of distributed local manufacturing, optimization of facility locations, and the integration of sustainability metrics analysis into decision-making processes. These measures contribute to the improvement of supply chain sustainability and performance. Originality/value: This article makes a significant contribution to the field by proposing a DLM-based PCSC model that incorporates facility siting. The employment of a hybrid simulation model presents an integrated approach to assessing supply chains. In addition, it expands the measurement of sustainability metrics and provides insights to enhance the sustainability and efficiency of PCSCs.
Jie Zhang 0076, Yan Hong 0002, Song Guo 0001, Xianyi Zeng
IEEE Trans. Ind. Informatics3
2023 Towards Unbiased Training in Federated Open-world Semi-supervised Learning
abstract
Federated Semi-supervised Learning (FedSSL) has emerged as a new paradigm for allowing distributed clients to collaboratively train a machine learning model over scarce labeled data and abundant unlabeled data. However, existing works for FedSSL rely on a closed-world assumption that all local training data and global testing data are from seen classes observed in the labeled dataset. It is crucial to go one step further: adapting FL models to an open-world setting, where unseen classes exist in the unlabeled data. In this paper, we propose a novel Federatedopen-world Semi-Supervised Learning (FedoSSL) framework, which can solve the key challenge in distributed and open-world settings, i.e., the biased training process for heterogeneously distributed unseen classes. Specifically, since the advent of a certain unseen class depends on a client basis, the locally unseen classes (exist in multiple clients) are likely to receive differentiated superior aggregation effects than the globally unseen classes (exist only in one client). We adopt an uncertainty-aware suppressed loss to alleviate the biased training between locally unseen and globally unseen classes. Besides, we enable a calibration module supplementary to the global aggregation to avoid potential conflicting knowledge transfer caused by inconsistent data distribution among different clients. The proposed FedoSSL can be easily adapted to state-of-the-art FL methods, which is also validated via extensive experiments on benchmarks and real-world datasets (CIFAR-10, CIFAR-100 and CINIC-10).
Jie Zhang 0076, Xiaosong Ma, Song Guo 0001, Wenchao Xu 0001
ICML1
2023 Towards Fairer and More Efficient Federated Learning via Multidimensional Personalized Edge Models
abstract
Federated learning (FL) is an emerging technique that trains massive and geographically distributed edge data while maintaining privacy. However, FL has inherent challenges in terms of fairness and computational efficiency due to the rising heterogeneity of edges, and thus usually results in sub-optimal performance in recent state-of-the-art (SOTA) solutions. In this paper, we propose a Customized Federated Learning (CFL) system to eliminate FL heterogeneity from multiple dimensions. Specifically, CFL tailors personalized models from the specially designed global model for each client jointly guided by an online trained model-search helper and a novel aggregation algorithm. Extensive experiments demonstrate that CFL has full-stack advantages for both FL training and edge reasoning and significantly improves the SOTA performance w.r.t. model accuracy (up to 7.2% in the non-heterogeneous environment and up to 21.8% in the heterogeneous environment), efficiency, and FL fairness.
Jingcai Guo, Jie Zhang 0076, Song Guo 0001, Weizhan Zhang
IJCNN3
2023 Prophet: Conflict-Free Sharding Blockchain via Byzantine-Tolerant Deterministic Ordering
abstract
Sharding scales throughput by splitting blockchain nodes into parallel groups. However, different shards’ independent and random scheduling for cross-shard transactions results in numerous conflicts and aborts, since cross-shard transactions from different shards may access the same account. A deterministic ordering can eliminate conflicts by determining a global order for transactions before processing, as proved in the database field. Unfortunately, due to the intertwining of the Byzantine environment and information isolation among shards, there is no trusted party able to predetermine such an order for cross-shard transactions. To tackle this challenge, this paper proposes Prophet, a conflict-free sharding blockchain based on Byzantine-tolerant deterministic ordering. It first depends on untrusted self-organizing coalitions of nodes from different shards to pre-execute cross-shard transactions for prerequisite information about ordering. It then determines a trusted global order based on stateless ordering and post-verification for pre-executed results, through shard cooperation. Following the order, the shards thus orderly execute and commit transactions without conflicts. Prophet orchestrates the pre-execution, ordering, and execution processes in the sharding consensus for minimal overhead. We rigorously prove the determinism and serializability of transactions under the Byzantine and sharded environment. An evaluation of our prototype shows that Prophet improves the throughput by 3.11× and achieves nearly no aborts on 1 million Ethereum transactions compared with state-of-the-art sharding.
Zicong Hong, Song Guo 0001, Enyuan Zhou, Wuhui Chen, Jinwen Liang, Jie Zhang 0076, Albert Y. Zomaya
INFOCOM7
2023 SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models
abstract
Test-time adaptation (TTA) is a special and practical setting in unsupervised domain adaptation, which allows a pre-trained model in a source domain to adapt to unlabeled test data in another target domain. To avoid the computation-intensive backbone fine-tuning process, the zero-shot generalization potentials of the emerging pre-trained vision-language models (e.g., CLIP, CoOp) are leveraged to only tune the run-time prompt for unseen test domains. However, existing solutions have yet to fully exploit the representation capabilities of pre-trained models as they only focus on the entropy-based optimization and the performance is far below the supervised prompt adaptation methods, e.g., CoOp. In this paper, we propose SwapPrompt, a novel framework that can effectively leverage the self-supervised contrastive learning to facilitate the test-time prompt adaptation. SwapPrompt employs a dual prompts paradigm, i.e., an online prompt and a target prompt that averaged from the online prompt to retain historical information. In addition, SwapPrompt applies a swapped prediction mechanism, which takes advantage of the representation capabilities of pre-trained models to enhance the online prompt via contrastive learning. Specifically, we use the online prompt together with an augmented view of the input image to predict the class assignment generated by the target prompt together with an alternative augmented view of the same image. The proposed SwapPrompt can be easily deployed on vision-language models without additional requirement, and experimental results show that it achieves state-of-the-art test-time adaptation performance on ImageNet and nine other datasets. It is also shown that SwapPrompt can even achieve comparable performance with supervised prompt adaptation methods.
Xiaosong Ma, Jie Zhang 0076, Song Guo 0001, Wenchao Xu 0001
NeurIPS2
2023 Opponent Modeling Based Dynamic Resource Trading for UAV-Assisted Edge Computing
abstract
This paper proposes a dynamic resource trading scheme in unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) network. A UAV-assisted MEC server adaptively adjusts its trajectory to sell the computation offloading services to the mobile users (MUs), where the MUs have stochastic task arrivals. In this context, we formulate the sequential resource trading problem as a stochastic Stackelberg game, which is composed of two stages for each trading round. In the first stage, the self-interested UAV jointly optimizes its trajectory and service price to maximize its long-term profits. In the second stage, the non-cooperative MUs optimize their binary offloading decisions to minimize the average task processing delay and service payment. However, it is challenging to obtain the equilibrium across the fully decentralized agents with constantly evolving and tightly coupled policies, where each agent is confronted with a non-stationary environment. To solve this problem, we propose an opponent modeling based double deep Q learning (OM-DDQN) algorithm, where each agent adopts opponent modeling to effectively predict the trading strategies of other agents in the network. Simulation results demonstrate that, compared with the baseline algorithms, the proposed algorithm can achieve a win-win resource trading outcome that not only enhances the UAV's profit but also reduces the MUs' costs.
Jinxiang Bai, Zhe Wang 0005, Jun Li 0004, Long Shi 0001, Jie Zhang 0076, Kang Wei 0004, Hengtao He
VTC Fall5
2023 Towards Data-Independent Knowledge Transfer in Model-Heterogeneous Federated Learning
abstract
Federated Distillation (FD) extends classic Federated Learning (FL) to a more general training framework that enables model-heterogeneous collaborative learning by Knowledge Distillation (KD) across multiple clients and the server. However, existing KD-based algorithms usually require a set of shared input samples for each client to produce soft-prediction for distillation. Worse still, such a manual selection is accompanied by careful deliberations or prior information on clients’ private data distribution, which is not in line with the privacy-preserving characteristic of classic FL. In this paper, we propose a novel training framework to achieve data-independent knowledge transfer by properly designing a distributed generative adversarial network (GAN) between the server and clients that can synthesize shared feature representations to facilitate the FD training. Specifically, we deploy a generator on the server and reuse each local model as a federated discriminator to form a lightweight efficient distributed GAN that can automatically synthesize simulated global feature representations for distillation. Moreover, since the synthesized feature representations are usually more faithful and homologous with global data distribution, faster and better training convergence can be obtained. Extensive experiments on different tasks and heterogeneous models demonstrate the effectiveness of the proposed framework on model accuracy and communication overhead.
Jie Zhang 0076, Song Guo 0001, Jingcai Guo, Deze Zeng, Jingren Zhou 0001, Albert Y. Zomaya
IEEE Trans. Computers1
2023 From Deterioration to Acceleration: A Calibration Approach to Rehabilitating Step Asynchronism in Federated Optimization
abstract
In the setting of federated optimization, where a global model is aggregated periodically, step asynchronism occurs when participants conduct model training by efficiently utilizing their computational resources. It is well acknowledged that step asynchronism leads to objective inconsistency under non-i.i.d. data, which degrades the model’s accuracy. To address this issue, we propose a new algorithmFedaGrac, which calibrates the local direction to a predictive global orientation. Taking advantage of the estimated orientation, we guarantee that the aggregated model does not excessively deviate from the global optimum while fully utilizing the local updates of faster nodes. We theoretically prove thatFedaGracholds an improved order of convergence rate than the state-of-the-art approaches and eliminates the negative effect of step asynchronism. Empirical results show that our algorithm accelerates the training and enhances the final accuracy.
Feijie Wu, Song Guo 0001, Haozhao Wang, Haobo Zhang 0002, Zhihao Qu, Jie Zhang 0076
IEEE Trans. Parallel Distributed Syst.6
2023 RuleDRL: Reliability-Aware SFC Provisioning With Bounded Approximations in Dynamic Environments
abstract
As a key enabling technology for 5G, network function virtualization abstracts services into software-based service function chains (SFCs), facilitating mission-critical services with high-reliability requirements. However, it is challenging to cost-effectively provide reliable SFCs in dynamic environments due to delayed rewards caused by future SFC requests, limited infrastructure resources, and heterogeneity in hardware and software reliability. Although deep reinforcement learning (DRL) can effectively capture delayed rewards in dynamic environments, its trial-and-error exploration in a vast solution space with massive infeasible solutions may lead to frequent constraint violations and traps in poor local optima. To address these challenges, we propose a RuleDRL algorithm that combines the capability of DRL to capture delayed rewards and the strength of rule-based schemes to explore high-quality solutions without violating constraints. Specifically, we first formulate the reliable SFC provision problem as an integer nonlinear programming problem, which is proven to be NP-hard. Then, we jointly design DRL and rule-based schemes that are coupled to make the final decision and establish a bounded approximation ratio in general cases. Extensive trace-driven simulations show that RuleDRL can save the total cost by up to 65.67% and improve the SFC acceptance ratio by up to 82%, compared to the state-of-the-art solution.
Yue Zeng 0002, Zhihao Qu, Song Guo 0001, Bin Tang 0002, Jing Li 0093, Jie Zhang 0076
IEEE Trans. Serv. Comput.7
2022 Layer-wised Model Aggregation for Personalized Federated Learning
abstract
Personalized Federated Learning (pFL) not only can capture the common priors from broad range of distributed data, but also support customized models for heterogeneous clients. Researches over the past few years have applied the weighted aggregation manner to produce personalized models, where the weights are determined by calibrating the distance of the entire model parameters or loss values, and have yet to consider the layer-level impacts to the aggregation process, leading to lagged model convergence and inadequate personalization over non-IID datasets. In this paper, we propose a novel pFL training framework dubbed Layer-wised Personalized Federated learning (pFedLA) that can discern the importance of each layer from different clients, and thus is able to optimize the personalized model aggregation for clients with heterogeneous data. Specifically, we employ a dedicated hyper-network per client on the server side, which is trained to identify the mutual contribution factors at layer granularity. Meanwhile, a parameterized mechanism is introduced to update the layer-wised aggregation weights to progressively exploit the inter-user similarity and realize accurate model personalization. Extensive experiments are conducted over different models and learning tasks, and we show that the proposed methods achieve significantly higher performance than state-of-the-art pFL methods.
Xiaosong Ma, Jie Zhang 0076, Song Guo 0001, Wenchao Xu 0001
CVPR2
2022 Sign bit is enough: a learning synchronization framework for multi-hop all-reduce with ultimate compression
abstract
Traditional one-bit compressed stochastic gradient descent can not be directly employed in multi-hop all-reduce, a widely adopted distributed training paradigm in network-intensive high-performance computing systems such as public clouds. According to our theoretical findings, due to the cascading compression, the training process has considerable deterioration on the convergence performance. To overcome this limitation, we implement a sign-bit compression-based learning synchronization framework, Marsit. It prevents cascading compression via an elaborate bit-wise operation for unbiased sign aggregation and its specific global compensation mechanism for mitigating compression deviation. The proposed framework retains the same theoretical convergence rate as non-compression mechanisms. Experimental results demonstrate that Marsit reduces up to 35% training time while preserving the same accuracy as training without compression.
Feijie Wu, Shiqi He, Song Guo 0001, Zhihao Qu, Haozhao Wang, Weihua Zhuang, Jie Zhang 0076
DAC7
2022 Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning
abstract
It witnesses that the collaborative learning (CL) systems often face the performance bottleneck of limited bandwidth, where multiple low-end devices continuously generate data and transmit intermediate features to the cloud for incremental training. To this end, improving the communication efficiency by reducing traffic size is one of the most crucial issues for realistic deployment. Existing systems mostly compress features at pixel level and ignore the characteristics of feature structure, which could be further exploited for more efficient compression. In this paper, we take new insights into implementing scalable CL systems through a hierarchical compression on features, termed Stripe-wise Group Quantization (SGQ). Different from previous unstructured quantization methods, SGQ captures both channel and spatial similarity in pixels, and simultaneously encodes features in these two levels to gain a much higher compression ratio. In particular, we refactor feature structure based on inter-channel similarity and bound the gradient deviation caused by quantization, in forward and backward passes, respectively. Such a double-stage pipeline makes SGQ hold a sublinear convergence order as the vanilla SGD-based optimization. Extensive experiments show that SGQ achieves a higher traffic reduction ratio by up to 15.97 times and provides 9.22 times image processing speedup over the uniform quantized training, while preserving adequate model accuracy as FP32 does, even using 4-bit quantization. This verifies that SGQ can be applied to a wide spectrum of edge intelligence applications.
Qihua Zhou, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Jiewei Zhang, Tao Guo 0004, Zhenda Xu, Zhihao Qu
NeurIPS4
2022 Adaptive Federated Learning on Non-IID Data With Resource Constraint
abstract
Federated learning (FL) has been widely recognized as a promising approach by enabling individual end-devices to cooperatively train a global model without exposing their own data. One of the key challenges in FL is the non-independent and identically distributed (Non-IID) data across the clients, which decreases the efficiency of stochastic gradient descent (SGD) based training process. Moreover, clients with different data distributions may cause bias to the global model update, resulting in a degraded model accuracy. To tackle the Non-IID problem in FL, we aim to optimize the local training process and global aggregation simultaneously. For local training, we analyze the effect of hyperparameters (e.g., the batch size, the number of local updates) on the training performance of FL. Guided by the toy example and theoretical analysis, we are motivated to mitigate the negative impacts incurred by Non-IID data via selecting a subset of participants and adaptively adjust their batch size. A deep reinforcement learning based approach has been proposed to adaptively control the training of local models and the phase of global aggregation. Extensive experiments on different datasets show that our method can improve the model accuracy by up to 30 percent, as compared to the state-of-the-art approaches.
Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Deze Zeng, Yufeng Zhan, Rajendra Akerkar
IEEE Trans. Computers1
2022 Adaptive Vertical Federated Learning on Unbalanced Features
abstract
Most of the existing FL systems focus on a data-parallel architecture where training data are partitioned by samples among several parties. In some real-life applications, however, partitioning by features is also of practical relevance and the number of features is usually unbalanced among parties. The corresponding learning framework is referred to as Vertical Federated Learning (VFL). Though some pioneering work focused on VFL, the convergence properties of VFL on unbalanced features, especially when parties conduct different numbers of local updates concerning heterogeneous computational capabilities are still unknown. In this article, we propose a new learning framework to improve the training efficiency of VFL on unbalanced features. Given the number of features and the computational capability owned by each party, our thorough theoretical analysis exhibits that the number of local updates conducted by each party has a great effect on the convergence rate and the computational complexity, both of which jointly determine the overall training efficiency in an interrelated and sophisticated way. Based on our theoretical findings, we formulate an optimization problem and derive the optimal solution by selecting an adaptive number of local training rounds for each party. Extensive experiments on various datasets and models demonstrate that our approach significantly improves the training efficiency of VFL.
Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Deze Zeng, Haozhao Wang, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2021 Parameterized Knowledge Transfer for Personalized Federated Learning
abstract
In recent years, personalized federated learning (pFL) has attracted increasing attention for its potential in dealing with statistical heterogeneity among clients. However, the state-of-the-art pFL methods rely on model parameters aggregation at the server side, which require all models to have the same structure and size, and thus limits the application for more heterogeneous scenarios. To deal with such model constraints, we exploit the potentials of heterogeneous model settings and propose a novel training framework to employ personalized models for different clients. Specifically, we formulate the aggregation procedure in original pFL into a personalized group knowledge transfer training algorithm, namely, KT-pFL, which enables each client to maintain a personalized soft prediction at the server side to guide the others' local training. KT-pFL updates the personalized soft prediction of each client by a linear combination of all local soft predictions using a knowledge coefficient matrix, which can adaptively reinforce the collaboration among clients who own similar data distribution. Furthermore, to quantify the contributions of each client to others' personalized training, the knowledge coefficient matrix is parameterized so that it can be trained simultaneously with the models. The knowledge coefficient matrix and the model parameters are alternatively updated in each round following the gradient descent way. Extensive experiments on various datasets (EMNIST, Fashion_MNIST, CIFAR-10) are conducted under different settings (heterogeneous models and data distributions). It is demonstrated that the proposed framework is the first federated learning paradigm that realizes personalized model training via parameterized group knowledge transfer while achieving significant performance gain comparing with state-of-the-art algorithms.
Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Haozhao Wang, Wenchao Xu 0001, Feijie Wu
NeurIPS1
2020 Dual-view Attention Networks for Single Image Super-Resolution
abstract
One non-negligible flaw of the convolutional neural networks (CNNs) based single image super-resolution (SISR) models is that most of them are not able to restore high-resolution (HR) images containing sufficient high-frequency information. Worse still, as the depth of CNNs increases, the training easily suffers from the vanishing gradients. These problems hinder the effectiveness of CNNs in SISR. In this paper, we propose the Dual-view Attention Networks to alleviate these problems for SISR. Specifically, we propose the local aware (LA) and global aware (GA) attentions to deal with LR features in unequal manners, which can highlight the high-frequency components and discriminate each feature from LR images in the local and global views, respectively. Furthermore, the local attentive residual-dense (LARD) block that combines the LA attention with multiple residual and dense connections is proposed to fit a deeper yet easy to train architecture. The experimental results verified the effectiveness of our model compared with other state-of-the-art methods.
Jingcai Guo, Shiheng Ma, Jie Zhang 0076, Qihua Zhou, Song Guo 0001
ACM Multimedia3
2019 Multi-Path Routing Oriented Flow Statistics Collection in Software Defined Networks
abstract
In Software Defined Networks (SDNs), one of the key tasks in control plane is the monitoring and measurement of the whole network. A typical SDN consists of a set of switches and a logically centralized controller responsible for network state monitoring and flow scheduling, to which low cost and efficient flow statistics collection plays an important role. However, existing flow statistics collection methods mainly focus on single-path routing and cannot accurately capture the flow statistics in the case of multi-path routing (MPR), which is widely used in modern networks. In this paper, we are motivated to propose a Multi-Path oriented Flow Statistics Collection (MFSC) strategy to minimize the total communication cost for flow statistics collection. The problem is first formulated into an integer linear programming (ILP) form. After analyzing the complexity of this problem, we present a relaxation-based algorithm with an approximation factor p, where p is the maximum number of switches passed by each flow. The experiment results demonstrate that our proposed algorithm can reduce the total communication cost by over 14% compared with the traditional solutions.
Jie Zhang 0076, Song Guo 0001, Deze Zeng, Zhihao Qu
ICPADS1
2018 Stochastic Scheduling Towards Cost Efficient Network Function Virtualization in Edge Cloud
abstract
Network Function Virtualization (NFV) emerges as a promising technology to increase the network flexibility, customizability and efficiency by softwarizing traditional dedicated hardware based functions to virtualized network functions. The prosperous potential of edge cloud makes it an ideal platform to host the network functions. From the perspective of network service providers, an inevitable concern is how to reduce the overall cost for renting various resources from infrastructure providers. In this paper, unlike existing related studies assuming a preknown network traffic demand, we alternatively consider a practical case without any prior knowledge. We investigate how to dynamically minimize the overall operational cost with joint consideration of packet scheduling, network function management and resource allocation. The tradeoff between the queue backlog and overall cost is analyzed using a Lyapunov optimization framework. A backpressure based online scheduling algorithm is proposed and its efficiency is extensively evaluated by trace-driven simulations.
Deze Zeng, Jie Zhang 0076, Lin Gu 0002, Song Guo 0001
SECON2
2017 Minimize Coflow Completion Time via Joint Optimization of Flow Scheduling and Processor Placement
abstract
The recent progress in big data has inspired lots of data- parallel applications deployed in the datacenters. Although how to optimize the data flow scheduling in datacenters has been extensively studied, traditional per-flow based optimizations usually do not perform well in dealing with the transferring of a collection of parallel flows, i.e., coflow. Consequently, how to schedule the coflow towards various objectives, e.g., minimizing the coflow completion time, has attracted much attention recently. We notice that existing coflow scheduling studies usually suggest a fixed destination for each coflow. Taking the advantage of virtualization technology, we argue that the destination can be flexibly placed in the cloud. Therefore, it is essential to jointly optimize the coflow scheduling and data processor placement. In this paper, we are motivated to investigate the problem of coflow completion time minimization with joint consideration of coflow scheduling and data processor placement. We first formally describe the problem into a mixed integer non-linear programming (MINLP) problem. By linearizing the MINLP, we further propose a relaxation based heuristic algorithm. Via extensive simulation studies, the high efficiency of our heuristic algorithm is validated.
Deze Zeng, Jie Zhang 0076, Lin Gu 0002, Peng Li 0017, Hong Yao
GLOBECOM2
2017 Joint Optimization of Virtual Function Migration and Rule Update in Software Defined NFV Networks
abstract
Emerging technologies such as Software-Defined Networks (SDN) and Network Function Virtualization (NFV) promise to address cost reduction and flexibility in network operation while enabling innovative network service delivery. To catch up with the time- varying traffic demands, the network changes frequently. We should come up with a sequence of instructions to manipulate the starting network into the goal network, while preserving the network semantics correctness (e.g., freedom of loops, bandwidth guaranteeing). In this case, how to migrate the virtual network functions (VNF) and update the flow forwarding rules efficiently is an important and challenging problem. In this paper, we are motivated to address the migration of VNF and flow update rule problem with joint consideration of migration cost and update delay. The problem is first formulated into a mixed integer non-linear programming (MINLP). By linearizing and relaxing the MINLP, we then present a polynomial-time two-stage heuristic algorithm. The high efficiency of our algorithm is extensively validated by simulation based studies by the fact that it performs much closer to the optimal solution.
Jie Zhang 0076, Deze Zeng, Lin Gu 0002, Hong Yao, Muzhou Xiong
GLOBECOM1
2017 Which DRM grade could BYOD users employ? A differentiated DRM service between the cloud and mobile devices
abstract
The idea of employees leveraging their personal mobile devices for their work (Bring Your Own Device, or BYOD) is becoming increasingly popular in recent years. As BYOD users will use various digital goods (such as cloud services and mobile software) for their work and personal purposes via the same mobile devices, it brings serious security risks into both the cloud and mobile devices. Generally, the BYOD users would employ digital rights management (DRM) to control and manage the execution of digital goods. However, the security requirements for using the digital goods for work and personal tasks are very different, and conventional unified cloud-based DRM services lack the flexibility to satisfy the BYOD users' demands on diversified security levels. In this paper, we regard the security of digital goods as a metric to differentiate the DRM service into multiple grades. We propose a differentiated DRM service to increase the security flexibility of digital goods, which allows BYOD users to choose their preferred DRM grades to maximize their utility. Moreover, the differentiated DRM service can increase the benefit of service providers (SPs) even when the SPs competes with others, and thus, it becomes a dominant strategy for the SPs.
Jie Zhang 0076, Wei Lou
IWQoS1
2017 It Can Drain Out Your Energy: An Energy-Saving Mechanism Against Packet Overhearing in High Traffic Wireless LANs
abstract
Energy efficiency is a critical issue of wireless devices. As the packets are broadcast to the devices in the wireless transmission media, all active neighboring devices have to spend their energy receiving the packets though the packets are not addressed to them, which is called as the packet overhearing problem. The real-world traffic trace analysis reveals that the energy cost on the packet overhearing accounts for the majority of the devices' energy inefficiency in high traffic wireless local area networks (WLANs). In this paper, we propose a novel sample-address sample-duration (SASD) scheme to solve the energy inefficiency of the packet overhearing problem. By adding a new SASD header, which contains the critical information, in front of the data packet at the physical layer, the SASD enables the devices to discern the required information in the energy-saving downclocking mode. Consequently, the non-destination devices of the packet can switch to the sleeping mode to avoid the packet overhearing problem. We demonstrate the feasibility of the SASD through hardware experiments and evaluate its energy-saving performance through ns-2 simulations. The results show that the SASD can greatly outperform the existing approaches in the high traffic WLAN scenario.
Junmei Yao, Jie Zhang 0076, Wei Lou
IEEE Trans. Mob. Comput.3
2016 On Eliminating the Exposed Terminal Problem Using Signature Detection
abstract
Wireless networks are propelled to improve the network throughput effectively to face the challenge of sustaining the rapid growth of data traffic and the high density of wireless nodes. Exposed terminals are a main source in wireless networks that degrades the network throughput performance through excessively avoiding interferences and forbidding concurrencies. To combat the exposed terminal problem and exploit the concurrent transmissions in wireless networks, we present the design of Interference Resistant Multiple Access (IRMA) in this paper, which can achieve higher throughput compared to the 802.11 standard. By observing that nodes in current protocols waste transmission opportunities in two different scenarios, IRMA exploits the concurrency in two aspects. IRMA proposes a signature detection method in the physical layer to combat control frames' collisions, thus exploits the concurrency at the transmitter side. IRMA also designs a new NAV update scheme in the MAC layer to differentiate the interference ranges of different transmission links, thus exploits the concurrency of all non-interfering links. Experimental results based on USRP2 demonstrate the feasibility of the signature detection method, and simulations based on ns-2 show that IRMA outperforms the 802.11 standard and other protocols significantly.
Junmei Yao, Jie Zhang 0076, Wei Lou
IEEE Trans. Mob. Comput.3
2015 On Rule Placement for Multi-path Routing in Software-Defined Networks
Jie Zhang 0076, Deze Zeng, Lin Gu 0002, Hong Yao
CollaborateCom1
2015 MIO: Enhancing Wireless Communications Security Through Physical Layer Multiple Inter-Symbol Obfuscation
abstract
Communications security is a critical and increasingly challenging issue in wireless networks. A well-known approach for achieving information-theoretic secrecy relies on deploying artificial noises to blind the intruders' interception in the physical layer. However, this approach requires a static channel condition for the transmitter and receiver to generate and offset the controllable artificial noise, which can hardly be implemented in real wireless environments. In this paper, we explore the feasibility of symbol obfuscation to defend against the passive eavesdropping attack and fake packet injection attack during the wireless communications. We propose a multiple inter-symbol obfuscation (MIO) scheme, which utilizes a set of artificial noisy symbols (symbols key) to obfuscate the original data symbols in the physical layer. MIO can effectively enhance the wireless communications security. On the one hand, an eavesdropper, without knowing the artificial noisy symbols, cannot correctly decrypt the obfuscated symbols from the eavesdropped packets. On the other hand, a legitimate receiver can easily check the integrity of the symbols key and then reject the fake packets from the received packets. The security analysis reveals that, without considering the initial key, the MIO scheme can achieve information-theoretic secrecy against the passive eavesdropping attack and computational secrecy against the fake packet injection attack. Moreover, we have implemented our approach in a USRP2 testbed and conducted simulations with Simulink tools to validate the effectiveness of MIO in enhancing wireless communications security.
Wei Lou, Jie Zhang 0076, Hailun Tan
IEEE Trans. Inf. Forensics Secur.3
2014 Community Clinic: Economizing Mobile Cloud Service Cost via Cloudlet Group
abstract
The explosive growth of mobile applications causes the mobile traffic to easily exceed the capacity of the cloud service due to the bandwidth limits of last mile connections to the cloud and legacy backhauls to macrocells' base stations. It degrades mobile applications' quality of service since the mobile devices have to spend more time and thus consume more battery power for data transmissions. It also enforces the cloud provider to put a huge investment to update its infrastructure and the induced cost is inevitably borne by all mobile users. To resolve this issue, in this paper we propose a so-called community clinic solution, which embeds the cloudlet group between the cloud and mobile users, to cut down the cost introduced by the massive deployment of the cloud's data centers and save the battery power consumed by the mobile devices. We firstly show that the mobile devices can consume less energy by choosing the service provided by the cloudlet group. We then model the system with and without the cloudlet group as two types of supply chain and prove that the cloudlet group can increase the cloud's profit without putting additional cost on mobile users. We also propose the real-time group-buying auction for the cloudlet group to promote its service to its nearby mobile users with a lower price and maximize its profit. The community clinic can result in a win-win-win outcome among the cloud, cloudlet group and mobile users. Numerical experiments are further conducted to demonstrate the effectiveness of our scheme.
Jie Zhang 0076, Wei Lou
MASS1
2014 On eliminating energy inefficiency of the packet overhearing problem in high traffic wireless LANs
abstract
Energy efficiency is known to be a critical issue of wireless devices. Our analysis reveals that the energy cost on overhearing unnecessary packet transmissions accounts for the majority of unnecessary energy consumption of wireless devices in high traffic wireless LANs. In this paper, we propose a novel SASD mechanism to eliminate the energy inefficiency of the packet overhearing problem. By adding a new header that carries some information in front of the normal packet frame at the PHY layer, SASD enables the wireless devices to discern the required information under the energy-saving downclocking mode so that the devices, whenever not addressed, can switch to the sleeping mode to save energy. Our simulation results show that SASD can effectively save the energy cost by avoiding the packet overhearing in high traffic wireless LANs.
Jie Zhang 0076, Wei Lou
SECON2
2012 Symbol-level detection: A new approach to silencing hidden terminals
abstract
Hidden terminals are typical interference sources that can significantly reduce the throughput of a wireless network if it adopts the CSMA/CA MAC protocol. The RTS/CTS mechanism is a well-known solution to this hidden terminal problem. However, it only works well under the assumption that all hidden terminals can decode the CTS packets correctly. In the real world, the CTS packets might not be correctly received all the time due to either the CTS packets are unable to be decoded at remote hidden terminals or the CTS packets are collided with other packets at the hidden terminals. Both of these drawbacks can make the standard RTS/CTS mechanism fail to silence all hidden terminals, and deteriorate the throughput of the wireless network. In this paper, we present the RTS/S-CTS mechanism, a novel symbol-level detection mechanism that combats these two drawbacks. The RTS/S-CTS frames make slight changes to the standard RTS/CTS frames, and can be compatible with the standard 802.11 MAC layer. We design the symbol-level detection decoder (SLDD) and NAV decision algorithm that enable the S-CTS frame to be correctly detected from collisions and by remote hidden terminals. We build a testbed of RTS/S-CTS with GNURadio/USRP2 software radio to demonstrate its feasibility and run ns-2 simulations to evaluate its performance. The results show that the RTS/S-CTS can achieve up to 63% throughput improvement in the random topology network scenario compared with the standard RTS/CTS.
Jie Zhang 0076, Junmei Yao, Wei Lou
ICNP2
2011 The digital rights management game in peer-to-peer streaming systems
abstract
In this paper we model the digital rights management (DRM) for peer-to-peer streaming (P2PS) systems as a game. We construct the DRM game from both content service provider (CSP) and user's aspects, and propose a design of DRM policy based on homogeneous peers and homogeneous digital goods, which gets the maximal utility for the CSP as well as the criterion whether the DRM is fit for a P2PS system. Another sort of games in this paper consider how a peer deals with digital goods with regard to various situations in P2PS systems with DRM, together with the CSP's response to the peer's actions. We construct different games to avoid three notorious misbehaviors of peers: freeriding, jailbreaking and whitewashing. We take examples to show how these games work in P2PS systems with DRM and how equilibria are established in these games. Numerical experiments are conducted to demonstrate the effectiveness of the strategies devised from these games.
Jie Zhang 0076, Wei Lou
INFOCOM1