Guogang Zhu

dblp:305/8757 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0002-6381-1420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DiffDVC: Accurate Event Detection for Dense Video Captioning via Diffusion Models
abstract
Dense video captioning (DVC) aims to describe multiple events within a video, and its performance is greatly affected by the accuracy of video event detection. Video event detection involves predicting the proposal boundaries (start and end times) and the classification score of each event in a video. Recently, a few methods have applied diffusion models originally designed for image object detection to detect events in DVC. These methods add noise to the ground-truth event proposal boundaries, and subsequently learn the denoising process. However, these methods often overlook the fundamental differences between videos and images. We observe that, whereas in images the important information for object classification is normally around the boundaries of the ground-truth boxes, in videos the key information for event classification is typically centered in the middle of ground-truth event proposals. As a result, the classification module in these existing diffusion models becomes insensitive to boundary changes introduced by the added noise, leading to sub-optimal performance. This paper introduces DiffDVC, an innovative diffusion model for DVC. The core of DiffDVC is a boundary-sensitive detector. The detector increases the sensitivity of the classification module to boundary changes by focusing on frames within a specific range around the start and end times of noisy event proposals. Additionally, this range is dynamically adjusted to suit different event proposals. Comprehensive experiments on ActivityNet-1.3, ActivityNet Captions, and YouCook2 datasets show DiffDVC achieving superior performance.
Wei Chen 0109, Jianwei Niu 0002, Xuefeng Liu 0001, Shaojie Tang 0001, Guogang Zhu
AAAI6
2025 Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation
abstract
Federated Learning (FL) faces challenges due to data heterogeneity, which limits the global model’s performance across diverse client distributions. Personalized Federated Learning (PFL) addresses this by enabling each client to process an individual model adapted to its local distribution. Many existing methods assume that certain global model parameters are difficult to train effectively in a collaborative manner under heterogeneous data. Consequently, they localize or fine-tune these parameters to obtain personalized models. In this paper, we reveal that both the feature extractor and classifier of the global model are inherently strong, and the primary cause of its suboptimal performance is the mismatch between local features and the global classifier. Although existing methods alleviate this mismatch to some extent and improve performance, we find that they either (1) fail to fully resolve the mismatch while degrading the feature extractor, or (2) address the mismatch only post-training, allowing it to persist during training. This increases inter-client gradient divergence, hinders model aggregation, and ultimately leaves the feature extractor suboptimal for client data. To address this issue, we propose FedPFT, a novel framework that resolves the mismatch during training using personalized prompts. These prompts, along with local features, are processed by a shared self-attention-based transformation module, ensuring alignment with the global classifier. Additionally, this prompt-driven approach offers strong flexibility, enabling task-specific prompts to incorporate additional training objectives (\eg, contrastive learning) to further enhance the feature extractor. Extensive experiments show that FedPFT outperforms state-of-the-art methods by up to 5.07%, with further gains of up to 7.08% when collaborative contrastive learning is incorporated.
Xinghao Wu, Xuefeng Liu 0001, Jianwei Niu 0002, Guogang Zhu, Mingjia Shi, Shaojie Tang 0001, Jing Yuan 0002
NeurIPS4
2025 The Diversity Bonus: Learning From Dissimilar Clients in Personalized Federated Learning
abstract
Personalized federated learning (PFL) allows clients to collaboratively train their personalized models to handle situations where data from different clients are not independent and identically distributed (non-IID). Previous PFL research implicitly assumes that clients benefit most from those with similar data distributions. Correspondingly, methods such as personalized weight aggregation assign higher weights to similar clients during aggregation. We pose a question: can a client benefit from other clients with dissimilar data distributions, and if so, how? This question is particularly relevant in scenarios with a high degree of non-IID, where clients have widely different distributions, and learning from only similar clients will result in a loss of knowledge from many other clients. We note that when dealing with clients with similar distributions, current methods tend to enforce their models to be close in the parameter space. It is reasonable to conjecture that a client can benefit from dissimilar clients if we allow their models to depart from each other. Based on this idea, we propose DiversiFed, which allows each client to learn from clients with diversified distribution. DiversiFed pushes personalized models of clients with dissimilar distributions apart in the parameter space while pulling together those with similar distributions. In addition, to achieve the above effect without using prior knowledge of distribution, we design a loss function that leverages model similarity to determine the degree of attraction and repulsion between any two models. Experiments on benchmark and medical datasets show that DiversiFed can outperform the state-of-the-art (SOTA) methods by up to 3.19%.
Xinghao Wu, Jianwei Niu 0002, Xuefeng Liu 0001, Guogang Zhu, Shaojie Tang 0001, Wanyu Lin, Jiannong Cao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Take Your Pick: Enabling Effective Distributed Learning Within Low-Dimensional Feature Space
abstract
Personalized federated learning (PFL) is a popular distributed learning framework that allows clients to have different models and has many applications where clients' data are in different domains, including autonomous driving, traffic surveillance, and medical diagnosis. The typical model of a client in PFL features a global encoder trained by all clients to extract universal features from the raw data and personalized layers (e.g., a classifier) trained using the client's local data. Nonetheless, due to the differences between the data distributions of different clients (also known as, domain gaps), the universal features produced by the global encoder largely encompass numerous components irrelevant to a certain client's local task. Some recent PFL methods address the above problem by personalizing specific parameters within the encoder. However, these methods encounter substantial challenges attributed to the high dimensionality and nonlinearity of neural network parameter space. In contrast, the feature space exhibits a lower dimensionality, providing greater intuitiveness and interpretability as compared to the parameter space. To this end, we propose a novel PFL framework named FedPick. FedPick achieves PFL within the low-dimensional feature space by adaptively selecting task-relevant features for each client from the features generated by the global encoder based on its local data distribution. It presents a more accessible and interpretable implementation of PFL compared to those methods working in the parameter space. Extensive experimental results on multiple cross-domain datasets show that FedPick can effectively select task-relevant features for each client and improve model performance in cross-domain FL.
Guogang Zhu, Xuefeng Liu 0001, Shaojie Tang 0001, Jianwei Niu 0002, Xinghao Wu, Jiaxing Shen, Wanyu Lin
IEEE Trans. Neural Networks Learn. Syst.1
2024 BeyondVision: An EMG-driven Micro Hand Gesture Recognition Based on Dynamic Segmentation
Nana Wang 0002, Jianwei Niu 0002, Xuefeng Liu 0001, Dongqin Yu, Guogang Zhu, Xinghao Wu, Mingliang Xu 0001, Hao Su 0001
IJCAI5
2024 Estimating before Debiasing: A Bayesian Approach to Detaching Prior Bias in Federated Semi-Supervised Learning
Guogang Zhu, Xuefeng Liu 0001, Xinghao Wu, Shaojie Tang 0001, Jianwei Niu 0002, Hao Su 0001
IJCAI1
2024 Enabling Collaborative Test-Time Adaptation in Dynamic Environment via Federated Learning
abstract
Deep learning models often suffer performance degradation when test data diverges from training data. Test-Time Adaptation (TTA) aims to adapt a trained model to the test data distribution using unlabeled test data streams. In many real-world applications, it is quite common for the trained model to be deployed across multiple devices simultaneously. Although each device can execute TTA independently, it fails to leverage information from the test data of other devices. To address this problem, we introduce Federated Learning (FL) to TTA to facilitate on-the-fly collaboration among devices during test time. The workflow involves clients (i.e., the devices) executing TTA locally, uploading their updated models to a central server for aggregation, and downloading the aggregated model for inference. However, implementing FL in TTA presents many challenges, especially in establishing inter-client collaboration in dynamic environment, where the test data distribution on different clients changes over time in different manners. To tackle these challenges, we propose a server-side Temporal-Spatial Aggregation (TSA) method. TSA utilizes a temporal-spatial attention module to capture intra-client temporal correlations and inter-client spatial correlations. To further improve robustness against temporal-spatial heterogeneity, we propose a heterogeneity-aware augmentation method and optimize the module using a self-supervised approach. More importantly, TSA can be implemented as a plug-in to TTA methods in distributed environments. Experiments on multiple datasets demonstrate that TSA outperforms existing methods and exhibits robustness across various levels of heterogeneity. The code is available at https://github.com/ZhangJiayuan-BUAA/FedTSA.
Jiayuan Zhang 0001, Xuefeng Liu 0001, Guogang Zhu, Jianwei Niu 0002, Shaojie Tang 0001
KDD4
2024 Decoupling General and Personalized Knowledge in Federated Learning via Additive and Low-rank Decomposition
abstract
To address data heterogeneity, the key strategy of Personalized Federated Learning (PFL) is to decouple general knowledge (shared among clients) and client-specific knowledge, as the latter can have a negative impact on collaboration if not removed. Existing PFL methods primarily adopt a parameter partitioning approach, where the parameters of a model are designated as one of two types: parameters shared with other clients to extract general knowledge and parameters retained locally to learn client-specific knowledge. However, as these two types of parameters are put together like a jigsaw puzzle into a single model during the training process, each parameter may simultaneously absorb both general and client-specific knowledge, thus struggling to separate the two types of knowledge effectively. In this paper, we introduce FedDecomp, a simple but effective PFL paradigm that employs parameter additive decomposition to address this issue. Instead of assigning each parameter of a model as either a shared or personalized one, FedDecomp decomposes each parameter into the sum of two parameters: a shared one and a personalized one, thus achieving a more thorough decoupling of shared and personalized knowledge compared to the parameter partitioning method. In addition, as we find that retaining local knowledge of specific clients requires much lower model capacity compared with general knowledge across all clients, we let the matrix containing personalized parameters be low rank during the training process. Moreover, a new alternating training strategy is proposed to further improve the performance. Experimental results across multiple datasets and varying degrees of data heterogeneity demonstrate that FedDecomp outperforms state-of-the-art methods up to 4.9%. The code is available at https://github.com/XinghaoWu/FedDecomp
Xinghao Wu, Xuefeng Liu 0001, Jianwei Niu 0002, Haolin Wang 0002, Shaojie Tang 0001, Guogang Zhu, Hao Su 0001
ACM Multimedia6
2024 DualFed: Enjoying both Generalization and Personalization in Federated Learning via Hierachical Representations
abstract
In personalized federated learning (PFL), it is widely recognized that achieving both high model generalization and effective personalization poses a significant challenge due to their conflicting nature. As a result, existing PFL methods can only manage a trade-off between these two objectives. This raises an interesting question: Is it feasible to develop a model capable of achieving both objectives simultaneously? Our paper presents an affirmative answer, and the key lies in the observation that deep models inherently exhibit hierarchical architectures, which produce representations with various levels of generalization and personalization at different stages. A straightforward approach stemming from this observation is to select multiple representations from these layers and combine them to concurrently achieve generalization and personalization. However, the number of candidate representations is commonly huge, which makes this method infeasible due to high computational costs. To address this problem, we propose DualFed, a new method that can directly yield dual representations correspond to generalization and personalization respectively, thereby simplifying the optimization task. Specifically, DualFed inserts a personalized projection network between the encoder and classifier. The pre-projection representations are able to capture generalized information shareable across clients, and the post-projection representations are effective to capture task-specific information on local clients. This design minimizes the mutual interference between generalization and personalization, thereby achieving a win-win situation. Extensive experiments show that DualFed can outperform other FL methods. Code is available at https://github.com/GuogangZhu/DualFed.
Guogang Zhu, Xuefeng Liu 0001, Jianwei Niu 0002, Shaojie Tang 0001, Xinghao Wu, Jiayuan Zhang 0001
ACM Multimedia1
2024 Learning by imitating the classics: Mitigating class imbalance in federated learning via simulated centralized learning
Guogang Zhu, Xuefeng Liu 0001, Jianwei Niu 0002, Yucheng Wei, Shaojie Tang 0001, Jiayuan Zhang 0001
Expert Syst. Appl.1
2024 Aligning Before Aggregating: Enabling Communication Efficient Cross-Domain Federated Learning via Consistent Feature Extraction
abstract
Cross-domain federated learning (FL), where data on local clients come from different domains, is a common case of FL. In such a cross-domain case, features extracted from the raw data of different clients deviate from each other in the feature space, leading to a so-called feature shift. This phenomenon can reduce feature discrimination and degrade the performance of the learned model. However, most existing FL methods are not specifically designed for the cross-domain setting. In this article, we propose a novel cross-domain FL method named AlignFed. In AlignFed, each client model consists of a personalized feature extractor and a shared lightweight classifier. The feature extractor maps the features to a consistent space by aligning them to identical global target points. Inspired by recent studies in contrastive learning, AlignFed regards points that are uniformly distributed on the hypersphere as global target points. It then pushes features toward global target points of their corresponding classes and away from those of other classes to improve feature discrimination. The shared classifier aggregates knowledge across clients over the consistent feature space, which can mitigate performance degradation caused by feature shift while reducing communication cost. We conduct convergence analysis and perform extensive experiments to evaluate AlignFed.
Guogang Zhu, Xuefeng Liu 0001, Shaojie Tang 0001, Jianwei Niu 0002
IEEE Trans. Mob. Comput.1
2023 Bold but Cautious: Unlocking the Potential of Personalized Federated Learning through Cautiously Aggressive Collaboration
abstract
Personalized federated learning (PFL) reduces the impact of non-independent and identically distributed (non-IID) data among clients by allowing each client to train a personalized model when collaborating with others. A key question in PFL is to decide which parameters of a client should be localized or shared with others. In current mainstream approaches, all layers that are sensitive to non-IID data (such as classifier layers) are generally personalized. The reasoning behind this approach is understandable, as localizing parameters that are easily influenced by non-IID data can prevent the potential negative effect of collaboration. However, we believe that this approach is too conservative for collaboration. For example, for a certain client, even if its parameters are easily influenced by non-IID data, it can still benefit by sharing these parameters with clients having similar data distribution. This observation emphasizes the importance of considering not only the sensitivity to non-IID data but also the similarity of data distribution when determining which parameters should be localized in PFL. This paper introduces a novel guideline for client collaboration in PFL. Unlike existing approaches that prohibit all collaboration of sensitive parameters, our guideline allows clients to share more parameters with others, leading to improved model performance. Additionally, we propose a new PFL method named FedCAC, which employs a quantitative metric to evaluate each parameter’s sensitivity to non-IID data and carefully selects collaborators based on this evaluation. Experimental results demonstrate that FedCAC enables clients to share more parameters with others, resulting in superior performance compared to state-of-the-art methods, particularly in scenarios where clients have diverse distributions. The code is integrated into our FL training framework: https://github.com/kxzxvbk/Fling.
Xinghao Wu, Xuefeng Liu 0001, Jianwei Niu 0002, Guogang Zhu, Shaojie Tang 0001
ICCV4
2022 ChannelFed: Enabling Personalized Federated Learning via Localized Channel Attention
abstract
One vital challenge in federated learning (FL) is the statistical heterogeneity of data in different clients, which negatively affects the performance of the finally obtained model. One common approach to address this problem, called as personalized federated learning (PFL), is to train a personalized model for each client. A key design issue in PFL-based methods is determining which parts of the model should be personalized for each client. For example, one popular method in PFL is to personalize the batch normalization layers. In this paper, we propose ChannelFed, a new PFL-based method which personalizes the channel attention module. ChannelFed is designed based on the following observation: Channel attention assigns different weights to channels for different classes of data, which can be utilized to exploit knowledge of heterogeneous data from different clients. By keeping the channel attention module localized, ChannelFed enables clients to concentrate on client-specific channels. ChannelFed implements normalization across samples in the channel attention module to better fit for statistical heterogeneity scenarios. Experiments on CIFAR-10, Fashion-MNIST, and CIFAR-100 datasets demonstrate that ChannelFed outperforms other PFL methods under statistical heterogeneity scenarios.
Kaiyu Zheng, Xuefeng Liu 0001, Guogang Zhu, Xinghao Wu, Jianwei Niu 0002
GLOBECOM3
2022 Aligning before Aggregating: Enabling Cross-domain Federated Learning via Consistent Feature Extraction
abstract
Federated learning (FL) is an emerging machine learning paradigm where multiple distributed clients collaboratively train a model without centrally collecting their raw data. In FL setting, it is a common case that the data on local clients come from different domains, e.g., photos taken by different mobile phones can vary in intensity and contrast due to the difference of imaging parameters. In such a cross-domain case, features extracted from data of different clients deviate from each other in the feature space, leading to the so-called feature shift. The feature shift can reduce the discrimination of features and degrade the performance of the learned model. However, most existing FL methods are not particularly designed for cross-domain setting. In this paper, we propose a novel cross-domain FL method, named AlignFed. In AlignFed, the model on each client is separated to a personalized feature extractor and a shared classifier. The former extracts consistent features among clients by aligning features of different clients to some specific points in the feature space. The latter aggregates the knowledge across clients over the consistent feature space, which can mitigate the performance degradation caused by the feature shift in cross-domain FL. We conduct experiments on common-used multi-domain datasets, including Digits-Five, Office-Caltech10, and DomainNet. The experimental results demonstrate that AlignFed can outperform the state-of-art FL methods.
Guogang Zhu, Xuefeng Liu 0001, Shaojie Tang 0001, Jianwei Niu 0002
ICDCS1