EDBT 2026 Demo / reviewers in the wild / expert
He Li 0054
dblp:05/4746-54
· DBLP profile ↗
27ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-8469-8260ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Real-World Holistic Privacy-Preserving Person Re-IdentificationabstractReal-world person re-identification (Re-ID) systems are susceptible to malicious attacks, leading to the leakage of pedestrian images and the Re-ID model, posing severe threats to the privacy of both system owners and pedestrians. Existing privacy-preserving person re-identification (PPPR) methods fail to simultaneously resist data leakage, model leakage, and data & model leakage while compromising the normal functionality of Re-ID systems. In this paper, we begin with an in-depth analysis of prior methodologies and identify the gap between existing works and the ideal PPPR paradigm. Inspired by the concept of "Let the invisible perturbation become the system trigger", we propose SHIELD, a pioneering and comprehensive two-stage privacy-preserving framework. To resist data leakage, we propose a self-supervised method for Protected Dataset Generation in the first stage, which obviates the dependence on identity labels and ensures image quality. To resist model leakage without compromising the normal retrieval accuracy, we propose Original Feature Deconstruction and Protected Feature Alignment to train the system model with paired protected and original images. Extensive experiments substantiate that SHIELD significantly outperforms existing PPPR methods, offering robust and holistic protection for Re-ID systems while maintaining decent retrieval accuracy for authorized users. The code will be released soon. Qianxiang Meng, He Li 0054, Min Cao 0005, Mang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | NightReID: A Large-Scale Nighttime Person Re-Identification BenchmarkabstractPerson re-identification (Re-ID) is crucial for intelligent surveillance systems, facilitating the identification of individuals across multiple camera views. While significant advancements have been made for daytime scenarios, ensuring reliable Re-ID performance during nighttime remains a significant challenge. Given the cost and limited accessibility of infrared cameras, we investigate a critical question: Can RGB cameras be effectively utilized for accurate Re-ID during nighttime? To address this, we introduce NightReID, a large-scale RGB Re-ID dataset collected from a real-world nighttime surveillance system. NightReID includes 1,500 identities and over 53,000 images, capturing diverse scenes with complex lighting and adverse weather conditions. This rich dataset provides a valuable benchmark for advancing nighttime Re-ID research. Moreover, we propose the Enhancement, Denoising, and Alignment (EDA) framework with two novel modules to enhance nighttime Re-ID performance. First, an unsupervised Image Enhancement and Denoising (IED) method is designed to improve the quality of nighttime images, preserving critical details while removing noise without requiring paired ground truth. Second, we introduce Data Distribution Alignment (DDA) through statistical priors, aligning the distributions between pre-training data and nighttime data to mitigate domain shift. Extensive experiments on multiple nighttime Re-ID datasets demonstrate the significance of NightReID and validate the efficacy, flexibility, and applicability of the EDA framework. Weijian Ruan, He Li 0054, Mang Ye |
AAAI | 3 |
| 2025 | FedSPA: Generalizable Federated Graph Learning under Homophily HeterogeneityabstractFederated Graph Learning (FGL) has emerged as a solution to address real-world privacy concerns and data silos in graph learning, which relies on Graph Neural Networks (GNNs). Nevertheless, the homophily level discrepancies within the local graph data of clients, termed homophily heterogeneity, significantly degrade the generalizability of a global GNN. Existing research ignores this issue and suffers from unpromising collaboration. In this paper, we propose FedSPA, an effective framework that addresses homophily heterogeneity from the perspectives of homophily conflict and homophily bias. In the first place, the homophily conflict arises when training on inconsistent homophily levels across clients. Correspondingly, we propose Subgraph Feature Propagation Decoupling (SFPD), thereby achieving collaboration on unified homophily levels across clients. To further address homophily bias, we design Homophily Bias-Driven Aggregation (HBDA) which emphasizes clients with lower biases. It enables the adaptive adjustment of each client contribution to the global GNN based on its homophily bias. The superiority of FedSPA is validated through extensive experiments. The code is available at https://github.com/OakleyTan/FedSPA. Zihan Tan, Guancheng Wan, Wenke Huang 0003, He Li 0054, Guibin Zhang, Carl Yang 0001, Mang Ye |
CVPR | 4 |
| 2025 | Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identificationabstractPerson re-identification (ReID) is the task of matching individuals across different camera views. Existing approaches typically employ neural networks to extract discriminative features, ranking gallery images based on their similarities to probe images. While effective, these methods are often enhanced through re-ranking, a post-processing step that refines initial retrieval results without requiring additional model training. However, current re-ranking methods mostly rely on k-nearest neighbor search to extract similar images that might have the same identity as the query, which is time-consuming with a high computation burden, limiting their applications in reality. We rethink the effect of the k-nearest neighbor search and introduce the Chebyshev’s Theorem-guided Graph Re-ranking (Cheb-GR) method, which adopts the adaptive neighbor search guided by Chebyshev’s Theorem over the k-nearest neighbor search for efficient neighbor selection. Our method leverages graph convolution operations to refine image features and achieve robust re-ranking, leading to enhanced retrieval performance. Furthermore, we provide a theoretical analysis based on Chebyshev’s Inequality to elucidate the factors contributing to the strong performance of the proposed method. Our method significantly reduces the computation costs while maintaining relatively strong performance. Through extensive experiments in both general and cross-domain settings, we demonstrate the effectiveness of Cheb-GR and its potential for real-world applications. Jinxi Yang, He Li 0054, Bo Du 0001, Mang Ye |
CVPR | 2 |
| 2025 | Learn from Downstream and Be Yourself in Multimodal Large Language Models Fine-TuningabstractMultimodal Large Language Model (MLLM) has demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets. Fine-tuning MLLM has become a common practice to improve performance on specific downstream tasks. However, during fine-tuning, MLLM often faces the risk of forgetting knowledge acquired during pre-training, which can result in a decline in generalization abilities. To balance the trade-off between generalization and specialization, we propose measuring the parameter importance for both pre-trained and fine-tuning distributions, based on frozen pre-trained weight magnitude and accumulated fine-tuning gradient values. We further apply an importance-aware weight allocation strategy, selectively updating relatively important parameters for downstream tasks. We conduct empirical evaluations on both image captioning and visual question-answering tasks using various MLLM architectures. The comprehensive experimental analysis demonstrates the effectiveness of the proposed solution, highlighting the efficiency of the crucial modules in enhancing downstream specialization performance while mitigating generalization degradation in MLLM Fine-Tuning. Wenke Huang 0003, Jian Liang 0003, Zekun Shi, Didi Zhu, Guancheng Wan, He Li 0054, Bo Du 0001, Dacheng Tao, Mang Ye |
ICML | 6 |
| 2025 | Be Confident: Uncovering Overfitting in MLLM Multi-Task TuningabstractFine-tuning Multimodal Large Language Models (MLLMs) in multi-task learning scenarios has emerged as an effective strategy for achieving cross-domain specialization. However, multi-task fine-tuning frequently induces performance degradation on open-response datasets. We posit that free-form answer generation primarily depends on language priors, and strengthening the integration of visual behavioral cues is critical for enhancing prediction robustness. In this work, we propose Noise Resilient Confidence Alignment to address the challenge of open-response overfitting during multi-task fine-tuning. Our approach prioritizes maintaining consistent prediction patterns in MLLMs across varying visual input qualities. To achieve this, we employ Gaussian perturbations to synthesize distorted visual inputs and enforce token prediction confidence alignment towards the normal visual branch. By explicitly linking confidence calibration to visual robustness, this method reduces over-reliance on language priors. We conduct extensive empirical evaluations across diverse multi-task downstream settings via popular MLLM architectures. The comprehensive experiment demonstrates the effectiveness of our method, showcasing its ability to alleviate open-response overfitting while maintaining satisfying multi-task fine-tuning performance. Wenke Huang 0003, Jian Liang 0003, Guancheng Wan, Didi Zhu, He Li 0054, Jiawei Shao, Mang Ye, Bo Du 0001, Dacheng Tao |
ICML | 5 |
| 2025 | Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on fine-tuning or manual annotations, which are resource-intensive and limit scalability. In this work, we focus on improving the ability of MLLMs to capture emotions during the inference phase. Specifically, MLLMs encounter two main issues: they struggle to distinguish between semantically similar emotions, leading to misclassification, and they are overwhelmed by redundant or irrelevant visual information, which distracts from key emotional cues. To address these, we propose Sharpening Emotion Perception in MLLMs (SEPM), which incorporates a Confidence-Guided Coarse-to-Fine Inference framework to refine emotion classification by guiding the model through simpler tasks. Additionally, SEPM employs Focus-on-Emotion Visual Augmentation to reduce visual redundancy by directing the attention of models to relevant emotional cues in images. Experimental results demonstrate that SEPM significantly improves MLLM performance on emotion-related tasks, providing a resource-efficient and scalable solution for emotion recognition. Yiyang Fang, Jian Liang 0001, Wenke Huang 0003, He Li 0054, Kehua Su, Mang Ye |
ICML | 4 |
| 2025 | EAGLES: Towards Effective, Efficient, and Economical Federated Graph Learning via Unified SparsificationabstractFederated Graph Learning (FGL) has gained significant attention as a privacy-preserving approach to collaborative learning, but the computational demands increase substantially as datasets grow and Graph Neural Network (GNN) layers deepen. To address these challenges, we propose $\textbf{EAGLES}$, a unified sparsification framework. EAGLES applies client-consensus parameter sparsification to generate multiple unbiased subnetworks at varying sparsity levels, reducing the need for iterative adjustments and mitigating performance degradation. In the graph structure domain, we introduced a dual-expert approach: a $\textit{graph sparsification expert}$ uses multi-criteria node-level sparsification, and a $\textit{graph synergy expert}$ integrates contextual node information to produce optimal sparse subgraphs. Furthermore, the framework introduces a novel distance metric that leverages node contextual information to measure structural similarity among clients, fostering effective knowledge sharing. We also introduce the $\textbf{Harmony Sparsification Principle}$, EAGLES balances model performance with lightweight graph and model structures. Extensive experiments demonstrate its superiority, achieving competitive performance on various datasets, such as reducing training FLOPS by 82\% $\downarrow$ and communication costs by 80\% $\downarrow$ on the ogbn-proteins dataset, while maintaining high performance. Zitong Shi, Guancheng Wan, Wenke Huang 0003, Guibin Zhang, He Li 0054, Carl Yang 0001, Mang Ye |
ICML | 5 |
| 2025 | S2FGL: Spatial Spectral Federated Graph LearningabstractFederated Graph Learning (FGL) combines the privacy-preserving capabilities of Federated Learning (FL) with the strong graph modeling capability of Graph Neural Networks (GNNs). Current research addresses subgraph-FL from the structural perspective, neglecting the propagation of graph signals on the spatial and spectral domains of the structure. From a spatial perspective, subgraph-FL introduces edge disconnections between clients, leading to disruptions in label signals and a degradation in the semantic knowledge of the global GNN. From a spectral perspective, spectral heterogeneity causes inconsistencies in signal frequencies across subgraphs, which makes local GNNs overfit the local signal propagation schemes. As a result, spectral client drift occurs, undermining global generalizability. To tackle the challenges, we propose a global knowledge repository to mitigate the challenge of poor semantic knowledge caused by label signal disruption. Furthermore, we design a frequency alignment to address spectral client drift. The combination of Spatial and Spectral strategies forms our framework $S^2$FGL. Extensive experiments on multiple datasets demonstrate the superiority of $S^2$FGL. The code is available at https://github.com/Wonder7racer/S2FGL.git. Zihan Tan, Suyuan Huang 0003, Guancheng Wan, Wenke Huang 0003, He Li 0054, Mang Ye |
ICML | 5 |
| 2025 | Pixel-wise Divide and Conquer for Federated Vessel SegmentationabstractAccurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federated learning settings due to vessel morphology inconsistency and vessel-background imbalance. The former limits the ability of models to capture fine-grained vessels, while the latter overemphasizes background pixels and biases the model towards them. To address these challenges, we propose a novel method named Federated Vessel-Aware Calibration (FVAC), which leverages global uncertainty to provide differentiated guidance for clients, focusing on pixels of various morphologies that are difficult to distinguish. Furthermore, we introduce a foreground-background decoupling alignment strategy that utilizes more stable and balanced global features to mitigate semantic drift caused by vessel-background imbalance in local clients. Comprehensive experiments confirm the effectiveness of our method Wenke Huang 0003, Zhihao Wang 0002, Zekun Shi, He Li 0054, Mang Ye, Bo Du 0001, Yongchao Xu |
IJCAI | 5 |
| 2025 | Prototype-guided Knowledge Propagation with Adaptive Learning for Lifelong Person Re-identificationabstractLifelong Person Re-identification (LReID) is essential in dynamic camera networks, which continually adapts to new environments while preserving previously acquired knowledge. Existing LReID techniques often preserve samples from past datasets to maintain old knowledge, potentially leading to privacy risks. While prototype-based methods offer privacy advantages, current approaches primarily focus on adjusting classifiers for image classification tasks, neglecting representation biases between old and new identities in person re-identification. This study introduces a novel Prototype-guided Knowledge Propagation (PKP) method, which mitigates discrepancies in similar identity images between old and new tasks by guiding prototype construction through triplet loss constraints. Additionally, to address disparities between prototypes and the updated feature extractor, an Adaptive Parameter Evolution (APE) strategy is proposed. APE optimizes the integration of the old and new models by assessing the importance of the new tasks, dynamically selecting the most pertinent parameters for updates according to their contribution to the current task. Extensive experiments on the LReID benchmark demonstrate that our approach surpasses state-of-the-art prototype-based LReID methods in terms of mAP and rank-1 accuracy. Code is available at https://github.com/joyner-7/IJCAI2025-PKA. Zhijie Lu, Wuxuan Shi, He Li 0054, Mang Ye |
IJCAI | 3 |
| 2025 | MLLMs Meet Person Re-identificationabstractPerson re-identification (Re-ID) models have achieved remarkable advancements with the advent of deep learning. However, their performance often degrades in diverse scenarios, such as variations in viewing angles, lighting conditions, and environmental changes. These limitations arise from the difficulty in generalizing across multiple factors, including environments and subject appearances. Multimodal Large Language Models (MLLMs) offer a promising alternative to address these challenges by leveraging generalized knowledge, as demonstrated in biometric tasks like face and iris recognition. This study explores the Re-ID capabilities of MLLMs by comparatively evaluating six representative MLLMs on the most challenging scenarios, including angle variation, illumination differences, clothing changes, image corruption, and visually fine-grained scenarios in Re-ID. We find that GPT-4o outperforms other MLLMs in handling angle variation, illumination differences, corruption resistance, and fine-grained detail disturbances, demonstrating high accuracy and robustness in challenging Re-ID scenarios. However, further optimization is required for robustness against illumination variation, corruption handling, and fine-grained identification across all tested MLLMs. Additionally, the Re-ID performance of MLLMs can be improved by applying several prompt templates. Our research suggests potential directions for integrating MLLMs into Re-ID systems to enhance performance and robustness, underscoring their promising potential in this field. Mengying Duan, He Li 0054, Mang Ye |
ACM Multimedia | 2 |
| 2025 | MoodAngels: A Retrieval-augmented Multi-agent Framework for Psychiatry DiagnosisabstractThe application of AI in psychiatric diagnosis faces significant challenges, including the subjective nature of mental health assessments, symptom overlap across disorders, and privacy constraints limiting data availability. To address these issues, we present MoodAngels, the first specialized multi-agent framework for mood disorder diagnosis. Our approach combines granular-scale analysis of clinical assessments with a structured verification process, enabling more accurate interpretation of complex psychiatric data. Complementing this framework, we introduce MoodSyn, an open-source dataset of 1,173 synthetic psychiatric cases that preserves clinical validity while ensuring patient privacy. Experimental results demonstrate that MoodAngels outperforms conventional methods, with our baseline agent achieving 12.3\% higher accuracy than GPT-4o on real-world cases, and our full multi-agent system delivering further improvements. Together, these contributions provide both an advanced diagnostic tool and a critical research resource for computational psychiatry, bridging important gaps in AI-assisted mental health assessment. Mengxi Xiao, Ben Liu 0002, He Li 0054, Jimin Huang, Qianqian Xie, Xiaofen Zong, Mang Ye, Min Peng 0002 |
NeurIPS | 3 |
| 2025 | DKDR: Dynamic Knowledge Distillation for Reliability in Federated LearningabstractFederated Learning (FL) has demonstrated a promising future in privacy-friendly collaboration but it faces the data heterogeneity problem. Knowledge Distillation (KD) can serve as an effective method to address this issue. However, challenges arise from the unreliability of existing distillation methods in multi-domain scenarios.
Prevalent distillation solutions primarily aim to fit the distributions of the global model directly by minimizing forward Kullback-Leibler divergence (KLD). This results in significant bias when the outputs of the global model are multi-peaked, which indicates the unreliability of the distillation pathway. Meanwhile, cross-domain update conflicts can notably reduce the accuracy of the global model (teacher model) in certain domains, reflecting the unreliability of the teacher model in these domains.
In this work, we propose DKDR (Dynamic Knowledge Distillation for Reliability in Federated Learning), which dynamically assigns weights to forward and reverse KLD based on knowledge discrepancies. This enables clients to fit the outputs from the teacher precisely. Moreover, we use knowledge decoupling to identify domain experts, thus clients can acquire reliable domain knowledge from experts. Empirical results from single-domain and multi-domain image classification tasks demonstrate the effectiveness of the proposed method and the efficiency of its key modules. The code is available at https://github.com/YueyangYuan/DKDR. Yueyang Yuan, Wenke Huang 0003, Frank Wan, Kaiqi Guan, He Li 0054, Mang Ye |
NeurIPS | 5 |
| 2025 | Self-knowledge distillation with dimensional history knowledge
Wenke Huang 0003, Mang Ye, Zekun Shi, He Li 0054, Bo Du 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Revisiting federated learning with label skew: an over-confidence perspective
Mang Ye, Wenke Huang 0003, Zekun Shi, He Li 0054, Bo Du 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | A Novel Manifold Optimization Algorithm With the Dual Function and a Fuzzy Valuation StepabstractFuzzy mathematical theory is widely used, fuzzy optimization is a branch of fuzzy mathematical theory, the significant application area is artificial intelligence in computer science, especially machine learning (deep learning) and pattern recognition. Fuzzy mathematics, especially fuzzy optimization, has become a bridge between the manifold optimization theory and deep learning applications, which is an essential theoretical foundation. The manifold optimization algorithm employs the projection method, which is unstable. In order to resolve the problem, in this article, the theory and methodology of manifold optimization concerning real and complex spaces is fully considered. Our primary focus is on the Riemannian manifold, where a groundbreaking optimization algorithm with the dual function and a fuzzy valuation step is proposed. To accelerate the convergence and enhance the stability of the optimization algorithm, a novel learning rate is present, which is referred as bivariate gradual learning rate warm-up. A comprehensive analysis of its convergence rates is conducted in various scenarios and the experiments results substantiate our discoveries, and demonstrate the correctness and effectiveness of our devised algorithm. Youfa Liu, He Li 0054, Jingui Zou |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | Kindle Federated Generalization With Domain Specialized and Invariant KnowledgeabstractFederated learning, hailed as a privacy-preserving collaboration paradigm, has garnered significant attention in research circles. Typically, it involves multiple clients collaborating to integrate multi-party knowledge, facilitating the learning of a shared global model with decentralized local data. Despite the popularity of federated learning, the surge in approaches addressing various realistic challenges has highlighted a critical issue. The aggregated model may struggle to capture diverse domain knowledge across participants, leading to limited performance in cross-client domain scenarios. Furthermore, the incorporation of knowledge from participating parties can hinder generalization on out-of-client distributions. To comprehensively address this challenge, we dissect federated generalization into two dimensions: the participating domain and the unseen domain. In this paper, we propose a novel solution incorporating domain-specialized and invariant experts. These experts are designed to faithfully represent individual domain characteristics and different domain universality. Additionally, we introduce a pioneering test-time expert aggregation strategy that utilizes prediction consistency metrics to aggregate different experts, specifically tailored for handling agnostic testing distributions. Empirical results validate that our proposed methodology significantly enhances federated performance on both cross-client and out-of-client generalization under different scenarios and with various related methods. A comprehensive ablation study demonstrates the effectiveness of the proposed modules. Wenke Huang 0003, Mang Ye, Zekun Shi, He Li 0054, Bo Du 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Concept-Based Lesion Aware Transformer for Interpretable Retinal Disease DiagnosisabstractExisting deep learning methods have achieved remarkable results in diagnosing retinal diseases, showcasing the potential of advanced AI in ophthalmology. However, the black-box nature of these methods obscures the decision-making process, compromising their trustworthiness and acceptability. Inspired by the concept-based approaches and recognizing the intrinsic correlation between retinal lesions and diseases, we regard retinal lesions as concepts and propose an inherently interpretable framework designed to enhance both the performance and explainability of diagnostic models. Leveraging the transformer architecture, known for its proficiency in capturing long-range dependencies, our model can effectively identify lesion features. By integrating with image-level annotations, it achieves the alignment of lesion concepts with human cognition under the guidance of a retinal foundation model. Furthermore, to attain interpretability without losing lesion-specific information, our method employs a classifier built on a cross-attention mechanism for disease diagnosis and explanation, where explanations are grounded in the contributions of human-understandable lesion concepts and their visual localization. Notably, due to the structure and inherent interpretability of our model, clinicians can implement concept-level interventions to correct the diagnostic errors by simply adjusting erroneous lesion predictions. Experiments conducted on four fundus image datasets demonstrate that our method achieves favorable performance against state-of-the-art methods while providing faithful explanations and enabling concept-level interventions. Our code is publicly available at https://github.com/Sorades/CLAT. Chi Wen, Mang Ye, He Li 0054 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | All in One Framework for Multimodal Re-Identification in the WildabstractIn Re-identification (ReID), recent advancements yield noteworthy progress in both unimodal and cross-modal re-trieval tasks. However, the challenge persists in developing a unified framework that could effectively handle varying multimodal data, including RGB, infrared, sketches, and textual information. Additionally, the emergence of large-scale models shows promising performance in various vision tasks but the foundation model in ReID is still blank. In response to these challenges, a novel multimodal learning paradigm for ReID is introduced, referred to as All-in-One (AIO), which harnesses a frozen pre-trained big model as an encoder, enabling effective multimodal re-trieval without additional fine-tuning. The diverse multi-modal data in AIO are seamlessly tokenized into a unified space, allowing the modality-shared frozen encoder to extract identity-consistent features comprehensively across all modalities. Furthermore, a meticulously crafted ensemble of cross-modality heads is designed to guide the learning trajectory. AIO is the first framework to perform all-in-one ReID, encompassing four commonly used modali-ties. Experiments on cross-modal and multimodal ReID reveal that AIO not only adeptly handles various modal data but also excels in challenging contexts, showcasing exceptional performance in zero-shot and domain generalization scenarios. Code will be available at: https://github.com/lihe404/AIO. He Li 0054, Mang Ye, Ming Zhang 0019, Bo Du 0001 |
CVPR | 1 |
| 2024 | Self-Driven Entropy Aggregation for Byzantine-Robust Heterogeneous Federated LearningabstractFederated learning presents massive potential for privacy-friendly collaboration. However, the performance of federated learning is deeply affected by byzantine attacks, where malicious clients deliberately upload crafted vicious updates. While various robust aggregations have been proposed to defend against such attacks, they are subject to certain assumptions: homogeneous private data and related proxy datasets. To address these limitations, we propose Self-Driven Entropy Aggregation (SDEA), which leverages the random public dataset to conduct Byzantine-robust aggregation in heterogeneous federated learning. For Byzantine attackers, we observe that benign ones typically present more confident (sharper) predictions than evils on the public dataset. Thus, we highlight benign clients by introducing learnable aggregation weight to minimize the instance-prediction entropy of the global model on the random public dataset. Besides, with inherent data heterogeneity in federated learning, we reveal that it brings heterogeneous sharpness. Specifically, clients are optimized under distinct distribution and thus present fruitful predictive preferences. The learnable aggregation weight blindly allocates high attention to limited ones for sharper predictions, resulting in a biased global model. To alleviate this problem, we encourage the global model to offer diverse predictions via batch-prediction entropy maximization and conduct clustering to equally divide honest weights to accommodate different tendencies. This endows SDEA to detect Byzantine attackers in heterogeneous federated learning. Empirical results demonstrate the effectiveness. Wenke Huang 0003, Zekun Shi, Mang Ye, He Li 0054, Bo Du 0001 |
ICML | 4 |
| 2024 | Parameter Disparities Dissection for Backdoor Defense in Heterogeneous Federated LearningabstractBackdoor attacks pose a serious threat to federated systems, where malicious clients optimize on the triggered distribution to mislead the global model towards a predefined target. Existing backdoor defense methods typically require either homogeneous assumption, validation datasets, or client optimization conflicts. In our work, we observe that benign heterogeneous distributions and malicious triggered distributions exhibit distinct parameter importance degrees. We introduce the Fisher Discrepancy Cluster and Rescale (FDCR) method, which utilizes Fisher Information to calculate the degree of parameter importance for local distributions. This allows us to reweight client parameter updates and identify those with large discrepancies as backdoor attackers. Furthermore, we prioritize rescaling important parameters to expedite adaptation to the target distribution, encouraging significant elements to contribute more while diminishing the influence of trivial ones. This approach enables FDCR to handle backdoor attacks in heterogeneous federated learning environments. Empirical results on various heterogeneous federated scenarios under backdoor attacks demonstrate the effectiveness of our method. Wenke Huang 0003, Mang Ye, Zekun Shi, Guancheng Wan, He Li 0054, Bo Du 0001 |
NeurIPS | 5 |
| 2024 | Federated Learning for Generalization, Robustness, Fairness: A Survey and BenchmarkabstractFederated learning has emerged as a promising paradigm for privacy-preserving collaboration among different parties. Recently, with the popularity of federated learning, an influx of approaches have delivered towards different realistic challenges. In this survey, we provide a systematic overview of the important and recent developments of research on federated learning. First, we introduce the study history and terminology definition of this area. Then, we comprehensively review three basic lines of research: generalization, robustness, and fairness, by introducing their respective background concepts, task settings, and main challenges. We also offer a detailed overview of representative literature on both methods and datasets. We further benchmark the reviewed methods on several well-known datasets. Finally, we point out several open issues in this field and suggest opportunities for further research. Wenke Huang 0003, Mang Ye, Zekun Shi, Guancheng Wan, He Li 0054, Bo Du 0001, Qiang Yang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Rethinking Federated Learning with Domain Shift: A Prototype ViewabstractFederated learning shows a bright promise as a privacy-preserving collaborative learning technique. However, prevalent solutions mainly focus on all private data sampled from the same domain. An important challenge is that when distributed data are derived from diverse domains. The private model presents degenerative performance on other domains (with domain shift). Therefore, we expect that the global model optimized after the federated learning process stably provides generalizability performance on multiple domains. In this paper, we propose Federated Proto-types Learning (FPL) for federated learning under domain shift. The core idea is to construct cluster prototypes and unbiased prototypes, providing fruitful domain knowledge and a fair convergent target. On the one hand, we pull the sample embedding closer to cluster prototypes belonging to the same semantics than cluster prototypes from distinct classes. On the other hand, we introduce consistency regularization to align the local instance with the respective unbiased prototype. Empirical results on Digits and Office Caltech tasks demonstrate the effectiveness of the proposed solution and the efficiency of crucial modules. Wenke Huang 0003, Mang Ye, Zekun Shi, He Li 0054, Bo Du 0001 |
CVPR | 4 |
| 2022 | Pyramidal Transformer with Conv-Patchify for Person Re-identificationabstractThe robust and discriminative feature extraction is the key component in person re-identification (Re-ID). The major weakness of conventional convolution neural network (CNN) based methods is that they cannot extract long-range information from diverse parts, which can be alleviated by recently developed Transformers. Existing vision Transformers show their power on various vision tasks. However, they (i) cannot address translation problems and different viewpoints; (ii) cannot capture detailed features to discriminate people with a similar appearance. In this paper, we propose a powerful Re-ID baseline built on top of the pyramidal transformer with conv-patchify operation, termed PTCR, which inherits the advantages of both CNN and Transformer. The pyramidal structure captures multi-scale fine-grained features, while the conv-patchify enhances the robustness against translation. Moreover, we additionally design two novel modules to improve the robust feature learning. A Token Perception module augments the patch embeddings to enhance the robustness against perturbation and viewpoint changes, while the Auxiliary Embedding module integrates the auxiliary information (cam ID, pedestrian attributes, etc) to reduce feature bias caused by non-visual factors. Our method is validated through extensive experiments to show its superior performance with abundant ablation studies. Notably, without re-ranking, we achieve 98.0% Rank-1 on Market-1501 and 88.6% Rank-1 on MSMT17, significantly outperforming the counterparts. The code is available at: https://github.com/lihe404/PTCR He Li 0054, Mang Ye, Cong Wang 0039, Bo Du 0001 |
ACM Multimedia | 1 |
| 2022 | Collaborative Refining for Person Re-Identification With Label NoiseabstractExisting person re-identification (Re-ID) methods usually rely heavily on large-scale thoroughly annotated training data. However, label noise is unavoidable due to inaccurate person detection results or annotation errors in real scenes. It is extremely challenging to learn a robust Re-ID model with label noise since each identity has very limited annotated training samples. To avoid fitting to the noisy labels, we propose to learn a prefatory model using a large learning rate at the early stage with a self-label refining strategy, in which the labels and network are jointly optimized. To further enhance the robustness, we introduce an online co-refining (CORE) framework with dynamic mutual learning, where networks and label predictions are online optimized collaboratively by distilling the knowledge from other peer networks. Moreover, it also reduces the negative impact of noisy labels using a favorable selective consistency strategy. CORE has two primary advantages: it is robust to different noise types and unknown noise ratios; it can be easily trained without much additional effort on the architecture design. Extensive experiments on Re-ID and image classification demonstrate that CORE outperforms its counterparts by a large margin under both practical and simulated noise settings. Notably, it also improves the state-of-the-art unsupervised Re-ID performance under standard settings. Code is available at https://github.com/mangye16/ReID-Label-Noise. Mang Ye, He Li 0054, Bo Du 0001, Jianbing Shen, Ling Shao 0001, Steven C. H. Hoi |
IEEE Trans. Image Process. | 2 |
| 2021 | WePerson: Learning a Generalized Re-identification Model from All-weather Virtual DataabstractThe aim of person re-identification (Re-ID) is retrieving a person of interest across multiple non-overlapping cameras. Re-ID has gained significantly increased advancement in recent years. However, real data annotation is costly and model generalization ability is hindered by the lack of large-scale and diverse data. To address this problem, we propose a Weather Person pipeline that can generate a synthesized Re-ID dataset with different weather, scenes, and natural lighting conditions automatically. The pipeline is built on the top of a game engine which contains a digital city, weather and lighting simulation system, and various character models with manifold dressing. To train a generalizable Re-ID model from the large-scale virtual WePerson dataset, we design an adaptive sample selection strategy to close the domain gap and avoid redundancy. We also design an informative sampling method for a mini-batch sampler to accelerate the learning process. In addition, an efficient training method is introduced by adopting instance normalization to capture identity invariant components from various appearances. We evaluate our pipeline using direct transfer on 3 widely-used real-world benchmarks, achieving competitive performance without any real-world image training. This dataset starts the attempt to evaluate diverse environmental factors in a controllable virtual engine, which provides important guidance for future generalizable Re-ID model design. Notably, we improve the current state-of-the-art accuracy from 38.5% to 46.4% on the challenging MSMT17 dataset. Dataset and code are available at https://github.com/lihe404/WePerson https://github.com/lihe404/WePerson. He Li 0054, Mang Ye, Bo Du 0001 |
ACM Multimedia | 1 |