EDBT 2026 Demo / reviewers in the wild / expert
Mingrui Lao
dblp:222/4779
· DBLP profile ↗
33ranked-venue papers
8as first author
33since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | JAIL: Adaptive multi-turn jailbreak attacks reveal limitations of LLM safety alignment
Yunhao Feng, Mingrui Lao, Yishan Li, Yuxiang Xie, Yanming Guo |
Expert Syst. Appl. | 2 |
| 2026 | Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningabstractIn real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks. Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Yanming Guo, Huiping Zhuang |
AAAI | 6 |
| 2026 | EDAgent: A Concurrent Orchestration Method for Collaborative LLM Multi-Agent System
Zongyang Yuan, Zechang Zhang, Qinbin Li, Lailong Luo, Deke Guo, Mingrui Lao |
ICDCS | 6 |
| 2026 | OTKD: A general knowledge distillation pipeline for object tracking
Yongqi Pan, Lailong Luo, Hanlin Tan, Mingrui Lao, Yuxuan Liang 0002 |
Expert Syst. Appl. | 5 |
| 2026 | Overcoming semantic manifold deviation for robust multimodal violence detection with incomplete modality
Jianan Zhu, Yanming Guo, Lai Kang, Yirun Ruan, Mingrui Lao |
Expert Syst. Appl. | 6 |
| 2026 | Closed-loop correction reprogramming for fine-grained visual prompting
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001 |
Neural Networks | 4 |
| 2026 | FSKD: A few-shot knowledge distillation framework for object tracking
Yongqi Pan, Lailong Luo, Mingrui Lao, Qianzhen Zhang, Xianqiang Zhu |
Pattern Recognit. | 4 |
| 2026 | PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001 |
Pattern Recognit. | 6 |
| 2025 | Multi-Modal Entities Matter: Benchmarking Multi-Modal Entity AlignmentabstractMulti-modal entity alignment (MMEA) is a long-standing task that aims to discover identical entities between different multi-modal knowledge graphs (MMKGs). However, most of the existing MMEA datasets consider the multi-modal data as the attributes of textual entities, while neglecting the correlations among the multi-modal data and do not fit in the real-world scenarios well. In response, in this work, we establish a novel yet practical MMEA dataset, i.e. NMMEA, which models multi-modal data (e.g., images) equally as textual entities in the MMKG. Due to the introduction of multi-modal data, NMMEA poses new challenges to existing MMEA solutions, i.e., heterogeneous structural representation learning and cross-modal alignment inference. Hence, we put forward a simple yet effective solution, CrossEA, which can effectively learn the structural information of entities by considering both intra-modal and cross-modal relations, and further infer the similarity of different types of entity pairs. Extensive experiments validate the significance of NMMEA, where CrossEA can achieve superior performance in contrast to competitive methods on the proposed dataset. Guanchen Xiao, Weixin Zeng, Shiqi Zhang 0011, Mingrui Lao, Xiang Zhao 0002 |
COLING | 4 |
| 2025 | Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake Detection
Xueyi Zhang 0001, Peiyin Zhu, Zhiyuan Yan 0002, Jikang Cheng, Mingrui Lao, Siqi Cai 0002, Yanming Guo |
ICCV | 6 |
| 2025 | Boosting Adversarial Robustness Through Structure-Guided Adversarial Distillation
Yanming Guo, Chengsi Du, Dengjin Li, Mingrui Lao |
ICIC (2) | 6 |
| 2025 | Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake DetectionabstractThe rapid evolution of deepfake techniques presents dual challenges for detection models: adapting to continuously shifting attack distributions while retaining previously learned knowledge. Although recent continual deepfake detection methods have made progress, they often rely on replay-based training, which limits scalability and deployment. Meanwhile, the task structure of deepfake detection offers a unique opportunity that remains under-explored: it is inherently a binary classification problem with a fixed label space, where the main difficulty lies in distributional drift rather than class expansion. This insight enables the modeling of each incremental distribution shift as a dedicated expert, focusing on specific forgery patterns. To this end, we propose a novel analytically driven, replay-free continual detection framework that eliminates the need for iterative gradient updates. In this framework, task-specific experts are constructed via closed-form ridge regression, requiring only a single forward pass and ensuring non-interference with previous tasks. To enhance the model's capacity for fine-grained forgery recognition, we introduce a lightweight Forgery-Aware Residual Enhancer (FARE). At inference, an Uncertainty-Guided Expert Selection module (UGES) dynamically routes each sample to the most confident expert, which does not require prior knowledge of the attack type. The proposed framework achieves a favorable trade-off between efficiency, privacy, and generalization. It achieves state-of-the-art performance across four benchmark datasets, with an average accuracy of 91.82% and only 1.78% forgetting. Notably, it improves cross-forgery generalization by 9.28% on unseen forgery types, demonstrating strong generalization. Xueyi Zhang 0001, Peiyin Zhu, Jinping Sui, Xiaoda Yang, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Jun Tang 0001 |
ACM Multimedia | 6 |
| 2025 | EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingabstractEvent cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading. Xueyi Zhang 0001, Jialu Sun, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001 |
ACM Multimedia | 7 |
| 2025 | TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient ProjectionabstractPrompt learning has emerged as an efficient adaptation paradigm for vision-language models (VLMs), yet it remains highly vulnerable to label noise, which limits its real-world applicability. We propose TrustCLIP, a noise-robust prompt tuning framework that leverages the inherent semantic structure of CLIP through two key components: Semantic Label Verification (SLV) and Trust-aligned Gradient Projection (TGP). SLV defines a semantic trust boundary based on CLIP's zero-shot predictions to identify reliable samples for standard supervised training. For uncertain samples, TGP projects their gradients into a trust-aligned subspace constructed from the gradients of clean samples, thereby preserving semantically aligned learning signals while suppressing noise-induced optimization drift. Unlike prior approaches, TrustCLIP doesn't require additional parameters, loss reweighting, or uncertainty estimation. Extensive experiments on 7 benchmark datasets with both synthetic and real-world noisy labels demonstrate that TrustCLIP consistently outperforms state-of-the-art methods in terms of both robustness and transferability. Xueyi Zhang 0001, Peiyin Zhu, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Haizhou Li 0001 |
ACM Multimedia | 5 |
| 2025 | Learning from Peers: Collaborative Ensemble Adversarial Training
Dengjin Li, Yanming Guo, Yuxiang Xie, Jiangming Chen, Mingrui Lao |
PRCV (1) | 7 |
| 2025 | Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality MissingabstractMultimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies. Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001 |
SIGIR | 1 |
| 2025 | FCAT: Federated causal adversarial training
Yunhao Feng, Yanming Guo, Mingrui Lao, Yishan Li, Yuxiang Xie |
Knowl. Based Syst. | 3 |
| 2024 | Boosting Adversarial Robustness Distillation Via Hybrid Decomposed KnowledgeabstractAdversarial Robust Distillation (ARD) has emerged as a potent defense mechanism tailored to small models against adversarial threats. However, mainstream ARD methods typically exploit teachers’ response as the transferred knowledge, while neglecting the analysis of involved target-related knowledge to mitigate adversarial attacks. Furthermore, these methods primarily focus on logits-level distillation, which overlook the features-level knowledge in teacher models. In this paper, we introduce a novel Hybrid Decomposed Distillation (HDD) approach, which attempts to identify the vital knowledge against adversarial threats through dual-level distillation. Specifically, we first seek to separate the predictions of teacher model into target-related and target-unrelated knowledge for flexible yet efficient logits-level distillation. Besides, to further boost the distillation efficacy, HDD leverages the channel correlations to decompose intermediate features into highly and less relevant components. Extensive experiments on two benchmarks demonstrate that our HDD achieves superior performance in both clean accuracy and robustness, in contrast to current state-of-the-art methods. Mingrui Lao, Yanming Guo |
ICASSP | 2 |
| 2024 | Balanced Confidence Calibration for Graph Neural NetworksabstractThis paper delves into the confidence calibration in prediction when using Graph Neural Networks (GNNs), which has emerged as a notable challenge in the field. Despite their remarkable capabilities in processing graph-structured data, GNNs are prone to exhibit lower confidence in their predictions than what the actual accuracy warrants. Recent advances attempt to address this by minimizing prediction entropy to enhance confidence levels. However, this method inadvertently risks leading to over-confidence in model predictions. Our investigation in this work reveals that most existing GNN calibration methods predominantly focus on the highest logit, thereby neglecting the entire spectrum of prediction probabilities. To alleviate this limitation, we introduce a novel framework called Balanced Calibrated Graph Neural Network (BCGNN), specifically designed to establish a balanced calibration between over-confidence and under-confidence in GNNs' prediction. To theoretically support our proposed method, we further demonstrate the mechanism of the BCGNN framework in effective confidence calibration and significant trustworthiness improvement in prediction. We conduct extensive experiments to examine the developed framework. The empirical results show our method's superior performance in predictive confidence and trustworthiness, affirming its practical applicability and effectiveness in real-world scenarios. Hao Yang 0042, Min Wang 0034, Cheems Wang, Mingrui Lao, Yun Zhou 0001 |
KDD | 4 |
| 2024 | Maximizing Feature Distribution Variance for Robust Neural NetworksabstractThe security of Deep Neural Networks (DNNs) has proven to be critical for their applicabilities in real-world scenarios. However, DNNs are well-known to be vulnerable against adversarial attacks, such as adding artificially designed imperceptible magnitude perturbation to the benign input. Therefore, adversarial robustness is essential for DNNs to defend against malicious attacks. Stochastic Neural Networks (SNNs) have recently shown effective performance on enhancing adversarial robustness by injecting uncertainty into models. Nevertheless, existing SNNs are still limited for adversarial defense, as their insufficient representation capability from the fixed uncertainty. In this paper, to elevate feature representation capability of SNNs, we propose a novel yet practical stochastic neural network that maximizes feature distribution variance (MFDV-SNN). In addition, we provide theoretical insights to support the adversarial resistance of MFDV, which primarily derived from the stochastic noise we injected into DNNs. Our research demonstrates that by gradually increasing the level of stochastic noise in a DNN, the model naturally becomes more resistant to input perturbations. Since adversarial training is not required, MFDV-SNN does not compromise clean data accuracy and saves up to 7.5 times computation time. Extensive experiments on various attacks demonstrate that MFDV-SNN improves adversarial robustness significantly compared to other methods. Hao Yang 0042, Min Wang 0034, Zhengfei Yu, Zhi Zeng 0001, Mingrui Lao, Yun Zhou 0001 |
ACM Multimedia | 5 |
| 2024 | MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental TasksabstractClass-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy. Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 5 |
| 2024 | PD-Refiner: An Underlying Surface Inheritance Refiner with Adaptive Edge-Aware Supervision for Point Cloud DenoisingabstractPoint clouds from real-world scenarios inevitably contain complex noise, significantly impairing the accuracy of downstream tasks. To tackle this challenge, cascading encoder-decoder architecture has become a conventional technical route to iterative denoise. However, circularly feeding the output of denoiser as its input again involves the re-extraction of underlying surface, leading to unstable denoising process and over-smoothed geometric details. To address these issues, we propose a novel denoising paradigm dubbed PD-Refiner that employs a single encoder to model the underlying surface. Then, we leverage several lightweight hierarchical Underlying Surface Inheritance Refiners (USIRs) to inherit and strengthen it, thereby avoiding the re-extraction from the intermediate point cloud. Furthermore, we design adaptive edge-aware supervision to improve the edge awareness of the USIRs, allowing for the adjustment of the denoising preferences from global structure to local details. The results demonstrate that our method not only achieves state-of-the-art performance in terms of denoising stability and efficacy, but also enhances edge clarity and point cloud uniformity. Xueyi Zhang 0001, Xianghu Yue, Mingrui Lao, Tao Jiang 0062, Fubo Zhang, Longyong Chen |
ACM Multimedia | 4 |
| 2024 | Language Without Borders: A Dataset and Benchmark for Code-Switching Lip ReadingabstractLip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR. Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001 |
NeurIPS | 2 |
| 2024 | EIOA: A computing expectation-based influence evaluation method in weighted hypergraphs
Qingtao Pan, Jun Tang 0001, Zhaolin Lv, Yirun Ruan, Tianyuan Yv, Mingrui Lao |
Inf. Process. Manag. | 9 |
| 2023 | COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models are not only prone to over-exploit prevalent language bias, but also suffer from additional joint-modal biases caused by the shortcut relations between textual-auditory/visual co-occurrences and dominated answers. In this paper, we propose a COllabrative CAusal (COCA) Regularization to remedy this more challenging issue of data biases. Specifically, a novel Bias-centered Causal Regularization (BCR) is proposed to alleviate specific shortcut biases by intervening bias-irrelevant causal effects, and further introspect the predictions of AVQA models in counterfactual and factual scenarios. Based on the fact that the dominated bias impairing model robustness for different samples tends to be different, we introduce a Multi-shortcut Collaborative Debiasing (MCD) to measure how each sample suffers from different biases, and dynamically adjust their debiasing concentration to different shortcut correlations. Extensive experiments demonstrate the effectiveness as well as backbone-agnostic ability of our COCA strategy, and it achieves state-of-the-art performance on the large-scale MUSIC-AVQA dataset. Mingrui Lao, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
AAAI | 1 |
| 2023 | Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationabstractVisual Question Answering (VQA) has achieved significant success over the last few years, while most studies focus on training a VQA model on a stationary domain (e.g., a given dataset). In real-world application scenarios, however, these methods are often inefficient because VQA systems are always supposed to extend their knowledge and meet the ever-changing demands of users. In this paper, we introduce a new and challenging multi-domain lifelong VQA task, dubbed MDL-VQA, which encourages the VQA model to continuously learn across multiple domains while mitigating the forgetting on previously-learned domains. Furthermore, we propose a novel replay-free Self-Critical Distillation (SCD) framework tailor-made for MDL-VQA, which alleviates forgetting issue via transferring previous-domain knowledge from teacher to student models. First, we propose to introspect the teacher's understanding over original and counterfactual samples, thereby creating informative instance-relevant and domain-relevant knowledge for logits-based distillation. Second, on the side of feature-based distillation, we propose to introspect the reasoning behavior of student model to establish the harmful domain-specific knowledge acquired in current domain, and further leverage the metric learning strategy to encourage student to learn useful knowledge in new domain. Extensive experiments demonstrate that SCD framework outperforms state-of-the-art competitors with different training orders. Mingrui Lao, Nan Pu, Yu Liu 0012, Zhun Zhong, Erwin M. Bakker, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 1 |
| 2023 | FedVQA: Personalized Federated Visual Question Answering over Heterogeneous ScenesabstractThis paper presents a new setting for visual question answering (VQA) called personalized federated VQA (FedVQA) that addresses the growing need for decentralization and data privacy protection. FedVQA is both practical and challenging, requiring clients to learn well-personalized models on scene-specific datasets with severe feature/label distribution skews. These models then collaborate to optimize a generic global model on a central server, which is desired to generalize well on both seen and unseen scenes without sharing raw data with the server and other clients. The primary challenge of FedVQA is that, client models tend to forget the global knowledge initialized from central server during the personalized training, which impairs their personalized capacity due to the potential overfitting issue on local data. This further leads to divergence issues when aggregating distinct personalized knowledge at the central server, resulting in an inferior generalization ability on unseen scenes. To address the challenge, we propose a novel federated pairwise preference preserving (FedP3) framework to improve personalized learning via preserving generic knowledge under FedVQA constraints. Specifically, we first design a differentiable pairwise preference (DPP) to improve knowledge preserving by formulating a flexible yet effective global knowledge. Then, we introduce a forgotten-knowledge filter (FKF) to encourage the client models to selectively consolidate easily-forgotten knowledge. By aggregating the DPP and the FKF, FedP3 coordinates the generic and the personalized knowledge to enhance the personalized ability of clients and generalizability of the server. Extensive experiments show that FedP3 consistently surpasses the state-of-the-art in FedVQA task. Mingrui Lao, Nan Pu, Zhun Zhong, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 1 |
| 2023 | Dual selective knowledge transfer for few-shot classificationabstractAbstract Few-shot learning aims at recognizing novel visual categories from very few labelled examples. Different from the existing few-shot classification methods that are mainly based on metric learning or meta-learning, in this work we focus on improving the representation capacity of feature extractors. For this purpose, we propose a new two-stage dual selective knowledge transfer (DSKT) framework, to guide models towards better optimization. Specifically, we first exploit an improved multi-task learning approach to train a feature extractor with robust representation capability as a teacher model. Then, we design an effective dual selective knowledge distillation method, which enables the student model to selectively learn knowledge from the teacher model and current samples, thereby improving the student model’s ability to generalize on unseen classes. Extensive experimental results show that our DSKT achieves competitive performances on four well-known few-shot classification benchmarks. Nan Pu, Mingrui Lao, Erwin M. Bakker, Michael S. Lew |
Appl. Intell. | 3 |
| 2023 | Lifelong Fine-Grained Image RetrievalabstractFine-grained image retrieval has been extensively explored in a zero-shot manner. A deep model is trained on the seen part and then evaluated the generalization performance on the unseen part. However, this setting is infeasible for many real-world applications since (1) the retrieval dataset can be non-fixed so that new data are added constantly, and (2) data samples of the seen categories are also common in practice and are important for evaluation. In this paper, we explore lifelong fine-grained image retrieval (LFGIR), which learns continuously on a sequence of new tasks with data from different datasets. We first use knowledge distillation to minimize catastrophic forgetting on old tasks. Training continuously on different datasets causes large domain shifts between the old and new tasks while image retrieval is sensitive to even small shifts in the features. This tends to weaken the effectiveness of knowledge distillation by the frozen teacher. To mitigate the impact of domain shifts, we use the network inversion method to generate images of the old tasks. In addition, we design an on-the-fly teacher which transfers knowledge captured on a new task to the student to improve better generalization performance, thereby achieving a better balance between old and new tasks in the end. We name the whole framework as Dual Knowledge Distillation (DKD), whose efficacy is demonstrated by extensive experimental results on sequential tasks including seven datasets. Wei Chen 0072, Haoyang Xu, Nan Pu, Yu Liu 0012, Mingrui Lao, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 5 |
| 2022 | VQA-BC: Robust Visual Question Answering Via Bidirectional ChainingabstractCurrent VQA models are suffering from the problem of overdependence on language bias, which severely reduces their robustness in real-world scenarios. In this paper, we analyze VQA models from the view of forward/backward chaining in the inference engine, and propose to enhance their robustness via a novel Bidirectional Chaining (VQA-BC) framework. Specifically, we introduce a backward chaining with hardnegative contrastive learning to reason from the consequence (answers) to generate crucial known facts (question-related visual region features). Furthermore, to alleviate the overconfident problem in answer prediction (forward chaining), we present a novel introspective regularization to connect forward and backward chaining with label smoothing. Extensive experiments verify that VQA-BC not only effectively overcomes language bias on out-of-distribution dataset, but also alleviates the over-correct problem caused by ensemble-based method on in-distribution dataset. Compared with competitive debiasing strategies, our method achieves state-of-the-art performance to reduce language bias on VQA-CP v2 dataset. Mingrui Lao, Yanming Guo, Wei Chen 0072, Nan Pu, Michael S. Lew |
ICASSP | 1 |
| 2021 | A Language Prior Based Focal Loss for Visual Question AnsweringabstractAccording to current research, one of the major challenges in Visual Question Answering (VQA) models is the overdependence on language priors (and neglect of the visual modality). VQA models tend to predict answers only based on superficial correlations between the first few words in question and frequency of related answer candidates. To address this issue, we propose a novel Language Prior based Focal Loss (LP-Focal Loss) by rescaling the standard cross entropy loss. Specifically, we employ a question-only branch to capture the language biases for each answer candidate based on the corresponding question input. Then, the LP-Focal Loss dynamically assigns lower weights to biased answers when computing the training loss, thereby reducing the contribution of more-biased instances in the train split. Extensive experiments show that the LP-Focal Loss can be generally applied to common baseline VQA models, and achieves significantly better performance on the VQA-CP v2 dataset, with an overall 18% accuracy boost over benchmark models. Mingrui Lao, Yanming Guo, Yu Liu 0012, Michael S. Lew |
ICME | 1 |
| 2021 | From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringabstractMost Visual Question Answering (VQA) models are faced with language bias when learning to answer a given question, thereby failing to understand multimodal knowledge simultaneously. Based on the fact that VQA samples with different levels of language bias contribute differently for answer prediction, in this paper, we overcome the language prior problem by proposing a novel Language Bias driven Curriculum Learning (LBCL) approach, which employs an easy-to-hard learning strategy with a novel difficulty metric Visual Sensitive Coefficient (VSC). Specifically, in the initial training stage, the VQA model mainly learns the superficial textual correlations between questions and answers (easy concept) from more-biased examples, and then progressively focuses on learning the multimodal reasoning (hard concept) from less-biased examples in the following stages. The curriculum selection of examples on different stages is according to our proposed difficulty metric VSC, which is to evaluate the difficulty driven by the language bias of each VQA sample. Furthermore, to avoid the catastrophic forgetting of the learned concept during the multi-stage learning procedure, we propose to integrate knowledge distillation into the curriculum learning framework. Extensive experiments show that our LBCL can be generally applied to common VQA baseline models, and achieves remarkably better performance on the VQA-CP v1 and v2 datasets, with an overall 20% accuracy boost over baseline models. Mingrui Lao, Yanming Guo, Yu Liu 0012, Wei Chen 0072, Nan Pu, Michael S. Lew |
ACM Multimedia | 1 |
| 2021 | Multi-stage hybrid embedding fusion network for visual question answering
Mingrui Lao, Yanming Guo, Nan Pu, Wei Chen 0072, Yu Liu 0012, Michael S. Lew |
Neurocomputing | 1 |