EDBT 2026 Demo / reviewers in the wild / expert
Feifei Zhang 0001
dblp:53/1871-1
· DBLP profile ↗
44ranked-venue papers
15as first author
32since 2021 · last 2026
0000-0002-8153-9977ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 14 first-author · 27 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OAD-Promoter: Enhancing Zero-Shot VQA Using Large Language Models with Object Attribute DescriptionabstractLarge Language Models (LLMs) have become a crucial tool in Visual Question Answering (VQA) for handling knowledge-intensive questions in few-shot or zero-shot scenarios. However, their reliance on massive training datasets often causes them to inherit language biases during the acquisition of knowledge. This limitation imposes two key constraints on existing methods: (1) LLM predictions become less reliable due to bias exploitation, and (2) despite strong knowledge reasoning capabilities, LLMs still struggle with out-of-distribution (OOD) generalization. To address these issues, we propose Object Attribute Description Promoter (OAD-Promoter), a novel approach for enhancing LLM-based VQA by mitigating language bias and improving domain-shift robustness. OAD-Promoter comprises three components: the Object-concentrated Example Generation (OEG) module, the Memory Knowledge Assistance (MKA) module, and the OAD Prompt. The OEG module generates global captions and object-concentrated samples, jointly enhancing visual information input to the LLM and mitigating bias through complementary global and regional visual cues. The MKA module assists the LLM in handling OOD samples by retrieving relevant knowledge from stored examples to support questions from unseen domains. Finally, the OAD Prompt integrates the outputs of the preceding modules to optimize LLM inference. Experiments demonstrate that OAD-Promoter significantly improves the performance of LLM-based VQA methods in few-shot or zero-shot settings, achieving new state-of-the-art results. Quanxing Xu, Ling Zhou 0005, Feifei Zhang 0001, Rubing Huang, Jinyu Tian 0001 |
AAAI | 3 |
| 2026 | Duplex Rewards Optimization for Test-Time Composed Image RetrievalabstractComposed Image Retrieval (CIR) combines the reference image with text to retrieve the intended target image. Recently, zero-shot CIR has gained significant attention by eliminating the need for labeled triplets required in supervised CIR. However, it inevitably demands additional training corpus, storage, and computational resources, limiting its applicability in real-world scenarios. Inspired by advancements in Test-Time Adaptation (TTA), we propose a Test-Time CIR setting named TT-CIR, which aims to efficiently adapt models to unlabeled test samples while reducing resource consumption. Within the TT-CIR setting, we identify that naively introducing existing TTA methods (e.g., reward-based) into CIR faces two vital challenges: 1) Modification-restricted reward pool, which limits the exploration of semantically relevant candidate rewards; 2) Conservative knowledge feedback, which inhibits the adaptability of rewards to the current data distribution. To address these challenges, we propose a test-time reinforcement learning framework that integrates a Counterfactual-guided Multinomial Sampling (CMS) strategy and a Duplex Rewards Modeling (DRM) module. The CMS explores a candidate reward pool that is visually similar and semantically relevant to the given query, while the DRM generates stable and adaptive duplex rewards to guide model adaptation. Extensive experiments demonstrate the superiority and adaptability of our method over existing approaches. Haoliang Zhou, Feifei Zhang 0001, Changsheng Xu |
AAAI | 2 |
| 2026 | ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 4 |
| 2026 | Refined generation-based framework for consistent and reliable visual question answering
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 4 |
| 2026 | Modeling Semantic and Localization Uncertainty for Weakly Supervised Temporal Action LocalizationabstractThe objective of weakly supervised temporal action localization (WTAL) is to accurately identify the temporal intervals of actions using only video-level annotations for training. Existing cross-modal WTAL methods integrate vision-language models to provide rich semantic supervision, aiming to alleviate the inherent supervision limitations in weakly supervised scenarios. However, it is crucial to acknowledge that contemporary cross-modal methods incorporate textual information simultaneously, which inevitably introduces uncertainty in the alignment of cross-modal semantics. Moreover, previous approaches typically output deterministic temporal localization results, while neglecting to evaluate the predictive uncertainty and confidence of localization results. To address the above issues, we propose a novel Modeling Semantic and Localization Uncertainty (MSLU) framework for WTAL, which can simultaneously model the semantic uncertainty in cross-modal representations and the uncertainty of localization results to achieve more precise and robust temporal action localization. Specifically, we propose the Probabilistic Semantic Uncertainty Modeling (PSUM) module, which utilizes probabilistic encoding to capture diverse cross-modal feature representations, effectively mitigating semantic ambiguity in feature alignment. In addition, we propose the Uncertainty-guided Localization Estimation (ULE) module, which leverages evidential deep learning to estimate predictive uncertainty of localization results in weakly supervised scenarios. Through extensive experiments on benchmark datasets including THUMOS14, ActivityNet1.2, ActivityNet1.3, and FineAction, our framework demonstrates superior performance compared to existing state-of-the-art methods. The empirical results validate the effectiveness of simultaneously modeling both semantic and localization uncertainty. Yuxiang Shao, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | PLMAS: Adaptive Sample Selection for Prompting LLMs in Knowledge-Based Visual Question AnsweringabstractWith the rapid advancement of large-scale model technology, Visual Question Answering (VQA)—a core subfield of multimodal research—increasingly relies on these models to address complex challenges. This trend is especially evident in Knowledge-based VQA (KB-VQA), which requires integrating external knowledge. While most studies approach KB-VQA using explicit or implicit knowledge bases, recent studies employ in-context learning to guide large language models (LLMs) with implicit knowledge (e.g., PICa and Prophet). However, existing sample selection strategies for in-context learning are oversimplified and fail to adequately leverage the tacit knowledge encoded within LLMs. To address this limitation, we propose an adaptive sample selection strategy that integrates triple similarity calculations (question-image, question-caption, and question-pre-answer) and dynamically assembles the most relevant samples using weighted combinations, thereby effectively activating the large model’s implicit knowledge. To evaluate the performance of our proposed approach, we conducted experiments on benchmark datasets. Results demonstrate that our method (PLMAS) achieves state-of-the-art performance on both the OK-VQA and A-OKVQA datasets. Quanxing Xu, Ling Zhou 0005, Feifei Zhang 0001, Rubing Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Enhancing Image Captioning through Bridging Image-Text Gap and Reducing HallucinationsabstractWhile autoregressive models have achieved remarkable success in image captioning, their slow inference speed limits their applicability in real-time scenarios. Non-autoregressive methods provide a promising alternative for faster caption generation; however, they still encounter significant challenges. In particular, they struggle to capture complex content and abstract concepts necessary for producing semantically rich and accurate captions, which hinders the bridging of the image–text gap. Moreover, the generation process often leads to object hallucination—instances where incorrect or non-existent objects are described, resulting in captions that misalign with the actual visual content. To address these issues, we propose the Vision-Text Semantic Reconstruction and Contrast (VTSRC) mechanism, which consists of two key modules. The first module is the Visual-Text Reconstruction Network (VRN), which reconstructs visual representations into textual space, enriching captions with contributive and complex semantics to bridge the image-text gap. The second module is the Visual Contrastive Generation (VCG), which leverages visual uncertainty to contrast distributions, recalibrating the model’s output and significantly reducing the incidence of hallucination, thereby generating coherent linguistic representations. Extensive evaluations demonstrate that our approach markedly improves the creation of semantically rich image captions, considerably reducing the frequency of hallucinations while maintaining high descriptive accuracy. Experimental results demonstrate that VTSRC achieves competitive performance on the challenging MSCOCO image captioning dataset, reaching the best CIDEr score of 133.9% on the COCO-caption Karpathy split to date. Feifei Zhang 0001, Lingkai Ran, Caixia Song, Ling Zhou 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | Concise Object-word Visuals as Effective Cues for Visual Question AnsweringabstractIn Visual Question Answering (VQA) , both the image and its accompanying question serve as the primary sources of information for the model. Conventional approaches typically rely heavily on dense visual representations for reasoning and answer prediction. However, when the visual and textual modalities are imbalanced or semantically misaligned, such disparities hinder effective multimodal learning and inference. To address this issue, we propose a multimodal information adjustment method, the Visual Text Information Adjuster (ViTA) . ViTA investigates the impact of embedding textual cues within images on the VQA process and promotes cross-modal balance to improve accuracy. Specifically, since image content often dominates over question content, ViTA adjusts the balance by either masking visual information or augmenting it with object-word visual cues directly embedded in the image. Experimental results validate our hypothesis and further demonstrate that ViTA can serve as an effective data augmentation strategy, yielding measurable improvements across multiple VQA models. The code will be released at https://github.com/xqx23/ViTA . Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Rubing Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | QG-STR: Training-Time Optimized Question-Guided Scene Text Recognition via Visual Question AnsweringabstractScene Text Spotting (STS) aims to transcribe text embedded in natural images, typically encompassing Scene Text Detection (STD) and Scene Text Recognition (STR) . Advances in image understanding have made end-to-end text spotting increasingly viable. Concurrently, multimodal research has highlighted the potential of vision-language reasoning tasks, such as Visual Question Answering (VQA) . To leverage multimodal reasoning for STR, we propose a training-time question-guided STR framework that integrates VQA, termed Question-Guided STR (QG-STR) . The framework unifies STR, Visual Question Generation (VQG) , and VQA within a single architecture, enabling multimodal reasoning to enhance text-spotting performance. Specifically, visual understanding and logical reasoning are used as supervisory signals during training to improve text recognition accuracy and boost end-to-end text spotting. QG-STR is model-agnostic and compatible with diverse STR and VQA architectures, employing question guidance solely as a training-time supervision mechanism. During inference, the STR module functions independently without requiring external questions. Extensive experiments on Total-Text , ICDAR2015 , ICDAR2013 , and CTW1500 validate the effectiveness of QG-STR. Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Rubing Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and ApproachabstractVisual Question Answering (VQA) is a multifaceted task that integrates computer vision and natural language processing to produce textual answers from images and questions. Existing VQA benchmarks predominantly adhere to a closed-set paradigm, limiting their ability to address arbitrary, unseen answers, and thus falling short in real-world scenarios. To address this limitation, we introduce the Open-Vocabulary Visual Question Answering (OVVQA) benchmark, specifically designed to evaluate models under open-world conditions by assessing their performance on both base classes (seen, common answers) and novel classes (unseen, rare answers). In conjunction with this benchmark, we propose a model-agnostic Causal Adapter to combat the inherent bias found in current VQA tasks. Our approach leverages front-door adjustment to enhance causal reasoning, significantly improving model performance on novel categories while maintaining accuracy on base classes. Additionally, we introduce an adaptive transfer loss to facilitate the transfer of more knowledge from the pretrained model to our OVVQA task. Extensive experiments across multiple datasets validate the superiority of our method over existing state-of-the-art approaches, demonstrating its robust generalization and adaptability in open-world VQA scenarios. Feifei Zhang 0001, Changsheng Xu |
AAAI | 1 |
| 2025 | Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering
Feifei Zhang 0001, Changsheng Xu |
ICCV | 1 |
| 2025 | Diff-ZsVQA: Zero-shot Visual Question Answering with Frozen Large Language Models Using Diffusion ModelabstractVisual Question Answering (VQA) methods leveraging Large Language Models (LLMs) aim to enhance performance in few/zero-shot scenarios. While the results attained by these approaches were outstanding, there remains scope for further enhancement. Given the remarkable capabilities demonstrated by Diffusion Models (DMs), we recognize that the DMs can potentially improve the performance of VQA by optimizing the generation of captions. Furthermore, existing approaches for prompt construction neglect the influence of non-original questions and generated question-answer (QA) pairs, which leads to adverse effects on the inference. This paper proposes a novel framework called Diff used Z ero- s hot VQA, shortly Diff-ZsVQA, which innovatively incorporates a powerful DM into the LLM-based VQA pipeline for image-to-text converting. Moreover, to reduce the impact of non-original questions and generated QA pairs, we devise an Original-Question-Centric (OQC) prompt whose examples’ questions are identical while contexts are diverse. We first construct initial prompts to formulate answer candidates, then the final answer is selected among options in answer heuristics via OQC prompting. Compared with previous LLM-based VQA methods, the proposed architecture is simpler and it brings a higher efficiency to predictions in zero-shot VQA. Extensive experiments demonstrate that Diff-ZsVQA with OQC prompt achieves competitive performance with higher inference speed than most existing methods. Quanxing Xu, Yuhao Tian, Ling Zhou 0005, Feifei Zhang 0001, Rubing Huang |
Expert Syst. Appl. | 5 |
| 2025 | Text-Video Knowledge Guided Prompting for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to localize action instances with only video-level labels for supervision. Recent methods convert category labels to natural language through prompting and utilize pre-trained vision-language models to generate text representation from natural language for supervision. This is because natural language can provide more prosperous and generalized semantic supervision to compensate for the lack of supervision in weakly supervised scenarios. However, it should be noted that current prompting methods face limitations in generating dynamic prompts that adapt to each video, which leads to difficulties in accurately aligning text and video representations. In this work, we propose a novel Text-Video Knowledge Guided Prompting (TVKP) framework for WTAL, which generates video-aware prompts based on text-video knowledge to enhance semantic alignment between text and video representations and introduce more video-related external category labels to enrich semantic supervision. We introduce the video-aware prompting (VAP) module to learn text-video knowledge from the joint distribution of text and video representations to generate video-aware text representation. Meanwhile, to make VAP more effectively learn text-video knowledge, a text-video contrastive loss is proposed to ensure semantic consistency between text and video representations. In addition, we propose the external knowledge prompting (EKP) module to introduce more video-related text labels from an external knowledge base to enrich prompts for accurate semantic alignment. Extensive experiments are conducted on three public datasets, THUMOS14, ActivityNet1.2, and ActivityNet1.3, demonstrating that our approach outperforms state-of-the-art methods. Yuxiang Shao, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Dual Uncertainty-Aware Correspondence Adapting and Retaining for Continual Composed Image RetrievalabstractRecent research in continual learning has primarily focused on unimodal tasks, with limited attention to multimodal tasks such as Composed Image Retrieval (CIR). In this paper, we establish a novel Continual CIR setting named C2IR to simulate the ever-change retrieval demands in the real world. Using the C2IR setting, we identify two significant challenges: intra-task correspondence uncertainty, which hinders the model's ability to manage noisy query-target pair correspondences; and inter-task drift uncertainty, which impedes the model's consistent understanding of relationships, exacerbating catastrophic forgetting across continual tasks. To address these challenges, we propose a Dual Uncertainty-aware Correspondence Adapting and Retaining (U2CAR) framework for C2IR, which leverages uncertainty learning to acquire and consolidate composed correspondence. To ensure reliable composed correspondence inference in each task, we introduce an Uncertainty-based Correspondence Reasoning (UCR) module that estimates and refines the uncertainty in query-target correspondence. Besides, to mitigate catastrophic forgetting of previous tasks, we design an Uncertainty-guided Re-parameterization (URep) paradigm that consolidates valuable composed correspondence knowledge based on the uncertainty variance across various tasks. Extensive experimental results illustrate that our U2CAR significantly outperforms existing methods, demonstrating the robust adaptability and anti-forgetting capabilities of the proposed approach. Haoliang Zhou, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2025 | CACP: Covariance-Aware Cross-Domain Prototypes for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to reduce domain shifts / discrepancies between source and target domains, improving the source domain model's generalization ability to the target domain. Recently, prototypical methods, which primarily use single-source or single-target domain prototypes as category centers to aggregate features from both domains, have achieved competitive performance in this task. However, due to large domain shifts, single-source domain prototypes have finite generalization ability and not all source domain knowledge is conducive to model generalization. Single-target domain prototypes are noisy because they are prematurely initialized with all features filtered by pseudo labels, which causes error accumulation in the prototypes. To address these issues, we propose a covariance-aware cross-domain prototypes method (CACP) to achieve robust domain adaptation. We propose to use both domain prototypes to dynamically rectify pseudo labels in the target domain, effectively reducing the recognition difficulty of hard target domain samples and narrowing the gap between features of the same category in both domains. In addition, to further generalize the model to the target domain, we propose two modules based on covariance correlation, FSPC (Features Selection by Prototypes Covariances) and WSPC (Weighting Source by Prototypes Coefficients), to learn discriminative characteristics. FSPC selects highly correlated features to update target domain prototypes online, denoising and enhancing discriminativeness between categories. WSPC utilizes the correlation coefficients between target domain prototypes and source domain features to weight each point in the source domain, eliminating the information interference from the source domain. In particular, CACP achieves excellent performance on the GTA5$\to$Cityscapes and SYNTHIA$\to$Cityscapes tasks with minimal computational resources and time. Yanbing Xue, Feifei Zhang 0001, Xianbin Wen, Zan Gao 0002, Shengyong Chen |
IEEE Trans. Multim. | 3 |
| 2024 | Overcoming the Pitfalls of Vision-Language Model for Image-Text Retrieval
Feifei Zhang 0001, Sijia Qu, Fan Shi 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2024 | Multi-modal Knowledge-Enhanced Fine-Grained Image Classification
Suyan Cheng, Feifei Zhang 0001, Haoliang Zhou, Changsheng Xu |
PRCV (5) | 2 |
| 2024 | BIVL-Net: Bidirectional Vision-Language Guidance for Visual Question Answering
Feifei Zhang 0001 |
PRCV (3) | 2 |
| 2024 | Scene-text aware cross-modal retrieval based on semantic matching (ChinaMM2024)
Suyan Cheng, Feifei Zhang 0001 |
Multim. Syst. | 2 |
| 2024 | NExT-OOD: Overcoming Dual Multiple-Choice VQA BiasesabstractIn recent years, multiple-choice Visual Question Answering (VQA) has become topical and achieved remarkable progress. However, most pioneer multiple-choice VQA models are heavily driven by statistical correlations in datasets, which cannot perform well on multimodal understanding and suffer from poor generalization. In this paper, we identify two kinds of spurious correlations, i.e., a Vision-Answer bias (VA bias) and a Question-Answer bias (QA bias). To systematically and scientifically study these biases, we construct a new video question answering (videoQA) benchmark NExT-OOD in OOD setting and propose a graph-based cross-sample method for bias reduction. Specifically, the NExT-OOD is designed to quantify models' generalizability and measure their reasoning ability comprehensively. It contains three sub-datasets including NExT-OOD-VA, NExT-OOD-QA, and NExT-OOD-VQA, which are designed for the VA bias, QA bias, and VA&QA bias, respectively. We evaluate several existing multiple-choice VQA models on our NExT-OOD, and illustrate that their performance degrades significantly compared with the results obtained on the original multiple-choice VQA dataset. Besides, to mitigate the VA bias and QA bias, we explicitly consider the cross-sample information and design a contrastive graph matching loss in our approach, which provides adequate debiasing guidance from the perspective of whole dataset, and encourages the model to focus on multimodal contents instead of spurious statistical regularities. Extensive experimental results illustrate that our method significantly outperforms other bias reduction strategies, demonstrating the effectiveness and generalizability of the proposed approach. Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | CEPrompt: Cross-Modal Emotion-Aware Prompting for Facial Expression RecognitionabstractFacial expression recognition (FER) remains a challenging task due to the ambiguity and subtlety of expressions. To address this challenge, current FER methods predominantly prioritize visual cues while inadvertently neglecting the potential insights that can be gleaned from other modalities. Recently, vision-language pre-training (VLP) models integrated textual cues as guidance, culminating in a powerful multi-modal solution that has proven effective for a range of computer vision tasks. In this paper, we propose a Cross-Modal Emotion-Aware Prompting (CEPrompt) framework for FER based on VLP models. To make VLP models sensitive to expression-relevant visual discrepancies, we devise an Emotion Conception-guided Visual Adapter (EVA) to capture the category-specific appearance representations with emotion conception guidance. Moreover, knowledge distillation is employed to prevent the model from forgetting the pre-trained category-invariant knowledge. In addition, we design a Conception-Appearance Tuner (CAT) to facilitate the interaction of multi-modal information via cooperatively tuning between emotion conception and appearance prompts. In this way, semantic information about emotion text conception is infused directly into facial appearance images, thereby enhancing a comprehensive and precise understanding of expression-related facial details. Quantitative and qualitative experiments show that our CEPrompt outperforms state-of-the-art approaches on three real-world FER datasets. The code is available athttps://github.com/HaoliangZhou/CEPrompt. Haoliang Zhou, Shucheng Huang, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Snippet-to-Prototype Contrastive Consensus Network for Weakly Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize action instances from untrimmed videos with only video-level labels. Due to the lack of frame-wise annotations, most methods embrace a localization-by-classification paradigm. However, the large supervision gap between classification and localization hinders models from obtaining accurate snippet-wise classification sequences and action proposals. We propose a snippet-to-prototype contrastive consensus network (SPCC-Net) to simultaneously generate feature-level and label-level supervision information to narrow the supervision gap between classification and localization. Specifically, the network adopts a two-stream framework incorporating the optical flow and fusion streams to fully leverage the motion and complementary information from multiple modalities. Firstly, the snippet-to-prototype contrast module is executed within each stream to learn prototypes for all categories and contrast them with action snippets to guarantee intra-class compactness and inter-class separability of snippet features. Secondly, for generating accurate label-level supervision information through complementary information of multimodal features, the multi-modality consensus module ensures not only category consistency through knowledge distillation but also semantic consistency through contrastive learning. Finally, we introduce the auxiliary multiple instance learning (MIL) loss to alleviate the issue that existing MIL-based methods only localize sparse discriminative snippets. Extensive experiments are conducted on two public datasets, THUMOS-14 and ActivityNet-1.3, to demonstrate the superior performance of our method over state-of-the-art methods. Yuxiang Shao, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | VQACL: A Novel Visual Question Answering Continual Learning SettingabstractResearch on continual learning has recently led to a variety of work in unimodal community, however little attention has been paid to multimodal tasks like visual question answering (VQA). In this paper, we establish a novel VQA Continual Learning setting named VQACL, which contains two key components: a dual-level task sequence where visual and linguistic data are nested, and a novel composition testing containing new skill-concept combinations. The former devotes to simulating the ever-changing multimodal datastream in real world and the latter aims at measuring models' generalizability for cognitive reasoning. Based on our VQACL, we perform in-depth evaluations of five well-established continual learning methods, and observe that they suffer from catastrophic forgetting and have weak generalizability. To address above issues, we propose a novel representation learning method, which leverages a sample-specific and a sample-invariant feature to learn representations that are both discriminative and generalizable for VQA. Furthermore, by respectively extracting such representation for visual and textual input, our method can explicitly disentangle the skill and concept. Extensive experimental results illustrate that our method significantly outperforms existing models, demonstrating the effectiveness and compositionality of the proposed approach. The code is available at https://github.com/zhangxi1997/VQACL. Feifei Zhang 0001, Changsheng Xu |
CVPR | 2 |
| 2023 | Reducing Vision-Answer Biases for Multiple-Choice VQAabstractMultiple-choice visual question answering (VQA) is a challenging task due to the requirement of thorough multimodal understanding and complicated inter-modality relationship reasoning. To solve the challenge, previous approaches usually resort to different multimodal interaction modules. Despite their effectiveness, we find that existing methods may exploit a new discovered bias (vision-answer bias) to make answer prediction, leading to suboptimal VQA performances and poor generalization. To solve the issues, we propose a Causality-based Multimodal Interaction Enhancement (CMIE) method, which is model-agnostic and can be seamlessly incorporated into a wide range of VQA approaches in a plug-and-play manner. Specifically, our CMIE contains two key components: a causal intervention module and a counterfactual interaction learning module. The former devotes to removing the spurious correlation between the visual content and the answer caused by the vision-answer bias, and the latter helps capture discriminative inter-modality relationships by directly supervising multimodal interaction training via an interactive loss. Extensive experimental results on three public benchmarks and one reorganized dataset show that the proposed method can significantly improve seven representative VQA models, demonstrating the effectiveness and generalizability of the CMIE. Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2022 | Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalabstractComposed Query Based Image Retrieval (CQBIR) aims at searching images relevant to a composed query, i.e., a reference image together with a modifier text. Compared with conventional image retrieval, which takes a single image or text to retrieve desired images, CQBIR encounters more challenges as it requires not only effective semantic correspondence between the heterogeneous query and target, but also synergistic understanding of the composed query. To establish robust CQBIR model, four critical types of relational information can be included, i.e., cross-modal, intra-sample, inter-sample, and cross-sample relationships. Pioneer studies mainly exploit parts of the information, which are hard to make them enhance and complement each other. In this paper, we propose a comprehensive relationship reasoning network by fully exploring the four types of information for CQBIR, which mainly includes two key designs. First, we introduce a memory-augmented cross-modal attention module, in which the representation of the composed query is augmented by considering the cross-modal relationship between the reference image and the modification text. Second, we design a multi-scale matching strategy to optimize our network, aiming at harnessing information from the intra-sample, inter-sample, and cross-sample relationships. To the best of our knowledge, this is the first work to fully explore the four pieces of relationships in a unified deep model for CQBIR. Comprehensive experimental results on five standard benchmarks demonstrate that the proposed method performs favorably against state-of-the-art models. Feifei Zhang 0001, Ming Yan 0008, Ji Zhang 0011, Changsheng Xu |
ACM Multimedia | 1 |
| 2022 | Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
Ling Zhou 0005, Qirong Mao, Xiaohua Huang 0002, Feifei Zhang 0001, Zhihong Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | Joint Expression Synthesis and Representation Learning for Facial Expression RecognitionabstractFacial expression recognition (FER) is a challenging task due to the large appearance variations and the lack of sufficient training data. Conventional deep approaches either learn a good representation through deep models or synthesize images automatically to enlarge the training set. In this paper, we perform both tasks jointly and propose an end-to-end deep model for simultaneous facial expression recognition and facial image synthesis. The proposed model is based on Generative Adversarial Network (GAN) and enjoys several merits. First, the facial image synthesis and facial expression recognition tasks can boost their performance for each other via the unified model. Second, paired images are not required in our facial image synthesis network, which makes the proposed model much more general and flexible. Meanwhile, the generated facial images largely expand the training set and ease the overfitting problem in our FER task. Third, different expressions are encoded in a disentangled manner in a latent space, which enables us to synthesize facial images with arbitrary expressions by exchanging certain parts of their latent identity features. Quantitative and qualitative evaluations on both controlled and in-the-wild FER benchmarks (Multi-PIE, MMI, and RAF-DB) demonstrate the effectiveness of our proposed method on both facial image synthesis and facial expression recognition task. Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Geometry Sensitive Cross-Modal Reasoning for Composed Query Based Image RetrievalabstractComposed Query Based Image Retrieval (CQBIR) aims at retrieving images relevant to a composed query containing a reference image with a requested modification expressed via a textual sentence. Compared with the conventional image retrieval which takes one modality as query to retrieve relevant data of another modality, CQBIR poses great challenge over the semantic gap between the reference image and modification text in the composed query. To solve the challenge, previous methods either resort to feature composition that cannot model interactions in the query or explore inter-modal attention while ignoring the spatial structure and visual-semantic relationship. In this paper, we propose a geometry sensitive cross-modal reasoning network for CQBIR by jointly modeling the geometric information of the image and the visual-semantic relationship between the reference image and modification text in the query. Specifically, it contains two key components: a geometry sensitive inter-modal attention module (GS-IMA) and a text-guided visual reasoning module (TG-VR). The GS-IMA introduces the spatial structure into the inter-modal attention in both implicit and explicit manners. The TG-VR models the unequal semantics not included in the reference image to guide further visual reasoning. As a result, our method can learn effective feature for the composed query which does not exhibit literal alignment. Comprehensive experimental results on three standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art methods. Feifei Zhang 0001, Mingliang Xu 0001, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2022 | Weakly-Supervised Facial Expression Recognition in the Wild With Noisy DataabstractFacial expression recognition (FER) has attracted much attention in recent years due to its wide applications. While some progress has been achieved thanks to the emergence of deep learning, the challenge occasioned by pose variations remains. Therefore, most conventional approaches mainly perform FER under laboratory-controlled environment, and the FER in-the-wild has received relatively less attention. To implement the FER in-the-wild, the pose-invariant expression recognition model would be a possible solution but for a paucity of training data. Sufficient training data with reliable expression labels on FER tasks typically are unavailable. This paper devotes to addressing the problem of how to model pose variations in facial images, and how to leverage noisy data in the web to boost the FER performance. The proposed model is implemented in an end-to-end weakly supervised manner and enjoys several merits. First, the proposed model utilizes massive noisy labeled data to boost the performance of the FER classifier trained on a small set of clean labels. Second, we offer a novel pose modeling network to adaptively capture the discrepancy in the deep representation space of facial images under different head poses, and consequently, the pose-invariant expression representations can be learned in our model. Last, to exploit the reliable information in the noisy data, we formulate a noise modeling network, which is capable of learning the mapping from feature space to the residuals between clean labels and noisy labels. We validate the proposed approach on four public FER benchmarks: AffectNet, RAF-DB, SFEW, and BU-3DFE. Extensive experiments show that the proposed method performs favorably against state-of-the-art methods. Feifei Zhang 0001, Mingliang Xu 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2022 | Explicit Cross-Modal Representation Learning for Visual Commonsense ReasoningabstractGiven a question about an image, Visual Commonsense Reasoning (VCR) needs to provide not only a correct answer, but also a rationale to justify the answer. VCR is a challenging task due to the requirement of proper semantic alignment and reasoning between the image and linguistic expression. Recent approaches offer a great promise by exploring holistic attention mechanisms or graph-based networks, but most of them do implicit reasoning and ignore the semantic dependencies among the linguistic expression. In this paper, we propose a novel explicit cross-modal representation learning network for VCR by incorporating syntactic information into the visual reasoning and natural language understanding. The proposed method enjoys several merits. First, based on a two-branch neural module network, we can do explicit crossmodal reasoning guided by the high-level syntactic structure of linguistic expression. Second, the semantic structure of the linguistic expression is incorporated into a syntactic GCN to facilitate language understanding. Third, our explicit crossmodal representation learning network can provide a traceable reasoning-flow, which offers visible fine-grained evidence of the answer and rationale. Quantitative and qualitative evaluations on the public VCR dataset demonstrate that our approach performs favorably against state-of-the-art methods. The full code for our work is available in the supplementary material. Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2022 | Tell, Imagine, and Search: End-to-end Learning for Composing Text and Image to Image RetrievalabstractComposing Text and Image to Image Retrieval ( CTI-IR ) is an emerging task in computer vision, which allows retrieving images relevant to a query image with text describing desired modifications to the query image. Most conventional cross-modal retrieval approaches usually take one modality data as the query to retrieve relevant data of another modality. Different from the existing methods, in this article, we propose an end-to-end trainable network for simultaneous image generation and CTI-IR . The proposed model is based on Generative Adversarial Network (GAN) and enjoys several merits. First, it can learn a generative and discriminative feature for the query (a query image with text description) by jointly training a generative model and a retrieval model. Second, our model can automatically manipulate the visual features of the reference image in terms of the text description by the adversarial learning between the synthesized image and target image. Third, global-local collaborative discriminators and attention-based generators are exploited, allowing our approach to focus on both the global and local differences between the query image and the target image. As a result, the semantic consistency and fine-grained details of the generated images can be better enhanced in our model. The generated image can also be used to interpret and empower our retrieval model. Quantitative and qualitative evaluations on three benchmark datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Feifei Zhang 0001, Mingliang Xu 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Multi-Level Counterfactual Contrast for Visual Commonsense ReasoningabstractGiven a question about an image, a Visual Commonsense Reasoning (VCR) model needs to provide not only a correct answer, but also a rationale to justify the answer. It is a challenging task due to the requirements of diverse visual content understanding, abstract language comprehending, and complicated inter-modality relationship reasoning. To solve above challenges, previous methods either resort to holistic attention mechanism or explore transformer-based model with pre-training, which, however, cannot perform comprehensive understanding and usually suffer from heavy computing burden. In this paper, we propose a novel multi-level counterfactual contrastive learning network for VCR by jointly modeling the hierarchical visual contents and the inter-modality relationships between the visual and linguistic domains. The proposed method enjoys several merits. First, with sufficient instance-level, image-level, and semantic-level contrastive learning, our model can extract discriminative features and perform comprehensive understanding for the image and linguistic expressions. Second, taking advantage of counterfactual thinking, we can generate informative factual and counterfactual samples for contrastive learning, resulting in stronger perception ability of our model. Third, an auxiliary contrast module is incorporated into our method to directly optimize the answer prediction in VCR, which further facilitates the representation learning. Extensive experiments on the VCR dataset demonstrate that our approach performs favorably against the state-of-the-arts. Feifei Zhang 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2020 | Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalabstractCross-model retrieval has attracted much attention in recent years due to its wide applications. Conventional approaches usually take one modality as query to retrieve relevant data of another modality. In this paper, we devote to an emerging task in cross-modal retrieval, Composing Text and Image to Image Retrieval (CTI-IR), which aims at retrieving images relevant to a query image with text describing desired modifications to the query image. Compared with conventional cross-modal retrieval, the new task is particularly useful for the retrieval that the query image does not perfectly match the user's expectations. Generally, the CTI-IR involves two underlying problems: how to manipulate visual features of the query image specified by the text, and how to model the modality gap between the query and target. Most previous methods focus on solving the second problem. In this paper, we aim to deal with both problems simultaneously in a unified model. Specifically, the proposed method is based on the graph attention network and adversarial learning network, which enjoys several merits. First, the query image and the modification text are constructed in a relation graph for learning text-adaptive representations. Second, semantic contents from the text are injected into the visual features through graph attention. Third, an adversarial loss is incorporated into the conventional cross-modal retrieval loss to learn more discriminative modality invariant representations for CTI-IR. Extensive experiments on three benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods. Feifei Zhang 0001, Mingliang Xu 0001, Qirong Mao, Changsheng Xu |
ACM Multimedia | 1 |
| 2020 | Geometry Guided Pose-Invariant Facial Expression RecognitionabstractDriven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. However, most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifier for each pose. Different from existing methods, this paper proposes an end-to-end deep learning model that allows to simultaneous facial image synthesis and pose-invariant facial expression recognition by exploiting shape geometry of the face image. The proposed model is based on generative adversarial network (GAN) and enjoys several merits. First, given an input face and a target pose and expression designated by a set of facial landmarks, an identity-preserving face can be generated through guiding by the target pose and expression. Second, the identity representation is explicitly disentangled from both expression and pose variations through the shape geometry delivered by facial landmarks. Third, our model can automatically generate face images with different expressions and poses in a continuous way to enlarge and enrich the training set for the FER task. Our approach is demonstrated to perform well when compared with state-of-the-art algorithms on both controlled and in-the-wild benchmark datasets including Multi-PIE, BU-3DFE, and SFEW. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2020 | A Unified Deep Model for Joint Facial Expression Recognition, Face Synthesis, and Face AlignmentabstractFacial expression recognition, face synthesis, and face alignment are three coherently related tasks and can be solved in a joint framework. To achieve this goal, in this paper, we propose a novel end-to-end deep learning model by exploiting the expression code, geometry code and generated data jointly for simultaneous pose-invariant facial expression recognition, face image synthesis, and face alignment. The proposed deep model enjoys several merits. First, to the best of our knowledge, this is the first work to address these three tasks jointly in a unified deep model to complement and enhance each other. Second, the proposed model can effectively disentangle the global and local identity representation from different expression and geometry codes. As a result, it can automatically generate facial images with different expressions under arbitrary geometry codes. Third, these three tasks can further boost their performance for each other via our model. Extensive experimental results on three standard benchmarks demonstrate that the proposed deep model performs favorably against state-of-the-art methods on the three tasks. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2019 | Unpaired Images based Generator Architecture for Facial Expression RecognitionabstractFacial expression recognition (FER) is a challenging task due to the lack of sufficient training data. Most conventional approaches usually rotate or flip the images for data augmentation. More recently, numerous methods synthesize images automatically by using Generative Adversarial Network (GAN). However, paired images are always required in these methods. Different from existing methods, in this paper, we propose an end-to-end deep learning model for simultaneous facial expression synthesis and facial expression recognition. In our method, paired images are not required, which makes the proposed model much more flexible and general. Furthermore, different expressions are encoded in a disentangled manner in a latent space, which enables us to generate facial images with arbitrary expressions by exchanging certain parts of their latent identity features. Finally, the facial expression synthesis and facial expression recognition tasks can further boost their performance for each other via our model. Quantitative and qualitative evaluations on both controlled and in-the-wild datasets demonstrate that the proposed method performs favorably against state-of-the-art methods. Feifei Zhang 0001, Changsheng Xu |
VCIP | 2 |
| 2018 | Joint Pose and Expression Modeling for Facial Expression RecognitionabstractFacial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we propose an end-to-end deep learning model by exploiting different poses and expressions jointly for simultaneous facial image synthesis and pose-invariant facial expression recognition. The proposed model is based on generative adversarial network (GAN) and enjoys several merits. First, the encoder-decoder structure of the generator can learn a generative and discriminative identity representation for face images. Second, the identity representation is explicitly disentangled from both expression and pose variations through the expression and pose codes. Third, our model can automatically generate face images with different expressions under arbitrary poses to enlarge and enrich the training set for FER. Quantitative and qualitative evaluations on both controlled and in-the-wild datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
CVPR | 1 |
| 2018 | Facial Expression Recognition in the Wild: A Cycle-Consistent Adversarial Attention Transfer ApproachabstractFacial expression recognition (FER) is a very challenging problem due to different expressions under arbitrary poses. Most conventional approaches mainly perform FER under laboratory controlled environment. Different from existing methods, in this paper, we formulate the FER in the wild as a domain adaptation problem, and propose a novel auxiliary domain guided Cycle-consistent adversarial Attention Transfer model (CycleAT) for simultaneous facial image synthesis and facial expression recognition in the wild. The proposed model utilizes large-scale unlabeled web facial images as an auxiliary domain to reduce the gap between source domain and target domain based on generative adversarial networks (GAN) embedded with an effective attention transfer module, which enjoys several merits. First, the GAN-based method can automatically generate labeled facial images in the wild through harnessing information from labeled facial images in source domain and unlabeled web facial images in auxiliary domain. Second, the class-discriminative spatial attention maps from the classifier in source domain are leveraged to boost the performance of the classifier in target domain. Third, it can effectively preserve the structural consistency of local pixels and global attributes in the synthesized facial images through pixel cycle-consistency and discriminative loss. Quantitative and qualitative evaluations on two challenging in-the-wild datasets demonstrate that the proposed model performs favorably against state-of-the-art methods. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Ling-Yu Duan, Changsheng Xu |
ACM Multimedia | 1 |
| 2018 | Cascaded Multi-level Transformed Dirichlet Process for Multi-pose Facial Expression RecognitionabstractAs an essential way of human emotional behavior understanding, facial expression recognition (FER) has been studied extensively in recent years. However, the existing methods of FER are typically based on near-frontal face data. High-recognition accuracy for multi-pose FER continues to be a challenge. In this paper, we present a novel cascaded multi-level Transformed Dirichlet Process (cml-TDP) model for multi-pose FER. The top-level structure of the cml-TDP model has been carefully designed to make coarse-to-fine prediction, and the outputs of the model are fused for robust and accurate estimation at each level. There are three primary merits to cml-TDP. First, pose is explicitly introduced into cml-TDP so that separate training and parameter tuning for each pose is not required. Second, cml-TDP describes an image by its detected positions and appearance features to implicitly construct geometric constraints. Third, cml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. The proposed model has been evaluated on two benchmark databases, BU-3DFE and RAFD, and achieved 79.33% and 75.00% FER accuracy on these two datasets, respectively, which has outperformed current state-of-the-art FER methods. Qirong Mao, Feifei Zhang 0001, Liangjun Wang, Sidian Luo, Ming Dong 0001 |
Comput. J. | 2 |
| 2018 | Affective rating ranking based on face images in arousal-valence dimensional spaceabstractIn dimensional affect recognition, the machine learning methods, which are used to model and predict affect, are mostly classification and regression. However, the annotation in the dimensional affect space usually takes the form of a continuous real value which has an ordinal property. The aforementioned methods do not focus on taking advantage of this important information. Therefore, we propose an affective rating ranking framework for affect recognition based on face images in the valence and arousal dimensional space. Our approach can appropriately use the ordinal information among affective ratings which are generated by discretizing continuous annotations. Specifically, we first train a series of basic cost-sensitive binary classifiers, each of which uses all samples relabeled according to the comparison results between corresponding ratings and a given rank of a binary classifier. We obtain the final affective ratings by aggregating the outputs of binary classifiers. By comparing the experimental results with the baseline and deep learning based classification and regression methods on the benchmarking database of the AVEC 2015 Challenge and the selected subset of SEMAINE database, we find that our ordinal ranking method is effective in both arousal and valence dimensions. Guopeng Xu, Haitang Lu, Feifei Zhang 0001, Qirong Mao |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | Spatially Coherent Feature Learning for Pose-Invariant Facial Expression RecognitionabstractFeature learning has enjoyed much attention and achieved good performance in recent studies of image processing. Unlike the required training conditions often assumed there, far less labeled data is available for training emotion classification systems. In addition, current feature learning is typically performed on an entire face image without considering the dependency between features. These approaches ignore the fact that faces are structured and the neighboring features are dependent. Thus, the learned features lack the power to describe visually coherent facial images. Our method is therefore designed with the goal of simplifying the problem domain by removing expression-irrelevant factors from the input images, with a key region-based mechanism, which is an effort to reduce the amount of data required to effectively train the feature-learning methods. Meanwhile, we can construct geometric constraints between the key regions and its detected positions. To this end, we introduce a Spatially Coherent featurelearning method for Pose-invariant Facial Expression Recognition (SC-PFER). In our model, we first perform face frontalization through a 3D pose-normalization technique, which could normalize poses while preserving the identity information through synthesizing frontal faces for facial images with arbitrary views. Subsequently, we select a sequence of key regions around 51 key points in the synthetic frontal face images for efficient unsupervised feature learning. Finally, we introduce a linkage structure over the learning-based features and the corresponding geometry information of each key region to encode the dependencies of the regions. Our method, on the whole, does not require training multiple models for each specific pose and avoids separating training and parameter tuning for each pose. The proposed framework has been evaluated on two benchmark databases, BU-3DFE and SFEW, for pose-invariant Facial Expression Recognition (FER). The experimental results demonstrate that our algorithm outperforms current state-of-the-art FER methods. Specifically, our model achieves an improvement of 1.72% and 1.11% FER accuracy, on average, on BU-3DFE and SFEW, respectively. Feifei Zhang 0001, Qirong Mao, Xiangjun Shen, Yongzhao Zhan 0001, Ming Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2016 | Domain adaptation for speech emotion recognition by sharing priors between related source and target classesabstractIn speech emotion recognition (SER), speech data is usually captured from different scenarios, which often leads to significant performance degradation due to the inherent mismatch between training and test set. To cope with this problem, we propose a domain adaptation method called Sharing Priors between Related Source and Target classes (SPRST) based on a two-layer neural network. The classifier parameters, namely the weights of the second layer, are imposed the common priors between the related classes, so that the classes with few labeled data in target domain can borrow knowledge from the related classes in source domain. The method is evaluated on the INTERSPEECH 2009 Emotion Challenge two-class task. Experimental results show that our approach significantly improves the performance when only a small number of target labeled instances are available. Qirong Mao, Wentao Xue, Qiyu Rao, Feifei Zhang 0001, Yongzhao Zhan 0001 |
ICASSP | 4 |
| 2016 | Multi-pose Facial Expression Recognition Using Transformed Dirichlet ProcessabstractDriven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. In this paper, we propose a novel graphical model, multi-level Transformed Dirichlet Process (ml-TDP), for multi-pose FER. In our approach, pose is explicitly introduced into ml-TDP so that separate training and parameter tuning for each pose is not required. In addition, ml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially-coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. Extensive experimental result on benchmark facial expression databases shows the superior performance of ml-TDP. Feifei Zhang 0001, Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001 |
ACM Multimedia | 1 |
| 2016 | Pose-robust feature learning for facial expression recognition
Feifei Zhang 0001, Qirong Mao, Jianping Gou, Yongzhao Zhan 0001 |
Frontiers Comput. Sci. | 1 |