Fan Qi

dblp:228/1390 · DBLP profile ↗
← Back
27ranked-venue papers
15as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 13 first-author · 18 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Multi-Instance Generation With Janus-Pro-Driven Prompt Parsing
abstract
Recent advances in text-guided diffusion models have revolutionized conditional image generation, yet they struggle to synthesize complex scenes with multiple objects due to imprecise spatial grounding and limited scalability. We address these challenges through two key modules: 1) Janus-Pro-driven Prompt Parsing, a prompt-layout parsing module that bridges text understanding and layout generation via a compact 1B-parameter architecture, and 2) MIGLoRA, a parameter-efficient plug-in integrating Low-Rank Adaptation (LoRA) into UNet (SD1.5) and DiT (SD3) backbones. MIGLoRA is capable of preserving the base model’s parameters and ensuring plug-and-play adaptability, minimizing architectural intrusion while enabling efficient fine-tuning. To support a comprehensive evaluation, we create DescripBox and DescripBox-1024, benchmarks that span diverse scenes and resolutions. The proposed method achieves state-of-the-art performance on COCO and LVIS benchmarks while maintaining parameter efficiency, demonstrating superior layout fidelity and scalability for open-world synthesis. Here is our project page: https://github.com/FanQi-AI/MIGLoRA.
Fan Qi, Mengxian Li, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2026 Fed-DiffLoRA: Personalized Federated Style Transfer for T2I Diffusion Models
abstract
Low-Rank Adaptation (LoRA) merging enables efficient customization of T2I diffusion models; nevertheless, centralized aggregation raises serious privacy concerns. While federated LoRA adaptations mitigate these risks, diffusion models present unique challenges: structural heterogeneity and vulnerability to member inference attacks. To overcome these limitations, we propose Fed-DiffLoRA, a privacy-preserving framework that securely aggregates LoRA adapters. At the client level, we disentangle user-specific adaptations into two orthogonal subspaces: content LoRAs preserving semantic fidelity and style LoRAs encoding stylistic features, isolating sensitive attributes from stylistic components. We design a learnable aggregation operator that dynamically optimizes cross-client style LoRA fusion based on the semantic vectors of clients' content LoRAs, achieving high-fidelity style blending while suppressing client-identifiable patterns. We also provide theoretical guarantees for convergence and privacy. Extensive experiments validate the effectiveness of the proposed approach, achieving substantial reductions in attack success rates and consistently high stylization fidelity.
Fan Qi, Xiaoshan Yang, Changsheng Xu
IEEE Trans. Image Process.1
2025 Customized Condition Controllable Generation for Video Soundtrack
abstract
Recent advances in latent diffusion models (LDMs) have enabled data-driven paradigms for video soundtrack generation, improving multimodal alignment capabilities. However, current two-stage frameworks—which separately optimize audio-visual correspondence and conditional audio synthesis—fundamentally limit joint modeling of dynamic acoustic properties. In this paper, we propose a novel framework for generating video soundtracks that simultaneously produces music and sound effect tailored to the video content. Our method incorporates a Contrastive Visual-Sound-Music pretraining process that maps these modalities into a unified feature space, enhancing the model’s ability to capture intricate audio dynamics. We design Spectrum Divergence Masked Attention for Unet to differentiate between the unique characteristics of sound effect and music. We utilize Score-guided Noise Iterative Optimization to provide musicians with customizable control during the generation process. Extensive evaluations on the FilmScoreDB and SymMV&HIMV datasets demonstrate that our approach significantly outperforms state-of-the-art baselines in both subjective and objective assessments, highlighting its potential as a robust tool for video soundtrack generation.
Fan Qi, Kunsheng Ma, Changsheng Xu
CVPR1
2025 Rethinking the Temperature for Federated Heterogeneous Distillation
abstract
Federated Distillation (FedKD) relies on lightweight knowledge carriers like logits for efficient client-server communication. Although logit-based methods have demonstrated promise in addressing statistical and architectural heterogeneity in federated learning (FL), current approaches remain constrained by suboptimal temperature calibration during knowledge fusion. To address these limitations, we propose ReT-FHD, a framework featuring: 1) Multi-level Elastic Temperature, which dynamically adjusts distillation intensities across model layers, achieving optimized knowledge transfer between heterogeneous local models; 2) Category-Aware Global Temperature Scaling that implements class-specific temperature calibration based on confidence distributions in global logits, enabling personalized distillation policies; 3) Z-Score Guard, a blockchain-verified validation mechanism mitigating 44% of label-flipping and model poisoning attacks. Evaluations across diverse benchmarks with varying model/data heterogeneity demonstrate that the ReT-FHD achieves significant accuracy improvements over baseline methods while substantially reducing communication costs compared to existing approaches. Our work establishes that properly calibrated logits can serve as self-sufficient carriers for building scalable and secure heterogeneous FL systems.
Fan Qi, Daxu Shi, Chuokun Xu, Changsheng Xu
ICML1
2025 Granular Music Attribute Transformation with Proximal Policy Optimization Adapters for Diffusion Model
abstract
The rapid development of music diffusion models has provided diverse paths for music creation transformations. However, existing methods still lack continuous strength regulation over stylistic attributes-specifically, they cannot achieve scalable adjustment of intensity (e.g., smooth transitions between ''gentle'' and ''intense'' jazz) while preserving spectral-temporal coherence. To address this, we propose RLScale-LoRA, a two-stage finetuning framework built on a structurally modified low-rank adaptation (LoRA) architecture with scale layers. In Stage 1, we finetune the modified LoRA to specialize in capturing attribute-aware latent spaces on unseen/seen music data. Stage 2 trains lightweight scale layers via proximal policy optimization (PPO), where reward functions enforce intermediate spectral-temporal state stability. Therefore, our RLScale-LoRA achieves precise, continuous music attribute transformations. Extensive experiments on Mtg-Jamendo and MedleyMD-Prompts datasets demonstrate RLScale-LoRA's superiority in granularity and coherence.
Kunsheng Ma, Fan Qi, Changsheng Xu
ACM Multimedia2
2025 FORGET ME: Federated Unlearning for Face Generation Models
abstract
Federated face generation technology leverages decentralized private data to achieve high-quality face synthesis. However, regulations such as the GDPR confer users the right to be forgotten, necessitating the removal of contributions from specific clients in the global model. Existing generation model unlearning methods are primarily designed for centralized environments and are inadequate for addressing the constraints of data privacy storage and limited client computational resources in federated settings. To address this gap, we propose F2GU, the first federated unlearning framework specifically tailored for face generation models, enabling the effective removal of contributions associated with specific clients (identities) while ensuring privacy. Our proposed Generation Trajectory Redirection method dynamically guides the generation trajectory away from target identities, thereby effectively eliminating contributions from specific clients. Additionally, we devise a Mirroring-guided Trajectory Optimization strategy that constructs a mirror projection utilizing the retained client trajectory origins to ensure the generative capabilities of the model are preserved post-unlearning. We conduct extensive experiments on two mainstream face generation models (GAN and Diffusion Model) across three different datasets.The results indicate that our method demonstrates superior performance in both the success rate of identity unlearning and the preservation of generation quality. The code can be available at https://github.com/FanQi-AI/FFGU.
Fan Qi, Zixin Zhang 0004, Changsheng Xu
ACM Multimedia1
2025 Fine-tuning Bias Neurons for Fair Text-to-Image Generation
abstract
Diffusion Models (DMs) have revolutionized Text-to-Image (T2I) generation, yet inherent dataset biases often result in skewed representations across demographics, perpetuating stereotypes and social inequities. Existing debiasing approaches primarily focus on the text processing component, overlooking the intricate biases in the diffusion model's U-Net architecture. This paper presents a novel approach to addressing these biases through a causal analysis of bias disentanglement within the U-Net architecture. We introduce the Contrast Neuron Sensitivity Metric, which enables precise identification of neurons sensitive to bias, allowing for targeted interventions. Our debiasing paradigm fine-tunes these identified neurons with a combination of distribution and semantic loss, requiring only 0.2M parameters to be adjusted, which is far less than prior methods. Experiments show that our method effectively removes gender and race biases and maintains the diversity distribution of images. It enables both absolute fairness and relative adjustments by modifying target attribute distributions (e.g., young:old = 7:3). Furthermore, our approach is scalable, allowing simultaneous fine-tuning across multiple biases, and achieves good bias reduction even with non-templated prompts. The code is available on https://github.com/FanQi-AI/Debias.
Fan Qi, Changsheng Xu, Huaiwen Zhang
ACM Multimedia1
2025 One-shot Multimodal Federated Learning via Diverse Synthetic Feature Optimization
abstract
One-shot Federated Learning (FL) enables collaborative model training through a single round of communication, thereby significantly reducing communication overhead. However, in multimodal settings, client data heterogeneity and modality heterogeneity make it difficult to accurately learn the global multimodal representation within a single communication round. We propose Diverse Synthetic Feature Optimization (DSFO) for one-shot MFL framework that achieves global modality consensus via a single communication round. DSFO introduces: (1) Dynamic Probabilistic Scheduling Model Queue (DPSMQ), which ensures local multimodal feature-level distillation quality through diversity-aware model sampling from the queue, directly transmitting synthetic features can significantly reduces the risk of misleading knowledge transfer and (2) Pareto-Optimal Global Feature Consensus (POGFC), a server-side multi-objective optimization that extracts maximally representative synthetic features per modality-category pair to mitigate data heterogeneity, supported by theoretical guarantees in the appendix. Experiments on multimodal datasets demonstrate that DSFO outperforms current methods while significantly reducing communication overhead.
Fan Qi, Zixin Zhang 0004, Huaiwen Zhang
MMAsia1
2025 Scaling research aim identification: Language models for classifying scientific and societal-oriented studies
abstract
Abstract The classification of research according to its aims has been a longstanding focus in the fields of quantitative science studies and R&D statistics. Since 1963, the Organization for Economic Co‐operation and Development (OECD) has employed a classical distinction among basic, applied, and experimental research. Building on this framework, our previous work highlighted the utility of differentiating between scientific and societal progress as two primary research objectives. This distinction enabled the quantitative analysis of scientific publication abstracts and the development of an automated method for large‐scale classification. In the current study, we systematically evaluate text classification techniques, including traditional text mining models, classification tools, BERT‐based language models, and decoder‐only large language models (LLMs) such as ChatGPT. Our findings show that the fine‐tuned GPT‐4o‐mini model performs the best among single‐model approaches. However, traditional and BERT‐based models outperform in certain fine‐grained classification tasks. Leveraging majority voting strategies to incorporate their strengths yields performance comparable to closed‐source GPT models. A case study on 10 biomedical journals further validates the method, demonstrating strong alignment between journal scopes, model predictions, and outputs generated by the fine‐tuned GPT‐4o‐mini model. These results highlight the robustness and practical effectiveness of the proposed methodology for nuanced research aim classification.
Mengjia Wu, Gunnar Sivertsen, Lin Zhang 0004, Fan Qi, Yi Zhang 0095
J. Assoc. Inf. Sci. Technol.4
2025 Active Supervised Cross-Modal Retrieval
abstract
Supervised Cross-Modal Retrieval (SCMR) achieves significant performance with the supervision provided by substantial label annotations of multi-modal data. However, the requirement for large annotated multi-modal datasets restricts the use of supervised cross-modal retrieval in many practical scenarios. Active Learning (AL) has been proposed to reduce labeling costs while improving performance in various label-dependent tasks, in which the most informative unlabeled samples are selected for labeling and training. Directly exploiting the existing AL methods for supervised cross-modal retrieval may not be a good idea since they only focus on the uncertainty within each modality, ignoring the inter-modality relationship within the text-image pairs. Furthermore, existing methods focus exclusively on the informativeness of data during sample selection, leading to a biased, homogenized set where selected samples often contain nearly identical semantics and are densely distributed in a region of the feature space. Persistent training with such biased data selections can disturb multi-modal representation learning and substantially degrade the retrieval performance of SCMR. In this work, we propose an Active Supervised Cross-Modal Retrieval (ASCMR) framework, which effectively identifies informative multi-modal samples and generates unbiased sample selections. In particular, we propose a probabilistic multi-modal informativeness estimation that captures both the intra-modality and inter-modality uncertainty of multi-modal pairs within a unified representation. To ensure unbiased sample selection, we introduce a density-aware budget allocation strategy that constrains the active learning objective of maximizing the informativeness of selection with a novel semantic density regularization term. The proposed methods are evaluated on three widely used benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr, demonstrating our effectiveness in significantly reducing the annotation cost while outperforming other baselines of active learning strategies. We could achieve over 95% of the fully supervised model's performance by only utilizing 6%, 3%, and 4% active selected samples for MS-COCO, NUS-WIDE, and MIRFlickr, respectively.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Knowledge-Guided Label Distribution Calibration for Federated Affective Computing
abstract
The federated learning (FL) paradigm can significantly solve the rising public concern about data privacy in affective computing. However, conventional FL methods perform poorly due to the uniqueness of the task, as the personalized emotion data vary from client to client. To resolve the privacy-utility paradox, this work proposes a framework that largely improves federated affective computing (FAC) via calibrating the global feature space and communicating privacy-agnostic auxiliary information. The framework consists of two components: first, an emotion hemisphere (EH) representation structure is proposed, which utilizes emotional prior knowledge to unify the emotion global feature space of different clients. Second, the server uses the normalized parameter importance matrix to guide the model aggregation. It retains crucial parameters for individual local models, thereby alleviating the slow convergence problem in the global model caused by the skewed label distribution. The proposed framework yields significant performance gains, and extensive experiments on three emotion datasets demonstrate the effectiveness and the practicality of our approach.
Zixin Zhang 0004, Fan Qi, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.2
2024 SignGen: End-to-End Sign Language Video Generation with Latent Diffusion
Fan Qi, Huaiwen Zhang, Changsheng Xu
ECCV (53)1
2024 FedVAD: Enhancing Federated Video Anomaly Detection with GPT-Driven Semantic Distillation
Fan Qi, Ruijie Pan, Huaiwen Zhang, Changsheng Xu
ECCV (53)1
2024 Enhancing Storage and Computational Efficiency in Federated Multimodal Learning for Large-Scale Models
abstract
The remarkable generalization of large-scale models has recently gained significant attention in multimodal research. However, deploying heterogeneous large-scale models with different modalities under Federated Learning (FL) to protect data privacy imposes tremendous challenges on clients' limited computation and storage. In this work, we propose M$^2$FedSA to address the above issue. We realize modularized decomposition of large-scale models via Split Learning (SL) and only retain privacy-sensitive modules on clients, alleviating storage overhead. By freezing large-scale models and introducing two specialized lightweight adapters, the models can better focus on task-specific knowledge and enhance modality-specific knowledge, improving the model's adaptability to different tasks while balancing efficiency. In addition, M$^2$FedSA further improves performance by transferring multimodal knowledge to unimodal clients at both the feature and decision levels, which leverages the complementarity of different modalities. Extensive experiments on various multimodal classification tasks validate the effectiveness of our proposed M$^2$FedSA. The code is made available publicly at https://github.com/M2FedSA/M-2FedSA.
Zixin Zhang 0004, Fan Qi, Changsheng Xu
ICML2
2024 Cross-Modal Meta Consensus for Heterogeneous Federated Learning
abstract
In the evolving landscape of federated learning (FL), the integration of multimodal data presents both unprecedented opportunities and significant challenges. Existing works fall short of meeting the growing demand for systems that can efficiently handle diverse tasks and modalities in rapidly changing environments. We propose a meta learning strategy tailored for Multimodal Federated Learning in a multitask setting, which harmonizes intra-modal and inter-modal feature spaces through the Cross-Modal Meta Consensus. This innovative approach enables seamless integration and transfer of knowledge across different data types, enhancing task personalization within modalities and facilitating effective cross-modality knowledge sharing. Additionally, we introduce Gradient Consistency-based Clustering for multimodal convergence, specifically designed to resolve conflicts at meta-initialization points arising from diverse modality distributions, supported by theoretical guarantees. Our approach, evaluated as M3Fed on five federated datasets, with at most four modalities and four downstream tasks, demonstrates strong performance across diverse data distributions, affirming its effectiveness in Multimodal Federated Learning.
Fan Qi, Zixin Zhang 0004, Changsheng Xu
ACM Multimedia2
2024 T3RD: Test-Time Training for Rumor Detection on Social Media
abstract
With the increasing number of news uploaded to the internet daily, rumor detection has garnered significant attention in recent years. Existing rumor detection methods excel on familiar topics with sufficient training data (high resource) collected from the same domain. However, when facing emergent events or rumors propagated in different languages, the performance of these models is significantly degraded, due to the lack of training data and prior knowledge (low resource). To tackle this challenge, we introduce the Test-Time Training for Rumor Detection (T^3RD) to enhance the performance of rumor detection models on low-resource datasets. Specifically, we introduce self-supervised learning (SSL) as an auxiliary task in the test-time training. It consists of global and local contrastive learning, in which the global contrastive learning focuses on obtaining invariant graph representations and the local one focuses on acquiring invariant node representations. We employ the auxiliary SSL tasks for both the training and test-time training phases to mine the intrinsic traits of test samples and further calibrate the trained model for these test samples. To mitigate the risk of distribution distortion in test-time training, we introduce feature alignment constraints aimed at achieving a balanced synergy between the knowledge derived from the training set and the test samples. The experiments conducted on the two widely used cross-domain datasets demonstrate that the proposed model achieves a new state-of-the-art in performance. Our code is available at https://github.com/social-rumors/T3RD.
Huaiwen Zhang, Xinxin Liu 0016, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
WWW5
2024 A Versatile Multimodal Learning Framework for Zero-Shot Emotion Recognition
abstract
Multi-modal Emotion Recognition (MER) aims to identify various human emotions from heterogeneous modalities. With the development of emotional theories, there are more and more novel and fine-grained concepts to describe human emotional feelings. Real-world recognition systems often encounter unseen emotion labels. To address this challenge, we propose a versatile zero-shot MER framework to refine emotion label embeddings for capturing inter-label relationships and improving discrimination between labels. We integrate prior knowledge into a novel affective graph space that generates tailored label embeddings capturing inter-label relationships. To obtain multimodal representations, we disentangle the features of each modality into egocentric and altruistic components using adversarial learning. These components are then hierarchically fused using a hybrid co-attention mechanism. Furthermore, an emotion-guided decoder exploits label-modal dependencies to generate adaptive multimodal representations guided by emotion embeddings. We conduct extensive experiments with different multimodal combinations, including visual-acoustic and visual-textual inputs, on four datasets in both single-label and multi-label zero-shot settings. Results demonstrate the superiority of our proposed framework over state-of-the-art methods.
Fan Qi, Huaiwen Zhang, Xiaoshan Yang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2023 AffectFAL: Federated Active Affective Computing with Non-IID Data
abstract
Federated affective computing, which deploys traditional affective computing in a distributed framework, achieves a trade-off between privacy and utility, and offers a wide variety of applications in business and society. However, the expensive annotation cost of obtaining reliable emotion labels at the local client remains a barrier to the effective use of local emotional data. Therefore, we propose a federated active affective paradigm to improve the performance of federated affective computing with a limited annotation budget on the client. A major challenge in federated active learning is the inconsistency between the active sampling goals of global and local models, particularly in scenarios with Non-IID data across clients, which exacerbates the problem. To address the above challenge, we propose AffectFAL, a federated active affective computing framework. It incorporates a Preference-aware Group Aggregation module, which obtains global models representing the different emotional preferences among clients. We also devise a tailored De-biased Federated Active Sampling strategy with an improved vote entropy, facilitating class balancing of labeled samples and alleviating the problem of sampling goals inconsistency between the global and local models. We evaluate AffectFAL on diverse benchmarks (image, video and physiological signal) and experimental settings for affective computing. Thorough comparisons with other active sampling strategies demonstrate our method's advantages in affective computing for Non-IID federated learning.
Zixin Zhang 0004, Fan Qi, Changsheng Xu
ACM Multimedia2
2023 C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal Data
abstract
Massive numbers of new images are uploaded to the internet every day. However, existing cross-modal retrieval (CMR) approaches struggle to accommodate this continuously growing data. The prevalent practice involves periodically retraining or fine-tuning a new model based on the accumulated data, which in turn invalidates billions of indexed features extracted by the previous model and incurs another substantial computational cost to extract new features for the entire data archive. Is it possible to develop a retrieval model that effectively captures the knowledge of upcoming sessions while preserving the discriminative power of features extracted in previous sessions? In this paper, we propose an online continual learning setup, OC-CMR, to formalize the data-incremental growth challenge faced by cross-modal retrieval systems. It consists of two key settings: 1) Similar to the real-world scenarios, the streaming multi-modal data arrives once per session; 2) Consider the computational costs, each instance of archived data has its feature extracted only once and by its corresponding model in its session. Based on our OC-CMR, we perform in-depth evaluations of state-of-the-art cross-modal retrieval methods and observe that they suffer from representational shift and collapse due to the catastrophic forgetting. To address this issue, we propose the Continual Cross-Modal Retrieval (C2MR) approach, which learns a shared common space not only across modalities but also sessions and maintains relationships between samples from distinct sessions via cross-modal relational coherence and semantic representation coordination. We construct two new benchmarks by adapting MS-COCO and Flickr30K datasets to the OC-CMR setting, providing a more challenging evaluation framework for CMR tasks. Experimental results demonstrate that our method effectively alleviates forgetting and significantly outperforms combinations of previous arts in cross-modal retrieval and continual learning.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
ACM Multimedia3
2023 Debiased Video-Text Retrieval via Soft Positive Sample Calibration
abstract
With the emergence of enormous videos on various video apps, semantic video-text retrieval has become a critical task for improving the user experience. The primary paradigm for video-text retrieval learns the semantic video-text representations in a common space by pulling the positive samples close to the query and pushing the negative samples away. However, in practice, the video-text datasets contain only the annotations of positive samples. The negative samples are randomly drawn from the entire dataset. There may exist soft positive samples, which are sampled as negatives but share the same semantics as positive samples. Indiscriminately enforcing the model to push all the negative samples away from the query leads to inaccurate supervision and then misleads the video-text feature representation learning. In this paper, we introduce debiased video-text retrieval objectives that calibrate the punishment of soft positive samples. In particular, we propose a novel uncertainty measure framework to estimate the credibility of negative samples for each instance. Then, the reliability of negative samples is used to find the soft positive samples and rescale their contribution within video-text retrieval losses, including triplet loss and contrastive loss. Experimental results on five widely used datasets demonstrate that our debiased video-text retrieval objectives achieve significant performance improvements and establish a new state-of-the-art.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Robust Video-Text Retrieval Via Noisy Pair Calibration
abstract
Video-text retrieval is a fundamental task in managing the emerging massive amounts of video data. The main challenge focuses on learning a common representation space for videos and queries where the similarity measurement can reflect the semantic closeness. However, existing video-text retrieval models may suffer from the following noise in the common space learning procedure: First, the video-text correspondences in positive pairs may not be exact matches. The crowdsourcing annotation for existing datasets leads to inevitable tagging noise for non-expert annotators. Second, the learning of video-text representation is based on the negative samples randomly sampled. Instances that are semantically similar to the query may be incorrectly categorized as negative samples. To alleviate the adverse impact of these noisy pairs, we propose a novel robust video-text retrieval method that protects the model from noisy positive and negative pairs by identifying and calibrating noisy pairs with their uncertainty score. In particular, we propose a noisy pair identifier, which divides the training dataset into noisy and clean subsets based on the estimated uncertainty of each pair. Then, with the help of uncertainties, we calibrate the two types of noisy pairs with an adaptive margin triplet loss and a weighted triplet loss function, respectively. To verify the effectiveness of our methods, we conduct extensive experiments on three widely used datasets. Experimental results show that the proposed robust video-text retrieval methods successfully identify and calibrate the noisy pairs and improve retrieval performance.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Multim.3
2022 Feeling Without Sharing: A Federated Video Emotion Recognition Framework Via Privacy-Agnostic Hybrid Aggregation
abstract
The explosion of video data brings new opportunities and challenges for emotion recognition. Video emotion applications have great commercial value, but the potential to involve illegal snooping on personal feelings has led to controversy over privacy protection. The federated learning (FL) paradigm can substantially address the growing public concerns about data privacy in video emotion recognition. However, conventional FL methods perform poorly due to the uniqueness of the task: the data are heterogeneous across clients induced by emotional label skew and cross-culture expression differences. To mitigate the heterogeneous data, we propose EmoFed, a practical framework of federated learning video-based emotion recognition via multi-group clustering and privacy-agnostic hybrid aggregation. It yields a generically applicable and improved model while protecting privacy, which trains local models under group-aware personalized aggregation. To further encourage communicating comprehensive and privacy-agnostic information among clients, we upload model parameters of both the global layers and personalization layers to the server. We utilize the homomorphically encrypted method for personalization layers, which incurs no learning accuracy loss since no noise is added to the model updates during the encryption/decryption process. The proposed method works on video-based emotion recognition tasks to predict actors' emotional expressions and induced emotion by viewers. Extensive experiments and ablation studies on four benchmarks have demonstrated the efficacy and practicability of our method.
Fan Qi, Zixin Zhang 0004, Xianshan Yang, Huaiwen Zhang, Changsheng Xu
ACM Multimedia1
2022 A unified framework for multi-modal federated learning
Baochen Xiong, Xiaoshan Yang, Fan Qi, Changsheng Xu
Neurocomputing3
2021 Zero-shot Video Emotion Recognition via Multimodal Protagonist-aware Transformer Network
abstract
Recognizing human emotions from videos has attracted significant attention in numerous computer vision and multimedia applications, such as human-computer interaction and health care. It aims to understand the emotional response of humans, where candidate emotion categories are generally defined by specific psychological theories. However, with the development of psychological theories, emotion categories become increasingly diverse and fine-grained, samples are also increasingly difficult to collect. In this paper, we investigate a new task of zero-shot video emotion recognition, which aims to recognize rare unseen emotions. Specifically, we propose a novel multimodal protagonist-aware transformer network, which is composed of two branches: one is equipped with a novel dynamic emotional attention mechanism and a visual transformer to learn better visual representations; the other is an acoustic transformer for learning discriminative acoustic representations. We manage to align the visual and acoustic representations with semantic embeddings of fine-grained emotion labels through jointly mapping them into a common space under a noise contrastive estimation objective. Extensive experimental results on three datasets demonstrate the effectiveness of the proposed method.
Fan Qi, Xiaoshan Yang, Changsheng Xu
ACM Multimedia1
2021 Emotion Knowledge Driven Video Highlight Detection
abstract
This paper addresses video highlight detection which aims to select a small subset of frames according to user's major or special interest. The performances of conventional methods highly depend on large-scale manually labeled training data which are time-consuming and labor-intensive to collect. To deal with this problem, we trace back to the original problem definition and find that whether a user is interested in a specific video segment heavily depends on human's subjective emotions. Leveraging this insight, we introduce an emotion knowledge driven video detection framework for modeling human's general emotion and inferencing highlight strength. Firstly, we obtain the concept-level representation of the video clip with a front-end network. The concepts are used as nodes to build an emotion-related knowledge graph, and their relationships in the graph are modeled via external public knowledge graphs. Then we adopt Siamese GCNs to model the dependencies between nodes in the graph and propagate messages along the edges. Finally, we compute the emotion-aware representation of the video clip based on the GCN layers and further use it to predict the highlight score. Our framework, including the front-end network, graph convolution layers and the highlight mapping network, can be trained in an end-to-end manner with the constraint of a ranking loss. Experiments on two benchmark datasets show that our proposed method performs favorably against the state-of-the-art methods.
Fan Qi, Xiaoshan Yang, Changsheng Xu
IEEE Trans. Multim.1
2020 Discriminative multimodal embedding for event classification
Fan Qi, Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu
Neurocomputing1
2018 A Unified Framework for Multimodal Domain Adaptation
abstract
Domain adaptation aims to train a model on labeled data from a source domain while minimizing test error on a target domain. Most of existing domain adaptation methods only focus on reducing domain shift of single-modal data. In this paper, we consider a new problem of multimodal domain adaptation and propose a unified framework to solve it. The proposed multimodal domain adaptation neural networks(MDANN) consist of three important modules. (1) A covariant multimodal attention is designed to learn a common feature representation for multiple modalities. (2) A fusion module adaptively fuses attended features of different modalities. (3) Hybrid domain constraints are proposed to comprehensively learn domain-invariant features by constraining single modal features, fused features, and attention scores. Through jointly attending and fusing under an adversarial objective, the most discriminative and domain-adaptive parts of the features are adaptively fused together. Extensive experimental results on two real-world cross-domain applications (emotion recognition and cross-media retrieval) demonstrate the effectiveness of the proposed method.
Fan Qi, Xiaoshan Yang, Changsheng Xu
ACM Multimedia1