Xuanbai Chen

dblp:279/6352 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-5777-781XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
abstract
Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of their face data with service providers. In this work, we introduce a novel method in which the face anti-spoofing model can be adapted by the client itself to a target domain at test time using only a small sample of data while keeping model parameters and training data inaccessible to the client. Specifically, we develop a prototype-based base model and an optimal transport-guided adaptor that enables adaptation in either a lightweight training or training-free fashion, without updating base model’s parameters. Furthermore, we propose geodesic mixup, an optimal transport-based synthesis method that generates augmented training data along the geodesic path between source prototypes and target data distribution. This allows training a lightweight classifier to effectively adapt to target-specific characteristics while retaining essential knowledge learned from the source domain. In cross-domain and cross-attack settings, compared with recent methods, our method achieves average relative improvements of 19.17% in HTER and 8.58% in AUC, respectively.
Zhuowei Li 0002, Tianchen Zhao, Xuanbai Chen, Alessandro Bergamo, Anil K. Jain 0001, Yifan Xing
CVPR6
2025 Model Diagnosis and Correction via Linguistic and Implicit Attribute Editing
abstract
How can we troubleshoot a deep visual model, i.e. understand why it makes certain mistakes and further take action to correct its behavior? We design a ${\mathbf{M}}$ odel ${\mathbf{D}}$ iagnosis and ${\mathbf{C}}$ orrection system (MDC), an automated framework that analyzes the pattern of errors, proposes candidate causes of attributes, conducts hypothesis testing via attribute editing, and ultimately generates counterfactual training samples to improve the performance of the model. Unlike previous methods, in addition to the linguistic attributes, our method also incorporates the analysis for implicit causal attributes, those cannot to be accurately described by natural language. To achieve this, we propose an image editing module capable of leveraging both implicit and linguistic attributes to generate counterfactual images depicting error patterns and further experimentally validate causality relationships. Lastly, we enrich the training set with synthetic samples depicting verified causal attributes and retrain the model, further boosting accuracy and robustness. Extensive experiments on both generalized and specialized domains demonstrate the superiority of MDC in model diagnosis and correction. Specifically, we achieve an average relative improvement of 62.01% in HTER for face security application over state-of-the-art methods.
Xuanbai Chen, Tianchen Zhao, Pietro Perona, Yifan Xing
CVPR1
2025 Salient Concept-Aware Generative Data Augmentation
abstract
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73\% and 6.5\% under conventional and long-tail settings, respectively.
Tianchen Zhao, Xuanbai Chen, Dongsheng An, Zhuowen Tu, Yifan Xing
NeurIPS2
2024 Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
abstract
Vanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional control - human-centric priors such as pose or depth maps - during the image generation phase. This paper explores the integration of these human-centric priors directly into the model fine-tuning stage, essentially eliminating the need for extra conditions at the inference stage. We realize this idea by proposing a human-centric alignment loss to strengthen human-related information from the textual prompts within the cross-attention maps. To ensure semantic detail richness and human structural accuracy during fine-tuning, we introduce scale-aware and step-wise constraints within the diffusion process, according to an indepth analysis of the cross-attention layer. Extensive experiments show that our method largely improves over state-of-the-art text-to-image models to synthesize high-quality human images based on user-written prompts. Project page: https://hcplayercvpr2024.github.io.
Junyan Wang 0001, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen, Hao Li 0030, Cheng Zhang 0014, Yang Song 0001
CVPR4
2024 An Aesthetic-Guided Multimodal Framework for Video Summarization
abstract
Video summarization aims to generate a concise yet informative version of a lengthy video for efficient viewing. Generally, humans can discern important shots using audiovisual information and aesthetic elements of the video. However, existing methods often rely solely on unimodal information or ignore aesthetic elements, resulting in sub-optimal summaries. To tackle these issues, we present an aesthetic-guided multimodal framework (AMF) for video summarization. Specifically, we design a visual-aesthetics encoder to extract diverse aesthetic elements, and jointly integrate with visual content to create a comprehensive feature representation. Furthermore, we construct a multimodal fusion module to leverage the complementary properties among modalities, focusing on the most important content for the desired summary. Extensive experiments on four popular datasets show that the proposed method outperforms various compared methods in commonly used evaluation metrics.
Jiehang Xie, Xuanbai Chen, Shao-Ping Lu
ICME2
2024 Video summarization via knowledge-aware multimodal deep networks
Jiehang Xie, Xuanbai Chen, Sicheng Zhao, Shao-Ping Lu
Knowl. Based Syst.2
2023 ITI-Gen: Inclusive Text-to-Image Generation
abstract
Text-to-image generative models often reflect the biases of the training data, leading to unequal representations of underrepresented groups. This study investigates inclusive text-to-image generative models that generate images based on human-written prompts and ensure the resulting images are uniformly distributed across attributes of interest. Unfortunately, directly expressing the desired attributes in the prompt often leads to sub-optimal results due to linguistic ambiguity or model misrepresentation. Hence, this paper proposes a drastically different approach that adheres to the maxim that "a picture is worth a thousand words". We show that, for some attributes, images can represent concepts more expressively than text. For instance, categories of skin tones are typically hard to specify by text but can be easily represented by example images. Building upon these insights, we propose a novel approach, ITI-Gen1, that leverages readily available reference images for Inclusive Text-to-Image GENeration. The key idea is learning a set of prompt embeddings to generate images that can effectively represent all desired attribute categories. More importantly, ITI-Gen requires no model fine-tuning, making it computationally efficient to augment existing text-to-image models. Extensive experiments demonstrate that ITI-Gen largely improves over state-of-the-art models to generate inclusive images from a prompt.
Cheng Zhang 0014, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, Fernando De la Torre
ICCV2
2023 Controllable 3D Generative Adversarial Face Model via Disentangling Shape and Appearance
abstract
3D face modeling has been an active area of research in computer vision and computer graphics, fueling applications ranging from facial expression transfer in virtual avatars to synthetic data generation. Existing 3D deep learning generative models (e.g., VAE, GANs) allow generating compact face representations (both shape and texture) that can model non-linearities in the shape and appearance space (e.g., scatter effects, specularities,..). However, they lack the capability to control the generation of subtle expressions. This paper proposes a new 3D face generative model that can decouple identity and expression and provides granular control over expressions. In particular, we propose using a pair of supervised auto-encoder and generative adversarial networks to produce high-quality 3D faces, both in terms of appearance and shape. Experimental results in the generation of 3D faces learned with holistic expression labels, or Action Unit (AU) labels, show how we can decouple identity and expression; gaining fine-control over expressions while preserving identity.1
Fariborz Taherkhani, Aashish Rai, Quankai Gao, Shaunak Srivastava, Xuanbai Chen, Fernando De la Torre, Steven Song, Aayush Prakash, Daeil Kim
WACV5
2023 Multimodal-Based and Aesthetic-Guided Narrative Video Summarization
abstract
Narrative videos usually illustrate the main content through multiple narrative information such as audios, video frames and subtitles. Existing video summarization approaches rarely consider the multiple dimensional narrative inputs, or ignore the impact of shots artistic assembly when directly applied to narrative videos. This paper introduces a multimodal-based and aesthetic-guided narrative video summarization method. Our method leverages multimodal information including visual content, subtitles and audio information through our specified key shots selection, subtitle summarization, and highlight extraction components. Furthermore, under the guidance of cinematographic aesthetic, we design a novel shots assembly module to ensure the shot content completeness and then assemble the selected shots into a desired summary. Besides, our method also provides the flexible specification for shots selection, to achieve which it automatically selects semantically related shots according to the user-designed text. By conducting a large number of quantitative experimental evaluations and user studies, we demonstrate that our method effectively preserves important narrative information of the original video, and it is capable of rapidly producing high-quality and aesthetic-guided narrative video summaries.
Jiehang Xie, Xuanbai Chen, Tianyi Zhang 0013, Shao-Ping Lu, Pablo César, Yulu Yang
IEEE Trans. Multim.2
2022 A Knowledge Augmented and Multimodal-Based Framework for Video Summarization
abstract
Video summarization aims to generate a compact version of a lengthy video that retains its primary content. In general, humans are gifted with producing a high-quality video summary, because they acquire crucial content through multiple dimensional information and own abundant background knowledge about the original video. However, existing methods rarely consider multichannel information and ignore the impact of external knowledge, resulting in the limited quality of the generated summaries. This paper proposes a knowledge augmented and multimodal-based video summarization method, termed KAMV, to address the problem above. Specifically, we design a knowledge encoder with a hybrid method consisting of generation and retrieval, to capture descriptive content and latent connections between events and entities based on the external knowledge base, which can provide rich implicit knowledge for better comprehending the video viewed. Furthermore, for the sake of exploring the interactions among visual, audio, implicit knowledge and emphasizing the content that is most relevant to the desired summary, we present a fusion module under the supervision of these multimodal information. By conducting extensive experiments on four public datasets, the results demonstrate the superior performance yielded by the proposed KAMV compared to the state-of-the-art video summarization approaches.
Jiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu Yang
ACM Multimedia2
2022 Emotional Semantics-Preserved and Feature-Aligned CycleGAN for Visual Emotion Adaptation
abstract
Thanks to large-scale labeled training data, deep neural networks (DNNs) have obtained remarkable success in many vision and multimedia tasks. However, because of the presence of domain shift, the learned knowledge of the well-trained DNNs cannot be well generalized to new domains or datasets that have few labels. Unsupervised domain adaptation (UDA) studies the problem of transferring models trained on one labeled source domain to another unlabeled target domain. In this article, we focus on UDA in visual emotion analysis for both emotion distribution learning and dominant emotion classification. Specifically, we design a novel end-to-end cycle-consistent adversarial model, called CycleEmotionGAN++. First, we generate an adapted domain to align the source and target domains on the pixel level by improving CycleGAN with a multiscale structured cycle-consistency loss. During the image translation, we propose a dynamic emotional semantic consistency loss to preserve the emotion labels of the source images. Second, we train a transferable task classifier on the adapted domain with feature-level alignment between the adapted and target domains. We conduct extensive UDA experiments on the Flickr-LDL and Twitter-LDL datasets for distribution learning and ArtPhoto and Flickr and Instagram datasets for emotion classification. The results demonstrate the significant improvements yielded by the proposed CycleEmotionGAN++ compared to state-of-the-art UDA approaches.
Sicheng Zhao, Xuanbai Chen, Xiangyu Yue 0001, Chuang Lin 0003, Pengfei Xu 0013, Ravi Krishna, Jufeng Yang, Guiguang Ding, Alberto L. Sangiovanni-Vincentelli, Kurt Keutzer
IEEE Trans. Cybern.2