Yuchen Liu 0006

dblp:69/10440-6 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-3096-448XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Vision-Language Efficient Tuning for Mitigating Catastrophic Forgetting in Multi-Modal Learning
Yuchen Liu 0006, Wenrui Dai, Xiaopeng Zhang 0008, Junni Zou, Qi Tian 0001, Hongkai Xiong
Int. J. Comput. Vis.2
2025 METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
Yuchen Liu 0006, Bowen Shi 0003, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ICCV1
2024 DomainFusion: Generalizing to Unseen Domains with Latent Diffusion Models
Yabo Chen, Yuchen Liu 0006, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ECCV (41)3
2024 Towards Unified Representation of Invariant-Specific Features in Missing Modality Face Anti-spoofing
Guanghao Zheng, Yuchen Liu 0006, Wenrui Dai, Junni Zou, Hongkai Xiong
ECCV (22)2
2024 MC-DiT: Contextual Enhancement via Clean-to-Clean Reconstruction for Masked Diffusion Models
abstract
Diffusion Transformer (DiT) is emerging as a cutting-edge trend in the landscape of generative diffusion models for image generation. Recently, masked-reconstruction strategies have been considered to improve the efficiency and semantic consistency in training DiT but suffer from deficiency in contextual information extraction. In this paper, we provide a new insight to reveal that noisy-to-noisy masked-reconstruction harms sufficient utilization of contextual information. We further demonstrate the insight with theoretical analysis and empirical study on the mutual information between unmasked and masked patches. Guided by such insight, we propose a novel training paradigm named MC-DiT for fully learning contextual information via diffusion denoising at different noise variances with clean-to-clean mask-reconstruction. Moreover, to avoid model collapse, we design two complementary branches of DiT decoders for enhancing the use of noisy patches and mitigating excessive reliance on clean patches in reconstruction. Extensive experimental results on 256$\times$256 and 512$\times$512 image generation on the ImageNet dataset demonstrate that the proposed MC-DiT achieves state-of-the-art performance in unconditional and conditional image generation with enhanced convergence speed.
Guanghao Zheng, Yuchen Liu 0006, Wenrui Dai, Junni Zou, Hongkai Xiong
NeurIPS2
2024 Source-Free Domain Adaptation With Domain Generalized Pretraining for Face Anti-Spoofing
abstract
Source-free domain adaptation (SFDA) shows the potential to improve the generalizability of deep learning-based face anti-spoofing (FAS) while preserving the privacy and security of sensitive human faces. However, existing SFDA methods are significantly degraded without accessing source data due to the inability to mitigate domain and identity bias in FAS. In this paper, we propose a novel Source-free Domain Adaptation framework for FAS (SDA-FAS) that systematically addresses the challenges of source model pre-training, source knowledge adaptation, and target data exploration under the source-free setting. Specifically, we develop a generalized method for source model pre-training that leverages a causality-inspired PatchMix data augmentation to diminish domain bias and designs the patch-wise contrastive loss to alleviate identity bias. For source knowledge adaptation, we propose a contrastive domain alignment module to align conditional distribution across domains with a theoretical equivalence to adaptation based on source data. Furthermore, target data exploration is achieved via self-supervised learning with patch shuffle augmentation to identify unseen attack types, which is ignored in existing SFDA methods. To our best knowledge, this paper provides the first full-stack privacy-preserving framework to address the generalization problem in FAS. Extensive experiments on nineteen cross-dataset scenarios show our framework considerably outperforms state-of-the-art methods.
Yuchen Liu 0006, Yabo Chen, Wenrui Dai, Mengran Gou, Chun-Ting Huang, Hongkai Xiong
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Promoting Semantic Connectivity: Dual Nearest Neighbors Contrastive Learning for Unsupervised Domain Generalization
abstract
Domain Generalization (DG) has achieved great success in generalizing knowledge from source domains to unseen target domains. However, current DG methods rely heavily on labeled source data, which are usually costly and unavailable. Since unlabeled data are far more accessible, we study a more practical unsupervised domain generalization (UDG) problem. Learning invariant visual representation from different views, i.e., contrastive learning, promises well semantic features for in-domain unsupervised learning. However, it fails in cross-domain scenarios. In this paper, we first delve into the failure of vanilla contrastive learning and point out that semantic connectivity is the key to UDG. Specifically, suppressing the intra-domain connectiv-ity and encouraging the intra-class connectivity help to learn the domain-invariant semantic information. Then, we propose a novel unsupervised domain generalization approach, namely Dual Nearest Neighbors contrastive learning with strong Augmentation (DN2A). Our DN2A leverages strong augmentations to suppress the intra-domain connectivity and proposes a novel dual nearest neighbors search strategy to find trustworthy cross domain neighbors along with in-domain neighbors to encourage the intra-class connectivity. Experimental results demonstrate that our DN2A outperforms the state-of-the-art by a large margin, e.g., 12.01% and 13.11 % accuracy gain with only 1% labels for linear evaluation on PACS and DomainNet, respectively.
Yuchen Liu 0006, Yabo Chen, Wenrui Dai, Junni Zou, Hongkai Xiong
CVPR1
2023 Adapting Shortcut with Normalizing Flow: An Efficient Tuning Framework for Visual Recognition
abstract
Pretraining followed by fine-tuning has proven to be effective in visual recognition tasks. However, fine-tuning all parameters can be computationally expensive, particularly for large-scale models. To mitigate the computational and storage demands, recent research has explored Parameter-Efficient Fine-Tuning (PEFT), which focuses on tuning a minimal number of parameters for efficient adaptation. Existing methods, however, fail to analyze the impact of the additional parameters on the model, resulting in an unclear and suboptimal tuning process. In this paper, we introduce a novel and effective PEFT paradigm, named SNF (Shortcut adaptation via Normalization Flow), which utilizes normalizing flows to adjust the shortcut layers. We highlight that layers without Lipschitz constraints can lead to error propagation when adapting to downstream datasets. Since modifying the over-parameterized residual connections in these layers is expensive, we focus on adjusting the cheap yet crucial shortcuts. Moreover, learning new information with few parameters in PEFT can be challenging, and information loss can result in label information degradation. To address this issue, we propose an information-preserving normalizing flow. Experimental results demonstrate the effectiveness of SNF. Specifically, with only 0.036M parameters, SNF surpasses previous approaches on both the FGVC and VTAB-1k benchmarks using ViT/B-16 as the backbone. The code is available at https://github.com/Wang-Yaoming/SNF
Bowen Shi 0003, Xiaopeng Zhang 0008, Jin Li 0057, Yuchen Liu 0006, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
CVPR5
2023 Learning Causal Representations for Generalizable Face Anti Spoofing
abstract
Generalization ability of face anti-spoofing has been widely concerned in recent years. Existing domain generalization methods use adversarial learning or metric learning to extract invariant features across domains but are proved to be flawed from causal views. The learned domain-invariant features cannot generalize well to unseen real scenarios in both theoretical and practical senses. In this paper, we propose a novel method that learns Causal Representations for Face Anti-Spoofing (CRFAS). We first model the data generation process of face anti-spoofing via Structural Casual Model (SCM) and reveal that only the causal feature is capable of generalizing to unseen domains. On such basis, we extract causal features via back- door adjustment without prior assumptions rather than learn domain-invariant features as existing methods. Furthermore, we employ the supervised contrastive loss to generate more realistic counterfactual features for backdoor adjustment and improve the generalization ability of learned causal features. Extensive experiments on six cross-dataset testing scenarios demonstrate that CRFAS outperforms the state-of-the-art face anti-spoofing methods in terms of HTER and AUC in generalizing to unseen domains.
Guanghao Zheng, Yuchen Liu 0006, Wenrui Dai, Junni Zou, Hongkai Xiong
ICASSP2
2023 Towards Unsupervised Domain Generalization for Face Anti-Spoofing
abstract
Generalizable face anti-spoofing (FAS) based on domain generalization (DG) has gained growing attention due to its robustness in real-world applications. However, these DG methods rely heavily on labeled source data, which are usually costly and hard to access. Comparably, unlabeled face data are far more accessible in various scenarios. In this paper, we propose the first Unsupervised Domain Generalization framework for Face Anti-Spoofing, namely UDG-FAS, which could exploit large amounts of easily accessible unlabeled data to learn generalizable features for enhancing the low-data regime of FAS. Yet without supervision signals, learning intrinsic live/spoof features from complicated facial information is challenging, which is even tougher in cross-domain scenarios due to domain shift. Existing unsupervised learning methods tend to learn identity-biased and domain-biased features as shortcuts, and fail to specify spoof cues. To this end, we propose a novel Split-Rotation-Merge module to build identity-agnostic local representations for mining intrinsic spoof cues and search the nearest neighbors in the same domain as positives for mitigating the identity bias. Moreover, we propose to search cross-domain neighbors with domain-specific normalization and merged local features to learn a domain-invariant feature space. To our best knowledge, this is the first attempt to learn generalized FAS features in a fully unsupervised way. Extensive experiments show that UDG-FAS significantly outperforms state-of-the-art methods on six diverse practical protocols.
Yuchen Liu 0006, Yabo Chen, Mengran Gou, Chun-Ting Huang, Wenrui Dai, Hongkai Xiong
ICCV1
2023 VioLET: Vision-Language Efficient Tuning with Collaborative Multi-modal Gradients
abstract
Parameter-Efficient Tuning (PET) has emerged as a leading advancement in both Natural Language Processing and Computer Vision, enabling efficient accommodation of downstream tasks without costly fine-tuning. However, most existing PET approaches are limited to uni-modal tuning, even for vision-language models like CLIP. We investigate this limitation and demonstrate that simultaneous tuning of the two modalities in such models leads to multi-modal forgetting and catastrophic performance degradation, particularly when generalizing to new classes. To address this issue, we propose a novel PET approach called VioLET (Vision Language Efficient Tuning) that utilizes collaborative multi-modal gradients to unlock the full potential of both modalities. Specifically, we incorporate an additional visual encoder without learnable parameters and use these two visual encoders to compute the gradients of the context parameters separately. When conflicts arise, we replace the original gradient with an orthogonal gradient. Extensive experiments are conducted on few-shot recognition and unseen class generalization tasks using ResNet-50 or ViT/B-16 as the backbone. VioLET consistently outperforms several state-of-the-art methods on 11 datasets, showcasing its superiority over existing PET approaches. The code is available at https://github.com/Wang-Yaoming/VioLET.
Yuchen Liu 0006, Xiaopeng Zhang 0008, Jin Li 0057, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ACM Multimedia2
2022 SdAE: Self-distillated Masked Autoencoder
Yabo Chen, Yuchen Liu 0006, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ECCV (30)2
2022 Source-Free Domain Adaptation with Contrastive Domain Alignment and Self-supervised Exploration for Face Anti-spoofing
Yuchen Liu 0006, Yabo Chen, Wenrui Dai, Mengran Gou, Chun-Ting Huang, Hongkai Xiong
ECCV (12)1
2022 Causal Intervention for Generalizable Face Anti-Spoofing
abstract
Generalizable face anti-spoofing (FAS) has drawn growing attention due to its robustness to unseen real scenarios. Existing domain generalization methods leverage adversarial learning or meta-learning to mitigate the domain bias and improve generalizability. However, these methods are heuristic and suffer from complicated min-max problems or cumbersome meta-updates. In this paper, we propose a simple yet effective Causal Intervention method for generalizable Face Anti-Spoofing, namely CIFAS. Firstly, we figure out the generalizability is undermined by a domain-aware confounder based on the structural causal model. Instantiating the confounder as the domain-specific factor, a domain embedding module is employed with Dirichlet mixup to obtain representative domain features. Consequently, we propose a novel backdoor adjustment model for causal intervention to capture the true causality and learn a robust FAS model. Our CIFAS is the first attempt to introduce causal learning into FAS. Extensive experiments on seven cross-dataset tests demonstrate that CIFAS outperforms the state-of-the-art methods.
Yuchen Liu 0006, Yabo Chen, Wenrui Dai, Junni Zou, Hongkai Xiong
ICME1
2021 Learning Latent Architectural Distribution in Differentiable Neural Architecture Search via Variational Information Maximization
abstract
Existing differentiable neural architecture search approaches simply assume the architectural distribution on each edge is independent of each other, which conflicts with the intrinsic properties of architecture. In this paper, we view the architectural distribution as the latent representation of specific data points. Then we propose Variational Information Maximization Neural Architecture Search (VIM-NAS) to leverage a simple yet effective convolutional neural network to model the latent representation, and optimize for a tractable variational lower bound to the mutual information between the data points and the latent representations. VIM-NAS automatically learns a nearly one-hot distribution from a continuous distribution with extremely fast convergence speed, e.g., converging with one epoch. Experimental results demonstrate VIM-NAS achieves state-of-the-art performance on various search spaces, including DARTS search space, NAS-Bench-1shot1, NAS-Bench-201, and simplified search spaces S1-S4. Specifically, VIM-NAS achieves a top-1 error rate of 2.45% and 15.80% within 10 minutes on CIFAR-10 and CIFAR-100, respectively, and a top-1 error rate of 24.0% when transferred to ImageNet.
Yuchen Liu 0006, Wenrui Dai, Junni Zou, Hongkai Xiong
ICCV2