EDBT 2026 Demo / reviewers in the wild / expert
Jiajiong Cao
dblp:218/7415
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-8311-5820ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin |
Pattern Recognit. Lett. | 5 |
| 2025 | EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsabstractThe area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial key points to drive images into videos, while they can yield satisfactory results, certain issues exist. For instance, methods driven solely by audios can be unstable at times due to the relatively weaker audio signal, while methods driven exclusively by facial key points, although more stable in driving, can result in unnatural outcomes due to the excessive control of key point information. In addressing the previously mentioned challenges, in this paper, we introduce a novel approach which we named EchoMimic. EchoMimic is concurrently trained using both audios and facial landmarks. Through the implementation of a novel training strategy, EchoMimic is capable of generating portrait videos not only by audios and facial landmarks individually, but also by a combination of both audios and selected facial landmarks. EchoMimic has been comprehensively compared with alternative algorithms across various public datasets and our collected dataset, showcasing superior performance in both quantitative and qualitative evaluations. The code and models are available on the project page. Jiajiong Cao, Zhiquan Chen, Chenguang Ma |
AAAI | 2 |
| 2025 | Efficient Video Face Enhancement with Enhanced Spatial-Temporal ConsistencyabstractAs a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-World online videos are often plagued by degradations such as blurring and quantization noise, due to the high compression ratio caused by high communication costs and limited transmission bandwidth. These degradations have a particularly serious impact on face videos because the human visual system is highly sensitive to facial details. Despite the significant advancement in video face enhancement, current methods still suffer from i) long processing time and ii) inconsistent spatial-temporal visual effects (e.g., flickering). This study proposes a novel and efficient blind video face enhancement method to overcome the above two challenges, restoring high-quality videos from their compressed low-quality versions with an effective de-flickering mechanism. In particular, the proposed method develops upon a 3D-VQGAN backbone associated with spatial-temporal codebooks recording high-quality portrait features and residual-based temporal information. We develop a two-stage learning framework for the model. In Stage I, we learn the model with a regularizer mitigating the codebook collapse problem. In Stage II, we learn two transformers to look up code from the codebooks and further update the encoder of low-quality videos. Experiments conducted on the VFHQ-Test dataset demonstrate that our method surpasses the current state-of-the-art blind face video restoration and de-flickering methods on both efficiency and effectiveness. Code is available at https://github.com/Dixin-Lab/BFVR-STC. Jiajie Teng, Jiajiong Cao, Chenguang Ma, Hongteng Xu, Dixin Luo |
CVPR | 3 |
| 2024 | SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
Weilong Chai, Jiajiong Cao, Zhiquan Chen, Changbao Wang, Chenguang Ma |
ECCV (43) | 3 |
| 2024 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006, Stephen J. Maybank |
Int. J. Comput. Vis. | 2 |
| 2023 | Nasty-SFDA: Source Free Domain Adaptation from a Nasty ModelabstractA challenging problem called Nasty Source Free Domain Adaptation (Nasty-SFDA) is proposed in this work, where only a nasty source model and unlabeled target samples are available for DA. Further, after DA, the target model is expected to be a nasty model. In order to deal with Nasty-SFDA, Nasty HypOthesis Transfer (NHOT) with an improved version of Information Maximization (IM) loss called Multi-Peak Constraints (MPC) and several Label Generation (LG) techniques is proposed. Experiments on four popular datasets show the superiority of NHOT for both Nasty-SFDA and SFDA. In addition, the target model obtained via NHOT is proven to be a nasty model. Jiajiong Cao, Yufan Liu 0001, Weiming Bai, Jingting Ding, Liang Li 0006 |
ICASSP | 1 |
| 2023 | Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action RecognitionabstractVideo action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performance, compared with state-of-the-art (SOTA) raw-domain-based (RD) methods. In order to solve the problem, we propose a cross-modality knowledge distillation method to force the CD model to learn the knowledge from the RD model. In particular, spatial knowledge and temporal knowledge are first constructed to align feature space between the raw domain and the compressed domain. Then, an adaptively multi-path knowledge learning scheme is presented to help the CD model learn in a more efficient way. Experiments verify the effectiveness of the proposed method in large-scale and small-scale datasets. Yufan Liu 0001, Jiajiong Cao, Weiming Bai, Bing Li 0001, Weiming Hu 0004 |
ICASSP | 2 |
| 2023 | Learning to Explore Distillability and Sparsability: A Joint Framework for Model CompressionabstractDeep learning shows excellent performance usually at the expense of heavy computation. Recently, model compression has become a popular way of reducing the computation. Compression can be achieved using knowledge distillation or filter pruning. Knowledge distillation improves the accuracy of a lightweight network, while filter pruning removes redundant architecture in a cumbersome network. They are two different ways of achieving model compression, but few methods simultaneously consider both of them. In this paper, we revisit model compression and define two attributes of a model: distillability and sparsability, which reflect how much useful knowledge can be distilled and how many pruned ratios can be obtained, respectively. Guided by our observations and considering both accuracy and model size, a dynamically distillability-and-sparsability learning framework (DDSL) is introduced for model compression. DDSL consists of teacher, student and dean. Knowledge is distilled from the teacher to guide the student. The dean controls the training process by dynamically adjusting the distillation supervision and the sparsity supervision in a meta-learning framework. An alternating direction method of multiplier (ADMM)-based knowledge distillation-with-pruning (KDP) joint optimization algorithm is proposed to train the model. Extensive experimental results show that DDSL outperforms 24 state-of-the-art methods, including both knowledge distillation and filter pruning methods. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006 |
ACCV (5) | 2 |
| 2022 | Self-supervised Face Anti-spoofing via Anti-contrastive Learning
Jiajiong Cao, Yufan Liu 0001, Jingting Ding, Liang Li 0006 |
PRCV (2) | 1 |
| 2019 | Knowledge Distillation via Instance Relationship GraphabstractThe key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including instance features, instance relationships and feature space transformation, while the latter two kinds of knowledge are neglected by previous methods. Firstly, the IRG is constructed to model the distilled knowledge of one network layer, by considering instance features and instance relationships as vertexes and edges respectively. Secondly, an IRG transformation is proposed to models the feature space transformation across layers. It is more moderate than directly mimicking the features at intermediate layers. Finally, hint loss functions are designed to force a student's IRGs to mimic the structures of a teacher's IRGs. The proposed method effectively captures the knowledge along the whole network via IRGs, and thus shows stable convergence and strong robustness to different network architectures. In addition, the proposed method shows superior performance over existing methods on datasets of various scales. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004, Yangxi Li, Yunqiang Duan |
CVPR | 2 |
| 2018 | FR-ANet: A Face Recognition Guided Facial Attribute Classification NetworkabstractIn this paper, we study the problem of facial attribute learning. In particular, we propose a Face Recognition guided facial Attribute classification Network, called FR-ANet. All the attributes share low-level features, while high-level features are specially learned for attribute groups. Further, to utilize the identity information, high-level features are merged to perform face identity recognition. The experimental results on CelebA and LFWA datasets demonstrate the promise of the FR-ANet. Jiajiong Cao, Yingming Li, Xi Li 0001, Zhongfei Zhang |
AAAI | 1 |
| 2018 | Partially Shared Multi-Task Convolutional Neural Network With Local Constraint for Face Attribute LearningabstractIn this paper, we study the face attribute learning problem by considering the identity information and attribute relationships simultaneously. In particular, we first introduce a Partially Shared Multi-task Convolutional Neural Network (PS-MCNN), in which four Task Specific Networks (TSNets) and one Shared Network (SNet) are connected by Partially Shared (PS) structures to learn better shared and task specific representations. To utilize identity information to further boost the performance, we introduce a local learning constraint which minimizes the difference between the representations of each sample and its local geometric neighbours with the same identity. Consequently, we present a local constraint regularized multitask network, called Partially Shared Multi-task Convolutional Neural Network with Local Constraint (PS-MCNN-LC), where PS structure and local constraint are integrated together to help the framework learn better attribute representations. The experimental results on CelebA and LFWA demonstrate the promise of the proposed methods. Jiajiong Cao, Yingming Li, Zhongfei Zhang |
CVPR | 1 |
| 2018 | Celeb-500K: A Large Training Dataset for Face RecognitionabstractIn this paper, we propose a large training dataset named Celeb-500K for face recognition, which contains 50M images from 500K persons. To better facilitate academic research, we clean Celeb-500K to obtain Celeb-500K-2R, which contains 25M aligned face images from 365K persons. Based on the developed dataset, we achieve state-of-the-art face recognition performance and reveal two important observations on face recognition study. First, metric learning methods have limited performance gain when the training dataset contains a large number of identities. Second, in order to develop an efficient training dataset, the number of identities is more important than the average image number of each identity from the perspective of face recognition performance. Extensive experimental results show the superiority of Celeb-500K and provide a strong support to the two observations. Jiajiong Cao, Yingming Li, Zhongfei Zhang |
ICIP | 1 |