EDBT 2026 Demo / reviewers in the wild / expert
Shan Jia
dblp:176/3600
· DBLP profile ↗
22ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0001-7503-8378ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modality-Agnostic Deepfakes DetectionabstractAs AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated.While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked.Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities.However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable.In such cases, audio-visual detection methods are less practical than two independent unimodal methods.Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector.In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms.To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) Jin Liu 0020, Jiao Dai, Xi Wang 0014, Shan Jia, Siwei Lyu, Jizhong Han |
IH&MMSec | 7 |
| 2024 | Exposing Text-Image Inconsistency Using Diffusion ModelsabstractIn the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based methods for text-image inconsistency can identify contextual inconsistencies but fail to provide explainable justifications for their decisions that humans can understand. Although more nuanced, human evaluation is impractical at scale and susceptible to errors. To address these limitations, this study introduces D-TIIL (Diffusion-based Text-Image Inconsistency Localization), which employs text-to-image diffusion models to localize semantic inconsistencies in text and image pairs. These models, trained on large-scale datasets act as ``omniscient" agents that filter out irrelevant information and incorporate background knowledge to identify inconsistencies. In addition, D-TIIL uses text embeddings and modified image regions to visualize these inconsistencies. To evaluate D-TIIL's efficacy, we introduce a new TIIL dataset containing 14K consistent and inconsistent text-image pairs. Unlike existing datasets, TIIL enables assessment at the level of individual words and image regions and is carefully designed to represent various inconsistencies. D-TIIL offers a scalable and evidence-based approach to identifying and localizing text-image inconsistency, providing a robust framework for future research combating misinformation. Mingzhen Huang, Shan Jia, Zhou Zhou 0009, Yan Ju, Jialing Cai, Siwei Lyu |
ICLR | 2 |
| 2024 | Exposing Lip-syncing Deepfakes from Mouth InconsistenciesabstractA lip-syncing deepfake is a digitally manipulated video in which a person’s lip movements are created convincingly using AI models to match altered or entirely new audio. Lipsyncing deepfakes are a dangerous type of deepfakes as the artifacts are limited to the lip region and more difficult to discern. In this paper, we describe a novel approach, LIP-syncing detection based on mouth INConsistency (LIPINC), for lip-syncing deepfake detection by identifying temporal inconsistencies in the mouth region. These inconsistencies are seen in the adjacent frames and throughout the video. Our model can successfully capture these irregularities and outperforms the state-of-the-art methods on several benchmark deepfake datasets. Code is available at https://github.com/skrantidatta/LIPINC. Soumyya Kanti Datta, Shan Jia, Siwei Lyu |
ICME | 2 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 2 |
| 2024 | ParallelEdits: Efficient Multi-Aspect Text-Driven Image Editing with Attention GroupingabstractText-driven image synthesis has made significant advancements with the development of diffusion models, transforming how visual content is generated from text prompts. Despite these advances, text-driven image editing, a key area in computer graphics, faces unique challenges. A major challenge is making simultaneous edits across multiple objects or attributes. Applying these methods sequentially for multi-attribute edits increases computational demands and efficiency losses.
In this paper, we address these challenges with significant contributions. Our main contribution is the development of ParallelEdits, a method that seamlessly manages simultaneous edits across multiple attributes. In contrast to previous approaches, ParallelEdits not only preserves the quality of single attribute edits but also significantly improves the performance of multitasking edits. This is achieved through innovative attention distribution mechanism and multi-branch design that operates across several processing heads.
Additionally, we introduce the PIE-Bench++ dataset, an expansion of the original PIE-Bench dataset, to better support evaluating image-editing tasks involving multiple objects and attributes simultaneously. This dataset is a benchmark for evaluating text-driven image editing methods in multifaceted scenarios. Mingzhen Huang, Jialing Cai, Shan Jia, Vishnu Suresh Lokhande, Siwei Lyu |
NeurIPS | 3 |
| 2024 | Improving Fairness in Deepfake DetectionabstractDespite the development of effective deepfake detectors in recent years, recent studies have demonstrated that biases in the data used to train these detectors can lead to disparities in detection accuracy across different races and genders. This can result in different groups being unfairly targeted or excluded from detection, allowing undetected deepfakes to manipulate public opinion and erode trust in a deepfake detection model. While existing studies have focused on evaluating fairness of deepfake detectors, to the best of our knowledge, no method has been developed to encourage fairness in deepfake detection at the algorithm level. In this work, we make the first attempt to improve deepfake detection fairness by proposing novel loss functions that handle both the setting where demographic information (e.g., annotations of race and gender) is available as well as the case where this information is absent. Fundamentally, both approaches can be used to convert many existing deepfake detectors into ones that encourages fairness. Extensive experiments on four deepfake datasets and five deepfake detectors demonstrate the effectiveness and flexibility of our approach in improving deep-fake detection fairness. Our code is available at https://github.com/littlejuyan/DF_Fairness. Yan Ju, Shu Hu 0001, Shan Jia, George H. Chen, Siwei Lyu |
WACV | 3 |
| 2024 | MDTL-NET: Computer-generated image detection based on multi-scale deep texture learning
Qiang Xu 0007, Shan Jia, Xinghao Jiang, Tanfeng Sun, Zhe Wang 0035, Hong Yan 0001 |
Expert Syst. Appl. | 2 |
| 2024 | GLFF: Global and Local Feature Fusion for AI-Synthesized Image DetectionabstractWith the rapid development of deep generative models (such as Generative Adversarial Networks and Diffusion models), AI-synthesized images are now of such high quality that humans can hardly distinguish them from pristine ones. Although existing detection methods have shown high performance in specific evaluation settings,e.g., on images from seen models or on images without real-world post-processing, they tend to suffer serious performance degradation in real-world scenarios where testing images can be generated by more powerful generation models or combined with various post-processing operations. To address this issue, we propose a Global and Local Feature Fusion (GLFF) framework to learn rich and discriminative representations by combining multi-scale global features from the whole image with refined local features from informative patches for AI-synthesized image detection. GLFF fuses information from two branches: the global branch to extract multi-scale semantic features and the local branch to select informative patches for detailed local artifacts extraction. Due to the lack of a synthesized image dataset simulating real-world applications for evaluation, we further create a challenging fake image dataset, named DeepFakeFaceForensics ($DF^{3}$), which contains 6 state-of-the-art generation models and a variety of post-processing techniques to approach the real-world scenarios. Experimental results demonstrate the superiority of our method to the state-of-the-art methods on the proposed$DF^{3}$dataset and three other open-source datasets. Yan Ju, Shan Jia, Jialing Cai, Haiying Guan, Siwei Lyu |
IEEE Trans. Multim. | 2 |
| 2022 | Text-Image De-Contextualization Detection Using Vision-Language ModelsabstractText-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization is a challenging problem in media forensics. Inspired by the recent advances in vision-language models with powerful relationship learning between images and texts, we leverage the vision-language models to the media de-contextualization detection task. Two popular models, namely CLIP and VinVL, are evaluated and compared on several news and social media datasets to show their performance in detecting image-text inconsistency in de-contextualization. We also summarize interesting observations and shed lights to the use of vision-language models in de-contextualization detection. Mingzhen Huang, Shan Jia, Ming-Ching Chang, Siwei Lyu |
ICASSP | 2 |
| 2022 | Model Attribution of Face-Swap Deepfake VideosabstractAI-created face-swap videos, commonly known as Deepfakes, have attracted wide attention as powerful impersonation attacks. Existing research on Deepfakes mostly focuses on binary detection to distinguish between real and fake videos. However, it is also important to determine the specific generation model for a fake video, which can help attribute it to the source for forensic investigation. In this paper, we fill this gap by studying the model attribution problem of Deepfake videos. We first introduce a new dataset with DeepFakes from Different Models (DFDM) based on several Autoencoder models. Specifically, five generation models with variations in encoder, decoder, intermediate layer, input resolution, and compression ratio have been used to generate a total of 6, 450 Deepfake videos based on the same input. Then we take Deepfakes model attribution as a multiclass classification task and propose a spatial and temporal attention based method to explore the differences among Deep-fakes in the new dataset. Experimental evaluation shows that most existing Deepfakes detection methods failed in Deep-fakes model attribution, while the proposed method achieved over 70% accuracy on the high-quality DFDM dataset1. Shan Jia, Xin Li 0005, Siwei Lyu |
ICIP | 1 |
| 2022 | Fusing Global and Local Features for Generalized AI-Synthesized Image DetectionabstractWith the development of the Generative Adversarial Networks (GANs) and DeepFakes, AI-synthesized images are now of such high quality that humans can hardly distinguish them from real images. It is imperative for media forensics to develop detectors to expose them accurately. Existing detection methods have shown high performance in generated images detection, but they tend to generalize poorly in the real-world scenarios, where the synthetic images are usually generated with unseen models using unknown source data. In this work, we emphasize the importance of combining information from the whole image and informative patches in improving the generalization ability of AI-synthesized image detection. Specifically, we design a two-branch model to combine global spatial information from the whole image and local informative features from multiple patches selected by a novel patch selection module. Multi-head attention mechanism is further utilized to fuse the global and local features. We collect a highly diverse dataset synthesized by 19 models with various objects and resolutions to evaluate our model. Experimental results demonstrate the high accuracy and good generalization ability of our method in detecting generated images. Our code is available at https://github.com/littlejuyan/FusingGlobalandLocal. Yan Ju, Shan Jia, Lipeng Ke, Hongfei Xue, Koki Nagano, Siwei Lyu |
ICIP | 2 |
| 2021 | Face spoofing detection under super-realistic 3D wax face attacks
Shan Jia, Chuanbo Hu, Xin Li 0005, Zhengquan Xu |
Pattern Recognit. Lett. | 1 |
| 2021 | 3D Face Anti-Spoofing With Factorized Bilinear CodingabstractWe have witnessed rapid advances in both face presentation attack models and presentation attack detection (PAD) in recent years. When compared with widely studied 2D face presentation attacks, 3D face spoofing attacks are more challenging because face recognition systems are more easily confused by the 3D characteristics of materials similar to real faces. In this work, we tackle the problem of detecting these realistic 3D face presentation attacks and propose a novel anti-spoofing method from the perspective of fine-grained classification. Our method, based on factorized bilinear coding of multiple color channels (namely MC_FBC), targets at learning subtle fine-grained differences between real and fake images. By extracting discriminative and fusing complementary information from RGB and YCbCr spaces, we have developed a principled solution to 3D face spoofing detection. A large-scale wax figure face database (WFFD) with both images and videos has also been collected as super realistic attacks to facilitate the study of 3D face presentation attack detection. Extensive experimental results show that our proposed method achieves the state-of-the-art performance on both our own WFFD and other face spoofing databases under various intra-database and inter-database testing scenarios. Shan Jia, Xin Li 0005, Chuanbo Hu, Guodong Guo, Zhengquan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Why current differential privacy schemes are inapplicable for correlated data publishing?
Hao Wang 0033, Zhengquan Xu, Shan Jia, Xu Zhang 0019 |
World Wide Web | 3 |
| 2020 | Face presentation attack detection in mobile scenarios: A comprehensive evaluation
Shan Jia, Guodong Guo, Zhengquan Xu, Qiangchang Wang |
Image Vis. Comput. | 1 |
| 2020 | A survey on 3D mask presentation attack detection and countermeasures
Shan Jia, Guodong Guo, Zhengquan Xu |
Pattern Recognit. | 1 |
| 2017 | A Case Study of Evaluation of Learners' Acceptance of AR_H2O2 System
Tao Wang 0037, Shan Jia, Jinrui Dai, Manli Lu, Xiaoru Xue, Su Cai, Feng-Kuang Chiang |
ICCE | 2 |
| 2017 | Designing a "Three Rings" Theory Framework for Electronic Schoolbag
Baoyuan Yin, Fati Wu, Shihua Huang, Shan Jia |
ICCE | 4 |
| 2017 | Cluster-Indistinguishability: A practical differential privacy mechanism for trajectory clusteringabstractAn important method of spatial-temporal data mining, trajectory clustering can mine valuable information in trajectories. However, cluster results without special sanitization pose serious threats to individual location privacy. Existing privacy preserving mechanisms for trajectory clustering still contend with the problems of narrow applicability, low-level utility, and difficulty in being applied to real scenarios. In this paper, we therefore propose a differential privacy preserving mechanism, Cluster-Indistinguishability, to support trajectory clustering. Firstly, a general model of typical trajectory clustering algorithms is given, and the definition of differential privacy is introduced according to the model. Then, we derive the probability density function of two-dimensional Laplace noise, which satisfies the above definition. Finally, we transform the noise from a Cartesian coordinate system to a Polar coordinate system to efficiently apply it in real scenarios. Experimental results show that Cluster-Indistinguishability has general applicability and better performance compared to existing methods. Hao Wang 0033, Zhengquan Xu, Shan Jia |
Intell. Data Anal. | 3 |
| 2017 | Motion-Adaptive Frame Deletion Detection for Digital Video ForensicsabstractThe detection of frame deletion forgery is of great significance in the field of video forensics. Existing approaches, however, are not applicable to video sequences with variable motion strengths. In addition, the impact of interfering frames has not been considered in these approaches. Our research aims to develop a motion-adaptive forensic method as well as to eliminate interfering frames. Through a study of the statistical characteristics of the most common interfering frames such as relocated I-frames, we develop a new fluctuation feature based on frame motion residuals to identify frame deletion points (FDPs). The fluctuation feature is further enhanced by an intra-prediction elimination procedure so that it can be adapted to sequences with various motion levels. The enhanced feature is measured using a moving window detector to identify the location of a FDP. Finally, a postprocessing procedure is proposed to eliminate the minor interferences of sudden lighting change, focus vibration, and frame jitter. Our experimental results demonstrate that for videos with variable motion strengths and different interfering frames, the true positive rate of the algorithm can reach 90% when the false alarm rate is 0.3%. Our proposed method could provide a foundation for many practical applications of video forensics. Chunhui Feng, Zhengquan Xu, Shan Jia, Yanyan Xu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A fast sub-pixel motion estimation algorithm for HEVCabstractMany fast integer-pixel motion estimation algorithms have been developed for the High Efficiency Video Coding Standard, however the speed of sub-pixel motion estimation still has room for improvement. A fast sub-pixel motion estimation algorithm is proposed in this paper to speed up the sub-pixel search process. First, the proposed scheme skips sub-pixel search process in smooth prediction units. Then a fast sub-pixel search algorithm based on texture direction analysis is proposed to further reduce the computational complexity of subpixel motion estimation. The simulation results show that compared with the Full Sub-pixel Search (FSPS), the encoding complexity of the whole motion estimation process can be reduced by an average of 40.9% with negligible coding performance loss. Shan Jia, Wenpeng Ding, Yunhui Shi |
ISCAS | 1 |
| 2016 | DCCP: an effective data placement strategy for data-intensive computations in distributed cloud computing systems
Tao Wang 0037, Shihong Yao, Zhengquan Xu, Shan Jia |
J. Supercomput. | 4 |