EDBT 2026 Demo / reviewers in the wild / expert
Zenan Shi
dblp:174/1709
· DBLP profile ↗
21ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-8554-4127ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit token modeling and hierarchical feature reasoning for multimodal fake news detection
Yixin Jia, Haipeng Chen 0002, Zenan Shi, Xun Yang 0001 |
Inf. Process. Manag. | 3 |
| 2026 | Bayesian perturbation-driven consistency regularization for semi-supervised medical image segmentation
Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Yongping Yang, Yu Wang 0112 |
Knowl. Based Syst. | 4 |
| 2026 | A multi-scale adaptive active selection strategy in semi-supervised medical image segmentation
Mingxiu Zhang, Zenan Shi, Jincai Song |
Multim. Syst. | 3 |
| 2026 | ADNet: Delving into generalizable deepfake detection via adaptive expert selection and discrepancy learning
Haipeng Chen 0002, Yixin Jia, Zenan Shi |
Pattern Recognit. | 3 |
| 2025 | Robustifying vision transformer for image forgery localization with multi-exit architectures
Zenan Shi |
Pattern Recognit. | 1 |
| 2025 | Customized Transformer Adapter With Frequency Masking for Deepfake DetectionabstractThe evolution of advanced artificial intelligence generated content approaches has heightened concerns about deepfake, due to the sophisticated forgeries and concealed appearances they produce. To this end, the pre-trained Vision Transformer (ViT) model has become a de facto choice for deepfake detection, thanks to its powerful learning capability. Despite favorable results achieved by existing ViT-based methods, they have inherent limitations that could result in suboptimal performance in scenarios with continuously evolving forgery techniques, such as overfitting to single forgery patterns or placing excessive emphasis on dominant forgery regions. In this paper, we propose CUTA, a simple yet effective deepfake detection paradigm that utilizes ViT adapters as the medium and fully exploits the spatial- and frequency-domain features of given images to overcome the limitations of existing methods. Specifically, CUTA focuses onfrequency domain maskingwithin the input space, which obscures parts of the high-frequency image to intensify the training challenge while preserving subtle forgery cues in the frequency domain to facilitate comprehensive forgery representations. Furthermore, we propose two task-customized modules within the ViT model, i.e., thetexture enhancement moduleand themulti-scale perceptron module, to seamlessly integrate local texture and rich contextual features. These two modules ensure an organic interaction between the task-specific forgery patterns and general semantic features within the pre-trained ViT framework. The experimental results on several publicly available benchmark datasets demonstrate CUTA’s superiority in performance, particularly showcasing its significant advantages in both cross-dataset and cross-manipulation scenarios. Zenan Shi, Haipeng Chen 0002, Yixin Jia, Wei Lu 0001, Xun Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Face Reconstruction-Based Generalized Deepfake Detection Model with Residual Outlook AttentionabstractWith the continuous development of deep counterfeiting technology, the information security in our daily life is under serious threat. While existing face forgery detection methods exhibit impressive accuracy when applied to datasets such as FaceForensics++ and Celeb-DF, they falter significantly when confronted with out-of-domain scenarios. This causes specialization of learned representations to known forgery patterns presented in the training set, rendering it difficult to detect forgeries with unknown patterns. To address this challenge, we propose a novel end-to-end Face Reconstruction-Based Generalized Deepfake Detection (FRG2D) model with Residual Outlook Attention (ROA) , which emphasizes the robust visual representations of genuine faces and discerns the subtle differences between authentic and manipulated facial images. Our methodology entails reconstructing authentic face images using an encoder–decoder architecture based on U-net, facilitating a deeper understanding of disparities between genuine and manipulated facial images. Furthermore, we integrate the convolutional block attention module (CBAM) and channel attention block (CAB) to selectively focus the network’s attention on salient features within real face images. Furthermore, we employ ROA to guide the network’s focus towards precise features within manipulated facial images. Simultaneously, the computed reconstruction differences obtained through ROA serves as the ultimate representation fed into the classifier for face forgery detection. Both the reconstruction and classification learning processes are optimized end-to-end. Through extensive experimentation, our model demonstrated a substantial improvement in deepfake detection across unknown domains, while maintaining a high accuracy within the known domain. Zenan Shi, Wenyu Liu 0013, Haipeng Chen 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Collaborative region-boundary interaction network for medical image segmentation
Na Ta 0009, Haipeng Chen 0002, Zenan Shi |
Multim. Tools Appl. | 5 |
| 2023 | Discrepancy-Guided Reconstruction Learning for Image Forgery DetectionabstractIn this paper, we propose a novel image forgery detection paradigm for boosting the model learning capacity on both forgery-sensitive and genuine compact visual patterns. Compared to the existing methods that only focus on the discrepant-specific patterns (\eg, noises, textures, and frequencies), our method has a greater generalization. Specifically, we first propose a Discrepancy-Guided Encoder (DisGE) to extract forgery-sensitive visual patterns. DisGE consists of two branches, where the mainstream backbone branch is used to extract general semantic features, and the accessorial discrepant external attention branch is used to extract explicit forgery cues. Besides, a Double-Head Reconstruction (DouHR) module is proposed to enhance genuine compact visual patterns in different granular spaces. Under DouHR, we further introduce a Discrepancy-Aggregation Detector (DisAD) to aggregate these genuine compact visual patterns, such that the forgery detection capability on unknown patterns can be improved. Extensive experimental results on four challenging datasets validate the effectiveness of our proposed method against state-of-the-art competitors. Zenan Shi, Haipeng Chen 0002, Long Chen 0016 |
IJCAI | 1 |
| 2023 | Semantic-agnostic progressive subtractive network for image manipulation detection and localization
Dengyun Xu, Xuanjing Shen, Zenan Shi, Na Ta 0009 |
Neurocomputing | 3 |
| 2023 | A complementary and contrastive network for stimulus segmentation and generalization
Na Ta 0009, Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Zhehao Liu |
Image Vis. Comput. | 5 |
| 2023 | FPF-Net: feature propagation and fusion based on attention mechanism for pancreas segmentation
Haipeng Chen 0002, Zenan Shi |
Multim. Syst. | 3 |
| 2023 | RB-Net: integrating region and boundary features for image manipulation localization
Dengyun Xu, Xuanjing Shen, Zenan Shi |
Multim. Syst. | 4 |
| 2023 | PL-GNet: Pixel Level Global Network for detection and localization of image forgeries
Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu |
Signal Process. Image Commun. | 1 |
| 2023 | Transformer-Auxiliary Neural Networks for Image Manipulation Localization by Operator InductionsabstractImage manipulation localization (IML), which seeks to accurately segment tampered regions that are artfully fastened into a normal image, is a fundamental yet challenging computer vision task. Despite that impressive results have been achieved by some progressive deep learning methods, they usually fail in capturing the subtle manipulation artifacts at different object scales, which are not competent to generate a perfect segmentation mask with complete and fine object structures. Besides, the problem of coarse boundaries also occurs frequently. To this end, in this paper, we propose a Transformer-Auxiliary by operator-induced neural Network (TANet) to localize forged regions for IML. Specifically, a stacked multi-scale transformer (SMT) branch is first introduced as a compensation for feature representations of the mainstream convolutional neural network branch. SMT can detect structured abnormalities of the input image at multi-levels by operating on patches of different sizes. Then TANet explicitly exploits an operator induction module (OIM) to excavate valuable and manipulated region-related boundary semantics to guide the representative learning of the mainstream branch. The OIM encourages the network to generate features that highlight object structure, thereby promoting precise boundary localization of forged regions. We conduct extensive experiments on various datasets and settings to validate the effectiveness of TANet. Results show that TANet outperforms the state-of-the-art methods by a large margin under widely-used evaluation metrics. Zenan Shi, Haipeng Chen 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Adaptive Multi-Order Graph Neural Networks for Human Motion PredictionabstractHuman motion prediction aims at capturing the hidden temporal correlations between historical motion and future poses. Various graph convolution networks have been presented for encoding the spatial dependencies between joints. Empirically, the crucial shortcoming of these methods is that they fail to extract enough spatially relevant information. In this paper, we propose an adaptive multi-order context fusion architecture that consists of two components. A novel message propagation module encodes the interaction between joints, while highlighting contexts from closely related joints. An adaptive aggregation module fuses various information from different-order joint features. Our model is evaluated on Human 3.6 Million dataset. Extensive experiments show that our method achieves state-of-the-art performance on short-term and long-term predictions. Pengxiang Su, Xuanjing Shen, Zenan Shi |
ICME | 3 |
| 2022 | PR-NET: Progressively-refined neural network for image manipulation localizationabstractCurrent deep learning-based image manipulation localization methods achieve impressive performance when rich spatial features and information are fully utilized. However, most of them suffer from the irrelevance of semantic awareness when identifying various manipulation categories. This leads to false alarms on recognizing forged regions. In this paper, we propose a Progressively-Refined Neural Network (PR-Net), to localize the tampered regions progressively under a coarse-to-fine workflow. Specifically, PR-Net is composed of a Feature Extractor (FE) that captures feature intrinsic correlations and a Mask Generation Module (MGM) with three refining generators. The FE takes a CNN to extract the image features and introduces an attention mechanism Convolution Block Attention Module (CBAM) to suppress the image content and guide the extractor in exploring the inconsistencies between the manipulated and authentic regions. The MGM comprises three generators where the Coarse Mask RR-Generator generates a localization result roughly, the Candidate Mask RR-Generator generates a possible tampered region according to the rough localization measure, and the Fine Mask RR-Generator produces the final prediction of manipulated regions. We also utilize the Rotated Residual (RR) structure to suppress the image content during the generative process. The extensive experimental results on four benchmark data sets (NIST16, COVER, CASIA v1.0, and In-The-Wild) demonstrate the superior performance of PR-Net compared with the state-of-the-art methods in localizing the manipulated regions. Zenan Shi, Chaoqun Chang, Haipeng Chen 0002, Xiaoyu Du 0002, Hanwang Zhang |
Int. J. Intell. Syst. | 1 |
| 2022 | Hybrid features and semantic reinforcement network for image forgery detection
Haipeng Chen 0002, Chaoqun Chang, Zenan Shi, Yingda Lyu |
Multim. Syst. | 3 |
| 2020 | Global Semantic Consistency Network for Image Manipulation DetectionabstractThis letter focuses on image manipulation detection which aims to recognize the manipulated regions under the contextual semantic information. Existing approaches usually overlook the semantic discrepancy between different levels of feature maps, and directly fuse (e.g., addition, or concatenation) them for detection. In this letter, we argue that the semantic gap is the main reason for the low effectiveness of feature fusion in manipulation predictions. To address this problem, we propose a Global Semantic Consistency Network (GSCNet) for image manipulation detection, which is based on an encoder-decoder structure. Specifically, to make GSCNet include more global texture information which has been empirically confirmed to be beneficial to manipulation detection, gram block is first deployed on each level of feature maps in the encoding stage. Based on that, bi-directional convolutional LSTM is further implemented on the decoding stage, such that feature maps of the same level have semantic consistency. Experimental results on NIST16, and CASIA v1.0 declare that GSCNet can accurately locate the manipulated regions. Furthermore, compared to the existing models, GSCNet can achieve new state-of-the-art results. Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu |
IEEE Signal Process. Lett. | 1 |
| 2018 | Image splicing detection based on Markov features in discrete octonion cosine transform domainabstractTo improve the poor robustness and low accuracy of the existing algorithms of image splicing detection, a novel passive image forgery detection method is proposed in this study, which is based on DOCT (discrete octonion cosine transform) and Markov. By introducing the octonion and DOCT, the colour information of six image channels (the RGB model and the HSI model) can be exhaustively extracted, which enhances the robustness of the algorithm. On the issue of improving the detection accuracy, the standard deviation is used to characterise the relationship of the colour information between the parts of DOCT coefficient matrix, and the K ‐fold cross‐validation is introduced to improve the identification performance of the classifier. The steps of the algorithm are as follows: Firstly, the 8 × 8 block DOCT transform is used to the original image to obtain parts of block DOCT coefficient. Secondly, the standard deviation is used to process the corresponding parts of all blocks of the image. Finally, the Markov feature vector of the DOCT coefficient is extracted and feds to the LIBSVM (a library for support vector machines). When using LIBSVM for classification, K ‐fold cross‐validation is executed to select the best parameter pairs. The experiment results demonstrate that the algorithm is superior to the other state‐of‐the‐art splicing detection methods. Hongda Sheng, Xuanjing Shen, Yingda Lyu, Zenan Shi, Shuyang Ma |
IET Image Process. | 4 |
| 2017 | Splicing image forgery detection using textural features based on the grey level co-occurrence matricesabstractTo further improve the detection rate with relatively low dimension feature vector, a novel passive splicing detection method using textural features based on the grey level co‐occurrence matrices, namely TF‐GLCM, is proposed in this study. In the TF‐GLCM, the GLCM are calculated based on the difference block discrete cosine transform arrays to capture the textural information and the spatial relationship between image pixels sufficiently. The discriminable properties contained in the GLCM are described by six textural features, which include two new introduced ones and four independent ones. In addition, the statistical moments mean Me and standard deviation SD of textural features are used instead of themselves as elements in feature vector to reduce the dimensionality of feature vector and computational complexity. A support vector machine is employed for classification purpose. Experimental results show that the TF‐GLCM achieves the detection rates of 98% on CASIA v1.0, and 97% on CASIA v2.0 with 96‐D feature vector. The detection rates benefit from the two new textural features. Meanwhile, the TF‐GLCM is superior to some state‐of‐the‐art methods with lower dimension feature vector. Xuanjing Shen, Zenan Shi, Haipeng Chen 0002 |
IET Image Process. | 2 |