Ziqi Sheng

dblp:210/3527 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework
abstract
Image forgery localization aims to precisely identify tampered regions within images, but it commonly depends on costly pixel-level annotations. To alleviate this annotation burden, weakly supervised image forgery localization (WSIFL) has emerged, yet existing methods still achieve limited localization performance as they mainly exploit intra-image consistency clues and lack external semantic guidance to compensate for insufficient supervision information. In this paper, we propose ViLaCo, a vision-language collaborative reasoning framework that introduces auxiliary semantic supervision derived from pre-trained vision-language models (VLMs), enabling accurate pixel-level localization using only image-level labels. Specifically, we first employ a vision-language feature modeling network to jointly extract textual semantics and visual features by leveraging pre-trained VLMs. Next, an adaptive vision-language reasoning network aligns these features through mutual interactions, producing semantically aligned representations. Subsequently, these representations are passed into dual prediction heads, where the coarse head performs image-level classification and the fine head generates pixel-level localization masks, allowing the coarse-grained task to provide guidance for the fine-grained localization. Moreover, a contrastive patch consistency module is introduced to cluster tampered features while separating authentic ones, facilitating more reliable forgery discrimination. Extensive experiments on multiple public datasets demonstrate that ViLaCo substantially outperforms existing WSIFL methods, achieving state-of-the-art performance in both detection and localization accuracy.
Ziqi Sheng, Junyan Wu, Wei Lu 0001, Jiantao Zhou 0001
AAAI1
2025 SUMI-IFL: An Information-Theoretic Framework for Image Forgery Localization with Sufficiency and Minimality Constraints
abstract
Image forgery localization (IFL) is a crucial technique for preventing tampered image misuse and protecting social safety. However, due to the rapid development of image tampering technologies, extracting more comprehensive and accurate forgery clues remains an urgent challenge. To address these challenges, we introduce a novel information-theoretic IFL framework named SUMI-IFL that imposes sufficiency-view and minimality-view constraints on forgery feature representation. First, grounded in the theoretical analysis of mutual information, the sufficiency-view constraint is enforced on the feature extraction network to ensure that the latent forgery feature contains comprehensive forgery clues. Considering that forgery clues obtained from a single aspect alone may be incomplete, we construct the latent forgery feature by integrating several orthogonal individual image features. Second, based on the information bottleneck, the minimality-view constraint is imposed on the feature reasoning network to achieve an accurate and concise forgery feature representation that counters the interference of task-unrelated features. Extensive experiments show the superior performance of SUMI-IFL to existing state-of-the-art methods, not only on in-dataset comparisons but also on cross-dataset comparisons.
Ziqi Sheng, Wei Lu 0001, Xiangyang Luo 0001, Jiantao Zhou 0001, Xiaochun Cao
AAAI1
2025 RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News Detection
abstract
Multimodal fake news detection aims to automatically identify real or fake news, thereby mitigating the adverse effects caused by such misinformation. Although prevailing approaches have demonstrated their effectiveness, challenges persist in cross-modal feature fusion and refinement for classification. To address this, we present a residual-aware compensation network with multi-granularity constraints (RaCMC) for fake news detection, that aims to sufficiently interact and fuse cross-modal features while amplifying the differences between real and fake news. First, a multiscale residual-aware compensation module is designed to interact and fuse features at different scales, and ensure both the consistency and exclusivity of feature interaction, thus acquiring high-quality features. Second, a multi-granularity constraints module is implemented to limit the distribution of both the news overall and the image-text pairs within the news, thus amplifying the differences between real and fake news at the news and feature levels. Finally, a dominant feature fusion reasoning module is developed to comprehensively evaluate news authenticity from the perspectives of both consistency and inconsistency. Experiments on three public datasets, including Weibo17, Politifact and GossipCop, reveal the superiority of the proposed method.
Xinquan Yu, Ziqi Sheng, Wei Lu 0001, Xiangyang Luo 0001, Jiantao Zhou 0001
AAAI2
2025 Exploring multi-scale forgery clues for stereo super-resolution image forgery localization
Ziqi Sheng, Chengxi Yin, Wei Lu 0001
Pattern Recognit.1
2025 DiRLoc: Disentanglement Representation Learning for Robust Image Forgery Localization
abstract
Deep Learning image forgery localization methods have achieved remarkable results but cannot maintain comparable performance when the forgery images are JPEG compressed, a format that is widely used in daily information transmission. The robustness against JPEG compression has become a bottleneck to the practical application of image forgery localization. To address this issue, a robust image forgery localization framework is proposed against the performance degradation caused by JPEG compression. Specifically, a cutting-edge progressive disentanglement strategy is proposed that incorporates coarse-grained image disentanglement to mitigate the detrimental effects of general JPEG compression, while harnessing the ability of fine-grained element disentanglement to separate multi-scale artifacts, thereby minimizing interference from content information. Moreover, the decision strategy is carefully designed to reinforce subtle signals from tampered areas, including artifacts fusion block reasoning multi-scale artifacts and dual attention block that learn more about forgery-related features. Extensive visualizations and experiments demonstrate that our method can achieve competitive performance in general JPEG-resistant image forgery localization, especially in the performance of generalization experiments.
Ziqi Sheng, Zuomin Qu, Wei Lu 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.1
2024 Deep generative network for image inpainting with gradient semantics and spatial-smooth attention
Ziqi Sheng, Cong Lin 0003, Wei Lu 0001, Long Ye
J. Vis. Commun. Image Represent.1
2024 Audio Multi-View Spoofing Detection Framework Based on Audio-Text-Emotion Correlations
abstract
In recent years, audio spoofing detection has received widespread attention for protecting personal privacy and social security. Despite the significant progress achieved in audio single-view spoofing detection, challenges remain with regard to addressing unknown spoofing attacks in realistic scenarios. To solve these challenging problems, in this paper, we introduce a novel audio multi-view spoofing detection framework (AMSDF), whose goal is to capture both intra-view and inter-view cues by measuring correlations within audio multi-view features (i.e., audio-emotion-text) for audio spoofing detection. In general, different view features are inherently interconnected in the real patterns, while they may present unnatural correlations in the spoofing patterns. Therefore, more discriminative cues can be mined by utilizing their complex interactions, which is beneficial to the audio spoofing detection task. To this end, an intra-view graph attention mechanism (IGAM) is first utilized to aggregate each intra-view node within the same view. Subsequently, a heterogeneous graph fusion module (HGFM) is applied to measure correlations within inter-view nodes, which are enhanced with a master node for comprehensive analysis purposes. Finally, a group-based readout scheme (GRS) is designed to capture and preserve the most distinctive cues by leveraging the strengths of different feature sets, thereby effectively distinguishing subtle differences between real and spoofing audio. The experimental results show that our proposed framework can achieve better performance than that of the state-of-the-art methods, especially in realistic scenarios. The code and pre-trained models are available athttps://github.com/ItzJuny/AMSDF.
Junyan Wu, Qilin Yin, Ziqi Sheng, Wei Lu 0001, Jiwu Huang, Bin Li 0011
IEEE Trans. Inf. Forensics Secur.3
2023 An Interpretable Image Tampering Detection Approach Based on Cooperative Game
abstract
In order to reply the potential security issues caused by the tampering of digital images, many image forensics approaches based on deep learning have been proposed in recent years. However, the interpretability of deep learning-based approaches has not been fully considered. In this paper, an interpretable image tampering detection approach is proposed. It consists of suspicious tampered region detection (STRD) module and cooperative game module. The STRD module, inspired by YOLO, combines the shallow-level and deep-level features to discriminate different tampering types of suspicious tampered regions in complex scenes, and also performs well in detecting small tampered regions. The prediction of STRD module could be extended to suspicious box and the payoff of cooperative game module. The cooperative game module utilizes the Shapley interaction index as the strategy to measure the information gain of image pixels. The Shapley interaction index disentangles the multi-order interaction between the pixels of the image, and we discover that image tampering mainly affects low-order interaction of the image. The final detection result is obtained by combining the pixels that contribute greatly to the payoff with the suspicious box. The proposed approach provides a new thought for the interpretability of image forensics and could be broadly applied to other digital image forensic approaches. Extensive experimental results have demonstrated the proposed approach outperforms SOTA approaches, which also has good interpretability and robustness.
Wei Lu 0001, Ziqi Sheng
IEEE Trans. Circuits Syst. Video Technol.3
2017 A Novel Power Allocation Method for Non-orthogonal Multiple Access in Cellular Uplink Network
abstract
In this paper, we propose a new power allocation method on non-orthogonal multiple access (NOMA) schemes in cellular uplink network. It is known that NOMA could be a promising candidate as a wireless access scheme for future 5G radio access technology. To enhance the spectrum efficiency, NOMA adopts a successive interference cancellation (SIC) receiver as the baseline receiver scheme for robust multiple access. With the goal of maximum total throughput for all users and make full use of bandwidth, an effective power allocation method should be utilized, in the conventional algorithm, it is a good way to get a satisfied performance such as average power allocation, water-filling method. We propose a method based on particle swarm optimization with genetic algorithm, which makes the power allocation by taking the advantage of channel state information (CSI) for each sub-channel; and can obtain a better performance comparing with the conventional method.
Ziqi Sheng, Xin Su 0002, Xuewu Zhang 0001
Intelligent Environments1