VLDB 2026 Research / reviewers in the wild / expert
Yun Cao 0001
dblp:40/2407-1
· DBLP profile ↗
51ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0003-3433-0764ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 30 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group AttentionabstractVision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual models. However, these approaches face high training and inference costs, as well as challenges in extracting visual details, effectively bridging across modalities. In this work, we propose a novel visual framework, MoCHA, to address these issues. Our framework integrates four vision backbones (i.e., CLIP, SigLIP, DINOv2 and ConvNeXt) to extract complementary visual features and is equipped with a sparse Mixture of Experts Connectors (MoECs) module to dynamically select experts tailored to different visual dimensions. To mitigate redundant or insufficient use of the visual information encoded by the MoECs module, we further design a Hierarchical Group Attention (HGA) with intra- and inter-group operations and an adaptive gating strategy for encoded visual features. We train MoCHA on two mainstream LLMs (e.g., Phi2-2.7B and Vicuna-7B) and evaluate their performance across various benchmarks. Notably, MoCHA outperforms state-of-the-art open-weight models on various tasks. For example, compared to CuMo (Mistral-7B), our MoCHA (Phi2-2.7B) presents outstanding abilities to mitigate hallucination by showing improvements of 3.25% in POPE and to follow visual instructions by raising 153 points on MME. Finally, ablation studies further confirm the effectiveness and robustness of the proposed MoECs and HGA in improving the overall performance of MoCHA. Yuqi Pang, Yun Cao 0001, Fan Rong |
AAAI | 3 |
| 2026 | CIF: A Constrained Inversion Framework for Reliable Message Extraction in Diffusion-Based Generative SteganographyabstractGenerative image steganography aims to conceal secret information in generated images without arousing suspicion. However, in practical scenarios involving high-capacity embedding or lossy transmission, existing methods still suffer from limited extraction accuracy. The main challenge lies in accurately recovering the secret-embedded latent vectors from stego images. To address this issue, we propose CIF, a constrained inversion framework designed to achieve accurate message extraction. Specifically, CIF mitigates dynamic structural errors by enforcing path consistency in the latent space and reduces numerical integration errors by adaptively selecting the integration order according to local trajectory stability. Experimental results show that our method reduces latent reconstruction error by more than 35% and achieves higher message extraction accuracy than existing approaches. Yuqi Qian, Yun Cao 0001, Meiyang Lv, Haocheng Fu |
IH&MMSec | 2 |
| 2026 | Temporal modality reliability and uncertainty-aware alignment for text-video retrievalabstractAbstract The rapid growth of multimodal content has introduced new challenges in cybersecurity, particularly in scenarios such as misinformation detection, multimedia forensics, and open-source intelligence. In these settings, verifying the consistency between textual descriptions and video content is critical, yet remains challenging due to noisy, incomplete, or even misleading multimodal signals. Text-to-video retrieval provides a important capability for cross-modal alignment. However, existing approaches often overlook a key issue: the usefulness of different modalities varies over time and depends on the query. In particular, audio signals can be informative in certain segments (e.g., speech) but misleading in others (e.g., background noise), making uniform fusion unreliable in security-critical scenarios. In this work, we propose Temporal Uncertainty-aware Retrieval (TUR), a unified framework that models text-video alignment from two complementary perspectives: temporal modality reliability and alignment uncertainty. TUR dynamically estimates the contribution of multimodal signals over time and adapts text representations according to cross-modal agreement, enabling more stable retrieval. Extensive experiments on MSR-VTT, DiDeMo, VATEX, and LSMDC demonstrate that TUR consistently outperforms prior methods. Further analysis shows that TUR achieves improved temporal grounding, more stable similarity estimation, and enhanced interpretability, which are desirable properties for security-sensitive applications. Yun Cao 0001, Hong Zhang 0005 |
Cybersecur. | 2 |
| 2025 | Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal ReasoningabstractAlthough Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large Language Models (MLLMs), however, is resource-intensive and constrained by various training limitations. In this paper, we propose the Modular-based Visual Contrastive Decoding (MVCD) framework to move this obstacle. Our framework leverages LLMs’ In-Context Learning (ICL) capability and the proposed visual contrastive-example decoding (CED), specifically tailored for this framework, without requiring any additional training. By converting visual signals into text and focusing on contrastive output distributions during decoding, we can highlight the new information introduced by contextual examples, explore their connections, and avoid over-reliance on prior encoded knowledge. MVCD enhances LLMs’ visual perception to make it see and reason over the input visuals. To demonstrate MVCD’s effectiveness, we conduct experiments with four LLMs across five question answering datasets. Our results not only show consistent improvement in model accuracy but well explain the effective components inside our decoding strategy. Our code will be available at https://github.com/Pbhgit/MVCD. Yuqi Pang, Haoqin Tu, Yun Cao 0001 |
ICASSP | 4 |
| 2025 | Object-Based Video Tampering Localization via Trace Consistency AnalysisabstractWith the rapid advancement of object-based video inpainting and splicing tampering techniques, the dissemination of malicious videos on the internet poses significant risks. Existing localization methods, however, exhibit limitations such as restriction to specific datasets, limited performance in detecting unknown forgeries, and lack of robustness when dealing with reprocessed videos. In this paper, we propose an effective video tampering localization network that comprehensively considers the common characteristics of inpainting and splicing and extracts more generalized features of forgery traces, significantly enhancing localization performance. Specifically, we design four modules to independently extract inherent difference features, including edge artifacts, pixel distribution, texture features, and frequency information. For feature fusion learning, we employ a two-stage approach: first, a Convolutional Neural Network (CNN)-based module is used to extract local features; then, a Vision Transformer (ViT)-based module is utilized to extract global correlation features. Experimental results demonstrate that our method significantly outperforms existing state-of-theart methods. Ablation studies verify the necessity of each feature component, and the two-stage feature fusion approach shows better performance compared to directly embedding a CNN into a ViT hybrid architecture. Pengfei Pei, Yun Cao 0001, Jinchuan Li, Yuqi Pang |
ICASSP | 2 |
| 2025 | Towards High-Capacity Provably Secure Steganography via Cascade Sampling
Meiyang Lv, Haocheng Fu, Xiaowei Yi, Hongxian Huang, Yun Cao 0001 |
ICICS (3) | 5 |
| 2025 | DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion ModelabstractRecent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk. Kangwei Liu 0003, Junwu Liu, Yun Cao 0001, Jinlin Guo, Xiaowei Yi |
ICME | 3 |
| 2025 | Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal SpaceabstractAudio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for comprehensive emotion manipulation, and (2) deterministic regression-based mapping that constrains the stochastic nature of emotional expressions and non-verbal behaviors, limiting the expressiveness of synthesized animations. To address these challenges, we present a diffusion-based framework for controllable expressive 3D facial animation. Our approach introduces two key innovations: (1) a FLAME-centered multimodal emotion binding strategy that aligns diverse modalities (text, audio, and emotion labels) through contrastive learning, enabling flexible emotion control from multiple signal sources, and (2) an attention-based latent diffusion model with content-aware attention and emotion-guided layers, which enriches motion diversity while maintaining temporal coherence and natural facial dynamics. Extensive experiments demonstrate that our method outperforms existing approaches across most metrics, achieving a 21.6% improvement in emotion similarity while preserving physiologically plausible facial dynamics. Project Page: https://kangweiiliu.github.io/Control_3D_Animation. Kangwei Liu 0003, Junwu Liu, Xiaowei Yi, Jinlin Guo, Yun Cao 0001 |
ICME | 5 |
| 2025 | Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for ForensicsabstractThe rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect.While multimodal Large Language Models (LLMs) have encoded rich world knowledge, they are not inherently tailored for combating AI-generated Content (AIGC) and struggle to comprehend local forgery details.In this work, we investigate the application of multimodal LLMs in forgery detection.We propose a framework capable of evaluating image authenticity, localizing tampered regions, providing evidence, and tracing generation methods based on semantic tampering clues.Our method demonstrates that the potential of LLMs in forgery analysis can be effectively unlocked through meticulous prompt engineering and the application of fewshot learning techniques.We conduct qualitative and quantitative experiments and show that GPT4V can achieve an accuracy of 92.1% in Autosplice and 86.3% in LaMa, which is competitive with state-of-the-art AIGC detection methods.We further discuss the limitations of multimodal LLMs in such tasks and propose potential improvements. Yun Cao 0001 |
IH&MMSec | 2 |
| 2025 | Triple-Stage Robust Audio Steganography Framework with AAC Encoding for Lossy Social Media ChannelsabstractRobust audio steganography has significant application value for secure communication, especially with the rise of social media platforms.However, the complexity of audio encoding and the distortions introduced by lossy channels have hindered the research in this field.This paper systematically analyzes the origins of this challenge, and evaluates the limitations of previous methods.Building on this foundation, we propose a triple-stage Robust Audio Steganography Framework (RASF), specifically designed for AAC encoding process.RASF consists three essential stages: Psy-Window Control to synchronize psychoacoustic model parameters, Robust Embedding Domain Construction to establish a robust embedding domain using stable quantized coefficients, and Error Correction to ensure reliable data recovery.Experiments demonstrate that the proposed framework achieves high capacity and strong robustness against compression.Notably, tests conducted on social media platforms reveal a very low bit error rate, enabling zero-bit-error transmission when combined with error-correcting codes.RASF addresses critical gaps in robust audio steganography, offering a practical solution for covert communication over lossy social media channels. Ziping Zhang, Jiamin Zeng, Xiaowei Yi, Yun Cao 0001 |
IH&MMSec | 5 |
| 2025 | ExpFormer: Cross-lingual One-shot Talking Head Generation via Enhanced 3D Expression ModelingabstractOne-shot audio-driven talking head generation suffers from significant limitations, primarily due to the complex relationship between speech and facial dynamics. Existing methods exhibit two major issues: (1) unrealistic expressions due to noisy 3D estimations in the lip region, and (2) inconsistent speaking styles, particularly evident in cross-lingual scenarios. To address these challenges, we propose ExpFormer, a novel 3DMM-based framework with two key innovations: (1) a lip movement enhancement strategy that reinforces lip region saliency during training, effectively mitigating 3D estimation errors in the lip region while preserving identity consistency, and (2) a transformer-based architecture with periodic positional encoding that captures both fine-grained lip synchronization and long-term speaking patterns, enabling natural facial animations across languages. To systematically evaluate cross-lingual generalization, we introduce a Mandarin Chinese dataset. Extensive experiments demonstrate that EXPFORMER significantly outperforms existing methods in visual quality, identity preservation, and lip synchronization across five different languages while achieving real-time performance. Kangwei Liu 0003, Xiaowei Yi, Junwu Liu, Yun Cao 0001 |
IJCNN | 4 |
| 2025 | DiffEmotionVC: A Dual-Granularity Disentangled Diffusion Framework for Any-to-Any Emotional Voice Conversion
Xiaosu Su, Xiaowei Yi, Yun Cao 0001 |
INTERSPEECH | 4 |
| 2025 | Text-Vision Embedding for Generalized Diffusion Generated Videos Detection
Jinchuan Li, Jinlin Guo, Yun Cao 0001, Kangwei Liu 0003 |
PRCV (13) | 3 |
| 2025 | DA-CCQ: Visually Explainable Image Forgery Localization via Difference Amplification and Cross-Clue Querying
Jinlin Guo, Yun Cao 0001, Jinchuan Li, Chengcheng Ma |
PRCV (12) | 3 |
| 2025 | NDCA: a neighboring block differences-based cost assignment method for robust video steganography on social networksabstractSocial network-based covert communication conceals the communication link between the sender and receiver, enabling one-to-many communication. Videos, due to their rich content and high embedding capacity, are ideal carriers for steganographic techniques. However, social networks typically apply lossy processing to uploaded videos, presenting significant challenges in constructing reliable covert communication channels. While prior research has proposed robust video steganographic methods, these approaches often rely on synchronization of robust regions to correctly extract hidden data. A major challenge arises when synchronization information is altered during lossy processing, complicating the accurate extraction of hidden data. To address this, a robust video steganographic framework is proposed. We then analyze the factors influencing the robustness of embedding units, including neighboring block differences, modulation types, and rate control modes. Based on this analysis, we introduce the Neighboring block Differences-based Cost Assignment (NDCA) method. Extensive experiments are conducted to demonstrate that the proposed framework and NDCA enhance robustness against lossy processing while maintaining high steganographic security. Furthermore, the robust video steganographic techniques based on the proposed framework and NDCA are broadly applicable to commonly used video encoders and rate control modes, enabling reliable covert communication on mainstream social networks. Hong Zhang 0005, Xinrui Xie, Yun Cao 0001 |
EURASIP J. Inf. Secur. | 4 |
| 2025 | Video Steganography With Optimized Robust Modulation Paths for Lossy ChannelsabstractSocial networks provide an ideal channel for covert communication due to their one-to-many broadcasting nature and the concealment of communication links. Videos, with their rich content and high embedding capacity, serve as suitable carriers for steganography. However, video transcoding performed by social networks often invalidates traditional steganographic methods. To address this challenge, we propose a novel frame work based on optimized robust modulation paths. Specifically, we analyze the influence of modulation types on the robustness of embedding units, introduce a cost assignment method to quantify the embedding impact, and develop an optimization strategy to identify robust modulation paths. Experimental results demon strate that the proposed method achieves an average bit error rate below 0.5% across mainstream social networks, outperforming state-of-the-art methods in terms of robustness while maintaining sufficient steganographic security. Hong Zhang 0005, Yun Cao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | OC-SAN: Unsupervised Deepfake Detection for Specific Individual Protection Based on Deep One-Class Classification
Yun Cao 0001, Yanfei Tong, Xin Liao 0001, Meineng Zhu |
PRCV (10) | 1 |
| 2024 | Unveiling tampering traces: Enhancing image reconstruction errors for visualization
Xianfeng Zhao, Yun Cao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Temporal Diversified Self-Contrastive Learning for Generalized Face Forgery DetectionabstractFace forgery detection receives widespread attention due to the great security threats arising from the development of face forgery technologies. Most existing works define it as a binary classification problem by modeling the spatial and temporal artifacts to distinguish real and fake videos. However, the detector tends to heavily rely on the binary labels and overfit method-specific forgery patterns of the training set, resulting in limited generalization ability. To mitigate this issue, we propose a Temporal Diversified Self-Contrastive Learning (TDSCL) framework, which guides the model to exploit generalized temporal inconsistencies for face forgery detection. Firstly, a Temporally Diversified Transformation (TDT) strategy is designed to create diverse training samples with multiple temporal scales. Subsequently, Short-term Self-contrastive Learning (STSC) and Long-term Self-contrastive Learning (LTSC) are proposed to perform temporal representations of the video at different temporal granularities to capture intrinsic and generalized forensics clues to expose fake videos, which can serve as auxiliary supervisions equipped with different backbones flexibly. Moreover, a Similarity-Guided Adaptive Fusion (SGAF) module is designed to adaptively reinforce the temporal inconsistencies for reliable classification. Extensive experiments verify that the proposed method achieves superior generalization ability over various state-of-the-art methods in different benchmark datasets. Rongchuan Zhang, Peisong He, Haoliang Li, Shiqi Wang 0001, Yun Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Generalizable Deep Video Inpainting Detection Based on Constrained Convolutional Neural Networks
Jinchuan Li, Xianfeng Zhao, Yun Cao 0001 |
IWDW | 3 |
| 2023 | VIFST: Video Inpainting Localization Using Multi-view Spatial-Frequency Traces
Pengfei Pei, Xianfeng Zhao, Jinchuan Li, Yun Cao 0001 |
PRICAI (3) | 4 |
| 2023 | Generalized Fake Image Detection Method Based on Gated Hierarchical Multi-Task LearningabstractRecently, the abuse of image generation techniques based on artificial intelligence has posed a great threat to the integrity of digital images. However, existing detection methods are hard to provide generalized detection capability of fake images generated by unseen models. To address this issue, we propose a generalized fake image detection framework based on gated hierarchical multi-task learning, which is supervised by well-designed forensics sub-tasks. Firstly, a global artifact learning task is constructed as binary classification with region masking augmentation. Besides, a block-wise spatial correlation learning task is designed by solving jigsaw puzzle cooperated with color jitter operations, which aims to explore common artifacts of various generators. Finally, a hierarchical multi-task learning paradigm is developed with multi-gate structures, which can adjust the importance of different forensics clues and jointly enhance detection performance. Extensive experiments have been conducted to evaluate the superiority of the proposed method on the open-set scenario with unseen generators Yanjiang Zhou, Peisong He, Weichuang Li, Yun Cao 0001, Xinghao Jiang |
IEEE Signal Process. Lett. | 4 |
| 2023 | Forensic Symmetry for DeepFakesabstractIn this paper, we propose a new DeepFakes forensics approach called forensic symmetry, which determines whether two symmetrical face patches contain the same or different natural features. To do this, we propose a multi-stream learning structure composed of two feature extractors. The first feature extractor obtains symmetry feature from the front face images. The second feature extractor obtains similarity feature from the side face images. Symmetry feature and similarity feature are collectively called natural feature. Forensic symmetry system maps the pair of symmetrical face patches into the angular hyperspace to quantify the difference of their natural features. The greater the difference of natural features, the higher the tamper probability of face images. The heuristic prediction algorithm is designed to compute the tamper probability of DeepFakes at video level. A series of experiments are carried out to evaluate the effectiveness of our proposed forensic symmetry system. Experimental results show that our approach is effective for DeepFakes detection under the scenarios of homologous detection, heterogeneous detection, and re- compression detection. Xianfeng Zhao, Yun Cao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | FMFCC-V: An Asian Large-Scale Challenging Dataset for DeepFake DetectionabstractThe abuse of DeepFake technique has raised enormous public concerns in recent years. Currently, the existing DeepFake datasets suffer some weaknesses of obvious visual artifacts, minimal Asian proportion, backward synthesis methods and short video length. To make up these weaknesses, we have constructed an Asian large-scale challenging DeepFake dataset to enable the training of DeepFake detection models and organized the accompanying video track of the first Fake Media Forensics Challenge of China Society of Image and Graphics (FMFCC-V). The FMFCC-V dataset is by far the first and the largest public available Asian dataset for DeepFake detection, which contains 38102 DeepFake videos and 44290 pristine videos, corresponding more than 23 million frames. The source videos in the FMFCC-V dataset are carefully collected from 83 paid individuals and all of them are Asians. The DeepFake videos are generated by four of the most popular face swapping methods. Extensive perturbations are applied to obtain a more challenging benchmark of higher diversity. The FMFCC-V dataset can lend powerful support to the development of more effective DeepFake detection methods. We contribute a comprehensive evaluation of six representative DeepFake detection methods to demonstrate the level of challenge posed by FMFCC-V dataset. Meanwhile, we provide a detailed analysis of the top submissions from the FMFCC-V competition. Xianfeng Zhao, Yun Cao 0001, Pengfei Pei, Jinchuan Li |
IH&MMSec | 3 |
| 2022 | Manipulated Face Detection and Localization Based on Semantic Segmentation
Xianfeng Zhao, Yun Cao 0001, Chengqiao Hu |
IWDW | 3 |
| 2022 | Visual Explanations for Exposing Potential Inconsistency of Deepfakes
Pengfei Pei, Xianfeng Zhao, Yun Cao 0001, Chengqiao Hu |
IWDW | 3 |
| 2021 | Exploiting Facial Symmetry to Expose DeepfakesabstractIn this paper, we introduce a new approach to detect synthetic portrait images and videos. Motivated by the observation that the symmetry of synthetic facial area would be easily broken, this approach aims to reveal the tampering trace by features learned from symmetrical facial regions. To do so, a two-stream learning framework is designed which uses a hard sharing Deep Residual Networks as the backbone network. The feature extractor maps the pair of symmetrical face patches to an angular distance indicating the difference of symmetry features. Extensive experiments are carried out to test the effectiveness in detecting synthetic portrait images and videos, and corresponding results show that our approach is effective even on heterogeneous data and re-compression data that were not used to train the detection model. Yun Cao 0001, Xianfeng Zhao |
ICIP | 2 |
| 2021 | A Multi-level Feature Enhancement Network for Image Splicing Localization
Yun Cao 0001, Xianfeng Zhao |
IWDW | 2 |
| 2021 | GIFMarking: The robust watermarking for animated GIF based deep learning
Xin Liao 0001, Yun Cao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Steganalysis of H.264/AVC Videos Exploiting Subtractive Prediction Error BlocksabstractTo cope with the abuse of steganography using H.264 videos, i.e., the dominant video format, as the carrier, this paper presents a steganalytic method which works well even in the scenario where both the training data and the prior knowledge of the test data are limited. As a key feature of H.264, intra prediction is incorporated to remove redundancies within one single frame by predicting the current block using previously coded blocks. Unlike in JPEG domain, the quantized discrete cosine transform (QDCT) coefficients in H.264 videos come from the prediction error (residual) blocks (PEBs) instead of the original pixel block, hence we suggest shifting the focal point from the spatial domain to the prediction error domain, i.e., the PEB domain. According to the traits of video coding, 3 types of subtractive PEB (SPEB) are defined to capture the inconsistency between correlated PEBs, and the differences between correlated SPEBs are modeled by first-order Markov chain. Then the so-called SUPERB (SUbtractive Prediction ERror Block) features are engineered by subsets of sample transition probability matrices for a steganalyzer. What's more, the features derived from IPM (Intra Prediction Mode) transition probabilities are also merged into SUPERB to improve detection ability. Extensive experiments are carried out from different aspects. Performance results demonstrate the effectiveness of SUPERB, particularly its essence of general applicability when the training and test data are of quite different attributes, which is more favorable for real-world applications. Yun Cao 0001, Hong Zhang 0005, Xianfeng Zhao, Xiaolei He |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Minimizing Embedding Impact for H.264 Steganography by Progressive Trellis CodingabstractThis paper proposes a novel coding strategy to achieve distortion minimization for H.264 steganography with quantized discrete cosine transform (QDCT) coefficients. Currently, with the help of syndrome-trellis codes (STCs), state-of-the-art image steganography embeds messages while minimizing a heuristically defined distortion function. However, this concept cannot be directly ported to steganography using compressed video as the cover media. According to the intra prediction principle, an H.264 QDCT coefficient block is predicted and coded based on previously encoded blocks, so even a slight embedding change will set off a chain reaction in the remaining cover blocks. Considering the cover block dependency, we make necessary changes to the standard trellis coding structure so as to be applicable for the joint compression embedding scenario. During the coding/embedding procedure, we maintain multiple contexts corresponding to possible optimal routes, and retrace each route periodically to determine how each cover block should be modified. After each modification, the remaining cover blocks, as well as their embedding costs, are re-evaluated, and each context is updated to reflect the embedding effect. In this way, the global optimality can be approached progressively in a block-by-block manner, so our proposed method is named progressive trellis coding (PTC). Extensive experiments have been conducted, and corresponding results show that the adoption of PTC brings about a significant gain in embedding performance. Yu Wang 0114, Yun Cao 0001, Xianfeng Zhao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | CEC: Cluster Embedding Coding for H.264 SteganographyabstractIn this letter, we first propose a novel coding method to design content-adaptive H.264 steganography with quantized discrete cosine transform (QDCT) coefficients. Currently, the state-of-the-art steganographic methods minimize a heuristically defined distortion function using STCs and embed messages in a pixel-wise manner. When this concept is directly applied to H.264 steganographic designing, there are two factors that hinder the improvement of empirical security: interaction of embedding changes, and the block-wise characteristics of video compression. Our proposed CEC scheme performs both message embedding and video compression in the block-wise manner where a cluster of elements in a cover block are modified simultaneously, and defines a joint distortion function on the cover block to reflect the features of rate-distortion optimization. Experimental results demonstrate the effectiveness of our proposed coding method. Yu Wang 0114, Yun Cao 0001, Xianfeng Zhao |
IEEE Signal Process. Lett. | 2 |
| 2019 | Light Multiscale Conventional Neural Network for MP3 Steganalysis
Jinghong Zhang, Xiaowei Yi, Xianfeng Zhao, Yun Cao 0001 |
IWDW | 4 |
| 2019 | Adversarial Learning for Constrained Image Splicing Detection and Localization Based on Atrous ConvolutionabstractConstrained image splicing detection and localization (CISDL), which investigates two input suspected images and identifies whether one image has suspected regions pasted from the other, is a newly proposed challenging task for image forensics. In this paper, we propose a novel adversarial learning framework to learn a deep matching network for CISDL. Our framework mainly consists of three building blocks. First, a deep matching network based on atrous convolution (DMAC) aims to generate two high-quality candidate masks, which indicate suspected regions of the two input images. In DMAC, atrous convolution is adopted to extract features with rich spatial information, a correlation layer based on a skip architecture is proposed to capture hierarchical features, and atrous spatial pyramid pooling is constructed to localize tampered regions at multiple scales. Second, a detection network is designed to rectify inconsistencies between the two corresponding candidate masks. Finally, a discriminative network drives the DMAC network to produce masks that are hard to distinguish from ground-truth ones. The detection network and the discriminative network collaboratively supervise the training of DMAC in an adversarial way. Besides, a sliding window-based matching strategy is investigated for high-resolution images matching. Extensive experiments, conducted on five groups of datasets, demonstrate the effectiveness of the proposed framework and the superior performance of DMAC. Xiaobin Zhu 0001, Xianfeng Zhao, Yun Cao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Cover Block Decoupling for Content-Adaptive H.264 SteganographyabstractThis paper makes the first attempt to achieve content-adaptive H.264 steganography with the quantised discrete cosine transform (QDCT) coefficients in intra-frames. Currently, state-of-the-art JPEG steganographic schemes embed their payload while minimizing a heuristically defined distortion. However, porting this concept to schemes of compressed videos remains an unsolved challenge. Because of H.264 intra prediction, the QDCT coefficient blocks are highly depended on their adjacent encoded blocks, and modifying one coefficient block will set off a chain reaction in the following cover blocks. Based on a thorough investigation into this problem, we propose two embedding strategies for cover block decoupling to inhibit the embedding interactions. With this methodology, the latest achievements in the JPEG domain are expected to be incorporated to construct H.264 steganographic schemes for better performances. Yun Cao 0001, Yu Wang 0114, Xianfeng Zhao, Meineng Zhu, Zhoujun Xu |
IH&MMSec | 1 |
| 2018 | Image Forgery Localization based on Multi-Scale Convolutional Neural NetworksabstractIn this paper, we propose to utilize Convolutional Neural Networks (CNNs) and the segmentation-based multi-scale analysis to locate tampered areas in digital images. First, to deal with color input sliding windows of different scales, we adopt a unified CNN architecture. Then, we elaborately design the training procedures of CNNs on sampled training patches. With a set of tampering detectors based on CNNs for different scales, a series of complementary tampering possibility maps can be generated. Last but not least, a segmentation-based method is proposed to fuse these maps and generate the final decision map. By exploiting the benefits of both the small-scale and large-scale analyses, the segmentation-based multi-scale analysis can lead to a performance leap in forgery localization of CNNs. Numerous experiments are conducted to demonstrate the effectiveness and efficiency of our method. Qingxiao Guan, Xianfeng Zhao, Yun Cao 0001 |
IH&MMSec | 4 |
| 2018 | Maintaining Rate-Distortion Optimization for IPM-Based Video Steganography by Constructing Isolated Channels in HEVCabstractThis paper proposes an effective intra-frame prediction mode (IPM)-based video steganography in HEVC to maintain rate-distortion optimization as well as improve empirical security. The unique aspect of this work and one that distinguishes it from prior art is that we capture the embedding impacts on neighboring prediction units, called inter prediction unit (inter-PU) embedding impacts caused by the predictive coding widespread employed in video coding standards, using a distortion measure. To avoid the emergence of neighboring IPMs mutually affecting each other within the same channel, three-layered isolated channels are established in terms of the property of IPM coding. According to theoretical analysis for embedding impacts on the current prediction unit, called intra prediction unit (intra-PU) embedding impacts on coding efficiency (both visual quality and compression efficiency), a novel distortion function purposely designed to discourage the embedding changes with impacts on adjacent channels is proposed to express the multi-level embedding impacts. Based on the defined distortion function, two-layered syndrome-trellis codes (STCs) are utilized in practical embedding implementation alternatively. Experimental results demonstrate that the proposed scheme outperforms other existing IPM-based video steganography in terms of rate-distortion optimization and empirical security. Yu Wang 0114, Yun Cao 0001, Xianfeng Zhao, Zhoujun Xu, Meineng Zhu |
IH&MMSec | 2 |
| 2017 | A Steganalytic Algorithm to Detect DCT-based Data Hiding Methods for H.264/AVC VideosabstractThis paper presents an effective steganalytic algorithm to detect Discrete Cosine Transform (DCT) based data hiding methods for H.264/AVC videos. These methods hide covert information into compressed video streams by manipulating quantized DCT coefficients, and usually achieve high payload and low computational complexity, which is suitable for applications with hard real-time requirements. In contrast to considerable literature grown up in JPEG domain steganalysis, so far there is few work found against DCT-based methods for compressed videos. In this paper, the embedding impacts on both spatial and temporal correlations are carefully analyzed, based on which two feature sets are designed for steganalysis. The first feature set is engineered as the histograms of noise residuals from the decompressed frames using 16 DCT kernels, in which a quantity measuring residual distortion is accumulated. The second feature set is designed as the residual histograms from the similar blocks linked by motion vectors between inter-frames. The experimental results have demonstrated that our method can effectively distinguish stego videos undergone DCT manipulations from clean ones, especially for those of high qualities. Yun Cao 0001, Xianfeng Zhao, Meineng Zhu |
IH&MMSec | 2 |
| 2017 | A Prediction Mode-Based Information Hiding Approach for H.264/AVC Videos Minimizing the Impacts on Rate-Distortion Optimization
Yu Wang 0114, Yun Cao 0001, Xianfeng Zhao, Linna Zhou |
IWDW | 2 |
| 2017 | Information Hiding Using CAVLC: Misconceptions and a Detection Strategy
Weike You, Yun Cao 0001, Xianfeng Zhao |
IWDW | 2 |
| 2017 | Segmentation Based Video Steganalysis to Detect Motion Vector ModificationabstractThis paper presents a steganalytic approach against video steganography which modifies motion vector (MV) in content adaptive manner. Current video steganalytic schemes extract features from fixed-length frames of the whole video and do not take advantage of the content diversity. Consequently, the effectiveness of the steganalytic feature is influenced by video content and the problem of cover source mismatch also affects the steganalytic performance. The goal of this paper is to propose a steganalytic method which can suppress the differences of statistical characteristics caused by video content. The given video is segmented to subsequences according to block’s motion in every frame. The steganalytic features extracted from each category of subsequences with close motion intensity are used to build one classifier. The final steganalytic result can be obtained by fusing the results of weighted classifiers. The experimental results have demonstrated that our method can effectively improve the performance of video steganalysis, especially for videos of low bitrate and low embedding ratio. Yun Cao 0001, Xianfeng Zhao |
Secur. Commun. Networks | 2 |
| 2017 | A Steganalytic Approach to Detect Motion Vector Modification Using Near-Perfect Estimation for Local OptimalityabstractThis paper presents a steganalytic approach against motion vector-based video steganography that does not depend on the detailed knowledge of embedding algorithms. In most state-of-the-art video coding standards, the motion vector is the result of block-based motion estimation using rate-distortion optimization. That is to say, each motion vector is locally optimal in a rate-distortion sense, and any modification will inevitably shift the motion vector from locally optimal to non-optimal. As a consequence, it is a very strong evidence of steganography if some motion vectors are found to be locally non-optimal. Based on this fact, the core of our method is an estimator to check the local optimality of motion vectors in a rate-distortion sense. We try to recover the necessary information used for motion vector decision that is lost during lossy compression, based on which a 36-D feature set is formed for training and classification. To demonstrate the effectiveness of the proposed approach, experiments are carried out in different settings. The corresponding results show that our approach has a wide applicability even at low embedding strengths. Particularly, the problem of cover source mismatch is largely alleviated, which indicates that the proposed approach is suitable to be used in situations where a very limited priori knowledge is available. Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Video Steganalysis Based on Centralized Error Detection in Spatial Domain
Yu Wang 0114, Yun Cao 0001, Xianfeng Zhao |
Inscrypt | 2 |
| 2016 | A Novel Embedding Distortion for Motion Vector-Based Steganography Considering Motion Characteristic, Local Optimality and Statistical DistributionabstractThis paper presents an effective motion vector (MV)-based steganography to cope with different steganalytic models. The main principle is to define a distortion scale expressing the multi-level embedding impact of MV modification. Three factors including motion characteristic of video content, MV's local optimality and statistical distribution are considered in distortion definition. For every embedding location, the contributions of three factors are dynamically adjusted according to MV's property. Based on the defined distortion function, two layered syndrome-trellis codes (STCs) are utilized to minimize the overall embedding impact in practical embedding implementation. Experimental results demonstrate that the proposed method achieves higher level of security compared with other existing MV-based approaches, especially for high quality videos. Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao |
IH&MMSec | 3 |
| 2016 | Data Hiding in H.264/AVC Video Files Using the Coded Block Pattern
Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao |
IWDW | 2 |
| 2016 | Motion vector-based video steganography with preserved local optimality
Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao |
Multim. Tools Appl. | 2 |
| 2015 | An adaptive detecting strategy against motion vector-based steganographyabstractThe goal of this paper is to improve the performance of the current video steganalysis in detecting motion vector (MV)-based steganography. It is noticed that many MV-based approaches embed secret bits in content adaptive manners. Typically, the modifications are applied only to qualified MVs, which implies that the number of modified MVs varies among frames after embedding. On the other hand, nearly all the current steganalytic methods ignore such uneven distribution. They divide the video into frame groups equally and calculate every single feature vector using all MVs within one group. For better classification performances, we suggest performing steganalysis also in an adaptive way. First, divide the video into groups with variable lengths according to frame dynamics. Then within each group, calculate a single feature vector using all suspicious MVs (MVs that are likely to be modified). The experimental results have shown the effectiveness of our proposed strategy. Yun Cao 0001, Xianfeng Zhao |
ICME | 2 |
| 2015 | Video Steganography Based on Optimized Motion Estimation PerturbationabstractIn this paper, a novel motion vector-based video steganographic scheme is proposed, which is capable of withstanding the current best statistical detection method. With this scheme, secret message bits are embedded into motion vector (MV) values by slightly perturbing their motion estimation (ME) processes. In general, two measures are taken for steganographic security (statistical undetectability) enhancement. First, the ME perturbations are optimized ensuring the modified MVs are still local optimal, which essentially makes targeted detectors ineffective. Secondly, to minimize the overall embedding impact under a given relative payload, a double-layered coding structure is used to control the ME perturbations. Experimental results demonstrate that the proposed scheme achieves a much higher level of security compared with other existing MV-based approaches. Meanwhile, the reconstructed visual quality and the coding efficiency are slightly affected as well. Yun Cao 0001, Hong Zhang 0005, Xianfeng Zhao |
IH&MMSec | 1 |
| 2015 | Video Steganalysis Based on Intra Prediction Mode Calibration
Yanbin Zhao, Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao |
IWDW | 3 |
| 2014 | Video steganography with perturbed macroblock partitionabstractIn this paper, with a novel data representation named macroblock partition mode, an effective steganography integrated with H.264/AVC compression is proposed. The main principle is to improve the steganographic security in two directions. First, to embed messages, an internal process of H.264 compression, i.e., the macroblock partition, is slightly perturbed, hence the compression compliance is ensured. Second, to minimize the embedding impact, a high efficient double-layered structure is deliberately designed. In the first layer, the syndrome-trellis codes (STCs) is utilized to perform adaptive embedding, and the costs in visual quality and compression efficiency are both considered to construct the distortion model. In the second layer, facilitated by the wet paper codes (WPCs), an expected 3-bit per change gain in embedding efficiency is obtained. Hong Zhang 0005, Yun Cao 0001, Xianfeng Zhao, Weiming Zhang 0001, Nenghai Yu |
IH&MMSec | 2 |
| 2012 | Video Steganalysis Exploiting Motion Vector Reversion-Based FeaturesabstractUnlike traditional image or video steganography in spatial/transform domain, motion vector (MV)-based methods target the internal dynamics of video compression and embed messages while performing motion estimation. However, we have noticed that some existing methods adopt nonoptimal selection rules and modify MVs in somewhat arbitrary manners which violate the encoding principles a lot. Aiming at these weaknesses, we design a calibration-based approach and propose MV reversion-based features for steganalysis. Experimental results demonstrate that the proposed features are very sensitive to the tendency of MV reversion during calibration and can be used to effectively detect some typical MV-based steganography even with low embedding rates. Yun Cao 0001, Xianfeng Zhao, Dengguo Feng |
IEEE Signal Process. Lett. | 1 |