EDBT 2026 Demo / reviewers in the wild / expert
Tanfeng Sun
dblp:52/2073 · also Tan-Feng Sun, TanFeng Sun
· DBLP profile ↗
88ranked-venue papers
2as first author
47since 2021 · last 2026
0000-0002-3253-5136ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 1 first-author · 21 since 2021Security and privacy · 25 · 12 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Does Adversarial Training Hurt Adversarial Robustness in Phase Translation of Training Data?
Ke Xu 0003, Xinghao Jiang, Tanfeng Sun, Zeyu Zhao 0006 |
ICIC (15) | 4 |
| 2026 | INN-RAE: Reversible adversarial examples based on invertible neural networks for facial protection
Zeyu Zhao 0006, Ke Xu 0003, Laijin Meng, Tanfeng Sun, Xinghao Jiang |
Expert Syst. Appl. | 4 |
| 2026 | Normalization-consistent data curation for generalizable deepfake detection
Shijie Hou, Xinghao Jiang, Ke Xu 0003, Qiang Xu 0007, Laijin Meng, Tanfeng Sun |
Neurocomputing | 6 |
| 2026 | Adaptive Learning With Augmentation Robustness Validation: Toward Generalizable Face Forgery DetectionabstractThe rapid advancement of facial manipulation technologies demands detection systems that can generalize to novel forgery techniques. We identify that Standard Learning, reliant on static datasets and uniform sampling, detrimentally biases models towards specific patterns tied to individual generation techniques, hindering their ability to learn general features. To overcome this, we introduce Adaptive Learning (AL) for face forgery detection, a cyclical framework that simultaneously refines both the detector model and the training data through dynamic sample selection and model optimization. AL’s efficacy hinges on identifying samples rich in generalizable forgery clues. Thus, we propose Augmentation Robustness Validation (ARV) as AL’s core purification engine. ARV exploits the stability of predictions across diverse semantic-preserving augmentations as a reliable proxy for general feature presence: samples that exhibit invariant predictions inherently contain robust manipulation traces. Integrating ARV with AL yields Adaptive Learning with Augmentation Robustness Validation (ALarv). ALarv strategically prioritizes stability-verified samples during iterative training cycles, progressively enhancing the model’s focus on transferable forensic features. Inspired by the architectural advantages of ConvNeXt, we incorporate it into ALarv, forming an effective method, ALarv-ConvNeXt. Extensive experiments demonstrate ALarv-ConvNeXt’s superior generalization performance, including emerging diffusion-based synthetic faces. Shijie Hou, Xinghao Jiang, Ke Xu 0003, Qiang Xu 0007, Tanfeng Sun, Laijin Meng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PSO-Based Closed Box Adversarial Patch Attack Against Face Recognition
Haotian Ma 0001, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Frame-Wise Detection of Fake Bitrate VVC Videos Based on Hierarchical Feature Mapping in Coding DomainabstractA prevalent video manipulation technique involves upsampling the bitrate parameter without modifying the underlying video content, often implemented under the pretext of improving viewer appeal and monetization potential. Such deceptive practices, which substitute fake bitrate specifications for genuine ones, mislead audiences and platforms while directly infringing copyright protections, constituting a form of manipulation that forensic experts classify as fake bitrate video manipulation. Addressing the issue of detecting fake bitrate Versatile Video Coding (VVC) videos, an algorithm based on hierarchical feature mapping in the coding domain is proposed. We first analyze the Coding Unit (CU) partitioning and the deblocking filtering of fake bitrate VVC videos during multiple encoding processes. Then, CU partitioning and deblocking decision-mode information in the coding domain are extracted during the decoding process. By combining the positional information, the CU Difference feature map (FCD) and the Deblocking Filtering Difference feature map (FDD) are obtained through a hierarchical mapping of the encoding feature, which are further enhanced by calculating the eccentric covariance matrix. After that, the enhanced feature maps are fed into a Dual-branch Difference Perception (DDP) module to obtain frame-level detection results. By comparing with existing algorithms, the experimental results demonstrate that the proposed algorithm achieves superior detection accuracy in different scenarios, validating its effectiveness and robustness. Qiang Xu 0007, Hao Wang 0247, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Toward Resisting Black-Box Attacks: A Robust Coverless Image Steganography Based on Hierarchical CID and Dual SIFT
Laijin Meng, Xinghao Jiang, Qiang Xu 0007, Zhongjie Mi, Shijie Hou, Tanfeng Sun |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | ResTNet: A ResNet-Transformer Network With Recompression Maps for Exposing Fake Bitrate Videos
Lizhi Xiong, Linsen Ding, Tanfeng Sun, Zhangjie Fu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | ASGA: Attention-Based Sparse Global Attack to Video Action Recognition
Zeyu Zhao 0006, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | SRAP: Robust and Transferable Self-Reversible Adversarial Patch for Image Privacy ProtectionabstractReversible adversarial examples offer adequate protection against malicious deep model identification and analysis. However, current methods still face challenges in terms of transferability and robustness, limiting their practical applicability. We introduce a novel technique for generating reversible adversarial examples utilizing Self-Reversible Adversarial Patch (SRAP) to address this. This approach significantly enhances the transferability and robustness of reversible adversarial examples against standard image processing techniques and adversarial defense methods. Specifically, we present a method for crafting adversarial patches that are small, non-overlapping, and adaptively integrated into specific regions. These adversarial patches are seamlessly combined with a reversible data-hiding technique that relies on prediction error expansion, resulting in adversarial examples with superior robustness and transferability. Experimental results indicate that our method achieves a remarkable transferability rate of up to 90% or higher between different models. Additionally, it exhibits strong robustness against image processing methods and adversarial defense strategies. Furthermore, our adversarial examples demonstrate an impressive attack success rate of 88% on commercial APIs, highlighting the effectiveness and practicality of our approach. Zeyu Zhao 0006, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | ExDA: Towards Universal Detection and Plug-and-Play Attribution of AI-Generated Ex-Regulatory ImagesabstractAs image-generative AI models become increasingly accessible to the public, the demand for content safety has surged. Although model developers have introduced alignment mechanisms to prevent the creation of threatening images, and extensive researches have been conducted on verifying the authenticity of AI-generated images, a significant number of ex-regulatory images have been discovered that fall into regulatory gaps. These images are neither covered by existing alignment mechanisms nor included in the scope of current detection methods. To address this, we introduce ExDA, a detection and attribution framework specifically designed for such ex-regulatory images. ExDA utilizes a frozen CLIP:ViT-L/14 as a visual feature extractor to extract rich and unbiased visual features, complemented by a text feature reduction layer to unify semantic styles. For obtaining highly discriminative features, ExDA introduces an SFS-ResNet network, where each basic layer is replaced with a meticulously designed Multi-Channel Margin Convolution (MMConv). Additionally, a plug-and-play multi-generation model attributor is integrated behind the detector. Given the lack of ex-regulatory images in existing public datasets, we constructed ExImage, a dataset containing 72,000 ex-regulatory images, to validate ExDA's effectiveness. Experiments show that ExDA achieves an average detection accuracy of 99.07% on ExImage, and demonstrating significant performance improvements of +5.73% and +10.36% on GenImage and high-challenge Chameleon datasets respectively in cross-datasets evaluation. Notably, ExDA also achieves excellent performance in attribution tasks, demonstrating its superior ability to identify the intrinsic fingerprints of generative models. Our code is available at https://github.com/mwp-create-wonders/ExDA. Wenpeng Mu, Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun |
ACM Multimedia | 5 |
| 2025 | Detection of fake bitrate videos based on high-frequency and deblocking filtering difference map features
Yikun Ao, Tanfeng Sun, Qiang Xu 0007 |
Neurocomputing | 2 |
| 2025 | MV-guided deformable convolution network for compressed video action recognition with P-frames
Yuting Mou, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
Neurocomputing | 4 |
| 2025 | A review of double compression detection for digital multimediaabstractThe rapid advancement of AI-driven multimedia manipulation has created an urgent need for more sophisticated digital forensics solutions. Current detection methods, while effective against specific tampering types, suffer from limited generalizability across diverse manipulation techniques. To address this challenge, researchers have developed Double Compression Detection (DCD) as a universal approach through compression-domain analysis. This review presents the comprehensive analysis of DCD techniques, systematically evaluating cutting-edge techniques for audio, image, and video content forensics. The pros and cons of existing DCD schemes are summarized for the first time from the perspective of generalization and effectiveness in this review. The emerging trends and fundamental limitations of existing researches are critically examined to guide future research directions in DCD. Tanfeng Sun, Qiang Xu 0007, Yueneng Wang |
Neurocomputing | 1 |
| 2025 | Detection of video transcoding from AVC to HEVC based on Intra Prediction Feature Maps
Yueneng Wang, Zhongjie Mi, Xinghao Jiang, Tanfeng Sun |
Neurocomputing | 4 |
| 2025 | Feature-Aware Transferable Adversarial Attacks on Visual Object TrackingabstractVisual object tracking is susceptible to adversarial attacks, posing significant security concerns for numerous application systems. Previous attack methods focused on white-box and untargeted attacks against response map. However, obtaining the tracking model in real-world scenarios is challenging, and the resulting adversarial trajectories are often unrealistic, making the attacks easily detectable. This paper proposes a Feature-aware Transferable Adversarial Patch (FTAP) that induces any black-box trackers to follow controllable and smooth trajectories. Tracker Following Assurance module is designed to manipulate bounding boxes to be valid and tightly align with the fake target. The movement of the tracker can be precisely controlled, resulting in adversarial trajectories stable and closely resemble natural trajectories, thereby reducing the risk of detection. The adversarial perturbation is generated solely from the initial template and applied to each frame. Consequently, the well-optimized generator can output universal adversarial patch capable of attacking any video without requiring additional computations. The intermediate layer features are corrupted to make the characteristics of the fake target closer to those of ground truth. Experimental results demonstrate that the proposed FTAP achieves state-of-the-art black-box attack performance and transferability across various tracker architectures. Mengdi Dong, Ke Xu 0003, Xinghao Jiang, Zeyu Zhao 0006, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | PLOVAD: Prompting Vision-Language Models for Open Vocabulary Video Anomaly DetectionabstractVideo anomaly detection (VAD) confronts significant challenges arising from data scarcity in real-world open scenarios, encompassing sparse annotations, labeling costs, and limitations on closed-set class definitions, particularly when scene diversity surpasses available training data. Although current weakly-supervised VAD methods offer partial alleviation, their inherent confinement to closed-set paradigms renders them inadequate in open-world contexts. Therefore, this paper explores open vocabulary video anomaly detection (OVVAD), leveraging abundant vision-related language data to detect and categorize both seen and unseen anomalies. To this end, we propose a robust framework, PLOVAD, designed to prompt tuning large-scale pretrained image-based vision-language models (I-VLMs) for the OVVAD task. PLOVAD consists of two main modules: the Prompting Module, featuring a learnable prompt to capture domain-specific knowledge and an anomaly-specific prompt crafted by a large language model (LLM) to capture semantic nuances and enhance generalization; and the Temporal Module, which integrates temporal information using graph attention network (GAT) stacking atop frame-wise visual features to address the transition from static images to videos. Extensive experiments on four benchmarks demonstrate the superior detection and categorization performance of our approach in the OVVAD task without bringing excessive parameters. Chenting Xu, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | HEVC Video Adversarial Samples Detection via Joint Features of Compression and Pixel DomainsabstractDeep learning models are currently under significant threat from adversarial attacks, while adversarial detection represents an effective means of countering such assaults. However, existing adversarial detection techniques are deficient in localizing video adversarial frames, leading to poor performance on sparse video adversarial attacks. This paper presents an approach for detecting adversarial perturbations in videos based on fusion features derived from the video compression and RGB domain. Our research begins by examining how the introduction of extensive non-natural noise during video adversarial attacks severely disrupts the spatial structure of individual frames and the motion information between frames. This disruption culminates in unnatural variations in the Coding Tree Units (CTU) partitioning during the HEVC video encoding process. Then meticulously mapping the positions and partitioning information of coding units (CU), predictive units (PU), and transformation units (TU) onto specific values and sizes, constituting the video’s Compression Domain Units (CDU) features. Finally, a dual-path network utilizing both the video’s CDU features and the decoded frames RGB features is employed for detecting video adversarial samples. Extensive experiments are conducted to verify the performance. The results show that the proposed scheme outperforms or rivals the state-of-the-art methods in video adversarial detection. Zeyu Zhao 0006, Yueneng Wang, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | A Robust Coverless Video Steganography Based on Two-Level DCT Features Against Video AttacksabstractCompared with traditional video steganography, coverless video steganography (CVS) can completely avoid being detected by steganalysis algorithms. Recently, the study of CVS has developed rapidly. However, it is still far from the theoretical maximum values in capacity, i.e., the theoretical limit is$2^\ell$for a hash sequence length of$\ell$. Besides, most existing CVS methods have only considered limited types of video attacks in robustness. In this paper, a novel coverless video steganography based on two-level discrete cosine transform (DCT) features is proposed. First, pre-processing is accomplished on the public video datasets. Then, two-level DCT features are calculated and the Coverless Video Database (CVD) is constructed by the K-means++ clustering algorithm. After that, the mapping table is established to map the secret segments to the CVD. Finally, each secret segment corresponds to a video sequence in the CVD by the mapping table to complete the process of information embedding and extraction. The proposed method first evaluates the robustness against the frame swapping attack, which is a common video attack. Experimental results show that the proposed method can achieve the theoretical maximum value in effective capacity and better robustness compared to the state-of-the-art works. Laijin Meng, Xinghao Jiang, Qiang Xu 0007, Tanfeng Sun |
IEEE Trans. Multim. | 4 |
| 2025 | Preemptive Defense Algorithm Based on Generalizable Black-Box Feedback Regulation Strategy Against Face-Swapping Deepfake ModelsabstractIn the previous efforts to counteract Deepfake, detection methods were most adopted, but they could only function after-effect and could not undo the harm. Preemptive defense has recently gained attention as an alternative, but such defense works have either limited their scenario to facial-reenactment Deepfake models or only targeted specific face-swapping Deepfake model. Motivated to fill this gap, we start by establishing the Deepfake scenario modeling and finding the scenario difference among categories, then move on to the face-swapping scenario setting overlooked by previous works. Based on this scenario, we first propose a novel Black-Box Penetrating Defense Process that enables defense against face-swapping models without prior model knowledge. Then we propose a novel Double-Blind Feedback Regulation Strategy to solve the reality problem of avoiding alarming distortions after defense that had previously been ignored, which helps conduct valid preemptive defense against face-swapping Deepfake models in reality. Experimental results in comparison with state-of-the-art defense methods are conducted against popular face-swapping Deepfake models, proving our proposed method valid under practical circumstances. Zhongjie Mi, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003, Qiang Xu 0007 |
IEEE Trans. Multim. | 3 |
| 2025 | Detection of HEVC Double Compression Based on Deep Representations of In-Loop Filtering and CU Depth MapsabstractIn the field of HEVC (High Efficiency Video Coding) double compression detection, relocated I-frame (RI frame) detection and original GOP size estimation are two significant problems for video forensics. However, little research explores the interconnection between the two problems, and effective methods to resolve them are still lacking. In this paper, a novel feature model called In-loop Filtering and CU Depth Map (IFCDM) is proposed to accurately detect RI frames, and the intrinsic correlation between RI frames and GOP structure is explored, which can be used for original GOP size estimation. Theoretical and statistical analysis of HEVC recompression process is first carried out. Then, sub-features of HEVC in-loop filtering modes and CU partition depth are extracted, and transformed into grey-scale maps to construct IFCDM. A neural network, consisting of tiny Vision Transformer and LSTM, is trained to learn spatial and temporal representations of input features, and further derive the RI frame detection results. Finally, an adaptive periodic analysis algorithm is designed, to integrate the RI frame detection results and estimate the original GOP size of recompressed videos. Experiments show that our method can outperform the existing state-of-the-art methods in both frame level and video level. Tanfeng Sun, Qiang Xu 0007, Ke Xu 0003, Xinghao Jiang |
IEEE Trans. Multim. | 2 |
| 2024 | Learning Spatio-Temporal Relations with Multi-Scale Integrated Perception for Video Anomaly DetectionabstractIn weakly supervised video anomaly detection, it has been verified that anomalies can be biased by background noise. Previous works attempted to focus on local regions to exclude irrelevant information. However, the abnormal events in different scenes vary in size, and current methods struggle to consider local events of different scales concurrently. To this end, we propose a multi-scale integrated perception (MSIP) learning approach to perceive abnormal regions of different scales simultaneously. In our method, a frame is partitioned into several groups of patches with varying scales, and a multi-scale patch spatial relation (MPSR) module is further proposed to model the inconsistencies among multi-scale patches. Specifically, we design a hierarchical graph convolution block in the MPSR module to improve the integration of patch features by implementing cross-scale feature learning. An existing clip temporal relation network is also introduced to enable spatio-temporal encoding in our model. Experiments show that our method achieves new state-of-the-art performance on the ShanghaiTech and competitive results on UCF-Crime benchmarks. Hongyu Ye, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICASSP | 4 |
| 2024 | Low-Quality Deepfake Video Detection Model Targeting Compression-Degraded Spatiotemporal Inconsistencies
Zhongjie Mi, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003, Qiang Xu 0007, Laijin Meng |
ICIC (9) | 3 |
| 2024 | Sparse Silhouette Jump: Adversarial Attack Targeted at Binary Image for Gait Privacy ProtectionabstractWith the widespread application of gait recognition technology, the issue of gait semantic security in videos has also attracted the attention of researchers. Its goal is to destroy the readability of data while preserving its semantic features. Thanks to the development of deep learning, some existing methods both domestically and internationally have made certain breakthroughs in recognition accuracy and visual effects. However, there is still significant room for improvement in balancing protection ability, visual effects, and computational complexity. This work is based on the ability of deep learning networks to extract gait identity features. In response to some problems in the current research field, we propose a gait privacy protection algorithm based on Sparse Silhouette Jump(SSJ), which draws on the idea of gradient descent in adversarial attacks and transfers adversarial noise to binary jumps to better adapt to binary graphs, while limiting the range of jumps from the perspective of spatial sparsity to balance the effectiveness and concealment of attacks. Experimental results have shown that our method achieves good effectiveness and concealment for various gait recognition models. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
TrustCom | 4 |
| 2024 | Compressed Video Action Recognition Based on Neural Video CompressionabstractCompressed video action recognition based on traditional codecs, like MPEG-4, H265, etc., has achieved remarkable progress with comparable performance to raw video action recognition. With the development of Neural Video Compression (NVC), action recognition based on NVC should be paid attention to and explored. Firstly, the encoded stream of NVC represents the high-dimension features of the neural network, which allows the features to be utilized for downstream tasks with less additional processing or even directly. Secondly, the high-dimension feature can not be understood by humans, which means the privacy of the raw video frames can be preserved. In this paper, we propose a novel model for compressed video action recognition based on NVC to explore the potential of NVC for action recognition. By introducing spatial and temporal co-attention (ST-CA), the spatial information from the reference frame feature and the temporal information from the motion vector and the residual feature are combined and complemented effectively. The proposed model achieves competitive performance with the traditional compressed video action recognition methods and the raw video action recognition methods on the HMDB-51 and UCF-101 datasets. Besides, the proposed model preserves the privacy of the raw video frames and has much less computational complexity than the raw video action recognition methods. Yuting Mou, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
TrustCom | 4 |
| 2024 | MDTL-NET: Computer-generated image detection based on multi-scale deep texture learning
Qiang Xu 0007, Shan Jia, Xinghao Jiang, Tanfeng Sun, Zhe Wang 0035, Hong Yan 0001 |
Expert Syst. Appl. | 4 |
| 2024 | A review of coverless steganography
Laijin Meng, Xinghao Jiang, Tanfeng Sun |
Neurocomputing | 3 |
| 2024 | Detection of HEVC double compression based on boundary effect of TU and non-zero DCT coefficient distribution
Tanfeng Sun, Qiang Xu 0007 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | A robust coverless video steganography based on maximum DC coefficients against video attacks
Laijin Meng, Xinghao Jiang, Zhaohong Li, Tanfeng Sun |
Multim. Tools Appl. | 5 |
| 2024 | Compressed Video Action Recognition With Dual-Stream and Dual-Modal TransformerabstractCompressed video action recognition offers the advantage of reducing decoding and inference time compared to the RGB domain. However, the compressed domain poses unique challenges with different types of frames (I-frames and P-frames). I-frames consistent with RGB are rich in frame information, but the redundant information may interfere with the recognition task. There are two modalities in P-frames, residual (R) and motion vector (MV). Although with less information, they can reflect the motion cue. To address these challenges and leverage the independent information from different frames and modalities, we propose a novel approach called Dual-Stream and Dual-Modal Transformer (DSDMT). Our approach consists of two streams: 1) The short-span P-frames stream contains temporal information. We propose the Dual-Modal Attention Module (DAM) to mine different modal variability in P-frames and complement the orthogonal feature vector. Besides, considering the sparsity of P-frames, we extract action features with Frame-level Patch Embedding (FPE) to avoid redundant computation. 2) The long-span I-frames stream extracts the global context feature of the entire video, including content and scene information. By fusing the global video context and local key-frame features, our model represents the action feature in terms of fine-grained and coarse-grained. We evaluated our proposed DSDMT on three public benchmarks with different scales: HMDB-51, UCF-101, and Kinetics-400. Ours achieve better performance with fewer Flops and lower latency. Our analysis shows that the independence and complements of the I-frames and P-frames extracted from the compressed video stream play a crucial role in action recognition. Yuting Mou, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun, Zepeng Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Robust Coverless Video Steganography Based on the Similarity of Inter-FramesabstractWith a deeper understanding of the security issues in steganography, coverless steganography has become a hotspot due to no modification to the carriers. However, the existing coverless video steganographic algorithms have considered a few types of video attacks. In this paper, a robust coverless video steganography based on the similarity of inter-frames is proposed. First, a public video database is selected and preprocessed to construct a Secret Communication Video Database (SCVD). The similarity score between the first and last frames is calculated for video sorting to utilize the temporal characteristics of videos. After that, the mapping table between the secret information and the SCVD is designed for both senders and receivers. Finally, each secret information segment can be represented by one video sequence in the SCVD according to the mapping table to accomplish the data hiding and extraction. Experimental results show that the proposed method performs much better in capacity, robustness, and security than the state-of-the-art methods. It is worth mentioning that the proposed method overcomes the security issue of transmitting a large amount of auxiliary information in coverless video steganographic algorithms. Laijin Meng, Xinghao Jiang, Tanfeng Sun, Zeyu Zhao 0006, Qiang Xu 0007 |
IEEE Trans. Multim. | 3 |
| 2023 | Deformable graph convolutional transformer for skeleton-based action recognition
Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
Appl. Intell. | 5 |
| 2023 | A Robust Coverless Image Steganography Based on an End-to-End Hash Generation ModelabstractRecently, coverless steganography algorithms have attracted increased research attention due to their ability to completely resist steganalysis algorithms. However, the existing algorithms do not attain the same robust balance against geometric and non-geometric attacks. In addition, most of the existing methods need to transmit some auxiliary information along with the stego-images, which increases the cost of the hidden information. In this paper, a robust coverless image steganography algorithm based on a hash generation model is proposed. Different from the existing methods, the hash sequences are generated by an end-to-end CNN model, where the input is the original images, and the output is the corresponding hash sequences. Therefore, no auxiliary information needs to be transmitted when hiding the secret information. Moreover, the attention mechanism and adversarial training are introduced to improve the robustness of the model. The loss function is redesigned to accommodate these operations. Finally, an index structure is built to enhance the mapping efficiency. The experimental results show that the proposed method possesses better robustness and security compared with the state-of-the-art coverless image steganography algorithms. Laijin Meng, Xinghao Jiang, Zhaohong Li, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Adaptive HEVC Steganography Based on Steganographic Compression Efficiency Degradation ModelabstractHigh Efficiency Video Coding (HEVC) places great emphasis on optimizing compression efficiency, where compression efficiency denotes file size ratio before and after compression. The current HEVC steganography is prone to cause degradation in compression efficiency. To analyze and avoid this problem, a Steganographic Compression Efficiency Degradation Model (SCEDM) is first proposed, which leverages the area ratio of different types of Coding Units (CU) as the distribution of block partitioning structure and combines with the K-L divergence to describe the compression efficiency degradation. By minimizing the output of the SCEDM, the degradation of compression efficiency caused by steganographies can be minimized. Besides, it is also proved that this minimizing process will not increase extra visual quality distortion. Based on this model, a novel adaptive steganography using HEVC intra block partitioning structure is proposed. This steganography consists of three parts: the CU Depth based Hierarchical Coding (CDHC) method, the structure merging strategy and the adaptive matching method. The CDHC method can convert secret binary bits to different block structures. The structure merging strategy improves the capacity, and the adaptive matching method minimizes the compression efficiency degradation according to the proposed SCEDM. The proposed steganography is further compared with state-of-the-art steganographies to confirm the effectiveness and advantages of the proposed model and steganography in compression efficiency, capacity, visual quality and resistance to video steganalysis. Xinghao Jiang, Zhaohong Li, Tanfeng Sun, Peisong He |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | An Anti-Steganalysis HEVC Video Steganography With High Performance Based on CNN and PU Partition ModesabstractThe steganography research of videos leads to excellent communication methods for transmitting secret message, and high efficiency video coding(HEVC) video is one popular steganographic carrier. This article proposes a prediction unit(PU) based wide residual-net steganography(PWRN) for HEVC videos. The visual quality distortion of modifying PUs is theoretically analyzed, which illustrates that modifying PUs only has a little negative effect on visual quality. Therefore, the data hiding method in this article allows to modify all types of PUs except for$2N\times 2N$to each other according to the secret data. In this way, high embedding efficiency is achieved, and the PU distributions in stego-videos can be kept similar to those of cover-videos, which is essential for resisting steganalysis. Meanwhile, a super-resolution convolutional neural network(CNN) with wide residual-net filter(WRNF) is proposed to replace the in-loop filter in HEVC for reconstructing I-pictures, which results in more precisely predicted P-pictures, and it further leads to less bitrate cost and better visual quality of stego-videos. The experimental results show that the proposed PWRN successfully resists the latest PU-targeted steganalysis algorithms, and compared with the state-of-the-art work, PWRN has achieved the lowest bitrate cost and the highest visual quality under the same capacity. Xinghao Jiang, Laijin Meng, Tanfeng Sun |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | Transferable Black-Box Attack Against Face Recognition With Spatial Mutable Adversarial PatchabstractDeep Neural Networks (DNNs) are vulnerable to adversarial patch attacks, which raises security concerns for face recognition systems using DNNs. Previous attack methods focus on the perturbation texture and generate adversarial patches with fixed shapes at random or pre-designed locations, which causes poor adversarial transferability. This paper proposes a Spatial Mutable Adversarial Patch (SMAP) method to generate a dynamic mutable patch to be injected into the face. In the proposed SMAP, the texture, position and shape of the patch are optimized simultaneously and the patch generation pipeline is end-to-end differentiable. Specifically, a Patch Location Selection Scheme is designed to find the critical patch position with the most significant influence on the target identity by the step-based gradient search. By innovatively bridging the pre-defined mask and the dynamic update of the patch, the patch position and shape are changed based on the affine transformation and sampling mechanism in each iteration, which maintains the importance of the injected patch to the adversarial objective. To evaluate the vulnerability of face recognition models, we explore more threatening impersonation attacks under the black-box setting and design a strict evaluation metric that aligns with the real-world scenario. Extensive experiments show that the proposed SMAP improves attack performance across various face recognition models and datasets. Moreover, SMAP achieves better transferability on commercial face recognition systems than existing methods. Haotian Ma 0001, Ke Xu 0003, Xinghao Jiang, Zeyu Zhao 0006, Tanfeng Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Multi-Channel HEVC Steganography by Minimizing IPM Steganographic DistortionsabstractCalibration is a common method for steganalysis, and Intra Prediction Mode (IPM) shift is a typical phenomenon used in calibration to detect video steganography. The current HEVC steganography lacks resistance to steganalysis based on this phenomenon because the new technology of HEVC introduces steganographic distortion in addition to providing more potential steganographic space. In this paper, an HEVC steganographic algorithm that resists IPM shift is proposed. First, we introduce the IPM shift in HEVC, and the previous H.264 steganalytic IPM shift feature is modeled and improved. By analyzing the HEVC encoding process, we found that modifying large-size blocks has a more significant impact on compression efficiency, while small ones are more sensitive to IPM optimality. Therefore, we perform the embedding channel division based on the block size and design the distortion function separately. In addition, we discover a unique IPM transition probability distribution in HEVC. According to our analysis, this unique distribution arises due to HEVC's MPM rules and the regularity of IPM direction. Modifying IPM in HEVC will change such distribution, thus, a mapping rule is designed based on this distribution to achieve a better embedding effect. Experimental results show that the channel division and proposed distortion function can effectively improve the overall performance. The proposed steganography outperforms the state-of-the-art steganography in resisting steganalysis, bitrate controlling, and visual quality. Xinghao Jiang, Zhaohong Li, Tanfeng Sun |
IEEE Trans. Multim. | 4 |
| 2022 | A Transformer-Based Cloth-Irrelevant Patches Feature Extracting Method for Long-Term Cloth-Changing Person Re-identification
Zepeng Wang 0002, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
CGI | 4 |
| 2022 | Research on video adversarial attack with long living cycleabstractIn recent years, the vulnerability of networks has attracted the attention of researchers. However, in these methods, the impact of video compression coding on the added adversarial perturbation, i.e., the robustness of the video adversarial example, is not considered. When an adversarial sample is just generated, its attack capability is the strongest. However, with multiple video encoding and video decoding in Internet transmission, the added adversarial disturbance will be continuously eliminated, eventually leading to the attack on the adversarial sample performance disappearing. We define this phenomenon as the decay of the lifetime of adversarial examples. We propose an adversarial attack method based on optimized integer space to resist this performance degradation. The robustness of anti-coding, the visual concealment, and the attack success rate are all considered during the attack process. In addition, we have also reduced the rounding loss caused by normalization in the deep neural network model process. The contributions of our methods are 1) We show the performance degradation caused by video compression coding on existing video adversarial attack methods, which seems an effective way for detecting of defending video adversarial examples. 2) A robust video adversarial attack method is proposed to resist video compression coding. The experiment shows that our method performs better on the robustness of anti-coding, visual concealment, and attack success rate. Zeyu Zhao 0006, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
UAI | 4 |
| 2022 | Dual-domain graph convolutional networks for skeleton-based action recognition
Ke Xu 0003, Zhongjie Mi, Xinghao Jiang, Tanfeng Sun |
Mach. Learn. | 5 |
| 2022 | Motion-Adaptive Detection of HEVC Double Compression With the Same Coding ParametersabstractHigh Efficiency Video Coding (HEVC) double compression detection is of prime significance in video forensics. However, double compression with the same parameters and video content with high motion displacement intensity have become two main factors that limit the performance of existing algorithms. To address these issues, a novel motion-adaptive algorithm is proposed in this paper. Firstly, the analysis of GOP structure in HEVC standard and the coding process of HEVC double compression are provided. Next, sub-features composed of fluctuation intensities of intra prediction modes and unstable Prediction Units (PUs) in normal Intra-Frames (I-frames) and optical flow in adaptive I-frames are exploited in our algorithm. Each sub-feature is extracted during the process of multiple decompression. We further combine these sub-features into a 27-dimensional detection feature, which is finally fed to the Support Vector Machine (SVM) classifier. By following a separation-fusion detection strategy, the experimental result shows that the proposed algorithm outperforms the existing state-of-the-art methods and demonstrates superior robustness to various motion displacement intensities and a wide variety of coding parameter settings. Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Gait Recognition Based on Local Graphical Skeleton Descriptor With Pairwise Similarity NetworkabstractGait recognition aims to identify a human through a walking sequence. It is a challenging task in computer vision since monocular camera loses most of the 3D information. Previous works described gait features with the contours of shape or the global geometrical characters of skeleton. So little work is researched on the local patterns of gait skeleton. In this paper, to resist the dress changes and speed changes, a Local Graphical Skeleton Descriptor (LGSD) is proposed to describe both the inner and intra local graphical patterns of a human gait skeleton. The gait features from the same or different identities are paired up and a Pairwise Similarity Network (PSN) is proposed to maximize the similarity of True matched pairs and minimize the similarity of False matched pairs. The contributions of our method are: 1) LGSD is proposed to describe human gait by computing four novel local geometrical patterns of skeleton sequences, which makes use of the intuitive cognition of gait based on the prior knowledge of mankind. 2) PSN is implemented by a two-stream CNN structure to build the gait model, which fused two popular gait recognition strategies. 3) The robustness of our method to dress changes and speed changes is proved on the public datasets. We have also achieved some state-of-the-art results on these datasets. The proposed method is examined on three public gait datasets which have RGB or infrared frames for evaluation: the CASIA-B dataset, the NLPR gait database, and the CASIA-C dataset. The performers in these datasets are walking under different views, speeds or dresses. The results are further compared with previous approaches to confirm the effectiveness and the advantages of our method. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Multim. | 3 |
| 2021 | SE_EDNet: A Robust Manipulated Faces Detection Algorithm
Chaoyang Peng, Lihong Yao, Tanfeng Sun, Xinghao Jiang, Zhongjie Mi |
CGI | 3 |
| 2021 | Gait Identification Based on Human Skeleton with Pairwise Graph Convolutional NetworkabstractVision based gait identification is an important research content in the field of biometrics. Most existing gait identification methods extract representations from gait videos and identify a probe gait by ranking the similarities between the probe gait and all the gallery gaits. Since human body skeletons convey significant static and dynamic information of gait, in this paper, an end-to-end gait identification network named Pairwise Graph Convolutional Network (PGCN) is proposed to capture gait feature from skeletons. The skeleton sequences are first mapped into gait graphs and the PGCN is constructed to learn graph representations. The proposed method is examined on CASIA-B dataset on identical-view and cross-dress cases. The contributions of our work are: 1) The PGCN gait identification model is proposed to extract robust gait representation from skeleton sequences. 2) Different fusion structures are compared to explore the best fusion strategy for gait representation. 3) Using skeleton data, our work outperforms previous methods on identical-view and cross-dress cases. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICME | 3 |
| 2021 | A HEVC Steganalysis Algorithm Based on Relationship of Adjacent Intra Prediction Modes
Henan Shi, Tanfeng Sun, Zhaohong Li |
IWDW | 2 |
| 2021 | Detection of HEVC double compression with non-aligned GOP structures via inter-frame quality degradation analysis
Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
Neurocomputing | 3 |
| 2021 | Detection of transcoded HEVC videos based on in-loop filtering and PU partitioning analyses
Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
Signal Process. Image Commun. | 3 |
| 2020 | Combined features for steganalysis against PU partition mode-based steganography in HEVC
Kuan Huang, Tanfeng Sun, Xinghao Jiang, Qianan Fang |
Multim. Tools Appl. | 2 |
| 2020 | Action Recognition Scheme Based on Skeleton Representation With DS-LSTM NetworkabstractSkeleton-based human action recognition has been a popular research field during the past few years. With the help of cameras equipping deep sensors, such as the Kinect, human action can be represented by a sequence of human skeleton data. Inspired by the skeleton descriptors based on Lie group, a spatial-temporal skeleton transformation descriptor (ST-STD) is proposed in this paper. The ST-STD describes the relative transformations of skeletons, including the rotation and translation during movement. It gives a comprehensive view of the skeleton in both spatial and temporal domain for each frame. To capture the temporal connections in the skeleton sequence, a denoising sparse long short term memory (DS-LSTM) network is proposed in this paper. The DS-LSTM is designed to deal with two problems in action recognition. First, to decrease the intra-class diversity, the spatial-temporal auto-encoder (STAE) is proposed in this paper to generate representations with higher abstractness. The denoising constraint and the sparsity constraint are applied on both spatial and temporal domain to enhance the robustness and to reduce action misalignment. Second, to model the action sequence, a three-layer LSTM structure is trained with STAE representations for temporal modeling and classification. The experiments are carried out on four popular datasets. The results show that our approach performs better than several existing skeleton-based action recognition methods, which prove the effectiveness of our method. Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Detection of HEVC Double Compression With the Same Coding Parameters Based on Analysis of Intra Coding Quality Degradation ProcessabstractThe emergence of the high-efficiency video coding (HEVC) standard enables people to enjoy high definition (HD) video content; meanwhile, HD videos, tamper detection has become a crucial issue and gradually aroused people's attention. The detection of double HEVC compressed videos with the same coding parameters is challenging since the recompression traces are inconspicuous. To deal with this issue, a novel method based on the intra prediction mode is proposed in this paper. First, the quality degradation mechanism is analyzed to facilitate the selection of classification features and the source of error in intra coding is fully considered to establish the equivalent error model. Second, the feature model of double HEVC compression detection, which is mainly based on the statistical feature of intra prediction mode, is proposed. Finally, the experiment is carried out in 720p and 1080p HEVC videos instead of low-resolution (CIF or QCIF) videos. Experimental results have demonstrated better efficiency of the proposed method in comparison to the state-of-the-art methods. Besides, the proposed method is more robust to various encoding configurations. Xinghao Jiang, Qiang Xu 0007, Tanfeng Sun, Bin Li 0011, Peisong He |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Video Anomaly Detection and Localization Based on an Adaptive Intra-Frame Classification NetworkabstractVideo anomaly detection and localization is still a challenging task in the computer vision field. Previous methods took this task as an outlier detection problem, which computed the deviation between the test samples and the normal patterns. In this paper, an adaptive intra-frame classification network (AICN) is proposed to transform this task to a multi-class classification problem. The contributions of our method are as follows. AICN is an end-to-end network for anomaly detection and localization. By using the motion convolutional layers and the shape convolutional layers, spatial-temporal features are extracted without resizing or splitting the frames before forward propagation. AICN enhances the adaptiveness of model. By using the adaptive region pooling layer and the intra-frame classifier, AICN is adaptive to frames with different resolutions and is easier to be applied on other scenes. AICN evaluates the abnormality of frames based on the intra-frame classification results. The intra-frame classification strategy reserves more connection information of sub-regions and makes the model outperform previous methods. The proposed method is examined on four public datasets with different background complexities and resolutions: UCSD Ped1 dataset, UCSD Ped2 dataset, Avenue dataset and Subway dataset. The results are further compared with previous approaches to confirm the effectiveness and the advantage of our method. Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Multim. | 2 |
| 2019 | Efficient Violence Detection Using 3D Convolutional Neural NetworksabstractAutomatically analyzing violent content in surveillance videos is of profound significance on many applications, ranging from Internet video filtration to public security protection. In this paper, we propose a deep learning model based on 3D convolutional neural networks, without using hand-crafted features or RNN architectures exclusively for encoding temporal information. The improved internal designs adopt compact but effective bottleneck units for learning motion patterns and leverage the DenseNet architecture to promote feature reusing and channel interaction, which is proved to be more capable of capturing spatiotemporal features and requires relatively fewer parameters. The performance of the proposed model is validated on three standard datasets in terms of recognition accuracy compared to other advanced approaches. Meanwhile, supplementary experiments are carried out to evaluate its effectiveness and efficiency. The final results demonstrate the advantages of the proposed model over the state-of-the-art methods in both recognition accuracy and computational efficiency. Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
AVSS | 3 |
| 2019 | 3D Gait Recognition Based on a CNN-LSTM Network with the Fusion of SkeGEI and DA FeaturesabstractGait recognition is a promising technology in biometrics in video surveillance applications for its characteristics of non-contact and uniqueness. With the popularization of the Kinect sensor, human gait can be recognized based on the 3D skeletal information. For exploiting raw depth data captured by Kinect device effectively, a novel gait recognition approach based on Skeleton Gait Energy Image (SkeGEI) and Relative Distance and Angle (DA) features fusion is proposed. They are fused in backward to complement each other for gait recognition. In order to maintain as much gait information as possible, a CNN-LSTM network is designed to extract the temporal-spatial deep feature information from SkeGEI and DA features. The experiments evaluated on three datasets show that our approach performs superior to most gait recognition approaches with multi-directional and abnormal patterns. Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
AVSS | 3 |
| 2019 | A Motion Vector-Based Steganographic Algorithm for HEVC with MTB Mapping Strategy
Mengyuan Guo, Tanfeng Sun, Xinghao Jiang, Ke Xu 0003 |
IWDW | 2 |
| 2019 | Gait Recognition with Clothing and Carrying Variations Based on GEI and CAPDS Features
Fengjia Yang, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
PRCV (2) | 3 |
| 2018 | A High Capacity HEVC Steganographic Algorithm Using Intra Prediction Modes in Multi-sized Prediction Blocks
Tanfeng Sun, Xinghao Jiang |
IWDW | 2 |
| 2018 | Anomaly detection based on Nearest Neighbor search with Locality-Sensitive B-tree
Maying Shen, Xinghao Jiang, Tanfeng Sun |
Neurocomputing | 3 |
| 2018 | Computer Graphics Identification Combining Convolutional and Recurrent Neural NetworksabstractIn this letter, a deep-learning-based pipeline is proposed to distinguish photographics (PGs) from computer-graphics (CGs) combining convolutional neural network (CNN) and recurrent neural network (RNN). In the preprocessing stage, the color space transformation and the Schmid filter bank are utilized to extract chrominance and luminance components, which suppress the irrelevant information of various image contents for the CG identification task. Then, a dual-path CNN architecture is designed to learn joint feature representations of local patches for exploiting their color and texture characteristics. To extract the global artifact, the directed acyclic graph RNN is applied to model the spatial dependence of local patterns. Finally, the output score of RNN is used to identify the input sample. The CG/PG dataset is constructed by collecting samples from the Internet. Experimental results show that the proposed framework can outperform state-of-the-art methods on identification ability of CGs, especially for images with low resolution. Peisong He, Xinghao Jiang, Tanfeng Sun, Haoliang Li |
IEEE Signal Process. Lett. | 3 |
| 2018 | Detection of Double Compression With the Same Coding Parameters Based on Quality Degradation Mechanism AnalysisabstractDetection of double compression with the same coding parameters is a very challenging problem in video forensics, since traces of recompression operations are extremely slight in this case. To solve this problem, we first analyze degradation mechanisms during recompression. It is observed that the video quality tends to become nearly unchanged after multiple recompressions with the same coding parameters. The degree of quality degradation is used to distinguish single and double compressed videos. This property can be described using the convergent tendency of video data to unchanged states after continuous recompressions. For MPEG videos, statistical features of rounding and truncation errors are extracted from the intra-coding process while macroblock-mode based features are obtained from the inter-coding process. The final feature is generated by concatenating these two sets of features to provide robust detection capability. Then, extracted features are fed to the SVM classifier to obtain the final detection result. In addition, aforementioned features are modified and extended to detect double compression on H.264 videos based on the unique coding techniques developed in the H.264 standard, such as intra-prediction. Several public available YUV sequences are used to construct double compression databases with three popular coding standards, including MPEG-2, MPEG-4, and H.264. In experiments, the proposed method outperforms several state-of-the-art methods for different compression qualities and rate control schemes. Experimental results demonstrate the proposed method has more robust detection capability of double compression under various encoding configurations. Xinghao Jiang, Peisong He, Tanfeng Sun, Shi-Lin Wang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Anomaly Detection Based on Stacked Sparse Coding With Intraframe Classification StrategyabstractAnomaly detection in videos is still a challenging task among the computer vision community. In this paper, an efficient anomaly detection method based on stacked sparse coding (SSC) with intraframe classification strategy is proposed. Each video is divided into blocks and the Foreground Interest Point (FIP) descriptor is proposed to describe the appearance and motion features for each block. The spatial-temporal features are then encoded with SSC. Specifically, the first stage of SSC encodes the spatial connections among blocks and the second stage of SSC encodes the temporal connections of all frame patches in each block. Finally, an intraframe classification strategy which uses the probabilistic outputs of SVM is proposed to evaluate the abnormality of each block. Contributions of this paper are listed as follows: 1) The FIP descriptor is proposed to describe the features of blocks, which reserves more spatial-temporal information. 2) The SSC encoding method encodes both the spatial and temporal connections of blocks, which makes the features more representative. 3) The intraframe classification strategy keeps the evaluation consistency among blocks and it helps to improve detection performance. The proposed method is examined on four public datasets with different background complexities and resolutions: UCSD Ped1 dataset, UCSD Ped2 dataset, Avenue dataset, and Subway dataset. The results are further compared with previous approaches to confirm the effectiveness and advantages of this method. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Multim. | 3 |
| 2017 | Anomaly Detection by Analyzing the Pedestrian Behavior and the Dynamic Changes of Behavior
Maying Shen, Xinghao Jiang, Tanfeng Sun |
ICIC (1) | 3 |
| 2017 | Double H.264 Compression Detection Scheme Based on Prediction Residual of Background Regions
Junjia Zheng, Tanfeng Sun, Xinghao Jiang, Peisong He |
ICIC (1) | 2 |
| 2017 | 3D human action recognition based on the Spatial-Temporal Moving Skeleton DescriptorabstractWith the popularization of the Kinect sensor, human actions can be recognized based on the 3D skeletal information. In this paper, the Spatial-Temporal Moving Skeleton Descriptor (STMSD) is proposed by the fusion of three complementary features which are the Relative Geometric Velocity (RGV) between body parts, Relative Joint Positions (RJP), and Joint Angles (JA). The STMSD descriptor gives a complete view of the body skeleton in space and time. Among the three features, the Relative Geometric Velocity (RGV) is first proposed in our work. Inspired by the relative geometry using the Lie group and the Lie algebra, RGV describes the variation rates of body transformations which include 3D rotations and translations. Then interpolation and normalization are applied in frame descriptors. After the temporal modeling, Principal Component Analysis (PCA) is utilized. Experimental results on three datasets show that our approach performs better than existing action recognition approaches, including skeleton-based and other types. Hongxian Yao, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
ICME | 3 |
| 2017 | Coding Efficiency Preserving Steganography Based on HEVC Steganographic Channel Model
Xinghao Jiang, Tanfeng Sun, Dawen Xu 0001 |
IWDW | 3 |
| 2017 | HEVC Double Compression Detection Based on SN-PUPM Feature
Qianyi Xu, Tanfeng Sun, Xinghao Jiang |
IWDW | 2 |
| 2017 | Detection of double compression in MPEG-4 videos based on block artifact measurement
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
Neurocomputing | 3 |
| 2017 | Frame-wise detection of relocated I-frames in double compressed H.264 videos based on convolutional neural network
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang, Bin Li 0011 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Two-Stream Dictionary Learning Architecture for Action RecognitionabstractIn this paper, a novel method based on the two-stream dictionary learning architecture for human action recognition is proposed. The architecture consists of interest patch (IP) detector and descriptor, two-stream dictionary models, and support vector machine (SVM) for classification. The novel IP detector combines a human detector and a contour detector to extract patches of interest on human contours. Then the IP descriptors are calculated in spatial stream and temporal stream separately. In each stream, a dictionary is trained for each action with the IP descriptors as an action model. In this way, measuring the similarity between an action sequence and an action model is transformed to reconstructing the IPs in this sequence with the model and computing the reconstruction error. For each action, an IP distribution histogram is constructed and the histogram is further used to train an SVM classifier in each stream. A score fusion method is applied to fuse the spatial and temporal SVM classification results to make a final decision. The proposed architecture is examined on four public data sets with different background complexities and camera motion conditions: Weizmann data set, KTH data set, Olympic sports data set, and HMDB51 data set. The results are further compared with state-of-the-art approaches in the experiment section to confirm the effectiveness of this architecture. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Detecting double MPEG compression with the same quantiser scale based on MBM featureabstractDetecting double MPEG compression is of prime significance in video forensics. However, existing methods are effective only when the primary compression and the secondary compression have different quantiser scales (QS). There is a lack of effective methods dealing with double MPEG compression with the same QS. In this paper, a novel method based on the statistical feature of macroblock mode (MBM) which consists of macroblock type and motion vector in P-frames is proposed to detect double MPEG compression with the same QS. The MBM statistical feature is extracted during multiple decoding procedures when the video is repeatedly compressed with the same QS for several times. Finally, the proposed feature is combined with the support vector machine (SVM) to classify the single MPEG compression and double MPEG compression. Experiments have demonstrated the effectiveness of the proposed method and the robustness to a wide range of QSs and different encoders. Jieyuan Chen, Xinghao Jiang, Tanfeng Sun, Peisong He, Shi-Lin Wang |
ICASSP | 3 |
| 2016 | Detecting Double H.264 Compression Based on Analyzing Prediction Residual Distribution
Tanfeng Sun, Xinghao Jiang, Peisong He, Shi-Lin Wang, Yun Q. Shi 0001 |
IWDW | 2 |
| 2016 | Double compression detection based on local motion vector field analysis in static-background videos
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Double Compression Detection in MPEG-4 Videos Based on Block Artifact Measurement with Variation of Prediction Footprint
Peisong He, Tanfeng Sun, Xinghao Jiang, Shi-Lin Wang |
ICIC (3) | 2 |
| 2015 | Human activity recognition based on pose points selectionabstractA novel method for human action recognition is proposed in this paper. Traditional spatial-temporal interest point detectors are easily affected by hair, face, shadow, clothes texture or the shake of camera. Inspired by the use of points distribution information, we propose a point selection method to select representative points (denoted by the “pose points”), which use HOG human detector and contour detector to select the points on human pose edges. The pose points carry both local gradient information and global pose information. 3D-SIFT scale selection method and novel descriptors called body scale and motion intensity feature are also studied. The descriptors calculate the width scale of different levels of human body and count motion intensity of activity in five directions. The descriptors combine spatial location with the moving intensity together and are used for further classification with SVMs. Experiments have been conducted on benchmark datasets and show better performance than previous methods, which achieved 99.1% on Weizmann dataset and 95.8% on KTH dataset. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICIP | 3 |
| 2014 | Exposing video inter-frame forgery based on velocity field consistencyabstractIn recent years, video forensics has become an important issue. Video inter-frame forgery detection is a significant branch of forensics. In this paper, a new algorithm based on the consistency of velocity field is proposed to detect video inter-frame forgery (i.e., consecutive frame deletion and consecutive frame duplication). The generalized extreme studentized deviate (ESD) test is applied to identify the forgery types and locate the manipulated positions in forged videos. Experiments show the effectiveness of our algorithm. Yuxing Wu, Xinghao Jiang, Tanfeng Sun, Wan Wang |
ICASSP | 3 |
| 2014 | Inter-frame Video Forgery Detection Based on Block-Wise Brightness Variance Descriptor
Tanfeng Sun, Yun Q. Shi 0001 |
IWDW | 2 |
| 2013 | Unsupervised feature learning using Markov deep belief networkabstractRecently, deep architectures, such as Deep Belief Network (DBN), have been used to learn features from unlabeled data. However, since DBN supports bi-directional inference and the units between two layers are fully connected, it is difficult to directly apply the traditional convolutional network to DBN, or scale DBN to fit the large images (e.g. 1024×768). In this paper, a new deep learning model, named Markov DBN (MDBN), is proposed to address these problems. This model employs a new way for DBN to reduce computational burden and handle large images. Markov sub-layers are also adopted to take the neighboring relationship of the inputs into consideration. To train MDBN, we devise Block Restricted Boltzmann Machine (BRBM) which chooses non-overlapping blocks as input. Furthermore, SIFT descriptor is employed to enable this model to learn translation, scaling and rotation invariant features. Experimental results on datasets Caltech-101 and Caltech-256 have demonstrated the superiority of our model. Dongyang Cheng, Tanfeng Sun, Xinghao Jiang, Shi-Lin Wang |
ICIP | 2 |
| 2013 | Identifying Video Forgery Process Using Optical Flow
Wan Wang, Xinghao Jiang, Shi-Lin Wang, Meng Wan, Tanfeng Sun |
IWDW | 5 |
| 2013 | Estimation of the primary quantization parameter in MPEG videosabstractThe advanced technology and sophisticated software have rendered audiovisual content exposed to forgery, inspiring the emergence of multimedia forensic research. Since video tampering may involve double compression, the analysis of compression history is of significance. In this paper, we consider the processing chains of two compression steps and propose an algorithm that aims at identifying the quantization parameter used in the previous coding process. The method relies on the fact that characteristic footprints can be observed under different relationships between quantization parameters of consecutive compression operations. Features are extracted from both Discrete Cosine Transform (DCT) coefficients and their differential counterparts to capture the statistical disturbance. Experimental results demonstrate the effectiveness of our method. Wan Wang, Xinghao Jiang, Shi-Lin Wang, Tanfeng Sun |
VCIP | 4 |
| 2013 | An adaptive video shot segmentation scheme based on dual-detection model
Xinghao Jiang, Tanfeng Sun, Jin Liu 0016, Juan Chao, Wensheng Zhang 0002 |
Neurocomputing | 2 |
| 2013 | Detection of Double Compression in MPEG-4 Videos Based on Markov StatisticsabstractWith the spread of powerful and easy-to-use video editing software, digital videos are exposed to various forms of tampering. Nowadays, a considerable proportion of surveillance systems and video cameras have built-in MPEG-4 codec. Therefore, the detection of double compression in MPEG-4 videos as a first step in video forensics research is of significance. In this paper, Markov based features are adopted to detect double compression artifacts, which imply that the original video may have been interpolated. The advantages and limitations of double MPEG-4 compression detection are analyzed. Experimental results have demonstrated that our scheme outperforms most existing methods. Xinghao Jiang, Wan Wang, Tanfeng Sun, Yun Q. Shi 0001, Shi-Lin Wang |
IEEE Signal Process. Lett. | 3 |
| 2012 | Exposing video forgeries by detecting MPEG double compressionabstractIn this paper, an improved video tampering detection model based on MPEG double compression is proposed. Double compression will import disturbance into Discrete Cosine Transform (DCT) coefficients, reflecting in the violation of the parametric logarithmic law for first digit distribution of quantized Alternating Current (AC) coefficients. A 12-D feature can be extracted from each group of pictures (GOP) and machine learning framework is adopted to enhance the detection accuracy. Furthermore, a novel approach with a serial Support Vector Machine (SVM) architecture to estimate original bit rate scale in doubly compressed video is proposed. Experiments demonstrate higher accuracy and effectiveness. Tanfeng Sun, Wan Wang, Xinghao Jiang |
ICASSP | 1 |
| 2012 | A Novel Video Inter-frame Forgery Model Detection Scheme Based on Optical Flow Consistency
Juan Chao, Xinghao Jiang, Tanfeng Sun |
IWDW | 3 |
| 2012 | A Robust Image Classification Scheme with Sparse Coding and Multiple Kernel Learning
Dongyang Cheng, Tanfeng Sun, Xinghao Jiang |
IWDW | 2 |
| 2011 | An Video Shot Segmentation Scheme Based on Adaptive Binary Searching and SIFT
Xinghao Jiang, Tanfeng Sun, Jin Liu 0016, Wensheng Zhang 0002, Juan Chao |
ICIC (2) | 2 |
| 2011 | A Drift Compensation Algorithm for H.264/AVC Video Robust Watermarking Scheme
Xinghao Jiang, Tanfeng Sun, Yun Q. Shi 0001 |
IWDW | 2 |
| 2011 | A Novel Horror Scene Detection Scheme on Revised Multiple Instance Learning Model
Xinghao Jiang, Tanfeng Sun, Shanfeng Zhang, Xiqing Chu, Chuxiong Shen, Jingwen Fan |
MMM (2) | 3 |
| 2011 | An automatic video content classification scheme based on combined visual features model with modified DAGSVM
Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
Multim. Tools Appl. | 2 |
| 2008 | A Novel Real-Time MPEG-2 Video Watermarking Scheme in Copyright Protection
Xinghao Jiang, Tanfeng Sun, Jianhua Li 0001, Ye Yun |
IWDW | 2 |