EDBT 2026 Demo / reviewers in the wild / expert
Xiangling Ding
dblp:165/1264
· DBLP profile ↗
42ranked-venue papers
15as first author
29since 2021 · last 2026
0000-0002-6581-4633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 14 since 2021Computer networks · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Security and privacy · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SF-CFNet: Spatial-frequency collaborative fusion detection network for AI-generated satellite images
Xiangling Ding, Huanghuang Deng |
Expert Syst. Appl. | 3 |
| 2026 | Deep video inpainting and video inpainting detection: A comprehensive survey from deep learning perspectiveabstractWith the advances of Deep learning, the field of video inpainting has also made significant progress recently, leading to the emergence of deep learning-based video inpainting, also known as Deep video inpainting. It learns the potential rules or feature distributions of the video dataset in a data-driven manner to complete the missing areas in the video from a spatial-temporal perspective. Its original goal is to recover damaged or lost parts of videos, but it is also used to maliciously remove target objects. As a result, the development of Deep video inpainting has brought negative effects and potential threats to the country, society, and individuals. Therefore, the detection of this issue has also attracted wide research interests in the field of information security. The primary objective of this article is to provide a comprehensive summary of Deep video inpainting and the corresponding detection methods. Specifically, we classify existing Deep video inpainting methods into different categories from the perspective of their designed deep learning module, including 3D convolution-based, optical flow-based, alignment-based, temporal shift-based, attention-based, and diffusion-based network models. Meanwhile, we also sort existing research on Deep video inpainting detection into four categories: spatial-domain, temporal-domain, frequency-domain, and hybrid-domain network models, starting from a network feature analysis perspective. In addition, we review their training objectives, loss functions, and common benchmark datasets. We present video-level and pixel-level evaluation metrics, conduct a qualitative and quantitative evaluation, and discuss the advantages and disadvantages of representative Deep video inpainting and their corresponding detection methods. Finally, potential future research directions have been outlined for Deep video inpainting and its detection methods. Jizhou Yao, Yinqian Deng, Xiangling Ding |
J. Inf. Secur. Appl. | 6 |
| 2026 | Secure HEVC video steganography using IPMs spatial distribution and transfer probability
Ramadhani R. Iddy, Gaobo Yang, Dewang Wang, Xiangling Ding, Senzota K. Semakuwa |
Multim. Tools Appl. | 4 |
| 2026 | FALCON-Net: Feature Aggregation of Local Patterns for AI-Generated Image DetectionabstractWith the rapid development of generative models, the visual quality of generated images has become almost indistinguishable from real images, which poses a huge challenge to content authenticity verification. A key limitation of existing detectors is their reliance on model-specific cues, resulting in poor generalization to unseen models. Based on the observation of local differences in the generated images, we found that the generated images lack device-specific sensor noise and unnatural pixel intensity variations caused by the oversimplified generation process. These discrepancies provide important forensic cues for distinguishing between real and generated images. We propose the Feature Aggregation for Localized Context and Noise Network (FALCON-Net), which leverages these discrepancies to enhance detection capabilities. FALCON-Net integrates two complementary modules to enhance detection capabilities: the Intrinsic Noise Pattern Isolation (INP) module isolates device-specific noise patterns by analyzing high-frequency features in the frequency domain, while the Local Variation Pattern (LVP) module models the complex relationships between local pixels to capture directional intensity variations and reveal unnatural regularities in generated images. By combining these sensor-level and local structural cues, FALCON-Net identifies fundamental generative inconsistencies, ensuring robustness to post-processing and strong generalization to unseen models. Extensive experimental results show that FALCON-Net achieves the state-of-the-art performance in detecting generated images and shows good generalization ability to unseen generative models. The code is available at https://github.com/humiaomiaohaha/FALCON-Net. Dengyong Zhang, Changsheng Chen 0001, Jin Wang 0001, Yun Song, Gaobo Yang, Xin Liao 0001, Xiangling Ding |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2026 | Bi-Level Routing Attention and Enhanced Spatial-Temporal Inconsistency Learning for Deep VFI Video DetectionabstractWith the maturation of Deep Learning-based Video Frame Interpolation (Deep VFI), the left spatial-temporal inconsistency in the synthesis process is greatly improved, which poses a challenge to the current VFI detector. This article presents a dual-stream identification network based on Bi-level Routing Attention and enhanced Spatial-Temporal inconsistency learning (BRA-ST) to address this challenge. Specifically, the spatial inconsistencies in Deep VFI are mainly reflected in their motion regions and moving object edges; thus, the high-pass filter is introduced to enhance them, facilitating the three-stage pyramid structure of BiFormer Blocks with bi-level routing attention in the frame-level stream to learn. To fully exploit the temporal inconsistencies in the Deep VFI video, the time-difference module in the time-level stream is superimposed with the ConvGRU to extract the temporally dependent features of continuous multiple frames. Additionally, the middle layer of the two streams interacts and aggregates with the channel attention, and then, their last layer adaptively merges from a whole and part perspective for the ultimate frame prediction. Finally, the experimental findings on a constructed dataset by the five most advanced Deep VFI methods indicate that the proposed BRA-ST achieved \(F_{\text{1Score}}\) of 99.73%, which is superior to the existing Deep VFI detectors, and further verify that the resolution of BRA-ST for different Deep VFI methods reached 78.55%. Our source codes and dataset are available at https://pan.baidu.com/s/1f05_gS0qu5G-SSIkd9F4Hw?pwd=j6t6 . Xiangling Ding, Yunyi Li, Gaobo Yang, Yubo Lang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Progressive Reverse Attention Network for image inpainting detection and localization
Jiyou Chen, Xiangling Ding, Gaobo Yang |
Comput. Vis. Image Underst. | 3 |
| 2025 | Spectral information guidance network for tampering localization of high-resolution satellite map
Xiangling Ding, Yuchen Nie |
Expert Syst. Appl. | 1 |
| 2025 | NG-RED:Nonconvex group-matrix residual denoising learning for image restoration
Yunyi Li, Huijuan Wu, Xiangling Ding |
Expert Syst. Appl. | 4 |
| 2025 | A styleGAN-based face de-morphing network for restoring accomplice's facial image
Juan Cai, Min Long 0003, Quantao Yao, Xiangling Ding |
Multim. Syst. | 5 |
| 2025 | Higher-order motion calibration and sparsity based outlier correction for video FRUC
Jiale He, Qunbing Xia, Gaobo Yang, Xiangling Ding |
Signal Process. Image Commun. | 4 |
| 2025 | Video Frame Interpolation via Fast Bidirectional 3D Correlation VolumeabstractRecently, there has been a growing demand for flow-based video frame interpolation methods, which introduce correlation volumes to supervise the correlation of bidirectional optical flows. However, they often overlook the symmetry of the bidirectional motion field by consuming substantial computational cost, which is reflected in the fact that these methods often require a long runtime. To address these issues, in this article, we propose a bidirectional 3D correlation volume which is suitable for video frame interpolation. By decomposing the 4D correlation volume into two 3D correlation volumes in the horizontal and vertical directions, we significantly enhance the model’s inference speed with a minor sacrifice compared to our baseline. Additionally, when handling 2K video frames, our method achieves several-fold improvement in inference speed compared to other methods which implied correlation volume. The code is available at https://github.com/famt0531 . Dengyong Zhang, Runqi Lou, Xiangling Ding, Xin Liao 0001, Gaobo Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Spatiotemporal Inconsistency Learning and Interactive Fusion for Deepfake Video DetectionabstractWith the rise of the metaverse, the rapid advancement of Deepfakes technology has become closely intertwined. Within the metaverse, individuals exist in digital form and engage in interactions, transactions, and communications through virtual avatars. However, the development of Deepfakes technology has led to the proliferation of forged information disseminated under the guise of users’ virtual identities, posing significant security risks to the metaverse. Hence, there is an urgent need to research and develop more robust methods for detecting deep forgeries to address these challenges. This article explores deepfake video detection by leveraging the spatiotemporal inconsistencies generated by deepfake generation techniques, thereby proposing the interactive spatiotemporal inconsistency learning and interactive fusion (ST-ILIF) detection method, which consists of phase-aware and sequence streams. The spatial inconsistencies exhibited in frames of deepfake videos are primarily attributed to variations in the structural information contained within the phase component of the Fourier domain. To mitigate the issue of overfitting the content information, a phase-aware stream is introduced to learn the spatial inconsistencies from the phase-based reconstructed frames. Additionally, considering that deepfake videos are generated frame by frame and lack temporal consistency between frames, a sequence stream is proposed to extract temporal inconsistency features from the spatiotemporal difference information between consecutive frames. Finally, through feature interaction and fusion of the two streams, the representation ability of intermediate and classification features is further enhanced. The proposed method, which was evaluated on four mainstream datasets, outperformed most existing methods, and extensive experimental results demonstrated its effectiveness in identifying deepfake videos. Our source code is available at https://github.com/qff98/Deepfake-Video-Detection . Dengyong Zhang, Xin Liao 0001, Feifan Qi, Gaobo Yang, Xiangling Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | One-Class HEVC Double Compression Detection with Same Coding ParametersabstractHigh Efficiency Video Coding (HEVC) double compression detection with the same coding parameters can be regarded as one principal procedure to analyze the integrity of HEVC-coded videos. Therefore, a one-class classification (OCC)-based hybrid heterogeneous network with a shallow convolutional neural network (CNN) and a six-node graph neural network (GNN), which only needs the pristine videos, is proposed in this paper. Concretely, the shallow CNN learns the subtle fluctuation of pixel values due to double compression, while the six-node GNN is developed to represent the local-global relationship of various divided patches and the number and position of zero-value pixels in a high-frequency component of the motion-aligned residual. Finally, the output vectors are fused into a specially designed semi-supervised Mahalanobis distance-based OCC to obtain the detection results. The experimental result demonstrates that the proposed method, which only learns features from single compressed videos, outperforms the existing state-of-the-art detection methods and other more complex OCC methods. Xiangling Ding |
ICME | 2 |
| 2024 | GAN-based adaptive cost learning for enhanced image steganography security
Dewang Wang, Gaobo Yang, Jiyou Chen, Xiangling Ding |
Expert Syst. Appl. | 4 |
| 2024 | Multi-scale noise-guided progressive network for image splicing detection and localization
Dengyong Zhang, Ningjing Jiang, Feng Li 0065, Xin Liao 0001, Gaobo Yang, Xiangling Ding |
Expert Syst. Appl. | 7 |
| 2024 | One-Class Hybrid Heterogeneous Network for Detecting HEVC Double Compression With the Same Coding ParametersabstractHigh Efficiency Video Coding (HEVC) is a recent yet increasing widely-used video coding standard, and double compression detection is usually an essential step to verify the integrity of HEVC-encoded videos. However, it is challenging due to fewer traces left by HEVC double compression with the same parameters. Moreover, existing full-supervised learning works for HEVC double compression detection are inefficient because they depend on large amounts of labeled pristine and forged videos, which are difficult to be collected. To address these issues, a One-Class Classification (OCC)-based hybrid heterogeneous network is proposed, which only needs the pristine videos. We first develop a modulation layer with both motion alignment and high-frequency preservation operations, which serves as an effective metric to evaluate the differences between those videos compressed once and twice. Then, a heterogeneous network with a shallow Convolutional Neural Network (CNN) and a six-node Graph Neural Network (GNN) is proposed. Specifically, the shallow CNN, which pays more attention to medium or fast-motion regions, learns from subtle fluctuations of pixel values caused by double compression, whereas GNN, which focuses on static and slow-motion regions, is developed to represent the local-global relationship of video patches and the distribution of zero-value pixels in the high-frequency components of motion-aligned residuals. Due to the motion-aware mechanism, the proposed approach only learns features from single compressed videos. Extensive experimental results show that the proposed approach outperforms the state-of-the-art full-supervised learning works and other more complex OCC works. Xiangling Ding, Yunyi Li |
IEEE Internet Things J. | 1 |
| 2024 | AFTLNet: An efficient adaptive forgery traces learning network for deep image inpainting localization
Xiangling Ding, Yingqian Deng, Wenyi Zhu |
J. Inf. Secur. Appl. | 1 |
| 2024 | A convolutional neural network based on noise residual for seam carving detection
Dengyong Zhang, Zhenyu Lv, Feng Li 0065, Xiangling Ding, Gaobo Yang |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | Contour-assistance-based video matting localization
Wenyi Zhu, Xiangling Ding, Zhang Chao, Yingqian Deng |
Multim. Syst. | 2 |
| 2024 | ERaL: Exceptional Regions-Aware Deep Video Interpolation LocalizationabstractDeep learning-based video frame interpolation (DVFI) can generate high frame-rate video sequences with high temporal consistency, usually producing visually plausible results. As DVFI can also be deployed for vicious video operations, it has misled users' visits and invalidates near-duplicate video detection. Therefore, it is urgent to locate the interpolated frames subjected to DVFI techniques. This letter investigates this issue by exploiting exceptional regions-aware localization (ERaL). In particular, we guide ERaL with an “inverted Z-shaped” network, which can better capture the position and intensity of exceptional regions regardless of the specific DVFI method, coming from the fact that the faked frame rate videos collapse even if any DVFI methods generate them, as ERaL only learns over original videos. Then, a hierarchical feature extraction is developed, integrating the feature enhancement, simplified transformer, and inverted residual feed-forward network, to produce a frame-wise localization of the interpolated frames for a given sequence. The proposed method is evaluated with counterfeited videos manipulated by three state-of-the-art DVFI approaches. Extensive experimental results demonstrate that the proposed method can effectively localize the interpolated frames, surpassing existing algorithms. Xiangling Ding, Dengyong Zhang, Gaobo Yang |
IEEE Signal Process. Lett. | 1 |
| 2023 | Video Frame Interpolation via Multi-scale Expandable Deformable ConvolutionabstractVideo frame interpolation is a challenging task in the video processing field. Benefiting from the development of deep learning, many video frame interpolation methods have been proposed, which focus on sampling pixels with useful information to synthesize each output pixel using their own sampling operation. However, these works have data redundancy limitations and fail to sample the correct pixel of complex motions. To solve these problems, we propose a new warping framework to sample called multi-scale expandable deformable convolution(MSEConv) which employs a deep fully convolutional neural network to estimate multiple small-scale kernel weights with different expansion degrees and adaptive weight allocation for each pixel synthesis. MSEConv covers most prevailing research methods as special cases of it, thus MSEConv is also possible to be transferred to existing works for performance improvement. To further improve the robustness of the whole network to occlusion, we also introduce a data preprocessing method for mask occlusion in video frame interpolation. Quantitative and qualitative experiments show that our method shows a robust performance comparable to or even superior to the state-of-the-art method. Our source code and visual comparable results are available at https://github.com/Pumpkin123709/MSEConv. Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Gaobo Yang |
IH&MMSec | 3 |
| 2023 | SRTNet: a spatial and residual based two-stream neural network for deepfakes detection
Dengyong Zhang, Xiangling Ding, Gaobo Yang, Feng Li 0065, Zelin Deng, Yun Song |
Multim. Tools Appl. | 3 |
| 2023 | L2BEC2: Local Lightweight Bidirectional Encoding and Channel Attention Cascade for Video Frame InterpolationabstractVideo frame interpolation (VFI) is of great importance for many video applications, yet it is still challenging even in the era of deep learning. Some existing VFI models directly exploit existing lightweight network frameworks, thus making synthesized in-between frames blurry and creating artifacts due to imprecise motion representation. The other existing VFI models typically depend on heavy model architectures with a large number of parameters, preventing them from being deployed on small terminals. To address these issues, we propose a local lightweight VFI network ( L 2 BEC 2 ) that leverages bidirectional encoding structure with channel attention cascade. Specifically, we improve visual quality by introducing a forward and backward encoding structure with channel attention cascade to better characterize motion information. Furthermore, we introduce a local lightweight strategy into the state-of-the-art Adaptive Collaboration of Flows (AdaCoF) model to simplify its model parameters. Compared with the original AdaCoF model, the proposed L 2 BEC 2 obtains performance gain at the cost of only one-third of the number of parameters and performs favorably against the state-of-the-art works on public datasets. Our source code is available at https://github.com/Pumpkin123709/LBEC.git . Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Yun Song, Gaobo Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Video Frame Interpolation via Local Lightweight Bidirectional Encoding with Channel Attention CascadeabstractDeep Neural Networks based video frame interpolation, synthesizing in-between frames given two consecutive neighboring frames, typically depends on heavy model architectures, preventing them from being deployed on small terminals. When directly adopting the lightweight network architecture from these models, the synthesized frames may suffer from poor visual appearance. In this paper, a lightweight-driven video frame interpolation network (L2BEC2) is proposed. Concretely, we first improve the visual appearance by introducing the bidirectional encoding structure with channel attention cascade to better characterize the motion information; then we further adopt the local network lightweight idea into the aforementioned structure to significantly eliminate its redundant parts of the model parameters. As a result, our L2BEC2performs favorably at the cost of only one third of the parameters compared with the state-of-the-art methods on public datasets. Our source code is available at https://github.com/Pumpkin123709/LBEC.git. Xiangling Ding, Pu Huang 0002, Dengyong Zhang, Xianfeng Zhao |
ICASSP | 1 |
| 2022 | DeepFake Videos Detection via Spatiotemporal Inconsistency Learning and Interactive FusionabstractWhile the rapid expansion of DeepFake generation techniques has arisen a serious impact on human society, the detection of DeepFake videos is challenging because of their highly plausible contents on each frame, which are not visually apparent. To address that, this paper proposes a two-stream method to capture the spatial-temporal inconsistency cues, and then interactively fuse them to detect DeepFake videos. Since the traces of spatial inconsistency in DeepFake video frames mainly appear in their structural information, which reflects by the phase component in the frequency domain, the proposed frame-level stream learns the spatial inconsistency from the phase-based reconstructed frames to avoid fitting the content information. Aiming at the problem that the temporal inconsistency in DeepFake videos might be ignored, the temporality-level stream is proposed to extract the temporal correlation feature by the temporal difference networks and stacked ConvGRU module on consecutive multiple frames. When interacted with channel attention in the intermediate layer of two streams, and adaptively fused with the discriminative features of two streams from a global-local perspective, our proposed method performs better than the state-of-the-art detection methods. Xiangling Ding, Dengyong Zhang |
SECON | 1 |
| 2022 | Forgery Detection Scheme of Deep Video Frame-rate Up-conversion Based on Dual-stream Multi-scale Spatial-temporal RepresentationabstractVideo Frame-Rate Up-Conversion (FRUC) is originally designed to produce the high frame-rate video by periodically inserting new frames between two adjacent frames. However, it can also be utilized to synthesize the faked high frame-rate videos or spliced videos for malicious intents. The existing FRUC detection methods can efficiently identify its occurrence by exploring the blurring effects or deformed structures left over from the traditional FRUC methods. But the video FRUC has been substantially improved in the past years, especially in this deep learning era. Deep learning-based video FRUC (Deep FRUC) weakens the visual traces of the traditional ones such that it is challenging for the current FRUC detectors. In this paper, we propose a forensics algorithm based on dual-stream multi-scale spatial-temporal representation for Deep FRUC. Specifically, we develop a multi-scale receptive field strategy, and attention scheme to learn hidden tampered trace and spatial-temporal representation, respectively. Besides, frame residual and noise residual attention streams are complementary learning from the spatial-temporal dimension, in which the former can capture contrast differences and tampered abnormal boundaries, while the latter can explode noise inconsistencies between original and tampered frames. Experimental results show that the proposed algorithm can effectively achieve the best detection accuracy compared with the existing FRUC forensics works under the Deep FRUC video datasets. Xiangling Ding, Dengyong Zhang |
TrustCom | 2 |
| 2022 | Robust detection of dehazed images via dual-stream CNNs with adaptive feature fusion
Jiyou Chen, Gaobo Yang, Xiangling Ding, Zhiqing Guo |
Comput. Vis. Image Underst. | 3 |
| 2021 | Detection of Deep Video Frame Interpolation via Learning Dual-Stream Fusion CNN in the Compression DomainabstractDeep learning-based Video Frame Interpolation (Deep VFI) diminishes the visual traces of the conventional one such that it is challenging for the current VFI detectors. Therefore, it is necessary to identify the presence of deep interpolated frames (DIF) in a video. This paper proposed a hybrid neural network to localize the DIF by learning spatio-temporal representations from the residual and motion vector information in the compression domain. Firstly, the residual and motion vector of motion regions are maintained by an intra-prediction constraints. Then, inherent tampering traces are further highlighted through subtracting the estimate of the residual or motion vector by virtue of residual modulation or MV refinement network. Finally, an attention-based dual-stream network is designed to jointly learn discriminative representations from the enhancement traces. Deep VFI video datasets created by the state-of-the-art deep VFI methods, have been evaluated, and extensive experimental results clearly demonstrate that our approach can achieve state-of-the-art performance compared with conventional methods. Xiangling Ding, Yifeng Pan, Jiyou Chen, Gaobo Yang, Yimao Xiong |
ICME | 1 |
| 2021 | Localization of Deep Video Inpainting Based on Spatiotemporal Convolution and Refinement NetworkabstractDeep learning-based video inpainting can fill the missing or undesired regions with spatial-temporal consistent contents without obvious visually distortion. Although the original purpose of deep inpainting is to repair flawed videos, it can also be adopted for malicious purposes, e.g., removal of specific objects. Therefore, automatically locating the inpainted regions is a challenging task in video forensics. This paper proposes a new forensic refinement framework to localize the deep inpainted regions by considering the spatial-temporal viewpoint. Firstly, we design a spatiotemporal convolution to suppress redundancy for highlighting deep inpainting traces. Then, a detection module is constructed with four concatenated ResNet blocks, and two upsampling layers to achieve a rough location map. Finally, a modified U-net based refinement module is developed for the pixel-wise localization map. Deep inpaiting video datasets created by the state-of-the-art deep inpainting method, have been evaluated, and extensive experimental results clearly demonstrate the efficacy of the proposed approach. Xiangling Ding, Yifeng Pan, Kui Luo, Yanming Huang, Junlin Ouyang, Gaobo Yang |
ISCAS | 1 |
| 2020 | Identification of Frame-Rate Up-Conversion Based on Spatial-Temporal Edge and Occlusion with Convolutional Neural NetworkabstractFrame-Rate Up-Conversion (FRUC) is a frame-based video manipulation where synthesized frames are periodic inserted into two consecutive frames to increase the frame-rate of videos. Thus, it can be employed by counterfeiters for falsifying high frame rate videos or splicing videos with different frame-rates. Automatically identifying the adopted FRUC techniques from a serial of faked frame-rate videos while achieving a stable recognition rate is a challenging task in video forensics. We propose a new forensic framework, named Spatial-Temporal Edge-Occlusion Identification of FRUC (STEO-IoF), by considering the capabilities of both clues: spatial-temporal edge and occlusion maps, to guide a forensic model using prior knowledge transferred from the existing steganalysis convolutional neural network (CNN) model for the recognition of adopted FRUC techniques. Three open FRUC softwares and four representative FRUC techniques have been evaluated, and experimental results clearly show that the efficacy of the proposed approach. Xiangling Ding, Yanming Huang |
ISCAS | 1 |
| 2020 | Forgery detection of motion compensation interpolated frames based on discontinuity of optical flow
Xiangling Ding, Yanming Huang, Yue Li 0016, Jiale He |
Multim. Tools Appl. | 1 |
| 2020 | Seam-Carved Image Tampering Detection Based on the Cooccurrence of Adjacent LBPsabstractSeam carving has been widely used in image resizing due to its superior performance in avoiding image distortion and deformation, which can maliciously be used on purpose, such as tampering contents of an image. As a result, seam-carving detection is becoming crucially important to recognize the image authenticity. However, existing methods do not perform well in the accuracy of seam-carving detection especially when the scaling ratio is low. In this paper, we propose an image forensic approach based on the cooccurrence of adjacent local binary patterns (LBPs), which employs LBP to better display texture information. Specifically, a total of 24 energy-based, seam-based, half-seam-based, and noise-based features in the LBP domain are applied to the seam-carving detection. Moreover, the cooccurrence features of adjacent LBPs are combined to highlight the local relationship between LBPs. Besides, SVM after training is adopted for feature classification to determine whether an image is seam-carved or not. Experimental results demonstrate the effectiveness in improving the detection accuracy with respect to different scaling ratios, especially under low scaling ratios. Dengyong Zhang, Feng Li 0065, Arun Kumar Sangaiah, Xiangling Ding |
Secur. Commun. Networks | 5 |
| 2020 | Self-learning residual model for fast intra CU size decision in 3D-HEVC
Yue Li 0016, Ningbo Zhu, Gaobo Yang, Yapei Zhu, Xiangling Ding |
Signal Process. Image Commun. | 5 |
| 2020 | Spatio-temporal Saliency-based Motion Vector Refinement for Frame Rate Up-conversionabstractA spatio-temporal saliency-based frame rate up-conversion (FRUC) approach is proposed, which achieves better quality of interpolated frames and invalidates existing texture variation-based FRUC detectors. A spatio-temporal saliency model is designed to select salient frames. After obtaining initial motion vector field by texture- and color-based bilateral motion estimation, two motion vector refining (MVR) schemes are adopted for high and low saliency frames to hierarchically refine the motion vectors, respectively. To produce high-quality interpolated frames, image enhancement are performed for salient frames after frame interpolation. Due to distinct MVR schemes, there are different degrees of texture information in interpolated frames. Some edge and texture information is supplemented into salient frames as post-processing, which can invalidate existing texture variation-based FRUC detectors. Experimental results show that the proposed approach outperforms state-of-the-art works in both objective and subjective qualities of interpolated frames, and achieves the purpose of FRUC anti-forensics. Jiale He, Gaobo Yang, Xiangling Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | An Efficient ECG Denoising Method Based on Empirical Mode Decomposition, Sample Entropy, and Improved Threshold FunctionabstractThe electrocardiogram (ECG) signal can easily be affected by various types of noises while being recorded, which decreases the accuracy of subsequent diagnosis. Therefore, the efficient denoising of ECG signals has become an important research topic. In the paper, we proposed an efficient ECG denoising approach based on empirical mode decomposition (EMD), sample entropy, and improved threshold function. This method can better remove the noise of ECG signals and provide better diagnosis service for the computer-based automatic medical system. The proposed work includes three stages of analysis: (1) EMD is used to decompose the signal into finite intrinsic mode functions (IMFs), and according to the sample entropy of each order of IMF following EMD, the order of IMFs denoised is determined; (2) the new threshold function is adopted to denoise these IMFs after the order of IMFs denoised is determined; and (3) the signal is reconstructed and smoothed. The proposed method solves the shortcoming of discarding the first-order IMF directly in traditional EMD denoising and proposes a new threshold denoising function to improve the traditional soft and hard threshold functions. We further conduct simulation experiments of ECG signals from the MIT-BIH database, in which three types of noise are simulated: white Gaussian noise, electromyogram (EMG), and power line interference. The experimental results show that the proposed method is robust to a variety of noise types. Moreover, we analyze the effectiveness of the proposed method under different input SNR with reference to improving SNR ( SNR imp ) and mean square error ( MSE ), then compare the denoising algorithm proposed in this paper with previous ECG signal denoising techniques. The results demonstrate that the proposed method has a higher SNR imp and a lower MSE . Qualitative and quantitative studies demonstrate that the proposed algorithm is a good ECG signal denoising method. Dengyong Zhang, Feng Li 0065, Shang Tian, Jin Wang 0001, Xiangling Ding, Rongrong Gong |
Wirel. Commun. Mob. Comput. | 6 |
| 2019 | Detection of motion compensated frame interpolation via motion-aligned temporal difference
Xiangling Ding, Yue Li 0016, Jiale He, Gaobo Yang |
Multim. Tools Appl. | 1 |
| 2019 | Robust Localization of Interpolated Frames by Motion-Compensated Frame Interpolation Based on an Artifact Indicated Map and Tchebichef MomentsabstractMotion-compensated frame interpolation (MCFI), a frame-interpolation technique to increase the motion continuity of low frame-rate video, can be utilized by counterfeiters for faking high bitrate video or splicing videos with different frame rates. For existing MCFI detectors, their performances are degraded under real-world scenarios such as H.264/AVC compression, noise, or blur. To address this issue, a robust MCFI detector is proposed to locate interpolated frames. By analyzing the distribution of residual energies within interpolated frames, we observe that there exist strong correlations between artifact regions and high residual energies. Thus, an artifact indicated map is introduced to select candidate artifact regions. Then, Tchebichef moments (TMs) are exploited to characterize the blurring effects or deformed structures among these regions. Specifically, the mean value of absolute high-order TMs of selected regions is used to model these temporal inconsistencies. Finally, a sliding window is adopted to locate interpolated frames, which are further refined by three post-processing operations. Chrominance information is also integrated with luminance information for robust identification of interpolated frames. Extensive experimental results show that compared with the state-of-the-art MCFI detectors, the proposed approach is more robust for compressed videos under various real-world scenarios. Xiangling Ding, Ningbo Zhu, Leida Li, Yue Li 0016, Gaobo Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Identification of Motion-Compensated Frame Rate Up-Conversion Based on Residual SignalsabstractMotion-compensated frame rate up-conversion (MC-FRUC) is originally presented to increase the motion continuity of low frame rate videos by periodically inserting new frames, which improves the viewing experience. However, MC-FRUC can also be exploited to fake high frame rate videos or splice two videos with different frame rates for malicious purposes. A blind forensics approach is proposed for the identification of various MC-FRUC techniques. A theoretical model is first built for residual signal, which is exploited as tampering trace for blind forensics. The identification of various MC-FRUC techniques is then converted into a problem of discriminating the differences of residual signals among them. A pre-classifier is designed to suppress the side effects of original frames and static interpolated frames in candidate videos. Then, spatial and temporal Markov statistics features are extracted from the residual signals inside the interpolated frames for MC-FRUC identification. Five open MC-FRUC softwares and six representative MC-FRUC techniques have been tested, and experimental results show that the proposed approach can effectively locate interpolated frames and further identify the adopted MC-FRUC technique for both uncompressed videos and compressed videos with high perceptual qualities. Xiangling Ding, Gaobo Yang, Ran Li 0003, Yue Li 0016, Xingming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Probability Model-Based Early Merge Mode Decision for Dependent Views Coding in 3D-HEVCabstractAs a 3D extension to the High Efficiency Video Coding (HEVC) standard, 3D-HEVC was developed to improve the coding efficiency of multiview videos. It inherits the prediction modes from HEVC, yet both Motion Estimation (ME) and Disparity Estimation (DE) are required for dependent views coding. This improves coding efficiency at the cost of huge computational costs. In this article, an early Merge mode decision approach is proposed for dependent texture views and dependent depth maps coding in 3D-HEVC based on priori and posterior probability models. First, the priori probability model is established by exploiting the hierarchical and interview correlations from those previously encoded blocks. Second, the posterior probability model is built by using the Coded Block Flag (CBF) of the current coding block. Finally, the joint priori and posterior probability model is adopted to early terminate the Merge mode decision for both dependent texture views and dependent depth maps coding. Experimental results show that the proposed approach saves 45.2% and 30.6% encoding time on average for dependent texture views and dependent depth maps coding while maintaining negligible loss of coding efficiency, respectively. Yue Li 0016, Gaobo Yang, Yapei Zhu, Xiangling Ding, Rongrong Gong |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Design of new scan orders for perceptual encryption of H.264/AVC videosabstractIn this study, a perceptual encryption algorithm is proposed for H.264/AVC video to enhance the scrambling effect and encryption space. Six new scan orders are designed for H.264/AVC encoder by analysing the energy distribution of discrete cosine transform coefficients. They are proven to have similar performance as the conventional zigzag scan order and its symmetrical scan order. These six new scan orders are combined with two existing scan orders to design a scan‐order based perceptual encryption algorithm. Specifically, video encryption is achieved more specifically by randomly selecting one scan order from the eight scan orders with a security key, and the sign bit flipping of DC coefficients is also incorporated to further increase the encryption space. Experimental results show that the proposed approach has the advantages of both low bitrate increase and low computational cost. Furthermore, it is more flexible and has stronger security than the existing scan‐order based video encryption schemes. Xiangling Ding, Yingzhuo Deng, Gaobo Yang, Yun Song, Dajiang He, Xingming Sun |
IET Inf. Secur. | 1 |
| 2017 | Unimodal Stopping Model-Based Early SKIP Mode Decision for High-Efficiency Video CodingabstractHigh-efficiency video coding (HEVC) can greatly improve coding efficiency compared with the prior video coding standard H.264/AVC by adopting advanced hierarchical coding structures such as coding unit (CU), prediction unit (PU), and transform unit. For each CU, an exhaustive mode decision strategy is adopted to achieve the best rate distortion (RD) cost, which simultaneously results in enormous computational complexity. In this paper, an early SKIP mode decision algorithm is proposed for the HEVC encoder to speed up the process of mode decision. Each CU size is categorized into either rare used or frequent used by exploiting the correlation of CU depth, which is estimated from the temporally colocated CUs. For the rare-used CU size, the SKIP mode is directly selected as the optimal mode and the remaining mode decision process is early terminated. For the frequent-used CU size, a unimodal stopping model is designed for its early SKIP mode decision by exploiting both hierarchical mode structure and RD cost property. Experimental results show that the proposed early SKIP mode decision method achieves average 58.5% and 54.8% encoding time savings, while the Bjontegaard Delta bit rate only increases average 0.8% and 0.8% for various test sequences under the random access and the low delay B conditions, respectively. Yue Li 0016, Gaobo Yang, Yapei Zhu, Xiangling Ding, Xingming Sun |
IEEE Trans. Multim. | 4 |
| 2015 | An efficient forgery detection algorithm for object removal by exemplar-based image inpainting
Zaoshan Liang, Gaobo Yang, Xiangling Ding, Leida Li |
J. Vis. Commun. Image Represent. | 3 |