Dengyong Zhang

dblp:165/8955 · DBLP profile ↗
← Back
49ranked-venue papers
20as first author
37since 2021 · last 2026
0000-0002-2789-2980ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 18 since 2021Security and privacy · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Computer networks · 5 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 GAMD-Net: A Geometric-Aware Micro-Defect Network for Precision Inspection of Power Transmission Infrastructures
Yubin Tang, Yuxu Peng, Dengyong Zhang, Lei Wang 0143
ICIC (21)4
2026 DocForensic: A VMamba-Wavelet Cross-Attention Network for Document Image Tampering Localization
abstract
The proliferation of sophisticated image editing tools has significantly reduced the barrier to document image forgery, necessitating robust tampering localization techniques. Existing deep learning methods often struggle to balance global semantic understanding with fine-grained texture capture efficiently. This paper proposes the DocForensic, a VMamba-Wavelet Cross-Attention Network for document image tampering localization. We introduce a Visual State Space Model (VSSM) module that models long-range dependencies with linear complexity via state space sequences, enabling a comprehensive understanding of the global semantic context. Simultaneously, a Wavelet Transformation Multi-Scale Convolution (WTMC) module isolates and enhances high-frequency sub-bands to capture local texture inconsistencies. A Bidirectional Cross-Modality Attention (CMA) module fuses these features adaptively using forward and reverse attention. Extensive experiments show that our network achieves state-of-the-art performance and exhibits superior robustness against common image processing attacks such as JPEG compression and noise.
Dengyong Zhang
ICMR3
2026 Multi-Scale Perception and Channel-Gated Fusion Network for Aerial Image Detection
Zhixiang Zheng, Pei Yi, Dengyong Zhang, Yuxu Peng, Lei Wang 0143
ICMR4
2026 FALCON-Net: Feature Aggregation of Local Patterns for AI-Generated Image Detection
abstract
With the rapid development of generative models, the visual quality of generated images has become almost indistinguishable from real images, which poses a huge challenge to content authenticity verification. A key limitation of existing detectors is their reliance on model-specific cues, resulting in poor generalization to unseen models. Based on the observation of local differences in the generated images, we found that the generated images lack device-specific sensor noise and unnatural pixel intensity variations caused by the oversimplified generation process. These discrepancies provide important forensic cues for distinguishing between real and generated images. We propose the Feature Aggregation for Localized Context and Noise Network (FALCON-Net), which leverages these discrepancies to enhance detection capabilities. FALCON-Net integrates two complementary modules to enhance detection capabilities: the Intrinsic Noise Pattern Isolation (INP) module isolates device-specific noise patterns by analyzing high-frequency features in the frequency domain, while the Local Variation Pattern (LVP) module models the complex relationships between local pixels to capture directional intensity variations and reveal unnatural regularities in generated images. By combining these sensor-level and local structural cues, FALCON-Net identifies fundamental generative inconsistencies, ensuring robustness to post-processing and strong generalization to unseen models. Extensive experimental results show that FALCON-Net achieves the state-of-the-art performance in detecting generated images and shows good generalization ability to unseen generative models. The code is available at https://github.com/humiaomiaohaha/FALCON-Net.
Dengyong Zhang, Changsheng Chen 0001, Jin Wang 0001, Yun Song, Gaobo Yang, Xin Liao 0001, Xiangling Ding
IEEE Trans. Inf. Forensics Secur.1
2025 GC-ConsFlow: Leveraging Optical Flow Residuals and Global Context for Robust Deepfake Detection
abstract
The rapid development of Deepfake technology has enabled the generation of highly realistic manipulated videos, posing severe social and ethical challenges. Existing Deepfake detection methods primarily focused on either spatial or temporal inconsistencies, often neglecting the interplay between the two or suffering from interference caused by natural facial motions. To address these challenges, we propose the global context consistency flow (GC-ConsFlow), a novel dual-stream framework that effectively integrates spatial and temporal features for robust Deepfake detection. The global grouped context aggregation module (GGCA), integrated into the global context-aware frame flow stream (GCAF), enhances spatial feature extraction by aggregating grouped global context information, enabling the detection of subtle, spatial artifacts within frames. The flow-gradient temporal consistency stream (FGTC), rather than directly modeling the residuals, it is used to improve the robustness of temporal feature extraction against the inconsistency introduced by unnatural facial motion using optical flow residuals and gradient-based features. By combining these two streams, GC-ConsFlow demonstrates the effectiveness and robustness in capturing complementary spatiotemporal forgery traces. Extensive experiments show that GC-ConsFlow outperforms existing state-of-the-art methods in detecting Deepfake videos under various compression scenarios.
Dengyong Zhang, Jingyang Meng
ICME3
2025 Wavelet Convolution and Multi-Scale Attention Network for Image Tampering Localization
abstract
Conventional tampering localization only extracts features in the image domain, which makes it hard to capture the subtle tampering traces. In this paper, we propose a wavelet convolution and multi-scale attention network (WCMA-Net) for image tampering detection and localization, in which a wavelet convolution module (WCM) branch and a multi-scale attention module (MSAM) branch are integrated following the backbone. In the WCM branch, wavelet decomposition is utilized to enhance high-frequency details and enhance the detection of subtle tampering traces. In the MSAM branch, a multi-scale attention operation is employed to extract global and local features, which are then combined according to their similarity to capture long-range dependencies among pixels. Finally, an adaptive weight strategy is employed to fuse the features from both branches for binary pixel-level tampering mask prediction. Experimental results on various public datasets demonstrate that the proposed method achieves superior precise pixel-level image tampering localization over state-of-the-art methods. Codes and models are available at https://github.com/csust-sonie/WCMA-Net.
Yun Song, Yaoyao Xu, Dengyong Zhang, Miaohui Wang
ICME5
2025 MDC-Net: Multi-dimensional Cross-Domain Collaborative Network for Image Manipulation Localization
Yawen Wei, Shengxin Cai, Dengyong Zhang
SecureComm (5)4
2025 Robust face forgery detection integrating local texture and global texture information
abstract
Facial forgery technology is advancing rapidly, leading to significant social security concerns. In recent years, as forgery technologies and types continue to emerge, many methods struggle to strike a balance between accuracy and robustness. Most existing methods rely on CNN to extract high-quality forged face clues but often overlook inherent forgery traces. Consequently, they may overfit the training dataset and perform poorly on data from diverse sources or be subjected to various post-processing operations. To address this challenge, we propose leveraging multi-scale texture information to expose subtle artifacts in RGB space. To achieve this, we devise a two-stream detection architecture that integrates texture features and RGB features. We analyze forgery traces using global large texture information and local detailed texture information separately. Additionally, we design a feature pyramid to fuse these two texture features and employ an attention mechanism to enhance the features of both streams. By examining forgery traces from multiple perspectives, we have developed an adaptive feature fusion module to facilitate interactive feature fusion between the two streams. We conduct extensive experiments on various benchmark datasets and compare our method with recent state-of-the-art (SOTA) methods to demonstrate its effectiveness. Our code will be provided at https://github.com/hryyyy/MST .
Rongrong Gong, Ruiyi He, Dengyong Zhang, Arun Kumar Sangaiah, Mohammed J. F. Alenazi
EURASIP J. Inf. Secur.3
2025 Face Forgery Detection via Multi-Scale and Multi-Domain Features Fusion
abstract
ABSTRACT Deepfake, as a popular form of visual forgery technique on the Internet, poses a serious threat to individuals' data privacy and security. In consumer electronics, fraudulent schemes leveraging Deepfake technology are widespread, making it urgent to safeguard users' data privacy and security. However, many Deepfake detection methods based on Convolutional Neural Networks (CNNs) struggle to achieve satisfactory performance on mainstream datasets, especially with heavily compressed images. Observing that tampered images leave traces in the frequency domain, which are imperceptible to the naked eye but detectable through spectrum analysis, this study proposes a novel face forgery detection framework integrating spatial and frequency domain features. The framework introduces three innovative modules: the cross‐attention fusion module (CAFM), the guided attention module (GAM), and the multi‐scale feature fusion module (MSFFM), Specifically, CAFM combines spatial and frequency‐domain features through cross‐attention to enhance feature interaction. GAM generates attention maps to refine the integration of spatial and frequency features, while MSFFM fuses multi‐scale hierarchical features to capture both global and local tampering artifacts. These modules collectively improve the richness and discrimination of the extracted features, contributing to the overall detection performance. The proposed method demonstrates its effectiveness and superiority in forgery detection tasks, achieving a 3.9% average improvement in AUC compared to the state‐of‐the‐art method GocNet [1] on FaceForensics++ (FF++) and WildDeepfake datasets. Extensive experiments further validate the effectiveness of our approach.
Rongrong Gong, Dengyong Zhang, Arun Kumar Sangaiah, Mohammed J. F. Alenazi
IET Image Process.3
2025 DFS-Net: StyleGAN2-based dual feature separation for face De-Morphing
Ming Long, Fei Peng 0001, Dengyong Zhang
Multim. Tools Appl.5
2025 YOLO-DC: Integrating deformable convolution and contextual fusion for high-performance object detection
Dengyong Zhang, Chuanzhen Xu, Lei Wang 0143
Signal Process. Image Commun.1
2025 Generalizing Face Forgery Detection by Suppressed Texture Network With Two-Branch Convolution
abstract
With the development of Internet technology, deepfake (DF) videos can spread rapidly through online platforms, providing a new way of cyberbullying by generating nude pictures of female victims and using their faces to generate pornographic movies, which bring potential harm to individuals, society, and the country. Recently, there have been some really impressive results with DF detection models. These models have shown excellent outstanding performance when they are trained and tested using data from the same dataset. However, detecting DF remains difficult when the data comes from challenging datasets. To address this issue, this article aims to enhance the model's generalization by taking full advantage of the learning and representation capabilities of convolutional neural networks (CNNs) to adaptively suppress image texture information and catch deeper and more universal forgery features. Specifically, we introduce the texture suppression module (TSM) as a first step to suppress image content while simultaneously revealing the differences between authentic and tampered regions. Then, we carefully designed the cross stream interaction module (CSIM) and the cross stream mix block (CSMB) module to fully exploit the extracted forgery traces. Our proposed model has demonstrated superior generalization performance in extensive experiments.
Dengyong Zhang, Daijie Li, Arun Kumar Sangaiah, Feng Li 0065, Zelin Deng, Chengcheng Wu
IEEE Trans. Comput. Soc. Syst.1
2025 Efficient Hierarchical Feature Collaboration Transformer for Image Inpainting
abstract
Existing image inpainting methods face limitations in detail restoration. Although transformer-based models have made certain progress recently, the lack of hierarchical feature interaction and insufficient consideration of the importance of features at different network levels lead to semantic ambiguity in image reconstruction. To enhance the visual quality and accuracy of image inpainting, we adopt a multi-level feature fusion approach and propose a novel, efficient hierarchical feature collaboration transformer (HFCT). Our approach comprises two modules: dual stream gated feature fusion (DSGF) and region-separated attention module (RSAM), effectively capturing features at different levels of the network and enhancing inter-level information exchange. The DSGF module uses soft gating to fuse primary and advanced features, strengthening the connection from local to global consistency and reducing artifacts. The RSAM module resolves attention isolation issues in feature fusion through region-separated attention, strengthening the understanding of feature relationships, capturing more image semantics, and improving restoration accuracy. Extensive experiments on the Paris StreetView, CelebA-HQ, and Places2 benchmark datasets demonstrate that our proposed method achieves superior image inpainting quality compared to several state-of-the-art inpainting algorithms.
Dengyong Zhang, Nuo Fu, Xin Liao 0001, Hengfu Yang, Gaobo Yang
IEEE Trans. Multim.1
2025 Video Frame Interpolation via Fast Bidirectional 3D Correlation Volume
abstract
Recently, there has been a growing demand for flow-based video frame interpolation methods, which introduce correlation volumes to supervise the correlation of bidirectional optical flows. However, they often overlook the symmetry of the bidirectional motion field by consuming substantial computational cost, which is reflected in the fact that these methods often require a long runtime. To address these issues, in this article, we propose a bidirectional 3D correlation volume which is suitable for video frame interpolation. By decomposing the 4D correlation volume into two 3D correlation volumes in the horizontal and vertical directions, we significantly enhance the model’s inference speed with a minor sacrifice compared to our baseline. Additionally, when handling 2K video frames, our method achieves several-fold improvement in inference speed compared to other methods which implied correlation volume. The code is available at https://github.com/famt0531 .
Dengyong Zhang, Runqi Lou, Xiangling Ding, Xin Liao 0001, Gaobo Yang
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Spatiotemporal Inconsistency Learning and Interactive Fusion for Deepfake Video Detection
abstract
With the rise of the metaverse, the rapid advancement of Deepfakes technology has become closely intertwined. Within the metaverse, individuals exist in digital form and engage in interactions, transactions, and communications through virtual avatars. However, the development of Deepfakes technology has led to the proliferation of forged information disseminated under the guise of users’ virtual identities, posing significant security risks to the metaverse. Hence, there is an urgent need to research and develop more robust methods for detecting deep forgeries to address these challenges. This article explores deepfake video detection by leveraging the spatiotemporal inconsistencies generated by deepfake generation techniques, thereby proposing the interactive spatiotemporal inconsistency learning and interactive fusion (ST-ILIF) detection method, which consists of phase-aware and sequence streams. The spatial inconsistencies exhibited in frames of deepfake videos are primarily attributed to variations in the structural information contained within the phase component of the Fourier domain. To mitigate the issue of overfitting the content information, a phase-aware stream is introduced to learn the spatial inconsistencies from the phase-based reconstructed frames. Additionally, considering that deepfake videos are generated frame by frame and lack temporal consistency between frames, a sequence stream is proposed to extract temporal inconsistency features from the spatiotemporal difference information between consecutive frames. Finally, through feature interaction and fusion of the two streams, the representation ability of intermediate and classification features is further enhanced. The proposed method, which was evaluated on four mainstream datasets, outperformed most existing methods, and extensive experimental results demonstrated its effectiveness in identifying deepfake videos. Our source code is available at https://github.com/qff98/Deepfake-Video-Detection .
Dengyong Zhang, Xin Liao 0001, Feifan Qi, Gaobo Yang, Xiangling Ding
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Multi-scale noise-guided progressive network for image splicing detection and localization
Dengyong Zhang, Ningjing Jiang, Feng Li 0065, Xin Liao 0001, Gaobo Yang, Xiangling Ding
Expert Syst. Appl.1
2024 ADFF: Adaptive de-morphing factor framework for restoring accomplice's facial image
abstract
Abstract Morphing attacks (MAs) pose a substantial security threat to the Automatic Border Control (ABC) system. While a few morphing attack detection (MAD) methods have been proposed, the face morphing accomplice's facial restoration has not received sufficient attention. Due to the inability to foresee the morphing factor used for a particular morphed image, selecting the appropriate de‐morphing factor becomes a challenging problem in the restoration of the accomplice's facial image. If the morphing factor cannot be chosen reasonably, achieving the desired restoration effect is difficult. Therefore, this paper presents an adaptive de‐morphing factor framework (ADFF) architecture for restoring the accomplice's facial image. By exploiting the morphed images stored in the electronic passport system and the real‐time captured criminal's images, ADFF can effectively restore the accomplice's facial image. Experimental results and analysis show that ADFF can significantly reduce the security threats of MAs on ABC.
Min Long 0003, Fei Peng 0001, Dengyong Zhang
IET Image Process.5
2024 A convolutional neural network based on noise residual for seam carving detection
Dengyong Zhang, Zhenyu Lv, Feng Li 0065, Xiangling Ding, Gaobo Yang
J. Vis. Commun. Image Represent.1
2024 ESRL: efficient similarity representation learning for deepfake detection
Dengyong Zhang, Zhiqing Guo, Dewang Wang, Gaobo Yang
Multim. Tools Appl.2
2024 ERaL: Exceptional Regions-Aware Deep Video Interpolation Localization
abstract
Deep learning-based video frame interpolation (DVFI) can generate high frame-rate video sequences with high temporal consistency, usually producing visually plausible results. As DVFI can also be deployed for vicious video operations, it has misled users' visits and invalidates near-duplicate video detection. Therefore, it is urgent to locate the interpolated frames subjected to DVFI techniques. This letter investigates this issue by exploiting exceptional regions-aware localization (ERaL). In particular, we guide ERaL with an “inverted Z-shaped” network, which can better capture the position and intensity of exceptional regions regardless of the specific DVFI method, coming from the fact that the faked frame rate videos collapse even if any DVFI methods generate them, as ERaL only learns over original videos. Then, a hierarchical feature extraction is developed, integrating the feature enhancement, simplified transformer, and inverted residual feed-forward network, to produce a frame-wise localization of the interpolated frames for a given sequence. The proposed method is evaluated with counterfeited videos manipulated by three state-of-the-art DVFI approaches. Extensive experimental results demonstrate that the proposed method can effectively localize the interpolated frames, surpassing existing algorithms.
Xiangling Ding, Dengyong Zhang, Gaobo Yang
IEEE Signal Process. Lett.4
2024 Face Forgery Detection via Multi-Feature Fusion and Local Enhancement
abstract
With the rapid growth of Internet technology, security concerns have risen, particularly with the prevalence of Deepfakes, a popular visual forgery technique. Therefore, there is necessary to research more powerful methods to detect Deepfakes. However, many Convolutional Neural Networks-based detection methods struggle with cross-database performance, often overfitting to specific color textures. We observe that image noises can weaken the influence of color textures and expose the forgery traces in the noise domain. This is because tampering techniques, when altering face images, disrupt the consistency of feature distribution in the noise space. And the forgery traces in the noise space are complementary to the tampering artifacts present in the image space information. Therefore, we propose a novel face forgery detection network that combines spatial domain and noise domain. Our Dual Feature Fusion Module and Local Enhancement Attention Module contribute to more comprehensive feature representations, enhancing our method’s discriminative ability. Experimental results demonstrate superior performance compared to existing methods on mainstream datasets. https://github.com/jhchen1998/DeepfakeDetection.
Dengyong Zhang, Xin Liao 0001, Feng Li 0065, Gaobo Yang
IEEE Trans. Circuits Syst. Video Technol.1
2023 Deep Reinforcement Learning Based Load Balancing for Heterogeneous Traffic in Datacenter Networks
Jinbin Hu 0001, Wangqing Luo, Yi He 0017, Jin Wang 0001, Dengyong Zhang
ICA3PP (3)5
2023 Enabling Traffic-Differentiated Load Balancing for Datacenter Networks
Jinbin Hu 0001, Ying Liu 0064, Shuying Rao, Jing Wang 0209, Dengyong Zhang
ICA3PP (3)5
2023 Image Inpainting Forensics Algorithm Based on Dual-Domain Encoder-Decoder Network
Dengyong Zhang, En Tan, Feng Li 0065, Jing Wang 0209, Jinbin Hu 0001
ICA3PP (5)1
2023 Video Frame Interpolation via Multi-scale Expandable Deformable Convolution
abstract
Video frame interpolation is a challenging task in the video processing field. Benefiting from the development of deep learning, many video frame interpolation methods have been proposed, which focus on sampling pixels with useful information to synthesize each output pixel using their own sampling operation. However, these works have data redundancy limitations and fail to sample the correct pixel of complex motions. To solve these problems, we propose a new warping framework to sample called multi-scale expandable deformable convolution(MSEConv) which employs a deep fully convolutional neural network to estimate multiple small-scale kernel weights with different expansion degrees and adaptive weight allocation for each pixel synthesis. MSEConv covers most prevailing research methods as special cases of it, thus MSEConv is also possible to be transferred to existing works for performance improvement. To further improve the robustness of the whole network to occlusion, we also introduce a data preprocessing method for mask occlusion in video frame interpolation. Quantitative and qualitative experiments show that our method shows a robust performance comparable to or even superior to the state-of-the-art method. Our source code and visual comparable results are available at https://github.com/Pumpkin123709/MSEConv.
Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Gaobo Yang
IH&MMSec1
2023 LincRNA ZNF529-AS1 inhibits hepatocellular carcinoma via FBXO31 and predicts the prognosis of hepatocellular carcinoma patients
abstract
BACKGROUND: Invasion and metastasis of hepatocellular carcinoma (HCC) is still an important reason for poor prognosis. LincRNA ZNF529-AS1 is a recently identified tumour-associated molecule that is differentially expressed in a variety of tumours, but its role in HCC is still unclear. This study investigated the expression and function of ZNF529-AS1 in HCC and explored the prognostic significance of ZNF529-AS1 in HCC. METHODS: Based on HCC information in TCGA and other databases, the relationship between the expression of ZNF529-AS1 and clinicopathological characteristics of HCC was analysed by the Wilcoxon signed-rank test and logistic regression. The relationship between ZNF529-AS1 and HCC prognosis was evaluated by Kaplan‒Meier and Cox regression analyses. The cellular function and signalling pathways involved in ZNF529-AS1 were analysed by GO and KEGG enrichment analysis. The relationship between ZNF529-AS1 and immunological signatures in the HCC tumour microenvironment was analysed by the ssGSEA algorithm and CIBERSORT algorithm. HCC cell invasion and migration were investigated by the Transwell assay. Gene and protein expression were detected by PCR and western blot analysis, respectively. RESULTS: ZNF529-AS1 was differentially expressed in various types of tumours and was highly expressed in HCC. The expression of ZNF529-AS1 was closely correlated with the age, sex, T stage, M stage and pathological grade of HCC patients. Univariate and multivariate analyses showed that ZNF529-AS1 was significantly associated with poor prognosis of HCC patients and could be an independent prognostic indicator of HCC. Immunological analysis showed that the expression of ZNF529-AS1 was correlated with the abundance and immune function of various immune cells. Knockdown of ZNF529-AS1 in HCC cells inhibited cell invasion and migration and inhibited the expression of FBXO31. CONCLUSION: ZNF529-AS1 could be a new prognostic marker for HCC. FBXO31 may be the downstream target of ZNF529-AS1 in HCC.
Wan-liang Sun, Shuo Shuo Ma, Guanru Zhao, Dengyong Zhang
BMC Bioinform.7
2023 A data augmentation framework by mining structured features for fake face image detection
Zhiqing Guo, Gaobo Yang, Dewang Wang, Dengyong Zhang
Comput. Vis. Image Underst.4
2023 Rethinking gradient operator for exposing AI-enabled face forgeries
Zhiqing Guo, Gaobo Yang, Dengyong Zhang
Expert Syst. Appl.3
2023 SRTNet: a spatial and residual based two-stream neural network for deepfakes detection
Dengyong Zhang, Xiangling Ding, Gaobo Yang, Feng Li 0065, Zelin Deng, Yun Song
Multim. Tools Appl.1
2023 From depth-aware haze generation to real-world haze removal
Jiyou Chen, Gaobo Yang, Dengyong Zhang
Neural Comput. Appl.4
2023 L2BEC2: Local Lightweight Bidirectional Encoding and Channel Attention Cascade for Video Frame Interpolation
abstract
Video frame interpolation (VFI) is of great importance for many video applications, yet it is still challenging even in the era of deep learning. Some existing VFI models directly exploit existing lightweight network frameworks, thus making synthesized in-between frames blurry and creating artifacts due to imprecise motion representation. The other existing VFI models typically depend on heavy model architectures with a large number of parameters, preventing them from being deployed on small terminals. To address these issues, we propose a local lightweight VFI network ( L 2 BEC 2 ) that leverages bidirectional encoding structure with channel attention cascade. Specifically, we improve visual quality by introducing a forward and backward encoding structure with channel attention cascade to better characterize motion information. Furthermore, we introduce a local lightweight strategy into the state-of-the-art Adaptive Collaboration of Flows (AdaCoF) model to simplify its model parameters. Compared with the original AdaCoF model, the proposed L 2 BEC 2 obtains performance gain at the cost of only one-third of the number of parameters and performs favorably against the state-of-the-art works on public datasets. Our source code is available at https://github.com/Pumpkin123709/LBEC.git .
Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Yun Song, Gaobo Yang
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Video Frame Interpolation via Local Lightweight Bidirectional Encoding with Channel Attention Cascade
abstract
Deep Neural Networks based video frame interpolation, synthesizing in-between frames given two consecutive neighboring frames, typically depends on heavy model architectures, preventing them from being deployed on small terminals. When directly adopting the lightweight network architecture from these models, the synthesized frames may suffer from poor visual appearance. In this paper, a lightweight-driven video frame interpolation network (L2BEC2) is proposed. Concretely, we first improve the visual appearance by introducing the bidirectional encoding structure with channel attention cascade to better characterize the motion information; then we further adopt the local network lightweight idea into the aforementioned structure to significantly eliminate its redundant parts of the model parameters. As a result, our L2BEC2performs favorably at the cost of only one third of the parameters compared with the state-of-the-art methods on public datasets. Our source code is available at https://github.com/Pumpkin123709/LBEC.git.
Xiangling Ding, Pu Huang 0002, Dengyong Zhang, Xianfeng Zhao
ICASSP3
2022 DeepFake Videos Detection via Spatiotemporal Inconsistency Learning and Interactive Fusion
abstract
While the rapid expansion of DeepFake generation techniques has arisen a serious impact on human society, the detection of DeepFake videos is challenging because of their highly plausible contents on each frame, which are not visually apparent. To address that, this paper proposes a two-stream method to capture the spatial-temporal inconsistency cues, and then interactively fuse them to detect DeepFake videos. Since the traces of spatial inconsistency in DeepFake video frames mainly appear in their structural information, which reflects by the phase component in the frequency domain, the proposed frame-level stream learns the spatial inconsistency from the phase-based reconstructed frames to avoid fitting the content information. Aiming at the problem that the temporal inconsistency in DeepFake videos might be ignored, the temporality-level stream is proposed to extract the temporal correlation feature by the temporal difference networks and stacked ConvGRU module on consecutive multiple frames. When interacted with channel attention in the intermediate layer of two streams, and adaptively fused with the discriminative features of two streams from a global-local perspective, our proposed method performs better than the state-of-the-art detection methods.
Xiangling Ding, Dengyong Zhang
SECON3
2022 Forgery Detection Scheme of Deep Video Frame-rate Up-conversion Based on Dual-stream Multi-scale Spatial-temporal Representation
abstract
Video Frame-Rate Up-Conversion (FRUC) is originally designed to produce the high frame-rate video by periodically inserting new frames between two adjacent frames. However, it can also be utilized to synthesize the faked high frame-rate videos or spliced videos for malicious intents. The existing FRUC detection methods can efficiently identify its occurrence by exploring the blurring effects or deformed structures left over from the traditional FRUC methods. But the video FRUC has been substantially improved in the past years, especially in this deep learning era. Deep learning-based video FRUC (Deep FRUC) weakens the visual traces of the traditional ones such that it is challenging for the current FRUC detectors. In this paper, we propose a forensics algorithm based on dual-stream multi-scale spatial-temporal representation for Deep FRUC. Specifically, we develop a multi-scale receptive field strategy, and attention scheme to learn hidden tampered trace and spatial-temporal representation, respectively. Besides, frame residual and noise residual attention streams are complementary learning from the spatial-temporal dimension, in which the former can capture contrast differences and tampered abnormal boundaries, while the latter can explode noise inconsistencies between original and tampered frames. Experimental results show that the proposed algorithm can effectively achieve the best detection accuracy compared with the existing FRUC forensics works under the Deep FRUC video datasets.
Xiangling Ding, Dengyong Zhang
TrustCom3
2022 Pancreas Co-segmentation based on dynamic ROI extraction and VGGU-Net
Zhe Liu 0004, Yuqing Song 0001, Dengyong Zhang, Yan Zhu 0018, Deqi Yuan, Qingsong Gan, Victor S. Sheng
Expert Syst. Appl.6
2022 Spectral Reweighting and Spectral Similarity Weighting for Sparse Hyperspectral Unmixing
abstract
Sparse unmixing separates the pixel of hyperspectral images into a collection of pure spectral signatures and the associated fractional coefficients with a complete spectral library as a priori, avoiding the drawback of inaccurate extraction of endmember information from the original hyperspectral image. As a state-of-the-art sparse unmixing method, fast multiscale spatial regularization unmixing algorithm (MUA) consists of two procedures, concerning on the approximation image domain and the original domain, respectively. However, it ignores the inter-superpixel correlation of the original domain that each superpixel only involves a small number of spectral signatures, and ignores the spectral variability of the approximate image domain. We address these two issues by introducing two different weighting factors to enhance the unmixing result. The effectiveness of our proposed algorithm is demonstrated by the experimental results on both synthetic and real hyperspectral data. The code and datasets of this letter can be found at https://github.com/wangtaowei11/Unmixing-Algorithm.
Dengyong Zhang, Taowei Wang, Shujun Yang, Yuheng Jia, Feng Li 0065
IEEE Geosci. Remote. Sens. Lett.1
2021 A Saliency Detection and Gram Matrix Transform-Based Convolutional Neural Network for Image Emotion Classification
abstract
Using the convolutional neural network (CNN) method for image emotion recognition is a research hotspot of deep learning. Previous studies tend to use visual features obtained from a global perspective and ignore the role of local visual features in emotional arousal. Moreover, the CNN shallow feature maps contain image content information; such maps obtained from shallow layers directly to describe low-level visual features may lead to redundancy. In order to enhance image emotion recognition performance, an improved CNN is proposed in this work. Firstly, the saliency detection algorithm is used to locate the emotional region of the image, which is served as the supplementary information to conduct emotion recognition better. Secondly, the Gram matrix transform is performed on the CNN shallow feature maps to decrease the redundancy of image content information. Finally, a new loss function is designed by using hard labels and probability labels of image emotion category to reduce the influence of image emotion subjectivity. Extensive experiments have been conducted on benchmark datasets, including FI (Flickr and Instagram), IAPSsubset, ArtPhoto, and Abstract. The experimental results show that compared with the existing approaches, our method has a good application prospect.
Zelin Deng, Qiran Zhu, Pei He, Dengyong Zhang, Yuansheng Luo
Secur. Commun. Networks4
2020 Local and nonlocal constraints for compressed sensing video and multi-view image recovery
Yun Song, Dengyong Zhang, Qiang Tang 0006, Sheng Tang, Kun Yang 0001
Neurocomputing2
2020 An efficient tensor completion method via truncated nuclear norm
Yun Song, Jie Li 0002, Dengyong Zhang, Qiang Tang 0006, Kun Yang 0001
J. Vis. Commun. Image Represent.4
2020 Detecting seam carved images using uniform local binary patterns
Dengyong Zhang, Gaobo Yang, Feng Li 0065, Jin Wang 0001, Arun Kumar Sangaiah
Multim. Tools Appl.1
2020 Seam-Carved Image Tampering Detection Based on the Cooccurrence of Adjacent LBPs
abstract
Seam carving has been widely used in image resizing due to its superior performance in avoiding image distortion and deformation, which can maliciously be used on purpose, such as tampering contents of an image. As a result, seam-carving detection is becoming crucially important to recognize the image authenticity. However, existing methods do not perform well in the accuracy of seam-carving detection especially when the scaling ratio is low. In this paper, we propose an image forensic approach based on the cooccurrence of adjacent local binary patterns (LBPs), which employs LBP to better display texture information. Specifically, a total of 24 energy-based, seam-based, half-seam-based, and noise-based features in the LBP domain are applied to the seam-carving detection. Moreover, the cooccurrence features of adjacent LBPs are combined to highlight the local relationship between LBPs. Besides, SVM after training is adopted for feature classification to determine whether an image is seam-carved or not. Experimental results demonstrate the effectiveness in improving the detection accuracy with respect to different scaling ratios, especially under low scaling ratios.
Dengyong Zhang, Feng Li 0065, Arun Kumar Sangaiah, Xiangling Ding
Secur. Commun. Networks1
2020 An Efficient ECG Denoising Method Based on Empirical Mode Decomposition, Sample Entropy, and Improved Threshold Function
abstract
The electrocardiogram (ECG) signal can easily be affected by various types of noises while being recorded, which decreases the accuracy of subsequent diagnosis. Therefore, the efficient denoising of ECG signals has become an important research topic. In the paper, we proposed an efficient ECG denoising approach based on empirical mode decomposition (EMD), sample entropy, and improved threshold function. This method can better remove the noise of ECG signals and provide better diagnosis service for the computer-based automatic medical system. The proposed work includes three stages of analysis: (1) EMD is used to decompose the signal into finite intrinsic mode functions (IMFs), and according to the sample entropy of each order of IMF following EMD, the order of IMFs denoised is determined; (2) the new threshold function is adopted to denoise these IMFs after the order of IMFs denoised is determined; and (3) the signal is reconstructed and smoothed. The proposed method solves the shortcoming of discarding the first-order IMF directly in traditional EMD denoising and proposes a new threshold denoising function to improve the traditional soft and hard threshold functions. We further conduct simulation experiments of ECG signals from the MIT-BIH database, in which three types of noise are simulated: white Gaussian noise, electromyogram (EMG), and power line interference. The experimental results show that the proposed method is robust to a variety of noise types. Moreover, we analyze the effectiveness of the proposed method under different input SNR with reference to improving SNR ( SNR imp ) and mean square error ( MSE ), then compare the denoising algorithm proposed in this paper with previous ECG signal denoising techniques. The results demonstrate that the proposed method has a higher SNR imp and a lower MSE . Qualitative and quantitative studies demonstrate that the proposed algorithm is a good ECG signal denoising method.
Dengyong Zhang, Feng Li 0065, Shang Tian, Jin Wang 0001, Xiangling Ding, Rongrong Gong
Wirel. Commun. Mob. Comput.1
2018 A robust forgery detection algorithm for object removal by exemplar-based image inpainting
Dengyong Zhang, Zaoshan Liang, Gaobo Yang, Qingguo Li, Leida Li, Xingming Sun
Multim. Tools Appl.1
2017 Detection of image seam carving by using weber local descriptor and local binary patterns
Dengyong Zhang, Qingguo Li, Gaobo Yang, Leida Li, Xingming Sun
J. Inf. Secur. Appl.1
2017 Detecting image seam carving with low scaling ratio using multi-scale spatial and spectral entropies
Dengyong Zhang, Ting Yin, Gaobo Yang, Leida Li, Xingming Sun
J. Vis. Commun. Image Represent.1
2017 Residual domain dictionary learning for compressed sensing video recovery
Yun Song, Gaobo Yang, Hongtao Xie 0001, Dengyong Zhang, Xingming Sun
Multim. Tools Appl.4
2015 Detecting seam carving based image resizing using local binary patterns
Ting Yin, Gaobo Yang, Leida Li, Dengyong Zhang, Xingming Sun
Comput. Secur.4
2015 Detection of seam carving-based video retargeting using forensics hash
abstract
Abstract Seam carving is a content‐aware multimedia retargeting technique to adaptively resize multimedia data for different display sizes. However, it can also be used to remove objects from digital object or video for malicious purposes. In this paper, a forensics hash‐based tampering detection and localization approach is proposed for seam carving‐based video retargeting. It extracts the invariant Speeded‐up Robust Feature points from every spatiotemporal image to represent the matching surface, and the relative position change of the neighboring matching surface is used to build the forensic hash in a compact and scalable way. Experimental results show that the proposed forensics approach can effectively estimate the exact amount and rough locations of deleted seam carving surfaces. It achieves desirable detection performance even when there are frames deleted. If the hash length is reasonably increased, it can estimate the rough location and exact amount of deleted frames. Moreover, the built forensics hash is of good robustness, scalability, and compactness. Copyright © 2014 John Wiley & Sons, Ltd.
Wei Fei, Gaobo Yang, Leida Li, Dengyong Zhang
Secur. Commun. Networks5
2013 Attribute-based knowledge transfer learning for human pose estimation
Feng Li 0065, Shuren Zhou, Jianming Zhang 0003, Dengyong Zhang, Lingyun Xiang
Neurocomputing4