EDBT 2026 Demo / reviewers in the wild / expert
Gaobo Yang
dblp:57/5520
· DBLP profile ↗
84ranked-venue papers
1as first author
54since 2021 · last 2026
0000-0003-2734-659XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 32 since 2021Security and privacy · 17 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 16 · 14 since 2021Computer networks · 8 · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained AnnotationsabstractImage manipulation localization (IML) faces a fundamental trade-off between minimizing annotation cost and achieving fine-grained localization accuracy. Existing fully-supervised IML methods depend heavily on dense pixel-level mask annotations, which limits scalability to large datasets or real-world deployment. In contrast, the majority of existing weakly-supervised IML approaches are based on image-level labels, which greatly reduce annotation effort but typically lack precise spatial localization. To address this dilemma, we propose BoxPromptIML, a novel weakly-supervised IML framework that effectively balances annotation cost and localization performance. Specifically, we propose a coarse region annotation strategy, which can generate relatively accurate manipulation masks at lower cost. To improve model efficiency and facilitate deployment, we further design an efficient lightweight student model, which learns to perform fine-grained localization through knowledge distillation from a fixed teacher model based on the Segment Anything Model (SAM). Moreover, inspired by the human subconscious memory mechanism, our feature fusion module employs a dual-guidance strategy that actively contextualizes recalled prototypical patterns with real-time observational cues derived from the input. Instead of passive feature extraction, this strategy enables a dynamic process of knowledge recollection, where long-term memory is adapted to the specific context of the current image, significantly enhancing localization accuracy and robustness. Extensive experiments across both in-distribution and out-of-distribution datasets show that BoxPromptIML outperforms or rivals fully-supervised models, while maintaining strong generalization, low annotation cost, and efficient deployment characteristics. Zhiqing Guo, Dongdong Xi, Gaobo Yang |
AAAI | 4 |
| 2026 | Uncovering and Mitigating Destructive Multi-Embedding Attacks in Deepfake Proactive ForensicsabstractWith the rapid evolution of deepfake technologies and the wide dissemination of digital media, personal privacy is facing increasingly serious security threats. Deepfake proactive forensics, which involves embedding imperceptible watermarks to enable reliable source tracking, serves as a crucial defense against these threats. Although existing methods show strong forensic ability, they rely on an idealized assumption of single watermark embedding, which proves impractical in real-world scenarios. In this paper, we formally define and demonstrate the existence of Multi-Embedding Attacks (MEA) for the first time. When a previously protected image undergoes additional rounds of watermark embedding, the original forensic watermark can be destroyed or removed, rendering the entire proactive forensic mechanism ineffective. To address this vulnerability, we propose a general training paradigm named Adversarial Interference Simulation (AIS). Rather than modifying the network architecture, AIS explicitly simulates MEA scenarios during fine-tuning and introduces a resilience-driven loss function to enforce the learning of sparse and stable watermark representations. Our method enables the model to maintain the ability to extract the original watermark correctly even after a second embedding. Extensive experiments demonstrate that our plug-and-play AIS training paradigm significantly enhances the robustness of various existing methods against MEA. Lixin Jia, Zhiqing Guo, Yunfeng Diao, Dan Ma 0003, Gaobo Yang |
AAAI | 6 |
| 2026 | Beyond Fully Supervised Pixel Annotations: Scribble-Driven Weakly-Supervised Framework for Image Manipulation LocalizationabstractDeep learning-based image manipulation localization (IML) methods have achieved remarkable performance in recent years, but typically rely on large-scale pixel-level annotated datasets. To address the challenge of acquiring high-quality annotations, some recent weakly supervised methods utilize image-level labels to segment manipulated regions. However, the performance is still limited due to insufficient supervision signals. In this study, we explore a form of weak supervision that improves the annotation efficiency and detection performance, namely scribble annotation supervision. We re-annotated mainstream IML datasets with scribble labels and propose the first scribble-based IML (Sc-IML) dataset. Additionally, we propose the first scribble-based weakly supervised IML framework. Specifically, we employ self-supervised training with a structural consistency loss to encourage the model to produce consistent predictions under multi-scale and augmented inputs. In addition, we propose a prior-aware feature modulation module (PFMM) that adaptively integrates prior information from both manipulated and authentic regions for dynamic feature adjustment, further enhancing feature discriminability and prediction consistency in complex scenes. We also propose a gated adaptive fusion module (GAFM) that utilizes gating mechanisms to regulate information flow during feature fusion, guiding the model toward emphasizing potential tampered regions. Finally, we propose a confidence-aware entropy minimization loss. This loss dynamically regularizes predictions in weakly annotated or unlabeled regions based on model uncertainty, effectively suppressing unreliable predictions. Experimental results show that our method outperforms existing fully supervised approaches in terms of average performance both in-distribution and out-of-distribution. Guofeng Yu, Zhiqing Guo, Yunfeng Diao, Dan Ma 0003, Gaobo Yang |
AAAI | 6 |
| 2026 | Misalignment-tolerant perceptual similarity metric for full reference image dehazing quality assessment
Jiyou Chen, Gaobo Yang, Wenqi Ren |
Expert Syst. Appl. | 4 |
| 2026 | Face forgery detection via identification of evident tampered regions and multi-view analysisabstractIn AI-synthesized faces, there usually exist prominent natural features, which poses a huge challenge for face forgery detection. In this work, we propose a Region-Aware Deep Neural Network (RDNN). RDNN calculates the tampering possibility of each face region based on the features learned from each region and selects the region with the highest tampering possibility as the detection result. Then, a new Latent Cue Capture Loss (LCCL) is designed to train RDNN to capture those fake face features ignored by traditional loss functions. Besides, by leveraging RDNN to locate forgeries, we propose a deepfake detection strategy namely RDNN-based Multi-Perspective Deepfake Detection (RMDD), to keep the advantages of RDNN while improving the detection robustness. Specifically, RMDD uses RDNN to locate suspected forgeries in the original and horizontally flipped faces, and mines local features in the vicinity of these suspected forgeries. Finally, the detection result is acquired by integrating the above detection clues. Ablation experiments verify the ability of RDNN to locate manipulated traces and the contribution of each component in RMDD. Moreover, experimental results demonstrate that RMDD has excellent detection accuracy and generalization ability. Hanling Zhang, Gaobo Yang |
Neurocomputing | 3 |
| 2026 | Deepfake detection with dual-mode swin transformer: Multi-scale feature learning and local ambiguity mitigation
Gaobo Yang, Hanling Zhang |
J. Inf. Secur. Appl. | 2 |
| 2026 | Secure HEVC video steganography using IPMs spatial distribution and transfer probability
Ramadhani R. Iddy, Gaobo Yang, Dewang Wang, Xiangling Ding, Senzota K. Semakuwa |
Multim. Tools Appl. | 2 |
| 2026 | Follow your prompts: Controllable image dehazing via latent space manipulation
Jiyou Chen, Gaobo Yang, Wenqi Ren |
Pattern Recognit. | 4 |
| 2026 | Learned lossless medical image compression via dual transform and subimage-wise auto-regression
Ruixiao Guo, Gaobo Yang |
Signal Process. Image Commun. | 4 |
| 2026 | DTBF: Combining Local Statistical Artifacts and Concept Alignment for Synthetic Image Detection
ShaoWei Weng, Lifang Yu, Gaobo Yang, Pei-Wei Tsai |
IEEE Signal Process. Lett. | 4 |
| 2026 | Beyond Direct Embedding: Secure Separable Latent Space Watermarking for Anti-Screen-ShootingabstractExisting anti-screen-shooting watermarking methods embed watermarks on either the server or client side. Server-side embedding incurs high computational and communication overhead under concurrent requests, while client-side methods risk watermark interception during transmission and require additional encrypted channels. To address these limitations, we propose an end-to-end separable watermarking framework (SepWater) that exploits latent space representations. By decoupling server-side watermark embedding from client-side image generation, SepWater enhances security by preventing leaks of both the original content and the watermark, while also reducing transmission costs. On the server side, a dedicated encoder processes the watermark information, while a frozen pre-trained encoder handles the original image. We then fuse their outputs into a compact latent vector for transmission to the client. On the client side, a frozen VQGAN decoder reconstructs the watermarked image directly. In addition, we propose a local residual attention loss, combined with other image quality constraints, to produce watermarked images with high capture resistance and visual fidelity. Furthermore, to improve the robustness of the SepWater, two noise modes are simulated that contain eye protection noise and lightweight edge grayscale deviation noise. Experiments show that SepWater outperforms state-of-the-art methods in withstanding screen-shooting distortions, optimizing communication efficiency, and scaling under high concurrency, making it suitable for practical deployment. The source code is released at https://github.com/CVhnu/SepWater. Jiyou Chen, Xiyang Xie, Dewang Wang, Gaobo Yang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | WaveGuard: Robust Deepfake Detection and Source Tracing via Dual-Tree Complex Wavelet and Graph Neural NetworksabstractDeepfake technology has great potential in the field of media and entertainment, but it also brings serious risks, including privacy disclosure and identity fraud. To counter these threats, proactive forensic methods have become a research hotspot by embedding invisible watermark signals to build active protection schemes. However, existing methods are vulnerable to watermark destruction under malicious distortions, which leads to insufficient robustness. Moreover, embedding strong signals may degrade image quality, making it challenging to balance robustness and imperceptibility. Although watermarked images look natural, their underlying structures are often different from the original images, which is ignored by traditional watermarking methods. To address these issues, this paper proposes a proactive watermarking framework called WaveGuard, which explores frequency domain embedding and graph-based structural consistency optimization. In this framework, the watermark is embedded into the high-frequency sub-bands by dual-tree complex wavelet transform (DT-CWT) to enhance the robustness against distortions and deepfake forgeries. By leveraging joint sub-band correlations and selected sub-band combinations, the framework enables robust source tracing and semi-robust deepfake detection. To enhance imperceptibility, we propose a Structural Consistency Graph Neural Network (SC-GNN) that constructs graph representations of the original and watermarked images to ensure structural consistency and reduce perceptual artifacts. Experimental results show that the proposed method performs exceptionally well in face swap and face replay tasks. The code has been published at https://github.com/vpsg-research/WaveGuard. Ziyuan He, Zhiqing Guo, Gaobo Yang, Yunfeng Diao, Dan Ma 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prototype Memory-Based Neighboring Feature Fusion Network for Image Manipulation LocalizationabstractImage manipulation localization (IML) aims to segment manipulated regions in suspicious images. However, most existing methods rely solely on intrinsic features extracted from the input image and passively model local or global inconsistencies, making it difficult to accurately delineate manipulated regions with ambiguous boundaries. To address these challenges, we propose a prototype memory-based neighboring feature fusion network (PNF-Net), which is inspired by a biological memory mechanism. PNF-Net simulates selective preference by learning manipulation-trace prototypes as memory priors, thereby guiding representation learning toward consistent and discriminative manipulation cues. Specifically, we propose a memory-guided localization module (MLM) that models the consistencies and anomalies between manipulated regions and the background as memory priors, enabling precise localization. We then propose a neighboring feature interaction module (NFIM) that preserves fine-grained details from neighboring shallow features, enhances global semantics from neighboring deep features, and effectively fuses them. Finally, a verification fusion module (VFM) is designed to enrich contextual semantics and improve the completeness and accuracy of localization results. Extensive experiments on multiple benchmark datasets show that our PNF-Net outperforms most state-of-the-art IML models. Our code is available on https://github.com/vpsg-research/PNF-Net. Zhiqing Guo, Changtao Miao, Wenzhong Yang, Gaobo Yang, Xin Liao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | FALCON-Net: Feature Aggregation of Local Patterns for AI-Generated Image DetectionabstractWith the rapid development of generative models, the visual quality of generated images has become almost indistinguishable from real images, which poses a huge challenge to content authenticity verification. A key limitation of existing detectors is their reliance on model-specific cues, resulting in poor generalization to unseen models. Based on the observation of local differences in the generated images, we found that the generated images lack device-specific sensor noise and unnatural pixel intensity variations caused by the oversimplified generation process. These discrepancies provide important forensic cues for distinguishing between real and generated images. We propose the Feature Aggregation for Localized Context and Noise Network (FALCON-Net), which leverages these discrepancies to enhance detection capabilities. FALCON-Net integrates two complementary modules to enhance detection capabilities: the Intrinsic Noise Pattern Isolation (INP) module isolates device-specific noise patterns by analyzing high-frequency features in the frequency domain, while the Local Variation Pattern (LVP) module models the complex relationships between local pixels to capture directional intensity variations and reveal unnatural regularities in generated images. By combining these sensor-level and local structural cues, FALCON-Net identifies fundamental generative inconsistencies, ensuring robustness to post-processing and strong generalization to unseen models. Extensive experimental results show that FALCON-Net achieves the state-of-the-art performance in detecting generated images and shows good generalization ability to unseen generative models. The code is available at https://github.com/humiaomiaohaha/FALCON-Net. Dengyong Zhang, Changsheng Chen 0001, Jin Wang 0001, Yun Song, Gaobo Yang, Xin Liao 0001, Xiangling Ding |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2026 | Bi-Level Routing Attention and Enhanced Spatial-Temporal Inconsistency Learning for Deep VFI Video DetectionabstractWith the maturation of Deep Learning-based Video Frame Interpolation (Deep VFI), the left spatial-temporal inconsistency in the synthesis process is greatly improved, which poses a challenge to the current VFI detector. This article presents a dual-stream identification network based on Bi-level Routing Attention and enhanced Spatial-Temporal inconsistency learning (BRA-ST) to address this challenge. Specifically, the spatial inconsistencies in Deep VFI are mainly reflected in their motion regions and moving object edges; thus, the high-pass filter is introduced to enhance them, facilitating the three-stage pyramid structure of BiFormer Blocks with bi-level routing attention in the frame-level stream to learn. To fully exploit the temporal inconsistencies in the Deep VFI video, the time-difference module in the time-level stream is superimposed with the ConvGRU to extract the temporally dependent features of continuous multiple frames. Additionally, the middle layer of the two streams interacts and aggregates with the channel attention, and then, their last layer adaptively merges from a whole and part perspective for the ultimate frame prediction. Finally, the experimental findings on a constructed dataset by the five most advanced Deep VFI methods indicate that the proposed BRA-ST achieved \(F_{\text{1Score}}\) of 99.73%, which is superior to the existing Deep VFI detectors, and further verify that the resolution of BRA-ST for different Deep VFI methods reached 78.55%. Our source codes and dataset are available at https://pan.baidu.com/s/1f05_gS0qu5G-SSIkd9F4Hw?pwd=j6t6 . Xiangling Ding, Yunyi Li, Gaobo Yang, Yubo Lang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Progressive Reverse Attention Network for image inpainting detection and localization
Jiyou Chen, Xiangling Ding, Gaobo Yang |
Comput. Vis. Image Underst. | 4 |
| 2025 | SPNet: Seam carving detection via spatial-phase learning
Jiyou Chen, Zhi Lv, Ge Jiao, Gaobo Yang |
J. Inf. Secur. Appl. | 5 |
| 2025 | Enhancing image steganography security via universal adversarial perturbations
Dewang Wang, Gaobo Yang |
Multim. Tools Appl. | 4 |
| 2025 | Higher-order motion calibration and sparsity based outlier correction for video FRUC
Jiale He, Qunbing Xia, Gaobo Yang, Xiangling Ding |
Signal Process. Image Commun. | 3 |
| 2025 | You Only Need Clear Images: Self-Supervised Single Image DehazingabstractImage hazing refers to adding haze to a clear image, which is important for improving the data amount and diversity of synthetic hazy images that are required to train deep image dehazing models. However, existing image hazing works generate hazy images from a given clear image with a single transmission map. This violates the fact that hazy images are diverse for a natural scene at different times. The domain shift issue between synthetic and real-world hazy images constrains the robustness of deep dehazing models when dealing with real-world hazy images. In this work, we propose an unsupervised haze generation work to synthesize multiple hazy images with diverse haze distributions from a clear image, which requires only an atmospheric scattering model without extra labeling information. Instead of estimating a transmission map from a clear image, we propose to customize the transmission maps by redefining the transmission function. In such a controllable way, hazy images with diverse haze distributions are generated, which avoids the labor-intensive collection of paired data and alleviates the common domain-shift issue of deep image dehazing. Incorporating the unsupervised hazy images generator, we also construct a generalizable self-supervised image dehazing (SSID) framework, where deep image dehazing models can be trained without any human annotations. Extensive experiments on real-world hazy images show that the proposed approach is superior to state-of-the-art unsupervised dehazing works, and achieves competitive performance with the supervised works. Moreover, the proposed SSID framework can be easily generalized to the existing deep dehazing models, greatly improving dehazing robustness on real-world hazy images. Jiyou Chen, Wenqi Ren, Qunbing Xia, Gaobo Yang |
IEEE Trans. Multim. | 5 |
| 2025 | Efficient Hierarchical Feature Collaboration Transformer for Image InpaintingabstractExisting image inpainting methods face limitations in detail restoration. Although transformer-based models have made certain progress recently, the lack of hierarchical feature interaction and insufficient consideration of the importance of features at different network levels lead to semantic ambiguity in image reconstruction. To enhance the visual quality and accuracy of image inpainting, we adopt a multi-level feature fusion approach and propose a novel, efficient hierarchical feature collaboration transformer (HFCT). Our approach comprises two modules: dual stream gated feature fusion (DSGF) and region-separated attention module (RSAM), effectively capturing features at different levels of the network and enhancing inter-level information exchange. The DSGF module uses soft gating to fuse primary and advanced features, strengthening the connection from local to global consistency and reducing artifacts. The RSAM module resolves attention isolation issues in feature fusion through region-separated attention, strengthening the understanding of feature relationships, capturing more image semantics, and improving restoration accuracy. Extensive experiments on the Paris StreetView, CelebA-HQ, and Places2 benchmark datasets demonstrate that our proposed method achieves superior image inpainting quality compared to several state-of-the-art inpainting algorithms. Dengyong Zhang, Nuo Fu, Xin Liao 0001, Hengfu Yang, Gaobo Yang |
IEEE Trans. Multim. | 6 |
| 2025 | Video Frame Interpolation via Fast Bidirectional 3D Correlation VolumeabstractRecently, there has been a growing demand for flow-based video frame interpolation methods, which introduce correlation volumes to supervise the correlation of bidirectional optical flows. However, they often overlook the symmetry of the bidirectional motion field by consuming substantial computational cost, which is reflected in the fact that these methods often require a long runtime. To address these issues, in this article, we propose a bidirectional 3D correlation volume which is suitable for video frame interpolation. By decomposing the 4D correlation volume into two 3D correlation volumes in the horizontal and vertical directions, we significantly enhance the model’s inference speed with a minor sacrifice compared to our baseline. Additionally, when handling 2K video frames, our method achieves several-fold improvement in inference speed compared to other methods which implied correlation volume. The code is available at https://github.com/famt0531 . Dengyong Zhang, Runqi Lou, Xiangling Ding, Xin Liao 0001, Gaobo Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Spatiotemporal Inconsistency Learning and Interactive Fusion for Deepfake Video DetectionabstractWith the rise of the metaverse, the rapid advancement of Deepfakes technology has become closely intertwined. Within the metaverse, individuals exist in digital form and engage in interactions, transactions, and communications through virtual avatars. However, the development of Deepfakes technology has led to the proliferation of forged information disseminated under the guise of users’ virtual identities, posing significant security risks to the metaverse. Hence, there is an urgent need to research and develop more robust methods for detecting deep forgeries to address these challenges. This article explores deepfake video detection by leveraging the spatiotemporal inconsistencies generated by deepfake generation techniques, thereby proposing the interactive spatiotemporal inconsistency learning and interactive fusion (ST-ILIF) detection method, which consists of phase-aware and sequence streams. The spatial inconsistencies exhibited in frames of deepfake videos are primarily attributed to variations in the structural information contained within the phase component of the Fourier domain. To mitigate the issue of overfitting the content information, a phase-aware stream is introduced to learn the spatial inconsistencies from the phase-based reconstructed frames. Additionally, considering that deepfake videos are generated frame by frame and lack temporal consistency between frames, a sequence stream is proposed to extract temporal inconsistency features from the spatiotemporal difference information between consecutive frames. Finally, through feature interaction and fusion of the two streams, the representation ability of intermediate and classification features is further enhanced. The proposed method, which was evaluated on four mainstream datasets, outperformed most existing methods, and extensive experimental results demonstrated its effectiveness in identifying deepfake videos. Our source code is available at https://github.com/qff98/Deepfake-Video-Detection . Dengyong Zhang, Xin Liao 0001, Feifan Qi, Gaobo Yang, Xiangling Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Efficient object detector via dynamic prior and dynamic feature fusionabstractAbstract Sparse R-CNN is a new paradigm of object detection, which predicts objects in a sparse way. However, there are some limitations in Sparse R-CNN. One is the presence of weak prior information caused by fixed learnable proposal boxes and features across different images, necessitating excessive iterations for the model to refine its predictions; the other is the inadequate exploitation of multi-scale information, leading to the sub-optimal detection performance. Thus, building upon Sparse R-CNN, we propose an efficient detector that incorporates dynamic prior and dynamic feature fusion, called $D^{2}$-Det. In particular, for the dynamic prior part, a prior information generator module dynamically generates proposal features and boxes as the dynamic prior for different images to alleviate the inference-inefficient iterative refinement process of predictions, and we further propose the class scores decoupling method to reduce the computation overhead. Furthermore, for the dynamic feature fusion part, we develop a novel lightweight multi-scale feature fusion module, which dynamically aggregates features from all layers for each proposal box, enabling adaptive feature fusion and improving detection precision by nearly 2 AP. Experiments show that $D^{2}$-Det can achieve 46.6 AP on COCO 2017 with fewer computations for the backbone ResNet50, surpassing most of the state-of-the-art detectors. Zhili Zhou 0001, Gaobo Yang, Q. M. Jonathan Wu |
Comput. J. | 4 |
| 2024 | GAN-based adaptive cost learning for enhanced image steganography security
Dewang Wang, Gaobo Yang, Jiyou Chen, Xiangling Ding |
Expert Syst. Appl. | 2 |
| 2024 | Multi-scale noise-guided progressive network for image splicing detection and localization
Dengyong Zhang, Ningjing Jiang, Feng Li 0065, Xin Liao 0001, Gaobo Yang, Xiangling Ding |
Expert Syst. Appl. | 6 |
| 2024 | Improving image steganography security via ensemble steganalysis and adversarial perturbation minimization
Dewang Wang, Gaobo Yang, Zhiqing Guo, Jiyou Chen |
J. Inf. Secur. Appl. | 2 |
| 2024 | A convolutional neural network based on noise residual for seam carving detection
Dengyong Zhang, Zhenyu Lv, Feng Li 0065, Xiangling Ding, Gaobo Yang |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | ESRL: efficient similarity representation learning for deepfake detection
Dengyong Zhang, Zhiqing Guo, Dewang Wang, Gaobo Yang |
Multim. Tools Appl. | 5 |
| 2024 | A two-stage fake face image detection algorithm with expanded attention
Hanling Zhang, Gaobo Yang, Zhiqing Guo, Jiyou Chen |
Multim. Tools Appl. | 3 |
| 2024 | ERaL: Exceptional Regions-Aware Deep Video Interpolation LocalizationabstractDeep learning-based video frame interpolation (DVFI) can generate high frame-rate video sequences with high temporal consistency, usually producing visually plausible results. As DVFI can also be deployed for vicious video operations, it has misled users' visits and invalidates near-duplicate video detection. Therefore, it is urgent to locate the interpolated frames subjected to DVFI techniques. This letter investigates this issue by exploiting exceptional regions-aware localization (ERaL). In particular, we guide ERaL with an “inverted Z-shaped” network, which can better capture the position and intensity of exceptional regions regardless of the specific DVFI method, coming from the fact that the faked frame rate videos collapse even if any DVFI methods generate them, as ERaL only learns over original videos. Then, a hierarchical feature extraction is developed, integrating the feature enhancement, simplified transformer, and inverted residual feed-forward network, to produce a frame-wise localization of the interpolated frames for a given sequence. The proposed method is evaluated with counterfeited videos manipulated by three state-of-the-art DVFI approaches. Extensive experimental results demonstrate that the proposed method can effectively localize the interpolated frames, surpassing existing algorithms. Xiangling Ding, Dengyong Zhang, Gaobo Yang |
IEEE Signal Process. Lett. | 5 |
| 2024 | LDFnet: Lightweight Dynamic Fusion Network for Face Forgery Detection by Integrating Local Artifacts and Global Texture InformationabstractFace forgery detection has become a new research hotspot. Though existing detection works have achieved impressive performance, they are difficult to achieve a proper trade-off between detection accuracy and model complexity. To solve this problem, we design some low-complexity modules and construct a lightweight dynamic fusion network (LDFnet) to achieve high accuracy and lightweight face forgery detection. Firstly, we regard significant local visual artifacts as a correct semantic feature needed for detection. A spatial group-wise enhance (SGE) module is introduced as a supervision to suppress possible noise and capture local artifacts. Secondly, we design a manipulation trace extraction block (TraceBlock), which can replace vanilla convolution to achieve global inference, thus capturing the texture information in the global scope. Based on TraceBlock, we construct a global texture representation (GTR) network to extract global manipulation features hierarchically. Finally, we design a dynamic fusion mechanism (DFM) to fully fuse local and global clues, and dynamically generate a more discriminating feature representation. Extensive experimental results show that the proposed LDFnet is significantly superior to the previous detection works on some popular face forgery datasets, such as FF++, DFDC, CelebDF and HFF. In particular, LDFnet only uses 963k model parameters and 801M FLOPs, which is far lower than the calculation cost of face forgery detection based on large model, and achieves better detection results. Zhiqing Guo, Wenzhong Yang, Gaobo Yang, Keqin Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Face Forgery Detection via Multi-Feature Fusion and Local EnhancementabstractWith the rapid growth of Internet technology, security concerns have risen, particularly with the prevalence of Deepfakes, a popular visual forgery technique. Therefore, there is necessary to research more powerful methods to detect Deepfakes. However, many Convolutional Neural Networks-based detection methods struggle with cross-database performance, often overfitting to specific color textures. We observe that image noises can weaken the influence of color textures and expose the forgery traces in the noise domain. This is because tampering techniques, when altering face images, disrupt the consistency of feature distribution in the noise space. And the forgery traces in the noise space are complementary to the tampering artifacts present in the image space information. Therefore, we propose a novel face forgery detection network that combines spatial domain and noise domain. Our Dual Feature Fusion Module and Local Enhancement Attention Module contribute to more comprehensive feature representations, enhancing our method’s discriminative ability. Experimental results demonstrate superior performance compared to existing methods on mainstream datasets. https://github.com/jhchen1998/DeepfakeDetection. Dengyong Zhang, Xin Liao 0001, Feng Li 0065, Gaobo Yang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Constructing New Backbone Networks via Space-Frequency Interactive Convolution for Deepfake DetectionabstractThe serious concerns over the negative impacts of Deepfakes have attracted wide attentions in the community of multimedia forensics. The existing detection works achieve deepfake detection by improving the traditional backbone networks to capture subtle manipulation traces. However, there is no attempt to construct new backbone networks with different structures for Deepfake detection by improving the internal feature representation of convolution. In this work, we propose a novel Space-Frequency Interactive Convolution (SFIConv) to efficiently model the manipulation clues left by Deepfake. To obtain high-frequency features from tampering traces, a Multichannel Constrained Separable Convolution (MCSConv) is designed as the component of the proposed SFIConv, which learns space-frequency features via three stages, namely generation, interaction and fusion. In addition, SFIConv can replace the vanilla convolution in any backbone networks without changing the network structure. Extensive experimental results show that seamlessly equipping SFIConv into the backbone network greatly improves the accuracy for Deepfake detection. In addition, the space-frequency interaction mechanism does benefit to capturing common artifact features, thus achieving better results in cross-dataset evaluation. Our code will be available athttps://github.com/EricGzq/SFIConv. Zhiqing Guo, Zhenhong Jia, Dewang Wang, Gaobo Yang, Nikola K. Kasabov |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Image Dehazing Assessment: A Real-World Dataset and a Haze Density-Aware CriteriaabstractFull-reference image dehazing quality assessment (FR-IDQA) evaluates the visual quality of a dehazed image by measuring its differences with a clear reference. The existing FR-IDQA methods are not convincing due to the lack of well-aligned datasets of hazy and clear image pairs and the limited hand-crafted features make it difficult to simulate the complicated perception by the human visual system (HVS). In this work, we build a real-world image dataset, namely RW-Haze, which comprises natural hazy images and their well-aligned clear references. Each clear image is paired with several hazy images with diverse haze levels from slight to heavy. Meanwhile, the existing FR-IDQA works evaluate the dehazed image quality in a global manner, without considering local haze distributions in the original hazy image. Actually, the perceived haze in a natural hazy image is not uniformly distributed, and the haze density varies with scene depth. Based on this priori observation, we design a haze density-aware convolutional neural network (CNN), namely DehIQA, for FR-IDQA. It adopts transfer learning to alleviate the issue of lacking sufficient labeled data. Specifically, we divide image dehazing assessment into two tasks. The source task is to classify unpaired clear and hazy images, which enforces the deep network to learn haze-related features. The target task is image quality assessment, which is achieved by transferring the trained model for the source task to the target task. Considering the fact that the perceived distortion in a dehazed image is also not uniform, we present a haze density-aware mechanism into DehIQA, which assigns different weights for different local regions in a dehazed image in terms of the dark channel of the original hazy image. Extensive experimental results show that DehIQA outperforms the state-of-the-art (SOTA) works on the benchmark dataset and achieves better consistency with human perceptions. Jiyou Chen, Gaobo Yang, Dewang Wang, Xin Liao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Enhancing Adversarial Embedding based Image Steganography via Clustering Modification DirectionsabstractImage steganography is a technique used to conceal secret information within cover images without being detected. However, the advent of convolutional neural networks (CNNs) has threatened the security of image steganography. Due to the inherent properties of adversarial examples, adding perturbations to stego images can mislead the CNN-based image steganalysis, but it also easily leads to some errors when extracting secret information. Recently, some adversarial embedding methods have been proposed for improving image steganography security. In this work, we aim at furthering enhance the security of adversarial embedding-based image steganography by exploiting the strong correlation between adjacent pixels. Specifically, we divide the cover image into four non-overlapping parts for four-stage information embedding. During the adversarial embedding process, we cluster the modification directions of adjacent pixels and select only those with relatively larger amplitudes of gradients and smaller embedding costs to update their original embedding costs. Experimental results demonstrate that our proposed method can effectively fool targeted steganalyzers and outperform state-of-the-art techniques under different scenarios. Dewang Wang, Gaobo Yang, Zhiqing Guo, Jiyou Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Video Frame Interpolation via Multi-scale Expandable Deformable ConvolutionabstractVideo frame interpolation is a challenging task in the video processing field. Benefiting from the development of deep learning, many video frame interpolation methods have been proposed, which focus on sampling pixels with useful information to synthesize each output pixel using their own sampling operation. However, these works have data redundancy limitations and fail to sample the correct pixel of complex motions. To solve these problems, we propose a new warping framework to sample called multi-scale expandable deformable convolution(MSEConv) which employs a deep fully convolutional neural network to estimate multiple small-scale kernel weights with different expansion degrees and adaptive weight allocation for each pixel synthesis. MSEConv covers most prevailing research methods as special cases of it, thus MSEConv is also possible to be transferred to existing works for performance improvement. To further improve the robustness of the whole network to occlusion, we also introduce a data preprocessing method for mask occlusion in video frame interpolation. Quantitative and qualitative experiments show that our method shows a robust performance comparable to or even superior to the state-of-the-art method. Our source code and visual comparable results are available at https://github.com/Pumpkin123709/MSEConv. Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Gaobo Yang |
IH&MMSec | 5 |
| 2023 | A data augmentation framework by mining structured features for fake face image detection
Zhiqing Guo, Gaobo Yang, Dewang Wang, Dengyong Zhang |
Comput. Vis. Image Underst. | 2 |
| 2023 | Rethinking gradient operator for exposing AI-enabled face forgeries
Zhiqing Guo, Gaobo Yang, Dengyong Zhang |
Expert Syst. Appl. | 2 |
| 2023 | Robust detection of seam carving with low ratio via pixel adjacency subtraction and CNN-based transfer learning
Jiyou Chen, Gaobo Yang |
J. Inf. Secur. Appl. | 3 |
| 2023 | Flexible revocation in ciphertext-policy attribute-based encryption with verifiable ciphertext delegation
Gaobo Yang, Wen Dong 0011 |
Multim. Tools Appl. | 2 |
| 2023 | SRTNet: a spatial and residual based two-stream neural network for deepfakes detection
Dengyong Zhang, Xiangling Ding, Gaobo Yang, Feng Li 0065, Zelin Deng, Yun Song |
Multim. Tools Appl. | 4 |
| 2023 | From depth-aware haze generation to real-world haze removal
Jiyou Chen, Gaobo Yang, Dengyong Zhang |
Neural Comput. Appl. | 2 |
| 2023 | Exposing Deepfake Face Forgeries With Guided ResidualsabstractFor Deepfake detection, residual-based features can preserve tampering traces and suppress irrelevant image content. However, inappropriate residual prediction brings side effects on detection accuracy. Meanwhile, residual-domain features are easily affected by some image operations such as lossy compression. Most existing works exploit either spatial-domain or residual-domain features, which are fed into the backbone network for feature learning. Actually, both types of features are mutually correlated. In this work, we propose an adaptive fusion based guided residuals network (AdapGRnet), which fuses spatial-domain and residual-domain features in a mutually reinforcing way, for Deepfake detection. Specifically, we present a fine-grained manipulation trace extractor (MTE), which is a key module of AdapGRnet. Compared with the prediction-based residuals, MTE can avoid the potential bias caused by inappropriate prediction. Moreover, an attention fusion mechanism (AFM) is designed to selectively emphasize feature channel maps and adaptively allocate the weights for two streams. Experimental results show that AdapGRnet achieves better detection accuracies than the state-of-the-art works on four public fake face datasets including HFF, FaceForensics++, DFDC and CelebDF. Especially, AdapGRnet achieves an accuracy up to 96.52% on the HFF-JP60 dataset, which improves about 5.50%. That is, AdapGRnet achieves better robustness than the existing works. Zhiqing Guo, Gaobo Yang, Jiyou Chen, Xingming Sun |
IEEE Trans. Multim. | 2 |
| 2023 | L2BEC2: Local Lightweight Bidirectional Encoding and Channel Attention Cascade for Video Frame InterpolationabstractVideo frame interpolation (VFI) is of great importance for many video applications, yet it is still challenging even in the era of deep learning. Some existing VFI models directly exploit existing lightweight network frameworks, thus making synthesized in-between frames blurry and creating artifacts due to imprecise motion representation. The other existing VFI models typically depend on heavy model architectures with a large number of parameters, preventing them from being deployed on small terminals. To address these issues, we propose a local lightweight VFI network ( L 2 BEC 2 ) that leverages bidirectional encoding structure with channel attention cascade. Specifically, we improve visual quality by introducing a forward and backward encoding structure with channel attention cascade to better characterize motion information. Furthermore, we introduce a local lightweight strategy into the state-of-the-art Adaptive Collaboration of Flows (AdaCoF) model to simplify its model parameters. Compared with the original AdaCoF model, the proposed L 2 BEC 2 obtains performance gain at the cost of only one-third of the number of parameters and performs favorably against the state-of-the-art works on public datasets. Our source code is available at https://github.com/Pumpkin123709/LBEC.git . Dengyong Zhang, Pu Huang 0002, Xiangling Ding, Feng Li 0065, Yun Song, Gaobo Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | RW-HAZE: A Real-World Benchmark Dataset to Evaluate Quantitatively Dehazing AlgorithmsabstractMost existing image dehazing approaches are able to achieve desirable results whose differences are too subtle for people to qualitatively judge. Therefore, it is important to adopt quantitative assessment on real-world hazy images. However, many dehazing works have not been quantitatively evaluated on real-world hazy images due to the lack of appropriate real-world datasets. In this work, we attempt to address the issue and present a well-aligned real-world benchmark dataset, namely RW-Haze, for image dehazing evaluation, which had been lacking for a long period of time. It contains 210 pairs of well-aligned haze-free images and hazy images with distinct haze densities, which were captured from six cities in China by fixed cameras. To the best of our knowledge, RW-Haze is the first real-world dataset that is made up of well-aligned image pairs of haze-free and hazy images with diverse haze levels. We select 13 state-of-the-art single image dehazing works for making comprehensive evaluations among them on RW-Haze dataset. Experimental results show that there still exist rich rooms for image dehazing research to improve its robustness on natural hazy images, especially dense haze scenes. Jiyou Chen, Gaobo Yang |
ICIP | 4 |
| 2022 | Robust detection of dehazed images via dual-stream CNNs with adaptive feature fusion
Jiyou Chen, Gaobo Yang, Xiangling Ding, Zhiqing Guo |
Comput. Vis. Image Underst. | 2 |
| 2022 | HDNet: A dual-stream network with progressive fusion for image hazing detection
Jiyou Chen, Gaobo Yang, Zhiqing Guo |
J. Inf. Secur. Appl. | 2 |
| 2022 | Detecting GAN-generated face images via hybrid texture and sensor noise based features
Gaobo Yang |
Multim. Tools Appl. | 3 |
| 2021 | Detection of Deep Video Frame Interpolation via Learning Dual-Stream Fusion CNN in the Compression DomainabstractDeep learning-based Video Frame Interpolation (Deep VFI) diminishes the visual traces of the conventional one such that it is challenging for the current VFI detectors. Therefore, it is necessary to identify the presence of deep interpolated frames (DIF) in a video. This paper proposed a hybrid neural network to localize the DIF by learning spatio-temporal representations from the residual and motion vector information in the compression domain. Firstly, the residual and motion vector of motion regions are maintained by an intra-prediction constraints. Then, inherent tampering traces are further highlighted through subtracting the estimate of the residual or motion vector by virtue of residual modulation or MV refinement network. Finally, an attention-based dual-stream network is designed to jointly learn discriminative representations from the enhancement traces. Deep VFI video datasets created by the state-of-the-art deep VFI methods, have been evaluated, and extensive experimental results clearly demonstrate that our approach can achieve state-of-the-art performance compared with conventional methods. Xiangling Ding, Yifeng Pan, Jiyou Chen, Gaobo Yang, Yimao Xiong |
ICME | 5 |
| 2021 | Localization of Deep Video Inpainting Based on Spatiotemporal Convolution and Refinement NetworkabstractDeep learning-based video inpainting can fill the missing or undesired regions with spatial-temporal consistent contents without obvious visually distortion. Although the original purpose of deep inpainting is to repair flawed videos, it can also be adopted for malicious purposes, e.g., removal of specific objects. Therefore, automatically locating the inpainted regions is a challenging task in video forensics. This paper proposes a new forensic refinement framework to localize the deep inpainted regions by considering the spatial-temporal viewpoint. Firstly, we design a spatiotemporal convolution to suppress redundancy for highlighting deep inpainting traces. Then, a detection module is constructed with four concatenated ResNet blocks, and two upsampling layers to achieve a rough location map. Finally, a modified U-net based refinement module is developed for the pixel-wise localization map. Deep inpaiting video datasets created by the state-of-the-art deep inpainting method, have been evaluated, and extensive experimental results clearly demonstrate the efficacy of the proposed approach. Xiangling Ding, Yifeng Pan, Kui Luo, Yanming Huang, Junlin Ouyang, Gaobo Yang |
ISCAS | 6 |
| 2021 | Fake face detection via adaptive manipulation traces extraction network
Zhiqing Guo, Gaobo Yang, Jiyou Chen, Xingming Sun |
Comput. Vis. Image Underst. | 2 |
| 2021 | Blind detection of glow-based facial forgery
Zhiqing Guo, Lipin Hu, Gaobo Yang |
Multim. Tools Appl. | 4 |
| 2021 | Dual-Tree Complex Wavelet Transform-Based Direction Correlation for Face Forgery DetectionabstractWith the rapid development of face synthesis techniques, things are going from bad to worse as high-quality fake face images are unnoticeable by human eyes, which has brought serious public confidence and security problems. Thus, effective detection of face image forgeries is in urgent need. We observe that some subtle artificial artifacts in spatial domain can be easily recognized in transformation domain, and most facial features have an inherent directional correlation, and generative models would ruffle this kind of distribution pattern. Inspired by this, we propose a two-stream dual-tree complex wavelet-based face forgery network (DCWNet) to expose face image forgeries. Specifically, dual-tree complex wavelet transform is exploited to obtain six directional features (±75°, ±45°, ±15°) of different frequency components from original images, and a direction correlation extraction (DCE) block is presented to capture the direction correlation. Then, the direction pattern-aware clues and the original image are taken as two complementary network inputs. We also explore how specific frequency components work in face forgery detection and propose a new multiscale channel attention mechanism for features fusion. The experimental results prove that the proposed DCWNet outperforms the state-of-the-art methods in open datasets such as FaceForensics++ and achieves high robustness against lossy image compression. Shichao Gao, Gaobo Yang |
Secur. Commun. Networks | 3 |
| 2020 | Identification of various image retargeting techniques using hybrid features
Gaobo Yang |
J. Inf. Secur. Appl. | 4 |
| 2020 | Audio style transfer using shallow convolutional networks and random filters
Jiyou Chen, Gaobo Yang, Manimaran Ramasamy |
Multim. Tools Appl. | 2 |
| 2020 | Detecting seam carved images using uniform local binary patterns
Dengyong Zhang, Gaobo Yang, Feng Li 0065, Jin Wang 0001, Arun Kumar Sangaiah |
Multim. Tools Appl. | 2 |
| 2020 | Self-learning residual model for fast intra CU size decision in 3D-HEVC
Yue Li 0016, Ningbo Zhu, Gaobo Yang, Yapei Zhu, Xiangling Ding |
Signal Process. Image Commun. | 3 |
| 2020 | Spatio-temporal Saliency-based Motion Vector Refinement for Frame Rate Up-conversionabstractA spatio-temporal saliency-based frame rate up-conversion (FRUC) approach is proposed, which achieves better quality of interpolated frames and invalidates existing texture variation-based FRUC detectors. A spatio-temporal saliency model is designed to select salient frames. After obtaining initial motion vector field by texture- and color-based bilateral motion estimation, two motion vector refining (MVR) schemes are adopted for high and low saliency frames to hierarchically refine the motion vectors, respectively. To produce high-quality interpolated frames, image enhancement are performed for salient frames after frame interpolation. Due to distinct MVR schemes, there are different degrees of texture information in interpolated frames. Some edge and texture information is supplemented into salient frames as post-processing, which can invalidate existing texture variation-based FRUC detectors. Experimental results show that the proposed approach outperforms state-of-the-art works in both objective and subjective qualities of interpolated frames, and achieves the purpose of FRUC anti-forensics. Jiale He, Gaobo Yang, Xiangling Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Detection of motion compensated frame interpolation via motion-aligned temporal difference
Xiangling Ding, Yue Li 0016, Jiale He, Gaobo Yang |
Multim. Tools Appl. | 5 |
| 2019 | Robust Localization of Interpolated Frames by Motion-Compensated Frame Interpolation Based on an Artifact Indicated Map and Tchebichef MomentsabstractMotion-compensated frame interpolation (MCFI), a frame-interpolation technique to increase the motion continuity of low frame-rate video, can be utilized by counterfeiters for faking high bitrate video or splicing videos with different frame rates. For existing MCFI detectors, their performances are degraded under real-world scenarios such as H.264/AVC compression, noise, or blur. To address this issue, a robust MCFI detector is proposed to locate interpolated frames. By analyzing the distribution of residual energies within interpolated frames, we observe that there exist strong correlations between artifact regions and high residual energies. Thus, an artifact indicated map is introduced to select candidate artifact regions. Then, Tchebichef moments (TMs) are exploited to characterize the blurring effects or deformed structures among these regions. Specifically, the mean value of absolute high-order TMs of selected regions is used to model these temporal inconsistencies. Finally, a sliding window is adopted to locate interpolated frames, which are further refined by three post-processing operations. Chrominance information is also integrated with luminance information for robust identification of interpolated frames. Extensive experimental results show that compared with the state-of-the-art MCFI detectors, the proposed approach is more robust for compressed videos under various real-world scenarios. Xiangling Ding, Ningbo Zhu, Leida Li, Yue Li 0016, Gaobo Yang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | A robust forgery detection algorithm for object removal by exemplar-based image inpainting
Dengyong Zhang, Zaoshan Liang, Gaobo Yang, Qingguo Li, Leida Li, Xingming Sun |
Multim. Tools Appl. | 3 |
| 2018 | Identification of Motion-Compensated Frame Rate Up-Conversion Based on Residual SignalsabstractMotion-compensated frame rate up-conversion (MC-FRUC) is originally presented to increase the motion continuity of low frame rate videos by periodically inserting new frames, which improves the viewing experience. However, MC-FRUC can also be exploited to fake high frame rate videos or splice two videos with different frame rates for malicious purposes. A blind forensics approach is proposed for the identification of various MC-FRUC techniques. A theoretical model is first built for residual signal, which is exploited as tampering trace for blind forensics. The identification of various MC-FRUC techniques is then converted into a problem of discriminating the differences of residual signals among them. A pre-classifier is designed to suppress the side effects of original frames and static interpolated frames in candidate videos. Then, spatial and temporal Markov statistics features are extracted from the residual signals inside the interpolated frames for MC-FRUC identification. Five open MC-FRUC softwares and six representative MC-FRUC techniques have been tested, and experimental results show that the proposed approach can effectively locate interpolated frames and further identify the adopted MC-FRUC technique for both uncompressed videos and compressed videos with high perceptual qualities. Xiangling Ding, Gaobo Yang, Ran Li 0003, Yue Li 0016, Xingming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Probability Model-Based Early Merge Mode Decision for Dependent Views Coding in 3D-HEVCabstractAs a 3D extension to the High Efficiency Video Coding (HEVC) standard, 3D-HEVC was developed to improve the coding efficiency of multiview videos. It inherits the prediction modes from HEVC, yet both Motion Estimation (ME) and Disparity Estimation (DE) are required for dependent views coding. This improves coding efficiency at the cost of huge computational costs. In this article, an early Merge mode decision approach is proposed for dependent texture views and dependent depth maps coding in 3D-HEVC based on priori and posterior probability models. First, the priori probability model is established by exploiting the hierarchical and interview correlations from those previously encoded blocks. Second, the posterior probability model is built by using the Coded Block Flag (CBF) of the current coding block. Finally, the joint priori and posterior probability model is adopted to early terminate the Merge mode decision for both dependent texture views and dependent depth maps coding. Experimental results show that the proposed approach saves 45.2% and 30.6% encoding time on average for dependent texture views and dependent depth maps coding while maintaining negligible loss of coding efficiency, respectively. Yue Li 0016, Gaobo Yang, Yapei Zhu, Xiangling Ding, Rongrong Gong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Design of new scan orders for perceptual encryption of H.264/AVC videosabstractIn this study, a perceptual encryption algorithm is proposed for H.264/AVC video to enhance the scrambling effect and encryption space. Six new scan orders are designed for H.264/AVC encoder by analysing the energy distribution of discrete cosine transform coefficients. They are proven to have similar performance as the conventional zigzag scan order and its symmetrical scan order. These six new scan orders are combined with two existing scan orders to design a scan‐order based perceptual encryption algorithm. Specifically, video encryption is achieved more specifically by randomly selecting one scan order from the eight scan orders with a security key, and the sign bit flipping of DC coefficients is also incorporated to further increase the encryption space. Experimental results show that the proposed approach has the advantages of both low bitrate increase and low computational cost. Furthermore, it is more flexible and has stronger security than the existing scan‐order based video encryption schemes. Xiangling Ding, Yingzhuo Deng, Gaobo Yang, Yun Song, Dajiang He, Xingming Sun |
IET Inf. Secur. | 3 |
| 2017 | Detection of image seam carving by using weber local descriptor and local binary patterns
Dengyong Zhang, Qingguo Li, Gaobo Yang, Leida Li, Xingming Sun |
J. Inf. Secur. Appl. | 3 |
| 2017 | Detecting image seam carving with low scaling ratio using multi-scale spatial and spectral entropies
Dengyong Zhang, Ting Yin, Gaobo Yang, Leida Li, Xingming Sun |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Lossless visible watermarking based on adaptive circular shift operation for BTC-compressed images
Nur Mohammad, Xingming Sun, Hengfu Yang, Jianping Yin, Gaobo Yang, Mingfang Jiang |
Multim. Tools Appl. | 5 |
| 2017 | Residual domain dictionary learning for compressed sensing video recovery
Yun Song, Gaobo Yang, Hongtao Xie 0001, Dengyong Zhang, Xingming Sun |
Multim. Tools Appl. | 2 |
| 2017 | Fast CU size decision and mode decision algorithm for intra prediction in HEVC
Yun Song, Ye Zeng, Xueyu Li, Biye Cai, Gaobo Yang |
Multim. Tools Appl. | 5 |
| 2017 | Detecting video frame rate up-conversion based on frame-level analysis of average texture variation
Min Xia 0002, Gaobo Yang, Leida Li, Ran Li 0003, Xingming Sun |
Multim. Tools Appl. | 2 |
| 2017 | Unimodal Stopping Model-Based Early SKIP Mode Decision for High-Efficiency Video CodingabstractHigh-efficiency video coding (HEVC) can greatly improve coding efficiency compared with the prior video coding standard H.264/AVC by adopting advanced hierarchical coding structures such as coding unit (CU), prediction unit (PU), and transform unit. For each CU, an exhaustive mode decision strategy is adopted to achieve the best rate distortion (RD) cost, which simultaneously results in enormous computational complexity. In this paper, an early SKIP mode decision algorithm is proposed for the HEVC encoder to speed up the process of mode decision. Each CU size is categorized into either rare used or frequent used by exploiting the correlation of CU depth, which is estimated from the temporally colocated CUs. For the rare-used CU size, the SKIP mode is directly selected as the optimal mode and the remaining mode decision process is early terminated. For the frequent-used CU size, a unimodal stopping model is designed for its early SKIP mode decision by exploiting both hierarchical mode structure and RD cost property. Experimental results show that the proposed early SKIP mode decision method achieves average 58.5% and 54.8% encoding time savings, while the Bjontegaard Delta bit rate only increases average 0.8% and 0.8% for various test sequences under the random access and the low delay B conditions, respectively. Yue Li 0016, Gaobo Yang, Yapei Zhu, Xiangling Ding, Xingming Sun |
IEEE Trans. Multim. | 2 |
| 2016 | Detecting video frame-rate up-conversion based on periodic properties of edge-intensity
Gaobo Yang, Xingming Sun, Leida Li |
J. Inf. Secur. Appl. | 2 |
| 2016 | No-Reference Image Blur Assessment Based on Discrete Orthogonal MomentsabstractBlur is a key determinant in the perception of image quality. Generally, blur causes spread of edges, which leads to shape changes in images. Discrete orthogonal moments have been widely studied as effective shape descriptors. Intuitively, blur can be represented using discrete moments since noticeable blur affects the magnitudes of moments of an image. With this consideration, this paper presents a blind image blur evaluation algorithm based on discrete Tchebichef moments. The gradient of a blurred image is first computed to account for the shape, which is more effective for blur representation. Then the gradient image is divided into equal-size blocks and the Tchebichef moments are calculated to characterize image shape. The energy of a block is computed as the sum of squared non-DC moment values. Finally, the proposed image blur score is defined as the variance-normalized moment energy, which is computed with the guidance of a visual saliency model to adapt to the characteristic of human visual system. The performance of the proposed method is evaluated on four public image quality databases. The experimental results demonstrate that our method can produce blur scores highly consistent with subjective evaluations. It also outperforms the state-of-the-art image blur metrics and several general-purpose no-reference quality metrics. Leida Li, Weisi Lin, Xuesong Wang 0001, Gaobo Yang, Khosro Bahrami, Alex Chichung Kot |
IEEE Trans. Cybern. | 4 |
| 2015 | Detecting seam carving based image resizing using local binary patterns
Ting Yin, Gaobo Yang, Leida Li, Dengyong Zhang, Xingming Sun |
Comput. Secur. | 2 |
| 2015 | Compressed sensing image reconstruction using intra prediction
Yun Song, Yanfei Shen, Gaobo Yang |
Neurocomputing | 4 |
| 2015 | An efficient forgery detection algorithm for object removal by exemplar-based image inpainting
Zaoshan Liang, Gaobo Yang, Xiangling Ding, Leida Li |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Robust visual tracking based on structured sparse representation model
Hanling Zhang, Gaobo Yang |
Multim. Tools Appl. | 3 |
| 2015 | Detection of seam carving-based video retargeting using forensics hashabstractAbstract Seam carving is a content‐aware multimedia retargeting technique to adaptively resize multimedia data for different display sizes. However, it can also be used to remove objects from digital object or video for malicious purposes. In this paper, a forensics hash‐based tampering detection and localization approach is proposed for seam carving‐based video retargeting. It extracts the invariant Speeded‐up Robust Feature points from every spatiotemporal image to represent the matching surface, and the relative position change of the neighboring matching surface is used to build the forensic hash in a compact and scalable way. Experimental results show that the proposed forensics approach can effectively estimate the exact amount and rough locations of deleted seam carving surfaces. It achieves desirable detection performance even when there are frames deleted. If the hash length is reasonably increased, it can estimate the rough location and exact amount of deleted frames. Moreover, the built forensics hash is of good robustness, scalability, and compactness. Copyright © 2014 John Wiley & Sons, Ltd. Wei Fei, Gaobo Yang, Leida Li, Dengyong Zhang |
Secur. Commun. Networks | 2 |
| 2014 | Complexity scalable intra-prediction mode decision algorithm for mobile video applicationsabstractThe full search scheme employed in H.264/AVC significantly improves the coding performance, but it also introduces a very high computational complexity which limits the applications in resource‐constrained mobile devices. In this study, the authors firstly present a discretisation total variation and orientation gradient‐based hierarchical intra‐prediction mode decision method for mobile video applications. By shrinking the candidate mode set in the rate–distortion optimisation (RDO) process, the proposed algorithm reduces the computational complexity and power consumption of the encoder. Furthermore, they extend the hierarchical algorithm to a complexity scalable version in which the coding complexity is measured on five levels by reserving various numbers of modes for RDO. Experimental results demonstrate that the proposed mode decision algorithm reduces the coding complexity significantly with negligible performance degradation and the proposed complexity scalable algorithm is effective and efficient for mobile video application. Yun Song, Jizhen Long, Kun Yang 0001, Gaobo Yang |
IET Commun. | 4 |
| 2014 | Referenceless Measure of Blocking Artifacts by Tchebichef Kernel AnalysisabstractThis letter presents a Referenceless quality Measure of Blocking artifacts (RMB) using Tchebichef moments. It is based on the observation that Tchebichef kernels with different orders have varying abilities to capture blockiness. In a block manner, high-odd-order moments are computed to score the blocking artifacts. The blockiness scores are further weighted to incorporate the characteristic of Human Visual System (HVS), which is achieved by classifying the blocks into smooth and textured. Experimental results and comparisons demonstrate the advantage of the proposed method. Leida Li, Hancheng Zhu, Gaobo Yang, Jiansheng Qian |
IEEE Signal Process. Lett. | 3 |
| 2012 | Detecting Removed Object from Video with Stationary Background
Leida Li, Gaobo Yang, Guozhang Hu |
IWDW | 4 |
| 2012 | A robust hashing algorithm based on SURF for video copy detection
Gaobo Yang |
Comput. Secur. | 1 |
| 2011 | A Cauchy distribution based video watermark detection for H.264/AVC in DCT domainabstractCompared with Generalized Gaussian distribution (GGD), Cauchy distribution is superior to describe the statistical distribution of the Intra-coded DCT coefficients in H.264/AVC For the bipolar additive watermark in H.264/AVC video stream, a Cauchy distribution based detection algorithm is proposed by ternary hypothesis testing. Experimental results show that the proposed approach can achieve more than 80% on average for the accuracy of watermark detection. Lina Chen, Gaobo Yang, Anthony Tung Shuen Ho |
ISCAS | 2 |