Sheng Li 0006

dblp:23/3439-6 · DBLP profile ↗
← Back
91ranked-venue papers
9as first author
76since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 70 · 4 first-author · 59 since 2021Artificial intelligence and machine learning · 16 · 16 since 2021Security and privacy · 10 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neural Representations for Animated GIFs
Gaozhi Liu, Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian
ICMR3
2026 Toward Detecting Hidden Functionalities in Deep Learning Models
abstract
Deep functionality hiding is an emerging technique that embeds confidential or sensitive functions within seemingly benign deep learning models (DLMs), which perform ordinary machine learning tasks. This enables such models to execute covert tasks while remaining undetected. Despite the rapid progress in deep functionality hiding, countermeasures remain unexplored. In this paper, we propose Distribution Offset Analysis (DOA), a novel method for detecting hidden functionalities in DLMs. Our key insight is that the weight distribution of a benign DLM typically follows a Gaussian distribution, whereas a container DLM with hidden functionalities exhibits notable statistical deviations from this Gaussian pattern. In our methodology, we first compute the distributional distance (i.e.,offsets) between the model's weights and an ideal Gaussian distribution. We then fuse these offsets with weight features into a unified representation, which is subsequently used to train a meta-classifier for hidden functionality detection. Through extensive experiments, we demonstrate the effectiveness of the proposed DOA method, which achieves an average detection rate of over 87% against existing state-of-the-art deep functionality hiding techniques.
Guobiao Li, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
IEEE Signal Process. Lett.2
2026 On Cover Independent Deep Neural Network Steganography
abstract
Steganography aims to hide secret data into a cover media by subtle perturbation. Recently, deep neural network (DNN) based steganography has attracted a lot of research interests. Most of the existing DNN-based steganographic schemes design the secret encoder and decoder in an asymmetrical manner, the performance of which is heavily dependent on the content of cover media. In addition, they are weak in preventing unauthorized data extraction from the stego-media (i.e., the media with hidden data). To address these two issues, we propose a novel DNN steganographic framework termed the Cover Independent DNN Steganography (CIDS). In our CIDS, we take advantage of the powerful generative models to obtain AI-generated cover media for both the sender and receiver according to a key. Then, we propose a pair of symmetrical secret encoder and decoder to conduct a bijective transformation between the secrets and perturbations. While the perturbation is combined with the cover media to produce the stego-media. Such a secret encoding/decoding strategy focuses only on the secrets and is independent of the cover media. We also propose to encrypt the perturbations to effectively prevent unauthorized data extraction. Comprehensive experiments demonstrate the advantage of our CIDS over the state-of-the-art approaches for image steganography.
Guobiao Li, Sheng Li 0006, Zicong Luo, Zhenxing Qian, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.2
2026 Adversarial Diffusion Model: Generating High-Quality and Undetectable Images From Scratch
abstract
Diffusion models have made tremendous progress in generating visually realistic images. However, these images are statistically different from the real images, which could be accurately classified by carefully designed detectors. To evade detection, researchers have proposed various adversarial example generation schemes for AI-generated images. Despite the progress, most of these schemes have the tendency to post-process the images and the distortion is inevitable. In this paper, we propose an Adversarial Diffusion Model (ADM), which is able to directly generate high quality and undetectable images from scratch on top of a pre-trained stable diffusion model. In the ADM, an adversarial denoising U-Net is proposed for searching an adversarial latent. This latent is helpful for generating a prompt consistent adversarial example which is able to deceive the detector. Then, we propose a latent compensation module to make the adversarial examples have a similar reconstruction error to that of the real images. We further propose an adversarial decoder to minimize the difference between the high-frequency components of the real and adversarial examples. Comprehensive experiments are carried out to demonstrate the advantages of our ADM in generating adversarial AI-generated images.
Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Embedding Robust Watermarking into Pattern to Protect the Copyright of Ceramic Artifacts
abstract
Ceramic artworks with elegant patterns present enormous collectible value and profits. To claim the copyright, the builder usually pastes their conspicuous stamp on the bottom or side of the ceramic artworks, which inevitably affects the external image of the artwork. In addition, the stamp is weak in resisting forgery attacks due to its visible nature. To address the above issues, we propose in this paper a novel framework for embedding invisible watermarking into patterns of the ceramic artworks. In the framework, a template-based watermarking embedding scheme is designed to map the watermark to an invisible template, which is added to the ceramic pattern to create its watermarked version. A distortion layer is further proposed to model the distortion of ceramic patterns in the ceramic manufacturing process, where a color-halftoning and an adaptive brightness adjustment strategy are developed to counter the print and firing operations that introduce the most significant distortions. Finally, a deep decoder is learned to extract the watermarking from the distorted pattern. Various experiments have been conducted to demonstrate the advantage of our proposed method for protecting the copyright of the ceramic artworks, which provides reliable watermark extraction accuracy without the need for a conspicuous stamp.
Yuliang Xue, Guobiao Li, Zhenxing Qian, Sheng Li 0006, Chunlei Bao
AAAI5
2025 Physical Marker: Revealing Invisible Hyperlinks Hidden in Printed Trademarks
abstract
Embedding links in brand logos is a promising technology, which allows consumers to access the online information of products by capturing physical logo images. Previous physical data hiding methods primarily embed data within cover media in a global manner, making them ineffective for processing brand logos in vector graphics format with a transparent background. To address this issue, we propose in this paper a novel physical deep hiding scheme for invisibly embedding links in printed trademarks. Specifically, the encoder embeds links only into the area of the brand logo under the constraints of a mask, which is generated from the transparency information of the logo image. A background variation distortion is introduced into the distortion layer that approximate practical logo print-camera environments, such that the decoder could be learnt to retrieve the link from the camera-captured logo with various backgrounds. A feature prompt subspace modulator is further proposed and employed in the encoder to enhance the invisibility of the encoded logo pattern and in the decoder to boost hyperlink extraction accuracy. Various experiments have been conducted to demonstrate the advantage of our proposed method for embedding links in printed brand logos, which provides reliable extraction accuracy under both simulated and real scenarios.
Yuliang Xue, Guobiao Li, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
AAAI5
2025 Watermarking One for All: A Robust Watermarking Scheme Against Partial Image Theft
abstract
The proliferation of digital images on the Internet has provided unprecedented convenience, but also poses significant risks of malicious theft and misuse. Digital watermarking has long been researched as an effective tool for copyright protection. However, it often falls short when addressing partial image theft, a common yet little-researched issue in practical applications. Most existing schemes typically require the entire image as input to extract watermarks. However, in practice, malicious users often steal only a portion of the image to create new content. The stolen portion can have arbitrary shape or content, being fused with a new background and may have undergone geometric transformations, making it challenging for current methods to extract correctly. To address the issues above, we propose WOFA (Watermarking One for All), a robust watermarking scheme against partial image theft. First of all, we define the entire process of partial image theft and construct a dataset accordingly. To gain robustness against partial image theft, we then design a comprehensive distortion layer that incorporates the process of partial image theft and several common distortions in channel. For easier network convergence, we employ a multi-level network structure on the basis of the commonly used embedder-distortion layer-extractor architecture and adopt a progressive training strategy. Abundant experiments demonstrate that our superior performance in the scenario of partial image theft, offering a more reliable solution for protecting digital images against unauthorized use in practical use.
Gaozhi Liu, Silu Cao, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006, Wanli Peng
CVPR5
2025 SyncGuard: Robust Audio Watermarking Capable of Countering Desynchronization Attacks
abstract
Audio watermarking has been widely applied in copyright protection and source tracing. However, due to the inherent characteristics of audio signals, watermark localization and resistance to desynchronization attacks remain significant challenges. In this paper, we propose a learning-based scheme named SyncGuard to address these challenges. Specifically, we design a frame-wise broadcast embedding strategy to embed the watermark in arbitrary-length audio, enhancing time-independence and eliminating the need for localization during watermark extraction. To further enhance robustness, we introduce a meticulously designed distortion layer. Additionally, we employ dilated residual blocks in conjunction with dilated gated blocks to effectively capture multi-resolution time-frequency features. Extensive experimental results show that SyncGuard efficiently handles variable-length audio segments, outperforms state-of-the-art methods in robustness against various attacks, and delivers superior auditory quality.
Zhenliang Gan, Xiaoxiao Hu, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ECAI3
2025 Diffusion Model Is a Good Steganalyzer: Magnifying Subtle Perturbations in Image Data
abstract
Digital steganography embeds secret messages into images via invisible modifications, posing challenges for steganalysis, which seeks to detect these alterations by analyzing subtle shifts in image distributions. Previous steganalysis efforts primarily focus on enhancing the steganographic signal while suppressing image semantic content, such as through high-pass filtering. However, these empirically designed methods often lack theoretical underpinnings, exhibiting reduced detection accuracy, particularly at low embedding capacities. To address these limitations, we propose an innovative steganalysis approach that transforms images into pure Gaussian noise representations, actively amplifying the subtle distribution shifts introduced by the steganographic processes in spatial images. This paper pioneers the application of diffusion models to magnify steganographic signals, proposing a new paradigm for further research. Specifically, we iteratively perform forward steps of the probability flow in diffusion models to diminish semantic information. By utilizing the natural spreading properties of the diffusion process, we have theoretically validated the efficacy of each forward step in amplifying differences in noise patterns between cover and stego samples. These magnified differences can be easily captured by a simple classifier—a two-layer MLP. Extensive experiments demonstrate the effectiveness of our method, highlighting detection accuracy gains of 10% to 20% under standard conditions and an average 6.7% increase at low embedding rates compared to existing schemes.
Xiaoxiao Hu, Jiaqi Jin, Shengjiu Dai, Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian
ECAI4
2025 Filtering Resistant Large Language Model Watermarking via Style Injection
abstract
The exorbitant cost of training Large Language Models (LLMs) makes it essential to protect the models from illegal copying and unauthorized usage. Recent attempts at LLM protection utilize black-box watermarking schemes, which embed distinctive input-output mapping (i.e., trigger set) directly into the models. However, most of them construct trigger inputs by injecting abnormal characters into normal text, which can easily be filtered out by unauthorized users, leading to a failure in watermark verification. In this paper, we propose a novel filtering-resistant LLM watermarking scheme, which takes advantage of imperceptible text styles to trigger the watermark. To achieve this, we adopt a trigger generation network to transform normal text into stylized sentences, which are assigned a specific watermarking label to build the trigger set. We then fine-tune the LLMs on both the trigger sets and clean samples for watermark embedding and performance stabilization. To boost watermark accuracy, we further propose a feature separation loss term to distinguish between normal and trigger inputs. Experimental results indicate the effectiveness of our proposed scheme for resisting the filtering attack.
Zhaojun Guo, Guobiao Li, Junqiang Huang, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006
ICASSP6
2025 MoFRR: Mixture of Diffusion Models for Face Retouching Restoration
Qichao Ying, Zhenxing Qian, Sheng Li 0006, Runqi Zhang, Xinpeng Zhang 0001
ICCV4
2025 Texture-Aware Neural Radiance Fields Watermarking for Resisting Feature-Modulation Surrogate Model Attacks
abstract
The exorbitant cost of training Neural Radiance Fields (NeRF) makes it essential to protect them from illegal copying and unauthorized usage. Recent attempts at NeRF protection utilize watermarking schemes, which subtly alter NeRFs such that their rendered images contain watermarks that can be extracted by a decoder. In this paper, we investigate the robustness of existing NeRF watermarking schemes against surrogate model attacks, which train a surrogate NeRF using the input-output pairs of a victim NeRF. By proposing a Feature-Modulation Surrogate Model Attack (FM-SMA), we successfully crack most of the existing NeRF watermarking schemes. As a remedy, we further propose Texture-Aware Neural Radiance Fields Watermarking (TA-NFW), which resists FM-SMA attacks by embedding watermarks within the textures of NeRF-rendered results. Various experiments have been conducted to demonstrate the vulnerability of existing NeRF watermarking methods and the robustness of TA-NFW against the proposed FM-SMA attack.
Yuliang Xue, Guobiao Li, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICME5
2025 Safe-BVAR: Text-to-Image Generative Watermarking for Bitwise Visual AutoRegressive Model
abstract
Bitwise Vision AutoRegressive (BVAR) Model, as a distinguished source of young blood, has been taking the lead in the track of text-to-image synthesis, which at the same time raises legal and ethnic concerns such as copyright and authenticity. However, existing methods mainly focus on watermarking within diffusion models, which rely on the distinctive attributes of diffusion steps and cannot be directly transferred to new circumstances. To this end, we propose Safe-BVAR, the first watermark framework to embed bit strings during image generation in BVAR. Our study discovers the local similarity of the inferenced latent feature and the element-wise robustness of image autoencoder. Therefore, combined with the residual-accumulative nature of BVAR, we propose a novel Late Stage Residual Implanter to embed watermark and extract the information based on Local Contextual Extractor. Furthermore, we propose a Distributed Rotational Arranger to enhance watermark against local distortions. Our method is training-free and plug-and-play. Meanwhile, it can be easily applied to flexible-sized images. We evaluate the robustness and invisibility of the watermark, showing that it can resist common image attacks and cast inappreciable influence on the image.
Shengjiu Dai, Xiujian Liang, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia3
2025 DynMark: A Robust Watermarking Solution for Dynamic Screen Content with Small-size Screenshot Support
Changyu Rao, Gaozhi Liu, Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian
ACM Multimedia3
2025 Learning Discrepant Transformations for Face Privacy Protection
abstract
Online face recognition systems usually store face features in the server database for authentication, which are vulnerable to face reconstruction attacks. Various face privacy protection approaches have been proposed to address this issue, where transformation-based schemes are shown to be promising. However, the existing transformation-based schemes are all hand-crafted approaches which are difficult to balance the privacy protection and face recognition. In this paper, we propose to learn a set of discrepant convolutional neural networks (DCNNs) to protect the privacy of face features. We randomly split the original face features into different sub-features. Each of the DCNNs transforms an original sub-feature into a protected one. We adopt appropriate strategies to make the DCNNs as diverse as possible to improve the ability of our protected features to resist different face reconstruction attacks, where a face recognition loss and a privacy protection loss are designed for training. The former ensures that the protected feature can be matched directly using the existing face recognizers, while the latter incorporates a shadow face reconstruction model to interrupt the correlation between the protected features and the face images. Experimental results demonstrate the advantage of our method over existing schemes for face privacy protection. Our protected features can be accurately matched using existing face recognizers, which are capable of resisting both black-box and white-box face reconstruction attacks.
Chenda Wei, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ACM Multimedia4
2025 Towards Generalized Physical Occlusion Detection On Documents
abstract
Fake document detection is an important area in image forensics. Most of the existing techniques focus on the detection of digitally forged documents. In this paper, we look into the forgery of generalized physical occlusion, which is a simple and effective strategy to generate fake document images. We propose an Adversary Decomposition Network (ADDNet) to effectively extract generalized physical occlusion features from various types of documents, where two adversarial classifiers are designed and trained for feature decomposition. On top of the ADDNet, we further propose a lightweight Document Adapter (DA) for flexible and scalable fake document detection, which works well when we encounter a new type of document with limited samples for fine-tuning. To facilitate the research, we newly construct a dataset for physical occlusion detection on different types of documents. Various experiments are carried out to demonstrate the advantage of our proposed scheme over the existing schemes for physical occlusion detection, especially when the document is unseen or has limited samples in training.
Yiang Zhu, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ACM Multimedia4
2025 A Key-Driven Framework for Identity-Preserving Face Anonymization
Guang Hua 0001, Sheng Li 0006, Guorui Feng
NDSS3
2025 Improved Generative Steganography Based on Diffusion Model
abstract
The rapid growth of generative models has led to a new direction in steganography called generative steganography (GS). It allows message-to-image generation without the need for a carrier image. Recently, generative steganography methods have been proposed using generative adversarial networks (GANs) and Flow models. On the one hand, methods that use GANs to generate stego images struggle to fully recover the hidden message because the networks are not reversible. On the other hand, methods based on Flow encounter a problem where the images they create might not look real, mainly because the network has limitations in being reversible. Diffusion models fulfill network reversibility while generating high-quality images. However, the framework of existing diffusion models is reversible, but hidden message recovery is not perfectly reversible, resulting in the recovered message being similar but not exactly the same as the hidden one. Existing diffusion models are typically trained for one-directional image generation tasks, so they face some problems when dealing with bi-directional steganography tasks. If pre-trained diffusion models are directly used to generate stego images, exact secret data extraction through the diffusion process cannot be achieved. In this paper, we present an improved generative steganography based on the diffusion model (GSD), which conceals secret data in the frequency domain of random noise to enhance the security and accuracy of steganography, and re-trains the denoising diffusion implicit model (DDIM) for steganography, called the StegoDiffusion. During training StegoDiffusion, random noise is injected into the clean natural images and then trained through the forward diffusion process to obtain the re-trained StegoDiffusion. Our proposed GSD scheme achieves a 100% extraction accuracy for hidden secret data with a payload of 1 bit-per-pixel (bpp) in a single channel, and generates high-quality stego images in PNG format.
Ping Wei 0004, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
IEEE Trans. Circuits Syst. Video Technol.5
2025 Conditional Flow-Based Generative Steganography
abstract
Generative steganography (GS) is a novel data-hiding technique that generates stego images directly from secret data without using cover images, which is different from traditional steganography. However, existing steganography methods have shortcomings in terms of hiding capacity, extraction accuracy, and diversity of stego images. To address these limitations, we propose a high-performance Conditional Flow-based Generative Steganography (CFGS). First, to achieve exact extraction of secret data in high-capacity scenarios, we hide secret data in the frequency domain to resist the impact of stego image distortion. In addition, to enhance the diversity of stego images, we introduce a novel conditional generative flow model (C-Flow) to generate stego images, which consists of two newly designed layers, the Conditional Attention-based Affine Coupling layer and the Conditional Invertible Norm layer. C-Flow can accurately guide the visual content of stego images through different conditions, enhancing the diversity of stego images. Our approach is the first GS method capable of conditional guidance of stego image visual content, and achieves extraction accuracy of hidden secret data equal to or close to 100% for payloads up to 1 bit-per-pixel (bpp). Extensive experiments demonstrate that our proposed approach outperforms state-of-the-art GS methods.
Ping Wei 0004, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006, Chuan Qin 0001
IEEE Trans. Dependable Secur. Comput.5
2024 Purified and Unified Steganographic Network
abstract
Steganography is the art of hiding secret data into the cover media for covert communication. In recent years, more and more deep neural network (DNN)-based steganographic schemes are proposed to train steganographic networks for secret embedding and recovery, which are shown to be promising. Compared with the handcrafted steganographic tools, steganographic networks tend to be large in size. It raises concerns on how to imperceptibly and effectively transmit these networks to the sender and receiver to facilitate the covert communication. To address this issue, we propose in this paper a Purified and Unified Steganographic Network (PUSNet). It performs an ordinary machine learning task in a purified network, which could be triggered into steganographic networks for secret embedding or recovery using different keys. We formulate the construction of the PUSNet into a sparse weight filling problem to flexibly switch between the purified and steganographic networks. We further instantiate our PUSNet as an image denoising network with two steganographic networks concealed for secret image embedding and recovery. Comprehensive experiments demonstrate that our PUSNet achieves good performance on secret image embedding, secret image recovery, and image denoising in a single architecture. It is also shown to be capable of imperceptibly carrying the steganographic networks in a purified network. Code is available at https://github.com/albblgb/PUSNet
Guobiao Li, Sheng Li 0006, Zicong Luo, Zhenxing Qian, Xinpeng Zhang 0001
CVPR2
2024 Engaging Live Video Comments Generation
abstract
Multimodal Dialogue agents are often required to respond to conversation history using both textual and visual content. Even though current dialogue studies predominantly strive to generate natural texts or images, they fall short in considering the relevance of multimodal responses within a dialogue context, consequently confining agents from making prudent choices based on multiple alternatives and their associated relevance scores for decision-making. In this paper, we present a bidirectional multimodal dialogue framework that skillfully combines the forward generation of multiple text and image response candidates with reverse selection guided by relevance scores evaluated on dialogue context, facilitating agents in selecting the most suitable multimodal responses. Specifically, the forward generation aspect of our framework leverages a stage-wise approach, first producing textual replies and composite visual descriptions from the dialogue context, followed by the generation of visual responses aligned with the descriptions. In the reverse selection process, visual responses are translated into tangible descriptive texts that, in conjunction with textual responses, are inversely tied back to the dialogue context for relevance assessment, assigning a reference score to each multimodal response candidate to assist the intelligent agent in making informed decisions. Experimental outcomes demonstrate that our proposed bidirectional dialogue response framework markedly elevates performance in both automatic and human evaluations, yielding a range of contextually fitting multimodal responses for selection.
Ge Luo 0003, Junqiang Huang, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia5
2024 Cover-separable Fixed Neural Network Steganography via Deep Generative Models
abstract
Image steganography is the process of hiding secret data in a cover image by subtle perturbation. Recent studies show that it is feasible to use a fixed neural network for data embedding and extraction. Such Fixed Neural Network Steganography (FNNS) demonstrates favorable performance without the need for training networks, making it more practical for real-world applications. However, the stego-images generated by the existing FNNS methods exhibit high distortion, which is prone to be detected by steganalysis tools. To deal with this issue, we propose a Cover-separable Fixed Neural Network Steganography, namely Cs-FNNS. In Cs-FNNS, we propose a Steganographic Perturbation Search (SPS) algorithm to directly encode the secret data into an imperceptible perturbation, which is combined with an AI-generated cover image for transmission. Through accessing the same deep generative models, the receiver could reproduce the cover image using a pre-agreed key, to separate the perturbation in the stego-image for data decoding. such an encoding/decoding strategy focuses on the secret data and eliminates the disturbance of the cover images, hence achieving a better performance. We apply our Cs-FNNS to the steganographic field that hiding secret images within cover images. Through comprehensive experiments, we demonstrate the superior performance of the proposed method in terms of visual quality and undetectability. Moreover, we show the flexibility of our Cs-FNNS in terms of hiding multiple secret images for different receivers. Code is available at https://github.com/albblgb/Cs-FNNS
Guobiao Li, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia2
2024 Are handcrafted filters helpful for attributing AI-generated images?
abstract
Recently, a vast number of image generation models have been proposed, which raises concerns regarding the misuse of these artificial intelligence (AI) techniques for generating fake images. To attribute the AI-generated images, existing schemes usually design and train deep neural networks (DNNs) to learn the model fingerprints, which usually requires a large amount of data for effective learning. In this paper, we aim to answer the following two questions for AI-generated image attribution, 1) is it possible to design useful handcrafted filters to facilitate the fingerprint learning? and 2) how we could reduce the amount of training data after we incorporate the handcrafted filters? We first propose a set of Multi-Directional High-Pass Filters (MHFs) which are capable to extract the subtle fingerprints from various directions. Then, we propose a Directional Enhanced Feature Learning network (DEFL) to take both the MHFs and randomly-initialized filters into consideration. The output of the DEFL is fused with the semantic features to produce a compact fingerprint. To make the compact fingerprint discriminative among different models, we propose a Dual-Margin Contrastive (DMC) loss to tune our DEFL. Finally, we propose a reference based fingerprint classification scheme for image attribution. Experimental results demonstrate that it is indeed helpful to use our MHFs for attributing the AI-generated images. The performance of our proposed method is significantly better than the state-of-the-art for both the closed-set and open-set image attribution, where only a small amount of images are required for training.
Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001, Athanasios V. Vasilakos
ACM Multimedia3
2024 Multi-view Feature Extraction via Tunable Prompts is Enough for Image Manipulation Localization
abstract
Deceptive images can quickly spread via social networking services, posing significant risks. The rapid progress in Image Manipulation Localization (IML) seeks to address this issue. However, the scarcity of public training datasets in the IML task directly hampers the performance of models. To address the challenge, we propose a Prompt-IML framework, which leverages the rich prior knowledge of pre-trained models by employing tunable prompts. Specifically, sets of tunable prompts enable the frozen pre-trained model to extract multi-view features, including spatial and high-frequency features. This approach minimizes redundant architecture for feature extraction across different views, resulting in reduced training costs. In addition, we develop a plug-and-play Feature Alignment and Fusion module that seamlessly integrates into the pre-trained models without additional structural modifications. The proposed module reduces noise and uncertainty in features through interactive processing. The experimental results showcase that our proposed method attains superior performance across 6 test datasets, demonstrating exceptional robustness.
Xuntao Liu, Yuzhou Yang, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
ACM Multimedia7
2024 From Covert Hiding To Visual Editing: Robust Generative Video Steganography
abstract
Traditional video steganography methods are based on modifying the covert space for embedding, whereas we propose an innovative approach that embeds secret message within semantic feature for steganography during the video editing process. Although existing traditional video steganography methods excel in balancing security and capacity, they lack adequate robustness against common distortions in online social networks (OSNs). In this paper, we propose an end-to-end robust generative video steganography network (RoGVSN), which achieves visual editing by modifying semantic feature of videos to embed secret message. We exemplify the face-swapping scenario as an illustration to demonstrate the visual editing effects. Specifically, we devise an adaptive scheme to seamlessly embed secret messages into the semantic features of videos through fusion blocks. Extensive experiments demonstrate the superiority of our method in terms of robustness, extraction accuracy, visual quality, and capacity.
Xueying Mao, Xiaoxiao Hu, Wanli Peng, Zhenliang Gan, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
ACM Multimedia7
2024 Emotion-Aware and Efficient Meme Sticker Dialogue Generation
abstract
Recent advances have emphasized the importance of meme stickers in open-domain dialogue systems.However, previous studies overlook the one-to-many issue that a single sticker could represent various emotions in different dialogue contexts.Additionally, they require retraining the model for new stickers which did not appear in previous training.To address the above issues, we propose in this paper an Emotion-Aware and Efficient Meme Sticker Dialogue generation framework.In the framework, we design an Emotion Adaptive Prompt to capture the emotional cues from the dialogue history, which is sent to an Emotion-Aware Fusion Decoder to guide the generation of text responses and to a meme sticker selector to choose the corresponding sticker.Furthermore, to improve the stickers' selection efficiency, we further incorporate the few-shot learning strategy into the proposed framework to avoid extensive model retraining for unseen meme stickers.Through extensive experiments, we demonstrate the superior performance of the proposed E 2 MSD compared to existing methods regarding the quality of response generation and the efficiency of meme sticker retrieval.
Zhaojun Guo, Junqiang Huang, Guobiao Li, Wanli Peng, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006
MMAsia7
2024 Disentangled Style Domain for Implicit z-Watermark Towards Copyright Protection
abstract
Text-to-image models have shown surprising performance in high-quality image generation, while also raising intensified concerns about the unauthorized usage of personal dataset in training and personalized fine-tuning. Recent approaches, embedding watermarks, introducing perturbations, and inserting backdoors into datasets, rely on adding minor information vulnerable to adversarial training, limiting their ability to detect unauthorized data usage. In this paper, we introduce a novel implicit Zero-Watermarking scheme that first utilizes the disentangled style domain to detect unauthorized dataset usage in text-to-image models. Specifically, our approach generates the watermark from the disentangled style domain, enabling self-generalization and mutual exclusivity within the style domain anchored by protected units. The domain achieves the maximum concealed offset of probability distribution through both the injection of identifier $z$ and dynamic contrastive learning, facilitating the structured delineation of dataset copyright boundaries for multiple sources of styles and contents. Additionally, we introduce the concept of watermark distribution to establish a verification mechanism for copyright ownership of hybrid or partial infringements, addressing deficiencies in the traditional mechanism of dataset copyright ownership for AI mimicry. Notably, our method achieves one-sample verification for copyright ownership in AI mimic generations. The code is available at: [https://github.com/Hlufies/ZWatermarking](https://github.com/Hlufies/ZWatermarking)
Junqiang Huang, Zhaojun Guo, Ge Luo 0003, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
NeurIPS5
2024 Stealthy Backdoor Attacks On Deep Point Cloud Recognization Networks
abstract
Abstract Deep neural networks are vulnerable to backdoor attacks. Previous backdoor attacks have mainly focused on images. Unlike images composed of regular pixels, 3D point clouds are composed of irregular three-dimensional XYZ coordinates, which are widely used in areas such as autonomous driving and 3D measurement. As many deep neural networks have been developed for processing 3D point clouds, these networks also face the risk of backdoor attacks. Nevertheless, backdoor attacks on 3D point clouds have rarely been investigated. This paper proposes a stealthy backdoor attack on point clouds in the physical world, aiming to generate trainable non-rigid deformations as backdoor patterns. Instead of directly adding backdoor patterns onto the point clouds, we deform the 3D space of the point clouds to a new space, ensuring that all point clouds have the same backdoor deformation. We use point cloud alignment to overcome the inconsistency of backdoor deformation caused by shifting and scaling in the physical world. We also propose a physical transformation layer to combat the physical transformations. Additionally, we propose mask contrast learning to eliminate pseudo backdoor patterns to make the network’s backdoor property stealthier. Extensive experiments indicate that the proposed method can achieve better attack success rates and stealthiness.
Le Feng, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
Comput. J.4
2024 Removing Watermarks For Image Processing Networks Via Referenced Subspace Attention
abstract
Abstract Deep neural network model extraction attack is the process of retraining a surrogate model based on the outputs of a target model with a given set of inputs. Such attacks are hard to defend for the sake of model owners’ interest. Recently, some work propose model watermarking scheme for image processing networks, which is able to prove the intellectual property of deep models even after the model extraction attack. This scheme makes sure that, once the target model (an image processing network) is watermarked, we can extract the watermark from the output of the surrogate model. In this paper, we propose a new model extraction attack scheme to fight against the latest method. Instead of directly using the output images of a target model, we propose to use their reconstructed versions for model retraining, where an asymmetrical UNet is proposed for image reconstruction. To thoroughly remove the watermarking traces, we propose and incorporate a referenced subspace attention module in the asymmetrical UNet, which removes the watermark by projecting the outputs of the target model into the subspaces of the reference image. Various experiments demonstrate the effectiveness of our attack.
Yuliang Xue, Yuhao Zhu 0006, Zhiying Zhu 0001, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
Comput. J.4
2024 Recent Advances in Deep Learning Model Security
Guorui Feng, Sheng Li 0006
Pattern Recognit. Lett.2
2024 Robust Image Steganography Against General Downsampling Operations With Lossless Secret Recovery
abstract
Resisting the operations in lossy channels is a challenge for image steganography. In this article, we propose a novel robust steganographic method to resist the image downsampling operation. Unlike the existing schemes, our method guarantees lossless secret recovery from the stego-image after general image downsampling operations, which considers the undetectability of the stego-images on both sides (sender and receiver) of the lossy channel. We first downsample the cover image to get its downsampled version on the receiver side, and select a set of embeddable pixels (i.e., the pixels that can be modified for data embedding) from the downsampled image. Then, we generate a stego-image on the receiver side (termed as the receiver stego-image) such that the distortion caused by the data embedding is minimized. Based on the receiver stego-image, we modify the cover image to produce the stego-image on the sender side (termed as the sender stego-image). The modification takes the embedding cost into account and makes sure that the downsampled version of the sender stego-image produces the same embeddable pixels as the receiver stego-image. Experimental results show that our method performs significantly better than the existing schemes in terms of robustness and undetectability for resisting general image downsampling operations.
Sheng Li 0006, Zichi Wang, Xiudong Zhang, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.1
2024 Establishing Robust Generative Image Steganography via Popular Stable Diffusion
abstract
Generative steganography, a novel paradigm in information hiding, has garnered considerable attention for its potential to withstand steganalysis. However, existing generative steganography approaches suffer from the limited visual quality of generated images and are challenging to apply to lossy transmissions in real-world scenarios with unknown channel attacks. To address these issues, this paper proposes a novel robust generative image steganography scheme, facilitating zero-shot text-driven stego image generation without the need for additional training or fine-tuning. Specifically, we employ the popular Stable Diffusion model as the backbone generative network to establish a covert transmission channel. Our proposed framework overcomes the challenges of numerical instability and perturbation sensitivity inherent in diffusion models. Adhering to Kerckhoff’s principle, we propose a novel mapping module based on dual keys to enhance robustness and security under lossy transmission conditions. Experimental results showcase the superior performance of our method in terms of extraction accuracy, robustness, security, and image quality.
Xiaoxiao Hu, Sheng Li 0006, Qichao Ying, Wanli Peng, Xinpeng Zhang 0001, Zhenxing Qian
IEEE Trans. Inf. Forensics Secur.2
2024 Multi-Source Style Transfer via Style Disentanglement Network
abstract
Despite the great success of deep neural networks for style transfer tasks, the entanglement of content and style in images leads to more style information not being captured. To tackle this problem, a novel style disentanglement network is proposed to transfer multi-source style elements. Specifically, we specialize in designing a learnable content style separation module, which can efficiently extract content and style components from images in the latent space. This method differs from the previous approaches by predefining content and style layers in the network. Under the condition of content and style separation, we continue to propose the multi-style swap module, which allows the content image to match more style elements. Additionally, by introducing alternate training strategies for the main and auxiliary decoders as well as style disentanglement loss, the stylized results look very similar to the original artworks. Experimental results demonstrate the superiority of our proposed method compared with existing schemes.
Sheng Li 0006, Zichi Wang, Xinpeng Zhang 0001, Guorui Feng
IEEE Trans. Multim.2
2023 Steganography of Steganographic Networks
abstract
Steganography is a technique for covert communication between two parties. With the rapid development of deep neural networks (DNN), more and more steganographic networks are proposed recently, which are shown to be promising to achieve good performance. Unlike the traditional handcrafted steganographic tools, a steganographic network is relatively large in size. It raises concerns on how to covertly transmit the steganographic network in public channels, which is a crucial stage in the pipeline of steganography in real world applications. To address such an issue, we propose a novel scheme for steganography of steganographic networks in this paper. Unlike the existing steganographic schemes which focus on the subtle modification of the cover data to accommodate the secrets. We propose to disguise a steganographic network (termed as the secret DNN model) into a stego DNN model which performs an ordinary machine learning task (termed as the stego task). During the model disguising, we select and tune a subset of filters in the secret DNN model to preserve its function on the secret task, where the remaining filters are reactivated according to a partial optimization strategy to disguise the whole secret DNN model into a stego DNN model. The secret DNN model can be recovered from the stego DNN model when needed. Various experiments have been conducted to demonstrate the advantage of our proposed method for covert communication of steganographic networks as well as general DNN models.
Guobiao Li, Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian
AAAI2
2023 Forward Creation, Reverse Selection: Achieving Highly Pertinent Multimodal Responses in Dialogue Contexts
abstract
Multimodal Dialogue agents are often required to respond to conversation history using both textual and visual content. Even though current dialogue studies predominantly strive to generate natural texts or images, they fall short in considering the relevance of multimodal responses within a dialogue context, consequently confining agents from making prudent choices based on multiple alternatives and their associated relevance scores for decision-making. In this paper, we present a bidirectional multimodal dialogue framework that skillfully combines the forward generation of multiple text and image response candidates with reverse selection guided by relevance scores evaluated on dialogue context, facilitating agents in selecting the most suitable multimodal responses. Specifically, the forward generation aspect of our framework leverages a stage-wise approach, first producing textual replies and composite visual descriptions from the dialogue context, followed by the generation of visual responses aligned with the descriptions. In the reverse selection process, visual responses are translated into tangible descriptive texts that, in conjunction with textual responses, are inversely tied back to the dialogue context for relevance assessment, assigning a reference score to each multimodal response candidate to assist the intelligent agent in making informed decisions. Experimental outcomes demonstrate that our proposed bidirectional dialogue response framework markedly elevates performance in both automatic and human evaluations, yielding a range of contextually fitting multimodal responses for selection.
Ge Luo 0003, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
CIKM4
2023 DRAW: Defending Camera-shooted RAW against Image Manipulation
abstract
RAW files are the initial measurement of scene radiance widely used in most cameras, and the ubiquitously-used RGB images are converted from RAW data through Image Signal Processing (ISP) pipelines. Nowadays, digital images are risky of being nefariously manipulated. Inspired by the fact that innate immunity is the first line of body defense, we propose DRAW, a novel scheme of defending images against manipulation by protecting their sources, i.e., camera-shooted RAWs. Specifically, we design a lightweight Multi-frequency Partial Fusion Network (MPF-Net) friendly to devices with limited computing resources by frequency learning and partial feature fusion. It introduces invisible watermarks as protective signal into the RAW data. The protection capability can not only be transferred into the rendered RGB images regardless of the applied ISP pipeline, but also is resilient to post-processing operations such as blurring or compression. Once the image is manipulated, we can accurately identify the forged areas with a localization network. Extensive experiments on several famous RAW datasets, e.g., RAISE, FiveK and SIDD, indicate the effectiveness of our method. We hope that this technique can be used in future cameras as an option for image protection, which could effectively restrict image manipulation at the source.
Xiaoxiao Hu, Qichao Ying, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICCV4
2023 Image Protection for Robust Cropping Localization and Recovery
abstract
Existing image cropping detection schemes ignore that recovering the cropped-out contents can unveil the purpose of the behaved cropping attack. This paper presents CLR-Net, a novel image protection scheme addressing the combined challenge of image Cropping Localization and Recovery. We first protect the original image by introducing imperceptible perturbations. Then, typical image post-processing attacks are simulated to erode the protected image. On the recipient’s side, we predict the cropping mask and recover the original image. Besides, we propose a novel Fine-Grained generative JPEG simulator (FG-JPEG) as well as a feature alignment network to improve the real-world robustness. Comprehensive experiments prove that the quality of the recovered image and the accuracy of crop localization are both satisfactory.
Qichao Ying, Hang Zhou 0007, Xiaoxiao Hu, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICME5
2023 WRAP: Watermarking Approach Robust Against Film-coating upon Printed Photographs
abstract
Recently, print-resist watermarking has attracted much interest. Many watermarking schemes have been proposed to achieve robustness against printing and camera-capturing. Though these studies have shown promising results overall, they overlook the scenario of film-coating photographs, which is a significant and common scenario in real-world. The film-coating process can introduce severe distortions to the original image and easily incapacitate the watermark. To address this issue, we propose WRAP, a novel Watermarking scheme Robust Against film-coating upon Printed photographs. We first construct a large dataset with 120,000 film-coating images to train a style-transfer-based film-coating simulation network. Based on the network, we propose a comprehensive distortion layer which includes film-coating simulation and common disturbances in the printing and camera-capturing process. With the distortion layer, the entire embedding and extraction network can be trained end-to-end to gain robustness against film-coating upon printed photographs. Extensive experiments demonstrate the superior performances of our model in terms of robustness and generalization capability. Our model outperforms state-of-the-art print-resist watermarking schemes when testing in film-coating scenario and achieves outstanding performance across various datasets, types of films, and cameras. To the best of our knowledge, we are the first to conduct research on digital watermarking in film-coating scenario.
Gaozhi Liu, Yichao Si, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006, Wanli Peng
ACM Multimedia5
2023 Securing Fixed Neural Network Steganography
abstract
Image steganography is the art of concealing secret information in images in a way that is imperceptible to unauthorized parties. Recent advances show that is possible to use a fixed neural network (FNN) for secret embedding and extraction. Such fixed neural network steganography (FNNS) achieves high steganographic performance without training the networks, which could be more useful in real-world applications. However, the existing FNNS schemes are vulnerable in the sense that anyone can extract the secret from the stego-image. To deal with this issue, we propose a key-based FNNS scheme to improve the security of the FNNS, where we generate key-controlled perturbations from the FNN for data embedding. As such, only the receiver who possesses the key is able to correctly extract the secret from the stego-image using the FNN. In order to improve the visual quality and undetectability of the stego-image, we further propose an adaptive perturbation optimization strategy by taking the perturbation cost into account. Experimental results show that our proposed scheme is capable of preventing unauthorized secret extraction from the stego-images. Furthermore, our scheme is able to generate stego-images with higher visual quality than the state-of-the-art FNNS scheme, especially when the FNN is a neural network for ordinary learning tasks.
Zicong Luo, Sheng Li 0006, Guobiao Li, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia2
2023 Deep Neural Network Watermarking against Model Extraction Attack
abstract
Deep neural network (DNN) watermarking is an emerging technique to protect the intellectual property of deep learning models. At present, many DNN watermarking algorithms have been proposed to achieve provenance verification by embedding identify information into the internals or prediction behaviors of the host model. However, most methods are vulnerable to model extraction attacks, where attackers collect output labels from the model to train a surrogate or a replica. To address this issue, we present a novel DNN watermarking approach, named SSW, which constructs an adaptive trigger set progressively by optimizing over a pair of symmetric shadow models to enhance the robustness to model extraction. Precisely, we train a positive shadow model supervised by the prediction of the host model to mimic the behaviors of potential surrogate models. Additionally, a negative shadow model is normally trained to imitate irrelevant independent models. Using this pair of shadow models as a reference, we design a strategy to update the trigger samples appropriately such that they tend to persist in the host model and its stolen copies. Moreover, our method could well support two specific embedding schemes: embedding the watermark via fine-tuning or from scratch. Our extensive experimental results on popular datasets demonstrate that our SSW approach outperforms state-of-the-art methods against various model extraction attacks in whether trigger set classification accuracy based or hypothesis test based verification. The results also show that our method is robust to common model modification schemes including fine-tuning and model compression.
Jingxuan Tan, Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
ACM Multimedia5
2023 On Physically Occluded Fake Identity Document Detection
abstract
Many online applications require the users to upload their identity documents for authentication. The fake identity document is one of the main threats which compromises the security and reliability of such online applications. Existing techniques focus on the detection of digitally forged identity documents, which neglect the impact of physical forgeries. In this paper, we look into the problem of detecting physically occluded fake identity documents, which can be easily generated without any image processing knowledge. We observe that the physical occlusions inevitably produce occluded boundaries on the document. To take the advantage, we propose an Occluded Boundary Representation Learning (OBRL) module to progressively learn the occluded boundary features. These are then fed into an Occluded Boundary Message Passing (OBMP) module to effectively diffuse the physical occlusion traces to enhance the backbone features for robust detection. We newly construct a Physically Occluded Fake ID Card image dataset (POID) for evaluation. Various experiments are conducted on the POID, where our scheme is able to achieve 99.6% of accuracy in detecting physically occluded fake ID card images with a mAP of over 85% to localize the occlusion regions.
Sheng Li 0006, Silu Cao, Rui Yang 0006, Jishen Zeng, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia2
2023 Rethinking Neural Style Transfer: Generating Personalized and Watermarked Stylized Images
abstract
Neural style transfer (NST) has attracted many research interests recent years. The existing NST schemes could only generate one stylized image from a content-style image pair. They are weak in creating diverse and personalized artistic styles. On the other hand, the stylized images could easily be stolen and illegally redistributed when shared online, which has not been addressed at all in the existing NST schemes. In this paper, we propose a personalized and watermark-guided style transfer network (PWST-Net) to tackle the aforementioned issues. Our PWST-Net could generate diverse stylized images from a content-style image pair using different personalization keys. Once the style transfer is done, our stylized images are with watermarks naturally embedded for copyright protection. We propose a novel style encoder in our PWST-Net to progressively generate the stylized images, which contains a Guided Fusion (GF) block and a Style Transformation (ST) block. The GF block generates a coarse stylized image based on a personalized direction field that is specific to a personalization key and the style image. The ST block refines the coarse stylized image into the final stylized image. It embeds a watermark into the deep feature space of the stylized image during the style transfer. To make the stylized images more diverse, we further propose a new personalization loss for training our PWST-Net. Various experiments demonstrate the effectiveness of our proposed method for generating personalized and watermarked stylized images, which also outperforms the state-of-the-art NST schemes in terms of artistic visual appearance.
Sheng Li 0006, Xinpeng Zhang 0001, Guorui Feng
ACM Multimedia2
2023 RetouchingFFHQ: A Large-scale Dataset for Fine-grained Face Retouching Detection
abstract
The widespread use of face retouching filters on short-video platforms has raised concerns about the authenticity of digital appearances and the impact of deceptive advertising. To address these issues, there is a pressing need to develop advanced face retouching techniques. However, the lack of large-scale and fine-grained face retouching datasets has been a major obstacle to progress in this field. In this paper, we introduce RetouchingFFHQ, a large-scale and fine-grained face retouching dataset that contains over half a million conditionally-retouched images. RetouchingFFHQ stands out from previous datasets due to its large scale, high quality, fine-grainedness, and customization. By including four typical types of face retouching operations and different retouching levels, we extend the binary face retouching detection into a fine-grained, multi-retouching type, and multi-retouching level estimation problem. Additionally, we propose a Multi-granularity Attention Module (MAM) as a plugin for CNN backbones for enhanced cross-scale representation learning. Extensive experiments using different baselines as well as our proposed method on RetouchingFFHQ show decent performance on face retouching detection.
Qichao Ying, Sheng Li 0006, Haisheng Xu, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia3
2023 VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal Contexts
abstract
Live video commenting, or "bullet screen," is a popular social style on video platforms. Automatic live commenting has been explored as a promising approach to enhance the appeal of videos. However, existing methods neglect the diversity of generated sentences, limiting the potential to obtain human-like comments. In this paper, we introduce a novel framework called "VCMaster" for multimodal live video comments generation, which balances the diversity and quality of generated comments to create human-like sentences. We involve images, subtitles, and contextual comments as inputs to better understand complex video contexts. Then, we propose an effective Hierarchical Cross-Fusion Decoder to integrate high-quality trimodal feature representations by cross-fusing critical information from previous layers. Additionally, we develop a Sentence-Level Contrastive Loss to enlarge the distance between generated and contextual comments by contrastive learning. It helps the model to avoid the pitfall of simply imitating provided contextual comments and losing creativity, encouraging the model to achieve more diverse comments while maintaining high quality. We also construct a large-scale multimodal live video comments dataset with 292,507 comments and three sub-datasets that cover nine general categories. Extensive experiments demonstrate that our model achieves a level of human-like language expression and remarkably fluent, diverse, and engaging generated comments compared to baselines.
Ge Luo 0003, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia4
2023 Unlabeled backdoor poisoning on trained-from-scratch semi-supervised learning
Le Feng, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
Inf. Sci.4
2023 Learning to Immunize Images for Tamper Localization and Self-Recovery
abstract
Digital images are vulnerable to nefarious tampering attacks such as content addition or removal that severely alter the original meaning. It is somehow like a person without protection that is open to various kinds of viruses. Image immunization (Imuge) is a technology of protecting the images by introducing trivial perturbation, so that the protected images are immune to the viruses in that the tampered contents can be auto-recovered. This paper presents Imuge+, an enhanced scheme for image immunization. By observing the invertible relationship between image immunization and the corresponding self-recovery, we employ an invertible neural network to jointly learn image immunization and recovery respectively in the forward and backward pass. We also introduce an efficient attack layer that involves both malicious tamper and benign image post-processing, where a novel distillation-based JPEG simulator is proposed for improved JPEG robustness. Our method achieves promising results in real-world tests where experiments show accurate tamper localization as well as high-fidelity content recovery. Additionally, we show superior performance on tamper localization compared to state-of-the-art schemes based on passive forensics.
Qichao Ying, Hang Zhou 0007, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Cross-Modal Text Steganography Against Synonym Substitution-Based Text Attack
abstract
Steganography has received massive attention from the information-hiding community due to its excellent security for covert communication systems. Existing work focuses on improving security on single-modal media while cross-modal media is less explored. However, cross-modal interaction has become a prevalent social manner on current social networks, which arises potential behavioral security issues of single-modal steganography. In this letter, we propose a novel text steganography to explore the practicability of cross-modal steganography. The proposed scheme is composed with image encoder, message encoder, language model, and message extractor networks, where the generated stego texts are semantically consistent with the input reference image. In addition, current generative text steganography schemes are vulnerable to text attack based on synonym substitution since these heuristic algorithms embed information by constructing a mapping between secret messages and candidate tokens. Thus, we design a text attack layer based on synonym substitution to further improve the robustness of generated stego text. Experiments illustrate the superior performance of the proposed cross-modal steganography scheme in terms of security and robustness.
Wanli Peng, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
IEEE Signal Process. Lett.4
2023 Text Steganalysis Based on Hierarchical Supervised Learning and Dual Attention Mechanism
abstract
Recent methods with deep neural networks for text steganalysis have succeeded in mining various feature representations. However, a limited number of studies have explicitly analyzed potential security issues of generative text steganography. Furthermore, current text steganalysis approaches lack detailed consideration in the intricate design of deep learning architectures tailored to these challenges. In this article, in order to tackle these problems, we first theoretically and empirically analyze the inevitable embedding distortions of generative text steganography at a semantic and statistical levels. In light of this, we then propose an innovative text steganalysis method based on hierarchical supervised learning and a dual attention mechanism. Concretely, to extract highly effective semantic features, the proposed method involves fine-tuning a BERT extractor through the hierarchical supervised learning that combines signals from multiple softmax classifiers, rather than relying solely on the final one. The mean and standard deviation values in the Gaussian distribution of cover and stego texts are then estimated using an encoder of variational autoencoders and used to capture features representing the statistical distortion of generative text steganography. Subsequently, we introduce a dual attention mechanism that dynamically fuses the semantic and statistical features, thereby creating discriminative feature representations essential for text steganalysis. The experimental results demonstrate that our proposed text steganalysis method surpasses the current state-of-the-art techniques across three distinct text steganalysis scenarios: specific text steganalysis, semi-blind text steganalysis, and blind text steganalysis.
Wanli Peng, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Image Sanitization in Online Social Networks: A General Framework for Breaking Robust Information Hiding
abstract
With the development of robust information hiding (RIH) approaches, secret messages can be extracted successfully from stego-data after transmission through lossy channels of online social networks (OSNs). To interrupt illegal covert communications in OSNs, some methods sanitize the uploaded images by image processing operations to destroy the hidden data that may exist. However, none of the existing methods takes the RIH methods that can resist scaling into consideration, while scaling is a common operation in OSNs. In this paper, we first propose a general framework for image sanitization in OSN platforms, which serves as a countermeasure against the RIH. By using such a framework, the secret messages embedded in the upload images can be removed and the quality of the sanitized image can be well maintained. Our framework contains two deep neural networks: Scaling-Net and SC-Net. The Scaling-Net is dedicated to the sanitization of oversized images while the SC-Net is designed for other images. To achieve a good image quality, we also propose a discriminator for adversarial training of the Scaling-Net and SC-Net. Experimental results on different datasets demonstrate that our proposed method outperforms the state-of-the-art methods. The source code and pretrained models are available at our code repository (https://github.com/zyzhu19/Image_Sanitization).
Zhiying Zhu 0001, Ping Wei 0004, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Patch Diffusion: A General Module for Face Manipulation Detection
abstract
Detection of manipulated face images has attracted a lot of interest recently. Various schemes have been proposed to tackle this challenging problem, where the patch-based approaches are shown to be promising. However, the existing patch-based approaches tend to treat different patches equally, which do not fully exploit the patch discrepancy for effective feature learning. In this paper, we propose a Patch Diffusion (PD) module which can be integrated into the existing face manipulation detection networks to boost the performance. The PD consists of Discrepancy Patch Feature Learning (DPFL) and Attention-Aware Message Passing (AMP). The DPFL effectively learns the patch features by a newly designed Pairwise Patch Loss (PPLoss), which takes both the patch importance and correlations into consideration. The AMP diffuses the patches through attention-aware message passing in a graph network, where the attentions are explicitly computed based on the patch features learnt in DPFL. We integrate our PD module into four recent face manipulation detection networks, and carry out the experiments on four popular datasets. The results demonstrate that our PD module is able to boost the performance of the existing networks for face manipulation detection.
Baogen Zhang, Sheng Li 0006, Guorui Feng, Zhenxing Qian, Xinpeng Zhang 0001
AAAI2
2022 Stealthy Backdoor Attack with Adversarial Training
abstract
Research shows that deep neural networks are vulnerable to back-door attacks. The backdoor network behaves normally on clean examples, but once backdoor patterns are attached to examples, back-door examples will be classified into the target class. In the previous backdoor attack schemes, backdoor patterns are not stealthy and may be detected. Thus, to achieve the stealthiness of backdoor patterns, we explore an invisible and example-dependent backdoor attack scheme. Specifically, we employ the backdoor generation network to generate the invisible backdoor pattern for each example, and backdoor patterns are not generic to each other. However, without other measures, the backdoor attack scheme cannot bypass the neural cleanse detection. Thus, we propose adversarial training to bypass neural cleanse detection. Experiments show that the proposed backdoor attack achieves a considerable attack success rate, invisibility, and can bypass the existing defense strategies.
Le Feng, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICASSP2
2022 Encryption Resistant Deep Neural Network Watermarking
abstract
Deep neural network (DNN) watermarking is one of the main techniques to protect the DNN. Although various DNN watermarking schemes have been proposed, none of them is able to resist the DNN encryption. In this paper, we propose an encryption resistent DNN watermarking scheme, which is able to resist the parameter shuffling based DNN encryption. Unlike the existing schemes which use the kernels separately for watermarking embedding, we propose to embed the watermark into the fused kernels to resist the parameter shuffling. We further propose a MappingNet to map the the fused kernels into a higher dimension to increase the watermarking capacity. The MappingNet and the DNN are jointly trained to conduct final watermark embedding. Experimental results indicate the effectiveness of our proposed scheme for resisting the DNN encryption.
Guobiao Li, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICASSP2
2022 Object-Oriented Backdoor Attack Against Image Captioning
abstract
Backdoor attack against image classification task has been widely studied and proven to be successful, while there exist few researches on backdoor attack against vision-language models. In this paper, we explore backdoor attack towards image captioning models by poisoning training data. Assuming the attacker has total access to the training dataset, and cannot intervene in model construction or training process. Specifically, a portion of benign training samples is randomly selected to be poisoned. Afterwards, considering that the captions are usually unfolded around objects in an image, we design an object-oriented method to craft poisons, which aims to modify pixel values by a slight range with the modification number proportional to the scale of the current detected object region. After training with the poisoned data, the attacked model behaves normally on benign images, but for poisoned images, the model will generate some sentences irrelevant to the given image. The attack controls the model behavior on specific test images without scarifying the generation performance on benign test images. Our method proves the weakness of image captioning models to backdoor attack and we hope this work can raise the awareness of defending against backdoor attack in the image captioning field.
Nan Zhong, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006
ICASSP5
2022 Image Steganalysis with Convolutional Vision Transformer
abstract
Recent research has shown that deep learning based methods offer more accurate detection for image steganalysis than the traditional detection paradigm based on rich media models. Existing network architectures based on deep learning, however, stack more and more convolutional layers to increase local receptive fields for image stegananlysis. Limited by hardware, the detector with several convolutional layers may not extract features of steganography images from a global perspective effectively. In this paper, we propose a Convolutional Vision Transformer for image stegananlysis, which can capture both local and global dependencies among noise features. In image processing phase, our network preserves CNN frame for its capacity of producing image noise residuals. Different from previous methods, we utilize the attention mechanism of vision transformer for feature extraction and classification. The proposed network is validated on two public image datasets (BOSSbase 1.01 and ALASKA #2). Experimental results demonstrate that our network performs well over fixed-size dataset and arbitrary-size dataset.
Ge Luo 0003, Ping Wei 0004, Shuwen Zhu, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006
ICASSP6
2022 Joint Learning for Addressee Selection and Response Generation in Multi-Party Conversation
abstract
A large number of multi-party conversation scenarios exist in social networks, which have been seldom studied in the field of human-machine conversation. In this paper, we study a novel task of joint learning for addressee selection and response generation in multi-party conversations. Systems are expected to select whom they address and generate the corresponding response. To solve it, we propose an end-to-end addressee selection and response generation (ASRG) model, containing an addressee selection module and a response generation module. In the selection module, we develop an addressee prediction attention scheme to obtain a unique context vector for each candidate, thereby calculating the probability of the candidate more accurately. In the generation module, we propose a Focus Transformer to generate responses. These two modules are jointly learnt to fully explore the correlations between addressee and response. Experimental results show ASRG remarkably outperforms baselines and generates relevant content for different addressees.
Sheng Li 0006, Ping Wei 0004, Ge Luo 0003, Xinpeng Zhang 0001, Zhenxing Qian
ICASSP2
2022 RWN: Robust Watermarking Network for Image Cropping Localization
abstract
Image cropping can be maliciously used to manipulate the layout of an image and alter the underlying meaning. Previous image cropping detection schemes only predict whether an image has been cropped, ignoring which part of the image is cropped. This paper presents a novel robust watermarking network for image cropping localization. We train an anti-cropping processor (ACP) that embeds a watermark into a target image. The visually indistinguishable protected image is then posted on the social network instead of the original image. At the recipient’s side, ACP extracts the watermark from the attacked image, and we conduct feature matching on the original and extracted watermark to locate the position of the cropping. We further extend our scheme to detect tampering attacks on the attacked image, and a simple yet efficient method (JPEG-Mixup) is proposed that noticeably improves the generalization of JPEG robustness. We demonstrate that our scheme is the first to provide high-accuracy and robust image cropping localization.
Qichao Ying, Xiaoxiao Hu, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICIP5
2022 Invertible Image Dataset Protection
abstract
The security of data storage is a big issue for companies. They must take effective steps to prevent valuable image datasets from being stolen for illegal commercial purposes. While data encryption is a common solution, it drastically down-grades the visual quality and therefore forbids common yet trivial use such as eye-checking without a decryption. We present a novel solution for dataset protection in this scenario by robustly and reversibly transform the images into adver-sarial images. An invertible Image Dataset Protection NET-work (IDP-Net) is developed to introduce slight and acceptable changes to the images within the dataset. The protected images can be published and circulated on the social networks instead of their original version. Malicious attackers can only observe the images but cannot train pirated models based on them. Meanwhile, IDP-Net ensures the performance of au-thorized models, namely, trusted users can revert the protection and retrieve the protected images to their original version. Therefore, the dataset can be stored within the pro-tected version alone to ensure safety. Extensive experiments demonstrate that IDP-Net can better protect the security of image dataset against defensive methods compared to previ-ous methods. Besides, the introduced distortion is acceptable and the original images can be reconstructed nearly error-free.
Kejiang Chen, Xianhan Zeng, Qichao Ying, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICME4
2022 Unlabeled Backdoor Poisoning in Semi-Supervised Learning
abstract
Different from supervised learning which requires all training examples to be labeled, Semi-Supervised Learning (SSL) learns from a few labeled training examples and a large number of unlabeled training examples. Recently, studies have shown that SSL is also vulnerable to backdoor attacks. However, their performance is poor. In this paper, we propose a novel unlabeled backdoor poisoning attack against SSL, where only poisoning unlabeled examples in the training set to inject the backdoor into the network. Specifically, our attack exploits the vulnerability of SSL algorithms in guessing pseudo labels of unlabeled examples. We propose a backdoor generation network to generate poisoned examples with both the backdoor property and misleading function, thus inducing the victim model itself to mislabel the poisoned examples as the target class and causing the backdoor to be injected. Our attack achieves favorable attack success rates on the SSL algorithm while bypassing backdoor defenses.
Le Feng, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICME2
2022 Generative Steganographic Flow
abstract
Generative steganography (GS) is a new data hiding manner, featuring direct generation of stego media from secret data. Existing GS methods are generally criticized for their poor performances. In this paper, we propose a novel flow based GS approach - Generative Steganographic Flow (GSF), which provides direct generation of stego images without cover image. We take the stego image generation and secret data recovery process as an invertible transformation, and build a reversible bijective mapping between input secret data and generated stego images. In the forward mapping, secret data is hidden in the input latent of Glow model to generate stego images. By reversing the mapping, hidden data can be extracted exactly from generated stego images. Furthermore, we propose a novel latent optimization strategy to improve the fidelity of stego images. Experimental results show our proposed GSF has far better performances than SOTA works.
Ping Wei 0004, Ge Luo 0003, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006
ICME6
2022 Generative Steganography Network
abstract
Steganography usually modifies cover media to embed secret data. A new steganographic approach called generative steganography (GS) has emerged recently, in which stego images (images containing secret data) are generated from secret data directly without cover media. However, existing GS schemes are often criticized for their poor performances. In this paper, we propose an advanced generative steganography network (GSN) that can generate realistic stego images without using cover images. We firstly introduce the mutual information mechanism in GS, which helps to achieve high secret extraction accuracy. Our model contains four sub-networks, i.e., an image generator (G), a discriminator (D), a steganalyzer (S), and a data extractor (E). D and S act as two adversarial discriminators to ensure the visual quality and security of generated stego images. E is to extract the hidden secret from generated stego images. The generator G is flexibly constructed to synthesize either cover or stego images with different inputs. It facilitates covert communication by concealing the function of generating stego images in a normal generator. A module named secret block is designed to hide secret data in the feature maps during image generation, with which high hiding capacity and image fidelity are achieved. In addition, a novel hierarchical gradient decay (HGD) skill is developed to resist steganalysis detection. Experiments demonstrate the superiority of our work over existing methods.
Ping Wei 0004, Sheng Li 0006, Xinpeng Zhang 0001, Ge Luo 0003, Zhenxing Qian
ACM Multimedia2
2022 Image Generation Network for Covert Transmission in Online Social Network
abstract
Online social networks have stimulated communications over the Internet more than ever, making it possible for secret message transmission over such noisy channels. In this paper, we propose a Coverless Image Steganography Network, called CIS-Net, that synthesizes a high-quality image directly conditioned on the secret message to transfer. CIS-Net is composed of four modules, namely, the Generation, Adversarial, Extraction, and Noise Module. The receiver can extract the hidden message without any loss even the images have been distorted by JPEG compression attacks. To disguise the behaviour of steganography, we collected images in the context of profile photos and stickers and train our network accordingly. As such, the generated images are more inclined to escape from malicious detection and attack. The distinctions from previous image steganography methods are majorly the robustness and losslessness against diverse attacks. Experiments over diverse public datasets have manifested the superior ability of anti-steganalysis.
Zhengxin You, Qichao Ying, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia3
2022 On Generating Identifiable Virtual Faces
abstract
Face anonymization with generative models have become increasingly prevalent since they sanitize private information by generating virtual face images, ensuring both privacy and image utility. Such virtual face images are usually not identifiable after the removal or protection of the original identity. In this paper, we formalize and tackle the problem of generating identifiable virtual face images. Our virtual face images are visually different from the original ones for privacy protection. In addition, they are bound with new virtual identities, which can be directly used for face recognition. We propose an Identifiable Virtual Face Generator (IVFG) to generate the virtual face images. The IVFG projects the latent vectors of the original face images into virtual ones according to a user specific key, based on which the virtual face images are generated. To make the virtual face images identifiable, we propose a multi-task learning objective as well as a triplet styled training strategy to learn the IVFG. We evaluate the performance of our virtual face images using different face recognizers on diffident face image datasets, all of which demonstrate the effectiveness of the IVFG for generate identifiable virtual face images.
Zhuowen Yuan, Zhengxin You, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001, Alex Chichung Kot
ACM Multimedia3
2022 HF-Defend: Defending Against Adversarial Examples Based on Halftoning
abstract
How to deal with the adversarial examples attracts a lot of interest recently. In this paper, we propose HF-Defend: a novel method to defend against the adversarial examples based on halftoning. Unlike the existing schemes, HF-Defend thoroughly removes the adversarial perturbations by transforming a 8-bit grayscale image (or one of the RGB channels in a color image) into a 1-bit halftoned image. To maintain the image quality and content, we propose a reconstruction module for the recovery of both the main contents and fine details of the image. In particular, we newly design a nonlinear low-pass filter to extract the main contents, and a FilterNet to establish a high-pass filter for the reconstruction of fine details. Experimental results demonstrate the advantage of our HF-Defend over the existing schemes.
Gaozhi Liu, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
MMSP2
2022 A novel hashing scheme via image feature map and 2D PCA
abstract
Abstract Hashing scheme is a high‐efficiency technique for processing massive images. Two critical metrics of the hashing scheme are discrimination and robustness, but most schemes do not get satisfied classification performance between them. This paper proposes a novel hashing scheme via image feature map and 2D PCA. First, the proposed scheme extracts local phase quantization (LPQ) features in the frequency domain and local ternary pattern (LTP) features in the spatial domain, and combines them to construct an image feature map. Second, the proposed scheme conducts dimension reduction via 2D PCA for learning features from the image feature map. Last, the learned features are compressed to generate the hash sequence. Performances are tested on open image datasets. The results demonstrate that the proposed scheme can make a good balance between discrimination and robustness. In addition, the classification and copy detection of the proposed scheme are both superior to those of some famous hashing schemes.
Xiaoping Liang, Zhenjun Tang, Sheng Li 0006, Chunqiang Yu, Xianquan Zhang
IET Image Process.3
2022 Robust backdoor injection with the capability of resisting network transfer
Le Feng, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
Inf. Sci.2
2022 An SVD-based screen-shooting resilient watermarking scheme
Biao Deng, Sheng Li 0006, Zhenxing Qian
Multim. Tools Appl.2
2022 Exploring Stable Coefficients on Joint Sub-Bands for Robust Video Watermarking in DT CWT Domain
abstract
Video watermarking on the dual tree-complex wavelet (DT CWT) domain is shown to be effective to offer high robustness. Existing DT CWT video watermarking schemes tend to use all the coefficients on the high-pass sub-bands of the DT CWT domain for watermark embedding and detection, which lack of investigating the correlations among different sub-bands and fail to explore the stable coefficients for robust watermarking. In this paper, we propose a novel DT CWT video watermarking scheme by exploring the stable coefficients on joint sub-bands. We first extract a set of candidate coefficients by applying block singular value decomposition (SVD) on the DT CWT domain. Then, we simulate the watermark embedding by modifying the candidate coefficients on each sub-band, from which we identity two pairs of strongly correlated sub-bands termed as the joint sub-bands. The watermark is eventually embedded by modifying the candidate coefficients of the joint sub-bands on a level which is adaptively chosen according to the video resolution. During the watermark detection, we identify and extract a set of stable coefficients from the candidate coefficients of the joint sub-bands to verify the ownership of the video. Extensive experiments demonstrate the advantage of our propose scheme over the latest DT CWT based schemes, which also performs better than the existing non-DT CWT transformed domain video watermarking schemes.
Wennan Huan, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 High-Capacity Framework for Reversible Data Hiding in Encrypted Image Using Pixel Prediction and Entropy Encoding
abstract
While the existing reserving room before encryption (RRBE) based reversible data hiding in encrypted image (RDHEI) schemes can achieve decent embedding capacity, the capacity of the existing vacating room by encryption (VRBE) based schemes is relatively low. To address this issue, this paper proposes a generalized framework for high-capacity RDHEI for both the RRBE and VRBE cases. First, an efficient embedding room generation algorithm (ERGA) is designed to produce large embedding room using pixel prediction and entropy encoding. Then, we propose two RDHEI schemes, one for RRBE, another for VRBE. In the RRBE scenario, the image owner generates the embedding room with ERGA and encrypts the preprocessed image using stream cipher with two encryption keys. Then, the data hider locates the embedding room and embeds the additional encrypted data. In the VRBE scenario, the cover image is encrypted by an improved block modulation and permutation encryption algorithm, where the spatial redundancy in the plain-text image is greatly preserved. Then, the data hider applies ERGA on the encrypted image to generate the embedding room and conducts data embedding. For both schemes, receivers with different authentication keys can conduct either error-free data extraction or error-free image recovery. The experimental results show that the two proposed schemes outperform many state-of-the-art RDHEI schemes. Besides, they can ensure high security level, where the original image can be hardly discovered from the encrypted version before or after data hiding by unauthorized users.
Yingqiang Qiu, Qichao Ying, Yuyan Yang, Huanqiang Zeng, Sheng Li 0006, Zhenxing Qian
IEEE Trans. Circuits Syst. Video Technol.5
2022 Multi-granularity Brushstrokes Network for Universal Style Transfer
abstract
Neural style transfer has been developed in recent years, where both performance and efficiency have been greatly improved. However, most existing methods do not transfer the brushstrokes information of style images well. In this article, we address this issue by training a multi-granularity brushstrokes network based on a parallel coding structure. Specifically, we first adopt the content parsing module to obtain the spatial distribution of content image and the smoothness of different regions. Then, different brushstrokes features are transformed by a multi-granularity style-swap module guided by the region content map. Finally, the stylized features of the two branches are fused to enhance the stylized results. The multi-granularity brushstrokes network is jointly supervised by a new multi-layer brushstroke loss and pre-existing loss. The proposed method is close to the artistic drawing process. In addition, we can control whether the color of the stylized results tend to be the style image or the content image. Experimental results demonstrate the advantage of our proposed method compare with the existing schemes.
Sheng Li 0006, Xinpeng Zhang 0001, Guorui Feng
ACM Trans. Multim. Comput. Commun. Appl.2
2021 On Generating JPEG Adversarial Images
abstract
Adversarial attacks slightly perturb the original image to fool deep neural networks (DNN). Various schemes have been proposed to generate uncompressed adversarial images, which are usually ineffective after being compressed during the transmission. In this paper, we propose to generate JPEG adversarial images directly from the DNN. Two adversarial rounding schemes, including fast rounding and iterative rounding, are proposed to produce quantized DCT coefficients of JPEG adversarial images. Both schemes use the gradients of adversarial images in the DCT domain to guide the rounding. In fast rounding, we propose a novel indicator to evaluate the importance of the DCT coefficients for adversarial attacks, where only those with high importance are adversarially rounded to reduce the distortion. In iterative rounding, we additionally incorporate a loss function to mea-sure the distortion caused by adversarial rounding. The experiments show that our schemes can obtain effective JPEG adversarial images with low distortion.
Mengte Shi, Sheng Li 0006, Zhao-Xia Yin, Xinpeng Zhang 0001, Zhenxing Qian
ICME2
2021 Reversible Privacy-Preserving Recognition
abstract
In this paper, we propose a novel reversible face privacy-preserving scheme. Before uploading facial images onto the cloud, we first cover the facial region with mosaic and train an encoder to generate protected images with original facial information embedded. We train another classifier with protected images for facial expression recognition and a decoder for recovering original facial images. On the cloud service, protected images provide little identity information to malicious attackers. For low-privileged users, they can use the provided classifier to do computer vision tasks with protected images. For authorized users, after content recovery, the nor-mal usage of the facial images will not be affected. Experimental results show that the proposed method is effective in facial images recovery. In addition, the protected images can maintain similar accuracy on typical computer vision tasks compared to the original images.
Zhengxin You, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICME2
2021 Fragile Neural Network Watermarking with Trigger Image Set
Renjie Zhu, Ping Wei 0004, Sheng Li 0006, Zhao-Xia Yin, Xinpeng Zhang 0001, Zhenxing Qian
KSEM3
2021 Diffusing the Liveness Cues for Face Anti-spoofing
abstract
Face anti-spoofing is an important step for secure face recognition. One of the main challenges is how to learn and build a general classifier that is able to resist various presentation attacks. Recently, the patch-based face anti-spoofing schemes are shown to be able to improve the robustness of the classifier. These schemes extract subtle liveness cues from small local patches independently, which do not fully exploit the correlations among the patches. In this paper, we propose a Patch-based Compact Graph Network (PCGN) to diffuse the subtle liveness cues from all the patches. Firstly, the image is encoded into a compact graph by connecting each node with its backward neighbors. We then propose an asymmetrical updating strategy to update the compact graph. Such a strategy aggregates the node based on whether it is a sender or receiver, which leads to better message-passing. The updated graph is eventually decoded for making the final decision. We conduct the experiments on four public databases with four intra-database protocols and eight cross-database protocols, the results of which demonstrate the effectiveness of our PCGN for face anti-spoofing.
Sheng Li 0006, Guorui Feng, Xinpeng Zhang 0001, Zhenxing Qian
ACM Multimedia1
2021 Destroying robust steganography in online social networks
Zhiying Zhu 0001, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
Inf. Sci.2
2021 Special issue on low complexity methods for multimedia security
Guorui Feng, Sheng Li 0006, Haoliang Li, Shujun Li 0001
Multim. Syst.2
2021 Detection of Spoofing Medium Contours for Face Anti-Spoofing
abstract
Face anti-spoofing is an important step for secure face recognition. In this paper, we target on building a general classifier to detect the face images with spoofing medium contours (termed as SMCs for simplicity). To this end, we consider the task of face anti-spoofing as the detection of SMCs from the image. We propose and train a Contour Enhanced Mask R-CNN (CEM-RCNN) model for the detection. This model detects the existence of the SMCs by incorporating the contour objectness which measures how likely an object contains the SMCs. The experimental results demonstrate the generality of the CEM-RCNN for identifying the face images with SMCs, which performs significantly better than the state-of-the-art on the cross-database scenario.
Sheng Li 0006, Xinpeng Zhang 0001, Haoliang Li, Alex Chichung Kot
IEEE Trans. Circuits Syst. Video Technol.2
2020 Diversity-Based Cascade Filters for JPEG Steganalysis
abstract
Steganalysis is a technique for detecting the existence of secret information hidden in digital media. In this paper, we propose a novel scheme for JPEG steganalysis. In this scheme, we first design the diverse base filters which are able to obtain the image residuals from various directions. Then, we propose a cascade filter generation strategy to construct a set of high order cascade filters from the base filters. We further select the cascade filters with the maximum diversity. The selected filters are convolved with the decompressed JPEG image to obtain residuals which capture the subtle embedding traces. The residuals, termed as the maximum diversity cascade filter residual, are eventually used to extract features to train an ensemble classifier for classification. The experiments are carried out on the detection of stego-images generated using common JPEG steganographic schemes, the results of which demonstrate the effectiveness of the proposed scheme for JPEG steganalysis.
Guorui Feng, Xinpeng Zhang 0001, Yanli Ren, Zhenxing Qian, Sheng Li 0006
IEEE Trans. Circuits Syst. Video Technol.5
2020 Key Based Artificial Fingerprint Generation for Privacy Protection
abstract
With the widespread use of biometrics recognition systems, it is of paramount importance to protect the privacy of biometrics. In this paper, we propose to protect the fingerprint privacy by the artificial fingerprint, which is generated based on three pieces of information, i) the original minutiae positions; ii) the artificial fingerprint orientation; and iii) the artificial minutiae polarities. To make it real-look alike and diverse, we propose to generate the artificial fingerprint orientation by a model taking both the global and local fingerprint orientation into account. Its parameters can be easily guided by an user specific key with simple constraints. The artificial minutiae polarities are generated from the same key, where a block based and a function based approach are proposed for the minutiae polarities generation. These information are properly integrated to form a real-look alike artificial fingerprint. It is difficult for the attacker to distinguish such a fingerprint from the real fingerprints. If it is stolen, the complete fingerprint minutiae feature will not be compromised, and we can generate a different artificial fingerprint using another key. Experimental results show that the artificial fingerprint can be recognized accurately.
Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian, Guorui Feng, Yanli Ren
IEEE Trans. Dependable Secur. Comput.1
2019 Towards Robust Image Steganography
abstract
Posting images on social network platforms is happening everywhere and every single second. Thus, the communication channels offered by various social networks have a great potential for covert communication. However, images transmitted through such channels will usually be JPEG compressed, which fails most of the existing steganographic schemes. In this paper, we propose a novel image steganography framework that is robust for such channels. In particular, we first obtain the channel compressed version (i.e., the channel output) of the original image. Secret data is embedded into the channel compressed original image by using any of the existing JPEG steganographic schemes, which produces the stego-image after the channel transmission. To generate the corresponding image before the channel transmission (termed the intermediate image), we propose a coefficient adjustment scheme to slightly modify the original image based on the stego-image. The adjustment is done such that the channel compressed version of the intermediate image is exactly the same as the stego-image. Therefore, after the channel transmission, secret data can be extracted from the stego-image with 100% accuracy. Various experiments are conducted to show the effectiveness of the proposed framework for image steganography robust to JPEG compression.
Jinyuan Tao, Sheng Li 0006, Xinpeng Zhang 0001, Zichi Wang
IEEE Trans. Circuits Syst. Video Technol.2
2019 Toward Construction-Based Data Hiding: From Secrets to Fingerprint Images
abstract
Data hiding usually involves the alteration of a cover signal for embedding a secret message. In this paper, we propose a construction based data hiding technique which transforms a secret message into a fingerprint image directly. Unlike the conventional data hiding techniques, this scheme does not need any cover signals to participate. Instead, it generates the fingerprint image based on a piece of hologram phase constructed from the secret message. The hologram phase consists of the spiral phase and the continuous phase. Firstly, we propose to map the secret message to a polynomial and encode it into a set of points with different polarities, from which the spiral phase is computed and constructed. Then, we construct the continuous phase by decomposing a fingerprint image synthetically generated. The spiral phase and the continuous phase are combined to form the hologram phase. This is eventually used to construct a fingerprint image in a common form such as a grayscale fingerprint image, a binary fingerprint image, or a thinned fingerprint image. The secret message can be extracted by detecting the encoded points in the constructed fingerprint. We conduct the experiments by constructing fingerprint images with ordinary sizes, the results show that the secret message can be extracted accurately. It is also difficult to detect the existence of secret message from the constructed fingerprint images.
Sheng Li 0006, Xinpeng Zhang 0001
IEEE Trans. Image Process.1
2018 Lossless data hiding in JPEG bitstream using alternative embedding
Yingqiang Qiu, Han He, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
J. Vis. Commun. Image Represent.4
2018 A novel auxiliary data construction scheme for reversible data hiding in JPEG images
Jinpeng Lv, Sheng Li 0006, Xinpeng Zhang 0001
Multim. Tools Appl.2
2016 On Branded Handbag Recognition
abstract
Manufacturing branded handbags is a big business in the fashion world. Shoppers’ feedback showing photos of their purchased handbags in social networks or blogs is important for branding purposes. In this paper, we deal with handbag recognition. It is a challenging problem due to the inter-class style similarity and the intra-class color variation. We focus on developing discriminative representations of handbag style and color. For handbag style representation, two supervised mid-level patch selection procedures are proposed to select discriminative patches, regarding individual classes and pairwise classes. We also propose a low-level complementary feature, extracted from texture-enhanced mid-level patches, to capture the fine details of the mid-level patches. For handbag color representation, we propose to extract dominant color features to handle the illumination changes. The performance of our proposed method is evaluated on a newly built branded handbag dataset. The results show that our method performs favorably in recognizing handbags, with around$10\%$improvement in accuracy when compared with the existing fine-grained or generic object recognition methods.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Multim.2
2015 Joint learning for image-based handbag recommendation
abstract
Fashion recommendation helps shoppers to find desirable fashion items, which facilitates online interaction and product promotion. In this paper, we propose a method to recommend handbags to each shopper, based on the handbag images the shopper has clicked. This is performed by Joint learning of attribute Projection and One-class SVM classification (JPO) based on the images of the shopper's preferred handbags. More specifically, for the handbag images clicked by each shopper, we project the original image feature space into an attribute space which is more compact. The projection matrix is learned jointly with a one-class SVM to yield a shopper-specific one-class classifier. The results show that the proposed JPO handbag recommendation performs favorably based on initial subject testing.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
ICME2
2015 DeepBag: Recognizing Handbag Models
abstract
In this paper, we address the problem of branded handbag recognition. It is a challenging problem due to the non-rigid deformation, illumination changes, and inter-class similarity. We propose a novel framework based on deep convolutional neural network (CNN). Concretely, we propose a new CNN model, called feature selective joint classification - regression CNN (FSCR-CNN). Its advantages lie in two folds: 1) it alleviates the illumination changes by a feature selection strategy to focus on the color- nondiscriminative features in the network learning, and 2) rather than only targeting on the hard label (i.e., the handbag model), it also incorporates a soft label (i.e., a distribution measuring the similarity between the ground truth model and all the models to be trained) to construct the loss function for training CNN, which leads to a better classifier for handbags with large inter-class similarity. We evaluate the performance of our framework on a newly built branded handbag dataset. The results show that it performs favorably for recognizing handbags with 94.48% in accuracy. We also apply the proposed FSCR-CNN model in recognizing other fine-grained objects with state-of-the-art CNN architectures, which is able to achieve over 5% improvement in accuracy.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Multim.2
2014 Complementary feature extraction for branded handbag recognition
abstract
Fine-grained object recognition aims at recognizing objects belonging to the same basic-level class such as dog, bird or fish, which is a challenging problem in computer vision. In this paper, we consider the problem of recognizing handbags that belong to a specific brand. In order to identify the subtle differences among handbags, we propose to enhance the handbag local structure pattern by using the Hölder exponent, and extract the feature from the enhanced handbag image to complement the feature extracted directly from the original handbag image. We term such two types of features as the complementary and original features. These features will then be fused by using Multiple Kernel Learning (MKL) for branded handbag recognition. We conduct the experiments on a newly built branded handbag dataset, the results of which demonstrate the effectiveness of the proposed complementary feature in recognizing the handbags.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
ICIP2
2013 Fingerprint Combination for Privacy Protection
abstract
We propose here a novel system for protecting fingerprint privacy by combining two different fingerprints into a new identity. In the enrollment, two fingerprints are captured from two different fingers. We extract the minutiae positions from one fingerprint, the orientation from the other fingerprint, and the reference points from both fingerprints. Based on this extracted information and our proposed coding strategies, a combined minutiae template is generated and stored in a database. In the authentication, the system requires two query fingerprints from the same two fingers which are used in the enrollment. A two-stage fingerprint matching process is proposed for matching the two query fingerprints against a combined minutiae template. By storing the combined minutiae template, the complete minutiae feature of a single fingerprint will not be compromised when the database is stolen. Furthermore, because of the similarity in topology, it is difficult for the attacker to distinguish a combined minutiae template from the original minutiae templates. With the help of an existing fingerprint reconstruction approach, we are able to convert the combined minutiae template into a real-look alike combined fingerprint. Thus, a new virtual identity is created for the two different fingerprints, which can be matched using minutiae-based fingerprint matching algorithms. The experimental results show that our system can achieve a very low error rate with FRR = 0.4% at FAR = 0.1%. Compared with the state-of-the-art technique, our work has the advantage in creating a better new virtual identity when the two different fingerprints are randomly chosen.
Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2012 An Improved Scheme for Full Fingerprint Reconstruction
abstract
Different fingerprint recognition systems store minutiae-based fingerprint templates differently. Some store them inside a small token; some can be found in a server database. As the minutiae template is very compact, many take it for granted that the template does not contain sufficient information for reconstructing the original fingerprint. This paper proposes a scheme to reconstruct a full fingerprint image from the minutiae points based on the amplitude and frequency modulated (AM-FM) fingerprint model. The scheme starts with generating a binary ridge pattern which has a similar ridge flow to that of the original fingerprint. The continuous phase is intuitively reconstructed by removing the spirals in the phase image estimated from the ridge pattern. To reduce the artifacts due to the discontinuity in the continuous phase, a refinement process is introduced for the reconstructed phase image, which is the combination of the continuous phase and the spiral phase (corresponding to the minutiae). Finally, the refined phase image is used to produce a thinned version of the fingerprint, from which a real-look alike gray-scale fingerprint image is reconstructed. The experimental results show that our proposed scheme performs better than the-state-of-the-art technique.
Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2011 A novel system for fingerprint privacy protection1
abstract
This paper proposes a novel system for protecting the fingerprint privacy without using a token or key. In the enrollment, two fingerprints are captured from two of an user's fingers. We extract the minutiae positions from one fingerprint, the orientation from the other fingerprint and the primary cores from both fingerprints. Based on these extracted information, a combined minutiae template is generated and stored in a database. In the authentication, the user needs to provide two query fingerprints from the same two fingers which are used in the enrollment. By storing the combined minutiae template, the complete minutiae feature of a single fingerprint will not be compromised when the database is stolen. Furthermore, because of the similarity in topology, it is also difficult for the attacker to distinguish our template from the minutiae of an original fingerprint. We evaluate the performance of our system over the FVC2002 DB2_A database. The results show that the False Rejection Rate of our system is 3% when the False Acceptance Rate is 0.01%.
Sheng Li 0006, Alex Chichung Kot
IAS1
2011 Privacy Protection of Fingerprint Database
abstract
A fingerprint authentication system for the privacy protection of the fingerprint template stored in a database is introduced here. The considered fingerprint data is a binary thinned fingerprint image, which will be embedded with some private user information without causing obvious abnormality in the enrollment phase. In the authentication phase, these hidden user data can be extracted from the stored template for verifying the authenticity of the person who provides the query fingerprint. A novel data hiding scheme is proposed for the thinned fingerprint template. This scheme does not produce any boundary pixel in the thinned fingerprint during data embedding. Thus, the abnormality caused by data hiding is visually imperceptible in the marked-thinned fingerprint. Compared with using existing binary image data hiding techniques, the proposed method causes the least abnormality for a thinned fingerprint without compromising the performance of the fingerprint identification.
Sheng Li 0006, Alex Chichung Kot
IEEE Signal Process. Lett.1
2010 Privacy protection of fingerprint database using lossless data hiding
abstract
In this paper, we introduce a fingerprint authentication system for protecting the privacy of the fingerprint template stored in a database. The template, which is a binary fingerprint image after thinning, will be embedded with private personal data in the user enrollment phase. In the user authentication phase, these hidden personal data can be extracted from the stored template for verifying the authenticity of the person who provides the query fingerprint. A novel lossless data hiding scheme is proposed for a thinned fingerprint. By adopting “embeddability criterion”, data is hidden into the template by just adding some boundary pixels in the template. These boundary pixels can be extracted and removed to reconstruct the original thinned fingerprint so that fingerprint matching accuracy is not affected. Compared with using existing binary image data hiding techniques, our scheme has a better performance for a thinned fingerprint.
Sheng Li 0006, Alex Chichung Kot
ICME1