VLDB 2026 Research / reviewers in the wild / expert
Nan Zhong
dblp:129/2343
· DBLP profile ↗
18ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised AI-Generated Image Detection: A Camera Metadata PerspectiveabstractThe proliferation of AI-generated imagery poses escalating challenges for multimedia forensics, yet many existing detectors depend on assumptions about the internals of specific generative models, limiting their cross-model applicability. We introduce a self-supervised approach for detecting AI-generated images that leverages camera metadata-specifically exchangeable image file format (EXIF) tags-to learn features intrinsic to digital photography. Our pretext task trains a feature extractor solely on camera-captured photographs by classifying categorical EXIF tags (e.g., camera model and scene type) and pairwise-ranking ordinal and continuous EXIF tags (e.g., focal length and aperture value). Using these EXIF-induced features, we first perform one-class detection by modeling the distribution of photographic images with a Gaussian mixture model and flagging low-likelihood samples as AI-generated. We then extend to binary detection that treats the learned extractor as a strong regularizer for a classifier of the same architecture, operating on high-frequency residuals from spatially scrambled patches. Extensive experiments across various generative models demonstrate that our EXIF-induced detectors substantially advance the state of the art, delivering strong generalization to in-the-wild samples and robustness to common benign image perturbations. Nan Zhong, Mian Zou, Zhenxing Qian, Xinpeng Zhang 0001, Baoyuan Wu, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Hiding Images in Diffusion Models by Editing Learned Score FunctionsabstractHiding data using neural networks (i.e., neural steganography) has achieved remarkable success across both discriminative classifiers and generative adversarial networks. However, the potential of data hiding in diffusion models remains relatively unexplored. Current methods exhibit limitations in achieving high extraction accuracy, model fidelity, and hiding efficiency due primarily to the entanglement of the hiding and extraction processes with multiple denoising diffusion steps. To address these, we describe a simple yet effective approach that embeds images at specific timesteps in the reverse diffusion process by editing the learned score functions. Additionally, we introduce a parameter-efficient fine-tuning method that combines gradient-based parameter selection with low-rank adaptation to enhance model fidelity and hiding efficiency. Comprehensive experiments demonstrate that our method extracts high-quality images at human-indistinguishable levels, replicates the original model behaviors at both sample and population levels, and embeds images orders of magnitude faster than prior methods. Besides, our method naturally supports multi-recipient scenarios through independent extraction channels. Yunqiao Yang 0001, Nan Zhong, Kede Ma |
CVPR | 3 |
| 2025 | Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated ImagesabstractThe prevalence of AI-generated images has evoked concerns regarding the potential misuse of image generation technologies. In response, numerous detection methods aim to identify AI-generated images by analyzing generative artifacts. Unfortunately, most detectors quickly become obsolete with the development of generative models. In this paper, we first design a low-level feature extractor that transforms spatial images into feature space, where different source images exhibit distinct distributions. The pretext task for the feature extractor is to distinguish between images that differ only at the pixel level. This image set comprises the original image as well as versions that have been subjected to varying levels of noise and subsequently denoised using a pre-trained diffusion model. We employ the diffusion model as a denoising tool rather than an image generation tool. Then, we frame the AI-generated image detection task as a one-class classification. We estimate the low-level intrinsic feature distribution of real photographic images and identify features that deviate from this distribution as indicators of AI-generated images. We evaluate our method against over 20 different generative models, including those in GenImage and DRCT-2M datasets. Extensive experiments demonstrate its effectiveness on AI-generated images produced not only by diffusion models but also by GANs, flow-based models, and their variants. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
CVPR | 1 |
| 2025 | Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
Mian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan, Kede Ma |
ICCV | 2 |
| 2025 | Watermarking Diffusion Models By Constructing Generative ClassifiersabstractDiffusion models have demonstrated outstanding performance in generating photorealistic images, garnering significant interest for their application in a wide range of downstream tasks. However, their powerful generative capabilities and increasing deployment raise serious concerns regarding copyright protection. While watermarking techniques for diffusion models have been proposed to address potential misuse, ensuring robustness against model weight modifications remains a persistent challenge. In this paper, we propose a robust watermarking method for diffusion models. Our approach leverages the diffusion model as a generative prior and constructs a generative classifier to facilitate an effective watermarking scheme. Drawing inspiration from fragile watermarking techniques for neural network classifiers—where trigger images are optimized to make predictions highly sensitive to weight changes—we invert this idea to enhance robustness. Specifically, we employ a bi-level optimization framework that jointly optimizes the trigger images and the parameters of the diffusion model. Experiments demonstrate the effectiveness of our proposed method in terms of watermark extraction accuracy, robustness, and model fidelity. Nan Zhong |
ICIP | 2 |
| 2025 | Fine-grained Prompt Screening: Defending Against Backdoor Attack on Text-to-Image Diffusion ModelsabstractText-to-image (T2I) diffusion models exhibit impressive generation capabilities in recently studies. However, they are vulnerable to backdoor attacks, where model outputs are manipulated by malicious triggers. In this paper, we propose a novel input-level defense method, called Fine-grained Prompt Screening (GrainPS). Our method is motivated by the phenomenon, i.e., Semantics Misalignment, where the backdoor trigger causes the inconsistency between the cross-attention projections of object words (the key words to determine the main content of the generated image) and their true semantics. In particular, we divide each prompt into pieces and conduct fine-grained analysis by examining the impact of the trigger on object words in the cross-attention layers rather than their global influence on the entire generated image. To assess the impact of each word on object words, we formulate "semantics alignment score'' as the metric with a carefully crafted detection strategy to identify the trigger. Therefore, our implementation can detect backdoor input prompts and localize of triggers simultaneously. Evaluations across four advanced backdoor attack scenarios demonstrate the effectiveness of our proposed defense method. Nan Zhong, Guobiao Li, Anda Cheng, Yinggui Wang, Zhenxing Qian, Xinpeng Zhang 0001 |
IJCAI | 2 |
| 2024 | Revocable Backdoor for Deep Model TradingabstractDeep models are being applied in numerous fields and have become a new important digital product. Meanwhile, previous studies have shown that deep models are vulnerable to backdoor attacks, in which compromised models return attacker-desired results when a trigger appears. Backdoor attacks severely break the trustworthiness of deep models. In this paper, we turn this weakness of deep models into a strength, and propose a novel revocable backdoor and deep model trading scenario. Specifically, we aim to compromise deep models without degrading their performance, meanwhile, we can easily detoxify poisoned models without re-training the models. We design specific mask matrices to manage the internal feature maps of the models. These mask matrices can be used to deactivate the backdoors. The revocable backdoor can be adopted in the deep model trading scenario. Sellers train models with revocable backdoors as a trial version. Buyers pay a deposit to sellers and obtain a trial version of the deep model. If buyers are satisfied with the trial version, they pay a final payment to sellers and sellers send mask matrices to buyers to withdraw revocable backdoors. We demonstrate the feasibility and robustness of our revocable backdoor by various datasets and network architectures. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
ECAI | 2 |
| 2024 | Sparse Backdoor Attack Against Neural NetworksabstractAbstract Recent studies show that neural networks are vulnerable to backdoor attacks, in which compromised networks behave normally for clean inputs but make mistakes when a pre-defined trigger appears. Although prior studies have designed various invisible triggers to avoid causing visual anomalies, they cannot evade some trigger detectors. In this paper, we consider the stealthiness of backdoor attacks from input space and feature representation space. We propose a novel backdoor attack named sparse backdoor attack, and investigate the minimum required trigger to induce the well-trained networks to make incorrect results. A U-net-based generator is employed to create triggers for each clean image. Considering the stealthiness of the trigger, we restrict the elements of the trigger between −1 and 1. In the aspect of the feature representation domain, we adopt an entanglement cost function to minimize the gap between feature representations of benign and malicious inputs. The inseparability of benign and malicious feature representations contributes to the stealthiness of our attack against various model diagnosis-based defences. We validate the effectiveness and generalization of our method by conducting extensive experiments on multiple datasets and networks. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
Comput. J. | 1 |
| 2024 | Backdoor attack detection via prediction trustworthiness assessment
Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
Inf. Sci. | 1 |
| 2023 | Multimodal Sentiment Analysis Based on Multiple Stacked Attention MechanismsabstractDeciphering sentiments or emotions in face-to-face human interactions is an inherent capability of human intelligence, and thus a natural goal of artificial intelligence. The proliferation of multimedia data in video sites gives rise to multimodal sentiment analysis in various applications and research fields such as movie and product review, opinion polling, and affective computing. In order to improve the performance of multimodal sentiment analysis task, this paper proposes a novel neural network with multiple stacked attention mechanism (MSAM) on multimodal data containing texts, video, and audio at an utterance level. We conduct experiments using two benchmark datasets, namely CMU Multi-modal Opinion-level Sentiment Intensity (CMU-MOSI) corpus, and CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus. Compared with a comprehensive set of state-of-the-art baselines, the evaluation results demonstrate the effectiveness of our proposed MSAM network. Yingshan Shen, Nan Zhong, Dongqing Song, Huijuan Hu, Dingju Zhu, Lihua Cai |
CSCWD | 3 |
| 2023 | Physical Invisible Backdoor Based on Camera ImagingabstractBackdoor attack aims to compromise a model, which returns an adversary-wanted output when a specific trigger pattern appears yet behaves normally for clean inputs. Current backdoor attacks require changing pixels of clean images, which results in poor stealthiness of attacks and increases the difficulty of the physical implementation. This paper proposes a novel physical invisible backdoor based on camera imaging without changing nature image pixels. Specifically, a compromised model returns a target label for images taken by a particular camera, while it returns correct results for other images. To implement and evaluate the proposed backdoor, we take shots of different objects from multi-angles using multiple smartphones to build a new dataset of 21,500 images. Conventional backdoor attacks work ineffectively with some classical models, such as ResNet18, over the above-mentioned dataset. Therefore, we propose a three-step training strategy to mount the backdoor attack. First, we design and train a camera identification model with the phone IDs to extract the camera fingerprint feature. Subsequently, we elaborate a special network architecture, which is easily compromised by our backdoor attack, by leveraging the attributes of the CFA interpolation algorithm and combining it with the feature extraction block in the camera identification model. Finally, we transfer the backdoor from the elaborated special network architecture to the classical architecture model via teacher-student distillation learning. Since the trigger of our method is related to the specific phone, our attack works effectively in the physical world. Experiment results demonstrate the feasibility of our proposed approach and robustness against various backdoor defences. Yusheng Guo, Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | Deep Neural Network Watermarking against Model Extraction AttackabstractDeep neural network (DNN) watermarking is an emerging technique to protect the intellectual property of deep learning models. At present, many DNN watermarking algorithms have been proposed to achieve provenance verification by embedding identify information into the internals or prediction behaviors of the host model. However, most methods are vulnerable to model extraction attacks, where attackers collect output labels from the model to train a surrogate or a replica. To address this issue, we present a novel DNN watermarking approach, named SSW, which constructs an adaptive trigger set progressively by optimizing over a pair of symmetric shadow models to enhance the robustness to model extraction. Precisely, we train a positive shadow model supervised by the prediction of the host model to mimic the behaviors of potential surrogate models. Additionally, a negative shadow model is normally trained to imitate irrelevant independent models. Using this pair of shadow models as a reference, we design a strategy to update the trigger samples appropriately such that they tend to persist in the host model and its stolen copies. Moreover, our method could well support two specific embedding schemes: embedding the watermark via fine-tuning or from scratch. Our extensive experimental results on popular datasets demonstrate that our SSW approach outperforms state-of-the-art methods against various model extraction attacks in whether trigger set classification accuracy based or hypothesis test based verification. The results also show that our method is robust to common model modification schemes including fine-tuning and model compression. Jingxuan Tan, Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006 |
ACM Multimedia | 2 |
| 2022 | Object-Oriented Backdoor Attack Against Image CaptioningabstractBackdoor attack against image classification task has been widely studied and proven to be successful, while there exist few researches on backdoor attack against vision-language models. In this paper, we explore backdoor attack towards image captioning models by poisoning training data. Assuming the attacker has total access to the training dataset, and cannot intervene in model construction or training process. Specifically, a portion of benign training samples is randomly selected to be poisoned. Afterwards, considering that the captions are usually unfolded around objects in an image, we design an object-oriented method to craft poisons, which aims to modify pixel values by a slight range with the modification number proportional to the scale of the current detected object region. After training with the poisoned data, the attacked model behaves normally on benign images, but for poisoned images, the model will generate some sentences irrelevant to the given image. The attack controls the model behavior on specific test images without scarifying the generation performance on benign test images. Our method proves the weakness of image captioning models to backdoor attack and we hope this work can raise the awareness of defending against backdoor attack in the image captioning field. Nan Zhong, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006 |
ICASSP | 2 |
| 2022 | Imperceptible Backdoor Attack: From Input Space to Feature RepresentationabstractBackdoor attacks are rapidly emerging threats to deep neural networks (DNNs). In the backdoor attack scenario, attackers usually implant the backdoor into the target model by manipulating the training dataset or training process. Then, the compromised model behaves normally for benign input yet makes mistakes when the pre-defined trigger appears. In this paper, we analyze the drawbacks of existing attack approaches and propose a novel imperceptible backdoor attack. We treat the trigger pattern as a special kind of noise following a multinomial distribution. A U-net-based network is employed to generate concrete parameters of multinomial distribution for each benign input. This elaborated trigger ensures that our approach is invisible to both humans and statistical detection. Besides the design of the trigger, we also consider the robustness of our approach against model diagnose-based defences. We force the feature representation of malicious input stamped with the trigger to be entangled with the benign one. We demonstrate the effectiveness and robustness against multiple state-of-the-art defences through extensive datasets and networks. Our trigger only modifies less than 1\% pixels of a benign image while the modification magnitude is 1. Our source code is available at https://github.com/Ekko-zn/IJCAI2022-Backdoor. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
IJCAI | 1 |
| 2021 | Undetectable Adversarial Examples Based on Microscopical RegularizationabstractRecent works have demonstrated that neural networks are vulnerable to adversarial examples. Although existing methods have achieved satisfactory attack success rates, most adversarial examples can be detected by statistical analysis and further removed. In previous methods, adversarial perturbations are added using adversarial loss and distance metrics, in which the positions of modified pixels are not considered. In this paper, we elaborate a microscopical regularization that introduces adversarial perturbations onto rich texture regions. The microscopical regularization is used to evaluate pixel-level differences between a normal image and its adversarial version. We further propose a novel optimization strategy of modification probability matrices to minimize the loss function that satisfies the restriction of L∞. Through extensive experiments, we show that our method can resist statistical analysis by a large margin and achieve better visual quality than others. The proposed microscopical regularization can also be combined with existing approaches to enhance the undetectability and robustness. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
ICME | 1 |
| 2021 | Deep Neural Network RetrievalabstractWith the rapid development of deep learning-based techniques, the general public can use a lot of "machine learning as a service" (MLaaS), which provides end-to-end machine learning solutions. Taking the image classification task as an example, users only need to update their dataset and labels to MLaaS without requiring the specific knowledge of machine learning or a concrete structure of the classifier. Afterward, MLaaS returns a well-trained classifier to them. In this paper, we explore a potential novel task named "deep neural network retrieval" and its application which helps MLaaS to save computation resources. MLaaS usually owns a huge amount of well-trained models for various tasks and datasets. If a user requires a task that is similar to the one having been finished previously, MLaaS can quickly retrieve a model rather than training from scratch. We propose a pragmatic solution and two different approaches to extract the semantic feature of DNN representing the function of DNN, which is analogous to the usage of word2vec in natural language processing. The semantic feature of DNN can be expressed as a vector by feeding some well-designed litmus images into the DNN or as a matrix by reversely constructing the most desired input of DNN. Both methods can consider the topological information and parameters of the DNN simultaneously. Extensive experiments, including multiple datasets and networks, also demonstrate the efficiency of our method and show the high accuracy of deep neural network retrieval. Nan Zhong, Zhenxing Qian, Xinpeng Zhang 0001 |
ACM Multimedia | 1 |
| 2021 | Batch Steganography via Generative NetworkabstractBatch steganography is a technique that hides information into multiple covers. To achieve a better performance on the security of data hiding, we propose a novel strategy of batch steganography using a generative network. In this method, the approaches of cover selection, payload allocation, and distortion evaluation are considered in the round. We define a quality metric to evaluate the distortion between the cover image and the stego. When training the generation function, we define an objective function containing two parts: the entropy loss and the steganalytic loss. While the entropy loss is used to represent the gap between the payload inside stego images and the entire embedding capacity, the steganalytic loss is used to assess the data embedding impact using the proposed quality metric. With back-propagation, we minimize the objective function to obtain an optimal solution. Accordingly, different payloads can be allocated to different images, and the ± 1 modification probability for pixels in each cover can be calculated. Finally, we embed information into the selected images by STC. Experimental results show that the proposed method achieves a better undetectability against modern steganalytic tools. Nan Zhong, Zhenxing Qian, Zichi Wang, Xinpeng Zhang 0001, Xiaolong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Defocus blur detection based on multiscale SVD fusion in gradient domain
Huimei Xiao, Wei Lu 0001, Nan Zhong, Yuileong Yeung, Junjia Chen, Wei Sun 0007 |
J. Vis. Commun. Image Represent. | 4 |