VLDB 2026 Research / reviewers in the wild / expert
Renshuai Tao
dblp:250/2018
· DBLP profile ↗
32ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0001-5695-2009ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dormant Backdoor: Weaponizing Model Finetuning for Feasible Backdoor Attacks Against Pretrained ModelsabstractAs the pretraining-finetuning paradigm becomes dominant in modern AI, the security of model supply chains faces new risks from backdoor attacks. Existing work primarily studies backdoors injected during pretraining and treats subsequent finetuning with clean data as a defense, while recent finetuning-activated attacks assume white-box access to the downstream data distribution, which is rarely realistic in practice. We introduce Dormant Backdoor, a finetuning-activated attack that requires no prior knowledge of downstream tasks. Instead of binding the backdoor to static input patterns, Dormant Backdoor exploits the universal dynamics of gradient-based optimization as a process-as-trigger mechanism. We formulate the attack as a bilevel optimization problem that simulates the victim's finetuning trajectory on proxy data, and jointly optimizes the poisoned model and trigger under lethality, utility, and stealth objectives. Before finetuning, the poisoned model remains behaviorally close to a clean model and can evade existing backdoor detectors; after finetuning, the same adaptation process reliably amplifies the backdoor on diverse downstream datasets and finetuning strategies. Our results reveal a previously underexplored class of process-as-trigger vulnerabilities and highlight the need for defenses that explicitly secure the model adaptation process. Ruitao Li, Jiakai Wang, Hairong Chen, Huihu Ding, Jinghan Zhou, Renshuai Tao |
AAAI | 6 |
| 2026 | RAIN: Redundancy-Aware Latent Injection for Quality-Preserving Image WatermarkingabstractDiffusion models have gained widespread adoption due to their ability to generate highly realistic images, yet their rapid proliferation also raises security and traceability concerns. To address issues of ownership verification and accountability, current watermarking techniques primarily focus on embedding information into the internal mechanisms of generative pipelines. Nevertheless, many existing methods inject watermarks directly into latent representations without adequately exploiting inherent redundancies or perceptual properties in latent space, leading to degraded image quality. In this work, we conduct a systematic analysis aimed at quantifying differentiated redundancies present within latent space, and further propose a novel Redundancy-Aware Latent Injection framework RAIN based on the above analysis. Specifically, a redundancy‑aware adaptive watermark fusion method is introduced to preserve image quality, which utilizes the differentiated redundancy distribution to guide adaptive watermark allocation in different perception tolerance regions. Moreover, a distribution alignment initialization strategy is designed to align the watermark’s initial distribution to the latent prior, reducing initialization bias and improving convergence efficiency. Comprehensive experimental evaluations demonstrate that RAIN achieves state-of-the-art performance by delivering superior perceptual quality under high-capacity watermarking scenarios. Yehan Sun, Chuangchuang Tan, Huan Liu 0001, Wenhao Ni, Renshuai Tao, Yao Zhao 0001 |
AAAI | 6 |
| 2026 | Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake DetectionabstractRecent deepfake detection studies often treat unseen sample detection as a ``zero-shot" task, training on images generated by known models but generalizing to unknown ones. A key real-world challenge arises when a model performs poorly on unknown samples, yet these samples remain available for analysis. This highlights that it should be approached as a ``few-shot" task, where effectively utilizing a small number of samples can lead to significant improvement. Unlike typical few-shot tasks focused on semantic understanding, deepfake detection prioritizes image realism, which closely mirrors real-world distributions. In this work, we propose the Few-shot Training-free Network (FTNet) for real-world few-shot deepfake detection. Simple yet effective, FTNet differs from traditional methods that rely on large-scale known data for training. Instead, FTNet uses only one fake sample from an evaluation set, mimicking the scenario where new samples emerge in the real world and can be gathered for use, without any training or parameter updates. During evaluation, each test sample is compared to the known fake and real samples, and it is classified based on the category of the nearest sample. We conduct a comprehensive analysis of AI-generated images from 29 different generative models and achieve a new SoTA performance, with an average improvement of 8.7% compared to existing methods. This work introduces a fresh perspective on real-world deepfake detection: when the model struggles to generalize on a few-shot sample, leveraging the failed samples leads to better performance. Shibo Yao, Renshuai Tao, Xiaolong Zheng 0001, Chao Liang 0001, Chunjie Zhang 0001 |
AAAI | 2 |
| 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item DetectionabstractTraining prohibited item detection models requires a large amount of X-ray security images, but collecting and annotating these images is time-consuming and laborious. To address data insufficiency, X-ray security image synthesis methods composite images to scale up datasets. However, previous methods primarily follow a two-stage pipeline, where they implement labor-intensive foreground extraction in the first stage and then composite images in the second stage. Such a pipeline introduces inevitable extra labor cost and is not efficient. In this paper, we propose a one-stage X-ray security image synthesis pipeline (Xsyn) based on text-to-image generation, which incorporates two effective strategies to improve the usability of synthetic images. The Cross-Attention Refinement (CAR) strategy leverages the cross-attention map from the diffusion model to refine the bounding box annotation. The Background Occlusion Modeling (BOM) strategy explicitly models background occlusion in the latent space to enhance imaging complexity. To the best of our knowledge, compared with previous methods, Xsyn is the first to achieve high-quality X-ray security image synthesis without extra labor cost. Experiments demonstrate that our method outperforms all previous methods with 1.2% mAP improvement, and the synthetic images generated by our method are beneficial to improve prohibited item detection performance across various X-ray security datasets and detectors. Code is available at https://github.com/pILLOW-1/Xsyn/. Jialong Sun, Hongguang Zhu, Weizhe Liu, Yunda Sun, Renshuai Tao, Yunchao Wei |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Practical and Flexible Backdoor Attack Against Deep Learning Models via Shell Code InjectionabstractRecently, backdoor attack, which aims to implant malicious logic into deep learning models (DLMs), has attracted so extensive research attention. Among them, the non-poisoning-based backdoor attack appears considerable development prospects owing to the posed threats against the DLMs-based artificial intelligence applications in cyberspace. However, previous non-poisoning- based backdoor attacks for DLMs are limited to the impractical attacking forms, resulting in certain weaknesses in both attacking complexity and attacking adaptability. To tackle the mentioned issues, this paper proposes a novel backdoor attack framework, namely the shell code injection (SCI), to perform backdoor attacks against DLMs with lower complexity and higher adaptability. Specifically, for alleviating the attacking complexity, we elaborate the logic-driven stealthy backdoor shell motivated by the biological behavior in nature,e.g., the camouflage and attack strategy of crabs. By introducing the trigger consistency verification and short-circuit code packaging strategies, the SCI misleads the victim models to output wrong predictions without training requirements according to the preset poisonous decision logic. For enhancing the attacking adaptability, we design the LLM-assisted adaptive attacking target code generation that consists of the model concept detection module and the attack target adjusting module. Since the attacking goals could be generated dynamically according to the aware victim model information and appointed attacker preset instructions, the SCI could achieve more flexible attacking performance. Extensive experiments are conducted to demonstrate that the proposed backdoor attack framework appears awesome attacking ability (almost 100% ASR) under various settings. Additionally, we provide a case study on combining the cyber attack with SCI, which also exhibits certain space for imagination of new-type backdoor attacks. Jiakai Wang, Renshuai Tao, Xianglong Liu 0001, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsabstractDiffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse. Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Renshuai Tao, Yongjun Xu 0001, Michele Magno |
AAAI | 8 |
| 2025 | Unsupervised Region-Based Image Editing of Denoising Diffusion ModelsabstractAlthough diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper, we propose a method to identify semantic attributes in the latent space of pre-trained diffusion models without any further training. By projecting the Jacobian of the targeted semantic region into a low-dimensional subspace which is orthogonal to the non-masked regions, our approach facilitates precise semantic discovery and control over local masked areas, eliminating the need for annotations. We conducted extensive experiments across multiple datasets and various architectures of diffusion models, achieving state-of-the-art performance. In particular, for some specific face attributes, the performance of our proposed method even surpasses that of supervised approaches, demonstrating its superior ability in editing local image properties. Zixiang Li, Yue Song 0002, Renshuai Tao, Xiaohong Jia 0002, Yao Zhao 0001, Wei Wang 0108 |
AAAI | 3 |
| 2025 | C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake DetectionabstractThis work focuses on AIGC detection to develop universal detectors capable of identifying various types of forgery images. Recent studies have found large pre-trained models, such as CLIP, are effective for generalizable deepfake detection along with linear classifiers. However, two critical issues remain unresolved: 1) understanding why CLIP features are effective on deepfake detection through a linear classifier; and 2) exploring the detection potential of CLIP. In this study, we delve into the underlying mechanisms of CLIP's detection capabilities by decoding its detection features into text and performing word frequency analysis. Our finding indicates that CLIP detects deepfakes by recognizing similar concepts. Building on this insight, we introduce Category Common Prompt CLIP, called C2P-CLIP, which integrates the category common prompt into the text encoder to inject category-related concepts into the image encoder, thereby enhancing detection performance. Our method achieves a 12.4% improvement in detection accuracy compared to the original CLIP. Chuangchuang Tan, Renshuai Tao, Huan Liu 0001, Guanghua Gu, Baoyuan Wu, Yao Zhao 0001, Yunchao Wei |
AAAI | 2 |
| 2025 | ODDN: Addressing Unpaired Data Challenges in Open-World Deepfake Detection on Online Social NetworksabstractDespite significant advances in deepfake detection, handling varying image quality, especially due to different compressions on online social networks (OSNs), remains challenging. Current methods succeed by leveraging correlations between paired images, whether raw or compressed. However, in open-world scenarios, paired data is scarce, with compressed images readily available but corresponding raw versions difficult to obtain. This imbalance, where unpaired data vastly outnumbers paired data, often leads to reduced detection performance, as existing methods struggle without corresponding raw images. To overcome this issue, we propose a novel approach named the open-world deepfake detection network (ODDN), which comprises two core modules: open-world data aggregation (ODA) and compression-discard gradient correction (CGC). ODA effectively aggregates correlations between compressed and raw samples through both fine-grained and coarse-grained analyses for paired and unpaired data, respectively. CGC incorporates a compression-discard gradient correction to further enhance performance across diverse compression methods in OSN. This technique optimizes the training gradient to ensure the model remains insensitive to compression variations. Extensive experiments conducted on 17 popular deepfake datasets demonstrate the superiority of the ODDN over SOTA baselines. Renshuai Tao, Manyi Le, Chuangchuang Tan, Huan Liu 0001, Haotong Qin, Yao Zhao 0001 |
AAAI | 1 |
| 2025 | Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?abstractTo detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-view imaging or insufficient sample diversity. To address these gaps, we introduce the Large-scale Dual-view X-ray (LDXray), which consists of 353,646 instances across 12 categories, providing a diverse and comprehensive resource for training and evaluating models. To emulate human intelligence in dual-view detection, we propose the Auxiliary-view Enhanced Network (AENet), a novel detection framework that leverages both the main and auxiliary views of the same object. The main-view pipeline focuses on detecting common categories, while the auxiliary-view pipeline handles more challenging categories using "expert models" learned from the main view. Extensive experiments on the LDXray dataset demonstrate that the dual-view mechanism significantly enhances detection performance, e.g., achieving improvements of up to +24.7% for the challenging category of umbrellas. Furthermore, our results show that AENet exhibits strong generalization across seven different detection models for X-ray Inspection1. Renshuai Tao, Yuzhe Guo, Hairong Chen, Li Zhang 0023, Xianglong Liu 0001, Yunchao Wei, Yao Zhao 0001 |
CVPR | 1 |
| 2025 | Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme TailoringabstractThere is a growing concern about adversarial attacks against automatic speech recognition (ASR) systems. Although research into targeted universal adversarial examples (AEs) has progressed, current methods are constrained by inefficient exploitation of audio features, demonstrating insufficient attack ability and robustness in the physical world. To solve this problem, we propose a phoneme-tailored attack (PTA) to improve the quality of the generated AEs. Specifically, to improve attack ability, we propose a Diverse Audio Composition Enrichment method, which enhances the utilization of audio features through phoneme-level slicing and recombination. To adapt AEs to complex environments, we propose a Natural Noise Pattern Guidance method to align AEs with natural noise patterns to improve their robustness. Experiments show that our method achieves an average accuracy of more than 72.34% and 98% with and without a norm constraint, and also demonstrates excellent performance in terms of generalization across datasets and resilience to MP3 compression. Yanqu Chen, Jiakai Wang, Renshuai Tao, Xianglong Liu 0001 |
ICASSP | 5 |
| 2025 | Unlocking the Potential of Lightweight Quantized Models for Deepfake DetectionabstractDeepfake detection is increasingly crucial due to the rapid rise of AI-generated content. Existing methods achieve high performance relying on computationally intensive large models, making real-time detection on resource-constrained edge devices challenging. Given that deepfake detection is a binary classification task, there is potential for model compression and acceleration. In this paper, we propose a low-bit quantization framework for lightweight and efficient deepfake detection. The Connected Quantized Block extracts common forgery features via the quantized path and retains method-specific textures through the shortcut connections. Additionally, the Shifted Logarithmic Redistribution Quantizer mitigates information loss in near-zero domains by unfolding the unbalanced activations, enabling finer quantization granularity. Comprehensive experiments demonstrate that this new framework significantly reduces 10.8x computational costs and 12.4x storage requirements while maintaining high detection performance, even surpassing SOTA methods using less than 5% FLOPs, paving the way for efficient deepfake detection in resource-limited scenarios. Renshuai Tao, Ziheng Qin, Chuangchuang Tan, Jiakai Wang, Wei Wang 0108 |
IJCAI | 1 |
| 2025 | Self-supervised Image Flicker Removal for Rolling-shutter CamerasabstractAlternating current (AC)-powered artificial lighting systems often introduce high-frequency flickering, which manifests as banding-pattern flickers in images captured by rolling-shutter cameras. Existing methods rely heavily on prior knowledge of lighting systems, synthetic data, or specialized hardware, limiting their practicality in real-world scenarios. To address these limitations, we propose a self-supervised framework for flicker removal using real-world data. Our method is grounded in a theoretical analysis demonstrating that flicker-induced luminance fluctuations follow a zero-mean distribution in the temporal domain, enabling the adoption of a self-supervised strategy. We design a modified U-Net architecture and introduce the Banding Exclusion (BE) loss to suppress residual artifacts along edges while preserving structural details. To support robust training and evaluation, we curate a comprehensive dataset comprising synthetic images generated via a physics-based flicker simulation and 160 real-world RAW sequences captured under flickering illumination. Experimental results demonstrate that our framework outperforms existing methods. Shuoxin Shan, Yakun Chang, Yujia Liu 0005, Renshuai Tao, Shikui Wei, Yao Zhao 0001 |
VCIP | 4 |
| 2025 | Dual Intention Escape: Penetrating and Toxic Jailbreak Attack against Large Language ModelsabstractRecently, the jailbreak attack, which generates adversarial prompts to bypass safety measures and mislead large language models (LLMs) to output harmful answers, has attracted extensive interest due to its potential to reveal the vulnerabilities of LLMs. However, ignoring the exploitation of the characteristics in intention understanding, existing studies could only generate prompts with weak attacking ability, failing to evade defenses (e.g., sensitive word detect) and causing malice(e.g., harmful outputs). Motivated by the mechanism in the psychology of human misjudgment, we propose a dual intention escape (DIE) jailbreak attack framework to generate more stealthy and toxic prompts to deceive LLMs to output harmful content. For stealthiness, inspired by the anchoring effect, we designed the Intention-anchored Malicious Concealment(IMC) module that hides the harmful intention behind a generated anchor intention by the recursive decomposition block and contrary intention nesting block. Since the anchor intention will be received first, the LLMs might pay less attention to the harmful intention and enter response status. For toxicity, we propose the Intention-reinforced Malicious Inducement (IMI) module based on the availability bias mechanism in a progressive malicious prompting approach. Due to the ongoing emergence of statements correlated to harmful intentions, the output content of LLMs will be closer to these more accessible intentions, i.e., more toxic. We conducted extensive experiments under black-box settings, supporting that DIE could achieve 100% ASR-R and 92.9% ASR-G against GPT3.5-turbo. Yanni Xue, Jiakai Wang, Zixin Yin, Yuqing Ma, Haotong Qin, Renshuai Tao, Xianglong Liu 0001 |
WWW | 6 |
| 2025 | LEDNet: a multimodal foundation model for robust deepfake detection
Renshuai Tao, Shijie Tang, Haotong Qin, Wei Wang 0108, Yunchao Wei, Yao Zhao 0001 |
Sci. China Inf. Sci. | 1 |
| 2025 | LS-PRISM: A layer-selective pruning method via low-rank approximation and sparsification for efficient large language model compression
Renshuai Tao, Hairong Chen, Yuzhe Guo, Jiakai Wang, Boying Wang, Yao Zhao 0001 |
Neural Networks | 1 |
| 2025 | SAGNet: Decoupling Semantic-Agnostic Artifacts From Limited Training Data for Robust Generalization in Deepfake DetectionabstractDeepfake detection presents a significant challenge, particularly when the available training data is constrained to a limited set of semantic categories—a common and realistic scenario. In deepfake detection, the training labels typically indicate whether an image is real or fake, without specifying the semantic content, such as object classes. Moreover, we cannot know in advance the object categories present in an image to be detected. Ideally, a deepfake detection model should perform consistently across different semantic categories during inference, irrespective of the content. However, existing methods often exhibit significant performance bias between seen and unseen classes, struggling to generalize effectively. To address this issue, we propose Semantic-AGnostic artifact Network (SAGNet), an innovative and efficient approach designed to decouple semantic-agnostic artifacts from content-specific distributions in the training data. Our method eliminates semantic-specific biases, ensuring that the model focuses on universal artifacts related to image authenticity rather than content-dependent features. By employing this decoupling strategy, SAGNet greatly enhances the model’s generalization capacity, even when trained on limited data. Remarkably, through experiments, we demonstrate that SAGNet achieves performance comparable to models trained with 10 times more data, despite being trained on only 2 classes (comparing SAGNet trained on 2 classes in Table I with Ojha [1] trained on 20 categories in Table IV). Furthermore, through extensive experiments, we show that SAGNet’s improvements are not only evident across different semantic categories but also extend to various generative methods, including multiple GAN-based and diffusion-based models. This cross-method generalization emphasizes SAGNet’s versatility and effectiveness in diverse generative scenarios. Overall, our method represents a significant advancement in deepfake detection, particularly in realistic situations where the training data is limited. The code is released at https://github.com/rstao-bjtu/SAGNet/. Renshuai Tao, Chuangchuang Tan, Huan Liu 0030, Jiakai Wang, Haotong Qin, Yakun Chang, Wei Wang 0108, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Exploring X-Ray Prohibited Item Detection From Long-Tailed Learning PerspectiveabstractExisting X-ray prohibited item detection methods primarily focus on boosting the detection performance of uniformly distributed items. However, in the real-world scenarios, various prohibited items exhibit the long-tailed distribution, thus posing the huge challenge to the detection task. To support this study, we introduce LTXRay, a dedicated X-ray benchmark that better assesses long-tailed prohibited item detection. LTXRay consists of 18,718 images from 12 common classes with an imbalance factor of 280.35. Meanwhile, we propose a novel Memory-Guided Learning Network(MGLNet) to develop baseline methods on LTXRay, which enhance the within-class diversity for the tail classes and consequentially improves long-tailed object detection. Specifically, we first introduce a frequency-based feature refinement module to extract discriminative contextual representations, then store the various instance features in the memory bank and dynamically generate the sample according to the historical features. Extensive experiments have been performed on the LTXRay to demonstrate the effectiveness of the proposed method. The experimental results indicate that the proposed method can consistently improve the performance of baseline methods. Boying Wang, Xiangfei Fang, Ruyi Ji, Renshuai Tao, Yaming Cao, Jing Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Reversible Column Disentangled Augmentation Tricks for Graph Contrastive Learning
Yuntai Ding, Renshuai Tao, Yifan Wang 0014, Chong Chen 0002, Xian-Sheng Hua 0001, Wei Ju 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Rethinking Depth Guided Reflection RemovalabstractWhen photographing through glass, reflections are often observed, which negatively impact the quality of the captured images or videos. In this article, we summarize and rethink depth guided reflection removal methods and, inspired by the human binocular vision system, investigate how to utilize depth for effective binocular video reflection removal. We propose an end-to-end learning-based reflection removal method that learns the transmission depth and designs a unified structure to achieve depth guided, cross-view, and cross-frame feature enhancement in a cascaded manner. Within the unified structure, different gating controllers are custom-designed to emphasize the direction of feature interaction. A dataset containing synthetic and real binocular mixture video dataset is built for network training and testing. Experimental results on both synthetic and real data from the proposed dataset demonstrate that the proposed method achieves superior performance in binocular video reflection removal. Lingzhi He, Yakun Chang, Runmin Cong, Hongyu Liu 0003, Renshuai Tao, Yao Zhao 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation
Yanni Xue, Haojie Hao, Jiakai Wang, Qiang Sheng 0001, Renshuai Tao, Pu Feng, Xianglong Liu 0001 |
IJCAI | 5 |
| 2023 | X-Adv: Physical Adversarial Object Attacks against X-ray Prohibited Item Detection
Aishan Liu, Jun Guo 0009, Jiakai Wang, Siyuan Liang 0004, Renshuai Tao, Wenbo Zhou 0004, Cong Liu 0006, Xianglong Liu 0001, Dacheng Tao |
USENIX Security Symposium | 5 |
| 2022 | Exploring Endogenous Shift for Cross-domain Detection: A Large-scale Benchmark and Perturbation Suppression NetworkabstractExisting cross-domain detection methods mostly study the domain shifts where differences between domains are often caused by external environment and perceivable for humans. However, in real-world scenarios (e.g., MRI medical diagnosis, X-ray security inspection), there still exists another type of shift, named endogenous shift, where the differences between domains are mainly caused by the intrinsic factors (e.g., imaging mechanisms, hardware components, etc.), and usually inconspicuous. This shift can also severely harm the cross-domain detection performance but has been rarely studied. To support this study, we contribute the first Endogenous Domain Shift (EDS) benchmark, X-ray security inspection, where the endogenous shifts among the domains are mainly caused by different X-ray machine types with different hardware parameters, wear degrees, etc. EDS consists of 14,219 images including 31,654 common instances from three domains (X-ray machines), with bounding-box annotations from 10 categories. To handle the endogenous shift, we further introduce the Perturbation Suppression Network (PSN), motivated by the fact that this shift is mainly caused by two types of perturbations: category-dependent and category-independent ones. PSN respectively exploits local prototype alignment and global adversarial learning mechanism to suppress these two types of perturbations. The comprehensive evaluation results show that PSN outperforms SOTA methods, serving a new perspective to the cross-domain research community. Renshuai Tao, Hainan Li, Yanlu Wei, Yifu Ding 0001, Bowei Jin, Hongping Zhi, Xianglong Liu 0001, Aishan Liu |
CVPR | 1 |
| 2022 | Defensive Patches for Robust Recognition in the Physical WorldabstractTo operate in real-world high-stakes environments, deep learning systems have to endure noises that have been con-tinuously thwarting their robustness. Data-end defense, which improves robustness by operations on input data in-stead of modifying models, has attracted intensive attention due to its feasibility in practice. However, previous data-end defenses show low generalization against diverse noises and weak transferability across multiple models. Motivated by the fact that robust recognition depends on both local and global features, we propose a defensive patch generation framework to address these problems by helping mod-els better exploit these features. For the generalization against diverse noises, we inject class-specific identifiable patterns into a confined local patch prior, so that defensive patches could preserve more recognizable features towards specific classes, leading models for better recognition under noises. For the transferability across multiple models, we guide the defensive patches to capture more global fea-ture correlations within a class, so that they could activate model-shared global perceptions and transfer better among models. Our defensive patches show great potentials to im-prove application robustness in practice by simply sticking them around target objects. Extensive experiments show that we outperform others by large margins (improve 20+ % accuracy for both adversarial and corruption robustness on average in the digital and physical world).11Our codes are available at https://github.com/nlsde-safety-team/DefensivePatch. Jiakai Wang, Zixin Yin, Aishan Liu, Renshuai Tao, Haotong Qin, Xianglong Liu 0001, Dacheng Tao |
CVPR | 5 |
| 2022 | Few-shot X-ray Prohibited Item Detection: A Benchmark and Weak-feature Enhancement NetworkabstractX-ray prohibited items detection of security inspection plays an important role in protecting public safety. It is a typical few-shot object detection (FSOD) task because some categories of prohibited items are highly scarce due to low-frequency appearance, e.g. pistols, which has been ignored by recent X-ray detection works. In contrast to most FSOD studies that rely on rich feature correlations from natural scenarios, the more practical X-ray security inspection usually faces the dilemma of only weak features learnable due to heavy occlusion, color fading, etc, which causes a severe performance drop when traditional FSOD methods are adopted. However, professional X-ray FSOD evaluation benchmarks and effective models of this scenario have been rarely studied in recent years. Therefore, in this paper, we propose the first X-ray FSOD dataset on the typical industrial X-ray security inspection scenario consisting of 12,333 images and 41,704 instances from 20 categories, which could benchmark and promote FSOD studies in such more challenging scenarios. Further, we propose the Weak-feature Enhancement Network (WEN) containing two core modules, i.e. Prototype Perception (PR) and Feature Reconciliation (FR), where PR first generates a prototype library by aggregating and extracting the basis feature from critical regions around instances, to generate the basis information for each category; FR then adaptively adjusts the impact intensity of the corresponding prototype and forces the model to precisely enhance the weak features of specific objects through the basis information. This mechanism is also effective in traditional FSOD tasks. Extensive experiments on X-ray FSOD and Pascal VOC datasets demonstrate that WEN outperforms other baselines in both X-ray and common scenarios. Renshuai Tao, Ziyang Wu, Cong Liu 0006, Aishan Liu, Xianglong Liu 0001 |
ACM Multimedia | 1 |
| 2021 | Diversifying Sample Generation for Accurate Data-Free QuantizationabstractQuantization has emerged as one of the most prevalent approaches to compress and accelerate neural networks. Recently, data-free quantization has been widely studied as a practical and promising solution. It synthesizes data for calibrating the quantized model according to the batch normalization (BN) statistics of FP32 ones and significantly relieves the heavy dependency on real training data in traditional quantization methods. Unfortunately, we find that in practice, the synthetic data identically constrained by BN statistics suffers serious homogenization at both distribution level and sample level and further causes a significant performance drop of the quantized model. We propose Diverse Sample Generation (DSG) scheme to mitigate the adverse effects caused by homogenization. Specifically, we slack the alignment of feature statistics in the BN layer to relax the constraint at the distribution level and design a layerwise enhancement to reinforce specific layers for different data samples. Our DSG scheme is versatile and even able to be applied to the state-of-the-art post-training quantization method like AdaRound. We evaluate the DSG scheme on the large-scale image classification task and consistently obtain significant improvements over various network architectures and quantization methods, especially when quantized to lower bits (e.g., up to 22% improvement on W4A4). Moreover, benefiting from the enhanced diversity, models calibrated with synthetic data perform close to those calibrated with real data and even outperform them on W4A4. Xiangguo Zhang, Haotong Qin, Yifu Ding 0001, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li 0001, Fengwei Yu, Xianglong Liu 0001 |
CVPR | 6 |
| 2021 | Towards Real-world X-ray Security Inspection: A High-Quality Benchmark And Lateral Inhibition Module For Prohibited Items DetectionabstractProhibited items detection in X-ray images often plays an important role in protecting public safety, which often deals with color-monotonous and luster-insufficient objects, resulting in unsatisfactory performance. Till now, there have been rare studies touching this topic due to the lack of specialized high-quality datasets. In this work, we first present a High-quality X-ray (HiXray) security inspection image dataset, which contains 102,928 common prohibited items of 8 categories. It is the largest dataset of high quality for prohibited items detection, gathered from the real-world airport security inspection and annotated by professional security inspectors. Besides, for accurate prohibited item detection, we further propose the Lateral Inhibition Module (LIM) inspired by the fact that humans recognize these items by ignoring irrelevant information and focusing on identifiable characteristics, especially when objects are overlapped with each other. Specifically, LIM, the elaborately designed flexible additional module, suppresses the noisy information flowing maximumly by the Bidirectional Propagation (BP) module and activates the most identifiable charismatic, boundary, from four directions by Boundary Activation (BA) module. We evaluate our method extensively on HiXray and OPIXray and the results demonstrate that it outperforms SOTA detection methods.1 Renshuai Tao, Yanlu Wei, Xiangjian Jiang, Hainan Li, Haotong Qin, Jiakai Wang, Yuqing Ma, Libo Zhang 0001, Xianglong Liu 0001 |
ICCV | 1 |
| 2021 | Multi-Pretext Attention Network For Few-Shot Learning With Self-SupervisionabstractFew-shot learning is an interesting and challenging study, which enables machines to learn from few samples like humans. Existing studies rarely exploit auxiliary information from large amount of unlabeled data. Self-supervised learning is emerged as an efficient method to utilize unlabeled data. Existing self-supervised learning methods always rely on the combination of geometric transformations for the single sample by augmentation, while seriously neglect the endogenous correlation information among different samples that is the same important for the task. In this work, we propose a Graph-driven Clustering (GC), a novel augmentation-free method for self-supervised learning, which does not rely on any auxiliary sample and utilizes the endogenous correlation information among input samples. Besides, we propose Multi-pretext Attention Network (MAN), which exploits a specific attention mechanism to combine the traditional augmentation-relied methods and our GC, adaptively learning their optimized weights to improve the performance and enabling the feature extractor to obtain more universal representations. We evaluate our MAN extensively on miniImageNet and tieredImageNet datasets and the results demonstrate that the proposed method outperforms the state-of-the-art (SOTA) relevant methods.1 Hainan Li, Renshuai Tao, Jun Li 0072, Haotong Qin, Yifu Ding 0001, Shuo Wang 0008, Xianglong Liu 0001 |
ICME | 2 |
| 2021 | Learning from Multimedia Data with Incomplete InformationabstractTraditional deep learning methods are based on the condition that the data is of high-quality, which means the data information is highly available. However, data in these scenes often have the characteristics of large background noise, lack of sample content, small target, serious occlusion and a small number of samples. The application of related tasks in real open scenarios is very important, so it is urgent to make full use of these incomplete information data accurately. Renshuai Tao |
IJCAI | 1 |
| 2021 | Sequential alignment attention model for scene text recognition
Yan Wu 0013, Renshuai Tao, Jiakai Wang, Haotong Qin, Aishan Liu, Xianglong Liu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Fast Nearest Subspace Search via Random Angular HashingabstractSubspaces frequently offer powerful representation in many tasks including recognition, retrieval, and optimization. In these tasks, the nearest subspaces (i.e., subspace-to-subspace search) often inevitably arise. Several studies in the literature have attempted to address this hard problem using techniques such as locality-sensitive hashing. Unfortunately, these subspace hashing methods are severely affected by poor scaling, with consequently high computational cost or unsatisfying accuracy, when the subspaces originally distribute with arbitrary dimensions. Accordingly, in this paper, we propose random angular hashing, a new and efficient type of locality-sensitive hashing, for linear subspaces of arbitrary dimension. The method we proposed preserves the angular distances among subspaces by randomly projecting their orthonormal basis and then encoding them with binary codes, meanwhile not only achieving fast computation but also maintaining a powerful collision probability. Moreover, its flexibility to easily get a balance between efficiency and accuracy in terms of performance. The extensive experimental results on tasks of face recognition, video de-duplication, and gesture recognition demonstrate that the proposed approach performs better than the state-of-the-art methods heavily, in terms of both accuracy and efficiency (up to 16× speedup). Yi Xu 0013, Xianglong Liu 0001, Binshuai Wang, Renshuai Tao, Ke Xia, Xianbin Cao 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Occluded Prohibited Items Detection: An X-ray Security Inspection Benchmark and De-occlusion Attention ModuleabstractSecurity inspection often deals with a piece of baggage or suitcase where objects are heavily overlapped with each other, resulting in an unsatisfactory performance for prohibited items detection in X-ray images. In the literature, there have been rare studies and datasets touching this important topic. In this work, we contribute the first high-quality object detection dataset for security inspection, named Occluded Prohibited Items X-ray (OPIXray) image benchmark. OPIXray focused on the widely-occurred prohibited item "cutter", annotated manually by professional inspectors from the international airport. The test set is further divided into three occlusion levels to better understand the performance of detectors. Furthermore, to deal with the occlusion in X-ray images detection, we propose the De-occlusion Attention Module (DOAM), a plug-and-play module that can be easily inserted into and thus promote most popular detectors. Despite the heavy occlusion in X-ray imaging, shape appearance of objects can be preserved well, and meanwhile different materials visually appear with different colors and textures. Motivated by these observations, our DOAM simultaneously leverages the different appearance information of the prohibited item to generate the attention map, which helps refine feature maps for the general detectors. We comprehensively evaluate our module on the OPIXray dataset, and demonstrate that our module can consistently improve the performance of the state-of-the-art detection methods such as SSD, FCOS, etc, and significantly outperforms several widely-used attention mechanisms. In particular, the advantages of DOAM are more significant in the scenarios with higher levels of occlusion, which demonstrates its potential application in real-world inspections. The OPIXray benchmark and our model are released at https://github.com/OPIXray-author/OPIXray. Yanlu Wei, Renshuai Tao, Zhangjie Wu, Yuqing Ma, Libo Zhang 0001, Xianglong Liu 0001 |
ACM Multimedia | 2 |