VLDB 2026 Research / reviewers in the wild / expert
Xiao Li 0028
dblp:66/2069-28
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-8992-4944ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Defending Against Patch-Based and Texture-Based Adversarial Attacks With Spectral DecompositionabstractAdversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch-based and texture-based attacks that are usually used to craft physical-world adversarial examples, posing real threats to security-critical applications such as person detection in surveillance and autonomous systems, because those attacks are physically realizable. Existing defense mechanisms face challenges in the adaptive attack setting, i.e., the attacks are specifically designed against them. In this paper, we propose Adversarial Spectrum Defense (ASD), a defense mechanism that leverages spectral decomposition via Discrete Wavelet Transform (DWT) to analyze adversarial patterns across multiple frequency scales. The multi-resolution and localization capability of DWT enables ASD to capture both high-frequency (fine-grained) and low-frequency (spatially pervasive) perturbations. By integrating this spectral analysis with the off-the-shelf Adversarial Training (AT) model, ASD provides a comprehensive defense strategy against both patch-based and texture-based adversarial attacks. Extensive experiments demonstrate that ASD+AT achieved state-of-the-art (SOTA) performance against various attacks, outperforming the APs of previous defense methods by 21.73%, in the face of strong adaptive adversaries specifically designed against ASD. Wei Zhang 0370, Xinyu Chang, Xiao Li 0028, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | PBCAT: Patch-Based Composite Adversarial Training Against Physically Realizable Attacks on Object DetectionabstractObject detection plays a crucial role in many security-sensitive applications. However, several recent studies have shown that object detectors can be easily fooled by physically realizable attacks, \eg, adversarial patches and recent adversarial textures, which pose realistic and urgent threats. Adversarial Training (AT) has been recognized as the most effective defense against adversarial attacks. While AT has been extensively studied in the $l_\infty$ attack settings on classification models, AT against physically realizable attacks on object detectors has received limited exploration. Early attempts are only performed to defend against adversarial patches, leaving AT against a wider range of physically realizable attacks under-explored. In this work, we consider defending against various physically realizable attacks with a unified AT method. We propose PBCAT, a novel Patch-Based Composite Adversarial Training strategy. PBCAT optimizes the model by incorporating the combination of small-area gradient-guided adversarial patches and imperceptible global adversarial perturbations covering the entire image. With these designs, PBCAT has the potential to defend against not only adversarial patches but also unseen physically realizable attacks such as adversarial textures. Extensive experiments in multiple settings demonstrated that PBCAT significantly improved robustness against various physically realizable attacks over state-of-the-art defense methods. Notably, it improved the detection accuracy by 29.7\% over previous defense methods under one recent adversarial texture attack. Xiao Li 0028, Wei Zhang 0370, Yingzhe He, Xiaolin Hu 0001 |
ICCV | 1 |
| 2025 | Efficient Neuron Segmentation in Electron Microscopy by Affinity-Guided QueriesabstractAccurate segmentation of neurons in electron microscopy (EM) images plays a crucial role in understanding the intricate wiring patterns of the brain. Existing automatic neuron segmentation methods rely on traditional clustering algorithms, where affinities are predicted first, and then watershed and post-processing algorithms are applied to yield segmentation results. Due to the nature of watershed algorithm, this paradigm has deficiency in both prediction quality and speed. Inspired by recent advances in natural image segmentation, we propose to use query-based methods to address the problem because they do not necessitate watershed algorithms. However, we find that directly applying existing query-based methods faces great challenges due to the large memory requirement of the 3D data and considerably different morphology of neurons. To tackle these challenges, we introduce affinity-guided queries and integrate them into a lightweight query-based framework. Specifically, we first predict affinities with a lightweight branch, which provides coarse neuron structure information. The affinities are then used to construct affinity-guided queries, facilitating segmentation with bottom-up cues. These queries, along with additional learnable queries, interact with the image features to directly predict the final segmentation results. Experiments on benchmark datasets demonstrated that our method achieved better results over state-of-the-art methods with a 2$\sim$3$\times$ speedup in inference. Code is available at https://github.com/chenhang98/AGQ. Hang Chen 0004, Chufeng Tang, Xiao Li 0028, Xiaolin Hu 0001 |
ICLR | 3 |
| 2025 | On the Importance of Backbone to the Adversarial Robustness of Object DetectorsabstractObject detection is a critical component of various security-sensitive applications, such as autonomous driving and video surveillance. However, existing object detectors are vulnerable to adversarial attacks, which poses a significant challenge to their reliability and security. Through experiments, first, we found that existing works on improving the adversarial robustness of object detectors give a false sense of security. Second, we found that adversarially pre-trained backbone networks were essential for enhancing the adversarial robustness of object detectors. We then proposed a simple yet effective recipe for fast adversarial fine-tuning on object detectors with adversarially pre-trained backbones. Without any modifications to the structure of object detectors, our recipe achieved significantly better adversarial robustness than previous works. Finally, we explored the potential of different modern object detector designs for improving adversarial robustness with our recipe and demonstrated interesting findings, which inspired us to design state-of-the-art (SOTA) robust detectors. Our empirical results set a new milestone for adversarially robust object detection. Code and trained checkpoints are available athttps://github.com/thu-ml/oddefense. Xiao Li 0028, Hang Chen 0004, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Language-Driven Anchors for Zero-Shot Adversarial RobustnessabstractDeep Neural Networks (DNNs) are known to be susceptible to adversarial attacks. Previous researches mainly fo-cus on improving adversarial robustness in the fully super-vised setting, leaving the challenging domain of zero-shot adversarial robustness an open question. In this work, we investigate this domain by leveraging the recent advances in large vision-language models, such as CLIP, to introduce zero-shot adversarial robustness to DNNs. We pro-pose LAAT, a Language-driven, Anchor-based Adversarial Training strategy. LAAT utilizes the features of a text en-coder for each category as fixed anchors (normalized feature embeddings) for each category, which are then employed for adversarial training. By leveraging the semantic consistency of the text encoders, LAAT aims to enhance the adversarial robustness of the image model on novel cate-gories. However, naively using text encoders leads to poor results. Through analysis, we identified the issue to be the high cosine similarity between text encoders. We then design an expansion algorithm and an alignment cross-entropy loss to alleviate the problem. Our experimental results demonstrated that LAAT significantly improves zero-shot adversarial robustness over state-of-the-art methods. LAAT has the potential to enhance adversarial robustness by large-scale multimodal models, especially when labeled data is unavailable during training. Code is available at https://github.com/LixiaoTHU/LAAT. Xiao Li 0028, Wei Zhang 0370, Zhanhao Hu, Bo Zhang 0010, Xiaolin Hu 0001 |
CVPR | 1 |
| 2024 | PartImageNet++ Dataset: Scaling Up Part-Based Models for Robust Recognition
Xiao Li 0028, Sitian Qin, Xiaolin Hu 0001 |
ECCV (71) | 1 |
| 2024 | Improve Adversarial Robustness of MNIST Classification via Topological Data Analysis
Xiao Li 0028, Sitian Qin, Xiaolin Hu 0001 |
ISNN | 2 |
| 2024 | Hiding from thermal imaging pedestrian detectors in the physical world
Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
Neurocomputing | 2 |
| 2024 | On the Privacy Effect of Data Enhancement via the Lens of MemorizationabstractMachine learning poses severe privacy concerns as it has been shown that the learned models can reveal sensitive information about their training data. Many works have investigated the effect of widely adopted data augmentation and adversarial training techniques, termed data enhancement in the paper, on the privacy leakage of machine learning models. Such privacy effects are often measured by membership inference attacks (MIAs), which aim to identify whether a particular example belongs to the training set or not. We propose to investigate privacy from a new perspective calledmemorization. Through the lens of memorization, we find that previously deployed MIAs produce misleading results as they are less likely to identify samples with higher privacy risks as members compared to samples with low privacy risks. To solve this problem, we deploy a recent attack that can capture individual samples’ memorization degrees for evaluation. Through extensive experiments, we unveil several findings about the connections between three essential properties of machine learning models, including privacy, generalization gap, and adversarial robustness. We demonstrate that the generalization gap and privacy leakage are less correlated than that of the previous results. Moreover, there is not necessarily a trade-off between adversarial robustness and privacy as stronger adversarial robustness does not make the model more susceptible to privacy attacks. Xiao Li 0028, Qiongxiu Li, Zhanhao Hu, Xiaolin Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Recognizing Object by Components With Human Prior Knowledge Enhances Adversarial Robustness of Deep Neural NetworksabstractAdversarial attacks can easily fool object recognition systems based on deep neural networks (DNNs). Although many defense methods have been proposed in recent years, most of them can still be adaptively evaded. One reason for the weak adversarial robustness may be that DNNs are only supervised by category labels and do not have part-based inductive bias like the recognition process of humans. Inspired by a well-known theory in cognitive psychology - recognition-by-components, we propose a novel object recognition model ROCK (Recognizing Object by Components with human prior Knowledge). It first segments parts of objects from images, then scores part segmentation results with predefined human prior knowledge, and finally outputs prediction based on the scores. The first stage of ROCK corresponds to the process of decomposing objects into parts in human vision. The second stage corresponds to the decision process of the human brain. ROCK shows better robustness than classical recognition models across various attack settings. These results encourage researchers to rethink the rationality of currently widely-used DNN-based object recognition models and explore the potential of part-based models, once important but recently ignored, for improving robustness. Xiao Li 0028, Ziqi Wang 0003, Bo Zhang 0010, Fuchun Sun 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Improving Image Segmentation with Boundary Patch Refinement
Xiaolin Hu 0001, Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Fooling Thermal Infrared Pedestrian Detectors in Real World Using Small BulbsabstractThermal infrared detection systems play an important role in many areas such as night security, autonomous driving, and body temperature detection. They have the unique advantages of passive imaging, temperature sensitivity and penetration. But the security of these systems themselves has not been fully explored, which poses risks in applying these systems. We propose a physical attack method with small bulbs on a board against the state of-the-art pedestrian detectors. Our goal is to make infrared pedestrian detectors unable to detect real-world pedestrians. Towards this goal, we first showed that it is possible to use two kinds of patches to attack the infrared pedestrian detector based on YOLOv3. The average precision (AP) dropped by 64.12% in the digital world, while a blank board with the same size caused the AP to drop by 29.69% only. After that, we designed and manufactured a physical board and successfully attacked YOLOv3 in the real world. In recorded videos, the physical board caused AP of the target detector to drop by 34.48%, while a blank board with the same size caused the AP to drop by 14.91% only. With the ensemble attack techniques, the designed physical board had good transferability to unseen detectors. Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
AAAI | 2 |
| 2021 | Look Closer To Segment Better: Boundary Patch Refinement for Instance SegmentationabstractTremendous efforts have been made on instance segmentation but the mask quality is still not satisfactory. The boundaries of predicted instance masks are usually imprecise due to the low spatial resolution of feature maps and the imbalance problem caused by the extremely low proportion of boundary pixels. To address these issues, we propose a conceptually simple yet effective post-processing refinement framework to improve the boundary quality based on the results of any instance segmentation model, termed BPR. Following the idea of looking closer to segment boundaries better, we extract and refine a series of small boundary patches along the predicted instance boundaries. The refinement is accomplished by a boundary patch refinement network at higher resolution. The proposed BPR framework yields significant improvements over the Mask R-CNN baseline on Cityscapes benchmark, especially on the boundary-aware metrics. Moreover, by applying the BPR framework to the "PolyTransform + SegFix" baseline, we reached 1stplace on the Cityscapes leaderboard. Code is available at https://github.com/tinyalpha/BPR. Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
CVPR | 3 |
| 2021 | Robust Logo Detection in E-Commerce Images by Data AugmentationabstractLogo detection is an important task in the intellectual property protection in e-commerce. In the paper, we introduce our solution for the ACM MM2021 Robust Logo Detection Grand Challenge. The competition requires the detection of logos (515 categories) in e-commerce images. This competition is challenged by long-tail distribution, small objects, and different types of noises. To overcome these challenges, we built a highly optimized and robust detector. We first tested many effective techniques for general object detection and then focused on data augmentation. We found that data augmentation was effective in improving the performance and robustness of logo detection. Based on the combination of these techniques, we achieved APs of 64.6% and 61.3% on the clean and noisy datasets respectively, which were improved by 8.1% and 19.5% relative to the official baseline. We ranked 5th among 36489 teams in the competition. Hang Chen 0004, Xiao Li 0028, Zefan Wang, Xiaolin Hu 0001 |
ACM Multimedia | 2 |