EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Li 0006
dblp:12/3242-6
· DBLP profile ↗
17ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-3873-8003ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RobNAS: Robust Neural Architecture Search for Point Cloud Adversarial DefenseabstractAs point clouds gain widespread application in fields such as autonomous driving and scene modeling, an increasing number of point cloud learning networks have emerged. As a result, research on 3D adversarial attacks and defenses has rapidly advanced. To the best of our knowledge, existing 3D defense methods primarily focus on enhancing the robustness of networks through point cloud processing or adversarial training, without attention given to network architecture. In this paper, we propose RobNAS, enhancing the adversarial robustness of point cloud classification networks by integrating adversarial training with Neural Architecture Search (NAS) from an architectural robustness perspective. Specifically, we first incorporate PGD-based adversarial training during the architecture search phase of RobNAS to obtain the most robust architecture. Subsequently, during the adversarial training phase, we introduce various adversarial examples to enhance the robustness of the model weights. Our experimental results demonstrate that our method achieves State-Of-The-Art (SOTA) performance. Furthermore, we aim to shed light on the promising potential of architectural robustness for learning robust point cloud representation. Shuoyang Sun, Hao Fang 0011, Bin Chen 0011, Jiawei Li 0006, Enze Huo, Shutao Xia |
ICASSP | 5 |
| 2025 | One Perturbation is Enough: On Generating Universal Adversarial Perturbations Against Vision-Language Pre-Training ModelsabstractVision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the multimodal alignment. However, previous studies have shown they are vulnerable to maliciously crafted adversarial samples. Despite recent success, these methods are generally instance-specific and require generating perturbations for each input sample. In this paper, we reveal that VLP models are also vulnerable to the instance-agnostic universal adversarial perturbation (UAP). Specifically, we design a novel Contrastive-training Perturbation Generator with Cross-modal conditions (C-PGC) to achieve the attack. In light that the pivotal multimodal alignment is achieved through the advanced contrastive learning technique, we devise to turn this powerful weapon against themselves, i.e., employ a malicious version of contrastive learning to train the C-PGC based on our carefully crafted positive and negative image-text pairs for essentially destroying the alignment relationship learned by VLP models. Besides, C-PGC fully utilizes the characteristics of Vision-and-Language (V+L) scenarios by incorporating both unimodal and cross-modal information as effective guidance. Extensive experiments show that C-PGC successfully forces adversarial samples to move away from their original area in the VLP model's feature space, thus essentially enhancing attacks across various victim models and V+L tasks. The GitHub repository is available at https://github.com/ffhibnese/CPGC_VLP_Universal_Attacks. Hao Fang 0011, Jiawei Kong 0001, Bin Chen 0011, Jiawei Li 0006, Shutao Xia, Ke Xu 0002 |
ICCV | 5 |
| 2024 | CAGEN: Controllable Anomaly Generator using Diffusion ModelabstractData augmentation has been widely applied in anomaly detection, which generates synthetic anomalous data for training. However, most existing anomaly augmentation methods focus on image-level cut-and-paste techniques, resulting in less realistic synthetic results, and are restricted to a few predefined patterns. In this paper, we propose our Controllable Anomaly Generator (CAGen) for anomaly data augmentation, which can generate high-quality images, and be flexibly controlled with text prompts. Specifically, our method fine-tunes a ControlNet model by using binary masks and textual prompts to control the spatial localization and style of generated anomalies. To further augment the resemblance between the generated features and normal samples, we propose a fusion method that integrates the generated anomalous features with the features of normal samples. Experiments on standard anomaly detection benchmarks show that the proposed data augmentation method significantly leads to a 0.4/3.1 improvement in the AUROC/AP metric. Bolin Jiang, Yuqiu Xie, Jiawei Li 0006, Naiqi Li, Yong Jiang 0001, Shutao Xia |
ICASSP | 3 |
| 2024 | GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process
Yuqiu Xie, Bolin Jiang, Jiawei Li 0006, Naiqi Li, Bin Chen 0011, Tao Dai 0001, Yuang Peng, Shutao Xia |
IJCAI | 3 |
| 2024 | IGSPAD: Inverting 3D Gaussian Splatting for Pose-agnostic Anomaly DetectionabstractPose-agnostic anomaly detection refers to the situation where the pose of test samples is inconsistent with the training dataset, allowing anomalies to appear at any position in any pose. We propose a novel method IGSPAD to address this challenge. Specifically, we employ 3D Gaussian splatting to represent the normal information from the training dataset. To accurately determine the pose of the test sample, we introduce an approach termed Inverting 3D Gaussian Splatting (IGS) to address the challenge of 6D pose estimation for anomalous images. The pose derived from IGS is utilized to render a normal image well-aligned with the test sample. Subsequently, the image encoder of the Segment Anything Model is employed to identify discrepancies between the rendered image and the test sample, predicting the location of anomalies. Experimental results on the MAD dataset demonstrate that the proposed method significantly surpasses the existing state-of-the-art method in terms of precision (from 97.8% to 99.7% at pixel level and from 90.9% to 98.0% at image level) and efficiency. Bolin Jiang, Yuqiu Xie, Jiawei Li 0006, Naiqi Li, Bin Chen 0011, Shutao Xia |
ACM Multimedia | 3 |
| 2023 | Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature DomainabstractBeyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed compression scenario, the existing method only implements patch matching at the image domain to solve the parallax problem caused by the difference in viewing points. However, the patch matching at the image domain is not robust to the variance of scale, shape, and illumination caused by the different viewing angles, and can not make full use of the rich texture information of the side information image. To resolve this issue, we propose Multi-Scale Feature Domain Patch Matching (MSFDPM) to fully utilizes side information at the decoder of the distributed image compression model. Specifically, MSFDPM consists of a side information feature extractor, a multi-scale feature domain patch matching module, and a multi-scale feature fusion network. Furthermore, we reuse inter-patch correlation from the shallow layer to accelerate the patch matching of the deep layer. Finally, we find that our patch matching in a multi-scale feature domain further improves compression rate by about 20% compared with the patch matching method at image domain. Yujun Huang, Bin Chen 0011, Shiyu Qin, Jiawei Li 0006, Yaowei Wang 0001, Tao Dai 0001, Shutao Xia |
AAAI | 4 |
| 2023 | Unsupervised Surface Anomaly Detection with Diffusion Probabilistic ModelabstractUnsupervised surface anomaly detection aims at discovering and localizing anomalous patterns using only anomaly-free training samples. Reconstruction-based models are among the most popular and successful methods, which rely on the assumption that anomaly regions are more difficult to reconstruct. However, there are three major challenges to the practical application of this approach: 1) the reconstruction quality needs to be further improved since it has a great impact on the final result, especially for images with structural changes; 2) it is observed that for many neural networks, the anomalies can also be well reconstructed, which severely violates the underlying assumption; 3) since reconstruction is an ill-conditioned problem, a test instance may correspond to multiple normal patterns, but most current reconstruction-based methods have ignored this critical fact. In this paper, we propose DiffAD, a method for unsupervised anomaly detection based on the latent diffusion model, inspired by its ability to generate high-quality and diverse images. We further propose noisy condition embedding and interpolated channels to address the aforementioned challenges in the general reconstruction-based pipeline. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MVTec dataset, especially in localization accuracy. Xinyi Zhang 0008, Naiqi Li, Jiawei Li 0006, Tao Dai 0001, Yong Jiang 0001, Shutao Xia |
ICCV | 3 |
| 2023 | Unsupervised Anomaly Detection with Local-Sensitive VQVAE and Global-Sensitive TransformersabstractUnsupervised anomaly detection (UAD) has been widely implemented in industrial and medical applications, which reduces the cost of manual annotation and improves efficiency in disease diagnosis. Recently, deep auto-encoder with its variants has demonstrated its advantages in many UAD scenarios. Training on the normal data, these models are expected to locate anomalies by producing higher reconstruction error for the abnormal areas than the normal ones. However, this assumption does not always hold because of the uncontrollable generalization capability. To solve this problem, we present LSGS, a method that builds on Vector Quantised-Variational Autoencoder (VQVAE) with a novel aggregated codebook and transformers with global attention. In this work, the VQVAE focus on feature extraction and reconstruction of images, and the transformers fit the manifold and locate anomalies in the latent space. Then, leveraging the generated encoding sequences that conform to a normal distribution, we can reconstruct a more accurate image for locating the anomalies. Experiments on various datasets demonstrate the effectiveness of the proposed method. Mingqing Wang, Jiawei Li 0006, Chengxiao Luo, Bin Chen 0011, Shutao Xia, Zhi Wang 0001 |
ICIP | 2 |
| 2022 | Multinomial random forest
Jiawang Bai, Yiming Li 0004, Jiawei Li 0006, Xue Yang 0003, Yong Jiang 0001, Shutao Xia |
Pattern Recognit. | 3 |
| 2021 | Attention on Attention Sparse Dense Convolutional Network for Financial Signal ProcessingabstractFinancial signal processing is a matter of great concern in FinTech. Traditionally, recurrent networks are often used to model time series, while the latest research shows that convolutional networks, especially temporal convolutional networks (TCNs), are also powerful and effective for a large number of sequence modeling tasks. The temporal convolutional network uses dilation convolution to expand the receptive field, resulting in very sparse connections in high network layers and no connection to neighbor points. Considered that short-term performance often has a more significant influence on the assets price movement, we suggest that TCNs are too sparse for financial signal processing. For a better solution, we propose a novel Attention on Attention Sparse Dense Convolutional Network (AoA-SDCN), which strengthens time decay characteristics by adding dense connections at close points. Moreover, we use the Attention on Attention mechanism to improve the performance further. Experimental results show that these techniques are effective for financial signal processing. The AoA-SDCN significantly outperforms state-of-the-art methods on Chinese commodity futures and stock datasets. Tianlei Zhu, Jiawei Li 0006, Xin-Ji Liu, Yong Jiang 0001, Shutao Xia |
ICASSP | 2 |
| 2020 | Multitask Deep Learning for Edge Intelligence Video Surveillance SystemabstractFrom the mutual empowerment of two high-speed development technologies: artificial intelligence and edge computing, we propose a tailored Edge Intelligent Video Surveillance (EIVS) system. It is a scalable edge computing architecture and uses multitask deep learning for relevant computer vision tasks. Due to the potential application of different surveillance devices are widely different, we adopt a smart IoT module to normalize the video data of different cameras, thus the EIVS system can conveniently found proper data for a specific task. In addition, the deep learning models can be deployed at every EIVS nodes, to make computer vision tasks on the normalized data. Meanwhile, due to the training and deploying of deep learning model are usually separated, for the related tasks in the same scenario, we propose to collaboratively train the depth learning models in a multitask paradigm on the cloud server. The simulation results on the publicly available datasets show that the system continuously supports intelligent monitoring tasks, has good scalability, and can improve performance through multitask learning. Jiawei Li 0006, Zhilong Zheng, Yiming Li 0004, Rubao Ma, Shutao Xia |
INDIN | 1 |
| 2020 | Multi-level Recognition on Falls from Activities of Daily LivingabstractThe falling accident is one of the largest threats to human health, which leads to broken bones, head injury, or even death. Therefore, automatic human fall recognition is vital for the Activities of Daily Living (ADL). In this paper, we try to define multi-level computer vision tasks for the visually observed fall recognition problem and study the methods and pipeline. We make frame-level labels for the fall action on several ADL datasets to test the methods and support the analysis. While current deep-learning fall recognition methods usually work on the sequence-level input, we propose a novel Dynamic Pose Motion (DPM) representation to go a step further, which can be captured by a flexible motion extraction module. Besides, a sequence-level fall recognition pipeline is proposed, which has an explicit two-branch structure for the appearance and motion feature, and has canonical LSTM to make temporal modeling and fall prediction. Finally, while current research only makes a binary classification on the fall and ADL, we further study how to detect the start time and the end time of a fall action in a video-level task. We conduct analysis experiments and ablation studies on both the simulated and real-life fall datasets. The relabelled datasets and extensive experiments form a new baseline on the recognition of falls and ADL. Jiawei Li 0006, Shutao Xia, Qianggang Ding |
ICMR | 1 |
| 2019 | Self-attentive Pyramid Network for Single Image De-raining
Taian Guo, Tao Dai 0001, Jiawei Li 0006, Shutao Xia |
ICONIP (1) | 3 |
| 2019 | Making Large Ensemble of Convolutional Neural Networks via Bootstrap Re-samplingabstractThe ensemble of Convolutional Neural Networks (CNNs) is known to be more accurate and robust than the component CNNs models. Along with the development of a fast training method, current research has managed to make an effective ensemble of several CNNs models and require no additional training cost. However, when the ensemble size of CNNs is further increased, it is hard to observe a corresponding performance enhancement. According to the generalization capability analysis of CNNs, this phenomenon can be explained by the oversaturation of model capacity and the close correlation among the component CNNs, especially when the CNNs are trained within the same dataset. To address this problem, we propose to train CNNs on re-sampled bootstrap datasets. Extensive experiments demonstrate the bootstrap re-sampling is effective for a large ensemble size (up to 80). Besides, benefiting from the usage of the bootstrap re-sampling technique, we can also have an unbiased estimate of the standard deviation of the ensemble output. Jiawei Li 0006, Xingchun Xiang, Tao Dai 0001, Shutao Xia |
VCIP | 1 |
| 2018 | Sure-Based Dual Domain Image DenoisingabstractRecently developed Dual Domain Image Denoising (DDID) algorithm is a simple version of block-matching 3D filtering (BM3D) by combining bilateral filter and frequency-based method. DDID and its invariants have achieved competitive results compared with state-of-the-art methods. However, this kind of methods share a common drawback: there are a few parameters of the algorithms that are data- and noise-dependent, and difficult to tune. In this paper, we propose to use Stein's unbiased risk estimate (SURE) to measure the mean square error (MSE) of the DDID algorithm for restoration of an image contaminated with additive white Gaussian noise. We derive an explicit expression for SURE value to optimize parameters without access to the noise-free signal. Experimental results demonstrate the effectiveness of the proposed parameter selection in term of both quantitative and qualitative metrics. Zhiya Xu, Tao Dai 0001, Li Niu 0002, Jiawei Li 0006, Qingtao Tang, Shutao Xia |
ICASSP | 4 |
| 2018 | Cyclic Annealing Training Convolutional Neural Networks for Image Classification with Noisy LabelsabstractNoisy labels modeling makes a convolutional neural network (CNN) more robust for the image classification problem. However, current noisy labels modeling methods usually require an expectation-maximization (EM) based procedure to optimize the parameters, which is computationally expensive. In this paper, we utilize a fast annealing training method to speed up the CNN training in every M-step. Since the training is repeated executed along the entire EM optimization path and obtain many local minimal CNN models from every training cycle, we name it as the Cyclic Annealing Training (CAT) approach. In addition to reducing the training time, CAT can further bagging all the local minimal CNN models at the test time to improve the performance of classification. We evaluate the proposed method on several image classification datasets with different noisy labels patterns, and the results show that our CAT approach outperforms state-of-the-art noisy labels modeling methods. Jiawei Li 0006, Tao Dai 0001, Qingtao Tang, Yeli Xing, Shutao Xia |
ICIP | 1 |
| 2018 | Portrait-Aware Artistic Style TransferabstractThe goal of artistic style transfer is to transfer the style of artistic works into photos. However, the performances of existing style transfer algorithms on portraits are not very satisfactory, because the synthetic photo is either not sufficiently stylized or distorted severely in the portrait domain (i.e., foreground), which limits the use of style transfer for portraits. In this paper, we propose a novel portrait-aware artistic style transfer algorithm, which treats foreground and background differently. Particularly, we separate the foreground from the background, and apply fine-grained style transfer to the background and coarse-grained style transfer to the entire image at the same time, so that the artistic style of entire image can be transferred with the details of the portrait well preserved. Extensive experiments demonstrate the effectiveness of our proposed method. Yeli Xing, Jiawei Li 0006, Tao Dai 0001, Qingtao Tang, Li Niu 0002, Shutao Xia |
ICIP | 2 |