VLDB 2026 Research / reviewers in the wild / expert
Zhishan Li
dblp:237/5725
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0003-2211-735XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decoupling While Coupling: Towards More Accurate Stereo Image Sand Removal Beyond CertaintyabstractStereo image sand removal is crucial to improve the perceptual quality for autonomous driving perception. Existing methods often fall short in accurately estimating the uncertainty inherent in degraded images, leading to suboptimal outcomes. To address this, we introduce a novel framework named Decoupling While Coupling(DWC). DWC pioneers the integration of inter-view uncertainty estimation, cross-view uncertainty-aware interaction and block-wise uncertainty representation for superior stereo image sand removal. For cross-view information interaction, we propose an Uncertainty-aware Cross-view Attentive Interaction module(UCAI) to cope with the lack of uncertainty estimation ability in the existing cross-view information interaction mechanism. For the uncertainty perception and information interaction within the inter-view, we propose a Distribution Modeling Coupling Block(DMCB), which transmits the representation of uncertainty between each backbone module. For block-wise uncertainty estimation, we use our proposed Uncertainty-aware Distribution Feature Modulator(UDFM) as the backbone of DWC to modulate the uncertainty inside the neural network itself. Extensive experimental validations on our proposed stereo image sand removal dataset SandST confirm the efficacy of DWC. Our method not only achieves higher PSNR and SSIM, but also exhibits enhanced robustness against various sand degrees and patterns. Bingcai Wei, Hui Liu 0065, Chuang Qian 0001, Yi Jia, Wangyu Wu, Zhishan Li |
ICASSP | 6 |
| 2025 | Optimal Control for Constrained Discrete-Time Nonlinear Systems Based on Safe Reinforcement LearningabstractThe state and input constraints of nonlinear systems could greatly impede the realization of their optimal control when using reinforcement learning (RL)-based approaches since the commonly used quadratic utility functions cannot meet the requirements of solving constrained optimization problems. This article develops a novel optimal control approach for constrained discrete-time (DT) nonlinear systems based on safe RL. Specifically, a barrier function (BF) is introduced and incorporated with the value function to help transform a constrained optimization problem into an unconstrained one. Meanwhile, the minimum of such an optimization problem can be guaranteed to occur at the origin. Then a constrained policy iteration (PI) algorithm is developed to realize the optimal control of the nonlinear system and to enable the state and input constraints to be satisfied. The constrained optimal control policy and its corresponding value function are derived through the implementation of two neural networks (NNs). Performance analysis shows that the proposed control approach still retains the convergence and optimality properties of the traditional PI algorithm. Simulation results of three examples reveal its effectiveness. Lingzhi Zhang, Lei Xie 0007, Yi Jiang 0007, Zhishan Li, Xueqin Amy Liu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Diffusion-Based Framework for Multi-Class Anomaly DetectionabstractReconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods might face challenges related to the preservation of image categories and pixel-wise structural integrity in the more practical multi-class setting. To solve the above problems, we propose a Difusion-based Anomaly Detection (DiAD) framework for multi-class anomaly detection, which consists of a pixel-space autoencoder, a latent-space Semantic-Guided (SG) network with a connection to the stable diffusion’s denoising network, and a feature-space pre-trained feature extractor. Firstly, The SG network is proposed for reconstructing anomalous regions while preserving the original image’s semantic information. Secondly, we introduce Spatial-aware Feature Fusion (SFF) block to maximize reconstruction accuracy when dealing with extensively reconstructed areas. Thirdly, the input and reconstructed images are processed by a pre-trained feature extractor to generate anomaly maps based on features extracted at different scales. Experiments on MVTec-AD and VisA datasets demonstrate the effectiveness of our approach which surpasses the state-of-the-art methods, e.g., achieving 96.8/52.6 and 97.2/99.0 (AUROC/AP) for localization and detection respectively on multi-class MVTec-AD dataset. Code will be available at https://lewandofskee.github.io/projects/diad. Haoyang He, Jiangning Zhang, Xuhai Chen, Zhishan Li, Xu Chen 0024, Yabiao Wang, Chengjie Wang 0001, Lei Xie 0007 |
AAAI | 5 |
| 2024 | Towards efficient filter pruning via adaptive automatic structure search
Xiaozhou Xu, Jun Chen 0023, Zhishan Li, Lei Xie 0007 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Adaptive multi-scale TF-net for high-resolution time-frequency representations
Tao Chen 0053, Zhishan Li, Lei Xie 0007 |
Signal Process. | 4 |
| 2024 | Toward Effective Traffic Sign Detection via Two-Stage Fusion Neural NetworksabstractAutomatic detection of traffic signs is crucial for Advanced Driving Assistance Systems (ADAS). Current two-stage approaches consist of a preliminary object detection step, where the traffic signs are categorized within broader families (e.g., speed limits), and then sub-classes (e.g., speed limit 40). However, these cascading methods fail to achieve satisfying performance, especially in more realistic driving scenarios where images are acquired under more challenging conditions. Under such conditions, the first-stage detection step is likely to provide inaccurate predictions, making the subsequent classification step useless. In this paper, we propose a simple yet effective two-stage fusion framework for traffic sign detection. Different from the previous cascading method, our framework directly predicts categories in the first-stage detection and fuse the two-stage category predictions to improves overall robustness. Besides, in order to filter the false detection boxes under low-resolution inputs, we also propose an effective post-processing method called Surrounding-Aware Non-Maximum Suppression (SA-NMS) as an alternative technique for the first-stage detection. After combining the above proposed methods, our framework obtains good detection performance. Experimental results on the widely used Tsinghua-Tencent 100K (TT100K) traffic sign dataset, which contains images of traffic signs collected under a variety of challenging conditions, show that the proposed framework outperforms current approaches in both accuracy and inference speed, achieving 89.7 mAP and 65 FPS for${608\times608}$low resolution images. Zhishan Li, Battista Biggio, Yifan He 0002, Haoran Cai, Fabio Roli, Lei Xie 0007 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Lightweight Cryptography Implementation for Internet of Things Network on FPGAabstractWith the development of modern communication technology, traditional Internet of Things systems could not provide sufficient support for large data flow, hardware resource, power consumption, and security during transmission. Especially when it comes to the Industrial Internet of Things (IIoT), where the limitations of hardware resource and power consumption are more strict; the requirements of network security and hardware security are much higher. This paper implements a User Datagram Protocol (UDP) communication system, which is encrypted with one of the Lightweight Cryptography, Xoodyak. We aim to design a lightweight encrypted communication platform. Compared with a public key encryption system, our implementation costs much less hardware resource and power consumption, while providing excellent side-channel attack (SCA) protection. Besides, when there is a large burst length of data flow, two asynchronous FIFO can restore those data respectively, so our design can maintain throughput and data integrity in extreme cases. Based on the above properties, our design is relatively ideal in IIoT encryption scenarios. Guanzhong Tian, Longhua Ma, Zhishan Li, Shanqi Liu |
SMC | 4 |
| 2023 | Towards accurate dense pedestrian detection via occlusion-prediction aware label assignment and hierarchical-NMS
Haoyang He, Zhishan Li, Guanzhong Tian, Lei Xie 0007, Shan Lu 0009 |
Pattern Recognit. Lett. | 2 |
| 2022 | CICC: Channel Pruning via the Concentration of Information and Contributions of Channels
Zhishan Li, Yingqing Yang, Lei Xie 0007, Yong Liu 0007, Longhua Ma, Shanqi Liu, Guanzhong Tian |
BMVC | 2 |
| 2022 | An Efficient Framework for Detection and Recognition of Numerical Traffic SignsabstractDue to the variety of categories and uneven distribution of available samples, automatic traffic sign detection and recognition is still a challenging task. For those categories with less training data, existing deep learning methods cannot achieve desirable performance, and the overall detection effect is not satisfactory as well. In this letter, we fully explore the relationship between different traffic signs with digital characters and transform the category objects into multi-level classes to alleviate the uneven distribution of samples. We design a lightweight two-stage object detection framework with high real-time performance. The first stage network is proposed to obtain the category groups of traffic signs, and then we construct another object detection network to identify the digital characters of the detected traffic signs. To make the prediction in the first stage more accurate, we put forward a boxes fusion algorithm in the post-processing process and a refine module to improve the recognition performance. Experimental results show that our approach possesses significantly improved performance compared with the latest object detection networks and other traffic sign detectors. Even some traffic signs that only exist in testset can also be recognized accurately by our method. Zhishan Li, Mingmu Chen, Yifan He 0002, Lei Xie 0007 |
ICASSP | 1 |
| 2022 | A Transformer-Based Object Detector with Coarse-Fine Crossing RepresentationsabstractTransformer-based object detectors have shown competitive performance recently. Compared with convolutional neural networks limited by the relatively small receptive fields, the advantage of transformer for visual tasks is the capacity to perceive long-range dependencies among all image patches, while the deficiency is that the local fine-grained information is not fully excavated. In this paper, we introduce the Coarse-grained and Fine-grained crossing representations to build an efficient Detection Transformer (CFDT). Specifically, we propose a local-global cross fusion module to establish the connection between local fine-grained features and global coarse-grained features. Besides, we propose a coarse-fine aware neck which enables detection tokens to interact with both coarse-grained and fine-grained features. Furthermore, an efficient feature integration module is presented for fusing multi-scale representations from different stages. Experimental results on the COCO dataset demonstrate the effectiveness of the proposed method. For instance, our CFDT achieves 48.1 AP with 173G FLOPs, which possesses higher accuracy and less computation compared with the state-of-the-art transformer-based detector ViDT. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/CFDT. Zhishan Li, Kai Han 0002, Jianyuan Guo, Yunhe Wang 0001 |
NeurIPS | 1 |
| 2022 | PBDE: an effective post-processing method based on box density for object detection
Zhishan Li, Baozhi Jia, Yifan He 0002, Lei Xie 0007 |
Appl. Intell. | 1 |
| 2021 | An Effective Face Anti-Spoofing Method via Stereo MatchingabstractVarious algorithms based on Convolutional Neural Network (CNN) have achieved great performance in the task of face anti-spoofing (FAS). However, the issue of most approaches is that the robustness in unknown scenes is not strong due to the different quality of attack images and environmental factors. In this letter, we propose a real-time face anti-spoofing method based on stereo matching. We input the left and right views of a pair of infrared face images into our proposed lightweight stereo matching network to get a disparity map. Then, we use the disparity map as input for the classification network to obtain face living information. Experimental results show that the proposed approach has significantly improved performance compared with the latest state-of-the-art face anti-spoofing methods. Considering the real-time requirement, our method is superior to most single-image-based models in inference time and far less than those based on other large scale stereo matching networks in computational complexity. The code is available athttps://github.com/lizhishan1997/StereoMatching_FAS. Zhishan Li, Jiayan Yuan, Baozhi Jia, Yifan He 0002, Lei Xie 0007 |
IEEE Signal Process. Lett. | 1 |
| 2020 | PBDE: An Effective Method for Filtering False Positive Boxes in Object DetectionabstractAn inevitable problem in the practice of object detection is the existence of false positive detection boxes. False detection boxes greatly reduce precision of object detection model and compromise the desired effect. In this paper, we propose a method named Prediction Box Density Evalution (PBDE). We summarize box density characteristics of true positive (TP) and false positive (FP) boxes to filter out a large number of FP boxes. After PBDE, we can obtain a significant improvement in precision with a low recall loss, and increase F1-score by up to 9 percentage when the confidence threshold is 0.1. The entire algorithm is carried out on the post-processing of object detection, and there is no need to change the original training method and network structure, which is of great practical importance. Zhishan Li, Baozhi Jia, Mingmu Chen, Shaokai Xu |
ICARCV | 1 |