EDBT 2026 Demo / reviewers in the wild / expert
Zhi Li 0080
dblp:43/3166-80
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-3729-5095ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly DetectionabstractPose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework methods require re-optimizing a new set of parameters for each query image, limiting their generalization and increasing computational burden. To overcome these limitations, we propose a novel method, Relative Pose Estimation for Pose-agnostic Anomaly Detection (RPE-PAD), which enhances both generalization and efficiency with a query-independent framework. Specifically, we propose a Random View Synthesis Scheme (RVSS) that generates new poses by adding Gaussian perturbations to the original poses, then renders the corresponding views to augment the dataset. To estimate the relative camera pose between two input images, we introduce an Iterative Relative Pose Refinement Network (IRPRN), which incorporates a hierarchical coarse-to-fine refinement strategy. Furthermore, we employ a Multi-Pair Training Strategy (MPTS) to train the proposed IRPRN, leveraging multiple image pairs to expand the relative pose transformation space during training. Extensive experiments demonstrate that our method achieves robust anomaly detection performance while significantly improving inference efficiency. Mengzan Qi, Rongkang Ma, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080 |
AAAI | 7 |
| 2025 | MPGB: Learning discriminative embeddings with multi-prototype and gradient balancing strategy for multi-modal 3D open world object detection
Liyan Ma, Zhi Li 0080, Tieyong Zeng |
Knowl. Based Syst. | 3 |
| 2025 | Auxiliary Representation Guided Network for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-identification aims to retrieve images of specific identities across modalities. To relieve the large cross-modality discrepancy, researchers introduce the auxiliary modality within the image space to assist modality-invariant representation learning. However, the challenge persists in constraining the inherent quality of generated auxiliary images, further leading to a bottleneck in retrieval performance. In this paper, we propose a novel Auxiliary Representation Guided Network (ARGN) to explore the potential of auxiliary representations, which are directly generated within the modality-shared embedding space. In contrast to the original visible and infrared representations, which contain information solely from their respective modalities, these auxiliary representations integrate cross-modality information by fusing both modalities. In our framework, we utilize these auxiliary representations as modality guidance to reduce the cross-modality discrepancy. First, we propose a High-quality Auxiliary Representation Learning (HARL) framework to generate identity-consistent auxiliary representations. The primary objective of our HARL is to ensure that auxiliary representations capture diverse modality information from both modalities while concurrently preserving identity-related discrimination. Second, guided by these auxiliary representations, we design an Auxiliary Representation Guided Constraint (ARGC) to optimize the modality-shared embedding space. By incorporating this constraint, the modality-shared embedding space is optimized to achieve enhanced intra-identity compactness and inter-identity separability, further improving the retrieval performance. In addition, to improve the robustness of our framework against the modality variation, we introduce a Part-based Adaptive Gaussian Module (PAGM) to adaptively extract discriminative information across modalities. Finally, extensive experiments are conducted to demonstrate the superiority of our method over state-of-the-art approaches on three VI-ReID datasets. Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080 |
IEEE Trans. Multim. | 6 |
| 2024 | Triple Feature Disentanglement for One-Stage Adaptive Object DetectionabstractIn recent advancements concerning Domain Adaptive Object Detection (DAOD), unsupervised domain adaptation techniques have proven instrumental. These methods enable enhanced detection capabilities within unlabeled target domains by mitigating distribution differences between source and target domains. A subset of DAOD methods employs disentangled learning to segregate Domain-Specific Representations (DSR) and Domain-Invariant Representations (DIR), with ultimate predictions relying on the latter. Current practices in disentanglement, however, often lead to DIR containing residual domain-specific information. To address this, we introduce the Multi-level Disentanglement Module (MDM) that progressively disentangles DIR, enhancing comprehensive disentanglement. Additionally, our proposed Cyclic Disentanglement Module (CDM) facilitates DSR separation. To refine the process further, we employ the Categorical Features Disentanglement Module (CFDM) to isolate DIR and DSR, coupled with category alignment across scales for improved source-target domain alignment. Given its practical suitability, our model is constructed upon the foundational framework of the Single Shot MultiBox Detector (SSD), which is a one-stage object detection approach. Experimental validation highlights the effectiveness of our method, demonstrating its state-of-the-art performance across three benchmark datasets. Haoan Wang, Shilong Jia, Tieyong Zeng, Guixu Zhang, Zhi Li 0080 |
AAAI | 5 |
| 2024 | Purified Distillation: Bridging Domain Shift and Category Gap in Incremental Object DetectionabstractIncremental Object Detection (IOD) simulates the dynamic data flow in real-world applications, which require detectors to learn new classes or adapt to new domains while retaining knowledge from previous tasks. Most existing IOD methods focus only on class incremental learning, assuming all data comes from the same domain. However, this is hardly achievable in practical applications, as images collected under different conditions often exhibit completely different characteristics, such as lighting, weather, style, etc. Class IOD methods suffer from performance degradation in these scenarios with domain shifts. To bridge domain shifts and category gaps in IOD, we propose Purified Distillation (PD), where we use a set of trainable queries to transfer the teacher's attention on old tasks to the student and adopt the gradient reversal layer to guide the student to learn the teacher's feature space structure from a micro perspective, which has not been extensively studied in previous works. Meanwhile, PD combines classification confidence with localization confidence to purify the most meaningful output nodes, so that the student model inherits a more comprehensive teacher knowledge. Extensive experiments across various IOD settings on six widely used datasets show that PD significantly outperforms state-of-the-art methods. Even after five steps of incremental learning, our method can preserve 60.6% mAP on the first task, while compared methods can only maintain up to 55.9%. Shilong Jia, Tingting Wu 0001, Yingying Fang, Tieyong Zeng, Guixu Zhang, Zhi Li 0080 |
ACM Multimedia | 6 |
| 2024 | WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution DiffusionabstractThe extraction of distribution from images with diverse weather conditions is crucial for enhancing the robustness of visual algorithms. When addressing image degradation caused by different weather, accurately perceiving the data distribution of weather-informed degradation becomes a fundamental challenge. However, given the highly stochastic nature, modelling weather distribution poses a formidable task. In this paper, we propose a novel multi-Weather distribution difFUsion blind restoration model, named WeaFU. Firstly, the model employs representation learning to map image distribution into a latent space. Subsequently, WeaFU utilizes a diffusion-based approach, with the assistance of Diffusion Distribution Generator (DDG), to perceive and extract corresponding weather distribution. This strategy ingeniously injects data distribution into the recovery process, significantly enhancing the robustness of the model in diverse weather scenarios. Finally, a Conditional Distribution-Aware Transformer (CDAT) is constructed to align the distribution information with pixels, thereby obtaining clear images. Extensive experiments on real and synthetic datasets demonstrate that WeaFU achieves superior performance. Bodong Cheng, Juncheng Li 0003, Jun Shi 0004, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-ResolutionabstractStereo image super-resolution aims to boost the performance of image super-resolution by exploiting the supplementary information provided by binocular systems. Although previous methods have achieved promising results, they did not fully utilize the information of cross-view and intra-view. To further unleash the potential of binocular images, in this letter, we propose a novel Transformer-based parallax fusion module called Parallax Fusion Transformer (PFT). PFT employs a Cross-view Fusion Transformer (CVFT) to utilize cross-view information and an Intra-view Refinement Transformer (IVRT) for intra-view feature refinement. Meanwhile, we adopted the Swin Transformer as the backbone for feature extraction and SR reconstruction to form a pure Transformer architecture called PFT-SSR. Extensive experiments and ablation studies show that PFT-SSR achieves competitive results and outperforms most SOTA methods. Source code is available at https://github.com/MIVRC/PFT-PyTorch. Hansheng Guo, Juncheng Li 0013, Guangwei Gao, Zhi Li 0080, Tieyong Zeng |
ICASSP | 4 |
| 2023 | Swin-ASNet: An Adaptive RGB-selection Network with Swin Transformer for Retinal Vessel SegmentationabstractThe retinal vasculature reflected by fundus images provides ophthalmologists with information to diagnose eye-related diseases. Therefore, the development of an accurate and automatic retinal vessel segmentation system is critical. Most deep learning-based methods directly use a color image or use a grayscale image simply transformed from the color image as input, and few models focus on the relationship between the three RGB channels. In this paper, we analyze the characteristics of the separated RGB channel images and propose an adaptive RGB-selection network with swin transformer (Swin-ASNet). The input of Swin-ASNet includes the original fundus image and three grayscale images separated from the color image. Our method can select the useful information adaptively through a designed adaptive selection aggregation module. In addition, we adopt the latest swin transformer as the backbone to extract strong features. To better fuse high and low-level features, we design a high-low interaction module, which applies a modified non-local operation under the graph convolution domain. The low-level features are injected into deep semantic information to enhance the detail representation. Experimental results show that our method can achieve state-of-the-art results in three public datasets, comparing with the existing methods. Qunchao Jin, Hongyu Hou, Guixu Zhang, Haoan Wang, Zhi Li 0080 |
ICME | 5 |
| 2023 | Fine-grained Learning for Visible-Infrared Person Re-identificationabstractVisible-Infrared Person Re-identification aims to retrieve specific identities from different modalities. In order to relieve the modality discrepancy, previous works mainly concentrate on aligning the distribution of high-level features, while disregarding the exploration of fine-grained information. In this paper, we propose a novel Fine-grained Information Exploration Network (FIENet) to implement discriminative representation, further alleviating the modality discrepancy. Firstly, we propose a Progressive Feature Aggregation Module (PFAM) to progressively aggregate mid-level features, and a Multi-Perception Interaction Module (MPIM) to achieve the interaction with diverse perceptions. Additionally, combined with PFAM and MPIM, more fine-grained information can be extracted, which is beneficial for FIENet to focus on discriminative human parts in both modalities effectively. Secondly, in terms of the feature center, we introduce an Identity-Guided Center Loss (IGCL) to supervise identity representation with intra-identity and inter-identity information. Finally, extensive experiments are conducted to demonstrate that our method achieves state-of-the-art performance. Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Zhi Li 0080 |
ICME | 5 |
| 2023 | Single image noise level estimation by artificial noise
Fang Li 0004, Faming Fang, Zhi Li 0080, Tieyong Zeng |
Signal Process. | 3 |
| 2023 | FEGNet: A Feedback Enhancement Gate Network for Automatic Polyp SegmentationabstractRegular colonoscopy is an effective way to prevent colorectal cancer by detecting colorectal polyps. Automatic polyp segmentation significantly aids clinicians in precisely locating polyp areas for further diagnosis. However, polyp segmentation is a challenge problem, since polyps appear in a variety of shapes, sizes and textures, and they tend to have ambiguous boundaries. In this paper, we propose a U-shaped model named Feedback Enhancement Gate Network (FEGNet) for accurate polyp segmentation to overcome these difficulties. Specifically, for the high-level features, we design a novel Recurrent Gate Module (RGM) based on the feedback mechanism, which can refine attention maps without any additional parameters. RGM consists of Feature Aggregation Attention Gate (FAAG) and Multi-Scale Module (MSM). FAAG can aggregate context and feedback information, and MSM is applied for capturing multi-scale information, which is critical for the segmentation task. In addition, we propose a straightforward but effective edge extraction module to detect boundaries of polyps for low-level features, which is used to guide the training of early features. In our experiments, quantitative and qualitative evaluations show that the proposed FEGNet has achieved the best results in polyp segmentation compared to other state-of-the-art models on five colonoscopy datasets. Qunchao Jin, Hongyu Hou, Guixu Zhang, Zhi Li 0080 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Graph Laplacian Regularized Spectral-Spatial-Sparse Unmixing for Hyperspectral ImageryabstractSparse unmixing aims at finding the optimal subset of endmembers in a spectral library to approximate the observed data, and has received increasing attention as it can circumvent the estimation of the endmember. In this paper, a graph Laplacian regularized spectral-spatial-sparse unmixing algorithm is proposed, namely, gLapS3U, incorporating the graph Laplacian regularization to consider the similarity between pixels of the whole image, and enforcing the spectral-spatial-sparse constraints to enhance the local spatial information as well as the sparsity of the abundance solution jointly. Experimental results on simulated and real data show the superiority of the proposed algorithm compared with state-of-the-art existing methods. Zhi Li 0080, Ruyi Feng, Yichang Shi, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng |
IGARSS | 1 |
| 2022 | CT image quality enhancement via a dual-channel neural network with jointing denoising and super-resolution
Hongyu Hou, Qunchao Jin, Guixu Zhang, Zhi Li 0080 |
Neurocomputing | 4 |
| 2022 | Quaternion-based weighted nuclear norm minimization for color image restoration
Chaoyan Huang, Zhi Li 0080, Yubing Liu, Tingting Wu 0001, Tieyong Zeng |
Pattern Recognit. | 2 |
| 2022 | Phase retrieval from incomplete data via weighted nuclear norm minimization
Zhi Li 0080, Ming Yan 0006, Tieyong Zeng, Guixu Zhang |
Pattern Recognit. | 1 |
| 2022 | Efficient Boosted DC Algorithm for Nonconvex Image Restoration with Rician NoiseabstractImage deblurring under Rician noise has attracted considerable attention in imaging science. Frequently appearing in medical imaging, Rician noise leads to an interesting nonconvex optimization problem, termed as the MAP-Rician model, which is based on the Maximum a Posteriori (MAP) estimation approach. As the MAP-Rician model is deeply rooted in Bayesian analysis, we want to understand its mathematical analysis carefully. Moreover, one needs to properly select a suitable algorithm for tackling this nonconvex problem to get the best performance. This paper investigates both issues. Indeed, we first present a theoretical result about the existence of a minimizer for the MAP-Rician model under mild conditions. Next, we aim to adopt an efficient boosted difference of convex functions algorithm (BDCA) to handle this challenging problem. Basically, BDCA combines the classical difference of convex functions algorithm (DCA) with a backtracking line search, which utilizes the point generated by DCA to define a search direction. In particular, we apply a smoothing scheme to handle the nonsmooth total variation (TV) regularization term in the discrete MAP-Rician model. Theoretically, using the Kurdyka--Lojasiewicz (KL) property, the convergence of the numerical algorithm can be guaranteed. We also prove that the sequence generated by the proposed algorithm converges to a stationary point with the objective function values decreasing monotonically. Numerical simulations are then reported to clearly illustrate that our BDCA approach outperforms some state-of-the-art methods for both medical and natural images in terms of image recovery capability and CPU-time cost. Tingting Wu 0001, Xiaoyu Gu, Zhi Li 0080, Jianwei Niu 0005, Tieyong Zeng |
SIAM J. Imaging Sci. | 4 |
| 2021 | Colour image segmentation based on a convex K-means approachabstractAbstract Image segmentation is a fundamental and challenging task in image processing and computer vision. The colour image segmentation is attracting more attention as the colour image provides more information than the grey image. A variational model based on a convex K‐means approach to segment colour images is proposed. The proposed variational method uses a combination of l 1 and l 2 regularizers to maintain edge information of objects in images while overcoming the staircase effect. Meanwhile, our one‐stage strategy is an improved version based on the smoothing and thresholding strategy, which contributes to improving the accuracy of segmentation. The proposed method performs the following steps. First, the colour set which can be determined by human or the K‐means method is specified. Second, a variational model to obtain the most appropriate colour for each pixel from the colour set via convex relaxation and lifting is used. The Chambolle–Pock algorithm and simplex projection are applied to solve the variational model effectively. Experimental results and comparison analysis demonstrate the effectiveness and robustness of the method. Tingting Wu 0001, Xiaoyu Gu, Jinbo Shao, Ruoxuan Zhou, Zhi Li 0080 |
IET Image Process. | 5 |