EDBT 2026 Demo / reviewers in the wild / expert
Weizhi Xian
dblp:277/6450
· DBLP profile ↗
19ranked-venue papers
4as first author
18since 2021 · last 2027
0000-0001-5137-3542ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Semantic-guided multi-feature fusion for underwater image quality assessment
Huayan Pu, Jun Luo 0006, Jielu Yan, Weizhi Xian, Xuekai Wei, Mingliang Zhou 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Image quality assessment: Unifying spatial and frequency distribution discrepancy in deep feature domains via Rényi divergence
Dongzi Wang 0003, Weizhi Xian, Jielu Yan, Xuekai Wei, Mingliang Zhou 0001, Sam Kwong |
Expert Syst. Appl. | 2 |
| 2026 | SAM-IAD: Injecting specific knowledge into SAM for industrial anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian, Xinyi Gong, Jianwen Han, Xian Tao |
Knowl. Based Syst. | 3 |
| 2026 | Neighborhood Attention-based Feature Reconstruction for Image Anomaly Detection and LocalizationabstractWith the advancement of machine vision technology, automated vision inspection systems are needed in broad quality control scenarios. This article proposes a neighborhood attention-based feature reconstruction method for image anomaly detection and localization (NAFRAD). To address the challenges of data scarcity, low visibility, and irregular defect shapes in unsupervised anomaly detection, we introduce a feature reconstruction framework that preserves high-level abstract features rather than focusing on pixel-level reconstruction. This approach enhances model robustness and generalizability by leveraging neighborhood attention (NA) mechanisms, which simultaneously capture local details and the global context through a sliding window strategy. The NA-based autoencoder reconstructs normal features by aggregating local inductive biases with translational equivariance, enabling precise anomaly localization. Extensive experiments on the MVTec Anomaly Detection (MVTec AD) dataset—comprising 15 categories with 5,354 images—demonstrate the superiority of NAFRAD. It achieves state-of-the-art performance with AUROC \({}_{I}\) = 99.02, AUROC \({}_{P}\) = 98.99, and AP = 79.40, outperforming existing methods by 3.6% in AP and 0.89% in AUROC \({}_{P}\) . The framework’s effectiveness is validated through ablation studies, visualization of feature reconstruction, and comparisons with eight leading unsupervised methods. The code is made public at https://github.com/Math-Computer/NAFRAD . Weizhi Xian, Yichi Chen 0002, Bin Chen 0022, Leong Hou U, Shiyou Liu, Yong Feng 0002, Mingliang Zhou 0001, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects SupervisionabstractOpen-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the background. To resolve this, we propose OV-DQUO, an Open-Vocabulary DETR with Denoising text Query training and open-world Unknown Objects supervision. Specifically, we introduce a wildcard matching method. This method enables the detector to learn from pairs of unknown objects recognized by the open-world detector and text embeddings with general semantics, mitigating the confidence bias between base and novel categories. Additionally, we propose a denoising text query training strategy. It synthesizes foreground and background query-box pairs from open-world unknown objects to train the detector through contrastive learning, enhancing its ability to distinguish novel objects from the background. We conducted extensive experiments on the OV-COCO and OV-LVIS benchmarks, achieving new state-of-the-art results of 45.6 AP50 and 39.3 mAP on novel categories, respectively. Bin Chen 0022, Bin Kang, Yulin Li 0003, Weizhi Xian, Yichi Chen 0002 |
AAAI | 5 |
| 2025 | A Dynamic Learning Strategy for Dempster-Shafer Theory with Applications in Classification and EnhancementabstractEffective modelling of uncertain information is crucial for quantifying uncertainty. Dempster–Shafer evidence (DSE) theory is a widely recognized approach for handling uncertain information. However, current methods often neglect the inherent a priori information within data during modelling, and imbalanced data lead to insufficient attention to key information in the model. To address these limitations, this paper presents a dynamic learning strategy based on nonuniform splitting mechanism and Hilbert space mapping. First, the framework uses a nonuniform splitting mechanism to dynamically adjust the weights of data subsets and combines the diffusion factor to effectively incorporate the data a priori information, thereby flexibly addressing uncertainty and conflict. Second, the conflict in the information fusion process is reduced by Hilbert space mapping. Experimental results on multiple tasks show that the proposed method significantly outperforms state-of-the-art methods and effectively improves the performance of classification and low-light image enhancement (LLIE) tasks. The code is available at https://anonymous.4open.science/r/Third-ED16. Mingliang Zhou 0001, Xuekai Wei, Weizhi Xian, Jielu Yan, Weijia Jia 0001 |
NeurIPS | 5 |
| 2025 | Continuous reinforcement learning via advantage value difference reward shaping: A proximal policy optimization perspective
Xuekai Wei, Weizhi Xian, Jielu Yan, Leong Hou U, Yong Feng 0002, Zhaowei Shang, Mingliang Zhou 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | No-Reference Image Quality Assessment: Exploring Intrinsic Distortion Characteristics via Generative Noise Estimation With MambaabstractIn the field of no-reference image quality assessment (NR-IQA), the visual masking effect has long been a challenging issue. Although existing methods attempt to alleviate the interference caused by masking by generating pseudoreference images, the quality of these images is often constrained by the accuracy and reconstruction capabilities of image restoration algorithms. This can introduce additional biases, thereby affecting the reliability of the evaluation results. To address this problem, we propose a novel generative “noise” estimation framework (GNE-Vim) that eliminates the need for pseudoreference images. Instead, it deeply decouples the distortion components from degraded images and performs quality-aware modelling of these components. During the training phase, the model leverages both reference images and distortion components to guide the learning of the true distortion distribution. In the inference phase, quality prediction is conducted directly on the basis of the decoupled distortion components, making the evaluation results more aligned with human subjective perception. The experimental results demonstrate that the proposed method achieves strong performance across datasets containing various types of distortions. The source code is publicly available at the following website: https://github.com/opencodelxt/GNE-Vim. Xuting Lan, Weizhi Xian, Mingliang Zhou 0001, Jielu Yan, Xuekai Wei, Jun Luo 0006, Weijia Jia 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DTSD: A Dual Teacher-Student-Based Discrimination Model for Anomaly DetectionabstractThe rapid development of computer vision technology for detecting anomalies in industrial products has received unprecedented attention. In this article, we propose a dual teacher–student-based discrimination (DTSD) model for anomaly detection, which combines the advantages of both embedding-based and reconstruction-based methods. First, the DTSD builds a dual teacher‒student architecture consisting of a pretrained teacher encoder with frozen parameters, a student encoder, and a student decoder. By distillation of knowledge from the teacher encoder, the two teacher‒student modules acquire the ability to capture both local and global anomaly patterns. Second, to address the issue of poor reconstruction quality faced by previous reconstruction-based approaches in some challenging cases, the model employs a feature bank that stores encoded features of normal samples. By incorporating template features from the feature bank, the student decoder receives explicit guidance to enhance the quality of reconstruction. Finally, a segmentation network is utilized to adaptively integrate multiscale anomaly information from the two teacher–student modules, thereby improving segmentation accuracy. Extensive experiments demonstrate that our method outperforms existing state-of-the-art approaches. The code of DTSD is publicly available at https://github.com/Math-Computer/DTSD . Weizhi Xian, Xuekai Wei, Jielu Yan, Yueting Huang, Kunyin Guo, Weijia Jia 0001, Mingliang Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Multi-Attribute Consistency Driven Visual Language Framework for Surface Defect DetectionabstractVisual Language Pre-training models encounter significant challenges stemming from the scarcity of data and the presence of ambiguous cues in industrial defect detection tasks. In this work, we propose a multi-attribute consistency-driven defect detection (MACD) framework to optimize text prompts in a coarse-to-fine trajectory. To bridge differences in domain knowledge, we build a structured attribute repository that contains descriptions of various defects’ inherent attributes. Based on this, we propose a multi-attribute consistency (MAC) module that can adequately model the global alignment between sentences with multiple attributes and defect images. Furthermore, we design a refined cross-alignment (RCA) module to determine the fine-grained correspondence between each attribute and the region within the image. Finally, the proposed method is experimentally validated on two benchmarks, resulting in significant performance improvements in a wide range of defective scenarios. Bin Kang, Bin Chen 0022, Weizhi Xian, Huifeng Chang |
ICME | 4 |
| 2024 | A Rate Control Scheme for VVC Intercoding Using a Linear ModelabstractVersatile video coding (VVC) aims to achieve high compression but also issues like varying content/network conditions. Existing rate control (RC) methods struggle to achieve optimal quality under these complex scenarios. This paper proposes a novel RC scheme for VVC based on a linear model. The Lagrange minimization multiplier is introduced under bit budget constraints, allowing optimized bit allocation. RC optimization is formulated as a convex solution, and is derived into the optimal quantization parameter (QP) for RC. Experimental analysis demonstrates the proposed linear model-based RC algorithm performances are better compared to other state-of-the-art methods due to their use of a linear model and optimal QP determination. Heqiang Wang, Xuekai Wei, Weizhi Xian, Jun Luo 0006, Huayan Pu, Zhigang Chu, Xin Wang 0051, Xueyong Xu, Chang Lu 0005, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2024 | Saliency and Depth-Aware Full Reference 360-Degree Image Quality AssessmentabstractWith the widespread adoption of virtual reality and 360-degree video, there is a pressing need for objective metrics to assess quality in this immersive panoramic format reliably. However, existing image quality assessment models developed for traditional fixed-viewpoint content do not fully consider the specific perceptual issues involved in 360-degree viewing. This paper proposes a 360-degree image full-reference quality assessment (FR-IQA) methodology based on a multi-channel architecture. The proposed 360-degree FR-IQA method further optimizes and identifies the distorted image quality using two easily obtained useful saliency and depth-aware image features. The convolutional neural network (CNN) is designed for training. Furthermore, the proposed method accounts for predicting user viewing behaviors within 360-degree images, which will further benefit the multi-channel CNN architecture and enable the weighted average pooling of the predicted FR-IQA scores. The performance is evaluated on publicly available databases to demonstrate the advantages brought by the proposed multi-channel model in performance evaluation and cross-database evaluation experiments, where it outperforms other state-of-the-art ones. Moreover, an ablation study exhibits good generalization ability and robustness. Xuekai Wei, Qunyue Huang, Bin Fang 0001, Lei Ouyang, Weizhi Xian, Jun Luo 0003, Huayan Pu, Xueyong Xu, Chang Lu 0005, Hao Nan, Xu Liu 0006, Yachao Li 0001, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2024 | An End-to-End Video Coding Method via Adaptive Vision TransformerabstractDeep learning-based video coding methods have demonstrated superior performance compared to classical video coding standards in recent years. The vast majority of the existing deep video coding (DVC) networks are based on convolutional neural networks (CNNs), and their main drawback is that since CNNs are affected by the size of the receptive field, they cannot effectively handle long-range dependencies and local detail recovery. Therefore, how to better capture and process the overall structure as well as local texture information in the video coding task is the core issue. Notably, the transformer employs a self-attention mechanism that captures dependencies between any two positions in the input sequence without being constrained by distance limitations. This is an effective solution to the problem described above. In this paper, we propose end-to-end transformer-based adaptive video coding (TAVC). First, we compress the motion vector and residuals through a compression network built on the vision transformer (ViT) and design the motion compensation network based on ViT. Second, based on the requirement of video coding to adapt to different resolution inputs, we introduce a position encoding generator (PEG) as adaptive position encoding (APE) to maintain its translation invariance across different resolution video coding tasks. The experiment shows that for multiscale structural similarity index measurement (MS-SSIM) metrics, this method exhibits significant performance gaps compared to conventional engineering codecs, such as [Formula: see text], [Formula: see text], and VTM-15.2. We also achieved a good performance improvement compared to the CNN-based DVC methods. In the case of peak signal-to-noise ratio (PSNR) evaluation metrics, TAVC also achieves good performance. Mingliang Zhou 0001, Zhaowei Shang, Huayan Pu, Jun Luo 0006, Xiaoxu Huang, Shilong Wang 0001, Huajun Cao, Xuekai Wei, Weizhi Xian |
Int. J. Pattern Recognit. Artif. Intell. | 10 |
| 2024 | Deep Dual-Stream Convolutional Neural Networks for Cardiac Image Semantic SegmentationabstractCardiac image segmentation is essential when applying biomedical informatics to improve industrial healthcare applications. To extract context and detailed information more efficiently and further improve cardiac image segmentation accuracy, we present a novel deep dual-stream convolutional neural network (CNN) for cardiac image semantic segmentation in this article. We use a body stream and a shape stream, respectively, in this method. First, in the body stream we propose integrating a gated fully fusion module to fuse multilevel features in the encoder and decoder paths. In addition, we integrate a feature aggregation module to extract the multiscale context. Second, in the shape stream, we propose using a gated shape CNN exploiting multilevel context to extract detailed information, such as boundary and shape features. Finally, we apply a multitask loss function to align the predicted masks with the ground truth labels. Our experiments on the public cardiac magnetic resonance image dataset show significant performance in the left and right ventricular cavities and myocardium compared to the state-of-the-art algorithms. Hengqi Hu, Bin Fang 0001, Yuting Ran, Xuekai Wei, Weizhi Xian, Mingliang Zhou 0001, Sam Kwong |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Perceptual Quality Analysis in Deep Domains Using Structure Separation and High-Order MomentsabstractImages are composed of “things” (i.e., structured objects) and “stuff” (i.e., textured surfaces), which have completely different effects on the human visual system (HVS). A good image quality assessment (IQA) method should fully consider the visual salience effects of image structures and the masking effects of image textures. In this article, we propose a perceptual quality analysis model using structure separation and high-order moments (SSHMPQA) in the deep domain. First, we use a total variation (TV) model to separate the perceptual structures in images from their deep feature maps, thereby maintaining meaningful object shapes with texture suppression and defining perceptual structure-aware distances in the deep domain. Then, we use the first- to fourth-order moments to calculate the mean, skewness and kurtosis of the probability distributions of the deep features. On this basis, we define a perceptual texture-aware distance in the deep domain. We then formulate the final model by solving a well-defined perceptual optimization problem. The proposed SSHMPQA model has good interpretability and is data-driven; moreover, the model does not require a complex and long training process because the optimization problem is convex and has an exact analytical solution. To verify the effectiveness of our model, comprehensive experiments are conducted. The experimental results show that the proposed model is superior to other state-of-the-art traditional and deep learning-based full-reference (FR) IQA methods. Weizhi Xian, Mingliang Zhou 0001, Bin Fang 0001, Tao Xiang 0001, Weijia Jia 0001, Bin Chen 0022 |
IEEE Trans. Multim. | 1 |
| 2024 | LGFDR: local and global feature denoising reconstruction for unsupervised anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian |
Vis. Comput. | 3 |
| 2022 | A content-oriented no-reference perceptual video quality assessment method for computer graphics animation videos
Weizhi Xian, Mingliang Zhou 0001, Bin Fang 0001, Sam Kwong |
Inf. Sci. | 1 |
| 2022 | An active contour model based on adaptively variable exponent combining Legendre polynomial for image segmentation
Jiajie Zhu 0001, Bin Fang 0001, Mingliang Zhou 0001, Futing Luo, Weizhi Xian, Gang Wang 0023 |
Multim. Tools Appl. | 5 |
| 2020 | Segmentation Algorithm of the Valid Region in Fisheye Images Using Edge and Region InformationabstractIn this paper, we propose a method to segment the valid region of fisheye images. First, we construct an objective function with three terms, which are the region driving term, the edge driving term and the length regularization term. Second, we minimize this objective function by a modified gradient descent method to find the best segmentation result. Our method can achieve valid region segmentation by making use of both region information and edge information. Experiments show that the proposed method can deal with blurred edges, halation noise and incomplete valid region problems. Tongxin Du, Bin Fang 0001, Mingliang Zhou 0001, Henjun Zhao, Weizhi Xian, Xuegang Wu |
ICIP | 5 |