Xudong Xie

dblp:67/1248 · DBLP profile ↗
← Back
37ranked-venue papers
15as first author
12since 2021 · last 2026
0000-0002-5932-1496ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 12 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GGEF: A Framework for Automatic Extraction of Gene-Drug Relationships from Cancer Guidelines
Liuzhen Su, Xudong Xie, Weilan Qian, Mincheng Li, Shunzheng Ma, Minhua Shao
DASFAA (6)2
2026 GeneDrug: A Retrieval Platform for Analyzing Gene-Drug Relations in Cancer
Liuzhen Su, Xudong Xie, Weilan Qian, Minhua Shao, Yungang He
DASFAA (6)2
2026 PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Xudong Xie, Minghui Liao, Wei Chen 0088, Xiang Bai
Int. J. Comput. Vis.1
2025 SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
abstract
Most previous scene text spotting methods rely on high-quality manual annotations to achieve promising performance. To reduce their expensive costs, we study semi-supervised text spotting (SSTS) to exploit useful information from unlabeled images. However, directly applying existing semi-supervised methods of general scenes to SSTS will face new challenges: 1) inconsistent pseudo labels between detection and recognition tasks, and 2) sub-optimal supervisions caused by inconsistency between teacher/student. Thus, we propose a new Semi-supervised framework for End-to-end Text Spotting, namely SemiETS that leverages the complementarity of text detection and recognition. Specifically, it gradually generates reliable hierarchical pseudo labels for each task, thereby reducing noisy labels. Meanwhile, it extracts important information in locations and transcriptions from bidirectional flows to improve consistency. Extensive experiments on three datasets under various settings demonstrate the effectiveness of SemiETS on arbitrary-shaped text. For example, it outperforms previous state-of-the-art SSL methods by a large margin on end-to-end spotting (+8.7%, +5.6%, and +2.6% H-mean under 0.5%, 1%, and 2% labeled data settings on Total-Text, respectively). More importantly, it still improves upon a strongly supervised text spotter trained with plenty of labeled data by 2.0%. Compelling domain adaptation ability shows practical potential. Moreover, our method demonstrates consistent improvement on different text spotters. Code will be available at https://github.com/DrLuo/SemiETS.
Dongliang Luo, Hanshen Zhu, Dingkang Liang, Xudong Xie, Xiang Bai
CVPR5
2025 Multi-Scenario Overlapping Text Segmentation with Depth Awareness
Xudong Xie, Xiang Bai
ICCV2
2025 MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
abstract
Scene text retrieval has made significant progress with the assistance of accurate text localization. However, existing approaches typically require costly bounding box annotations for training. Besides, they mostly adopt a customized retrieval strategy but struggle to unify various types of queries to meet diverse retrieval needs. To address these issues, we introduce Multi-query Scene Text retrieval with Attention Recycling (MSTAR), a box-free approach for scene text retrieval. It incorporates progressive vision embedding to dynamically capture the multi-grained representation of texts and harmonizes free-style text queries with style-aware instructions. Additionally, a multi-instance matching module is integrated to enhance vision-language alignment. Furthermore, we build the Multi-Query Text Retrieval (MQTR) dataset, the first benchmark designed to evaluate the multi-query scene text retrieval capability of models, comprising four query types and $16k$ images. Extensive experiments demonstrate the superiority of our method across seven public datasets and the MQTR dataset. Notably, MSTAR marginally surpasses the previous state-of-the-art model by 6.4\% in MAP on Total-Text while eliminating box annotation costs. Moreover, on the MQTR benchmark, MSTAR significantly outperforms the previous models by an average of 8.5\%. The code and datasets are available at \href{https://github.com/yingift/MSTAR}{https://github.com/yingift/MSTAR}.
Xudong Xie, Xiang Bai
NeurIPS2
2025 DG²-TCR: An Adaptive Clouds Removal Network for Optical Remote Sensing Images Using SAR-Driven Dual-Flow Fusion Guidance
abstract
Clouds in optical remote sensing images (ORSI) significantly limit image utilization. Traditional cloud removal methods using single or multi-temporal data sources struggle to ensure reliable reconstruction for thick cloud areas. Synthetic Aperture Radar (SAR) images are increasingly used to recover information obscured by clouds, but their performance in cloud-obscured regions is unstable. Therefore, an adaptive cloud removal network for remote sensing images, named DG2-TCR, is proposed based on SAR-driven dual-flow fusion guidance (DFG). DG2-TCR uses SAR and ORSI to construct DFG, including local spatial-spectral feature reconstruction (LSSFR) flow and global texture feature compensation (GTFC). LSSFR, driven by ORSI and SAR, efficiently extracts useful features in non-cloud areas and focuses on local information reconstruction using the designed spatial-spectral features inference reconstruction block (SSIRB). Based on SAR images, GTFC guides the compensation of global texture information. DFG can adaptively extract features and reconstruct missing information from local and global scales. The public SEN12MS-CR-TS dataset is divided into four sub-datasets with different coverage to evaluate the recovering capability in varying clouds. Experiments show that the PSNR, SSIM, RMSE, FID, and NCC indicator values on four sub-datasets and the SIMLE-CR dataset are better than the seven comparison methods. Furthermore, the ablation experiments show that the generalization and robustness of this proposed method on images with different cloud coverage are better than other comparison methods. Therefore, DG2-TCR can reliably recover information on cloud occlusions with various coverage and thickness, which is significant for cloud removal in practical applications.
Xianjun Gao, Jinhui Yang, Xudong Xie, Yuanwei Yang, Nan Wang 0038, Xinran Cao, Meilin Tan, Yuan Kou
IEEE Trans. Geosci. Remote. Sens.3
2024 WAS: Dataset and Methods for Artistic Text Segmentation
Xudong Xie, Yang Liu 0105, Xiang Bai
ECCV (55)1
2024 Progressive Evolution from Single-Point to Polygon for Scene Text
Linger Deng, Mingxin Huang, Xudong Xie, Xiang Bai
ICDAR (5)3
2024 ICDAR 2024 Competition on Artistic Text Recognition
Xudong Xie, Linger Deng
ICDAR (6)1
2022 Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition
Xudong Xie, Xiang Bai
ECCV (28)1
2021 Deep learning for predicting COVID-19 malignant progression
Cong Fang 0005, Song Bai 0001, Qianlan Chen, Yu Zhou 0016, Liming Xia, Lixin Qin, Shi Gong, Xudong Xie, Chunhua Zhou, Dandan Tu, Changzheng Zhang, Xiaowu Liu, Xiang Bai, Philip Torr 0001
Medical Image Anal.8
2020 Designing pulse-coupled neural networks with spike-synchronization-dependent plasticity rule: image segmentation and memristor circuit application
Xudong Xie, Shiping Wen 0001, Zheng Yan 0001, Tingwen Huang, Yiran Chen 0001
Neural Comput. Appl.1
2018 GST-memristor-based online learning neural networks
Shuixin Xiao, Xudong Xie, Shiping Wen 0001, Zhigang Zeng, Tingwen Huang, Jianhua Jiang
Neurocomputing2
2018 Memristor-based circuit implementation of pulse-coupled neural network with dynamical threshold generators
Xudong Xie, Shiping Wen 0001, Zhigang Zeng, Tingwen Huang
Neurocomputing1
2018 General memristor with applications in multilayer neural networks
Shiping Wen 0001, Xudong Xie, Zheng Yan 0001, Tingwen Huang, Zhigang Zeng
Neural Networks2
2017 Person re-identification by unsupervised video matching
Xiatian Zhu, Shaogang Gong, Xudong Xie, Jianming Hu, Kin-Man Lam 0001, Yisheng Zhong
Pattern Recognit.4
2015 Saliency detection based on singular value decomposition
Xudong Xie, Kin-Man Lam 0001, Jianming Hu, Yisheng Zhong
J. Vis. Commun. Image Represent.2
2015 Efficient saliency analysis based on wavelet transform and entropy theory
Xudong Xie, Kin-Man Lam 0001, Yisheng Zhong
J. Vis. Commun. Image Represent.2
2014 A SIFT-based mean shift algorithm for moving vehicle tracking
abstract
The classical mean shift algorithm is easy to pass into local maxima, which is caused by the lack of appropriate target model updating mechanism. In this paper, a SIFT-based mean shift algorithm is proposed, which can be used for continuous vehicle tracking in complex situations, such as the shape and the illumination of the vehicle object change. In our algorithm, the mean shift algorithm is utilized to determine the candidate target region, and then a judgment on the tracking effect is made according to the Bhattacharyya coefficient. If tracking fails, the candidate area is matched with the target model by SIFT feature, and a new track position is determined. Otherwise, the target model is periodically updated by SIFT feature matching, and the target model can be constantly updated according to the state change of the moving vehicle. In the scenes of moving vehicle target deformations, such as the variation of scale and illumination, the algorithm is tested and compared with other algorithms. The experimental results show that the proposed method can effectively track an object under the condition of varying illumination and shape deformation.
Xudong Xie, Yi Zhang 0029, Jianming Hu
Intelligent Vehicles Symposium2
2013 Color facial image denoising based on rpca and noisy pixel detection
abstract
In this paper, a novel approach for color facial-image denoising based on robust principal component analysis (RPCA) [1] in the L*a*b* color space and noisy pixel detection is proposed. Firstly, RPCA is employed for color facial-image recovery in the L*a*b* space. Then, the reconstructed image is used for noisy pixel detection. Finally, the denoised facial-image can be obtained. Experiments are conducted based on the AR database, where our proposed method is compared with several state-of-the-art image-denoising methods. Experimental results show that our method can achieve a better performance in terms of both quantitatively evaluation and visual quality.
Zhaojun Yuan, Xudong Xie, Kin-Man Lam 0001
ICASSP2
2012 Combination of global and local baseline-independent features for offline Arabic handwriting recognition
Xudong Xie, Kin-Man Lam 0001
ICPR2
2012 An efficient method for occluded face recognition
Xudong Xie, Kin-Man Lam 0001
ICPR2
2011 Color correction via robust reference selection and recovery using a low-rank matrix model
abstract
In this paper, we propose a method that can handle the color correction of a large collection of photographs simultaneously and automatically via robust reference selection. The method does not use any particular model to handle the errors on the photographs, but corrects all kinds of errors caused by changes of viewpoint, large illumination variations, gross pixel corruptions, and partial occlusions under a low-rank matrix model. Furthermore, our method uses the image pixel values directly in vector form, which preserves the spatial information, to obtain the matrix for color correction, unlike other statistics-based image-representation methods such as color histograms. Experiments verify that our method can achieve consistent and promising results on uncontrolled real photographs acquired from the Internet.
Dong Li 0028, Xudong Xie, Kin-Man Lam 0001
ICIP2
2009 Partially occluded face completion and recognition
abstract
This paper proposes a spectral graph based algorithm for face image repairing, which can improve the recognition performance on occluded faces. Our algorithm is called `guided label-learning', so named from graphical models, and can achieve a high-quality repairing of damaged or occluded faces. We apply our face repairing algorithm in order to produce completed faces, and then use face recognition to evaluate the performance of our algorithm. Experiment results show that, at most, a nearly 30%-increase in the recognition rate can be achieved for occluded faces with the use of our algorithm.
Yue Deng 0001, Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Qionghai Dai
ICIP3
2009 Elastic block set reconstruction for face recognition
abstract
In this paper, a novel face recognition algorithm named elastic block set reconstruction (EBSR) is proposed. In our method, the EBSR face is used to represent a set of training faces and to simulate different factors in a query image. An EBSR face is constructed by using the blocks from the training face images which best match to the blocks of the query image at the corresponding locations. The elastic local reconstruction (ELR) error is then used to evaluate how well a block pair matches, and the query image is classified based on the accumulated reconstruction error. The proposed method can effectively explore local information in the training set and deal with various conditions well. Also, the reconstruction error can be considered as a kind of dissimilarity measure, which gives a new approach to designing the training set so as to maximize robustness of recognition. Experiments show that consistent and promising results are obtained.
Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Zhigang Jin
ICIP2
2009 Gabor Boost Linear Discriminant Analysis for face recognition
abstract
This paper proposes an innovative algorithm named Gabor Boost Linear Discriminant Analysis (GBLDA) for face recognition. In our method, we want to estimate the distribution of high dimensional Gabor wavelet (GW) features in a low dimensional LDA subspace without computing the GW feature of an input image. The computational complexity can be reduced significantly. Hence, GBLDA is suitable for real-time applications. Experimental results show that our proposed method not only possesses the advantages of linear subspace analysis approaches such as low computational complexity, but also has the advantage of a high recognition performance in the Gabor based methods.
Xudong Xie, Qionghai Dai, Zhigang Jin
ICME2
2009 Facial expression recognition based on shape and texture
Xudong Xie, Kin-Man Lam 0001
Pattern Recognit.1
2008 Face recognition using anisotropic dual-tree complex wavelet packets
abstract
In this paper, we propose a novel face recognition method based on anisotropic dual-tree complex wavelet packets(ADT-CWP). 2-D dual-tree complex wavelet transform(DT-CWT) provides a geometrically oriented decomposition for image representation as well as shift invariance. By applying anisotropic wavelet packet decomposition on DT-CWT further, ADT-CWP can be used to extract facial features better, which turns out to benefit for face recognition. With adaptively assigning different weights to different wavelet subbands, consistent best performances can be obtained based on different face databases which are under different conditions, such as varying illuminations and expressions, compared to PCA and other face recognition methods, especially Gabor-based method. Furthermore, in addition to the consistent and promising classification performances, our proposed ADT-CWP-based method has a really low computational complexity.
YiGang Peng, Xudong Xie, Wenli Xu, Qionghai Dai
ICPR2
2008 Elastic shape-texture matching for human face recognition
Xudong Xie, Kin-Man Lam 0001
Pattern Recognit.1
2008 Face recognition using elastic local reconstruction based on a single face image
Xudong Xie, Kin-Man Lam 0001
Pattern Recognit.1
2007 An Effective Promoter Detection Method using the Adaboost Algorithm
Xudong Xie, Shuanhu Wu, Kin-Man Lam 0001, Hong Yan 0001
APBC1
2006 PromoterExplorer: an effective promoter identification method based on the AdaBoost algorithm
abstract
MOTIVATION: Promoter prediction is important for the analysis of gene regulations. Although a number of promoter prediction algorithms have been reported in literature, significant improvement in prediction accuracy remains a challenge. In this paper, an effective promoter identification algorithm, which is called PromoterExplorer, is proposed. In our approach, we analyze the different roles of various features, that is, local distribution of pentamers, positional CpG island features and digitized DNA sequence, and then combine them to build a high-dimensional input vector. A cascade AdaBoost-based learning procedure is adopted to select the most 'informative' or 'discriminating' features to build a sequence of weak classifiers, which are combined to form a strong classifier so as to achieve a better performance. The cascade structure used for identification can also reduce the false positive. RESULTS: PromoterExplorer is tested based on large-scale DNA sequences from different databases, including the EPD, DBTSS, GenBank and human chromosome 22. Experimental results show that consistent and promising performance can be achieved.
Xudong Xie, Shuanhu Wu, Kin-Man Lam 0001, Hong Yan 0001
Bioinform.1
2006 An efficient illumination normalization method for face recognition
Xudong Xie, Kin-Man Lam 0001
Pattern Recognit. Lett.1
2006 Gabor-based kernel PCA with doubly nonlinear mapping for face recognition with a single face image
abstract
In this paper, a novel Gabor-based kernel principal component analysis (PCA) with doubly nonlinear mapping is proposed for human face recognition. In our approach, the Gabor wavelets are used to extract facial features, then a doubly nonlinear mapping kernel PCA (DKPCA) is proposed to perform feature transformation and face recognition. The conventional kernel PCA nonlinearly maps an input image into a high-dimensional feature space in order to make the mapped features linearly separable. However, this method does not consider the structural characteristics of the face images, and it is difficult to determine which nonlinear mapping is more effective for face recognition. In this paper, a new method of nonlinear mapping, which is performed in the original feature space, is defined. The proposed nonlinear mapping not only considers the statistical property of the input features, but also adopts an eigenmask to emphasize those important facial feature points. Therefore, after this mapping, the transformed features have a higher discriminating power, and the relative importance of the features adapts to the spatial importance of the face images. This new nonlinear mapping is combined with the conventional kernel PCA to be called "doubly" nonlinear mapping kernel PCA. The proposed algorithm is evaluated based on the Yale database, the AR database, the ORL database and the YaleB database by using different face recognition methods such as PCA, Gabor wavelets plus PCA, and Gabor wavelets plus kernel PCA with fractional power polynomial models. Experiments show that consistent and promising results are obtained.
Xudong Xie, Kin-Man Lam 0001
IEEE Trans. Image Process.1
2005 Face recognition under varying illumination based on a 2D face shape model
Xudong Xie, Kin-Man Lam 0001
Pattern Recognit.1
2004 An efficient illumination compensation scheme for face recognition
abstract
This paper proposes a novel illumination compensation algorithm, which can compensate for the uneven illuminations on human faces and reconstruct face images of normal lighting conditions. A simple yet effective local contrast enhancement method, namely block-based histogram equalization (BHE), is first proposed. The resulting image processed by BHE is then compared with the original face image processed using histogram equalization (HE) to estimate the lighting category. Based on the category identified, a corresponding lighting compensation model is used to reconstruct an image that would visually be under normal illumination. In order to eliminate the influence of uneven illumination while retaining the shape information, a 2D face shape model is used. Experimental results show that, with the use of principal component analysis (PCA) for face recognition, the recognition rate can be improved by 53.3% to 56,1% when our proposed algorithm for lighting compensation is used.
Xudong Xie, Kin-Man Lam 0001
ICARCV1