VLDB 2026 Research / reviewers in the wild / expert
Qizhi Xu
dblp:44/6971
· DBLP profile ↗
50ranked-venue papers
7as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCH-Net: A hyperspectral object detection network with differential convolution and spectral gradient fusion
Ailin Niu, Xinyu Yan 0002, Jiuchen Chen, Xiaolin Han 0001, Qizhi Xu |
Pattern Recognit. | 6 |
| 2025 | Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large ImagesabstractGlobal contextual information and local detail features are essential for haze removal tasks. Deep learning models perform well on small, low-resolution images, but they encounter difficulties with large, high-resolution ones due to GPU memory limitations. As a compromise, they often resort to image slicing or downsampling. The former diminishes global information, while the latter discards high-frequency details. To address these challenges, we propose DehazeXL, a haze removal method that effectively balances global context and local feature extraction, enabling end-to-end modeling of large images on mainstream GPU hardware. Additionally, to evaluate the efficiency of global context utilization in haze removal performance, we design a visual attribution method tailored to the characteristics of haze removal tasks. Finally, recognizing the lack of benchmark datasets for haze removal in large images, we have developed an ultra-high-resolution haze removal dataset (8KDehaze) to support model training and testing. It includes 10000 pairs of clear and hazy remote sensing images, each sized at 8192 × 8192 pixels. Extensive experiments demonstrate that DehazeXL can infer images up to 10240 × 10240 pixels with only 21 GB of memory, achieving state-of-the-art results among all evaluated methods. The source code and experimental dataset are available at https://github.com/CastleChen339/DehazeXL. Jiuchen Chen, Xinyu Yan 0002, Qizhi Xu, Kaiqi Li |
CVPR | 3 |
| 2025 | Selection refines diagnosis: Mamba for acoustic weak fault diagnosis combining feature mode decomposition and selection
Shuchen Wang, Qizhi Xu, Hebin Liu |
Adv. Eng. Informatics | 2 |
| 2025 | Making transformer hear better: Adaptive feature enhancement based multi-level supervised acoustic signal fault diagnosis
Shuchen Wang, Qizhi Xu, Shunpeng Zhu, Biao Wang 0004 |
Expert Syst. Appl. | 2 |
| 2025 | DecloudFormer: Quest the key to consistent thin cloud removal of wide-swath multi-spectral images
Qizhi Xu, Kaiqi Li |
Pattern Recognit. | 2 |
| 2025 | CSFPR-RTDETR: Real-Time Small Object Detection Network for UAV Images Based on Cross-Spatial-Frequency Domain and Position RelationabstractSmall object detection in UAV images is one of the critical aspects for its widespread application. However, due to limited feature extraction for small object and complex backgrounds, there remain significant issues of missed detections and false alarms. This paper proposes a real-time small object detection network for UAV images based on cross spatial frequency domain and position relation (CSFPR-RTDETR). First, we propose a cross spatial frequency domain hybrid (CSFH) feature extraction network, which incorporates frequency domain processing based on the CSP network to effectively capture global contextual features and enhance the distinction between small objects and backgrounds. Second, we propose a position relation decoder that incorporates the two novel geometric priors: IoU and relative angle. Through rational characterization of spatial correlations, this design significantly strengthens the spatial perception capability of model, thereby improving the detection performance for densely distributed small objects. Finally, we design an efficient small-object high-frequency hybrid encoder, integrating the P2 detection head and proposing a mixed high-frequency enhancement fusion module (MHE-Fusion) to extract fine-grained high-frequency features of small objects, further boosting detection performance. Experimental results demonstrate that CSFPR-RTDETR achieves superior performance on the VisDrone, AI-TOD, and HIT-UAV datasets, with mAP50 metrics reaching 42.3%, 55.4% and 83.1% respectively, which is better than other SOTA models. Compared to RT-DETR, CSFPR-RTDETR reduces the parameters of the network by 29.1% while significantly enhancing detection performance: the mAP50 metrics reach notable improvements of 4.6%, 4.4%, and 1.5% on the three datasets, respectively. The source code is available at https://github.com/HuLei-JXNU/CSFPR-RTDETR. Lei Hu 0009, Jiwen Yuan, Bailiang Cheng, Qizhi Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | From Trail to Target: Efficient Infrared Moving Ship Detection via Dual-Head Supervision to Break the Slicing BarrierabstractMoving ship detection is vital for real-time maritime monitoring. Nevertheless, several challenges arise in this area: (i) Wide-area images often need to be sliced into patches to detect tiny targets, which is inefficient. (ii) The ships are small with almost no texture, leading to difficulties in accurate detection. (iii) The contrast between ships and ocean is relatively low, resulting in weak features. Although moving ships exhibit weak features, they often possess distinct wake trails. Capitalizing on this characteristic, we tailored a dual-head supervision network for moving ship detection. Initially, a dual-head supervision architecture is introduced to guide the model in using wake trails for target localization, thereby addressing the inefficiency caused by slicing. Subsequently, the background association head and target confirmation head are introduced to collaboratively enhance detection accuracy by leveraging inter-head attention mechanism. Finally, to address the issues of weak features, the dynamic feature enhancement module is embedded into backbone to boost the model’s feature extraction capability for moving targets. Experiments on GaoFen-1 dataset demonstrated that our method significantly improved the efficiency and performance of infrared moving ship detection and reached the state-of-the-art performance. Source codes will be available at https://github.com/KTqizhi/KTqizhi.github.io. Ziyang Kong, Qizhi Xu, Yuan Li 0037, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Toward Real-World Remote Sensing Image Super-Resolution: A New Benchmark and an Efficient ModelabstractSuper-resolution (SR) is a fundamental and crucial task in remote sensing. It can improve low-resolution (LR) remote sensing images and has potential benefits for downstream tasks such as remote sensing object detection and recognition. Existing remote sensing image SR (RSISR) methods are trained on simulated paired datasets, in which LR images are obtained by a simple and uniform (i.e., bicubic) degradation from corresponding high-resolution (HR) images. However, since this simulated degradation usually deviates from the real degradation, the performance of the trained model is limited when applied to real scenarios. To address this issue, we construct a novel real-world RSISR (RRSISR) dataset to model the real-world degradation, which exploits the imaging characteristics of the spectral camera to capture paired LR-HR images of the same scene. To ensure the precise alignment of the paired images, algorithms such as image registration and geometric correction are utilized. In addition, considering the vast amount of data involved in the RSISR task and its requirement for higher efficiency, we divide the image into patches with different restoration difficulties and propose a reference table-based patch exiting (RPE) method to efficiently reduce the computation of SR. Specifically, this method incorporates a predictor to estimate the performance of the current layer and a lookup table to decide whether to exit. Extensive experiments show that models trained on the proposed RRSISR dataset produce more realistic images than models with simulated datasets and generalize well to other satellites. We also demonstrate the efficiency of our RPE. Jia Wang 0038, Liuyu Xiang, Jiaochong Xu, Peipei Li 0002, Qizhi Xu, Zhaofeng He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Mitigating Texture Bias: A Remote Sensing Super-Resolution Method Focusing on High-Frequency Texture Reconstruction
Xinyu Yan 0002, Jiuchen Chen, Qizhi Xu, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Attention on the key modes: Machinery fault diagnosis transformers through variational mode decomposition
Hebin Liu, Qizhi Xu, Xiaolin Han 0001, Biao Wang 0004, Xiao-jian Yi 0001 |
Knowl. Based Syst. | 2 |
| 2024 | DTSSNet: Dynamic Training Sample Selection Network for UAV Object DetectionabstractObject detectors often struggle with accuracy and generalization when applied to aerial imagery, primarily due to the following challenges: 1) great scale variation of objects in aerial images: both extremely small and large objects are visible in the same image; and 2) an extreme imbalance of the training sample between positive and negative anchors: there are several positive ground truth (GT) anchors and an abundance of negative anchors. In this article, we propose a dynamic training sample selection network (DTSSNet) to solve the above-mentioned problems in two dimensions. An attention-enhanced feature module (AEFM) is proposed to enhance the basic features by focusing on both channel and semantic information related to targets. This module provides more valuable information for accurately classifying objects of different scales. To tackle the imbalance in training samples, this article implements a dynamic training sample selection (DTSS) module that divides the training samples based on GT information. This module dynamically selects samples, ensuring a more balanced representation of positive and negative anchors, leading to improved learning. Importantly, the combination of AEFM and DTSS does not introduce any additional computational costs. Experimental evaluations on the VisDrone2019-DET dataset demonstrate that DTSSNet outperforms base detectors and generic approaches. Furthermore, the effectiveness of DTSSNet is validated on the UAVDT benchmark dataset, where it achieves state-of-the-art performance. Chaoyang Liu, Wei Li 0032, Qizhi Xu, Hongbin Deng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Spectral Library-Based Spectral Super-Resolution Under Incomplete Spectral Coverage ConditionsabstractSpectral library based spectral super-resolution is an effective but challenging way to obtain high-spatial hyperspectral images from high-spatial multispectral images. However, the incomplete spectral coverage of spectral response functions makes it impossible to comprehensively sense the spectral information in the imaging model, thus greatly limits the performance of spectral super-resolution. To deal with this problem, a new spectral library based spectral super-resolution method under incomplete spectral coverage conditions is proposed in this paper. More specifically, a strategy for acquiring a typical set of spectra from the spectral library is proposed, trying to provide spectral observations under the incomplete spectral coverage conditions. Secondly, taking the typical set of spectra and the remaining spectral library as a priori, a new spectral super-resolution model is established under sparse and low-rank constraints. And then, the spectral dictionary is optimized utilizing the spectral information supplied by the prior spectral library. Finally, its corresponding coefficient matrix is optimized using the spatial information supplied by the multispectral image and the spectral similarity constraint on the typical spectra. Experimental results using different datasets with different spectral response functions show that, our proposed method outperforms other relative state-of-the-art methods in terms of both spectral reconstruction and spatial preservations. Xiaolin Han 0001, Wei Leng, Huan Zhang 0013, Wei Wang 0218, Qizhi Xu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TS-Track: Trajectory Self-Adjusted Ship Tracking for GEO Satellite Image Sequences via Multilevel Supervision ParadigmabstractAccurate and efficient ship tracking by geosynchronous orbit (GEO) satellites holds great significance for large-scale maritime surveillance. Nevertheless, ship tracking continues to grapple with a multitude of challenges as follows: 1) the targets are small and often obscured by cloud interference, leading to weakened features; 2) the contrasts between the ships and the background are relatively low, complicating the identification and tracking process; and 3) the frame-to-frame relative positioning accuracy is poor, posing difficulties in reflecting the actual movement trends of ships. In response to these challenges, we proposed TS-Track, a novel framework employing multilevel supervision paradigm to improve tracking performance. Initially, this framework restructured the tracking task into three key sub-modules: image enhancement, object tracking, and trajectory adjustment, inherently fostering a unified training protocol that naturally encompasses all components. Subsequently, a trajectory-based frame fusion strategy was proposed, utilizing consecutive three-frame images to enhance target features and produce consistent motion feature patterns; Last but not least, a trajectory adjustment network was developed to correct the position of ships during tracking, resulting in stable tracking trajectories, and reproduce the actual movement trends of ships. The experimental results on GaoFen-4 dataset validated that our method delivered a significant improvement in ship tracking and achieved state-of-the-art (SOTA) performance. Source codes are available athttps://github.com/KTqizhi/KTqizhi.github.io. Ziyang Kong, Qizhi Xu, Yuan Li 0037, Xiaolin Han 0001, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | DecloudNet: Cross-Patch Consistency is a Nontrivial Problem for Thin Cloud Removal From Wide-Swath Multispectral ImagesabstractCloud cover leads to great loss of spatial details in wide-swath multispectral images, and thus significantly affects their application value. Wide-swath images are huge in size and are usually cropped into patches before thin cloud removal. In addition, wide-swath images contain a rich variety of types and shapes of thin clouds, with each patch containing different clouds. However, most of the existing methods and datasets were primarily designed for natural image dehazing. These datasets had limited types and shapes of clouds. When applied to remote sensing image thin cloud removal task, these methods are unable to remove various types of clouds and lead to severe cross-patch color difference. To address this problem, a DecloudNet based on cross-patch consistency supervision was proposed. First, a multiscale cloud perception block (MCPB) with multisize convolutional kernels was proposed to enhance the network’s capability to extract clouds feature of different sizes. Second, a cross-patch consistency supervision was designed to reduce the network’s inconsistent cloud removal strength in different patches and remove cross-patch color difference when processing wide-swath images. Finally, a thin cloud simulation method based on Perlin noise, domain warping, and atmospheric scattering model was proposed to construct a high-quality declouding dataset containing clouds of multiple sizes and shapes, which can improve the performance of DecloudNet on different kinds of thin clouds. The DecloudNet and compared methods were tested for simulated thin cloud removal performance on images from QuickBird (QB), GaoFen-2 (GF2), and WorldView-2 (WV2) satellites, and for real thin cloud removal performance on wide-swath images from GaoFen-1 Wide Field of View (GF-1 WFV), GaoFen-1 (GF-1), and Earth-Observing-1 (EO-1) satellites. The experimental results demonstrated that DecloudNet outperformed the existing state-of-the-art (SOTA) methods. DecloudNet and cross-patch consistency supervision made it possible to perform thin cloud removal on wide-swath images of large size on most GPUs without worrying about graphics memory limitation. The source code and dataset are available athttps://github.com/N1rv4n4/DecloudNetthelink. Qizhi Xu, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MRF-Net: An Infrared Remote Sensing Image Thin Cloud Removal Method With the Intra-Inter Coherent ConstraintabstractThe usability of infrared remote sensing data is often compromised by thin cloud cover. To address this problem, we proposed the multiscale residual fusion network (MRF-Net) to remove thin cloud from infrared remote sensing imagery. Initially, we developed a thin cloud simulation method utilizing Perlin noise and affine transformation to generate high-fidelity thin cloud representations. Subsequently, to accurately discern and eliminate thin cloud from infrared images, we proposed MRF-Net. This model incorporates a multiscale feature fusion module (MSFFM) for extracting shallow features, a residual dense network module (RDNM) for in-depth feature extraction, a residual Swin transformer module (RSTM) for capturing global features, and attention mechanisms to selectively enhance target information. The Swin transformer, a hierarchical Transformer whose representation is computed with shifted windows, is employed to improve the efficiency of global feature extraction. Finally, we devised a combined loss function that accounts for both intrablock and interblock constraints to ensure de-clouding consistency across different image blocks. The intrablock constraint focuses on removing thin cloud within each image block, while the interblock constraint is designed to enhance the consistency of cloud removal between blocks. We have assembled a dataset comprising both simulated and real data to validate the efficiency of our proposed method. Experimental results have shown that our method effectively eliminates thin cloud and surpasses existing state-of-the-art methods. The source codes are available athttps://github.com/CastleChen339/MRF-Net. Qizhi Xu, Jiuchen Chen, Xinyu Yan 0002, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Multi-Level Supervised Network for Pansharpening to Reduce Color DistortionabstractDue to the inherent limitations of satellites, obtaining high-resolution multispectral (MS) images directly poses a challenge. Consequently, several pansharpening methods have been proposed to fuse panchromatic (Pan) images with MS images in order to generate high-resolution MS images. However, the resulting fused images often suffer from color distortion. To address this issue, we developed a multi-level supervised network aimed at minimizing color distortion. Our approach disassembled the pansharpening method into two models: an image generation module and a color optimization module. The image generation module was responsible for producing an initial fused image with rich texture, while the color optimization module focused on correcting the grey distribution of each band to achieve a high-fidelity fused image. Through experiments conducted on GaoFen-2, we have demonstrated significant improvements in reducing color distortion using our proposed method. Ziyang Kong, Qizhi Xu |
IGARSS | 3 |
| 2023 | A Joint Optimization Based Pansharpening via Subpixel-Shift DecompositionabstractPatch-based spatial dictionary has been widely used to fuse a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LMS) image under the framework of sparse representation. However, patch-based dictionary in the spatial domain is not sufficient to preserve spectral information, which may lead to large spectral distortion. To solve this problem, a new spectral dictionary based pansharpening method using subpixel-shift decomposition and joint optimization (termed as PANDA) is proposed. In this method, the model of pansharpening is formulated in a decomposed spectral domain under the sparse and low-rank constraint, as a joint optimization procedure of spectral dictionary and its coefficients. Specifically, a subpixel-shift decomposition is firstly constructed, to decompose the PAN image into a series of subimages with the same spatial resolution of the LMS image. Then, a new imaging model for the pansharpening problem of the LMS image and the decomposed PAN subimages is formulated, with sparse and low-rank constraints. And finally, a joint optimization procedure for the spectral dictionary and its coefficients are theoretically derived, using the spectral information provided by the LMS image and the spatial information provided by the entire PAN subimages, respectively. Experimental results on different datasets show that, the pansharpening performance of the proposed PANDA method outperforms the state-of-the-art methods in both spatial and spectral domains. Xiaolin Han 0001, Wei Leng, Qizhi Xu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Progressive Task-Based Universal Network for Raw Infrared Remote Sensing Imagery Ship DetectionabstractInfrared remote sensing images are becoming increasingly popular due to their superior penetration and resistance to light interference. However, challenges still remain when applying them in real-world applications: 1) raw infrared images suffer from severe stripes interference, and the preprocessing techniques used to obtain standard image products for subsequent detection tasks tend to be time-consuming, which fails to meet the application requirements; 2) current destriping techniques may inevitably weaken the local contrast between some objects and the local background since they need to consider the gray consistency of the overall image; 3) in low-resolution images, dim and small infrared targets are challenging to discriminate, resulting in high false alarms. To address these challenges, we proposed a progressive task-based universal network for raw infrared image ship detection while simultaneously removing stripes. First, we built an integrated network consisting of two components: the stripe denoising component (SDC) and the object detection component (ODC). We also designed a feedback loss adjustment mechanism to enhance the focus of the SDC on the target area. Second, a directed two-branch network was constructed for efficient stripe noise removal, including anx-direction branch for feature enhancement and ay-direction branch for grayscale smoothing. Finally, a parallel network with two labels was designed to extract the inherent features of the target and the background, as well as their relationship features, to achieve refined ship detection. We conducted experiments on a self-assembled dataset from the GaoFen-1 satellite to validate our approach. The experimental results demonstrated that the proposed method outperformed other state-of-the-art methods in infrared image ship detection. Yuan Li 0037, Qizhi Xu, Zhaofeng He 0001, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MULS-Net: A Multilevel Supervised Network for Ship Tracking From Low-Resolution Remote-Sensing Image SequencesabstractShip detection and tracking from remote sensing image sequences has become an increasingly important research point. However, there are still many challenges for ship tracking from low-resolution remote sensing image sequences: 1) the dim and small objects contain only a few shape and texture features, making it difficult to detect and track ships; 2) broken clouds often resemble ships, resulting in false tracking; 3) the ship may be occluded by clouds leading to missed tracking. To address these challenges, we proposed a novel multi-level supervision network for ship tracking from low-resolution remote sensing image sequences. First, we designed a gradient difference-guided object clarification network component to significantly improve the object saliency, which is also implemented based on the multi-frame correlation enhancement images to improve the feature strength of small targets in the input data. Second, to reduce the difficulty of completing complex tasks, a multi-level supervised network framework with multiple components was presented to achieve improving the target clarity, detecting targets and tracking targets step-by-step. Finally, to improve the trajectory integrity and tracking accuracy, a joint tracking method based on a low frame rate tracking criterion was proposed to control the state of target tracking module. The method was validated on a self-assembled dataset from the GaoFen-4 satellite. The experiment results show the stronger competitive and accuracy of the proposed method than other state-of-the-art object tracker. Yuan Li 0037, Qizhi Xu, Ziyang Kong, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SA-YOLO: The Saliency Adjusted Deep Network for Optical Satellite Image Ship DetectionabstractShip detection from remote sensing images plays an important role in military and civilian fields. However, since the small size of ship targets and the interference of cloud cover, this task still suffers from great missed detection and false-alarm. To tackle these problems, a Saliency Adjusted YOLO (SA-YOLO) for optical satellite image ship detection is developed. First, due to the fact that the ship in low resolution imagery can be regarded as a salient object, we designed a saliency guided dense sampling layer (SDSL) to improve the spatial sampling of small ship targets. Secondly, the saliency region-aware convolution (SAConv) strategy is designed to improve the representation capability of salient regions and increase the attention of network to these regions. We validated the proposed method using more than 2000 remote sensing images from GF-1 satellite. The experimental results demonstrated that the proposed method obtained a better detection performance than the state-of-the-art methods. Shuchen Wang, Hairan Sun, Yihang Zhu, Qizhi Xu |
IGARSS | 5 |
| 2022 | Optical Satellite Image Change Detection Via Transformer-Based Siamese NetworkabstractOptical satellite image change detection is essential to monitor the use of Earth's resources. Convolutional neural networks(CNN)-based methods exhibit excellent performance on change detection. As Transformers became the de-facto standard in the field of natural language processing(NLP), there were more and more methods based on it are proposed in computer vision, such as image classification, object detection, semantic segmentation and so on. Many proposed models based on vision Transformer(ViT) have surpassed the performance of CNN and show effectiveness and superiority. With the emergence of more and more applications of ViT in the field of image processing, it's advantages are gradually being explored. In terms of change detection, the CNN-based models have already shown great advantages over traditional methods. In view of current achievements of Transformer, we decided to apply Transformer to change detection in optical satellite image. Change detection of bitemporal images, we need to take two images as inputs. So we proposed a Siamese extensions of ViT networks which achieve the best results in tests on two open change detection datasets. Experimental results on real datasets show the effectiveness and the superiority of the proposed network. Qizhi Xu |
IGARSS | 4 |
| 2022 | Unsupervised Hyperspectral Pansharpening by Ratio Estimation and Residual Attention NetworkabstractMost deep learning-based hyperspectral pansharpening methods use the hyperspectral images (HSIs) as the ground truth. Training samples are usually obtained by blurring and downsampling the panchromatic image and HSI. However, the blurring and downsampling operation lose much spatial and spectral information. As a result, the model parameters trained by these reduced-resolution samples are unsuitable for fusing full-resolution images. To tackle this problem, we propose an unsupervised hyperspectral pansharpening method via ratio estimation (RE) and residual attention network (RE-RANet). The spatial and spectral information of the fused image are derived from the original panchromatic and HSI rather than reduced-resolution images. At first, we generate the initial ratio image using the ratio enhancement method. The initial ratio image is fine-tuned by the residual attention network (RANet) to generate a multichannel ratio image. Then, we inject the multichannel ratio image that contains spatial detail information into the HSI. Finally, the generated hyperspectral image is constrained by the spatial constraint loss and the spectral constraint loss. Experiments on the EO-1 and Chikusei datasets verify the effectiveness of the proposed method. Compared with other state-of-the-art approaches, our method performs well in qualitative visual effects and quantitative evaluation indicators. Jinyan Nie, Qizhi Xu, JunJun Pan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Hyperspectral Image Classification Based on Multiscale Spectral-Spatial Deformable NetworkabstractImage classification plays a fundamental role in hyperspectral image (HSI) analysis. Since the mixed pixels of the urban areas are generally more complex than other areas, the following two problems remain to be considered while dealing with urban HSI classification: 1) due to the fact that the spectral feature of different mixed pixels varies greatly in the same class, HSI classification of urban area is susceptible to the representativeness of the training samples and 2) since the urban area is densely packed with objects of different size, the comprehensive use of the spatial and spectral features to classify the objects is a difficult problem. To tackle these problems, HSI classification based on multiscale spectral–spatial deformable network (S2-DNet) is proposed. First, a$k$-means clustering method is adopted to cluster spectrum of each class, and representative samples are selected from the spectrum after clustering to reduce the impact of intraclass variation. Second, a spectral–spatial joint network is designed to extract the low-level features, including spectral features and spatial features. Third, the deformable network is introduced to extract high-level features of the object. Experimental results demonstrated that the proposed method outperformed the state-of-the-art methods on two widely used HSI data sets. Jinyan Nie, Qizhi Xu, JunJun Pan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | An Unsupervised SAR and Optical Image Fusion Network Based on Structure-Texture DecompositionabstractAlthough the unique advantages of optical and synthetic aperture radar (SAR) images promote their fusion, the integration of complementary features from the two types of data and their effective fusion remains a vital problem. To address that, a novel framework is designed based on the observation that the structure of SAR images and the texture of optical images look complementary. The proposed framework, named SOSTF, is an unsupervised end-to-end fusion network that aims to integrate structural features from SAR images and detailed texture features from optical images into the fusion results. The proposed method adopts the nest connect-based architecture, including an encoder network, a fusion part, and a decoder network. To maintain the structure and texture information of input images, the encoder architecture is utilized to extract multi-scale features from images. Then, we use the densely connected convolutional network (DenseNet) to perform feature fusion. Finally, we reconstruct the fusion image using a decoder network. In the training stage, we introduce a structure-texture decomposition model. In addition, a novel texture-preserving and structure-enhancing loss function are designed to train the DenseNet to enhance the structure and texture features of fusion results. Qualitative and quantitative comparisons of the fusion results with nine advanced methods demonstrate that the proposed method can fuse the complementary features of SAR and optical images more effectively. Yuanxin Ye, Wanchun Liu, Qizhi Xu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | COCO-Net: A Dual-Supervised Network With Unified ROI-Loss for Low-Resolution Ship Detection From Optical Satellite Image SequencesabstractLow-resolution ship detection from optical satellite image sequences is critical in high-orbit remote sensing satellite applications. However, it is still a difficult problem due to the following challenges: 1) the size of the ship is tiny in the low-resolution image; 2) the ship target is dim and the contrast with the background is low; 3) the interference of cloud and fog covering is complex and changeable. For these reasons, the targets are easily lost during the detection. In fact, the Clearer the Objects against to the background, the more Confident the Observers can detect it. In light of these considerations, we propose a COCO-Net to detect the small dynamic objects on low-resolution images in this paper. First, the multi-frame images are associated by introducing motion information as an effective compensation for small object features. Second, an integrated dual-supervised network that processes single-level tasks hierarchically is presented to adaptively enhance the input data quality of object detection without being limited by diverse scene disturbances. Third, a unified ROI-loss scheme that modulates the loss function of the first component by introducing ROI-masks from the second component is utilized to make the first component also work for object detection. In addition, we construct a new dataset for the small dynamic object detection based on the GaoFen-4 satellite imagery. Comprehensive experiments on a self-assembled dataset from the GaoFen-4 satellite show the superior performance of the proposed method compared to state-of-the-art object detectors. Qizhi Xu, Yuan Li 0037, Mingjin Zhang, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Ship Detection from Optical Remote Sensing Imagery Based on Scene Classification and Saliency-Tuned RetinanetabstractDue to the variation of ocean scene, such as thin cloud occlusion, thick cloud and sea surface clutter interferences, ship detection from optical remote sensing images often suffer from high false alarm rate. As a consequence, a ship detection method based on background classification and saliency-tuned RetinaNet is proposed to address this problem. First, the ocean scene is divided into four categories: thin cloud and fog occlusion scene, thick cloud interference scene, shore scene and calm sea scene by a designed scene classification module (SCM). Second, a multi-scale saliency feature fusion module (MSFM) is designed to provide more discriminative features for ship detection. In particular, by associating saliency maps and feature maps, our proposed MSFM can effectively suppress backgrounds noise. Finally, the MSFM are integrated into the rotation single-stage detector to effectively identify arbitrary-oriented ships from complex ocean scene. Extensive experiments on optical remote sensing dataset demonstrated that the proposed method can obtain better detection performance than the state-of-the-art methods and achieved a comparable detection speed. Ruoting Yin, Qizhi Xu, Yifang Ding |
IGARSS | 2 |
| 2021 | Gated Auxiliary Edge Detection Task for Road Extraction With Weight-Balanced LossabstractAutomated road extraction from very high-resolution (VHR) remote sensing imagery is important in many practical applications and has a long research history. Due to the diversity, narrowness, and sparsity of the road nature, extracting a full detailed road network remains a challenge, especially in the presence of interference. When applying semantic segmentation to deal with road extraction, U-Net-based architectures have achieved great progress through the use of dilated convolution or residual structure. However, the existing methods rarely focus on shape completeness and road continuity, and in fact, these are essential for road extraction. Inspirit by the multitask learning, in this letter, we present a novel road extraction architecture called gated auxiliary edge (GAE)-LinkNet with semantic segmentation as the main task and edge detection as the auxiliary task. With the proposed GatedBlocks, redundant features are filtered out and shape-relevant features stand out. Through the task loss weighing mechanism, these two tasks can work together seamlessly to make better use of the shape features. Experiments on a public road data set show that the proposed method is superior to state-of-the-art road extraction methods. Ruirui Li 0001, Bochuan Gao, Qizhi Xu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Automatic Clustering-Based Two-Branch CNN for Hyperspectral Image ClassificationabstractIt is observed that the great spectral variation in the same hyperspectral image (HSI) pixel class often leads to misclassification. To solve this problem, we have proposed an automatic clustering-based two-branch convolutional neural network (CNN): first, to reduce the intraclass spectral variation, the HSI pixels are automatically subdivided into smaller classes by clustering; second, in order to suppress the interference of spectral amplitude variation, the SincNet is introduced to capture the spectral pattern by giving more weight to the spectral shape; third, the DS-CNN with double directional strip convolution kernel is designed to extract spatial feature, so that specific contextual interactional features can be collected, especially in strip-shaped field-like roads and farmlands; finally, the spectral and spatial features extracted by the two branches are fused at fully connected layer to obtain an accurate classification. Extensive experiments demonstrated that the proposed method can obtain better classification performance than the state-of-the-art methods. Yuan Li 0037, Qizhi Xu, Wei Li 0032, Jinyan Nie |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | PSGAN: A Generative Adversarial Network for Remote Sensing Image Pan-SharpeningabstractThis article addresses the problem of remote sensing image pan-sharpening from the perspective of generative adversarial learning. We propose a novel deep neural network-based method named pansharpening GAN (PSGAN). To the best of our knowledge, this is one of the first attempts at producing high-quality pan-sharpened images with generative adversarial networks (GANs). The PSGAN consists of two components: a generative network (i.e., generator) and a discriminative network (i.e., discriminator). The generator is designed to accept panchromatic (PAN) and multispectral (MS) images as inputs and maps them to the desired high-resolution (HR) MS images, and the discriminator implements the adversarial training strategy for generating higher fidelity pan-sharpened images. In this article, we evaluate several architectures and designs, namely, two-stream input, stacking input, batch normalization layer, and attention mechanism to find the optimal solution for pan-sharpening. Extensive experiments on QuickBird, GaoFen-2, and WorldView-2 satellite images demonstrate that the proposed PSGANs not only are effective in generating high-quality HR MS images and superior to state-of-the-art methods but also generalize well to full-scale images. Qingjie Liu 0001, Huanyu Zhou, Qizhi Xu, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Pan-Sharpening with a CNN-Based Two Stage Ratio Enhancement MethodabstractWe propose a hybrid method combining the deep learning technique and the ratio enhancement (RE) method for pansharpening. The intuition behind is to utilize the deep learning technique to synthesize a panchromatic (PAN) image for the RE method to reduce the spectral distortion while keeping the spatial details. The method consists of two stages. First, the CNN synthesizer is optimized to generate the downsampled PAN image to guarantee the network have a good initialization. Second, CNN is integrated into the RE method and supervised by the ground truth multi-spectral (MS) to produce an ideal synthesized PAN for the RE method. We conduct experiments on various datasets and compare with widely used methods to demonstrate the superiority of the proposed method. Huanyu Zhou, Qingjie Liu 0001, Qizhi Xu, Yunhong Wang 0001 |
IGARSS | 3 |
| 2019 | R3-Net: A Deep Network for Multioriented Vehicle Detection in Aerial Images and VideosabstractVehicle detection is a significant and challenging task in aerial remote sensing applications. Most existing methods detect vehicles with regular rectangle boxes and fail to offer the orientation of vehicles. However, the orientation information is crucial for several practical applications, such as the trajectory and motion estimation of vehicles. In this paper, we propose a novel deep network, called a rotatable region-based residual network (R3-Net), to detect multioriented vehicles in aerial images and videos. More specially, R3-Net is utilized to generate rotatable rectangular target boxes in a half coordinate system. First, we use a rotatable region proposal network (R-RPN) to generate rotatable region of interests (R-RoIs) from feature maps produced by a deep convolutional neural network. Here, a proposed batch averaging rotatable anchor strategy is applied to initialize the shape of vehicle candidates. Next, we propose a rotatable detection network (R-DN) for the final classification and regression of the R-RoIs. In R-DN, a novel rotatable position-sensitive pooling is designed to keep the position and orientation information simultaneously while downsampling the feature maps of R-RoIs. In our model, R-RPN and R-DN can be trained jointly. We test our network on two open vehicle detection image data sets, namely, DLR 3K Munich Data set and VEDAI Data set, demonstrating the high precision and robustness of our method. In addition, further experiments on aerial videos show the good generalization capability of the proposed method and its potential for vehicle tracking in aerial videos. The demo video is available athttps://youtu.be/xCYD-tYudN0. Qingpeng Li, Lichao Mou, Qizhi Xu, Yun Zhang 0014, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Ship Detection From Thermal Remote Sensing Imagery Through Region-Based Deep ForestabstractShip detection from thermal remote sensing imagery is a challenging task because of cluttered scenes and variable appearances of ships. In this letter, we propose a novel detection algorithm named region-based deep forest (RDF) toward overcoming these existing issues. The RDF consists of a simple region proposal network and a deep forest ensemble. The region proposal network trained over gradient features robustly generates a small number of candidates that precisely cover ship targets in various backgrounds. The deep forest ensemble adaptively learns features from remote sensing data and discriminates real ships from region proposals efficiently. The training process of deep forest ensemble is efficient and users can control training cost according to computational resource available. Experimental results on numerous thermal satellite images demonstrate the superior performance of our method compared with state-of-the-art methods. Feng Yang 0002, Qizhi Xu, Bo Li 0006 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | The On-Orbit Noncloud-Covered Water Region Extraction for Ship Detection Based on Relative Spectral ReflectanceabstractFor most of the existing noncloud-covered water region (NCC water region) extraction methods, they are designed for on-ground ship detection. However, these methods could result in low accuracy and expensive computational costs. In this letter, an accurate and fast NCC water region extraction method is proposed for the typical on-orbit ship detection system with limited computing resources available. According to the relative spectral reflectance difference of water, land, and thick cloud, a triple-peak model is first constructed. Next, the parameters of the triple-peak model can be adaptively obtained based on the correlation of the previous adjacent image blocks acquired from the data stream of the panchromatic camera. Subsequently, the accurate NCC water region can be quickly extracted. Moreover, the proposed method is validated using a large amount of raw data captured by panchromatic satellite cameras. The experimental results on a Xilinx-5VFX130t field-programmable gate array imply that the proposed method performs well in the NCC water region extraction with the precision of 99.3% and recall of 99.5%, and it is suitable for on-orbit processing. Qizhi Xu, Cunguang Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Ship Detection From Optical Satellite Images Based on Saliency Segmentation and Structure-LBP FeatureabstractAutomatic ship detection from optical satellite imagery is a challenging task due to cluttered scenes and variability in ship sizes. This letter proposes a detection algorithm based on saliency segmentation and the local binary pattern (LBP) descriptor combined with ship structure. First, we present a novel saliency segmentation framework with flexible integration of multiple visual cues to extract candidate regions from different sea surfaces. Then, simple shape analysis is adopted to eliminate obviously false targets. Finally, a structure-LBP feature that characterizes the inherent topology structure of ships is applied to discriminate true ship targets. Experimental results on numerous panchromatic satellite images validate that our proposed scheme outperforms other state-of-the-art methods in terms of both detection time and detection accuracy. Feng Yang 0002, Qizhi Xu, Bo Li 0006 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Multi-sensor optical remote sensing image registration based on Line-Point InvariantabstractDue to the different imaging modalities and acquisition time, keypoint-based registration methods often suffer from false matches of keypoints while utilizing to register the optical remote sensing images from multi-sensors. In this paper, we proposed a novel method based on Line-Point Invariant for the multi-sensor image registration. First, the line segments of the images are extracted, and then the salient line segments are detected depending upon the adaptive confidence. Subsequently, conjugate salient lines between the two images are identified as the registration primitives by the probability relaxation labelling approach. Second, we obtain the SIFT keypoints of the images and establish the matches of the keypoints based on the Line-Point Invariant via dual matching. Consequently, false keypoint matches are greatly reduced and the correct match rate is significantly enhanced. The experiments conducted on various multi-sensor images demonstrate the effectiveness of the proposed method. Xianmin Wang, Qizhi Xu |
IGARSS | 2 |
| 2016 | An improved RANSAC based on the scale variation homogeneity
Yue Wang 0029, Qizhi Xu, Bo Li 0006, Hai-Miao Hu |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Automatic Change Detection in Synthetic Aperture Radar Images Based on PCANetabstractThis letter presents a novel change detection method for multitemporal synthetic aperture radar images based on PCANet. This method exploits representative neighborhood features from each pixel using PCA filters as convolutional filters. Thus, the proposed method is more robust to the speckle noise and can generate change maps with less noise spots. Given two multitemporal images, Gabor wavelets and fuzzy c-means are utilized to select interested pixels that have high probability of being changed or unchanged. Then, new image patches centered at interested pixels are generated and a PCANet model is trained using these patches. Finally, pixels in the multitemporal images are classified by the trained PCANet model. The PCANet classification result and the preclassification result are combined to form the final change map. The experimental results obtained on three real SAR image data sets confirm the effectiveness of the proposed method. Feng Gao 0005, Junyu Dong, Bo Li 0006, Qizhi Xu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Relative Radiometric Normalization for Multitemporal Remote Sensing Images by Hierarchical RegressionabstractThe existing relative radiometric normalization methods are insufficient to define the invariant pixels automatically, and the conventional methods do not perform well when the multitemporal images contain a lot of changes. Two types of changes should be particularly considered: one is caused by significant spectral differences due to change of ground objects, and the other is the pixels in the regions of misalignment caused by displacement due to differences in acquisition view angles and geometrical distortions. To automatically extract invariant pixels and reduce the influence of the changes, a hierarchical regression method is proposed to reduce the radiation difference for multitemporal images, which consists of extraction of the pseudo-invariant features (PIFs) and optimization of normalization parameters. A weighted regression based on spectral difference is proposed to automatically extract the PIFs, which can also suppress the negative effect of the first type of changes. In addition, a robust regression with gradient dependence is performed on the extracted PIFs to build the final relationship between the target image and the reference image, which can be robust for the second type of changes. Experimental results demonstrate that the proposed method has a better performance to normalize the target image. Qizhi Xu, Bo Li 0006 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Pansharpening based on an improved ratio enhancementabstractPansharpening technique is very important for many remote sensing applications. Many fusion algorithms have been proposed to pan-sharpen multispectral (MS) images. However, there are still some spatial or spectral distortion problems in fusion result. There are two major reasons: First, panchromatic (PAN) image contains some interference spectrum information similar to MS image which may cause color distortion in the fusion result. Second, MS image has some interference spatial information approximates PAN image which may cause spatial artifacts. It is difficult to simultaneous eliminate the interference information from PAN and MS images. To solve the above problems, the paper presents an improved pan-sharpen algorithm which integrates the advantages of the ratio enhancement method and Gaussian-fitting. The high-frequency information of each ithband of MS image and the low-frequency information of PAN image are extracted by Gaussian-fitting, and the information is synthesized into a group of low-resolution PAN images. Finally, each ithband of MS image is pan-sharpened by a ratio enhancement, in which the ratio is obtained by image division between the PAN image and the ithsynthesized low-resolution PAN image. Extensive experiments have been implemented on WorldView-2 images. Visual comparison and quantitative analysis demonstrated that the proposed method can achieve good performance in spatial and spectral fidelity. Xinzhi Li, Qizhi Xu, Feng Gao 0005, Lei Hu 0009 |
IGARSS | 2 |
| 2015 | Ship detection from optical satellite images based on visual search mechanismabstractAutomatic ship detection from high-resolution optical satellite images has attracted great interest in the wide applications of maritime security and traffic control. However, most of the popular methods have much difficulty in extracting targets without false alarms due to the variable appearances of ships and complicated background. In this paper, we propose a ship detection approach based on visual search mechanism to solve this problem. First, salient regions are extracted by a global contrast model fast and easily. Second, geometric properties and neighborhood similarity of targets are used for discriminating the ship candidates with ambiguous appearance effectively. Furthermore, we utilize the SVM algorithm to classify each image as including target(s) or not according to the LBP feature of each ship candidate. Extensive experiments validate our proposed scheme outperforms the state-of-the-art methods in terms of detection time and accuracy. Feng Yang 0002, Qizhi Xu, Feng Gao 0005, Lei Hu 0009 |
IGARSS | 2 |
| 2015 | Building change detection for high-resolution remotely sensed images based on a semantic dependencyabstractThe change of buildings is one of the most valuable information in the monitoring of land use for urban areas. Change detection technique based on multitemporal remote sensing image is an effective approach to obtain information of feature change. However, with the continuous improvement of resolution in remote sensing image, conventional change detection methods have much difficulty in exactly extracting building changes. One difficulty is the displacement for buildings between the Multitemporal remote sensing images due to the different view angles of sensors, another difficulty is the shadows, the above-mentioned difficulties are highlighted in the high-resolution remotely sensed Images. In this paper, a novel method for the detection of building changes from high-resolution images in urban areas is proposed, the candidate changed areas are obtained base on the spectral difference, and then a semantic dependency relation is integrated by a morphological building index technique and a shadow detection method to identify the real changes. The proposed method is evaluated with a pair of QuickBird images of Qingdao City, China. Experimental results demonstrate that the proposed method have a better performance to extract the building changes. Qizhi Xu, Feng Yang 0002, Lei Hu 0009 |
IGARSS | 2 |
| 2015 | Pansharpening Using Regression of Classified MS and Pan Images to Reduce Color DistortionabstractThe synthesis of low-resolution panchromatic (Pan) image is a critical step of ratio enhancement (RE) and component substitution (CS) pansharpening methods. The two types of methods assume a linear relation between Pan and multispectral (MS) images. However, due to the nonlinear spectral response of satellite sensors, the qualified low-resolution Pan image cannot be well approximated by a weighted summation of MS bands. Therefore, in some local areas, significant gray value difference exists between a synthetic Pan image and a high-resolution Pan image. To tackle this problem, the pixels of Pan and MS images are divided into several classes by$k$-means algorithm, and then multiple regression is used to calculate summation weights on each group of pixels. Experimental results demonstrate that the proposed technique can provide significant improvements on reducing color distortion. Qizhi Xu, Yun Zhang 0014, Bo Li 0006 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | An mean shift algorithm with adaptive bandwidth and weight selection for high spatial remotely sensed imagery segmentationabstractAn improved mean shift segmentation method featuring adaptive parameter selection is presented in this paper. We associate the bandwidths and weight for each point in a spatial-range feature space with boundary information in an image plane. Varying weight and bandwidth for each pixel are assigned according to a boundary map, which is obtained by learning multiple edge cues. We consider two groups of edge cues and two regressing modules, approach the cue combination as a supervised learning problem from the ground truth data (manually sketched boundary maps). From our preliminary results, the provided method can combine the top-down information got from regression models with the mean shift process and constrain over-clustering of pixels belonging different land objects. Qinling Dai, Leiguang Wang, Qizhi Xu, Yun Zhang 0014 |
IGARSS | 3 |
| 2014 | Ship Detection From Optical Satellite Images Based on Sea Surface AnalysisabstractAutomatic ship detection in high-resolution optical satellite images with various sea surfaces is a challenging task. In this letter, we propose a novel detection method based on sea surface analysis to solve this problem. The proposed method first analyzes whether the sea surface is homogeneous or not by using two new features. Then, a novel linear function combining pixel and region characteristics is employed to select ship candidates. Finally, Compactness and Length-width ratio are adopted to remove false alarms. Specifically, based on the sea surface analysis, the proposed method cannot only efficiently block out no-candidate regions to reduce computational time, but also automatically assign weights for candidate selection function to optimize the detection performance. Experimental results on real panchromatic satellite images demonstrate the detection accuracy and computational efficiency of the proposed method. Guang Yang 0024, Bo Li 0006, Shufan Ji, Feng Gao 0005, Qizhi Xu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | High-Fidelity Component Substitution Pansharpening by the Fitting of Substitution DataabstractDue to the difference of “mean information” between substitution component and substituted component, spectral distortion often occurs in component substitution (CS) pansharpening. In this paper, a data fitting scheme is adopted to improve spectral quality in image fusion based on well-established CS approach. A generalized CS framework that is capable of modeling any CS image fusion method is also presented. In this framework, instead of injecting detail information of panchromatic (Pan) image into substituted component, the data fitting strategy is designed to adjust the mean information of Pan image in the construction of substitution component. The data fitting scheme involves two matrix subtractions and one matrix convolution. It is fast in implementation and is effective to avoid the spectral distortion problem. Experimental results on a large number of Pan and multispectral images show that the improved CS methods have good performance on the spatial and spectral fidelity. Moreover, experiments carried out on large-size images also show an excellent running time performance of the proposed methods. Qizhi Xu, Bo Li 0006, Yun Zhang 0014 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | Multiscale Contour Extraction Using a Level Set Method in Optical Satellite ImagesabstractThis letter presents a novel coarse-to-fine level set method for contour extraction in optical satellite images. To distinguish objects from a background, the undecimated wavelet transform is firstly adopted to extract image features, and a homogeneity metric is defined to measure the variation of the features inside and outside contours. In addition, the weight distribution ratio is proposed to adaptively tune the relative weight of the features. Based on the homogeneity metric and the weight distribution ratio, a novel energy functional is developed to model a contour extraction problem, and in order to reduce the computation burden, a coarse-to-fine scheme is applied to progressively extract contours in finer scale, during which a contour position constraint is introduced to limit contours evolving in a small space around the candidate contours extracted in coarser scale. Extensive experiments have been carried out on optical satellite images to validate the proposed method. Qizhi Xu, Bo Li 0006, Zhaofeng He 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2003 | Automatic Segmentation and Recognition System for Handwritten Dates on Canadian Bank ChequesabstractThis paper describes a system being developed to recognizedate information handwritten on Canadian bankcheques. A segmentation based strategy is adopted in thissystem. In order to achieve high performances in terms ofefficiency and reliability, a knowledge-based module is proposedfor the date segmentation and a cursive month wordrecognition module is implemented based on a combinationof classifiers. The interaction between the segmentation andrecognition stages is properly established by using multi-hypothesesgeneration and evaluation modules. As a result,promising performance is obtained on a test set from a real-lifestandard cheque database. Qizhi Xu, Louisa Lam, Ching Y. Suen |
ICDAR | 1 |
| 2001 | A Knowledge-Based Segmentation System for Handwritten Dates on Bank ChequesabstractSegmenting handwritten date fields on bank cheque images into three subimages corresponding to the day, month and year is the first and critical step of our date recognition system. The paper describes a knowledge-based segmentation system, which introduces different kinds of knowledge at different segmentation stages to improve the performance. The knowledge includes information on the writing style, syntactic and semantic constraints, etc. Results have shown that the system is very effective compared with a previous structural feature based method. Qizhi Xu, Ching Y. Suen, Louisa Lam |
ICDAR | 1 |
| 2000 | Handwriting Recognition - The Last FrontiersabstractThe last frontiers of handwriting recognition are considered to have started in the last decade of the second millennium. The paper summarizes (a) the nature of the problem of handwriting recognition, (b) the state of the art of handwriting recognition at the turn of the new millennium, and (c) the results of CENPARMI researchers in automatic recognition of handwritten digits, touching numerals, cursive scripts, and dates formed by a mixture of the former 3 categories. Wherever possible, comparable results have been tabulated according to techniques used, databases, and performance. Aspects related to human generation and perception of handwriting are discussed. The extraction and usage of human knowledge, and their incorporation into handwriting recognition systems are presented. Challenges, aims, trends, efforts and possible rewards, and suggestions for future investigations are also included. Ching Y. Suen, Kye Kyung Kim, Qizhi Xu, Louisa Lam |
ICPR | 3 |
| 1999 | Automatic recognition of handwritten data on cheques - Fact or fiction?
Ching Y. Suen, Qizhi Xu, Louisa Lam |
Pattern Recognit. Lett. | 2 |