Xiangqian Wu 0002

dblp:56/4291-2 · DBLP profile ↗
← Back
71ranked-venue papers
17as first author
26since 2021 · last 2026
0000-0002-0956-8757ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 30 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 OSTE: Omni-Scene Text Editing with Latent Decoupling
Tonghua Su, Fuxiang Yang, Lei Fan 0007, Donglin Di, Zhongjie Wang 0003, Xiangqian Wu 0002
Comput. Vis. Image Underst.7
2026 Looking closer and smarter: Multi-scale progressive attention for visual text question answering
Xiangqian Wu 0002
Neurocomputing2
2026 Cluster-filter-based pseudo-label refinement for source-free domain adaptation fundus image segmentation
Yanqin Zhang, Ding Ma 0001, Xiangqian Wu 0002
Pattern Recognit.3
2026 EyeKey: Self-Supervised Keypoint Detection and Description Network Based on Local Feature Saliency for Retinal Image Global Registration
abstract
Retinal image registration (RIR) plays an important role in the diagnosis and long-term monitoring of retinal diseases. Retinal image global registration (RIGR) is usually the first step of RIR. Traditional methods often struggle to achieve robust keypoint detection and description when faced with high-resolution, fine-textured retinal images. Deep learning-based methods for this task have not been widely developed. Therefore, we propose a keypoint detection and description network based on local feature saliency, EyeKey, for RIGR. EyeKey uses the "Detect While Describing (DWD)" design. Specifically, two proposed UDPAM++ modules are embedded into the feature description network to enhance its feature description capability. Concurrently, these modules detect distinctive keypoints based on local feature saliency, combined with a Mapping Module featuring only three learnable parameters. Moreover, we achieve self-supervised feature description network training on high-resolution, fine-textured retinal images through the Random Local Hardest Example Mining strategy. Additionally, we realize robust unsupervised keypoint detection network training based on the High Matching Probability Defines Keypoints strategy and the proposed Cumulative Salient Keypoint Expansion, which, together with the DWD design, mutually reinforce the training of the keypoint detection and description network. Finally, combined with the feature-based RIGR pipeline, our method achieves outstanding performance while maintaining excellent inference speed on monomodal and multimodal RIGR evaluation datasets.
Yanchao Liang, Ding Ma 0001, Xiangqian Wu 0002
IEEE Trans. Image Process.3
2026 Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification
abstract
Vision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned with the aid of pre-trained image-text representations. Although such representations provide strong static semantic cues, they are often less effective in explicitly modeling target-state changes described by action-related phrases in natural language queries. As a result, dynamic linguistic cues, such as verbs and motion-related descriptions, may be insufficiently emphasized during cross-modal matching. To address this issue, we propose Self-Chained Dynamic Context Perception (SeDCP), a self-chained framework for explicit dynamic query modulation and language-guided visual refinement in TNL. Specifically, SeDCP consists of two coupled chains. First, the Forward Chain performs visual-evidence-guided dynamic query modulation by injecting trajectory-aware spatiotemporal cues into the language representation, thereby enhancing phrases that describe target-state changes. Second, the Backward Chain uses the dynamically enhanced query representation to refine visual spatiotemporal features, strengthening the alignment between language cues and target-state evolution. In addition, we introduce sequence-level matching rather than isolated pairwise matching to better exploit temporal dynamics, and design a Global-Local enhanced video Transformer to capture both long-range contextual dependencies and fine-grained target details. Extensive experiments on seven standard TNL benchmarks and an additional unseen $\mathrm {LaSOT}_{\mathrm {ext}}$ benchmark demonstrate that SeDCP consistently outperforms state-of-the-art methods and generalizes well to unseen categories and video characteristics.
Ding Ma 0001, Zexu Zhang, Xiangqian Wu 0002
IEEE Trans. Image Process.3
2025 Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image
abstract
Rapidly synthesizing text-relevant images has long been a significant challenge. Introducing pre-trained models into GANs can significantly enhance model performance, enabling the fast generation of high-quality images. Previous work has primarily focused on improving visual quality, with limited fine-grained research on text-image consistency. This often leads to issues such as the loss of details and concept confusion in the synthesized images. To address these challenges, we propose a novel image synthesis method combining GANs with pre-trained models. This method captures fine-grained visual concepts from text and constructs structured semantic prompts, which hierarchically guide the adjustment of visual features in the latent space. It improves image quality while significantly enhancing text-image consistency. Moreover, we introduce hard mining into GAN-based image generation tasks and propose a hard mining matching-aware loss, enabling the model to focus on the most challenging samples. This reduces the model’s dependency on batch size and training epochs, achieving good performance even under limited computational resources. Extensive experiments validate the effectiveness of our approach.
Shuheng Ge, Li Zhang 0025, Haoyu Xing, Xiangqian Wu 0002
ICASSP4
2025 Global-Local Aware Scene Text Editing
abstract
Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges.
Fuxiang Yang, Tonghua Su, Donglin Di, Xiangqian Wu 0002, Zhongjie Wang 0003, Lei Fan 0007
ICME5
2025 Source-free domain adaptation framework based on confidence constrained mean teacher for fundus image segmentation
Yanqin Zhang, Ding Ma 0001, Xiangqian Wu 0002
Neurocomputing3
2025 Joint Finger Valley Points-Free ROI Detection and Recurrent Layer Aggregation for Palmprint Recognition in Open Environment
abstract
Cooperative palmprint recognition, pivotal for civilian and commercial uses, stands as the most essential and broadly demanded branch in biometrics. These applications, often tied to financial transactions, require high accuracy in recognition. Currently, research in palmprint recognition primarily aims to enhance accuracy, with relatively few studies addressing the automatic and flexible palm region of interest (ROI) extraction (PROIE) suitable for complex scenes. Particularly, the intricate conditions of open environment, alongside the constraint of human finger skeletal extension limiting the visibility of Finger Valley Points (FVPs), render conventional FVPs-based PROIE methods ineffective. In response to this challenge, we propose an FVPs-Free Adaptive ROI Detection (FFARD) approach, which utilizes cross-dataset hand shape semantic transfer (CHSST) combined with the constrained palm inscribed circle search, delivering exceptional hand segmentation and precise PROIE. Furthermore, a Recurrent Layer Aggregation-based Neural Network (RLANN) is proposed to learn discriminative feature representation for high recognition accuracy in both open-set and closed-set modes. The Angular Center Proximity Loss (ACPLoss) is designed to enhance intra-class compactness and inter-class discrepancy between learned palmprint features. Overall, the combined FFARD and RLANN methods are proposed to address the challenges of palmprint recognition in open environment, collectively referred to as RDRLA. Experimental results on four palmprint benchmarks HIT-NIST-V1, IITD, MPD and BJTU_PalmV2 show the superiority of the proposed method RDRLA over the state-of-the-art (SOTA) competitors. The code of the proposed method is available athttps://github.com/godfatherwang2/RDRLA.
Tingting Chai, Ru Li 0002, Wei Jia 0001, Xiangqian Wu 0002
IEEE Trans. Inf. Forensics Secur.5
2025 PalmDiff: When Palmprint Generation Meets Controllable Diffusion Model
abstract
Due to its distinctive texture and intricate details, palmprint has emerged as a critical modality in biometric identity recognition. The absence of large-scale public palmprint datasets has substantially impeded the advancement of palmprint research, resulting in inadequate accuracy in commercial palmprint recognition systems. However, existing generative methods exhibit insufficient generalization, as the images they generate differ in specific ways from the conditional images. This paper proposes a method for generating palmprint images using a controllable diffusion model (PalmDiff), which addresses the issue of insufficient datasets by generating palmprint data, improving the accuracy of palmprint recognition. We introduce a diffusion process that effectively tackles the problems of excessive noise and loss of texture details commonly encountered in diffusion models. A linear attention mechanism is employed to enhance the backbone's expressive capacity and reduce the computational complexity. To this end, we proposed an ID loss function to enable the diffusion model to generate palmprint images under the same identical space consistently. PalmDiff is compared with other generation methods in terms of both image quality and the enhancement of palmprint recognition performance. Experiments show that PalmDiff performs well in image generation, with an FID score of 13.311 on MPD and 18.434 on Tongji. Besides, PalmDiff has significantly improved various backbones for palmprint recognition compared to other generation methods.
Tingting Chai, Zheng Zhang 0006, Miao Zhang 0035, Xiangqian Wu 0002
IEEE Trans. Image Process.5
2024 VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning
abstract
Achieving the optimal form of Visual Question Answering mandates a profound grasp of understanding, grounding, and reasoning within the intersecting domains of vision and language. Traditional VQA benchmarks have predom-inantly focused on simplistic tasks such as counting, visual attributes, and object detection, which do not necessitate intricate cross-modal information understanding and inference. Motivated by the need for a more comprehensive evaluation, we introduce a novel dataset comprising 23,781 questions derived from 10,124 image-text pairs. Specifically, the task of this dataset requires the model to align multimedia representations of the same entity to implement multi-hop reasoning between image and text and finally use natural language to answer the question. Furthermore, we evaluate this VTQA dataset, comparing the performance of both state-of-the-art VQA models and our proposed base-line model, the Key Entity Cross-Media Reasoning Network (KECMRN). The VTQA task poses formidable challenges for traditional VQA models, underscoring its intrinsic complexity. Conversely, KECMRN exhibits a modest improvement, signifying its potential in multimedia entity alignment and multi-step reasoning. Our analysis underscores the diversity, difficulty, and scale of the VTQA task compared to previous multimodal QA datasets. In conclusion, we anticipate that this dataset will serve as a pivotal resource for advancing and evaluating models proficient in multime-dia entity alignment, multi-step reasoning, and open-ended answer generation. Our dataset and code is available at https://visual-text-qa.github.io/
Xiangqian Wu 0002
CVPR2
2024 Do Keypoints Contain Crucial Information? Mining Keypoint Information to Enhance Cross-View Geo-Localization
abstract
Due to drastic view changes and different capturing times between images, extracting discriminative image-level features for cross-view geo-localization is challenging. Although recent works have achieved outstanding progress on cross-view geo-localization, the fine-grained information in images has not been fully explored in extracting image-level features. Inspired by the process of the human visual system to distinguish similar targets and the process of keypoint detection and description, we propose a framework called UDPA-Net, which guides the model to mine more favorable information for cross-view geolocalization by detecting keypoints. Specifically, we design a Unit Dot Product Attention Module (UDPAM) to discover remarkable keypoints automatically and guide the model to pay more attention to the salient regions. UDPA-Net introduces few parameters but yields significant performance gains and can be easily integrated into different networks. Our code is available at https://gitee.com/KerasLyc/UDPA-Net.
Yanchao Liang, Xiangqian Wu 0002
ICME2
2024 WSRFNet: Wavelet-Based Scale-Specific Recurrent Feedback Network for Diabetic Retinopathy Lesion Segmentation
Xiangqian Wu 0002
IJCAI2
2024 Parallel weight control based on policy gradient of relation refinement for cross-modal retrieval
Li Zhang 0025, Yahu Yang, Shuheng Ge, Guanghui Sun, Xiangqian Wu 0002
Eng. Appl. Artif. Intell.5
2023 Tracking by Natural Language Specification with Long Short-term Context Decoupling
abstract
The main challenge of Tracking by Natural Language Specification (TNL) is to predict the movement of the target object by giving two heterogeneous information, e.g., one is the static description of the main characteristics of a video contained in the textual query, i.e., long-term context; the other one is an image patch containing the object and its surroundings cropped from the current frame, i.e., the search area. Currently, most methods still struggle with the rationality of using those two information and simply fusing the two. However, the linguistic information contained in the textual query and the visual representation stored in the search area may sometimes be inconsistent, in which case the direct fusion of the two may lead to conflicts. To address this problem, we propose DecoupleTNL, introducing a video clip containing short-term context information into the framework of TNL and exploring a proper way to reduce the impact when visual representation is inconsistent with linguistic information. Concretely, we design two jointly optimized tasks, i.e., short-term context-matching and long-term context-perceiving. The context-matching task aims to gather the dynamic short-term context information in a period, while the context-perceiving task tends to extract the static long-term context information. After that, we design a long short-term modulation module to integrate both context information for accurate tracking. Extensive experiments have been conducted on three tracking benchmark datasets to demonstrate the superiority of DecoupleTNL.
Ding Ma 0001, Xiangqian Wu 0002
ICCV2
2023 VTQA2023: ACM Multimedia 2023 Visual Text Question Answering Challenge
abstract
The ideal form of Visual Question Answering requires understanding, grounding and reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most existing VQA benchmarks are limited to just picking the answer from a pre-defined set of options and lack attention to text. We present a new challenge with a dataset that contains 23,781 questions based on 10124 image-text pairs. Specifically, the task requires the model to align multimedia representations of the same entity to implement multi-hop reasoning between image and text and finally use natural language to answer the question. The aim of this challenge is to develop and benchmark models that are capable of multimedia entity alignment, multi-step reasoning and open-ended answer generation.
Tianli Zhao, Xiangqian Wu 0002
ACM Multimedia3
2023 Enhancing Feature Representation for Anomaly Detection via Local-and-Global Temporal Relations and a Multi-stage Memory
Ding Ma 0001, Xiangqian Wu 0002
PRCV (6)3
2023 Feature Refinement from Multiple Perspectives for High Performance Salient Object Detection
Congao Wang, Ding Ma 0001, Xiangqian Wu 0002
PRCV (12)4
2023 Capsule-Based Regression Tracking via Background Inpainting
abstract
Background cues play an accompanying role in most regression trackers, where they directly learn a mapping from dense sampling to soft label by giving a search area. In essence, the trackers need to identify a large amount of background information (i.e., other objects and distractor objects) under the circumstance of extreme target-background data imbalance. Therefore, we believe that it is more worth performing regression tracking depending on the informative background cues and using target cues as supplementary. To do this, we propose a capsule-based approach, referred to as CapsuleBI, which performs regression tracking based on a background inpainting network and a target-aware network. The background inpainting network explores the background representations by restoring the region of the target with all available scenes, and a target-aware network captures the target representations by focusing on the target itself only. To explore the subjects/distractors in the whole scene, we propose a global-guided feature construction module, which helps enhance the local features with global information. Both the background and target are encoded in capsules, which can model the relationships between objects or object parts in the background scene. Apart from this, the target-aware network assists the background inpainting network with a novel background-target routing algorithm that guides the background and target capsules to estimate the target location with multi-video relationships information precisely. Extensive experimental results show that the proposed tracker achieves favorably against state-of-the-art methods.
Ding Ma 0001, Xiangqian Wu 0002
IEEE Trans. Image Process.2
2022 QuadTreeCapsule: QuadTree Capsules for Deep Regression Tracking
abstract
Benefit from the capability of capturing part-to-whole relationships, Capsule Network has been successful in many vision tasks. However, their high computational complexity poses a significant obstacle to applying them to visual tracking, requiring fast inference. In this paper, we introduce the idea of QuadTree Capsules, which explores the property of part-to-whole relationships endowed by the Capsule Network by significantly reducing the computational complexity. We build capsule pyramids and select meaningful relationships in a coarse-to-fine manner, dubbed as QuadTreeCapsule. Specifically, the top K capsules with the highest activation values are selected, and routing is only calculated within the relevant regions corresponding to these top K capsules with a novel symmetric guided routing algorithm. Additionally, considering the importance of temporal relationships, a multi-spectral pose matrix attention mechanism is developed for more accurate spatio-temporal capsule assignments between two sets of capsules. Moreover, during online inference, we shift part of the spatio-temporal capsules long the temporal dimension, facilitating information exchanged among neighboring frames. Extensive experimentation has proved the effectiveness of our methodology, which achieves state-of-the-art results compared with other tracking methods on eight widely-used benchmarks. Our tracker runs at approximately 43 fps on GPU.
Ding Ma 0001, Xiangqian Wu 0002
ACM Multimedia2
2022 Multi-task framework based on feature separation and reconstruction for cross-modal retrieval
Li Zhang 0025, Xiangqian Wu 0002
Pattern Recognit.2
2022 Latent Space Semantic Supervision Based on Knowledge Distillation for Cross-Modal Retrieval
abstract
As an important field in information retrieval, fine-grained cross-modal retrieval has received great attentions from researchers. Existing fine-grained cross-modal retrieval methods made several improvements in capturing the fine-grained interplay between vision and language, failing to consider the fine-grained correspondences between the features in the image latent space and the text latent space respectively, which may lead to inaccurate inference of intra-modal relations or false alignment of cross-modal information. Considering that object detection can get the fine-grained correspondences of image region features and the corresponding semantic features, this paper proposed a novel latent space semantic supervision model based on knowledge distillation (L3S-KD), which trains classifiers supervised by the fine-grained correspondences obtained from an object detection model by using knowledge distillation for image latent space fine-grained alignment, and by the labels of objects and attributes for text latent space fine-grained alignment. Compared with existing fine-grained correspondence matching methods, L3S-KD can learn more accurate semantic similarities for local fragments in image-text pairs. Extensive experiments on MS-COCO and Flickr30K datasets demonstrate that the L3S-KD model consistently outperforms state-of-the-art methods for image-text matching.
Li Zhang 0025, Xiangqian Wu 0002
IEEE Trans. Image Process.2
2021 CapsuleRRT: Relationships-Aware Regression Tracking via Capsules
abstract
Regression tracking has gained more and more attention thanks to its easy-to-implement characteristics, while existing regression trackers rarely consider the relationships between the object parts and the complete object. This would ultimately result in drift from the target object when missing some parts of the target object. Recently, Capsule Network (CapsNet) has shown promising results for image classification benefits from its part-object relationships mechanism, while CapsNet is known for its high computational demand even when carrying out simple tasks. Therefore, a primitive adaptation of CapsNet to regression tracking does not make sense, since this will seriously affect speed of a tracker. To solve these problems, we first explore the spatial-temporal relationships endowed by the CapsNet for regression tracking. The entire regression framework, dubbed CapsuleRRT, consists of three parts. One is S-Caps, which captures the spatial relationships between the parts and the object. Meanwhile, a T-Caps module is designed to exploit the temporal relationships within the target. The response of the target is obtained by STCaps Learning. Further, a prior-guided capsule routing algorithm is proposed to generate more accurate capsule assignments for subsequent frames. Apart from this, the heavy computation burden in CapsNet is addressed with a knowledge distillation pose matrix compression strategy that exploits more tight and discriminative representation with few samples. Extensive experimental results show that CapsuleRRT performs favorably against state-of-the-art methods in terms of accuracy and speed.
Ding Ma 0001, Xiangqian Wu 0002
CVPR2
2021 Capsule-based Object Tracking with Natural Language Specification
abstract
Tracking with Natural-Language Specification (TNL) is a joint topic of understanding the vision and natural language with a wide range of applications. In previous works, the communication between two heterogeneous features of vision and language is mainly through a simple dynamic convolution. However, the performance of prior works is capped by the difficulty of linguistic variation of natural language in modeling the dynamically changing target and its surroundings. In the meanwhile, natural language and vision are firstly fused and then utilized for tracking, which is hard to model the query-focused context. Query-focused should pay more attention to context modeling to promote the correlation between these two features. To address these issues, we propose a capsule-based network, referred to as CapsuleTNL, which performs regression tracking with natural language query. In the beginning, the visual and textual input is encoded with capsules, which can not only establish the relationship between entities but also the relationship between the parts of the entity itself. Then, we devise two interaction routing modules, which consist of visual-textual routing module to reduce the linguistic variation of input query and textual-visual routing module to precisely incorporate query-based visual cues simultaneously. To validate the potential of the proposed network for visual object tracking, we evaluate our method on two large tracking benchmarks. The experimental evaluation demonstrates the effectiveness of our capsule-based network.
Ding Ma 0001, Xiangqian Wu 0002
ACM Multimedia2
2021 Conditioners for Adaptive Regression Tracking
Ding Ma 0001, Xiangqian Wu 0002
PRCV (1)2
2021 Distillation-Based Multi-exit Fully Convolutional Network for Visual Tracking
Ding Ma 0001, Xiangqian Wu 0002
PRCV (1)2
2019 Pyramid Feature Attention Network for Saliency Detection
abstract
Saliency detection is one of the basic challenges in computer vision. Recently, CNNs are the most widely used and powerful techniques for saliency detection, in which feature maps from different layers are always integrated without distinction. However, instinctively, the different feature maps of CNNs and the different features in the same maps should play different roles in saliency detection. To address this problem, a novel CNN named pyramid feature attention network (PFAN) is proposed to enhance the high-level context features and the low-level spatial structural features. In the proposed PFAN, a context-aware pyramid feature extraction (CPFE) module is designed for multi-scale high-level feature maps to capture the rich context features. A channel-wise attention (CA) model and a spatial attention (SA) model are respectively applied to the CPFE feature maps and the low-level feature maps, and then fused to detect salient regions. Finally, an edge preservation loss is proposed to get the accurate boundaries of salient regions. The proposed PFAN is extensively evaluated on five benchmark datasets and the experimental results demonstrate that the proposed network outperforms the state-of-the-art approaches under different evaluation metrics.
Xiangqian Wu 0002
CVPR2
2019 High Speed Recurrent Regression Network for Visual Tracking
abstract
For some recently released trackers, the spatial-temporal information of the target are processed separately, which is time consuming and inefficient for locating the target in the sequential data-videos. To solve this problem, we present a recurrent regression framework(RRNet), which leverages spatial and temporal information coherence on feature level simultaneously. The RRNet is composed of a regression network and a long short term memory network(LSTM). The regression network is learned on static image level for focusing on the spatial information of the target, and the whole framework is fine-tuned on videos by fixing the parameters of the regression network, which improves the per-frame regression by aggregation of recurrent prior. Especially, there is no need to online training for adapting the unseen targets. And the experimental results show that the proposed RRNet gets better performance than the compared trackers with a high speed (45 fps).
Ding Ma 0001, Xiangqian Wu 0002
ICME2
2019 Salient Object Detection Using Cascaded Convolutional Neural Networks and Adversarial Learning
abstract
Salient object detection has received much attention and achieved great success in last several years. It is still challenging to get clear boundaries and consistent saliencies, which can be considered as the structural information of salient objects. A popular solution is to conduct some post-processes (e.g., conditional random field (CRF)) to refine these structural information. In this paper, a novel cascaded convolutional neural networks (CNNs) based method is proposed to implicitly learn these structural information via adversarial learning for salient object detection (we termed the proposed method as CCAL). A cascaded CNNs model is first designed as a generator G, which consists of an encoder-decoder network for global saliency estimation and a deep residual network for local saliency refinement. It is hard to explicitly learn such structural information due to the limitation of frequently-used pixel-wise loss functions. Instead, a discriminator D is then designed to distinguish the real salient maps (i.e., ground truths) from the fake ones produced by G, based on which an adversarial loss is introduced to optimize G. G and D are trained in a fully end-to-end fashion by following the strategy of conditional generative adversarial networks to make G well learn the structural information. At last, G is able to produce high quality salient maps without requiring any post-process to fool D. Experimental results on eight benchmark datasets demonstrate the effectiveness and efficiency (about 17 fps on graphics processing unit (GPU)) of the proposed method for salient object detection.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Multim.2
2018 Multi-Scale Recurrent Tracking via Pyramid Recurrent Network and Optical Flow
Ding Ma 0001, Wei Bu, Xiangqian Wu 0002
BMVC3
2018 Dual-SVM tracker via Multiple Support Instance and LEVER Strategy
abstract
Visual tracking can be modeled as a binary classification problem, and the classic classifiersupport vector machine (SVM) based methods have been demonstrated encouraging performance in recent object tracking benchmarks. However, the performance of SVM is too sensitive to noisy training data during online update. In this paper, we propose an efficient dual-SVM based tracker to improve classification performance for visual tracking. The tracker proposed consists of two models: the holistic model and the part model. To learn the holistic model, the support instances are derived from the RMI-SVM trained in a deep feature space. As for the part model to highlight local structure of the target, a linear SVM is learned to further encode local details of the target, selecting candidate instances from the support instances by the confidence as input. To fuse the holistic model and the part model, we design a simple but efficient decision strategy (LEVER) to enforce the dual-SVM to focus on the target. The proposed LEVER is updated incrementally to capture changes of the appearance of the target. Extensive experimental results show that the proposed tracker performs favorably against state-of-the-art methods.
Ding Ma 0001, Wei Bu, Xiangqian Wu 0002
ICPR3
2018 Learning Collaborative Model for Visual Tracking
abstract
This paper proposes a robust visual tracking method by designing a collaborative model. The collaborative model employs a two-stage tracker and a HOG-based detector, which exploits both holistic and local information of the target. The two-stage tracker learns a linear classifier from the patches of original images and the HOG-based detector trains a linear discriminant analysis classifier with the object exemplar. Finally, a result decision making strategy is developed by considering both the original template and the appearance variations, making the tracker and the detector collaborate with each other. The proposed method has been evaluated on OTB-50, OTB-100 and Temple-Color datasets, and results demonstrate that the proposed method is able to effectively address the challenging cases such as scale variation and out-of-view and gets better performance than the state-of-the-art trackers.
Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002
ICPR5
2018 Segmentation-Guided Tracking with Prior Map Decision
abstract
For visual tracking, the target object is represented by an appearance model and the location of the target is estimated in each frame. Numerous tracking algorithms model the appearance of the target with a confidence score and rarely take into account the semantic information of the target. In this paper, we propose an efficient tracking algorithm that models the appearance of the target based on semantic segmentation. The overall architecture consists of two parts: the segmentation part and the tracking part. In the segmentation part, an attention model is employed, providing spatial highlights of the candidate region of the target. In the tracking part, the tracker is constructed by an online updated convolutional neural networks to identify the target in subsequent frames, taking advantage of the segmentation information of the target from the segmentation part. To enhance the performance of this architecture, we design an incremental updated prior map taking both the segmentation signal and the tracking signal into consideration. Extensive experiments on two benchmarks including OTB-50, OTB-100, and Temple-Color, show that the proposed method outperforms other trackers.
Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002
ICPR5
2018 Image Saliency Detection with Low-Level Features Enhancement
Xiangqian Wu 0002
PRCV (1)2
2018 Scene Text Detection Using Superpixel-Based Stroke Feature Transform and Deep Learning Based Region Classification
abstract
Scene text detection is a crucial step in end-to-end scene text recognition, a greatly challenging problem in computer vision. This paper proposes a novel scene text detection method that involves superpixel-based stroke feature transform (SSFT) and deep learning based region classification (DLRC). The SSFT is developed for candidate character region (CCR) extraction, which consists in partitioning an input image into several regions via superpixel-based clustering, removing most regions based on predefined criteria satisfied by the characters, and refining the remaining regions to obtain CCRs by computing a stroke width map. The character regions are identified from the CCRs using DLRC, in which several hand-crafted low-level features, i.e., color, texture, and geometric features, and some deep convolution neural network (CNN) based high-level features are first extracted from the regions, and then these features are fused by using two fully connected networks (FCNs) for region classification. In the DLRC step, the deep feature extraction CNN and the feature fusion FCNs are jointly trained. Next, the extracted character regions are merged to form candidate text regions, from which the final scene texts are detected. The proposed method is evaluated on three publicly available datasets: ICDAR2011, ICDAR2013, and street view text. It achieves F -measures of 0.876, 0.885, and 0.631, respectively, which demonstrate the effectiveness of the proposed scene text detection method.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Multim.2
2017 Salient Object Detection with Chained Multi-Scale Fully Convolutional Network
abstract
In this paper, we proposed a novel method for effective salient object detection by designing a chained multi-scale fully convolutional network (CMSFCN). CMSFCN contained multiple single-scale fully convolutional networks (SSFCNs), which were integrated successively by using chained connections and generated saliency prediction results from coarse to fine. The chained connections not only combined the saliency prediction result from previous SSFCN with the input image of current SSFCN, but also combined the intermediate features from previous SSFCN and current SSFCN. With these chained connections, the sequential SSFCNs in CMSFCN automatically learned complemental and discriminative features to improve the saliency predictions progressively. Therefore, after jointly training CMSFCN with an end-to-end manner, precise saliency prediction results were produced under a coarse-to-fine behaviour. Compared with seven state-of-the-art CNN based salient object detection approaches over five benchmark datasets, experimental results demonstrated the efficiency and effectiveness of CMSFCN.
Youbao Tang, Xiangqian Wu 0002
ACM Multimedia2
2017 Optic disc segmentation based on variational model with multiple energies
Baisheng Dai, Xiangqian Wu 0002, Wei Bu
Pattern Recognit.2
2017 Scene Text Detection and Segmentation Based on Cascaded Convolution Neural Networks
abstract
Scene text detection and segmentation are two important and challenging research problems in the field of computer vision. This paper proposes a novel method for scene text detection and segmentation based on cascaded convolution neural networks (CNNs). In this method, a CNN based text-aware candidate text region (CTR) extraction model (named detection network, DNet) is designed and trained using both the edges and the whole regions of text, with which coarse CTRs are detected. A CNN based CTR refinement model (named segmentation network, SNet) is then constructed to precisely segment the coarse CTRs into text to get the refined CTRs. With DNet and SNet, much fewer CTRs are extracted than with traditional approaches while more true text regions are kept. The refined CTRs are finally classified using a CNN based CTR classification model (named classification network, CNet) to get the final text regions. All of these CNN based models are modified from VGGNet-16. Extensive experiments on three benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance and greatly outperforms other scene text detection and segmentation approaches.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Image Process.2
2016 Saliency Detection via Combining Region-Level and Pixel-Level Predictions with CNNs
Youbao Tang, Xiangqian Wu 0002
ECCV (8)2
2016 Scene Text Detection via Edge Cue and Multi-features
abstract
Inspired by the fact that edge is an important cue to distinguish texts from background, we propose a novel scene text detection method via edge cue and multiple features, which has two main parts, i.e. candidate character region (CCR) extraction and region classification. For CCR extraction, the edges are first extracted from the input image, which are then broken and merged based on color features to form the final edge image. For each edge connected component, a number of image patches are extracted by translating and scaling its boundary rectangle to generate the CCRs. For region classification, the character regions are extracted from the CCRs by using a region classification technique, which extracts both the hand-designed low-level features and deep convolution neural network based high-level features of the regions for classification. And then the character regions are merged to form the candidate text regions, based on which the final text region are detected by using the region classification technique. The proposed method is evaluated on two latest ICDAR benchmark datasets and the experimental results demonstrate that the proposed method outperforms the state-of-the-art approaches of scene text detection.
Youbao Tang, Xiangqian Wu 0002
ICFHR2
2016 Text-Independent Writer Identification via CNN Features and Joint Bayesian
abstract
This paper proposes a novel method for offline text-independent writer identification by using convolutional neural network (CNN) and joint Bayesian, which consists of two stages, i.e. feature extraction and writer identification. In the stage of feature extraction, since a large number of data is essential to train an effective CNN model with high generalizability and the amount of handwriting is limited in writer identification, a data augmentation technique is first developed to generate thousands of handwriting images for each writer. Then a deep CNN network is designed to extract discriminative features to represent the properties of different writing styles, which is trained by using the generated handwriting images. In the stage of writer identification, the training dataset is used to train the CNN model for feature extraction and the joint Bayesian technique is employed to accomplish the task of writer identification based on the extracted CNN features. The proposed method is tested on two standard benchmark datasets, i.e. ICDAR2013 and CVL dataset. Experimental results demonstrate that the proposed method gets the best performance compared to the state-of-the-art approaches.
Youbao Tang, Xiangqian Wu 0002
ICFHR2
2016 Deeply-Supervised Recurrent Convolutional Neural Network for Saliency Detection
abstract
This paper proposes a novel saliency detection method by developing a deeply-supervised recurrent convolutional neural network (DSRCNN), which performs a full image-to-image saliency prediction. For saliency detection, the local, global, and contextual information of salient objects is important to obtain a high quality salient map. To achieve this goal, the DSRCNN is designed based on VGGNet-16. Firstly, the recurrent connections are incorporated into each convolutional layer, which can make the model more powerful for learning the contextual information. Secondly, side-output layers are added to conduct the deeply-supervised operation, which can make the model learn more discriminative and robust features by effecting the intermediate layers. Finally, all of the side-outputs are fused to integrate the local and global information to get the final saliency detection results. Therefore, the DSRCNN combines the advantages of recurrent convolutional neural networks and deeply-supervised nets. The DSRCNN model is tested on five benchmark datasets, and experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art saliency detection approaches on all test datasets.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
ACM Multimedia2
2016 Optic Disc Localization Using Directional Models
abstract
Reliable localization of the optic disc (OD) is important for retinal image analysis and ophthalmic pathology screening. This paper presents a novel method to automatically localize ODs in retinal fundus images based on directional models. According to the characteristics of retina vessel networks, such as their origin at the OD and parabolic shape of the main vessels, a global directional model, named the relaxed biparabola directional model, is first built. In this model, the main vessels are modeled by using two parabolas with a shared vertex and different parameters. Then, a local directional model, named the disc directional model, is built to characterize the local vessel convergence in the OD as well as the shape and the brightness of the OD. Finally, the global and the local directional models are integrated to form a hybrid directional model, which can exploit the advantages of the global and local models for highly accurate OD localization. The proposed method is evaluated on nine publicly available databases, and achieves an accuracy of 100% for each database, which demonstrates the effectiveness of the proposed OD localization method.
Xiangqian Wu 0002, Baisheng Dai, Wei Bu
IEEE Trans. Image Process.1
2015 Saliency Detection Based on Graph-Structural Agglomerative Clustering
abstract
This paper proposes a novel saliency detection method based on graph-structural agglomerative clustering (GSAC). In this method, a number of intermediate images with consecutive number of regions are firstly created by using GSAC to the input image with the maximum incremental path integral criterion. Then an initial salient map is computed based on the boundary connectivity of the regions in the intermediate images, with enforcing the early formed objects in the clustering process. Finally, the initial salient map is refined to get the final salient map by using the reconstruction errors of sparse coding and the object-bias prior. The experimental results demonstrate that the proposed method greatly outperforms the state-of-the-art approaches on two standard benchmark datasets.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
ACM Multimedia2
2015 Deformed Palmprint Matching Based on Stable Regions
abstract
Palmprint recognition (PR) is an effective technology for personal recognition. A main problem, which deteriorates the performance of PR, is the deformations of palmprint images. This problem becomes more severe on contactless occasions, in which images are acquired without any guiding mechanisms, and hence critically limits the applications of PR. To solve the deformation problems, in this paper, a model for non-linearly deformed palmprint matching is derived by approximating non-linear deformed palmprint images with piecewise-linear deformed stable regions. Based on this model, a novel approach for deformed palmprint matching, named key point-based block growing (KPBG), is proposed. In KPBG, an iterative M-estimator sample consensus algorithm based on scale invariant feature transform features is devised to compute piecewise-linear transformations to approximate the non-linear deformations of palmprints, and then, the stable regions complying with the linear transformations are decided using a block growing algorithm. Palmprint feature extraction and matching are performed over these stable regions to compute matching scores for decision. Experiments on several public palmprint databases show that the proposed models and the KPBG approach can effectively solve the deformation problem in palmprint verification and outperform the state-of-the-art methods.
Xiangqian Wu 0002, Qiushi Zhao
IEEE Trans. Image Process.1
2014 Text Line Segmentation Based on Matched Filtering and Top-Down Grouping for Handwritten Documents
abstract
This paper presents a novel text line segmentation method based on matched filtering and top-down grouping for handwritten documents. The proposed method consists of three distinct steps. Firstly, the foreground pixel density (FPD) of handwritten document image (HDI) is estimated, then FPD is used to decide the size of the generated filter which is the convolution of a band-shape filter and an isotropic LoG filter. Secondly, the centers of the text lines (CTLs) are extracted by performing filtering, binarizing, thinning and top-down grouping operation on HDI. Finally, the overlapping connected-components (OCCs) which travel through multiple text lines are separated, and then all OCCs are assigned to a label of CTLs by the nearest neighbor principle. The proposed method is tested on two public databases, and the experimental results show that the proposed method outperforms the state-of-the-art text line segmentation approaches in both of these databases.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
Document Analysis Systems2
2014 Exploiting diversity for optimizing margin distribution in ensemble learning
Qinghua Hu, Leijun Li, Xiangqian Wu 0002, Gerald Schaefer, Daren Yu
Knowl. Based Syst.3
2014 Exploration of classification confidence in ensemble learning
Leijun Li, Qinghua Hu, Xiangqian Wu 0002, Daren Yu
Pattern Recognit.3
2014 A SIFT-based contactless palmprint verification approach using iterative RANSAC and local palmprint descriptors
Xiangqian Wu 0002, Qiushi Zhao, Wei Bu
Pattern Recognit.1
2014 Offline Text-Independent Writer Identification Based on Scale Invariant Feature Transform
abstract
This paper proposes a novel offline text-independent writer identification method based on scale invariant feature transform (SIFT), composed of training, enrollment, and identification stages. In all stages, an isotropic LoG filter is first used to segment the handwriting image into word regions (WRs). Then, the SIFT descriptors (SDs) of WRs and the corresponding scales and orientations (SOs) are extracted. In the training stage, an SD codebook is constructed by clustering the SDs of training samples. In the enrollment stage, the SDs of the input handwriting are adopted to form an SD signature (SDS) by looking up the SD codebook and the SOs are utilized to generate a scale and orientation histogram (SOH). In the identification stage, the SDS and SOH of the input handwriting are extracted and matched with the enrolled ones for identification. Experimental results on six public data sets (including three English data sets, one Chinese data set, and two hybrid-language data sets) demonstrate that the proposed method outperforms the state-of-the-art algorithms.
Xiangqian Wu 0002, Youbao Tang, Wei Bu
IEEE Trans. Inf. Forensics Secur.1
2013 A level set method for very high resolution airborne sar image segmentation
abstract
This paper investigates the segmentation problem for very high resolution airborne synthetic aperture radar (SAR) images. In addition to the instinct speckles, these images show two extra characteristics: scene complexity and intensity inhomogeneity, which make segmentation more difficult. An unsupervised solution is proposed based on level set method. First, a new level set evolution method is put forward, it can get global minimum without initial contour, thus can handle complex images automatically. And the new evolution function also introduces the localizing idea from region-scalable-fitting (RSF) model to deal with the intensity inhomogeneity. Then the two segmentation results for background and targets are fused. The experimental results on real images demonstrate the effectiveness of the proposed method.
Siliang Sun, Junping Zhang, Bin Zou 0001, Xiangqian Wu 0002
ICIP4
2013 Contactless palmprint verification based on SIFT and iterative RANSAC
abstract
Contactless palmprint recognition is an effective technology for improving the user friendliness of palmprint recognition. The main challenge of contactless palmprint recognition is the intra-class variations due to the hand deformation caused by contactless image acquisition. Traditional palmprint feature extraction and matching algorithms usually require that the query image is well aligned with the gallery image, which cannot be always assured in contactless occasions. In this work, the scale invariant feature transform (SIFT) is applied for contactless palmprint feature extraction and matching. SIFT features are invariant to image rotation, translation, and scale variations, and hence are promising to solve the deformation problem. Moreover, an iterative RANSAC algorithm is proposed to refine the matched SIFT points. The iterative RANSAC retains more matched SIFT points than the traditional RANSAC algorithm despite the points comply with different transformation models. Experiments show a great improvement of verification accuracy on a public contactless palmprint database.
Qiushi Zhao, Xiangqian Wu 0002, Wei Bu
ICIP2
2013 Dynamic classifier ensemble using classification confidence
Leijun Li, Bo Zou 0001, Qinghua Hu, Xiangqian Wu 0002, Daren Yu
Neurocomputing4
2012 Retinal vessel segmentation via Iterative Geodesic Time Transform
Baisheng Dai, Wei Bu, Xiangqian Wu 0002, Yan Teng 0004
ICPR3
2010 Detection of microaneurysms using multi-scale correlation coefficients
Bob Zhang 0001, Xiangqian Wu 0002, Jane You, Qin Li 0001, Fakhri Karray
Pattern Recognit.2
2010 Retinopathy Online Challenge: Automatic Detection of Microaneurysms in Digital Color Fundus Photographs
abstract
The detection of microaneurysms in digital color fundus photographs is a critical first step in automated screening for diabetic retinopathy (DR), a common complication of diabetes. To accomplish this detection numerous methods have been published in the past but none of these was compared with each other on the same data. In this work we present the results of the first international microaneurysm detection competition, organized in the context of the Retinopathy Online Challenge (ROC), a multiyear online competition for various aspects of DR detection. For this competition, we compare the results of five different methods, produced by five different teams of researchers on the same set of data. The evaluation was performed in a uniform manner using an algorithm presented in this work. The set of data used for the competition consisted of 50 training images with available reference standard and 50 test images where the reference standard was withheld by the organizers (M. Niemeijer, B. van Ginneken, and M. D. Abràmoff). The results obtained on the test data was submitted through a website after which standardized evaluation software was used to determine the performance of each of the methods. A human expert detected microaneurysms in the test set to allow comparison with the performance of the automatic methods. The overall results show that microaneurysm detection is a challenging task for both the automatic methods as well as the human expert. There is room for improvement as the best performing system does not reach the performance of the human expert. The data associated with the ROC microaneurysm detection competition will remain publicly available and the website will continue accepting submissions.
Meindert Niemeijer, Bram van Ginneken, Michael J. Cree, Atsushi Mizutani, Gwenolé Quellec, Clara I. Sánchez, Bob Zhang 0001, Roberto Hornero, Mathieu Lamard, Chisako Muramatsu, Xiangqian Wu 0002, Guy Cazuguel, Jane You, Agustín Mayo, Qin Li 0001, Yuji Hatanaka, Béatrice Cochener, Christian Roux, Fakhri Karray, María García, Hiroshi Fujita 0001, Michael D. Abràmoff
IEEE Trans. Medical Imaging11
2009 Differential Feature Analysis for Palmprint Authentication
Xiangqian Wu 0002, Kuanquan Wang, Yong Xu 0001, David Zhang 0001
CAIP1
2009 Biometric cryptographic key generation based on city block distance
abstract
Information security is becoming increasingly important in our information driven society. Cryptography is one of the most effective ways to enhance information security. Biometrics based cryptographic key generation techniques, in which biometric features are used to generate cryptographic keys, have been developed to overcome the shortages of the traditional cryptographic methods. An essential issue of biometric cryptographic key generation is to remove the variance between biometric templates of genuine users. In previous works, error correction techniques are used to eliminate these variances. However, these techniques can only be used to remove errors in Hamming metric whereas many biometric templates are real valued vectors and cannot use Hamming distance to measure the similarity, which means that the error correction techniques can not be directly used to remove the variance between these biometric templates. In this paper, we proposed a novel biometric cryptographic framework based on city block distance. In the proposed framework, the real valued biometric feature vector is firstly quantized and then encoded into a binary string in such way that the city block distance between two feature vectors is converted to Hamming distance between two binary strings. After that, the error correction techniques are used to eliminate the errors between the strings of the genuine users. Finally, the error free string is hashed to form a cryptographic key. The experimental results conducted on face and palmprint biometrics demonstrate the effectiveness of the proposed framework.
Xiangqian Wu 0002, Kuanquan Wang, Yong Xu 0001
WACV1
2008 Tongue line extraction
abstract
Tongue line refers to the surface of the tongue covered with fissures or lines in deep or shallow shape and is one type of important features in clinical practice of Traditional Chinese Tongue Diagnosis (TCTD). However, it is hard to extract tongue lines completely due to the large variation of the widths of tongue lines and the strong noise caused by the rough surface of tongue and uneven illumination. In this paper, an improved wide line detector (WLD) is presented for tongue line extraction. Based on the characteristics of tongue lines, the original WLD is improved to avoid the undesired separation of a wide line and the influence of uneven lighting conditions. The proposed method has been tested on a total of 286 tongue line images and our experimental results demonstrate that the improved WLD significantly outperforms the original WLD for tongue line extraction by improving the TPR 16.5%, FPR 44.6% and PM 33.4%, respectively.
Laura Li Liu, David Zhang 0001, Ajay Kumar 0001, Xiangqian Wu 0002
ICPR4
2008 A cryptosystem based on palmprint feature
abstract
Biometric cryptography is a technique using biometric features to encrypt data, which can improve the security of the encrypted data and overcome the shortcomings of the traditional cryptography. This paper proposes a novel biometric cryptosystem based on palmprint features. In this system, the palmprint features, called DoG code, are extracted using Gaussian derivative filters. Then the Reed-Solomon error correcting technique and the logical XOR operation are employed to encrypt and decrypt the data. Experimental results show that this system can obtain a high security with a low false rejection rate.
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
ICPR1
2007 Fusion of Palmprint and Iris for Personal Authentication
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang, Ning Qi
ADMA1
2007 Recognize a Special Structure in Palmprint for Palm Medicine
abstract
Palm medicine is an important part of Tradition Chinese Medicine (TCM), which has been widely practiced in China and southeast country of Asia. Palmprint is composed of many lines and some special structures which imply a number of diseases. In this paper a fuzzy approach is proposed to recognize one of special structures in palmprint which is a key process in automated palm diagnosis system. Firstly, a palm image is preprocessed and all palmprint lines are extracted. Secondly, the extracted palm-lines are transformed to an undirected graph according to the connection of the points on the palm-lines. Thirdly, three features are extracted from this graph and their membership functions are defined. Finally, these three features are utilized to recognize one special structure which is called mi structure. Applying our approach to 200 palmprint images, the experimental results are encouraging.
Kuanquan Wang, Jing Liao 0013, Xiangqian Wu 0002, Henggui Zhang
CBMS3
2007 Automated Personal Authentication Using Both Palmprints
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
ICEC1
2006 Fusion of phase and orientation information for palmprint authentication
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang
Pattern Anal. Appl.1
2006 Palm line extraction and matching for personal authentication
abstract
The palm print is a new and emerging biometric feature for personal recognition. The stable line features or "palm lines", which are comprised of principal lines and wrinkles, can be used to clearly describe a palm print and can be extracted in low-resolution images. This paper presents a novel approach to palm line extraction and matching for use in personal authentication. To extract palm lines, a set of directional line detectors is devised, and then these detectors are used to extract these lines in different directions. To avoid losing the details of the palm line structure, these irregular lines are represented using their chain code. To match palm lines, a matching score is defined between two palm prints according to the points of their palm lines. The experimental results show that the proposed approach can effectively discriminate between palm prints even when the palm prints are dirty. The storage and speed of the proposed approach can satisfy the requirements of a real-time biometric system
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang
IEEE Trans. Syst. Man Cybern. Part A1
2005 Fusion of the Textural Feature and Palm-Lines for Palmprint Authentication
Xiangqian Wu 0002, Fengmiao Zhang, Kuanquan Wang, David Zhang 0001
ICIC (1)1
2005 Fusion of phase and orientation information for palmprint authentication
abstract
This paper presents a novel approach of palmprint authentication based on the fusion of the phase and orientation information. This approach is an improvement of a previous palmprint recognition method - fusioncode method (A. Kong and D. Zhang, 2004). In the proposed approach, the phase information (fusion-code) of a palmprint is extracted by using four 2-D Gabor filters with different orientations, and at the same time, the orientation information (called orientationcode) of the palmprint is also extracted. The fusioncode and the orientationcode are fused to make a new feature, called the palmprint phase orientation code (PPOC). At the matching stage, a modified Hamming distance is defined to measure the similarity of two PPOCs. This approach is tested on a palmprint database containing 7605 samples and the experimental results show that the PPOC approach greatly improves the performance of the fusioncode method.
Xiangqian Wu 0002, Kuanquan Wang, Fengmiao Zhang, David Zhang 0001
ICIP (2)1
2005 Wavelet Energy Feature Extraction and Matching for Palmprint Recognition
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
J. Comput. Sci. Technol.1
2004 A novel approach of palm-line extraction
abstract
Palm-lines, including the principal lines and wrinkles, can describe a palmprint clearly. This paper presents a novel approach of palm-line extraction for the online palmprints. This approach is composed of two stages: coarse-level extraction stage and fine-level extraction stage. In the first stage, morphological operations are used to extract palm-lines in different directions. In the second stage, for each extracted line, a recursive process is devised to further extract and trace the palm-line using the local information of the extracted part. Experimental results show that the proposed approach is suitable for palm-line extraction.
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
ICIG1
2004 Palmprint classification using principal lines
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang, Bo Huang 0003
Pattern Recognit.1
2003 Fisherpalms based palmprint recognition
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang
Pattern Recognit. Lett.1