VLDB 2026 Research / reviewers in the wild / expert
Yongsheng Gao 0001
dblp:22/6786-1
· DBLP profile ↗
175ranked-venue papers
9as first author
76since 2021 · last 2026
0000-0002-5382-5351ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 100 · 6 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 97 · 3 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 11 since 2021Security and privacy · 3Human-computer interaction and ubiquitous computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Time in Static ClassifiersabstractReal-world visual data rarely presents as isolated, static instances. Instead, it often evolves gradually over time through variations in pose, lighting, object state, or scene context. However, conventional classifiers are typically trained under the assumption of temporal independence, limiting their ability to capture such dynamics. We propose a simple yet effective framework that equips standard feedforward classifiers with temporal reasoning, all without modifying model architectures or introducing recurrent modules. At the heart of our approach is a novel Support-Exemplar-Query (SEQ) learning paradigm, which structures training data into temporally coherent trajectories. These trajectories enable the model to learn class-specific temporal prototypes and align prediction sequences via a differentiable soft-DTW loss. A multi-term objective further promotes semantic consistency and temporal smoothness. By interpreting input sequences as evolving feature trajectories, our method introduces a strong temporal inductive bias through loss design alone. This proves highly effective in both static and temporal tasks: it enhances performance on fine-grained and ultra-fine-grained image classification, and delivers precise, temporally consistent predictions in video anomaly detection. Despite its simplicity, our approach bridges static and temporal learning in a modular and data-efficient manner, requiring only a simple classifier on top of pre-extracted features. Xi Ding 0001, Lei Wang 0108, Piotr Koniusz, Yongsheng Gao 0001 |
AAAI | 4 |
| 2026 | Universal Facial Landmark Detection by Landmark-Clustering Relation-Reasoning Transformer
Jun Wan 0005, Yuanzhi Yao, Jiaxing Huang 0001, Xiaoying Ding, Lefei Zhang, Yongsheng Gao 0001, Dacheng Tao |
Int. J. Comput. Vis. | 6 |
| 2026 | Neuron Abandoning Attention Flow: Visual Explanation of Dynamics Inside CNN ModelsabstractIn this paper, we present a Neuron Abandoning Attention Flow (NAFlow) method to address the unsolved problem of visually explaining the attention evolution dynamics inside CNNs when making their classification decisions. A novel cascading neuron abandoning back-propagation algorithm is designed to precisely exclude the abandoned neurons on all intermediate layers inside a CNN model for the first time. Firstly, a Neuron Abandoning Back-Propagation module is proposed to generate Back-Propagation Feature Maps (BPFM) by using inverse function of the intermediate layers of CNN models, on which the neurons not used for decision-making are removed. Meanwhile, the cascading NA-BP modules calculate the tensors of importance coefficients which are linearly combined with the tensors of BPFMs to form the NAFlow. Secondly, to be able to visualize attention flow for similarity metric-based CNN models, a new channel contribution weights module is proposed to calculate the importance coefficients via Jacobian Matrix. Extensive evaluations demonstrate the effectiveness of the proposed NAFlow across eleven widely-used CNN models for various tasks of general image classification, contrastive learning classification, few-shot image classification, and image retrieval. Yongsheng Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A unified analysis on cross-architecture generalizability of coresetsabstractCoreset selection methods aim to identify a representative subset of training data that preserves competitive performance. However, mainstream coreset selection approaches are model-specific and assume they already have full information about the target model when the coreset is selected. This largely restricts the usefulness of coreset selection in practice. This work aims to fill that gap by formulating and investigating the problem of cross-architecture generalizability of coresets: we develop a unified theoretical framework that analyzes the upper bound of coreset selection objective functions, extend it to scenarios involving multiple downstream architectures, and provide an empirical analysis on cross-architecture coreset performance. Based on our findings, we propose a novel ensemble scoring method that aggregates multi-source knowledge to enhance cross-architecture generalizability. Our extensive experiments across thirteen architectures and six selection ratios provide comprehensive verification of our theoretical analysis. The source code is available at https://github.com/diqichen91/CACS.git . Diqi Chen, Jiajun Liu 0004, Frank de Hoog, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 6 |
| 2026 | DBCore: Shaping generalizable decision boundaries for coreset selectionabstractCoreset selection for classification often relies on assessing individual sample difficulty or importance, leading to sample-wise or range-based selection, but this can overlook the collective impact on model decision boundaries. Realizing that the representative power a coreset possesses is tightly associated with the decision boundaries a model can form on it, we propose a novel approach that directly optimizes the Decision Boundary (DB) formed by the selected coreset. Specifically, we ask: How can we collectively select samples to create a DB that is globally smoothed yet locally detailed, ensuring maximum generalizability and noise-resilience to the original dataset? To address this, we define two key objectives: (1) Global shape retention – The selected coreset should form a smoothed version of the original DB, preserving its overall structure and preventing overfitting; (2) Local detail preservation – While smoothing prevents overfitting, excessive smoothing risks losing critical nuances. Thus, the selection must also retain key points near the original DB to capture local complexities. We formulate these objectives as a convex quadratic optimization problem with linear constraints and solve it efficiently. Extensive evaluations demonstrate the consistent and substantial advantages of our method over the state-of-the-art coreset selection strategies. The source code is available at https://github.com/diqichen91/DBCore.git . Diqi Chen, Jiajun Liu 0004, Frank de Hoog, Wangzhi Xing, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 7 |
| 2026 | On learning denoisable student logitsabstractKnowledge Distillation (KD) aims to train a student model to mimic the behavior of a more powerful teacher model. In this paper, we reveal that through the lens of diffusion processes, student logits can be statistically treated as a noisy version of teacher logits, and KD helps reduce the noise level of student logits. This insight motivates us to design a framework leveraging KD to produce denoisable student logits that can be further recovered towards teacher logits via a reverse diffusion process. A key advantage of this approach is that the inference-diffusion process can occur in two physical locations and on separate devices, enabling a two-step and distributed inference process. The experimental results show that the derived denoisable student logits achieve comparable or even superior performance to standard KD’s, and the reverse diffusion process achieves a substantial improvement in accuracy, without needing the original image, thus preserving the privacy and security of the original data. Additionally, the logits can be further compressed before transmission, reducing the required bandwidth while achieving comparable overall performance. Diqi Chen, Yang Li 0184, Jiajun Liu 0004, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 6 |
| 2026 | SATE: Efficient knowledge distillation with implicit student-aware teacher ensemblesabstractRecent findings suggest that with the same teacher architecture, a fully converged or “stronger” checkpoint surprisingly leads to a worse student. This can be explained by the Information Bottleneck (IB) principle, as the features of a weaker teacher transfer more “dark” knowledge because they maintain higher mutual information with the inputs. Meanwhile, various works have shown that severe teacher-student structural disparity or capability mismatch often leads to worse student performance. To deal with these issues, we propose a generalizable and efficient Knowledge Distillation (KD) framework with implicit Student-Aware Teacher Ensembles (SATE). The SATE framework simultaneously trains a student network and a student-aware intermediate teacher as a learning companion. With the proposed co-training strategy, the intermediate teacher is trained gradually and forms implicit ensembles of weaker teachers along the learning process. Such a design enables the student model to retain more dark knowledge for better generalization ability. The proposed framework improves the training scheme in a plug-and-play way so that it can be applied to improve various classic and state-of-the-art KD methods on both intra-domain (up to 2.184 % ) and cross-domain (up to 7.358 % ) settings, under a diversified configurations on teacher-student architectures, and achieves a major efficient advantage over other generic frameworks. The code is available at https://github.com/diqichen91/SATE.git . Diqi Chen, Yang Li 0184, Jiajun Liu 0004, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 5 |
| 2026 | Adaptive feature selection-based feature reconstruction network for few-shot learning
Yaohui An, Tao Lei 0003, Junpo Yang, Zicheng Pan, Yongsheng Gao 0001, Changming Sun |
Pattern Recognit. | 8 |
| 2026 | SpectralKAN: Weighted Activation Distribution Kolmogorov-Arnold Network for Hyperspectral Image Change Detection
Xiaohan Yu 0001, Yongsheng Gao 0001, Jianjun Sha, Jian Wang 0138, Shiyong Yan, Yonggang Zhang 0001, Lianru Gao |
Pattern Recognit. | 3 |
| 2026 | Distance Learning-Based Prototypical Network With Multi-Domain Adaptation for Few-Shot Hyperspectral Medical Image ClassificationabstractHyperspectral imaging (HSI) holds immense potential for medical diagnostics by capturing tissue-specific spectral signatures that facilitate precise disease detection. However, effective HSI classification in clinical settings is hindered by two main challenges: (i) the severe lack of labelled medical HSI samples constrains model training. Prototypical networks, as a few-shot learning paradigm, have been adopted to address label scarcity. However, current Euclidean-based prototypical methods typically assume equal feature variance and spherical distributions, while ignoring intraclass covariance and spectral correlations; (ii) significant domain shifts across heterogeneous medical HSI datasets undermine model generalisation, impair multi-domain interpretability, and force expensive per-dataset retraining. To overcome these limitations, we propose a novel distance-learning-based prototypical network with multi-domain adaptation for few-shot hyperspectral medical image classification. First, by embedding a class-covariance-aware Mahalanobis metric within the prototypical block, our module adapts similarity measures to each class's intrinsic spectral-spatial covariance and scale variations, thereby enhancing prototype robustness under severe label scarcity and significantly reducing misclassification compared with existing few-shot networks. Secondly, we introduce the domain-aware adapter block designed to address domain shift and multi-domain variability by dynamically fusing shared spectral-spatial representations with domain-specific characteristics via spectral integration and switchable adapters. We undertook extensive experiments on three publicly available hyperspectral medical datasets: skin dermoscopy, multidimensional choledochal, and in-vivo brain dataset. Compared to state-of-the-art classifiers, the proposed method achieved excellent performance on all three datasets, paving the way for generalisable HSI solutions in clinical workflows and biomedical research. Favour Ekong, Jun Zhou 0001, Jing Wang 0062, Yongsheng Gao 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision TransformersabstractOver the past few years, vision transformers (ViTs) have consistently demonstrated remarkable performance across various visual recognition tasks. However, attempts to enhance their robustness have yielded limited success, mainly focusing on different training strategies, input patch augmentation, or network structural enhancements. These approaches often involve extensive training and fine-tuning, which are time-consuming and resource-intensive. To tackle these obstacles, we introduce a novel approach named Spatial Autocorrelation Token Analysis (SATA). By harnessing spatial relationships between token features, SATA enhances both the representational capacity and robustness of ViT models. This is achieved through the analysis and grouping of tokens according to their spatial autocorrelation scores prior to their input into the Feed-Forward Network (FFN) block of the self-attention mechanism. Importantly, SATA seamlessly integrates into existing pre-trained ViT baselines without requiring retraining or additional fine-tuning, while concurrently improving efficiency by reducing the computational load of the FFN units. Experimental results show that the baseline ViTs enhanced with SATA not only achieve a new state-of-the-art top-1 accuracy on ImageNet-1K image classification (94.9%) but also establish new state-of-the-art performance across multiple robustness benchmarks, including ImageNet-A (top-1=63.6%), ImageNet-R (top-1=79.2%), and ImageNet-C (mCE=13.6%), all without requiring additional training or fine-tuning of baseline models. Availability: https://github.com/nick-nikzad/SATA Nick Nikzad, Yongsheng Gao 0001, Jun Zhou 0001 |
CVPR | 3 |
| 2025 | Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained ModelsabstractContinual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in general continual learning, C-UFG remains an underexplored issue. This work establishes the first comprehensive C-UFG benchmark using massive soy leaf data. By analyzing recent pre-trained model (PTM) based continual learning methods on the proposed benchmark, we propose two simple yet effective PTM-based methods to boost the performance of VC-UFG and HC-UFG, respectively. On top of those, we integrate the two methods into one unified framework and propose the first unified model, Unic, that is capable of tackling the C-UFG problem where VC-UFG and HC-UFG co-exist in a single continual learning sequence. To understand the effectiveness of the proposed methods, we first evaluate the models on VC-UFG and HC-UFG challenges and then test the proposed Unic on a unified C-UFG challenge. Experimental results demonstrate the proposed methods achieve superior performance for C-UFG. The code is available at https://github.com/PatrickZad/unicufg. Pengcheng Zhang 0003, Xiaohan Yu 0001, Meiying Gu, Yongsheng Gao 0001, Xiao Bai 0001 |
IJCAI | 5 |
| 2025 | View-aware Decomposition and Unification for Fast Ground-to-Aerial Person SearchabstractGround-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall person search performance when training the model in a view-agnostic way. To address this, we propose a view-aware decomposition and unification (VADU) framework for ground-to-aerial person search. Specifically, we decompose the person search model to learn view-oriented modules for image feature encoding and person proposal generation. The data sampling and retrieval feature learning are also composed to cope with the decomposed model. This decomposition improves both person detection and discriminative feature learning within each view. On top of the decomposition, we propose view-aware unification to produce unified cross-view person features. Cross-view prototypical contrastive learning is introduced to enhance the unification between different views, enhancing model robustness to retrieve a target person in cameras of a different view. As the decomposed parts of the model are deployed on different devices for inference, this overall framework adds no extra computation cost in real-world applications. Extensive experiments demonstrate that the proposed method achieves superior person search performance and guarantees the efficiency of inference. The source code is available at https://github.com/QFWang-11/vadu. Qifei Wang, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Yongsheng Gao 0001 |
IROS | 5 |
| 2025 | Contrastive Lie Algebra Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself via contrastive learning has demonstrated great progress towards learning discriminative representation. Yet forcingly dividing highly similar categories at the representation level may over-guide the learned feature space, leading to overfitting in the ultra-FGVC tasks. To this end, this paper introduces CLA-Net, a novel contrastive Lie algebra learning framework to address this fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images at a second-order representation level. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLA-Net demonstrates strong performance on eight publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at: https://github.com/zichengpan/CLA-NET. Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0002, Qin Zhang 0011, Yongsheng Gao 0001 |
ACM Multimedia | 5 |
| 2025 | Graph Your Own PromptabstractWe propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of *self-prompting*, GCR enables the model to refine its internal structure using its own outputs. While deep networks learn rich representations, these often capture noisy inter-class similarities that contradict the model's predicted semantics. GCR addresses this issue by introducing parameter-free *Graph Consistency Layers* (GCLs) at arbitrary depths. Each GCL builds a batch-level feature similarity graph and aligns it with a global, class-aware masked prediction graph, derived by modulating softmax prediction similarities with intra-class indicators. This alignment enforces that feature-level relationships reflect class-consistent prediction behavior, acting as a *semantic regularizer* throughout the network. Unlike prior work, GCR introduces a multi-layer, cross-space graph alignment mechanism with adaptive weighting, where layer importance is learned from graph discrepancy magnitudes. This allows the model to prioritize semantically reliable layers and suppress noisy ones, enhancing feature quality without modifying the architecture or training procedure. GCR is model-agnostic, lightweight, and improves semantic structure across various networks and datasets. Experiments show that GCR promotes cleaner feature structure, stronger intra-class cohesion, and improved generalization, offering a new perspective on learning from prediction structure. Xi Ding 0001, Lei Wang 0108, Piotr Koniusz, Yongsheng Gao 0001 |
NeurIPS | 4 |
| 2025 | Pixel-Wise Shuffling with Collaborative Sparsity for Melanoma Hyperspectral Image ClassificationabstractHyperspectral imaging has emerged as a promising technology for medical image classification, particularly in skin cancer diagnosis. However, current methods face significant challenges in accurately and robustly classifying non-cancerous skin lesions, especially when melanoma lesions overlap with pigmented regions. Existing methods also lack sensitivity to spectral variations and accumulate excess redundant data, leading to inefficiencies, misclassifications, and overfitting while struggling to integrate spatial and spectral information effectively. To overcome these chal-lenges, we propose a novel method featuring collaborative sparse unmixing and an advanced pixel-wise shuffling approach with inter-similarity hybrid attention, aiming to improve the accuracy of skin cancer diagnosis in real-world scenarios. Experiments are conducted on a publicly available histology-verified dataset to evaluate the efficacy of the proposed method. The experimental results demonstrate that the proposed method can accurately classify melanoma lesions, even in cases where the lesions overlap with pig-mented regions. The findings indicate that the proposed method outperforms state-of-the-art methods by obtaining an overall accuracy of 73.34%, even when limited to 20% of the training data. The proposed approach has the potential to be a valuable tool for improving the diagnostic accuracy of skin cancer in clinical practice. Favour Ekong, Jun Zhou 0001, Kwabena Sarpong, Yongsheng Gao 0001 |
WACV | 4 |
| 2025 | A Generative Pretrained Transformer for Semi-Supervised Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSIs-CD) often faces the challenge of limited sample sizes, and labeling data is both time-consuming and labor-intensive. Foundation models leverage extensive unlabeled data for self-supervised generative pre-training, allowing the model to learn rich data representations. However, few models have been specifically designed for HSIs, and existing methods often rely on pre-training datasets that are limited to data from a small number of satellite sensors. This limitation affects generalization, especially when there are significant differences between data from different sensors. Moreover, the difference map (DMP) of bi-temporal HSIs is often used as input to the networks. While the DMP-based approach reduces FLOPs, it may lead to information loss compared to dual-branch networks. In this letter, we propose a mini-patch-based generative pre-trained spectral-spatial transformer (GPSST) for semi-supervised HSIs-CD. We begin by collecting public HSIs datasets and dividing them into thousands of patches. Each patch is then split into spectral-spatial tokens, with a portion of these tokens masked and used as input for the GPSST. We then design a spectral-spatial masked autoencoder (MAE) as the backbone of GPSST for self-supervised generative learning. Finally, we fine-tune the GPSST encoder using a small number of labeled patches and design a principal component analysis (PCA) branch to compensate for the information loss caused by the DMP. Our experiments demonstrate that GPSST outperforms existing methods, achieving superior accuracy in HSIs-CD. Jianjun Sha, Xiaohan Yu 0001, Yongsheng Gao 0001, Yonggang Zhang 0001, Xianhui Rong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | CoPISan: Contrastive Perceptual Inference and Sanity Checks for Concept-Based CNN ExplanationsabstractDespite the effectiveness of convolutional neural networks (CNNs) in visual categorization, the logic behind their predictions is not human-understandable. While existing concept-based explainability methods reveal what a CNN sees, there is a need to understand how a specific concept is chosen (rather than another concept) for a prediction, aligning more closely with human perception. To address this challenge, we propose a novel contrastive paradigm to bridge the critical gap in global concept discovery by leveraging contrasts from cognitive sciences for discriminative concept retrieval. A new multiple-case concept retrieval method is proposed for improved local understanding of (dis)similar classification cases. We argue that a contrastive paradigm for concept retrieval and sanity checks is essential to an explainer's trustworthiness and integrate these missing ingredients into state-of-the-art concept-based explanation frameworks to foster a better human understanding through contrast. The proposed Contrastive Perceptual Inference and Sanity Checks for Concept-based CNN Explanations (CoPISan) framework accelerates salient concept retrieval. It evaluates explainer trustworthiness via sanity checks conducted under Frontdoor and Poisoning adversarial attacks. Experimental results demonstrate CoPISan's encouraging performance, mitigating issues related to duplication, entanglement, diminishing returns, and ambiguity of concept explanations. CoPISan is motivated by cognition and perception, offers theoretical justification and resilience, and is computationally efficient. Ugochukwu Ejike Akpudo, Yongsheng Gao 0001, Jun Zhou 0001, Andrew Lewis 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Hierarchical Context Learning of object components for unsupervised semantic segmentationabstractUnsupervised Semantic Segmentation (USS) aims to learn semantically rich and dense representations without relying on labels. Recent advances in self-supervised learning have demonstrated the potential of pretrained vision transformers to capture patch-level semantic information, offering a promising direction to USS. However, existing methods face challenges in constructing a discriminative spatial token embedding space that consistently and effectively represents the well-structured semantic relationships among object components. Inspired by Edwin Hancock’s pioneer work on hierarchical pattern analysis, we highlight the critical role of hierarchical context to overcome this limitation. By modeling spatial relationships at multiple levels of granularity, hierarchical context helps align related object parts while distinguishing them across semantic groups. Based on this insight, we introduce Hierarchical Context Learning (HCL), a novel approach for USS that enhances semantic consistency by integrating hierarchical context. HCL incorporates a novel parallel multi-level vision transformer backbone to aggregate multi-level contextual information into object component tokens. To uncover the semantic structure of objects, we propose Momentum-based Global Foreground–Background Clustering (MoGoClustering) to cluster object components into coherent semantic groups and then calculate their semantic centroids. To enforce intra-group semantic consistency and maximize inter-group separation across spatial scales, we design a foreground–background-aware contrastive loss based on MoGoClustering. Our method achieves state-of-the-art performance on the COCO-Stuff and Pascal VOC datasets, demonstrating its ability to learn robust, context-aware, and discriminative object component semantics for USS. The code is available at: https://github.com/dbaofd/HCL . Dong Bao, Jun Zhou 0001, Gervase Tuxworth, Jue Zhang 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 5 |
| 2025 | Dynamic accumulated attention map for interpreting evolution of decision-making in vision transformer
Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2025 | Overcoming learning bias via Prototypical Feature Compensation for source-free domain adaptationabstractThe focus of Source-free Unsupervised Domain Adaptation (SFUDA) is to effectively transfer a well-trained model from the source domain to an unlabelled target domain. During the target domain adaptation , the source domain data is no longer accessible. Prevalent methodologies attempt to synchronize the data distributions between the source and target domains, utilizing pseudo-labels to impart categorical information, which has made some progress in improving the model’s performance. However, performance impairments persist due to the introduction of learning bias from the source model and the impact of noisy pseudo-labels generated for the target domain. In this research, we reveal that the central cause for feature misalignment during domain transition is the learning bias, which is generated by the discrepancy of information between source and target domain data . The source domain data may contain distinguishable features that do not appear on the target domain, which causes the pre-trained source model to fail to work during domain adaptation. To overcome the information discrepancy, we propose a Prototypical Feature Compensation (PFC) Network. The network extracts representative feature maps of the source domain. Then use them to minimize the discrepancy information in the target domain feature maps. This mechanism facilitates feature alignment across different domains, allowing the model to generate more accurate categorical data through pseudo-labelling. The experimental results and ablation studies demonstrate exceptional performance on three SFUDA datasets and provide evidence of the proposed PFC method’s ability to adjust the feature distribution of both source and target domain data, ensuring their overlap in the latent space. Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 4 |
| 2025 | UBSTrack: Unified Band Selection and Multimodel Ensemble for Hyperspectral Object TrackingabstractHyperspectral object tracking is notably challenging due to the high-dimensional nature of the data and the necessity of seamlessly integrating spectral, spatial and temporal information. Traditional methods often emphasize detection-based or tracking-based networks, each leveraging their inherent strengths but overlooking the potential advantages of a combined approach, leading to suboptimal performance in complex, real-world scenarios. Furthermore, this challenge is amplified by the variability of spectral bands across datasets, making the maintenance of consistent tracking performance complicated. To address these issues, we propose a novel, unified approach that merges adaptive band selection with a multi-model ensemble strategy. We introduce a local and global attention-based unified band selection (UBS) technique that identifies the most informative three bands from any dataset, significantly reducing data complexity while preserving critical spectral and spatial information. This UBS method employs spectral independence, allowing it to process hyperspectral video frames with any number of bands as input, ultimately generating a three-band pseudocolor image. This is coupled with a multi-model ensemble framework, utilizing a local and global attention-based appearance module. The module selects the optimal candidate by computing the similarity between the proposals generated by the base models and historical frames. Experimental results show that our approach, UBSTrack, achieves state-of-the-art performance, delivering robust and accurate tracking under different real-world challenging scenarios. The code of UBSTrack is available at the following link: source code. Jun Zhou 0001, Wangzhi Xing, Yongsheng Gao 0001, Kuldip K. Paliwal |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Session-Guided Attention in Continuous Learning With Few SamplesabstractFew-shot class-incremental learning (FSCIL) aims to learn from a sequence of incremental data sessions with a limited number of samples in each class. The main issues it encounters are the risk of forgetting previously learned data when introducing new data classes, as well as not being able to adapt the old model to new data due to limited training samples. Existing state-of-the-art solutions normally utilize pre-trained models with fixed backbone parameters to avoid forgetting old knowledge. While this strategy preserves previously learned features, the fixed nature of the backbone limits the model's ability to learn optimal representations for unseen classes, which compromises performance on new class increments. In this paper, we propose a novel SEssion-Guided Attention framework (SEGA) to tackle this challenge. SEGA exploits the class relationships within each incremental session by assessing how test samples relate to class prototypes. This allows accurate incremental session identification for test data, leading to more precise classifications. In addition, an attention module is introduced for each incremental session to further utilize the feature from the fixed backbone. As the session of the testing image is determined, we can fine-tune the feature with the corresponding attention module to better cluster the sample within the selected session. Our approach adopts the fixed backbone strategy to avoid forgetting the old knowledge while achieving novel data adaptation. Experimental results on three FSCIL datasets consistently demonstrate the superior adaptability of the proposed SEGA framework in FSCIL tasks. The code is available at: https://github.com/zichengpan/SEGA. Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | TraNCE: Transformative Nonlinear Concept Explainer for CNNsabstractConvolutional neural networks (CNNs) have succeeded remarkably in various computer vision tasks. However, they are not intrinsically explainable. While feature-level understanding of CNNs reveals where the models looked, concept-based explainability methods provide insights into what the models saw. However, their assumption of linear reconstructability of image activations fails to capture the intricate relationships within these activations. Their fidelity-only approach to evaluating global explanations also presents a new concern. For the first time, we address these limitations with the novel transformative nonlinear concept explainer (TraNCE) for CNNs. Unlike linear reconstruction assumptions made by existing methods, TraNCE captures the intricate relationships within the activations. This study presents three original contributions to the CNN explainability literature: 1) an automatic concept discovery mechanism based on variational autoencoders (VAEs). This transformative concept discovery process enhances the identification of meaningful concepts from image activations; 2) a visualization module that leverages the Bessel function to create a smooth transition between prototypical image pixels, revealing not only what the CNN saw but also what the CNN avoided, thereby mitigating the challenges of concept duplication as documented in previous works; and 3) a new metric, the faith score, integrates both coherence and fidelity for comprehensive evaluation of explainer faithfulness and consistency. Based on the investigations on publicly available datasets, we prove that a valid decomposition of a high-dimensional image activation should follow a nonlinear reconstruction, contributing to the explainer's efficiency. We also demonstrate quantitatively that, besides accuracy, consistency is crucial for the meaningfulness of concepts and human trust. The code is available at https://github.com/daslimo/TrANCE. Ugochukwu Ejike Akpudo, Yongsheng Gao 0001, Jun Zhou 0001, Andrew Lewis 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | DyCR: A Dynamic Clustering and Recovering Network for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims to continually learn novel data with limited samples. One of the major challenges is the catastrophic forgetting problem of old knowledge while training the model on new data. To alleviate this problem, recent state-of-the-art methods adopt a well-trained static network with fixed parameters at incremental learning stages to maintain old knowledge. These methods suffer from the poor adaptation of the old model with new knowledge. In this work, a dynamic clustering and recovering network (DyCR) is proposed to tackle the adaptation problem and effectively mitigate the forgetting phenomena on FSCIL tasks. Unlike static FSCIL methods, the proposed DyCR network is dynamic and trainable during the incremental learning stages, which makes the network capable of learning new features and better adapting to novel data. To address the forgetting problem and improve the model performance, a novel orthogonal decomposition mechanism is developed to split the feature embeddings into context and category information. The context part is preserved and utilized to recover old class features in future incremental learning stages, which can mitigate the forgetting problem with a much smaller size of data than saving the raw exemplars. The category part is used to optimize the feature embedding space by moving different classes of samples far apart and squeezing the sample distances within the same classes during the training stage. Experiments show that the DyCR network outperforms existing methods on four benchmark datasets. The code is available at: https://github.com/zichengpan/DyCR. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Self-Supervised Lie Algebra Representation Learning via Optimal Canonical MetricabstractLearning discriminative representation with limited training samples is emerging as an important yet challenging visual categorization task. While prior work has shown that incorporating self-supervised learning can improve performance, we found that the direct use of canonical metric in a Lie group is theoretically incorrect. In this article, we prove that a valid optimization measurement should be a canonical metric on Lie algebra. Based on the theoretical finding, this article introduces a novel self-supervised Lie algebra network (SLA-Net) representation learning framework. Via minimizing canonical metric distance between target and predicted Lie algebra representation within a computationally convenient vector space, SLA-Net avoids computing nontrivial geodesic (locally length-minimizing curve) metric on a manifold (curved space). By simultaneously optimizing a single set of parameters shared by self-supervised learning and supervised classification, the proposed SLA-Net gains improved generalization capability. Comprehensive evaluation results on eight public datasets show the effectiveness of SLA-Net for visual categorization with limited samples. Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0019, Yongsheng Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | EIANet: A Novel Domain Adaptation Approach to Maximize Class Distinction with Neural Collapse Principles
Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
BMVC | 3 |
| 2024 | Coherentice: Invertible Concept-Based Explainability Framework for CNNs beyond FidelityabstractIn their natural form, convolutional neural networks (CNNs) lack interpretability despite their effectiveness in visual categorization. Concept activation vectors (CAVs) offer human-interpretable quantitative explainability, utilizing feature maps from intermediate layers of CNNs. Current concept-based explainability methods assess explainer faithfulness primarily through Fidelity. However, relying solely on this metric has limitations. This study extends the Invertible Concept-based Explainer (ICE) to introduce a new ingredient measuring concept consistency. We propose the CoherentICE explainability framework for CNNs, expanding beyond Fidelity. Our analysis, for the first time, highlights that Coherence provides a more reliable faithfulness evaluation for CNNs, supported by empirical validations. Our findings emphasize that accurate concepts are meaningful only when consistently accurate and improve at deeper CNN layers. Ugochukwu Ejike Akpudo, Yongsheng Gao 0001, Jun Zhou 0001, Andrew Lewis 0004 |
ICME | 2 |
| 2024 | Multiscale Binary-Pattern Dependency: A Novel Co-Occurrence Texture Descriptor for Fine-Grained Leaf Image RetrievalabstractIn the research community of content-based image retrieval, great success has been achieved for leaf image retrieval in species. However, little progress has been made on the more challenging fine-grained leaf image retrieval (FGLIR) which focuses on subspecies/cultivars recognition. To address it, a novel co-occurrence local binary pattern (CoLBP), named Multi-scale Binary-Pattern Dependency (MBPD), is proposed in this study. Despite the potential of CoLBP in encoding contextual information among texture patterns, how to correlate LBPs to yield discriminative co-occurrence features remains an open issue. We introduce two new concepts, Axisymmetric Co-occurrence (ACO) and Cross-thresholding (CRT), into the design of CoLBP. The ACO produces CoLBP by sliding a pair of axisymmetric lines over an adaptive local patch to capture spatial axisymmetric relationship among LBPs. While the CRT correlates LBPs in intensity domain through exchanging their respective thresholds. Their combination is used to yield ACO-CRT local descriptors to encode the dependency information in both spatial and intensity domains. The ACO-CRT local descriptors of multi-scale and multi-position are aggregated into MBPD representation for efficient dissimilarity measure between two leaf images. Extensive experiments validate the superior performance of our method over the state-of-the-arts on FGLIR. Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
ICME | 3 |
| 2024 | SDePR: Fine-Grained Leaf Image Retrieval with Structural Deep Patch RepresentationabstractFine-grained leaf image retrieval (FGLIR) is a new unsupervised pattern recognition task in content-based image retrieval (CBIR). It aims to distinguish varieties/cultivars of leaf images within a certain plant species and is more challenging than general leaf image retrieval task due to the inherently subtle differences across different cultivars. In this study, we for the first time investigate the possible way to mine the spatial structure and contextual information from the activation of the convolutional layers of CNN networks for FGLIR. For achieving this goal, we design a novel geometrical structure, named Triplet Patch-Pairs Composite Structure (TPCS), consisting of three symmetric patch pairs segmented from the leaf images in different orientations. We extract CNN feature map for each patch in TPCS and measure the difference between the feature maps of the patch pair for constructing local deep self-similarity descriptor. By varying the size of the TPCS, we can yield multi-scale deep self-similarity descriptors. The final aggregated local deep self-similarity descriptors, named Structural Deep Patch Representation (SDePR), not only encode the spatial structure and contextual information of leaf images in deep feature domain, but also are invariant to the geometrical transformations. The extensive experiments of applying our SDePR method to the public challenging FGLIR tasks show that our method outperforms the state-of-the-art handcrafted visual features and deep retrieval models. Xin Chen 0100, Bin Wang 0041, Jinzheng Jiang, Kunkun Zhang, Yongsheng Gao 0001 |
ACM Multimedia | 5 |
| 2024 | Referring Human Pose and Mask Estimation In the WildabstractWe introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Bo Miao, Mingtao Feng, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
NeurIPS | 5 |
| 2024 | Separated Fan-Beam Projection with Gaussian Convolution for Invariant and Robust Butterfly Image Retrieval
Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2024 | Pseudo-set Frequency Refinement architecture for fine-grained few-shot class-incremental learningabstractFew-shot class-incremental learning was introduced to solve the model adaptation problem for new incremental classes with only a few examples while still remaining effective for old data. Although recent state-of-the-art methods make some progress in improving system robustness on common datasets, they fail to work on fine-grained datasets where inter-class differences are small. The problem is mainly caused by: (1) the overlapping of new data and old data in the feature space during incremental learning, which means old samples can be falsely classified as newly introduced classes and induce catastrophic forgetting phenomena; (2) lacking discriminative feature learning ability to identify fine-grained objects. In this paper, a novel Pseudo-set Frequency Refinement (PFR) architecture is proposed to tackle these problems. We design a pseudo-set training strategy to mimic the incremental learning scenarios so that the model can better adapt to novel data in future incremental sessions. Furthermore, separate adaptation tasks are developed by utilizing frequency-based information to refine the original features and address the above challenging problems. More specifically, the high and low-frequency components of the images are employed to enrich the discriminative feature analysis ability and incremental learning ability of the model respectively. The refined features are used to perform inter-class and inter-set analyses. Extensive experiments show that the proposed method consistently outperforms the state-of-the-art methods on four fine-grained datasets. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
Pattern Recognit. | 5 |
| 2024 | Re-abstraction and perturbing support pair network for few-shot fine-grained image classificationabstractThe goal of few-shot fine-grained image classification (FSFGIC) is to distinguish subordinate-level categories with subtle visual differences such as the species of bird and models of car with only a few samples. In this work, we argue that a designed network that has the ability to better distinguish feature descriptors of different categories will effectively improve the performance of FSFGIC. We propose a re-abstraction and perturbing support pair network (RaPSPNet) for FSFGIC. Specifically, we first design a feature re-abstraction embedding (FRaE) module which can not only effectively amplify the difference between the feature information from different categories but also better extract the feature information from images. Furthermore, a novel perturbing support pair (PSP) based similarity measure module is designed which evaluates the relationships of feature information among a query image and two different categories of support images (a support pair) at the same time for guiding the designed FRaE module to find salient feature information from the same category of query and support images and find distinguishable feature information from the different categories of query and support images. Extensive experiments on FSFGIC tasks demonstrate the superiority of the proposed methods over state-of-the-art benchmarks. Yali Zhao, Yongsheng Gao 0001, Changming Sun |
Pattern Recognit. | 3 |
| 2024 | Temporally Consistent Referring Video Object Segmentation With Hybrid MemoryabstractReferring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects. We propose an end-to-end R-VOS paradigm that explicitly models temporal instance consistency alongside the referring segmentation. Specifically, we introduce a novel hybrid memory that facilitates inter-frame collaboration for robust spatio-temporal matching and propagation. Features of frames with automatically generated high-quality reference masks are propagated to segment the remaining frames based on multi-granularity association to achieve temporally consistent R-VOS. Furthermore, we propose a new Mask Consistency Score (MCS) metric to evaluate the temporal consistency of video segmentation. Extensive experiments demonstrate that our approach enhances temporal consistency by a significant margin, leading to top-ranked performance on popular R-VOS benchmarks, i.e., Ref-YouTube-VOS (67.1%) and Ref-DAVIS17 (65.6%). The code is available athttps://github.com/bo-miao/HTR. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Mubarak Shah, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Hy-Tracker: A Novel Framework for Enhancing Efficiency and Accuracy of Object Tracking in Hyperspectral VideosabstractHyperspectral images, with their many spectral bands, provide a rich source of material information about an object that can be effectively used for object tracking. However, many trackers in this domain rely on detection-based techniques, which often perform suboptimally in challenging scenarios such as managing occlusions and distinguishing objects in cluttered backgrounds. This underperformance is primarily due to the presence of multiple spectral bands and the inability to leverage this abundance of data for effective tracking. Additionally, the scarcity of annotated hyperspectral videos and the absence of comprehensive temporal information exacerbate these difficulties, further limiting the effectiveness of current tracking methods. To address these challenges, this article introduces the novel Hy-Tracker framework, designed to bridge the gap between hyperspectral data and state-of-the-art object detection methods. Our approach leverages the strengths of YOLOv7 for object tracking in hyperspectral videos, enhancing both accuracy and robustness in complex scenarios. The Hy-Tracker framework comprises two key components. We introduce a hierarchical attention for band selection (HAS-BS) that selectively processes and groups the most informative spectral bands, thereby significantly improving detection accuracy. Additionally, we have developed a refined tracker that refines the initial detections by incorporating a classifier and a temporal network using gated recurrent units (GRUs). The classifier distinguishes similar objects, while the temporal network models temporal dependencies across frames for robust performance despite occlusions and scale variations (SVs). Experimental results on hyperspectral benchmark datasets demonstrate the effectiveness of Hy-Tracker in accurately tracking objects across frames and overcoming the challenges inherent in detection-based hyperspectral object tracking (HOT). Wangzhi Xing, Jun Zhou 0001, Yongsheng Gao 0001, Kuldip K. Paliwal |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | CTMNet: Enhanced Open-Pit Mine Extraction and Change Detection With a Hybrid CNN-Transformer Multitask NetworkabstractAutomatic open-pit mine extraction and change detection from high-resolution remote sensing images are of great importance to mineral resource management. However, the high spatial heterogeneity and spectral variations of mining area scenarios make these tasks challenging. Motivated by the strong correlation between the two tasks and their potential mutual benefits, this article presents a hybrid convolutional neural network (CNN)–Transformer multitask network (CTMNet). Constructed in an encoder-decoder manner, CTMNet has two sperate extraction paths (EPs) to localize the regions of interest for bi-temporal images, along with a change detection path (CDP) to identify discrepancies by differentiating the multiscale feature representations from the EPs. As the basic building block for the EP, a CNN-Transformer hybrid block is designed to enhance the global and local feature representation capacity. To cope with the variations in the bi-temporal images, we propose the feature alignment module for the CDP. A hard sample mining-based contrastive constraint loss is proposed to emphasize the contributions of hard samples to the training process. The experimental results on a collected open-pit mine extraction and change detection dataset (OMECSet) and two public datasets reveal the validity of the CTMNet when compared to the state-of-the-art methods. The OMECSet and the code of CTMNet have been made public available athttps://figshare.com/s/80519cb980ca54456447. Jianghe Xing, Jue Zhang 0001, Jun Li 0021, Yongsheng Gao 0001, Shouhang Du, Chengye Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Region Aware Video Object Segmentation With Deep Motion ModelingabstractCurrent semi-supervised video object segmentation (VOS) methods often employ the entire features of one frame to predict object masks and update memory. This introduces significant redundant computations. To reduce redundancy, we introduce a Region Aware Video Object Segmentation (RAVOS) approach, which predicts regions of interest (ROIs) for efficient object segmentation and memory storage. RAVOS includes a fast object motion tracker to predict object ROIs in the next frame. For efficient segmentation, object features are extracted based on the ROIs, and an object decoder is designed for object-level segmentation. For efficient memory storage, we propose motion path memory to filter out redundant context by memorizing the features within the motion path of objects. In addition to RAVOS, we also propose a large-scale occluded VOS dataset, dubbed OVOS, to benchmark the performance of VOS models under occlusions. Evaluation on DAVIS and YouTube-VOS benchmarks and our new OVOS dataset show that our method achieves state-of-the-art performance with significantly faster inference time, e.g., 86.1 J & F at 42 FPS on DAVIS and 84.4 J & F at 23 FPS on YouTube-VOS. Project page: ravos.netlify.app. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2024 | Distinctive Phase Interdependency Model for Retinal Vasculature Delineation in OCT-Angiography ImagesabstractAutomatic detection of retinal vasculature in optical coherence tomography angiography (OCTA) images faces several challenges such as the closely located capillaries, vessel discontinuity and high noise level. This paper introduces a new distinctive phase interdependency model to address these problems for delineating centerline patterns of the vascular network. We capture the inherent property of vascular centerlines by obtaining the inter-scale dependency information that exists between neighboring symmetrical wavelets in complex Poisson domain. In particular, the proposed phase interdependency model identifies vascular centerlines as the distinctive features that have high magnitudes over adjacent symmetrical coefficients whereas the coefficients caused by background noises are decayed rapidly along adjacent wavelet scales. The potential relationships between the neighboring Poisson coefficients are established based on the coherency of distinctive symmetrical wavelets. The proposed phase model is assessed on the OCTA-500 database (300 OCTA images + 200 OCT images), ROSE-1-SVC dataset (9 OCTA images), ROSE-1 (SVC+ DVC) dataset (9 OCTA images), and ROSE-2 dataset (22 OCTA images). The experiments on the clinically relevant OCTA images validate the effectiveness of the proposed method in achieving high-quality results. Our method produces average${F}_{{\textit {Scor}{e}}}$of 0.822, 0.782, and 0.779 on ROSE-1-SVC, ROSE-1 (SVC+ DVC), and ROSE-2 datasets, respectively, and the${F}_{{\textit {Scor}{e}}}$of 0.910 and 0.862 on OCTA_6mm and OCT_3mm datasets (OCTA-500 database), respectively, demonstrating its superior performance over the state-of-the-art benchmark methods. Mohsin Challoob, Yongsheng Gao 0001, Andrew Busch |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Fan-Beam Binarization Difference Projection (FB-BDP): A Novel Local Object Descriptor for Fine-Grained Leaf Image RetrievalabstractFine-grained leaf image retrieval (FGLIR) aims to search similar leaf images in subspecies level which involves very high interclass visual similarity and accordingly poses great challenges to leaf image description. In this study, we introduce a new concept, named fan-beam binarization difference projection (FB-BDP) to address this challenging issue. It is designed based on the theory of fan-beam projection (FBP) which is a mathematical tool originally used for computed tomographic reconstruction of objects and has the merits of capturing the inner structure information of objects in multiple directions and excellent ability to suppress image noise. However, few studies have been made to apply FBP to the description of texture patterns. Rather than calculating ray integrals over the whole object area, FB-BDP restricts its ray integrals calculated over local patches to guarantee the locality of the extracted features. By binarizing the intensity-differences between the off-center and center rays, FB-BDP enable its ray integrals insensitive to illumination change and more discriminative in the characterization of texture patterns. In additional, due to inheriting the merits of FBP, the proposed FB-BDP is superior over the existing local image descriptors by its invariance to scaling transformation, robustness to noise, and strong ability to capture direction and structure texture patterns. The results of extensive experiments on FGLIR show its higher retrieval accuracy over the benchmark methods, promising generalization power and strong complementarity to deep features. Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
ICCV | 3 |
| 2023 | Spectrum-guided Multi-granularity Referring Video Object SegmentationabstractCurrent referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perceive during the forward computation. This negatively affects the ability of segmentation kernels. To address the drift problem, we propose a Spectrum-guided Multi-granularity (SgMg) approach, which performs direct segmentation on the encoded features and employs visual details to further optimize the masks. In addition, we propose Spectrum-guided Cross-modal Fusion (SCF) to perform intra-frame global interactions in the spectral domain for effective multimodal representation. Finally, we extend SgMg to perform multi-object R-VOS, a new paradigm that enables simultaneous segmentation of multiple referred objects in a video. This not only makes R-VOS faster, but also more practical. Extensive experiments show that SgMg achieves state-of-the-art performance on four video benchmark datasets, outperforming the nearest competitor by 2.8% points on Ref-YouTube-VOS. Our extended SgMg enables multi-object R-VOS, runs about 3 faster while maintaining satisfactory performance. Code×is available at https://github.com/bo-miao/SgMg. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICCV | 3 |
| 2023 | CLE-ViT: Contrastive Learning Encoded Transformer for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself, e.g., predicting the rotation of a given image, has demonstrated great progress towards learning discriminative representation. Yet none of these works consider explicit supervision for learning mutual information at instance level. To this end, this paper introduces CLE-ViT, a novel contrastive learning encoded transformer, to address the fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLE-ViT demonstrates strong performance on 7 publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at https://github.com/Markin-Wang/CLEViT. Xiaohan Yu 0001, Jun Wang 0121, Yongsheng Gao 0001 |
IJCAI | 3 |
| 2023 | A Modulatory Elongated Model for Delineating Retinal Microvasculature in OCTA Images
Mohsin Challoob, Yongsheng Gao 0001, Andrew Busch |
MICCAI (7) | 2 |
| 2023 | SSFE-Net: Self-Supervised Feature Enhancement for Ultra-Fine-Grained Few-Shot Class Incremental LearningabstractUltra-Fine-Grained Visual Categorization (ultra-FGVC) has become a popular problem due to its great real-world potential for classifying the same or closely related species with very similar layouts. However, there present many challenges for the existing ultra-FGVC methods, firstly there are always not enough samples in the existing ultraFGVC datasets based on which the models can easily get overfitting. Secondly, in practice, we are likely to find new species that we have not seen before and need to add them to existing models, which is known as incremental learning. The existing methods solve these problems by Few-Shot Class Incremental Learning (FSCIL), but the main challenge of the FSCIL models on ultra-FGVC tasks lies in their inferior discrimination detection ability since they usually use low-capacity networks to extract features, which leads to insufficient discriminative details extraction from ultrafine-grained images. In this paper, a self-supervised feature enhancement for the few-shot incremental learning network (SSFE-Net) is proposed to solve this problem. Specifically, a self-supervised learning (SSL) and knowledge distillation (KD) framework is developed to enhance the feature extraction of the low-capacity backbone network for ultra-FGVC few-shot class incremental learning tasks. Besides, we for the first time create a series of benchmarks for FSCIL tasks on two public ultra-FGVC datasets and three normal finegrained datasets, which will facilitate the development of the Ultra-FGVC community. Extensive experimental results on public ultra-FGVC datasets and other state-of-the-art benchmarks consistently demonstrate the effectiveness of the proposed method. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
WACV | 4 |
| 2023 | Extracting optimal explanations for ensemble trees via automated reasoning
Gelin Zhang, Yanhong Huang, Jianqi Shi, Hadrien Bride, Jin Song Dong 0001, Yongsheng Gao 0001 |
Appl. Intell. | 7 |
| 2023 | ECFRNet: Effective corner feature representations network for image corner detection
Junfeng Jing, Chao Liu 0033, Yongsheng Gao 0001, Changming Sun |
Expert Syst. Appl. | 4 |
| 2023 | Background-Aware Band Selection for Object Tracking in Hyperspectral VideosabstractHyperspectral images contain many bands that can be used to obtain object material information for object tracking and remote sensing. Nevertheless, neighboring bands of hyperspectral images are often highly correlated, and a large number of bands increase the complexity of model learning. This issue is worsened by the shortage of labeled hyperspectral videos for fine-tuning pre-trained deep neural networks. To tackle these challenges, this paper introduces a novel background-aware band selection method to model spatial changes of an object and its corresponding local region, which is capable of selecting discriminative bands for object representation while reducing computational complexity. Specifically, the object and local region of each band is compared with other bands to obtain their dissimilarity scores. Guided by these scores, the top three bands are selected and form a three-channel image. This image is then fed into an object tracker. Experimental results demonstrate the efficiency and effectiveness of the proposed method on a benchmark hyperspectral object tracking dataset. Jun Zhou 0001, Yongsheng Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Image Feature Information Extraction for Interest Point Detection: A Comprehensive ReviewabstractInterest point detection is one of the most fundamental and critical problems in computer vision and image processing. In this paper, we carry out a comprehensive review on image feature information (IFI) extraction techniques for interest point detection. To systematically introduce how the existing interest point detection methods extract IFI from an input image, we propose a taxonomy of the IFI extraction techniques for interest point detection. According to this taxonomy, we discuss different types of IFI extraction techniques for interest point detection. Furthermore, we identify the main unresolved issues related to the existing IFI extraction techniques for interest point detection and any interest point detection methods that have not been discussed before. The existing popular datasets and evaluation standards are provided and the performances for fifteen state-of-the-art approaches are evaluated and discussed. Moreover, future research directions on IFI extraction techniques for interest point detection are elaborated. Junfeng Jing, Tian Gao 0004, Yongsheng Gao 0001, Changming Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Image Intensity Variation Information for Interest Point DetectionabstractInterest point detection methods are gaining more attention and are widely applied in computer vision tasks such as image retrieval and 3D reconstruction. However, there still exist two main problems to be solved: (1) from the perspective of mathematical representations, the differences among edges, corners, and blobs have not been convincingly explained and the relationships among the amplitude response, scale factor, and filtering orientation for interest points have not been thoroughly explained; (2) the existing design mechanism for interest point detection does not show how to accurately obtain intensity variation information on corners and blobs. In this paper, the first- and second-order Gaussian directional derivative representations of a step edge, four common genres of corners, an anisotropic-type blob, and an isotropic-type blob are analyzed and derived. Multiple interest point characteristics are discovered. The characteristics for interest points that we obtained help us describe the differences among edges, corners, and blobs, explain why the existing interest point detection methods with multiple scales cannot properly obtain interest points from images, and present novel corner and blob detection methods. Extensive experiments demonstrate the superiority of our proposed methods in terms of detection performance, robustness to affine transformations, noise, image matching, and 3D reconstruction. Changming Sun, Yongsheng Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | A Lie algebra representation for efficient 2D shape classification
Xiaohan Yu 0001, Yongsheng Gao 0001, Mohammed Bennamoun, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2023 | Mix-ViT: Mixing attentive vision transformer for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Jun Wang 0121, Yang Zhao 0019, Yongsheng Gao 0001 |
Pattern Recognit. | 4 |
| 2023 | Gait-Assisted Video Person RetrievalabstractVideo person retrieval aims at matching video clips of the same person across non-overlapping camera views, where video sequences contain more comprehensive information, e.g., temporal cues. How to extract useful temporal cues is the key to the success of a video person retrieval system. Gait, as a unique biometric modality indicating the way people walk, contains informative temporal information. To date, it is not clear how to fully utilize gait to boost the performance of video person retrieval. In this paper, to validate whether gait could help retrieve person in videos, we build a two-stream architecture, named appearance-gait network (AGNet), to jointly learn the appearance features and gait features from RGB video clips and silhouette video clips. We further explore how to fully utilize gait features to enhance the video feature representation. Specifically, we propose an appearance-gait attention module (AGA) to fuse a discriminative feature representation for the person retrieval task. Furthermore, to eliminate the requirement of silhouette video clips during inference, we propose a simple yet effective appearance-gait distillation module (AGD) which transfers the gait knowledge to appearance stream. As such, we are able to perform the enhanced video person retrieval without silhouette video clips, which makes the inference more flexible and practical. To the best of our knowledge, our work is the first to successfully introduce such appearance-gait knowledge distillation design for video person retrieval. We verify the effectiveness of the proposed methods on two large-scale challenging benchmarks of MARS and DukeMTMC-VideoReID. Extensive experiments demonstrate superior or comparable performance compared to the state-of-the-art methods while being much simpler. Source code is publicly available athttps://github.com/yangyangkiki/Gait-Assisted-Video-Reid. Yang Zhao 0019, Xiaohan Yu 0001, Chunlei Liu 0001, Yongsheng Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Separable Paravector Orientation Tensors for Enhancing Retinal VesselsabstractRobust detection of retinal vessels remains an unsolved research problem, particularly in handling the intrinsic real-world challenges of highly imbalanced contrast between thick vessels and thin ones, inhomogeneous background regions, uneven illumination, and complex geometries of crossing/bifurcations. This paper presents a new separable paravector orientation tensor that addresses these difficulties by characterizing the enhancement of retinal vessels to be dependent on a nonlinear scale representation, invariant to changes in contrast and lighting, responsive for symmetric patterns, and fitted with elliptical cross-sections. The proposed method is built on projecting vessels as a 3D paravector valued function rotated in an alpha quarter domain, providing geometrical, structural, symmetric, and energetic features. We introduce an innovative symmetrical inhibitory scheme that incorporates paravector features for producing a set of directional contrast-independent elongated-like patterns reconstructing vessel tree in orientation tensors. By fitting constraint elliptical volumes via eigensystem analysis, the final vessel tree is produced with a strong and uniform response preserving various vessel features. The validation of proposed method on clinically relevant retinal images with high-quality results, shows its excellent performance compared to the state-of-the-art benchmarks and the second human observers. Mohsin Challoob, Yongsheng Gao 0001, Andrew Busch, Mohammad Nikzad |
IEEE Trans. Medical Imaging | 2 |
| 2022 | A Generalized Kernel Risk Sensitive Loss for Robust Two-Dimensional Singular Value DecompositionabstractTwo-dimensional singular value decomposition (2DSVD) is an important dimensionality reduction algorithm which has inherent advantage in preserving the structure of 2D images. However, 2DSVD algorithm is based on the squared error loss, which may exaggerate the projection errors with the presence of outliers. To solve this problem, we propose a generalized kernel risk sensitive loss for measuring the projection error in 2DSVD, which automatically eliminates the outlier information during optimization. Since the proposed objective function is non-convex, a majorization-minimization algorithm is developed to efficiently solve it. Our method is rotational invariant and has intrinsic advantages in processing non-centered data. Experimental results on public databases demonstrate that the performance of the proposed method significantly outperforms several benchmark methods on different applications. Miaohua Zhang, Yongsheng Gao 0001, Jun Zhou 0001 |
ICASSP | 2 |
| 2022 | Pairwise Rotational-Difference LBP for Fine-Grained Leaf Image RetrievalabstractIn this study, we address the challenging issue of fine-grained leaf image retrieval which focuses on distinguishing different cultivars within the same species. We propose a novel local binary pattern, named pairwise rotation-difference LBP (PRDLBP), for the characterization of leaf image patterns. Different from the conventional LBP which measure the local grayscale contrast between the center pixel and its circular neighboring pixels, we consider the grayscale contrast between the circular neighboring pixels that are rotational symmetric about the center pixel. The proposed PRDLBP is a co-occurrence LBP feature representation which can not only encode spatially symmetric co-occurrence information, but also be inherently invariant to rotation. Its stronger discriminative power over the state-of-the-arts has been validated on two challenging fine-grained leaf image retrieval tasks, soybean cultivar identification and peanut cultivar identification. This work may attract considerable attention to fine-grained leaf image retrieval and advance the research of leaf image pattern identification from species to cultivars. Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
ICIP | 3 |
| 2022 | Self-Supervised Video Object Segmentation by Motion-Aware Mask PropagationabstractWe propose a self-supervised spatio-temporal matching method, coined Motion-Aware Mask Propagation (MAMP), for video object segmentation. MAMP leverages the frame reconstruction task for training without the need for annotations. During inference, MAMP builds a dynamic memory bank and propagates masks according to our proposed motion-aware spatio-temporal matching module, which is able to handle fast motion and long-term matching scenarios. Evaluation on DAVIS-2017 and YouTube-VOS datasets show that MAMP achieves state-of-the-art performance with stronger generalization ability compared to existing self-supervised methods, i.e., 4.2% higher mean$\mathcal{J}$&$\mathcal{F}$on DAVIS-2017 and 4.85% higher mean$\mathcal{J}$&$\mathcal{F}$on the unseen categories of YouTube-VOS than the nearest competitor. Moreover, MAMP performs at par with many supervised video object segmentation methods. Our code is available at: https://github.com/bo-miao/MAMP. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICME | 3 |
| 2022 | Leaf Vocabulary: Fine-Grained Leaf Image Retrieval Using Bag-of-Visual-Words RepresentationabstractThis paper addresses the issue of fine-grained leaf image retrieval (FGLIR) which focuses on differentiating between different leaf cultivars within the same species. We investigate a novel bag-of-visual-words approaches (BoVW) to FGLIR. Firstly, we treat each leaf boundary point as the key-point from which to spread a chord pair for measuring the local characteristics including shape, gray-level and gradient co-occurrence texture features of the leaf image. By varying the length of the chord, we obtain multiscale local features which are then used to form two local shape and texture feature vectors. Secondly, we separately collect all the local shape and texture vectors from the database images to learn a leaf shape vocabulary and a leaf texture vocabulary by k-means clustering algorithm. By mapping the two kinds of local feature vectors to visual words in their corresponding leaf vocabularies, we can represent each leaf image as two bags of visual words (one for shape, another for texture). Finally, we convert them into two visual-word vectors by counting the occurrence of each leaf visual word in the image and concatenate them as the final image representation. The proposed leaf vocabulary representation is applied to two challenging FGLIR tasks, soybean cultivar identification and peanut cultivar identification. The experimental results indicate its superior performance over the state-of-the-art leaf descriptors and show its potential to address the issue of FGLIR. Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
ICPR | 3 |
| 2022 | Distribution-Aware Margin Calibration for Semantic Segmentation in Images
Litao Yu, Zhibin Li 0002, Min Xu 0001, Yongsheng Gao 0001, Jiebo Luo 0001, Jian Zhang 0002 |
Int. J. Comput. Vis. | 4 |
| 2022 | Symmetric Binary Tree Based Co-occurrence Texture Pattern Mining for Fine-grained Plant Leaf Image Retrieval
Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2022 | SPARE: Self-supervised part erasing for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2022 | Learning discriminative region representation for person retrieval
Yang Zhao 0019, Xiaohan Yu 0001, Yongsheng Gao 0001, Chunhua Shen |
Pattern Recognit. | 3 |
| 2022 | Local R-Symmetry Co-Occurrence: Characterising Leaf Image Patterns for Identifying CultivarsabstractLeaf image recognition techniques have been actively researched for plant species identification. However it remains unclear whether analysing leaf patterns can provide sufficient information for further differentiating cultivars. This paper reports our attempt on cultivar recognition from leaves as a general very fine-grained pattern recognition problem, which is not only a challenging research problem but also important for cultivar evaluation, selection and production in agriculture. We propose a novel local R-symmetry co-occurrence method for characterising discriminative local symmetry patterns to distinguish subtle differences among cultivars. Through scalable and moving R-relation radius pairs, we generate a set of radius symmetry co-occurrence matrices (RsCoM)and their measures for describing the local symmetry properties of interior regions. By varying the size of the radius pair, the RsCoM measures local R-symmetry co-occurrence from global/coarse to fine scales. A new two-phase strategy of analysing the distribution of local RsCoM measures is designed to match the multiple scale appearance symmetry pattern distributions of similar cultivar leaf images. We constructed three leaf image databases, SoyCultivar, CottCultivar, and PeanCultivar, for an extensive experimental evaluation on recognition across soybean, cotton and peanut cultivars. Encouraging experimental results of the proposed method in comparison with the state-of-the-art leaf species recognition methods demonstrate the effectiveness of the proposed method for cultivar identification, which may advance the research in leaf recognition from species to cultivar. Bin Wang 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | An Attention-Based Lattice Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) with 3-D convolutional kernels are widely used for hyperspectral image (HSI) classification, which bring notable benefits in capturing joint spectral and spatial features. However, they suffer from poor computational efficiency, causing the low training/inference speed of the model. On the contrary, CNN-based methods with 1-D and 2-D kernels are efficient but mostly restricted to extracting either spectral or spatial features. Moreover, most CNN-based HSI classification frameworks are incapable of simultaneously taking advantage of residual and dense aggregations without over-allocating parameters or information loss for feature reusage. This article presents a novel attention-based lattice network (ALN) to overcome these shortcomings. The proposed 2-D lattice framework can effectively harness the advantages of residual and dense aggregations to achieve outstanding accuracy performance and computational efficiency simultaneously. Furthermore, the ALN employs a unique joint spectral–spatial attention mechanism to capture both spectral and spatial information effectively. In particular, a new pointwise spectral attention mechanism is adopted to fully capture spectral dependences for every pixel. Our extensive experimental investigation verifies the effectiveness and efficiency of the ALN architecture for HSI classification. Mohammad Nikzad, Yongsheng Gao 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Feature Fusion Vision Transformer for Fine-Grained Visual Categorization
Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
BMVC | 3 |
| 2021 | Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human PerformanceabstractDeep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on the basis of such human visual observation. In contrast, few research has been done for visual categorization at the ultra-fine-grained level, i.e., a granularity where even human experts can hardly identify the visual differences or are not yet able to give affirmative labels by inferring observed pattern differences. This paper reports our efforts towards mitigating this research gap. We introduce the ultra-fine-grained (UFG) image dataset, a large collection of 47,114 images from 3,526 categories. All the images in the proposed UFG image dataset are grouped into categories with different confirmed cultivar names. In addition, we perform an extensive evaluation of state-of-the-art fine-grained classification methods on the proposed UFG image dataset as comparative baselines. The proposed UFG image dataset and evaluation protocols is intended to serve as a benchmark platform that can advance research of visual classification from approaching human performance to beyond human ability, via facilitating benchmark data of artificial intelligence (AI) not to be limited by the labels of human intelligence (HI). The dataset is available online at https://githuh.com/XiaohanYu-GU/Ultra-FGVC. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
ICCV | 3 |
| 2021 | Fine-Grained Plant Leaf Image Retrieval Using Local Angle Co-occurrence HistogramsabstractLeaf image patterns have been actively researched for plant species recognition. However, as a very challenging fine-grained pattern identification issue, cultivar recognition in which the leaf image patterns usually have very subtle difference among cultivars has not yet received considerable attention in computer vision community. In this paper, a novel leaf image descriptor, named local angle co-occurrence histograms, is proposed for addressing this issue. It is a kind of co-occurrence descriptors that encoding both shape and texture features which make them more informative than the existing individual descriptors and co-occurrence features. A feature fusion scheme is proposed to integrate the handcrafted descriptors with deep learning features for further boosting the retrieval performance. The experimental results on the challenging soybean cultivar recognition and peanut cultivar recognition both indicate the superiority of the proposed method over the state-of-the-art methods on leaf image pattern characterization and validate the effectiveness of the proposed method for fine-grained leaf image retrieval. Xin Chen 0100, Jiawei You, Hui Tang 0003, Bin Wang 0041, Yongsheng Gao 0001 |
ICIP | 5 |
| 2021 | Mask Guided Attention For Fine-Grained Patchy Image ClassificationabstractIn this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies in two folds, ultra-fine-grained inter-category variances among objects and very few data available for training. This motivates us to consider employing more useful supervision signal to train a discriminative model within limited training samples. Specifically, the proposed MGA integrates a pre-trained semantic segmentation model that produces auxiliary supervision signal, i.e., patchy attention mask, enabling a discriminative representation learning. The patchy attention mask drives the classifier to filter out the insignificant parts of images (e.g., common features between different categories), which enhances the robustness of MGA for the fine-grained patchy image classification. We verify the effectiveness of our method on three publicly available patchy image datasets. Experimental results demonstrate that our MGA method achieves superior performance on three datasets compared with the state-of-the-art methods. In addition, our ablation study shows that MGA improves the accuracy by 2.25% and 2% on the SoyCultivarVein and BtfPIS datasets, indicating its practicality towards solving the fine-grained patchy image classification. Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
ICIP | 3 |
| 2021 | Attention-based Pyramid Dilated Lattice Network for Blind Image DenoisingabstractThough convolutional neural networks (CNNs) with residual and dense aggregations have obtained much attention in image denoising, they are incapable of exploiting different levels of contextual information at every convolutional unit in order to infer different levels of noise components with a single model. In this paper, to overcome this shortcoming we present a novel attention-based pyramid dilated lattice (APDL) architecture and investigate its capability for blind image denoising. The proposed framework can effectively harness the advantages of residual and dense aggregations to achieve a great trade-off between performance, parameter efficiency, and test time. It also employs a novel pyramid dilated convolution strategy to effectively capture contextual information corresponding to different noise levels through the training of a single model. Our extensive experimental investigation verifies the effectiveness and efficiency of the APDL architecture for image denoising as well as JPEG artifacts suppression tasks. Mohammad Nikzad, Yongsheng Gao 0001, Jun Zhou 0001 |
IJCAI | 2 |
| 2021 | Composite description based on color vector quantization and visual primary features for CBIR tasks
Muhammad Daud Abdullah Asif, Jing Wang 0062, Yongsheng Gao 0001, Jun Zhou 0001 |
Multim. Tools Appl. | 3 |
| 2021 | Dual subspace discriminative projection learning
Gregg Belous, Andrew Busch, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2021 | MaskCOV: A random mask covariance network for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 3 |
| 2021 | A unified weight learning and low-rank regression model for robust complex error modeling
Miaohua Zhang, Yongsheng Gao 0001, Jun Zhou 0001 |
Pattern Recognit. | 2 |
| 2021 | Learning deep part-aware embedding for person retrieval
Yang Zhao 0019, Chunhua Shen, Xiaohan Yu 0001, Hao Chen 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 5 |
| 2021 | Robust Tensor Decomposition for Image Representation Based on Generalized CorrentropyabstractTraditional tensor decomposition methods, e.g., two dimensional principal component analysis and two dimensional singular value decomposition, that minimize mean square errors, are sensitive to outliers. To overcome this problem, in this paper we propose a new robust tensor decomposition method using generalized correntropy criterion (Corr-Tensor). A Lagrange multiplier method is used to effectively optimize the generalized correntropy objective function in an iterative manner. The Corr-Tensor can effectively improve the robustness of tensor decomposition with the existence of outliers without introducing any extra computational cost. Experimental results demonstrated that the proposed method significantly reduces the reconstruction error on face reconstruction and improves the accuracies on handwritten digit recognition and facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
IEEE Trans. Image Process. | 2 |
| 2021 | Exploring Chromatic Aberration and Defocus Blur for Relative Depth Estimation From Monocular Hyperspectral ImageabstractThis article investigates spectral chromatic and spatial defocus aberration in a monocular hyperspectral image (HSI) and proposes methods on how these cues can be utilized for relative depth estimation. The main aim of this work is to develop a framework by exploring intrinsic and extrinsic reflectance properties in HSI that can be useful for depth estimation. Depth estimation from a monocular image is a challenging task. An additional level of difficulty is added due to low resolution and noises in hyperspectral data. Our contribution to handling depth estimation in HSI is threefold. Firstly, we propose that change in focus across band images of HSI due to chromatic aberration and band-wise defocus blur can be integrated for depth estimation. Novel methods are developed to estimate sparse depth maps based on different integration models. Secondly, by adopting manifold learning, an effective objective function is developed to combine all sparse depth maps into a final optimized sparse depth map. Lastly, a new dense depth map generation approach is proposed, which extrapolate sparse depth cues by using material-based properties on graph Laplacian. Experimental results show that our methods successfully exploit HSI properties to generate depth cues. We also compare our method with state-of-the-art RGB image-based approaches, which shows that our methods produce better sparse and dense depth maps than those from the benchmark methods. Ali Zia, Jun Zhou 0001, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Parameter-Efficient Deep Neural Networks With Bilinear ProjectionsabstractRecent research on deep neural networks (DNNs) has primarily focused on improving the model accuracy. Given a proper deep learning framework, it is generally possible to increase the depth or layer width to achieve a higher level of accuracy. However, the huge number of model parameters imposes more computational and memory usage overhead and leads to the parameter redundancy. In this article, we address the parameter redundancy problem in DNNs by replacing conventional full projections with bilinear projections (BPs). For a fully connected layer with D input nodes and D output nodes, applying BP can reduce the model space complexity fromO(D2) toO(2D), achieving a deep model with a sublinear layer size. However, the structured projection has a lower freedom of degree compared with the full projection, causing the underfitting problem. Therefore, we simply scale up the mapping size by increasing the number of output channels, which can keep and even boosts the model accuracy. This makes it very parameter-efficient and handy to deploy such deep models on mobile systems with memory limitations. Experiments on four benchmark data sets show that applying the proposed BP to DNNs can achieve even higher accuracies than conventional full DNNs while significantly reducing the model size. Litao Yu, Yongsheng Gao 0001, Jun Zhou 0001, Jian Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Deep Residual-Dense Lattice Network for Speech EnhancementabstractConvolutional neural networks (CNNs) with residual links (ResNets) and causal dilated convolutional units have been the network of choice for deep learning approaches to speech enhancement. While residual links improve gradient flow during training, feature diminution of shallow layer outputs can occur due to repetitive summations with deeper layer outputs. One strategy to improve feature re-usage is to fuse both ResNets and densely connected CNNs (DenseNets). DenseNets, however, over-allocate parameters for feature re-usage. Motivated by this, we propose the residual-dense lattice network (RDL-Net), which is a new CNN for speech enhancement that employs both residual and dense aggregations without over-allocating parameters for feature re-usage. This is managed through the topology of the RDL blocks, which limit the number of outputs used for dense aggregations. Our extensive experimental investigation shows that RDL-Nets are able to achieve a higher speech enhancement performance than CNNs that employ residual and/or dense aggregations. RDL-Nets also use substantially fewer parameters and have a lower computational requirement. Furthermore, we demonstrate that RDL-Nets outperform many state-of-the-art deep learning approaches to speech enhancement. Availability: https://github.com/nick-nikzad/RDL-SE. Mohammad Nikzad, Aaron Nicolson, Yongsheng Gao 0001, Jun Zhou 0001, Kuldip K. Paliwal, Fanhua Shang |
AAAI | 3 |
| 2020 | Patchy Image Structure Classification Using Multi-Orientation Region TransformabstractExterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal structure feature separately, and thus fail to function when classifying patchy image structures that have similar contours and flexible structures. To address above limitations, this paper proposes a novel Multi-Orientation Region Transform (MORT), which can effectively characterize both contour and structure features simultaneously, for patchy image structure classification. MORT is performed over multiple orientation regions at multiple scales to effectively integrate patchy features, and thus enables a better description of the shape in a coarse-to-fine manner. Moreover, the proposed MORT can be extended to combine with the deep convolutional neural network techniques, for further enhancement of classification accuracy. Very encouraging experimental results on the challenging ultra-fine-grained cultivar recognition task, insect wing recognition task, and large variation butterfly recognition task are obtained, which demonstrate the effectiveness and superiority of the proposed MORT over the state-of-the-art methods in classifying patchy image structures. Our code and three patchy image structure datasets are available at: https://github.com/XiaohanYu-GU/MReT2019. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
AAAI | 3 |
| 2020 | Quadratic Tensor Anisotropy Measures for Reliable Curvilinear Pattern Detection
Mohsin Challoob, Yongsheng Gao 0001 |
ACIVS | 2 |
| 2020 | A Local Flow Phase Stretch Transform for Robust Retinal Vessel Detection
Mohsin Challoob, Yongsheng Gao 0001 |
ACIVS | 2 |
| 2020 | A Novel Line Integral Transform for 2D Affine-Invariant Shape Retrieval
Bin Wang 0041, Yongsheng Gao 0001 |
ECCV (28) | 2 |
| 2020 | Gaussian Convolution Angles: Invariant Vein and Texture Descriptors for Butterfly Species IdentificationabstractIdentifying butterfly species by image patterns is a challenging task in computer vision and pattern recognition community due to many butterfly species having similar shape patterns with complex interior structures and considerable pose variation. In additional, geometrical transformation and illumination variation also make this task more difficult. In this paper, a novel image descriptor, named Gaussian convolution angle (GCA) is proposed for butterfly species classification. The proposed GCA projects the butterfly vein image function and intensity image function along a group of vectors that start from a common contour points and ends at the remaining contour points which results a group of vectors that capture the complex vein patterns and texture patterns of butterfly images. The Gaussian convolutions of different scales are conducted to the resulting vector functions to generate a multiscale GCA descriptors. The proposed GCA is not only invariant to geometrical transformation including rotation, scaling and translation, but also invariant to lighting change. The proposed method has been tested on a publicly available butterfly image dataset that has 832 samples of 10 species. It achieves a classification accuracy of 92.03% which is higher than the benchmark methods. Xin Chen 0100, Bin Wang 0041, Yongsheng Gao 0001 |
ICPR | 3 |
| 2020 | DE-Net: Dilated Encoder Network for Automated Tongue SegmentationabstractAutomated tongue recognition is a growing research field due to global demand for personal health care. Using mobile devices to take tongue pictures is convenient and of low cost for tongue recognition. It is particularly suitable for self-health evaluation of the public. However, images taken by mobile devices are easily affected by various imaging environment, which makes fine segmentation a more challenging task compared with those taken by specialized acquisition devices. Deep learning approaches are promising for tongue image segmentation because they have powerful feature learning and representation capability. However, the successive pooling operations in these methods lead to loss of information on image details, making them fail when segmenting low-quality images captured by mobile devices. To address this issue, we propose a dilated encoder network (DE-Net) to capture more high-level features and get high-resolution output for automated tongue image segmentation. In addition, we construct two tongue image datasets which contain images taken by specialized devices and mobile devices, respectively, to verify the effectiveness of the proposed method. Experimental results on both datasets demonstrate that the proposed method outperforms the state-of-the-art methods in tongue image segmentation. Hui Tang 0003, Bin Wang 0041, Jun Zhou 0001, Yongsheng Gao 0001 |
ICPR | 4 |
| 2020 | A robust matching pursuit algorithm using information theoretic learning
Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
Pattern Recognit. | 2 |
| 2020 | MobileFAN: Transferring deep hidden representation for face alignment
Yang Zhao 0019, Yifan Liu 0001, Chunhua Shen, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2020 | Double Graph Regularized Double Dictionary Learning for Image ClassificationabstractIn this paper, we present a novel double graph regularized double dictionary learning (DGRDDL) method for image classification. The proposed method jointly constructs a number of class-specific sub-dictionaries to capture the most discriminative features (class-specific information) of each class, and a class-shared dictionary to model the common patterns (class-shared information) shared by the images from different classes. A novel double graph regularization is proposed to correctly represent and differentiate these two types of information. Specifically, an intra-class similarity graph constraint is imposed on the representation coefficients over the class-specific dictionaries, and an inter-class similarity graph constraint is applied on the representation coefficients over the class-shared dictionary. In this way, the representations learned by the proposed DGRDDL method can correctly model the local similarity relationships of the class-specific and the class-shared information in images, respectively. Moreover, due to the differences between the intra-class and inter-class similarity graphs, the two types of information can be appropriately separated and captured by the learned dictionaries. We evaluate the performance of the proposed method on six public datasets and compared against those of seven benchmark methods. The experimental results demonstrate the effectiveness and superiority of the proposed method in image classification over the benchmark dictionary learning methods. Shengwu Xiong 0001, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Contour Covariance: A Fast Descriptor for ClassificationabstractThis paper presents a novel shape descriptor to effectively and efficiently characterize the local image statistics. The proposed descriptor, termed contour covariance (CC), characterizes covariance features driven by a moving point on the shape contour at multiple scales. To calculate the covariance matrices, three basic features including texture, intensity and distance map, are extracted from the object image. Based on coefficients of the obtained covariance matrices, the proposed CC descriptor is compact yet informative, as well as invariant to rotation, translation and scale. The experimental results on two databases demonstrate the superiority and efficiency of the proposed method among the state-of-the-art methods for shape classification. Xiaohan Yu 0001, Shengwu Xiong 0001, Yongsheng Gao 0001 |
ICIP | 3 |
| 2019 | Robust Sparse Learning Based on Kernel Non-Second Order MinimizationabstractPartial occlusions in face images pose a great problem for most face recognition algorithms due to the fact that most of these algorithms mainly focus on solving a second order loss function, e.g., mean square error (MSE), which will magnify the effect from occlusion parts. In this paper, we proposed a kernel non-second order loss function for sparse representation (KNS-SR) to recognize or restore partially occluded facial images, which both take the advantages of the correntropy and the non-second order statistics measurement. The resulted framework is more accurate than the MSE-based ones in locating and eliminating outliers information. Experimental results from image reconstruction and recognition tasks on publicly available databases show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 2 |
| 2019 | Kernel Mean P Power Error Loss for Robust Two-Dimensional Singular Value DecompositionabstractTraditional matrix-based dimensional reduction methods, e.g., two-dimensional principal component analysis (2DPCA) and two-dimensional singular value decomposition (2DSVD), minimize mean square errors (MSE), which is sensitive to outliers. To overcome this problem, in this paper we propose a new robust 2DSVD method based on the kernel mean p power error loss (KMPE-2DSVD). Different from the MSE and the correntropy based ones which are second order statistics based measurements, the KMPE-2DSVD is based on the non-second order statistics in the kernel space, and thus is more flexible in controlling the representation error. Experimental results show that the proposed method significantly improves the accuracy of facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 2 |
| 2019 | Decomposed-Excitatory Antagonistic Model for Reliable Retinal Vessel DetectionabstractA new retinal vessel detection method is introduced using an antagonistic model, consisting of the excitatory and inhibitory regions. It utilizes the model profile containing important image features such as lines and curves to detect retinal vessels. The excitatory area is analyzed into several subregions with multiple directions to decompose the local vessel structure, resulting in several profiles that include various vessel trees. Also, the strength of the inhibitory region is adaptively varied to differentiate vessel pixels from background while decreasing false detection. Then, each profile is linked to a multi-bound relaxed median operator for noise removal, and a logistic function for providing a soft transition between vessel and non-vessel regions. The final vessel tree is produced by fusing all profiles. The proposed method is evaluated on high-resolution and low-resolution fundus imaging databases (HRF and DRIVE), on which it outperforms state-of-the-art methods. Mohsin Challoob, Yongsheng Gao 0001 |
VCIP | 2 |
| 2019 | Gradient-Based Pooling for Convolutional Neural NetworksabstractPooling layers are an important part of convolutional neural networks (CNNs). They reduce the dimensionality of feature maps and pass salient information to subsequent layers. In this paper, we introduce a novel gradient-based feature pooling method that can down-sample feature maps while better preserving key information. This method considers the spatial gradient of the pixels within a pooling region as a key to select the most possible descriptive information in contrast to the current practice of existing methods that mostly rely on the pixel values. Extensive experiments on different benchmark image classification tasks and CNN architectures demonstrate that the proposed method achieves superior results over existing pooling approaches. Mohammad Nikzad, Yongsheng Gao 0001, Jun Zhou 0001 |
VCIP | 2 |
| 2019 | Learning discriminative subregions and pattern orders for facial gender classification
Zhong Chen 0003, Andrea Edwards, Yongsheng Gao 0001, Kun Zhang 0012 |
Image Vis. Comput. | 3 |
| 2019 | ST-CNN: Spatial-Temporal Convolutional Neural Network for crowd counting in videos
Yunqi Miao, Jungong Han, Yongsheng Gao 0001, Baochang Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | Conditional Random Field and Deep Feature Learning for Hyperspectral Image ClassificationabstractImage classification is considered to be one of the critical tasks in hyperspectral remote sensing image processing. Recently, a convolutional neural network (CNN) has established itself as a powerful model in classification by demonstrating excellent performances. The use of a graphical model such as a conditional random field (CRF) contributes further in capturing contextual information and thus improving the classification performance. In this paper, we propose a method to classify hyperspectral images by considering both spectral and spatial information via a combined framework consisting of CNN and CRF. We use multiple spectral band groups to learn deep features using CNN, and then formulate deep CRF with CNN-based unary and pairwise potential functions to effectively extract the semantic correlations between patches consisting of 3-D data cubes. Furthermore, we introduce a deep deconvolution network that improves the final classification performance. We also introduced a new data set and experimented our proposed method on it along with several widely adopted benchmark data sets to evaluate the effectiveness of our method. By comparing our results with those from several state-of-the-art models, we show the promising potential of our method. Fahim Irfan Alam, Jun Zhou 0001, Alan Wee-Chung Liew, Xiuping Jia, Jocelyn Chanussot, Yongsheng Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Chord Bunch Walks for Recognizing Naturally Self-Overlapped and Compound LeavesabstractEffectively describing and recognizing leaf shapes under arbitrary variations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing leaf shapes by walking and measuring along a bunch of chords that pass through the shape. A novel chord bunch walks (CBW) descriptor is developed through the chord walking behavior that effectively integrates the shape image function over the walked chord to reflect both the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chords to build a hierarchical framework for a coarse-to-fine description that can effectively characterize not only the subtle differences among leaf margin patterns but also the interior part of the shape contour formed inside a self-overlapped or compound leaf. Instead of using optimal correspondence based matching, a Log-Min distance that encourages one-to-one correspondences is proposed for efficient and effective CBW matching. The proposed CBW shape analysis method is invariant to rotation, scaling, translation, and mirror transforms. Five experiments, including image retrieval of compound leaves, image retrieval of naturally self-overlapped leaves, and retrieval of mixed leaves on three large scale datasets, are conducted. The proposed method achieved large accuracy increases with low computational costs over the state-of-the-art benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
IEEE Trans. Image Process. | 2 |
| 2018 | A Novel Multi-scale Invariant Descriptor Based on Contour and Texture for Shape Recognition
Jishan Guo, Yongsheng Gao 0001, Shengwu Xiong 0001 |
ACCV (4) | 3 |
| 2018 | Multi-Scale Piecewise Line Integral Strategy for Structure Integral TransformabstractStructure Integral transform (SIT) is a mathematical tool for invariant shape recognition. SIT has a superior ability to capture the interior structure information by integrating the shape image function over 2D dissecting structure, but the details on the dissecting structure are discarded. In this paper, for the first time, we propose a novel multi-scale piecewise line integral strategy for SIT which introduces the grayscale information. At a higher scale, the integral line on dissecting structure will be equally divided more times, which brings to a coarse to fine description for each integral line. The proposed multi-scale piecewise integral strategy has been validated to extend SIT from binary shape to grayscale shape successfully through experiments on Leeds' Butterflies dataset and Ponce Group's Butterflies dataset. Yongsheng Gao 0001, Jishan Guo, Shengwu Xiong 0001 |
ICIP | 3 |
| 2018 | Matching Pursuit Based on Kernel Non-Second Order MinimizationabstractThe orthogonal matching pursuit (OMP) is an important sparse approximation algorithm to recover sparse signals from compressed measurements. However, most MP algorithms are based on the mean square error(MSE) to minimize the recovery error, which is suboptimal when there are outliers. In this paper, we present a new robust OMP algorithm based on kernel non-second order statistics (KNS-OMP), which not only takes advantages of the outlier resistance ability of correntropy but also further extends the second order statistics based correntropy to a non-second order similarity measurement to improve its robustness. The resulted framework is more accurate than the second order ones in reducing the effect of outliers. Experimental results on synthetic and real data show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 2 |
| 2018 | Generative Adversarial Product QuantisationabstractProduct Quantisation (PQ) has been recognised as an effective encoding technique for scalable multimedia content analysis. In this paper, we propose a novel learning framework that enables an end-to-end encoding strategy from raw images to compact PQ codes. The system aims to learn both PQ encoding functions and codewords for content-based image retrieval. In detail, we first design a trainable encoding layer that is pluggable into neural networks, so the codewords can be trained in back-forward propagation. Then we integrate it into a Deep Convolutional Generative Adversarial Network (DC-GAN). In our proposed encoding framework, the raw images are directly encoded by passing through the convolutional and encoding layers, and the generator aims to use the codewords as constrained inputs to generate full image representations that are visually similar to the original images. By taking the advantages of the generative adversarial model, our proposed system can produce high-quality PQ codewords and encoding functions for scalable multimedia retrieval tasks. Experiments show that the proposed architecture GA-PQ outperforms the state-of-the-art encoding techniques on three public image datasets. Litao Yu, Yongsheng Gao 0001, Jun Zhou 0001 |
ACM Multimedia | 2 |
| 2017 | Retinal Vessel Segmentation Using Matched Filter with Joint Relative Entropy
Mohsin Challoob, Yongsheng Gao 0001 |
CAIP (1) | 2 |
| 2017 | Can Walking and Measuring Along Chord Bunches Better Describe Leaf Shapes?abstractEffectively describing and recognizing leaf shapes under arbitrary deformations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing shape by walking along a bunch of chords that pass through the shape to measure the regions trespassed. A novel chord bunch walks (CBW) descriptor is developed through the chord walking that effectively integrates the shape image function over the walked chord to reflect the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chord walks to build a hierarchical framework for a coarse-to-fine description. The proposed CBW descriptor is invariant to rotation, scaling, translation, and mirror transforms. Instead of using the expensive optimal correspondence based matching, an improved Hausdorff distance encoded correspondence information is proposed for efficient yet effective shape matching. In experimental studies, the proposed method obtained substantially higher accuracies with low computational cost over the benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
CVPR | 2 |
| 2017 | Robust Facial Landmark Localization Using LBP Histogram Correlation Based InitializationabstractFacial landmark localization on images with occlusions is an important and challenging task in many visual applications. Recently, the cascaded pose regression has attracted increasing attention, since it achieved superior performance in terms of facial landmark localization under occlusions. However, such approach is sensitive to initialization, where an improper initialization will decrease the performance sharply. In this paper, we propose a novel initialization method to get a robust initial shape by analysing correlation of Local Binary Patterns (LBP) histograms between the estimated face and training faces. The shape of the training face that is most correlated with the estimated face, will be selected as the initialization for the regression. The selected shape is closer to the real shape of the estimated face, which makes the landmark localization more accurate. Besides, in order to make the initial shape more robust to occlusions, we propose a boosted smart restarts technique by checking location and occlusion jointly instead of checking location only. We show that the proposed method significantly improves performance over existing landmark localization methods on the challenging dataset of COFW. The experimental results demonstrate that the proposed method reduces error by 11.9% and failure cases by 20.8% on COFW dataset. Moreover, it detects face occlusions with 85/40% precision/recall. Yiyun Pan, Junwei Zhou 0002, Yongsheng Gao 0001, Jianwen Xiang, Shengwu Xiong 0001, Yanchao Yang 0002 |
FG | 3 |
| 2017 | Robust Image Classification via Low-Rank Double Dictionary Learning
Shengwu Xiong 0001, Yongsheng Gao 0001 |
MMM (1) | 3 |
| 2017 | Surface geodesic pattern for 3D deformable texture matching
Farshid Hajati, Ali Cheraghian, Soheila Gheisari, Yongsheng Gao 0001, Ajmal Mian |
Pattern Recognit. | 4 |
| 2017 | Low-rank double dictionary learning from corrupted data for robust image classification
Shengwu Xiong 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2017 | Sparse 3D directional vertices vs continuous 3D curves: Efficient 3D surface matching and its application for single model face recognition
Xun Yu, Yongsheng Gao 0001, Jun Zhou 0001 |
Pattern Recognit. | 2 |
| 2017 | On the Sampling Strategy for Evaluation of Spectral-Spatial Methods in Hyperspectral Image ClassificationabstractSpectral-spatial processing has been increasingly explored in remote sensing hyperspectral image classification. While extensive studies have focused on developing methods to improve the classification accuracy, experimental setting and design for method evaluation have drawn little attention. In the scope of supervised classification, we find that traditional experimental designs for spectral processing are often improperly used in the spectral-spatial processing context, leading to unfair or biased performance evaluation. This is especially the case when training and testing samples are randomly drawn from the same image - a practice that has been commonly adopted in the experiments. Under such setting, the dependence caused by overlap between the training and testing samples may be artificially enhanced by some spatial information processing methods, such as spatial filtering and morphological operation. Such enhancement of dependence in return amplifies the classification accuracy, leading to an improper evaluation of spectral-spatial classification techniques. Therefore, the widely adopted pixel-based random sampling strategy is not always suitable to evaluate spectral-spatial classification algorithms, because it is difficult to determine whether the improvement of classification accuracy is caused by incorporating spatial information into classifier or by increasing the overlap between training and testing samples. To tackle this problem, we propose a novel controlled random sampling strategy for spectral-spatial methods. It can greatly reduce the overlap between training and testing samples and provides more objective and accurate evaluation. Jie Liang 0003, Jun Zhou 0001, Yuntao Qian, Lian Wen, Xiao Bai 0001, Yongsheng Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | Dynamic Texture Comparison Using Derivative Sparse Representation: Application to Video-Based Face RecognitionabstractVideo-based face, expression, and scene recognition are fundamental problems in human-machine interaction, especially when there is a short-length video. In this paper, we present a new derivative sparse representation approach for face and texture recognition using short-length videos. First, it builds local linear subspaces of dynamic texture segments by computing spatiotemporal directional derivatives in a cylinder neighborhood within dynamic textures. Unlike traditional methods, a nonbinary texture coding technique is proposed to extract high-order derivatives using continuous circular and cylinder regions to avoid aliasing effects. Then, these local linear subspaces of texture segments are mapped onto a Grassmann manifold via sparse representation. A new joint sparse representation algorithm is developed to establish the correspondences of subspace points on the manifold for measuring the similarity between two dynamic textures. Extensive experiments on the Honda/UCSD, the CMU motion of body, the YouTube, and the DynTex datasets show that the proposed method consistently outperforms the state-of-the-art methods in dynamic texture recognition, and achieved the encouraging highest accuracy reported to date on the challenging YouTube face dataset. The encouraging experimental results show the effectiveness of the proposed method in video-based face recognition in human-machine system applications. Farshid Hajati, Mohammad Tavakolian, Soheila Gheisari, Yongsheng Gao 0001, Ajmal Mian |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2016 | Tensor morphological profile for hyperspectral image classificationabstractThis paper proposes a novel multi-dimensional morphology descriptor, tensor morphology profile (TMP), for hyperspectral image classification. TMP is a general framework to extract the multi-dimensional structures in high-dimensional data. The nth-order morphology profile is proposed to work with the nth-order tensor, which can capture the inner high order structures. This is different with the traditional mathematical morphology operations which are usually limited to two-dimensional data. By treating hyperspectral images a tensor, it is possible to extend the morphology to high dimensional data so that the powerful morphological tools can be used to analyze the hyperspectral images with spectral-spatial information fused. Experimental results on two commonly used hyperspectral images show that the tensor morphological profile consistently performs better than the extended morphological profile for hyperspectral image classification. Jie Liang 0003, Jun Zhou 0001, Yongsheng Gao 0001 |
ICIP | 3 |
| 2016 | Discriminative dictionary pair learning from partially labeled dataabstractWhile conventional synthesis dictionary learning approaches have demonstrated tremendous success in various pattern recognition problems, the dictionary pair learning, i.e., jointly learning an analysis dictionary and a synthesis dictionary is still an open problem. Furthermore, the performance of traditional supervised dictionary learning methods is often limited by the amount of labeled training data. In this paper, we propose a novel dictionary pair learning model by utilizing both labeled and unlabeled data for analysis-synthesis dictionary training. In the dictionary learning phase, we integrate the unlabeled samples, whose labels are predicted through an entropy-based method, into their associated classes to increase the amount of the `labeled' data. This strategy promotes the discrimination power of both analysis dictionary and synthesis dictionary. Experimental evaluations on publicly available datasets demonstrate the usefulness of semi-supervised strategy and the effectiveness of the proposed method, especially in the case of limited number of the labeled samples. Shengwu Xiong 0001, Yongsheng Gao 0001 |
ICIP | 3 |
| 2016 | 3D face recognition under partial occlusions using radial stringsabstract3D face recognition with partial occlusions is a highly challenging problem. In this paper, we propose a novel radial string representation and matching approach to recognize 3D facial scans in the presence of partial occlusions. Here we encode 3D facial surfaces into an indexed collection of radial strings emanating from the nosetips and Dynamic Programming (DP) is then used to measure the similarity between two radial strings. In order to address the recognition problems with partial occlusions, a partial matching mechanism is established in our approach that effectively eliminates those occluded parts and finds the most discriminative parts during the matching process. Experimental results on the Bosphorus database demonstrate that the proposed approach yields superior performance on partially occluded data. Xun Yu, Yongsheng Gao 0001, Jun Zhou 0001 |
ICIP | 2 |
| 2016 | Robust tensor factorization using maximum correntropy criterionabstractTraditional tensor decomposition methods, e.g., two dimensional principle component analysis (2DPCA) and two dimensional singular value decomposition (2DSVD), minimize mean square errors (MSE) and are sensitive to outliers. In this paper, we propose a new robust tensor factorization method using maximum correntropy criterion (MCC) to improve the robustness of traditional tensor decomposition methods. A half-quadratic optimization algorithm is adopted to effectively optimize the correntropy objective function in an iterative manner. It can effectively improve the robustness of a tensor decomposition method to outliers without introducing any extra computational cost. Experimental results demonstrated that the proposed method significantly reduces the reconstruction error on face reconstruction and improves the accuracy rate on handwritten digit recognition. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, John La Salle, Junli Liang |
ICPR | 2 |
| 2016 | Face recognition using linear representation ensembles
Fumin Shen, Chunhua Shen, Yang Yang 0002, Yongsheng Gao 0001 |
Pattern Recognit. | 5 |
| 2016 | Nonnegative-Matrix-Factorization-Based Hyperspectral Unmixing With Partially Known EndmembersabstractHyperspectral unmixing is an important technique for estimating fractions of various materials from remote sensing imagery. Most unmixing methods make the assumption that no prior knowledge of endmembers is available before the estimation. This is, however, not true for some unmixing tasks for which part of the endmember signatures may be known in advance. In this paper, we address the hyperspectral unmixing problem with partially known endmembers. We extend nonnegative-matrix-factorization-based unmixing algorithms to incorporate prior information into their models. The proposed approach uses the spectral signature of known endmembers as a constraint, among others, in the unmixing model, and propagates the knowledge by an optimization process which minimizes the difference between the image data and the prior knowledge. Results on both synthetic and real data have validated the effectiveness of the proposed method and have shown that it has outperformed several state-of-the-art methods that use or do not use prior knowledge of endmembers. Jun Zhou 0001, Yuntao Qian, Xiao Bai 0001, Yongsheng Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Structure Integral Transform Versus Radon Transform: A 2D Mathematical Tool for Invariant Shape RecognitionabstractIn this paper, we present a novel mathematical tool, Structure Integral Transform (SIT), for invariant shape description and recognition. Different from the Radon Transform (RT), which integrates the shape image function over a 1D line in the image plane, the proposed SIT builds upon two orthogonal integrals over a 2D K -cross dissecting structure spanning across all rotation angles by which the shape regions are bisected in each integral. The proposed SIT brings the following advantages over the RT: 1) it has the extra function of describing the interior structural relationship within the shape which provides a more powerful discriminative ability for shape recognition; 2) the shape regions are dissected by the K -cross in a coarse to fine hierarchical order that can characterize the shape in a better spatial organization scanning from the center to the periphery; and 3) it is easier to build a completely invariant shape descriptor. The experimental results of applying SIT to shape recognition demonstrate its superior performance over the well-known Radon transform, and the well-known shape contexts and the polar harmonic transforms. Bin Wang 0041, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Delaunay-supported edges for image graphsabstractGraphs are a powerful and versatile data structure for pattern recognition. However, their flexibility brings inherent complexities for algorithms which seek to create or utilise graphs. In the context of computer vision, many feature detection algorithms can extract a suitable vertex set from an image. The creation of edges between these vertices presents new challenges, and effective methods for edge creation only exist for certain types of vertices such as points and regions. This paper presents a novel method for creating edges for image graphs, while supporting a wide array of vertex types. The presented method is principled, and its robustness is shown experimentally against a number of affine and projective transforms, as well as noise. Nicholas Dahm, Yongsheng Gao 0001, Terry Caelli, Horst Bunke |
ICIP | 2 |
| 2015 | Multi-scale bisector integrals: An invariant descriptor for accurate shape retrievalabstractA novel shape descriptor, termed multi-scale bisector integrals, is proposed in this paper. Different from the existing Radon transform based descriptors which integrate the shape image function over all the possible lines in its domain, the proposed method restrains the integrals only over a special class of lines, termed shape bisectors, for characterizing the essence of the shape. Integrating the shape image over the multi-orders of shape bisectors yields a multi-scale descriptor. The proposed descriptor is completely invariant to translation, scaling and rotation. The experimental results on the standard MEPG-7 CE-2 shape database demonstrate its superiority over the state-of-the-art approaches. Bin Wang 0041, Yongsheng Gao 0001 |
ICIP | 2 |
| 2015 | 3D Reconstruction from Hyperspectral Imagesabstract3D reconstruction from hyper spectral images has seldom been addressed in the literature. This is a challenging problem because 3D models reconstructed from different spectral bands demonstrate different properties. If we use a single band or covert the hyper spectral image to gray scale image for the reconstruction, fine structural information may be lost. In this paper, we present a novel method to reconstruct a 3D model from hyper spectral images. Our proposed method first generates 3D point sets from images at each wavelength using the typical structure from motion approach. A structural descriptor is developed to characterize the spatial relationship between the points, which allows robust point matching between two 3D models at different wavelength. Then a 3D registration method is introduced to combine all band-level models into a single and complete hyper spectral 3D model. As far as we know, this is the first attempt in reconstructing a complete 3D model from hyper spectral images. This work allows fine structural-spectral information of an object be captured and integrated into the 3D model, which can be used to support further research and applications. Ali Zia, Jie Liang 0003, Jun Zhou 0001, Yongsheng Gao 0001 |
WACV | 4 |
| 2015 | MARCH: Multiscale-arch-height description for mobile retrieval of leaf images
Bin Wang 0041, Douglas Brown, Yongsheng Gao 0001, John La Salle |
Inf. Sci. | 3 |
| 2015 | Efficient subgraph matching using topological node feature constraints
Nicholas Dahm, Horst Bunke, Terry Caelli, Yongsheng Gao 0001 |
Pattern Recognit. | 4 |
| 2014 | Face Recognition Using 3D Directional Corner PointsabstractIn this paper, we present a novel face recognition approach using 3D directional corner points (3D DCPs). Traditionally, points and meshes are applied to represent and match 3D shapes. Here we represent 3D surfaces by 3D DCPs derived from ridge and valley curves. Then we develop a 3D DCP matching method to compute the similarity of two different 3D surfaces. This representation, along with the similarity metric can effectively integrate structural and spatial information on 3D surfaces. The added information can provide more and better discriminative power for object recognition. It strengthens and improves the matching process of similar 3D objects such as faces. To evaluate the performance of our method for 3D face recognition, we have performed experiments on Face Recognition Grand Challenge v2.0 database (FRGC v2.0) and resulted in a rank-one recognition rate of 97.1%. This study demonstrates that 3D DCPs provides a new solution for 3D face recognition, which may also find its application in general 3D object representation and recognition. Xun Yu, Yongsheng Gao 0001, Jun Zhou 0001 |
ICPR | 2 |
| 2014 | Polygonal approximation using integer particle swarm optimization
Bin Wang 0041, Douglas Brown, Xiaozheng Zhang 0002, Yongsheng Gao 0001, Jie Cao 0001 |
Inf. Sci. | 5 |
| 2014 | Hierarchical String Cuts: A Translation, Rotation, Scale, and Mirror Invariant Descriptor for Fast Shape RetrievalabstractThis paper presents a novel approach for both fast and accurately retrieving similar shapes. A hierarchical string cuts (HSC) method is proposed to partition a shape into multiple level curve segments of different lengths from a point moving around the contour to describe the shape gradually and completely from the global information to the finest details. At each hierarchical level, the curve segments are cut by strings to extract features that characterize the geometric and distribution properties in that particular level of details. The translation, rotation, scale and mirror invariant HSC descriptor enables a fast metric based matching to achieve the desired high accuracy. Encouraging experimental results on four databases demonstrated that the proposed method can consistently achieve higher (or similar) retrieval accuracies than the state-of-the-art benchmarks with a more than 120 times faster speed. This may suggest a new way of developing shape retrieval techniques in which a high accuracy can be achieved by a fast metric matching algorithm without using the time-consuming correspondence optimisation strategy. Bin Wang 0041, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | 3D face recognition using topographic high-order derivativesabstractThis paper presents a novel feature, Topographic High-order Derivatives (THD) for 3D face recognition. THD is based on the high-order micro-pattern information extracted from face topography maps. Face topography maps are partitioned into polar sectors, and THDs are computed using directional highorder derivatives within the sectors. Local features are extracted by encoding directional high-order derivatives within polar neighborhoods. To evaluate the proposed method, we use Bosphorus and FRGC 3D face databases which include pose and expression changes. The performance of the proposed method is higher compared to the state-of-the-art benchmark approaches in 3D face recognition. Ali Cheraghian, Farshid Hajati, Ajmal Mian, Yongsheng Gao 0001, Soheila Gheisari |
ICIP | 4 |
| 2013 | Matching non-aligned objects using a relational string-graphabstractLocalising and aligning objects is a challenging task in computer vision that still remains largely unsolved. Utilising the syntactic power of graph representation, we define a relational string-graph matching algorithm that seeks to perform these tasks simultaneously. By matching the relations between vertices, where vertices represent high-level primitives, the relational string-graph is able to overcome the noisy and inconsistent nature of the vertices themselves. For each possible relation correspondence between two graphs, we calculate the rotation, translation, and scale parameters required to transform a relation into its counterpart. We plot these parameters in 4D space and use Gaussian mixture models and the expectation-maximisation algorithm to estimate the underlying parameters. Our method is tested on face alignment and recognition, but is equally (if not more) applicable for generic object alignment. Nicholas Dahm, Yongsheng Gao 0001, Terry Caelli, Horst Bunke |
ICIP | 2 |
| 2013 | Mobile plant leaf identification using smart-phonesabstractA novel shape description method is proposed for mobile retrieval of leaf images to aid in plant recognition. In this method, traveling the shape contour, the convexity and concavity properties of the arches of various levels are measured, respectively, to generate a multiscale shape descriptor. Its performance has been tested on two leaf datasets and the experimental results indicated higher recognition accuracies than the state-of-the-art approaches with a speed improvement of more than 170 times. The proposed method has been successfully applied to develop a prototype system of online plant leaf identification working on a consumer mobile platform. Bin Wang 0041, Douglas Brown, Yongsheng Gao 0001, John La Salle |
ICIP | 3 |
| 2013 | Face Recognition Using Ensemble String MatchingabstractIn this paper, we present a syntactic string matching approach to solve the frontal face recognition problem. String matching is a powerful partial matching technique, but is not suitable for frontal face recognition due to its requirement of globally sequential representation and the complex nature of human faces, containing discontinuous and non-sequential features. Here, we build a compact syntactic Stringface representation, which is an ensemble of strings. A novel ensemble string matching approach that can perform non-sequential string matching between two Stringfaces is proposed. It is invariant to the sequential order of strings and the direction of each string. The embedded partial matching mechanism enables our method to automatically use every piece of non-occluded region, regardless of shape, in the recognition process. The encouraging results demonstrate the feasibility and effectiveness of using syntactic methods for face recognition from a single exemplar image per person, breaking the barrier that prevents string matching techniques from being used for addressing complex image recognition problems. The proposed method not only achieved significantly better performance in recognizing partially occluded faces, but also showed its ability to perform direct matching between sketch faces and photo faces. Weiping Chen, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Fast and Effective Retrieval of Plant Leaf Shapes
Bin Wang 0041, Yongsheng Gao 0001 |
ACCV (2) | 2 |
| 2012 | Parametric Manifold of an Object under Different Viewing Directions
Xiaozheng Zhang 0002, Yongsheng Gao 0001, Terry Caelli |
ECCV (5) | 2 |
| 2012 | Locality-Regularized Linear Regression for face recognition
Douglas Brown, Yongsheng Gao 0001 |
ICPR | 3 |
| 2012 | Topological features and iterative node elimination for speeding up subgraph isomorphism detection
Nicholas Dahm, Horst Bunke, Terry Caelli, Yongsheng Gao 0001 |
ICPR | 4 |
| 2012 | Recognition of driving postures by multiwavelet transform and multilayer perceptron classifier
Chihang Zhao, Yongsheng Gao 0001, Jie He 0006 |
Eng. Appl. Artif. Intell. | 2 |
| 2012 | 2.5D face recognition using Patch Geodesic Moments
Farshid Hajati, Abolghasem A. Raie, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2012 | Heterogeneous Specular and Diffuse 3-D Surface Approximation for Face Recognition Across PoseabstractThis paper proposes a novel heterogeneous specular and diffuse (HSD) 3-D surface approximation which considers spatial variability of specular and diffuse reflections in face modelling and recognition. Traditional 3-D face modelling and recognition methods constrain human faces with either the Lambertian assumption or the homogeneity assumption, resulting in suboptimal shape and texture models. The proposed HSD approach allows both specular and diffuse reflectance coefficients to vary spatially to better accommodate surface properties of real human faces. From a small number of face images of a person under different lighting conditions, 3-D shape and surface reflectivity property are estimated using a localized stochastic optimization method. The resultant personalized 3-D face model is used to render novel gallery views under different poses for recognition across pose. The proposed approach is evaluated on both synthetic and real face datasets and benchmarked against the state-of-the-art approaches. Experimental results demonstrated that it can achieve a higher level of performances in modelling accuracy, algorithm reliability, and recognition accuracy, which suggests that face modelling and recognition beyond the Lambertian and homogeneity assumptions is a feasible and better solution towards pose-invariant face recognition. Xiaozheng Zhang 0002, Yongsheng Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Study on the BeiHang Keystroke Dynamics DatabaseabstractThis paper introduces a new BeiHang (BH) Keystroke Dynamics Database for testing and evaluation of biometric approaches. Different from the existing keystroke dynamics researches which solely rely on laboratory experiments, the developed database is collected from a real commercialized system and thus is more comprehensive and more faithful to human behavior. Moreover, our database comes with ready-to-use benchmark results of three keystroke dynamics methods, Nearest Neighbor classifier, Gaussian Model and One-Class Support Vector Machine. Both the database and benchmark results are open to the public and provide a significant experimental platform for international researchers in the keystroke dynamics area. Baochang Zhang 0001, Yao Cao, Sanqiang Zhao, Yongsheng Gao 0001, Jianzhuang Liu |
IJCB | 5 |
| 2011 | Local Kernel Feature Analysis (LKFA) for object recognition
Baochang Zhang 0001, Yongsheng Gao 0001 |
Neurocomputing | 2 |
| 2011 | Kernel Similarity Modeling of Texture Pattern Flow for Motion Detection in Complex BackgroundabstractThis paper proposes a novel kernel similarity modeling of texture pattern flow (KSM-TPF) for background modeling and motion detection in complex and dynamic environments. The texture pattern flow encodes the binary pattern changes in both spatial and temporal neighborhoods. The integral histogram of texture pattern flow is employed to extract the discriminative features from the input videos. Different from existing uniform threshold based motion detection approaches which are only effective for simple background, the kernel similarity modeling is proposed to produce an adaptive threshold for complex background. The adaptive threshold is computed from the mean and variance of an extended Gaussian mixture model. The proposed KSM-TPF approach incorporates machine learning method with feature extraction method in a homogenous way. Experimental results on the publicly available video sequences demonstrate that the proposed approach provides an effective and efficient way for background modeling and motion detection. Baochang Zhang 0001, Yongsheng Gao 0001, Sanqiang Zhao, Bineng Zhong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Recognizing Partially Occluded Faces from a Single Sample Per Class Using String-Based Matching
Weiping Chen, Yongsheng Gao 0001 |
ECCV (3) | 2 |
| 2010 | Recognizing face profiles in the presence of hairs/glasses interferencesabstractFacial profile provides a complementary structure of the face that is not present in frontal faces, which has been used in personal identification, face perception research and 3D face construction. In this paper, we present a novel local attributed string matching (LAStrM) approach to recognize face profiles in the presence of interferences. The conventional profile recognition algorithms heavily depend on the accuracy of the facial area cropping. However, in realistic scenarios the facial area may be difficult to localize due to interferences (e.g., glasses, hairstyles). The proposed approach is able to efficiently find the most discriminative local parts between face profiles addressing the recognition problem with interferences. Experimental results have shown that the proposed matching scheme is robust to interferences compared against several primary approaches using two profile image databases (Bern and FERET). It has potential capability for partially occluded shape classification. Weiping Chen, Yongsheng Gao 0001 |
ICARCV | 2 |
| 2010 | Pose-invariant 2.5D face recognition using Geodesic Texture WarpingabstractIn recent years, 3D face recognition has become a popular solution to deal with the problem of pose-invariant face recognition. The majority of 3D face data are, however, actually 2.5D which are sensitive to pose variations. This paper presents a novel Geodesic Texture Warping (GTW) solution for 2.5D pose-invariant face recognition. In this method, we use the geodesic distance computed on a 2.5D face scan to warp the texture of a rotated face to that of a frontal one to perform matching. A feasibility and effectiveness investigation for the proposed method is conducted using a wide range of experiments including samples with different face rotations. The encouraging experimental results demonstrate that the proposed method achieves much higher accuracy than the state-of-the-art method with a low computational cost. Farshid Hajati, Abolghasem A. Raie, Yongsheng Gao 0001 |
ICARCV | 3 |
| 2010 | Primitive-based 3D structure inference from a single 2D image for insect modeling: Towards an electronic field guide for insect identificationabstract3D insect models are useful to overcome viewing angle variations and self-occlusions in computer-assisted insect taxonomy for electronic field guides. The acquisition of 3D information is, however, unreliable due to the flexibility and small size of the insect bodies. This paper explores how to infer 3D insect models from a single 2D insect image, which will assist both insect description and identification. The 3D structure of the insect body is modeled from two geometric primitives, generalized cylinders and deformable ellipsoids. The primitives are fitted and warped based on both edge and medial axis constraints of the 2D image. Individualized 3D models are then built to approximate the insect structure. The proposed approach results in seemingly useful 3D insect models capable of representing the major morphological characteristics for a variety of insects with different body types. This method could be a helpful assistance for computer-assisted insect taxonomy and insect identification by entomologists and the public. Xiaozheng Zhang 0002, Yongsheng Gao 0001, Terry Caelli |
ICARCV | 2 |
| 2010 | A Novel Pose Invariant Face Recognition Approach Using a 2D-3D Searching StrategyabstractMany Face Recognition techniques focus on 2D-2D comparison or 3D-3D comparison, however few techniques explore the idea of cross-dimensional comparison. This paper presents a novel face recognition approach that implements cross-dimensional comparison to solve the issue of pose invariance. Our approach implements a Gabor representation during comparison to allow for variations in texture, illumination, expression and pose. Kernel scaling is used to reduce comparison time during the branching search, which determines the facial pose of input images. The conducted experiments prove the viability of this approach, with our larger kernel experiments returning 91.6% - 100% accuracy on a database comprised of both local data, and data from the USF Human ID 3D database. Nicholas Dahm, Yongsheng Gao 0001 |
ICPR | 2 |
| 2010 | Gender Classification Using Interlaced Derivative PatternsabstractAutomated gender recognition has become an interesting and challenging research problem in recent years with its potential applications in security industry and human-computer interaction systems. In this paper we present a novel feature representation, namely Interlaced Derivative Patterns (IDP), which is a derivative-based technique to extract discriminative facial features for gender classification. The proposed technique operates on a neighborhood around a pixel and concatenates the extracted regional feature distributions to form a feature vector. The experimental results demonstrate the effectiveness of the IDP method for gender classification, showing that the proposed approach achieves 29.6% relative error reduction compared to Local Binary Patterns (LBP), while it performs over four times faster than Local Derivative Patterns (LDP). Ameneh Shobeirinejad, Yongsheng Gao 0001 |
ICPR | 2 |
| 2010 | High-Order Circular Derivative Pattern for Image Representation and RecognitionabstractMicropattern based image representation and recognition, e.g. Local Binary Pattern (LBP), has been proved successful over the past few years due to its advantages of illumination tolerance and computational efficiency. However, LBP only encodes the first-order radial-directional derivatives of spatial images and is inadequate to completely describe the discriminative features for classification. This paper proposes a new Circular Derivative Pattern (CDP) which extracts high-order derivative information of images along circular directions. We argue that the high-order circular derivatives contain more detailed and more discriminative information than the first-order LBP in terms of recognition accuracy. Experimental evaluation through face recognition on the FERET database and insect classification on the NICTA Biosecurity Dataset demonstrated the effectiveness of the proposed method. Sanqiang Zhao, Yongsheng Gao 0001, Terry Caelli |
ICPR | 2 |
| 2010 | Performance Evaluation of Micropattern Representation on Gabor Features for Face RecognitionabstractFace recognition using micropattern representation has recently received much attention in the computer vision and pattern recognition community. Previous researches demonstrated that micropattern representation based on Gabor features achieves better performance than its direct usage on gray-level images. This paper conducts a comparative performance evaluation of micropattern representations on four forms of Gabor features for face recognition. Three evaluation rules are proposed and observed for a fair comparison. To reduce the high feature dimensionality problem, uniform quantization is used to partition the spatial histograms. The experimental results reveal that: 1) micropattern representation based on Gabor magnitude features outperforms the other three representations, and the performances of the other three are comparable; and 2) micropattern representation based on the combination of Gabor magnitude and phase features performs the best. Sanqiang Zhao, Yongsheng Gao 0001, Baochang Zhang 0001 |
ICPR | 2 |
| 2010 | Local Derivative Pattern Versus Local Binary Pattern: Face Recognition With High-Order Local Pattern DescriptorabstractThis paper proposes a novel high-order local pattern descriptor, local derivative pattern (LDP), for face recognition. LDP is a general framework to encode directional pattern features based on local derivative variations. The n(th)-order LDP is proposed to encode the (n-1)(th) -order local derivative direction variations, which can capture more detailed information than the first-order local pattern used in local binary pattern (LBP). Different from LBP encoding the relationship between the central point and its neighbors, the LDP templates extract high-order local information by encoding various distinctive spatial relationships contained in a given local region. Both gray-level images and Gabor feature images are used to evaluate the comparative performances of LDP and LBP. Extensive experimental results on FERET, CAS-PEAL, CMU-PIE, Extended Yale B, and FRGC databases show that the high-order LDP consistently performs much better than LBP for both face identification and face verification under various conditions. Baochang Zhang 0001, Yongsheng Gao 0001, Sanqiang Zhao, Jianzhuang Liu |
IEEE Trans. Image Process. | 2 |
| 2010 | General Retinal Vessel Segmentation Using Regularization-Based Multiconcavity ModelingabstractDetecting blood vessels in retinal images with the presence of bright and dark lesions is a challenging unsolved problem. In this paper, a novel multiconcavity modeling approach is proposed to handle both healthy and unhealthy retinas simultaneously. The differentiable concavity measure is proposed to handle bright lesions in a perceptive space. The line-shape concavity measure is proposed to remove dark lesions which have an intensity structure different from the line-shaped vessels in a retina. The locally normalized concavity measure is designed to deal with unevenly distributed noise due to the spherical intensity variation in a retinal image. These concavity measures are combined together according to their statistical distributions to detect vessels in general retinal images. Very encouraging experimental results demonstrate that the proposed method consistently yields the best performance over existing state-of-the-art methods on the abnormal retinas and its accuracy outperforms the human observer, which has not been achieved by any of the state-of-the-art benchmark methods. Most importantly, unlike existing methods, the proposed method shows very attractive performances not only on healthy retinas but also on a mixture of healthy and pathological retinas. Benson S. Y. Lam, Yongsheng Gao 0001, Alan Wee-Chung Liew |
IEEE Trans. Medical Imaging | 2 |
| 2009 | Textural Hausdorff Distance for wider-range tolerance to pose variation and misalignment in 2D face recognitionabstractThis paper addresses two critical but rarely concerned issues in 2D face recognition: wider-range tolerance to pose variation and misalignment. We propose a new Textural Hausdorff Distance (THD), which is a compound measurement integrating both spatial and textural features. The THD is applied to a Significant Jet Point (SJP) representation of face images, where a varied number of shape-driven SJPs are detected automatically from low-level edge map with rich information content. The comparative experiments conducted on publicly available FERET and AR face databases demonstrated that the proposed approach has a considerably wider range of tolerance against both in-depth head rotation and face misalignment. Sanqiang Zhao, Yongsheng Gao 0001 |
CVPR | 2 |
| 2009 | Recognition of expression variant faces from one sample image per enrolled subjectabstractDespite remarkable progress on face recognition, little attention has been given to robustly recognize expression variant faces from a single sample image per person. One way to deal with the recognition of faces under above conditions is by using local statistical approaches which appear to be more robust against variations in facial expression. In this paper, we propose a new weighted matching method based on our recent work of AWPPZMA to recognize expression variant faces when only one exemplar image per enrolled subject is available. The proposed weighting method gives more significance to those parts of the face with facial expression variations that change less compared to neutral face image and less significance to those parts that change more. In this contribution, we use the difference between local area in the input face and its corresponding local area in the neutral face image as a measure of observable structure changes. The encouraging experimental results demonstrate that the proposed method provides a new solution to the problem of robustly recognizing expression variant faces in single model databases. Hamidreza Rashidy Kanan, Yongsheng Gao 0001 |
ICIP | 2 |
| 2009 | Generalised ambient reflection models for Lambertian and Phong surfacesabstractAmbient reflection is widely present in many applications of computer graphics and image processing, which is traditionally modelled as a constant free from environmental factors. This paper reconsiders ambient reflection modelling of Lambertian and Phong surfaces and calculates it as reflection integrations of infinitesimal incident beams from the environment. It reveals that ambient reflection exists in variable forms for Lambertian surfaces of non-convex objects and Phong surfaces of all objects. For convex objects with Lambertian surfaces, ambient reflectance coefficient is actually the diffuse reflectance coefficient. Generalised ambient reflection models are proposed to calculate ambient reflection using the same reflection model as used in calculation of other reflections. Based on this analysis, new ambient reflection formulations of Lambertian and Phong surfaces are derived to enable efficient computations in computer graphics and image processing. Xiaozheng Zhang 0002, Yongsheng Gao 0001 |
ICIP | 2 |
| 2009 | Face recognition across pose: A review
Xiaozheng Zhang 0002, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2009 | Gabor feature constrained statistical model for efficient landmark localization and face recognition
Sanqiang Zhao, Yongsheng Gao 0001, Baochang Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2008 | Face recognition based on Gradient Gabor featureabstractIn this paper, a novel gradient Gabor (GGabor) filter is proposed to extract multi-scale and multi-orientation features to represent and classify faces. Gradient Gabor combines the derivative of Gaussian functions and the harmonic functions to capture the features in both spatial and frequency domains to deliver orientation and scale information. The spatial positions are combined into Gaussian derivatives which allows it to provide more stable information. An efficient Kernel Fisher analysis method is proposed to find multiple subspaces based on both GGabor magnitude and phase features, which is a local kernel mapping method to capture the structure information in faces. Experiments on two face databases, FRGC Version 1 and FRGC Version 2, are conducted to compare the performances of the Gabor and GGabor features, which show that GGabor can also be a powerful tool to model faces, and the Efficient Kernel Fisher classifier can improve the efficiency of the original kernel fisher method. Baochang Zhang 0001, Yongsheng Gao 0001, Yu Qiao 0001 |
ICIP | 2 |
| 2008 | Significant jet point for facial image representation and recognitionabstractGabor wavelet related feature extraction and classification is an important topic in image analysis and pattern recognition. Gabor features can be used either holistically or analytically. While holistic approaches involve significant computational complexity, existing analytic approaches require explicit correspondence of predefined feature points for classification. Different from these approaches, this paper presents a new analytic Gabor method for face recognition. The proposed method attaches Gabor features on a set of shape-driven sparse points to describe both geometric and textural information. Neither the number nor the correspondence of these points is needed. A variant of Hausdorff distance is employed to recognize faces. The experiments performed on AR database demonstrated that the proposed algorithm is effective to identify individuals in various circumstances, such as under expression and illumination changes. Sanqiang Zhao, Yongsheng Gao 0001 |
ICIP | 2 |
| 2008 | Sobel-LBPabstractThis paper presents a new Sobel-LBP, an extension of existing Local Binary Pattern (LBP), for facial image representation. The face image is filtered by Sobel operator to enhance the edge information. Sobel-LBP feature distributions are then extracted and concatenated into a spatial histogram to be used as a face descriptor. The proposed method is compared with the original LBP on both gray-level images and Gabor real and imaginary features for face recognition. The experimental results indicate that Sobel-LBP provides a significantly better performance than LBP under various conditions. Sanqiang Zhao, Yongsheng Gao 0001, Baochang Zhang 0001 |
ICIP | 2 |
| 2008 | Complex background modeling and motion detection based on Texture Pattern FlowabstractThis paper proposes a novel texture pattern flow (TPF) for complex background modeling and motion detection. The pattern flow is proposed to encode the binary pattern changes among the neighborhoods in the space-time domain. To model the distribution of the TPF, the TPF integral histograms are used to extract the discriminative features to represent the input video. Experimental results on the public videos testify the effectiveness of the proposed method in comparison to LBP and GMM based background modeling methods. Baochang Zhang 0001, Yongsheng Gao 0001, Bineng Zhong 0001 |
ICPR | 2 |
| 2008 | On averaging face images for recognition under pose variationsabstractRecently, psychological studies showed that averaging human face images greatly improves the performance of face recognition under various pose, illumination, expression, and/or aging conditions. This paper investigates quantitatively the mechanism of the face averaging process in face recognition specifically against pose variations. Facilitated with 3D face dataset, the process of face averaging is tested on face images free from human errors and misalignments. Single images are chosen as gallery and the averaged views as probe in all the experiments. Three different scenarios are experimented, i.e., identification using single gallery images, identification using different gallery images, and averaging using unbalanced range of input image. The experimental results show that the averaging process under pose variations is equivalent to generating a face view in an average pose and the improvement in face recognition is subject to the conditions that the gallery pose is close to the average probe pose. Xiaozheng Zhang 0002, Sanqiang Zhao, Yongsheng Gao 0001 |
ICPR | 3 |
| 2008 | Establishing point correspondence using multidirectional binary pattern for face recognitionabstractThis paper presents a new Multidirectional Binary Pattern (MBP) for face recognition. Different from most Local Binary Pattern (LBP) related approaches which cluster LBP occurrences from whole image or partitioned subimage patches and use single or concatenated histogram measurement for recognition, MBP is applied on a sparse set of shape-driven points. The new representation is designed for describing both global structure and local texture, and also significantly reduces the high dimensionality of LBP histogram description. Composed of binary patterns from multiple directions, MBP is capable of extracting more discriminative features than LBP. The experiments on face recognition demonstrated the effectiveness of the proposed algorithm against expression and lighting variations. Sanqiang Zhao, Yongsheng Gao 0001 |
ICPR | 2 |
| 2008 | A comparative evaluation of Average Face on holistic and local face recognition approachesabstractThis study focuses on a recent paper “100% Accuracy in Automatic Face Recognition” published on Science, in which an “Average Face” is proposed and claimed to be capable of dramatically improving performance of a face recognition system. To reveal its working mechanism, we perform the averaging process using pose-varied synthetic images generated from 3D face database and conduct a comparative study to observe its effectiveness on holistic and local face recognition approaches. Two representative methods, i.e. Eigenface and Local Binary Pattern (LBP) are employed to perform the experiments. It is interesting to find from our experiments that the performance of the “Average Face” is not independent of the face recognition approaches. Although face averaging increases the recognition accuracy of Eigenface method, it impairs the performance of LBP method. Sanqiang Zhao, Xiaozheng Zhang 0002, Yongsheng Gao 0001 |
ICPR | 3 |
| 2008 | Face recognition using adaptively weighted patch PZM array from a single exemplar image per person
Hamidreza Rashidy Kanan, Karim Faez, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2008 | Recognizing Rotated Faces From Frontal and Side Views: An Approach Toward Effective Use of Mugshot DatabasesabstractMug shot photography has been used to identify criminals by the police for more than a century. However, the common scenario of face recognition using frontal and side-view mug shots as gallery remains largely uninvestigated in computerized face recognition across pose. This paper presents a novel appearance-based approach using frontal and sideface images to handle pose variations in face recognition, which has great potential in forensic and security applications involving police mugshot databases. Virtual views in different poses are generated in two steps: 1) shape modelling and 2) texture synthesis. In the shape modelling step, a multilevel variation minimization approach is applied to generate personalized 3-D face shapes. In the texture synthesis step, face surface properties are analyzed and virtual views in arbitrary viewing conditions are rendered, taking diffuse and specular reflections into account. Appearance-based face recognition is performed with the augmentation of synthesized virtual views covering possible viewing angles to recognize probe views in arbitrary conditions. The encouraging experimental results demonstrated that the proposed approach by using frontal and side-view images is a feasible and effective solution to recognizing rotated faces, which can lead to a better and practical use of existing forensic databases in computerized human face-recognition applications. Xiaozheng Zhang 0002, Yongsheng Gao 0001, Maylor K. H. Leung |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | On transforming statistical models for non-frontal face verification
Conrad Sanderson, Samy Bengio, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2005 | Feature-Level Fusion in Personal IdentificationabstractThe existing studies of multi-modal and multi-view personal identification focused on combining the outputs of multiple classifiers at the decision level. In this study, we investigated the fusion at the feature level to combine multiple views and modals in personal identification. A new similarity measure is proposed, which integrates multiple 2D view features representing a visual identity of a 3D object seen from different viewpoints and from different sensors. The robustness to non-rigid distortions is achieved by the proximity correspondence manner in the similarity computation. The feasibility and capability of the proposed technique for personal identification were evaluated on multiple view human faces and palmprints. This research demonstrates that the feature-level fusion provides a new way to combine multiple modals and views for personal identification. Yongsheng Gao 0001, Michael Maggs |
CVPR (1) | 1 |
| 2005 | Fast Screening in Large Face Databases Using Merit-Based Dominant PointsabstractCurrent face identification approaches require computer systems to search through large quantity of face feature sets in the database and pick the ones that best match the features of an unknown input face. In this paper, a fast screening method for large face database searching is proposed. The method utilizes dominant points instead of edge maps as features for similarity measurement. A new formulation of Hausdorff distance is designed for merit-based dominant point matching. The screening experiments demonstrated that the proposed face screening method significantly improves the computational speed and the storage economy. It provides a very efficient way for large face databases searching and screening. Yongsheng Gao 0001 |
MMM | 1 |
| 2005 | Multilevel Quadratic Variation Minimization for 3D Face Modeling and Virtual View SynthesisabstractOne of the key remaining problems in face recognition is that of handling the variability in appearance due to changes in pose. One strategy is to synthesize virtual face views from real views. In this paper, a novel 3D face shape-modeling algorithm, Multilevel Quadratic Variation Minimization (MQVM), is proposed. Our method makes sole use of two orthogonal real views of a face, i.e., the frontal and profile views. By applying quadratic variation minimization iteratively in a coarse-to-fine hierarchy of control lattices, the MQVM algorithm can generate C²-smooth 3D face surfaces. Then realistic virtual face views can be synthesized by rotating the 3D models. The algorithm works properly on sparse constraint points and large images. It is much more efficient than single-level quadratic variation minimization. The modeling results suggest the validity of the MQVM algorithm for 3D face modeling and 2D face view synthesis under different poses. Xiaozheng Zhang 0002, Yongsheng Gao 0001, Maylor K. H. Leung |
MMM | 2 |
| 2005 | Robust visual similarity retrieval in single model face databases
Yongsheng Gao 0001, Yutao Qi |
Pattern Recognit. | 1 |
| 2004 | Single Model Face Database Retrieval by Directional Corner PointsabstractThis paper presents a new face image retrieval approach using directional corner points (DCPs). Though much work is being done on face similarity matching techniques, little attention is given to the design of face matching scheme suitable for visual retrieval in single model databases where accuracy, robustness to scale and environmental changes, and computational efficiency are three important issues to be concerned. This research demonstrates that the proposed DCP approach provides a new solution, which is both robust to scale and environmental changes, and efficient in computation, for retrieving human faces in single model databases. Yongsheng Gao 0001 |
MMM | 1 |
| 2003 | Noisy logo recognition using line segment Hausdorff distance
Maylor K. H. Leung, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2003 | Facial expression recognition from line-based caricaturesabstractThe automatic recognition of facial expression presents a significant challenge to the pattern analysis and man-machine interaction research community. Recognition from a single static image is particularly a difficult task. In this paper, we present a methodology for facial expression recognition from a single static image using line-based caricatures. The recognition process is completely automatic. It also addresses the computational expensive problem and is thus suitable for real-time applications. The proposed approach uses structural and geometrical features of a user sketched expression model to match the line edge map (LEM) descriptor of an input face image. A disparity measure that is robust to expression variations is defined. The effectiveness of the proposed technique has been evaluated and promising results are obtained. This work has proven the proposed idea that facial expressions can be characterized and recognized by caricatures. Yongsheng Gao 0001, Maylor K. H. Leung, Siu Cheung Hui, M. W. Tananda |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2002 | Motion analysis for face capturingabstractFace recognition is one of the most suitable biometric methods for supporting systems that require high security. The success of a face recognition system involves more than just the comparison algorithm. Face data acquisition is the first step towards developing face recognition systems. The accuracy of a face recognition system greatly depends on the face data acquired. This paper proposes the head-nodding method for face capturing using motion analysis. By analyzing the head-nodding motion of the user, the head-nodding method captures face images with consistent orientation to support face verification systems. In this paper, the design and implementation of the head-nodding method is presented. Preliminary experimental results are also presented. Siu Cheung Hui, Maylor K. H. Leung, Yongsheng Gao 0001 |
ICARCV | 4 |
| 2002 | 3D face modeling using image warping in pose-invariant face recognitionabstractOne of the key remaining problems in face recognition is that of handling the variability in appearance due to changes in pose. One strategy is to synthesize virtual face views from real views. In this paper, a new 3D face shape modeling algorithm - orthogonal reverse projection model (ORPM) is presented. Our method makes only use of two orthogonal real view of a face, i.e., the frontal and profile views. To obtain the 3D shape information of the face, a modified feature-based 2D image warping technique is employed to establish point correspondence between the frontal and profile views. Then the 3D face shape model is generated by orthogonal reverse projection using the corresponding point pairs. The virtual views can be synthesized by simply rotating the 3D face model. The modeling results suggest the validity of the ORPM algorithm for 3D face modeling and 2D face image synthesizing under different poses. Xiaozheng Zhang 0002, Yongsheng Gao 0001, Maylor K. H. Leung |
ICARCV | 2 |
| 2002 | Combined Classification of Multiple Views Using Facial CornersabstractThe profile view of a face provides a complementary structure that is not seen in the frontal view. The classification system combining both frontal and profile views of faces can improve the classification accuracy. And it would be more foolproof because it is difficult to fool the profile face identification by a mask. This paper proposes a new face recognition approach, which can be applied on both frontal and profile faces, to build a robust combined multiple view face identification system. The recognition employs a novel facial corner coding and matching method, and integrates the outline and interior facial parts in the profile matching. The proposed multiview modified Hausdorff distance fuses multiple views of faces to achieve an improved system performance. Yongsheng Gao 0001, Maylor K. H. Leung |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2002 | Face Recognition Using Line Edge MapabstractThe automatic recognition of human faces presents a significant challenge to the pattern recognition research community. Typically, human faces are very similar in structure with minor differences from person to person. They are actually within one class of "human face". Furthermore, lighting conditions change, while facial expressions and pose variations further complicate the face recognition task as one of the difficult problems in pattern analysis. This paper proposes a novel concept: namely, that faces can be recognized using a line edge map (LEM). The LEM, a compact face feature, is generated for face coding and recognition. A thorough investigation of the proposed concept is conducted which covers all aspects of human face recognition, i.e. face recognition under (1) controlled/ideal conditions and size variations, (2) varying lighting conditions, (3) varying facial expressions, and (4) varying pose. The system performance is also compared with the eigenface method, one of the best face recognition techniques, and with reported experimental results of other methods. A face pre-filtering technique is proposed to speed up the search process. It is a very encouraging to find that the proposed face recognition technique has performed better than the eigenface method in most of the comparison experiments. This research demonstrates that the LEM, together with the proposed generic line-segment Hausdorff distance measure, provides a new method for face coding and recognition. Yongsheng Gao 0001, Maylor K. H. Leung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Human face profile recognition using attributed string
Yongsheng Gao 0001, Maylor K. H. Leung |
Pattern Recognit. | 1 |
| 2002 | Line segment Hausdorff distance on face matching
Yongsheng Gao 0001, Maylor K. H. Leung |
Pattern Recognit. | 1 |