Yongxin Ge

dblp:77/8286 · DBLP profile ↗
← Back
50ranked-venue papers
8as first author
30since 2021 · last 2027
0000-0003-3266-1009ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2027 Semantic-aware prototype optimization for test-time adaptive semantic segmentation
Guangrui Li, Jianghao Zhou, Yongxin Ge
Inf. Sci.5
2026 Schema-free information extraction method based on dynamic structure generation from text content
Guanghui Chang, Kuo Chen, Yongxin Ge
Eng. Appl. Artif. Intell.4
2026 Corrigendum to "Schema-free information extraction method based on dynamic structure generation from text content" [Eng. Appl. Artif. Intell. 166 Part A (2026) 113645 ISSN 0952-1976]
Guanghui Chang, Kuo Chen, Yongxin Ge
Eng. Appl. Artif. Intell.4
2026 Adapting to the unknown: Class relational divergence for open set test-time adaptation
Guangrui Li 0005, Jianghao Zhou, Yongxin Ge
Knowl. Based Syst.3
2026 Efficient Pre-Trained Semantics Refinement for Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) is a fine-grained video understanding task that aims to ground the relevant video moments corresponding to given language queries. Most existing approaches utilize powerful Vision-Language Models (VLMs), augmented with additional network architectures or specialized modules to supplement temporal reasoning capabilities. Despite achieving impressive performance, these approaches tend to overlook a critical issue that pre-trained visual and textual representations are not specifically optimized for VTG. Particularly, such representations often suffer from inter-modal semantic misalignment referring to the inconsistency between different modalities, and intra-modal semantic confusion, caused by insufficient discriminability within visual modality. To address these limitations, we propose an efficient semantics refinement framework built upon pre-trained models, featuring two core components. First, the Modal Knowledge Bidirectional Propagation (MKBP) component promotes inter-modal semantic alignment via bidirectional enrichment of textual and visual semantics, exploiting their complementary strengths without introducing additional parameters. Second, the Content Context Contrast Learning (C3L) component alleviates intra-modal semantic confusion within visual modality by bringing query-specific visual features closer while separating irrelevant ones. Comprehensive experiments on six benchmark datasets demonstrate the superior performance of our proposed methods.
Yiqing Zhu, Yongxin Ge
IEEE Trans. Circuits Syst. Video Technol.4
2025 Image progressive steganography based on multi-frequency fusion deep network with dynamic sensing
Meng Li 0081, Jinhua Zhan, Yongxin Ge
Expert Syst. Appl.3
2025 Domain adaptation and knowledge distillation for lightweight pavement crack detection
Tianhao Xiao, Rong Pang, Chunhua Yang 0003, Zhimin Ruan, Yongxin Ge
Expert Syst. Appl.9
2025 Bridging asymmetry between image and video: Cross-modality knowledge transfer based on learning from video
Bingxin Zhou, Jianghao Zhou, Zhongming Chen, Long Deng, Yongxin Ge
Expert Syst. Appl.6
2025 Preserving knowledge from the source domain for cross-domain person re-identification
Yifeng Gou, Junyin Zhang, Yongxin Ge
Inf. Sci.4
2025 Self-supervised temporal adaptive learning for weakly-supervised temporal action localization
Jinrong Sheng, Jiaruo Yu, Yongxin Ge
Inf. Sci.5
2025 Instance-level feature bias calibration learning for text-to-image person re-identification
Yifeng Gou, Junyin Zhang, Yunnan Wang, Yongxin Ge
Knowl. Based Syst.5
2025 Summarized knowledge guidance for single-frame temporal action localization
Jinrong Sheng, Yongxin Ge
Pattern Recognit. Lett.3
2025 Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation
abstract
Semi-supervised learning based on consistency learning offers significant promise for enhancing medical image segmentation. Current approaches use copy-paste as an effective data perturbation technique to facilitate weak-to-strong consistency learning. However, these techniques often lead to a decrease in the accuracy of synthetic labels corresponding to the synthetic data and introduce excessive perturbations to the distribution of the training data. Such over-perturbation causes the data distribution to stray from its true distribution, thereby impairing the model's generalization capabilities as it learns the decision boundaries. We propose a weak-to-strong consistency learning framework that integrally addresses these issues with two primary designs: 1) it emphasizes the use of highly reliable data to enhance the quality of labels in synthetic datasets through cross-copy-pasting between labeled and unlabeled datasets; 2) it employs uncertainty estimation and foreground region constraints to meticulously filter the regions for copy-pasting, thus the copy-paste technique implemented introduces a beneficial perturbation to the training data distribution. Our framework expands the copy-paste method by addressing its inherent limitations, and amplifying the potential of data perturbations for consistency learning. We extensively validated our model using six publicly available medical image segmentation datasets across different diagnostic tasks, including the segmentation of cardiac structures, prostate structures, brain structures, skin lesions, and gastrointestinal polyps. The results demonstrate that our method significantly outperforms state-of-the-art models. For instance, on the PROMISE12 dataset for the prostate structure segmentation task, using only 10% labeled data, our method achieves a 15.31% higher Dice score compared to the baseline models. Our experimental code will be made publicly available at https://github.com/slhuang24/RCP4CL.
Senlong Huang, Yongxin Ge, Dongfang Liu, Mingjian Hong, Junhan Zhao, Alexander C. Loui
IEEE Trans. Image Process.2
2025 Counterfactual Bidirectional Co-Attention Transformer for Integrative Histology-Genomic Cancer Risk Stratification
abstract
Applying deep learning to predict patient prognostic survival outcomes using histological whole-slide images (WSIs) and genomic data is challenging due to the morphological and transcriptomic heterogeneity present in the tumor microenvironment. Existing deep learning-enabled methods often exhibit learning biases, primarily because the genomic knowledge used to guide directional feature extraction from WSIs may be irrelevant or incomplete. This results in a suboptimal and sometimes myopic understanding of the overall pathological landscape, potentially overlooking crucial histological insights. To tackle these challenges, we propose the CounterFactual Bidirectional Co-Attention Transformer framework. By integrating a bidirectional co-attention layer, our framework fosters effective feature interactions between the genomic and histology modalities and ensures consistent identification of prognostic features from WSIs. Using counterfactual reasoning, our model utilizes causality to model unimodal and multimodal knowledge for cancer risk stratification. This approach directly addresses and reduces bias, enables the exploration of 'what-if' scenarios, and offers a deeper understanding of how different features influence survival outcomes. Our framework, validated across eight diverse cancer benchmark datasets from The Cancer Genome Atlas (TCGA), represents a major improvement over current histology-genomic model learning methods. It shows an average 2.5% improvement in c-index performance over 18 state-of-the-art models in predicting patient prognoses across eight cancer types.
Zheyi Ji, Yongxin Ge, Chijioke Chukwudi, Kaicheng U, Sophia Meixuan Zhang, Yulong Peng, Junyou Zhu, Hossam Zaki, Xueling Zhang, Sen Yang 0006, Junhan Zhao
IEEE J. Biomed. Health Informatics2
2024 Prioritizing Test Cases Through Dual-Uncertainty Evaluating for Road Disease Detection System
abstract
Road damage detection system based on deep learning models has been extensively researched, capable of intelligently detecting damages in roads. However, to ensure reliability, a large number of images usually need to be collected and annotated to test the model. Collecting images is often a straightforward work, but annotating them is both time-consuming and costly. To minimize annotation costs during the testing phase, we can evaluate images using Test Prioritization method to select more meaningful samples for system testing. Yet, most mainstream work currently only evaluates cases based on classification result, neglecting the impact of localization. At the same time, these methods do not provide higher evaluation results to complex images that may detect errors or images where targets cannot be detected at all. To address these issues, we propose a new Test Prioritization tool, Dual-u, to evaluate road damage cases. It consists of two components: (1) Localization Uncertainty Evaluation, which evaluates the localization of damages while considering the value of the image, ensuring that complex images and images where targets cannot be detected also have higher priority; (2) Classification Uncertainty Evaluation, which evaluates the classification of damages. To evaluate Dual-u, we conduct extensive experiments using two object detection model architectures on multiple road disease datasets, and the results demonstrate that Dual-u outperforms existing methods in evaluating cases. Moreover, Dual-u can also be used to improve the accuracy and robustness of road disease detection models.
Tianhao Xiao, Zhimin Ruan, Yongxin Ge
APSEC7
2024 Efficient Dual-Confounding Eliminating for Weakly-supervised Temporal Action Localization
abstract
Weakly-supervised Temporal Action Localization (WTAL) following a localization-by-classification paradigm has achieved significant results, yet still grapples with confounding arising from ambiguous snippets. Previous works have attempted to distinguish these ambiguous snippets from action snippets without investigating the underlying causes of their formation, thus failing to effectively eliminate the bias on both action-context and action-content. In this paper, we revisit WTAL from the perspective of structural causal model to identify the true origins of confounding, and propose an efficient dual-confounding eliminating framework to alleviate these biases. Specifically, we construct a Substituted Confounder Set (SCS) to eliminate the confounding bias on action-context by leveraging the modal disparity between RGB and FLOW. Then, a Multi-level Consistency Mining (MCM) method is designed to mitigate the confounding bias on action-content by utilizing the consistency between discriminative snippets and corresponding proposals at both the feature and label levels. Notably, SCS and MCM could be seamlessly integrated into any two-stream models without additional parameters by Expectation-Maximization (EM) algorithm. Extensive experiments on two challenging benchmarks including THUMOS14 and ActivityNet-1.2 demonstrate the superior performance of our method.
Jinrong Sheng, Zhongming Chen, Yongxin Ge
ACM Multimedia5
2024 Temporal Contrastive and Spatial Enhancement Coarse Grained Network for Weakly Supervised Group Activity Recognition
Yongxin Ge
Eng. Appl. Artif. Intell.2
2024 Two-Stream Temporal Feature Aggregation Based on Clustering for Few-Shot Action Recognition
abstract
The metric learning paradigm has achieved notable success in few-shot action recognition; however, it faces unaddressed challenges. Specifically,(1)limited training data could impede the exploration of temporal action relations, and(2)precision would decline from the presence of outliers during the frame-level feature alignment. To address the challenges, we propose a two-stream temporal feature aggregation method based on clustering, incorporating a temporal augmentation module (TAM) and a feature aggregation module (FAM). The TAM adeptly integrates three consecutive grayscale frames into the original RGB frame through weighted summation, thereby addressing the color-related misguidance and enhancing the temporal information extraction. Meanwhile, the FAM employs clustering to aggregate the frame-level features into high semantic sub-actions and replaces the original features with cluster centers to mitigate the adverse impact of outliers on the model performance. Experimental results on benchmark datasets demonstrate the effectiveness of our method in few-shot action recognition. We validate our proposed approach by conducting comprehensive ablation experiments.
Long Deng, Bingxin Zhou, Yongxin Ge
IEEE Signal Process. Lett.4
2024 Error-Robust and Label-Efficient Deep Learning for Understanding Tumor Microenvironment From Spatial Transcriptomics
abstract
Spatial transcriptomics (ST) has become an important methodology in the analysis of the tumor microenvironment (TME) due to its ability to provide gene expression information with spatial resolution, enabling the identification and characterization of TME gene markers. Deep learning methods are proposed for analyzing spatial transcriptomic data for clustering the spatial regions of the TME based on gene expression. However, deep learning methods are often imposed by errors, which can impact the accuracy of gene expression quantification and TME gene identification. To address this issue, we propose a label-efficient method that utilizes curriculum learning and confidence learning to identify errors in graph deep learning when analyzing ST data. Our method explicitly incorporates the effect of noise in the learning process and employs probabilistic models or uncertainty estimates to represent the uncertainty in the data. Validated on human breast cancer ST data, we studied spatial gene expression in HER2-positive breast tumors using our method. The evaluation results suggest that the error quantification helps identify the noisy samples and subset the samples that results in more accurate gene expression quantification and TME gene identification. Additionally, there are biological insights obtained from the new subset formed by error samples. This error-robust deep learning method offers promising avenues for the analysis of spatial transcriptomic data, enabling accurate and label-efficient quantification of gene expression and identification of TME gene markers.
Jiake Leng, Yiming Cui 0002, Junhan Zhao, Yongxin Ge
IEEE Trans. Circuits Syst. Video Technol.6
2024 Self-Supervised Adversarial Learning for Domain Adaptation of Pavement Distress Classification
abstract
Pavement distress classification is crucial for the maintenance of highways. Although many methods for classifying pavement distress are available, they all assume that training and testing datasets are drawn from the same distribution. When we introduce a new unlabeled dataset with a different distribution, the performance of existing methods decreases considerably due to domain shift, motivating us to look beyond the supervised setting to utilize unlabeled datasets directly in training a model. Therefore, we develop a novel unsupervised domain adaptation (UDA) framework, namely, the Self-supervised Adversarial Network (SSAN) for the first time in this study to conduct multi-category pavement distress classification on an unlabeled target domain. In particular, SSAN leverages adversarial domain adaptation (ADA) thoughts to align the features of different domains. However, distress typically occupies a small Section of high-resolution pavement images. Consequently, aligning features directly is unreasonable because the aligning procedure is still dominated by background features instead of foreground features, which are the most useful information for classification. Therefore, we design a pretext module, called Self-supervised Learning for the Target domain (SLT), to mine foreground information. To validate our method, we use two challenging pavement crack datasets, namely, the Chonqing University Bituminous Pavement Disease Detection (CQU-BPDD) and the Chongqing University Bituminous Pavement Multi-label Disease Detection (CQU-BPMDD) datasets. Moreover, extensive experiments demonstrate that SSAN outperforms state-of-the-art UDA methods.
Yanwen Wu, Mingjian Hong, Sheng Huang 0001, Yongxin Ge
IEEE Trans. Intell. Transp. Syst.6
2023 Deep Domain Adaptation for Pavement Crack Detection
abstract
Deep learning-based pavement cracks detection methods often require large-scale labels with detailed crack location information to learn accurate predictions. In practice, however, crack locations are very difficult to be manually annotated due to various visual patterns of pavement crack. In this paper, we propose a Deep Domain Adaptation-based Crack Detection Network (DDACDN), which learns domain invariant features by taking advantage of the source domain knowledge to predict the multi-category crack location information in the target domain, where only image-level labels are available. Specifically, DDACDN first extracts crack features from both the source and target domain by a two-branch weights-shared backbone network. And in an effort to achieve the cross-domain adaptation, an intermediate domain is constructed by aggregating the three-scale features from the feature space of each domain to adapt the crack features from the source domain to the target domain. Finally, the network involves the knowledge of both domains and is trained to recognize and localize pavement cracks. To facilitate accurate training and validation for domain adaptation, we use two challenging pavement crack datasets CQU-BPDD and RDD2020. Furthermore, we construct a new large-scale Bituminous Pavement Multi-label Disease Dataset named CQU-BPMDD, which contains 38994 high-resolution pavement disease images to further evaluate the robustness of our model. Extensive experiments demonstrate that DDACDN outperforms state-of-the-art pavement crack detection methods in predicting the crack location on the target domain.
Chunhua Yang 0003, Sheng Huang 0001, Zhimin Ruan, Yongxin Ge
IEEE Trans. Intell. Transp. Syst.7
2022 Forcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action Localization
abstract
With video-level labels, weakly supervised temporal action localization (WTAL) applies a localization-by-classification paradigm to detect and classify the action in untrimmed videos. Due to the characteristic of classification, class-specific background snippets are inevitably mis-activated to improve the discriminability of the classifier in WTAL. To alleviate the disturbance of background, existing methods try to enlarge the discrepancy between action and background through modeling background snippets with pseudo-snippet-level annotations, which largely rely on artificial hypotheticals. Distinct from the previous works, we present an adversarial learning strategy to break the limitation of mining pseudo background snippets. Concretely, the background classification loss forces the whole video to be regarded as the background by a background gradient reinforcement strategy, confusing the recognition model. Reversely, the foreground(action) loss guides the model to focus on action snippets under such conditions. As a result, competition between the two classification losses drives the model to boost its ability for action modeling. Simultaneously, a novel temporal enhancement network is designed to facilitate the model to construct temporal relation of affinity snippets based on the proposed strategy, for further improving the performance of action localization. Finally, extensive experiments conducted on THUMOS14 and ActivityNet1.2 demonstrate the effectiveness of the proposed method.
Yongxin Ge, Jiaruo Yu, Zhongming Chen
ACM Multimedia2
2022 A channel-spatial-temporal attention-based network for vibration-based damage detection
Shiyun Liao, Jianxi Yang, Yongxin Ge
Inf. Sci.4
2022 Low-resolution assisted three-stream network for person re-identification
Jiahong Xie, Yongxin Ge, Junyin Zhang, Sheng Huang 0001, Feiyu Chen 0002, Hongxing Wang 0001
Vis. Comput.2
2021 Context Driven Network with Bayes for Weakly Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization (WTAL) aims to detect action instances from untrimmed videos. It may cause two problems, namely action incompleteness and background disturbance, due to only video-level class labels given. In this paper, we propose a context driven network with Bayes to alleviate the two problems, in which an attention module is used to predict coarse probability for each snippet, and then a Bayesian refinement module is designed to refine the coarse results by capturing the relationship between context snippets. Finally, the coarse and refined probabilities are combined as the inputs of the classifier for training. Quantitative and qualitative studies on two benchmark datasets, i.e., THUMOS’14 and ActivityNet 1.2, demonstrate that the proposed approach exceeds state-of-the-art methods.
Jiaruo Yu, Yongxin Ge, Zhongming Chen, Xiaolei Qin
ICME2
2021 Self-Supervised Pre-training on the Target Domain for Cross-Domain Person Re-identification
abstract
Most existing cluster-based cross-domain person re-identification (re-id) methods only pre-train the re-id model on the source domain. Unfortunately, the pre-trained model may not perform well on the target domain due to the large domain gap between source and target domains, which is harmful to the following optimization. In this paper, we propose a novel Self-supervised Pre-training method on the Target Domain (SPTD), which pre-trains the model on both the source and target domains in a self-supervised manner. Specifically, SPTD uses different kinds of data augmentation manners to simulate different intra-class changes and constraints the consistency between the augmented data distribution and the original data distribution. As a result, the pre-trained model involves some specific discriminative knowledge on the target domain and is beneficial to the following optimization. It is easy to combine the proposed SPTD with other cluster-based cross-domain re-id methods just by replacing the original pre-trained model with our pre-trained model. Comprehensive experiments on three widely used datasets, i.e. Market1501, DukeMTMC-ReID and MSMT17, demonstrate the effectiveness of SPTD. Especially, the final results surpass previous state-of-the-art methods by a large margin.
Junyin Zhang, Yongxin Ge, Xinqian Gu, Boyu Hua, Tao Xiang 0001
ACM Multimedia2
2021 Deep feature enhancing and selecting network for weakly supervised temporal action localization
Jiaruo Yu, Yongxin Ge, Xiaolei Qin, Sheng Huang 0001, Feiyu Chen 0002
J. Vis. Commun. Image Represent.2
2021 Deep snippet selective network for weakly supervised temporal action localization
Yongxin Ge, Xiaolei Qin, Dan Yang 0001, Martin Jägersand
Pattern Recognit.1
2021 Discriminative deep semi-nonnegative matrix factorization network with similarity maximization for unsupervised feature learning
Feiyu Chen 0002, Yongxin Ge, Sheng Huang 0001, Xiaohong Zhang 0002, Dan Yang 0001
Pattern Recognit. Lett.3
2021 Deep Variation Transformation Network for Foreground Detection
abstract
In existing literature, the distribution of pixel observations is analyzed with models designed for the video foreground detection task. However, it is possible that the background and foreground share similar observations, causing false detections. We propose a novel foreground detection method called Deep Variation Transformation Network (DVTN), focusing on analyzing the pixel variations instead of distributions. In particular, pixel variations are represented by a sequence of pixel observations, and DVTN is trained to transform the pixel variations into a new space, where the observations can be classified easily. Following this, the output of DVTN is utilized by a linear classifier to label pixels as foreground or background. As a result of the global analysis and the strong learning ability of DVTN, the proposed approach adaptively learns a good transformation from pixel variations to probabilities of labels to improve performance. Comprehensive experiments on several benchmark datasets demonstrate the superiority of our DVTN approach compared to both state-of-the-art deep learning and traditional methods, especially in scenes lacking texture and color information. Code is available at https://github.com/Zhangjunyin/DVTN.
Yongxin Ge, Junyin Zhang, Xinyu Ren, Chenqiu Zhao, Anup Basu
IEEE Trans. Circuits Syst. Video Technol.1
2020 Deep Selective Feature Learning for Action Recognition
abstract
Soft-attention mechanism has attracted a lot of attention in recent years due to its ability to capture the most discriminative image features for understanding actions. However, soft-attention tends to focus on fine-grained parts on images and ignores global information, which can lead to totally wrong classification results. To address this issue, we propose a novel deep selective feature learning network (DSFNet), which can automatically learn the feature maps with both fine-grained and global information. Specially, DSFNet is designed to have the ability to learn to adjust the actions for feature map selection by maximizing the cumulative discounted rewards. Moreover, the DSFNet is an easy-to-use extension of state-of-the-art base architectures of multiple tasks. Extensive experiments show that the proposed method has achieved superior performance on two standard action recognition benchmarks across still images (PPMI) and videos (HMDB51).
Yongxin Ge, Jinyuan Feng, Xiaolei Qin, Jiaruo Yu, Hui Yu 0001
ICME2
2020 Robust Bidirectional Generative Network For Generalized Zero-Shot Learning
abstract
In this work, we propose a novel generative approach named Robust Bidirectional Generative Network (RBGN) based on Conditional Generative Adversarial Network (CGAN) for Generalized Zero-shot Learning (GZSL). RBGN employs the adversarial attack to train a more rigorous discriminator, thus enhancing the generalizability and robustness of the feature generator under minimax strategy. Moreover, RBGN decodes the generated visual features back to their semantic representations to further improve the representational ability of generated visual features and alleviate the hubness problem. The experimental results of GZSL on four datasets, i.e. CUB, SUN, AWA1, AWA2, demonstrate that our model achieves competitive performance compared to state-of-the-art approaches and owns better generalizability to the unseen classes over conventional generative GZSL models. Further robustness analysis also validates the strong robustness of our model to the different types of semantic disturbance.
Sheng Huang 0001, Luwen Huangfu, Feiyu Chen 0002, Yongxin Ge
ICME5
2020 DTMMN: Deep transfer multi-metric network for RGB-D action recognition
Xiaolei Qin, Yongxin Ge, Jinyuan Feng, Dan Yang 0001, Feiyu Chen 0002, Sheng Huang 0001
Neurocomputing2
2020 Deep shape constrained network for robust face alignment
Yongxin Ge, Junyin Zhang, Min Chen 0016, Jiahong Xie, Dan Yang 0001
Pattern Recognit. Lett.1
2020 Spatial Enhancement and Temporal Constraint for Weakly Supervised Action Localization
abstract
Weakly supervised temporal action localization (WSTAL) is a practical but challenging issue in video understanding. However, most existing methods have to activate background snippets or deactivate action snippets in cases of no boundary annotations, which inevitably affects the localization of action instances. In this letter, we propose a spatial enhancement and temporal constraint (SETC) model to address this problem from three aspects. Specifically, we first propose a spatial enhancement module to enhance the discrimination of the extracted features. Then we leverage the instance sparse constraint to restrain the drastic fluctuation class activation sequence (CAS). Finally, we use the confidence connectivity enhancement to connect the snippets that are broken up by mistake. Experiments on THUMOS'14 and ActivityNet datasets validate the efficacy of SETC against existing state-of-the-art WSTAL algorithms.
Xiaolei Qin, Yongxin Ge, Hui Yu 0001, Feiyu Chen 0002, Dan Yang 0001
IEEE Signal Process. Lett.2
2019 Background Subtraction Based on Integration of Alternative Cues in Freely Moving Camera
abstract
Previous approaches to background subtraction in freely moving camera typically focus on improving the accuracy of motion estimation. In this paper, we propose that the accurate background subtraction is possible with the integration of alternative cues about foreground and background. We also put forward a novel background subtraction framework called the integration of foreground and background cues. Here, the foreground cues are extracted by the Gaussian mixture model compensated with image alignment, while the background cues are obtained from the spatiotemporal features filtered by the homography transformation. Subsequently, the integration is devised as a hierarchical competition procedure based on super-pixels under multiple levels with the underlying motivation to utilize the exclusiveness between these cues for the compensation of their corresponding defects. The result of competition between foreground and background cues in a particular super-pixel is used as the proximity, and the foreground is segmented by combining super-pixels with proximity under multiple levels. Comprehensive evaluations using standard benchmarks demonstrate the superiority of our work compared with the state-of-the-art.
Chenqiu Zhao, Aneeshan Sain, Ying Qu 0007, Yongxin Ge, Haibo Hu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2018 Deep Multi-Metric Learning for Person Re-Identification
abstract
In this paper, to exploit more discriminative information of the global-body and body-parts features, we present a novel deep multi-metric learning (DMML) network for person re-identification under the triplet framework. The main novelty of our learning framework lies in two aspects: 1) Unlike most existing metric learning-based approaches, which learn only one distance metric for comparison, our DMM-L method aims to learn different metrics for the global-body and body-parts features respectively by using convolutional neural network (CNN); 2) A new multi-metric loss function is proposed to train the DMML network, under which the distance of each negative pair is greater than that of each positive pair by a predefined margin, and the correlations of different metrics are maximized. Compared with the previous person re-identification methods that have shown state-of-the-art performances, our DMML approach can achieve competitive results on the challenging CUHK03, CUHKOl, VIPeR and iLIDS datasets.
Yongxin Ge, Xinqian Gu, Min Chen 0016, Hongxing Wang 0001, Dan Yang 0001
ICME1
2018 Joint Deep Learning for RGB-D Action Recognition
abstract
Recent approaches in RGB-based and depth-based human action recognition achieved outstanding performance respectively, which demonstrate the effectiveness of RGB and depth modalities for action classification, however it is infrequent to consider them both. Currently, available multimodal-based methods of action recognition suffer from some limitations, including non-end-to-end training, violent fusion and inefficiency. In this paper, we propose a novel joint deep learning (JDL) model which is capable of: 1) jointly optimizing the object of classification and feature extraction through a novel end-to-end two-stream deep learning model, 2) refining common-specific features via introducing the constraint of similarity loss in high-level, and 3) using 2D convolution kernel instead of 3D convolution kernel during feature extraction for gaining the efficiency. The experiments on two challenging datasets show the promising performance of our architecture.
Xiaolei Qin, Yongxin Ge, Liuwei Zhan, Guangrui Li 0003, Sheng Huang 0001, Hongxing Wang 0001, Feiyu Chen 0002
VCIP2
2018 Improved hypergraph regularized Nonnegative Matrix Factorization with sparse representation
abstract
As a commonly used data representation technique, Nonnegative Matrix Factorization (NMF) has received extensive attentions in the pattern recognition and machine learning communities over decades, since its working mechanism is in accordance with the way how the human brain recognizes objects. Inspired by the remarkable successes of manifold learning, more and more researchers attempt to incorporate the manifold learning into NMF for finding a compact representation ,which uncovers the hidden semantics and respects the intrinsic geometric structure simultaneously. Graph regularized Nonnegative Matrix Factorization (GNMF) is one of the representative approaches in this category. The core of such approach is the graph, since a good graph can accurately reveal the relations of samples which benefits the data geometric structure depiction. In this paper, we leverage the sparse representation to construct a sparse hypergraph for better capturing the manifold structure of data, and then impose the sparse hypergraph as a regularization to the NMF framework to present a novel GNMF algorithm called Sparse Hypergraph regularized Nonnegative Matrix Factorization (SHNMF). Since the sparse hypergraph inherits the merits of both the sparse representation and the hypergraph model, SHNMF enjoys more robustness and can better exploit the high-order discriminant manifold information for data representation . We apply our work to address the image clustering issue for evaluation. The experimental results on five popular image databases show the promising performances of the proposed approach in comparison with the state-of-the-art NMF algorithms.
Sheng Huang 0001, Hongxing Wang 0001, Yongxin Ge, Luwen Huangfu, Xiaohong Zhang 0002, Dan Yang 0001
Pattern Recognit. Lett.3
2017 Robust face alignment with cascaded coarse-to-fine auto-encoder network
abstract
In this paper, we present a novel face alignment method using a two-level cascaded auto-encoder networks (2-LCAN). In our framework, the first level auto-encoder networks generate rough facial landmarks locations by taking detected face images with low-resolution as inputs. The second level autoencoder networks are constructed by cascading several sub stacked auto-encoder networks (SSAN) in a coarse-to-fine manner. Each SSAN extracts SIFT features and local pixels features around current landmark positions, then fuses them together to further refine landmarks of different facial components with higher image resolutions. Finally, experimental results on LFPW and HELEN datasets demonstrate that our proposed method is significantly superior to the compared approaches both in accuracy and robustness.
Yongxin Ge, Mingjian Hong, Sheng Huang 0001, Dan Yang 0001
ICIP2
2016 Joint Local Regressors Learning for Face Alignment
Yongxin Ge, Mingjian Hong, Sheng Huang 0001, Dan Yang 0001
Neurocomputing1
2016 Discriminant Hyper-Laplacian Projections and its scalable extension for dimensionality reduction
Sheng Huang 0001, Dan Yang 0001, Yongxin Ge, Xiaohong Zhang 0002
Neurocomputing3
2015 Active appearance model search using partial least squares regression
abstract
A novel active appearance model (AAM) search algorithm based on partial least squares (PLS) regression is proposed. PLS models the relationship between independent (texture residuals) and dependent (error in the model parameters) variables in the training phase by extracting from independent and dependent variables a set of orthogonal factors called latent variables respectively which have the maximum covariance. During search, the parameter updates with the best predictive power are extracted from the texture residuals. On the other hand, PLS is well suited for the low observation-to-variable ratio context, where the sample covariance matrix is likely to be singular, which is very common in AAM. Experiments show that the proposed method has better performance than the original AAM and comparable performance to AAM search based on Canonical correlation analysis (CCA-AAM) in terms of convergence speed, whilst affording superior computational efficiency.
Yongxin Ge, Martin Jägersand, Dan Yang 0001
VCIP1
2015 Neighborhood repulsed correlation metric learning for kinship verification
abstract
In this paper, we propose a new neighborhood repulsed correlation metric learning (NRCML) method for kinship verification. While several metric learning algorithms have been proposed in recent years and some of them have successfully applied to kinship verification, most existing metric learning methods are developed based on the Euclidian similarity metric, which is not powerful enough to measure the similarity of face samples. To address this, we propose a NRCML method by using the correlation similarity measure to learn a discriminative distance metric, under which positive pairs are pulled as close as possible and negative pairs lying in a neighborhood are repulsed as far as possible, simultaneously. Experimental results are presented to show the effectiveness of the proposed method.
Haibin Yan, Xiuzhuang Zhou, Yongxin Ge
VCIP3
2015 Combined supervised information with PCA via discriminative component selection
Sheng Huang 0001, Dan Yang 0001, Yongxin Ge, Xiaohong Zhang 0002
Inf. Process. Lett.3
2015 Regularized local metric learning for person re-identification
Venice Erin Liong, Jiwen Lu, Yongxin Ge
Pattern Recognit. Lett.3
2013 Age estimation from human body images
abstract
In this paper, we investigate the problem of estimating human ages from full body images. To our best knowledge, this problem has not been formally addressed before possibly due to the great challenges and lacking of such publicly available datasets. However, estimating human ages at a distance has a number of potential applications, especially for visual surveillance in such places as supermarkets, airports, building entrances, and shopping malls. In this paper, we propose a new human age estimation approach from full body images with frontal or back views. Our contributions are three-fold. First, we collect a human body image dataset containing 1500 public figures or celebrities searched from the internet, as well as the age label information of each image. Second, we explore several widely used human local appearance feature descriptors with a regression model to estimate human ages from these body images. Lastly, we apply a multiview canonical correlation analysis (MCCA) method by making use of multiple feature descriptors to exploit complementary information to further improve the age estimation performance. Experimental results have clearly demonstrated the feasibility of using fully body images to estimate human age and the efficacy of our proposed approach.
Yongxin Ge, Jiwen Lu, Dan Yang 0001
ICASSP1
2013 Makeup-robust face verification
abstract
We investigate in this paper the problem of face verification in the presence of face makeups. To our knowledge, this problem has less formally addressed in the literature. A key challenge is how to increase the measured similarity between face images of the same person without and with makeups. In this paper, we propose a novel approach for makeup-robust face verification, by measuring correlations between face images in a meta subspace. The meta subspace is learned using canonical correlation analysis (CCA), with the objective that intra-personal sample correlations are maximized. Subsequently, discriminative learning with the support vector machine (SVM) classifier is applied to verify faces based on the low-dimensional features in the learned meta subspace. Experimental results on our dataset are presented to demonstrate the efficacy of our approach.
Junlin Hu 0001, Yongxin Ge, Jiwen Lu
ICASSP2
2013 Activity-based human identification
abstract
We investigate in this paper the problem of activity-based human identification. Different from most existing gait recognition methods where only human walking activity is considered and utilized for person identification, we aim to identify people from various activities such as eating, jumping, and weaving. For each video clip, we first extract binary human body masks by using background substraction, followed by computing the average energy image (AEI) features to represent each video clip. Then, a mapping is learned by applying an adaptive discriminant analysis (ADA) method to project AEI features into a low-dimensional subspace, such that the intra-class (activities performed by the same person) variations are minimized and the interclass (activities performed by different persons) are maximized, simultaneously. Moreover, interclass samples with large similarity difference are deemphasized and those with small difference are emphasized, such that more discriminative information can be used for recognition. Experimental results on three publicly available databases show the efficacy of our proposed approach.
Tzu-Yi Hung, Jiwen Lu, Junlin Hu 0001, Yap-Peng Tan, Yongxin Ge
ICASSP5
2013 Active appearance models using statistical characteristics of Gabor based texture representation
Yongxin Ge, Dan Yang 0001, Jiwen Lu, Xiaohong Zhang 0002
J. Vis. Commun. Image Represent.1