EDBT 2026 Demo / reviewers in the wild / expert
Xinge You
dblp:16/1184
· DBLP profile ↗
174ranked-venue papers
14as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 6 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 7 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Security and privacy · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mutually Causal Semantic Distillation Network for Zero-Shot Learning
Shiming Chen 0002, Shuhuang Chen, Guosen Xie, Xinge You |
Int. J. Comput. Vis. | 4 |
| 2026 | DH-MSVM: A hybrid algorithm for seeking quality support vectors in distributed learning
Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You |
Neural Networks | 5 |
| 2026 | Dual Adversarial Perturbations for Zero-Shot LearningabstractIn Zero-Shot Learning (ZSL), embedding-based methods learn a visual–semantic mapping that leverages the attribute knowledge of seen classes to predict the attributes of unseen classes, enabling knowledge transfer from seen to unseen classes. However, distributional discrepancies between seen and unseen classes introduce an inherent domain shift, and inter-class variations cause the same attribute to be expressed differently across categories. As a result, models trained on seen classes often struggle to accurately recognize attributes in unseen classes, limiting their generalization ability. To address these challenges, we propose DAPZSL, a dual adversarial perturbation framework that enhances the robustness of visual–semantic mappings through Feature-Level Adversarial Perturbation (FAP) and improves the model’s generalization ability via Weight-Level Adversarial Perturbation (WAP). Specifically, FAP generates semantically perturbed adversarial samples at the feature level, and incorporating these samples during training encourages the model to learn more robust visual–semantic mappings that are resilient to semantic variations, which improves attribute recognition on unseen classes. Meanwhile, WAP introduces adversarial perturbations into the model’s weight space, promoting a flatter loss landscape that alleviates overfitting to seen classes and enhances generalization. Extensive experiments on multiple benchmark datasets—including AWA2, SUN, and CUB—demonstrate that DAPZSL significantly outperforms existing ZSL models. Shiming Chen 0002, Guosen Xie, Chaojian Yu, Xinhua You, Qinmu Peng, Xinge You |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Few-Shot Object Detection via Spatial-Channel State Space ModelabstractDue to the limited training samples in few-shot object detection (FSOD), we observe that current methods may struggle to accurately extract effective features from each channel. Specifically, this issue manifests in two aspects: i) channels with high weights may not necessarily be effective, and ii) channels with low weights may still hold significant value. To handle this problem, we consider utilizing inter-channel correlation to ensure that the novel model can effectively highlight relevant channels and rectify incorrect ones, thereby strengthening channel quality. Since the channel sequence is also 1-dimensional, its similarity with the temporal sequence inspires us to take Mamba for modeling the correlation in the channel sequence Based on this concept, we propose the Spatial-Channel State Space Modeling (SCSM) module for spatial-channel-sequence modeling to accurately extract effective features from each channel. In SCSM, we design the Spatial Feature Modeling (SFM) module to ensure the quality of spatial feature representations. We then introduce the Channel State Modeling (CSM) module, which treats channels as a 1-dimensional sequence and take mamba to capture the correlation between channels. Extensive experiments on the VOC and COCO datasets show that the SCSM module enables the novel detector to improve the quality of channel feature representations and achieve state-of-the-art performance. Code is released at https://github.com/zhimengXin/SCSM. Zhimeng Xin, Tianxu Wu, Yixiong Zou, Shiming Chen 0002, Dingjie Fu, Xinge You |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | DHS-AE: A Distributed Support Vector Machine With Adaptive Regularization Parameters for Different Data DistributionsabstractIn distributed machine learning scenarios, the difference in data distribution among different nodes is a key issue that cannot be ignored. However, existing methods make it difficult to autonomously adjust model parameters for dynamically changing data distributions, leading to inflexible global decision boundaries with insufficient local adaptation. To address this problem, we propose a distributed hybrid support vector machine (SVM) based on the adaptive ensemble selection of regularization parameters, DHS-AE. The model utilizes the data structure information to cut the data space and thus identify data distribution characteristics. The SVM, integrated with regularization parameters that are adaptively determined within specific ranges, is utilized in the local subspace to enable real-time adjustment of decision boundaries in response to distribution changes, thereby further reducing the computational overhead. The generalization bound of DHS-AE is theoretically established using covering numbers, and the fast convergence speed and consistency are derived. In practical applications, we verify the excellent performance of the DHS-AE using a large number of real datasets. Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You |
IEEE Trans. Cybern. | 5 |
| 2026 | Domain-Adaptive Fuzzy Graph Diffusion Networks for Open-Set Cross-Domain Node ClassificationabstractFuzzy logic-based graph neural networks (FL-GNN) have recently garnered growing attention in node classification, which aims to enhance the ability of GNN in modeling uncertain relationships between nodes. However, existing FL-GNN typically assume that nodes in the source domain (training set) and target domain (test set) follow the identical data distribution and class sets. Real-world scenarios often exhibit significant distribution shifts and target domain even contains classes that were not present in the source domain, termed open-set cross-domain node classification (OSCD-NC), which seriously damages their superior performance. Thus, how to leverage the strong uncertain knowledge representation capacity of FL-GNN to learn a well-defined boundary between seen and unseen classes for improving OSCD-NC performance remains an open and underexplored research problem. In this paper, we propose an effective domain-adaptive fuzzy graph diffusion network (DFGDN) for OSCD-NC. Specifically, with the help of a fuzzy adjacency matrix, fuzzy graph diffusion networks are proposed to generate robust fuzzy node representations by adaptively enhancing feature collaboration between low-pass and high-pass graph filters. Then, a peer ($M$+1)-class classifier is introduced to learn a rough class boundary by measuring their class prediction probability difference for target domain. After that, the ($M$+1)-means clustering and decoder modules are simultaneously designed to discover more supervision guidance from target domain for learned class boundary optimization. Finally, we jointly optimize the above modules in an adversarial manner via classification loss, classifier discrepancy loss and mean squared error loss, which further improves the accuracy of the learned class boundary by pulling seen nodes from the source domain and target domain closer, and pushing unseen nodes away. Extensive experiments on three cross-domain data pairs and various openness rates demonstrate the effectiveness of the proposed DFGDN framework. Sichao Fu, Yanping Chen 0010, Songren Peng, Weihua Ou, Liangshuo Ning, Bin Zou 0002, Qinmu Peng, Xiaoyuan Jing, Xinge You |
IEEE Trans. Fuzzy Syst. | 9 |
| 2026 | DHL-FLD: A Distributed Hybrid Learning Based on Fisher Linear Discriminant for Data ClassificationabstractDistributed machine learning provides an efficient solution for large-scale data processing through parallel computing. However, current distributed learning relies on global or local paradigms and cannot adaptively adjust decision boundaries in complex data environments. To address this problem, we propose a Distributed Hybrid Learning algorithm based on Fisher Linear Discriminant (DHL-FLD). Specifically, DHL-FLD consists of a global pre-learning phase and a subspace local learning phase. On the one hand, the global pre-learning phase is designed to divide the data space, which can obtain the data structure information. On the other hand, the local learning phase dynamically adjusts and optimizes the decision boundaries, guided by the structural information and distributional properties of the data. Theoretically, we establish the generalization bound of DHL-FLD using the integral operator technique and verify the scalability and robustness of DHL-FLD. The effectiveness of DHL-FLD is demonstrated through extensive experiments on real datasets. Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsabstractIn web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG. Xingye Chen, Zhenbang Du, Yanyin Chen, Haohan Wang, Linkai Liu 0002, Jinyuan Zhao, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang |
WWW | 17 |
| 2025 | Unsupervised multiplex graph diffusion networks with multi-level canonical correlation analysis for multiplex graph representation learning
Sichao Fu, Qinmu Peng, Yange He, Baokun Du, Bin Zou 0002, Xiaoyuan Jing, Xinge You |
Sci. China Inf. Sci. | 7 |
| 2025 | Semantics-Conditioned Generative Zero-Shot Learning via Feature Refinement
Shiming Chen 0002, Ziming Hong, Xinge You, Ling Shao 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | GKA: Graph-guided knowledge association for fine-grained visual categorization
Yuetian Wang, Shuo Ye, Wenjin Hou, Duanquan Xu, Xinge You |
Neurocomputing | 5 |
| 2025 | GFIA: Generative Fault Image Analysis via vision-language model its application to train bogie transmission system
Xinge You |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Another Vertical View: A Hierarchical Network for Heterogeneous Trajectory Prediction via SpectrumsabstractWith the fast development of AI-related techniques, the applications of trajectory prediction are no longer limited to easier scenes and trajectories. More and more trajectories with different forms, such as coordinates, bounding boxes, and even high-dimensional human skeletons, need to be analyzed and forecasted. Among these heterogeneous trajectories, interactions between different elements within a frame of trajectory, which we call "Dimension-wise Interactions", would be more complex and challenging. However, most previous approaches focus mainly on a specific form of trajectories, and potential dimension-wise interactions are less concerned. In this work, we expand the trajectory prediction task by introducing the trajectory dimensionality $M$M, thus extending its application scenarios to heterogeneous trajectories. We first introduce the Haar transform as an alternative to the Fourier transform to better capture the time-frequency properties of each trajectory-dimension. Then, we adopt the bilinear structure to model and fuse two factors simultaneously, including the time-frequency response and the dimension-wise interaction, to forecast heterogeneous trajectories via trajectory spectrums hierarchically in a generic way. Experiments show that the proposed model outperforms most state-of-the-art methods on ETH-UCY, SDD, nuScenes, and Human3.6 M with heterogeneous trajectories, including 2D coordinates, 2D/3D bounding boxes, and 3D human skeletons. Beihao Xia, Conghao Wong, Duanquan Xu, Qinmu Peng, Xinge You |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Exploring sample relationship for few-shot classification
Xingye Chen, Wenxiao Wu, Li Ma 0005, Xinge You, Changxin Gao, Nong Sang, Yuanjie Shao |
Pattern Recognit. | 4 |
| 2025 | FN-NET: Adaptive data augmentation network for fine-grained visual categorization
Shuo Ye, Qinmu Peng, Yiu-Ming Cheung, Yu Wang 0106, Ziqian Zou, Xinge You |
Pattern Recognit. | 6 |
| 2025 | Adversarial Feature Training for Few-Shot Object DetectionabstractCurrently, most few-shot object detection (FSOD) methods apply the two-stage training strategy, which first requires training in abundant base classes and transfers the learned prior knowledge to the novel stage. However, due to the inherent imbalance between the base and novel classes, the trained model tends to have a bias toward recognizing novel classes as base ones when they are similar. To address this problem, we propose an adversarial feature training (AFT) strategy aimed at effectively calibrating the decision boundary between novel and base classes to alleviate classification confusion in FSOD. Specifically, we introduce the Classification Level Fast Gradient Sign Method (CL-FGSM), which leverages gradient information from the classifier module to generate adversarial samples with extra feature attention. By attacking the high-level features, we can create adversarial feature samples that are combined with clean high-level features in a suitable range of proportions. Such adversarial feature samples, generated by CL-FGSM, are then combined with clean high-level features in a suitable range of proportions to train the few-shot detector. By this, the novel model is forced to learn extra class-specific features that improve the robustness of the classifier to establish a correct decision boundary, which avoids confusion between base and novel classes in FSOD. Extensive experiments demonstrate that our proposed AFT strategy effectively calibrates the classification decision boundary to avoid classification confusion between base and novel classes and significantly improves the performance of FSOD. Our code is available athttps://github.com/wutianxu/AFT. Tianxu Wu, Zhimeng Xin, Shiming Chen 0002, Yixiong Zou, Xinge You |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | An Efficient Adversarial Attack on FCG-Based Android Malware Detection Systems
Heng Li 0008, Bang Wu 0002, Wei Yuan 0001, Cuiying Gao, Xinge You, Xiapu Luo |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Toward Disentangled and Controllable Deep Metric Learning With Human-Like Concept DecompositionabstractDeep metric learning (DML) has shown significant advancements in learning discriminative embeddings for images, playing a crucial role in various vision tasks. However, existing methods typically rely on deep neural networks to extract holistic embeddings, which are challenging to disentangle and interpret. To address this issue, we take inspiration from human cognition, where objects are decomposed into distinct concepts for better understanding. Specifically, we propose the concept metrics network (CMNs) to achieve disentangled and controllable DML. CMN begins by initializing learnable concept vectors to represent various visual concepts. These vectors are then associated with regional visual features via cross-attention mechanism, ensuring each vector corresponds to specific visual properties. Finally, the concept values, determined by their presence in the image, form the output embedding. Comprehensive experiments demonstrate that CMN effectively disentangles visual concepts, with each embedding dimension corresponding to a specific concept. Our method not only outperforms existing state-of-the-art methods in conventional DML application (i.e., image retrieval), but also enables more flexible and controllable application. The code is available at https://github.com/shchen0001/CMN. Shuhuang Chen, Shiming Chen 0002, Shuo Ye, Yuetian Wang, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Multilevel Contrastive Graph Masked Autoencoders for Unsupervised Graph-Structure LearningabstractUnsupervised graph-structure learning (GSL) which aims to learn an effective graph structure applied to arbitrary downstream tasks by data itself without any labels' guidance, has recently received increasing attention in various real applications. Although several existing unsupervised GSL has achieved superior performance in different graph analytical tasks, how to utilize the popular graph masked autoencoder to sufficiently acquire effective supervision information from the data itself for improving the effectiveness of learned graph structure has been not effectively explored so far. To tackle the above issue, we present a multilevel contrastive graph masked autoencoder (MCGMAE) for unsupervised GSL. Specifically, we first introduce a graph masked autoencoder with the dual feature masking strategy to reconstruct the same input graph-structured data under the original structure generated by the data itself and learned graph-structure scenarios, respectively. And then, the inter- and intra-class contrastive loss is introduced to maximize the mutual information in feature and graph-structure reconstruction levels simultaneously. More importantly, the above inter- and intra-class contrastive loss is also applied to the graph encoder module for further strengthening their agreement at the feature-encoder level. In comparison to the existing unsupervised GSL, our proposed MCGMAE can effectively improve the training robustness of the unsupervised GSL via different-level supervision information from the data itself. Extensive experiments on three graph analytical tasks and eight datasets validate the effectiveness of the proposed MCGMAE. Sichao Fu, Qinmu Peng, Bin Zou 0002, Duanquan Xu, Xiaoyuan Jing, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Filter Pruning Based on Information Capacity and IndependenceabstractFilter pruning has gained widespread adoption for the purpose of compressing and speeding up convolutional neural networks (CNNs). However, the existing approaches are still far from practical applications due to biased filter selection and heavy computation cost. This article introduces a new filter pruning method that selects filters in an interpretable, multiperspective, and lightweight manner. Specifically, we evaluate the contributions of filters from both individual and overall perspectives. For the amount of information contained in each filter, a new metric called information capacity is proposed. Inspired by the information theory, we utilize the interpretable entropy to measure the information capacity and develop a feature-guided approximation process. For correlations among filters, another metric called information independence is designed. Since the aforementioned metrics are evaluated in a simple but effective way, we can identify and prune the least important filters with less computation cost. We conduct comprehensive experiments on benchmark datasets employing various widely used CNN architectures to evaluate the performance of our method. For instance, on ILSVRC-2012, our method outperforms state-of-the-art methods by reducing floating-point operations (FLOPs) by 77.4% and parameters by 69.3% for ResNet-50 with only a minor decrease in an accuracy of 2.64%. Shuo Ye, Yufeng Shi 0003, Tianheng Hu, Qinmu Peng, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Sparse Additive Machine With the Correntropy-Induced LossabstractSparse additive machines (SAMs) have shown competitive performance on variable selection and classification in high-dimensional data due to their representation flexibility and interpretability. However, the existing methods often employ the unbounded or nonsmooth functions as the surrogates of 0-1 classification loss, which may encounter the degraded performance for data with outliers. To alleviate this problem, we propose a robust classification method, named SAM with the correntropy-induced loss (CSAM), by integrating the correntropy-induced loss (C-loss), the data-dependent hypothesis space, and the weighted -norm regularizer ( ) into additive machines. In theory, the generalization error bound is estimated via a novel error decomposition and the concentration estimation techniques, which shows that the convergence rate can be achieved under proper parameter conditions. In addition, the theoretical guarantee on variable selection consistency is analyzed. Experimental evaluations on both synthetic and real-world datasets consistently validate the effectiveness and robustness of the proposed approach. Peipei Yuan, Xinge You, Hong Chen 0004, Yingjie Wang 0007, Qinmu Peng, Bin Zou 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Multiplex Experts Governance Collaboration for Label Noise-Resistant Graph Representation LearningabstractRecently emerged label noise-resistant graph representation learning (LNR-GRL) has received increasing attention, which aims to enhance the generalization of graph neural networks (GNNs) in semi-supervised node classification with noisy and limited labels. Most of the existing LNR-GRL tend to introduce more complex sample selection strategies developed in nongraph areas to distinguish more noisy nodes to alleviate their misguidance. However, these proposed methods neglect the importance of inaccurate graph structure relationships rectification, and information collaboration between inaccurate graph structure relationships and noisy node label rectification in improving the quality of noisy node identification and its rectified node labels. To solve the above-mentioned issues, we propose a novel multiplex experts governance collaboration (MEGC) framework for LNR-GRL. Specifically, an unsupervised graph structure governance expert is first designed to rectify inaccurate graph structure relationships. Based on the rectified graph structure, a simple label noise governance expert is proposed to accurately identify noisy node labels and further improve the quality of noisy nodes’ rectified labels and unlabeled nodes’ pseudo-labels. Finally, the above-proposed governance experts can be effectively combined with GNNs to jointly guide their training via the introduced cross-view graph contrastive loss and cross-entropy loss, which can maximally limit the effect of noisy node labels and discover more effective supervision guidance from data itself for GNNs optimization. Extensive experiments on three benchmarks, two label noise types, four noise rates, and four training label rates demonstrate the superiority of the proposed method in comparison to the existing LNR-GRL methods. Sichao Fu, Qinmu Peng, Yiu-Ming Cheung, Yizhuo Xu, Bin Zou 0002, Xiaoyuan Jing, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2024 | Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot LearningabstractGenerative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes, which is an effective way to advance ZSL. However, existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype, which limit the generator only optimized on specific seen classes rather than characterizing each visual instance, resulting in poor generalizations (e.g., overfitting to seen classes). To address this issue, we propose a novel Visual-Augmented Dynamic Semantic prototype method (termed VADS) to boost the generator to learn accurate semantic-visual mapping by fully exploiting the visual-augmented knowledge into semantic conditions. In detail, VADS consists of two modules: (1) Visual-aware Domain Knowledge Learning module (VDKL) learns the local bias and global prior of the visual features (referred to as domain visual knowledge), which replace pure Gaussian noise to provide richer prior noise information; (2) VisionOriented Semantic Updation module (VOSU) updates the semantic prototype according to the visual representations of the samples. Ultimately, we concatenate their output as a dynamic semantic prototype, which serves as the condition of the generator. Extensive experiments demonstrate that our VADS achieves superior CZSL and GZSL performances on three prominent datasets and outperforms other state-of-the-art methods with averaging increases by 6.4%, 5.9% and 4.2% on SUN, CUB and AWA2, respectively. Wenjin Hou, Shiming Chen 0002, Shuhuang Chen, Ziming Hong, Xuetao Feng, Salman Khan 0001, Fahad Shahbaz Khan, Xinge You |
CVPR | 9 |
| 2024 | SocialCircle: Learning the Angle-based Social Interaction Representation for Pedestrian Trajectory PredictionabstractAnalyzing and forecasting trajectories of agents like pedestrians and cars in complex scenes has become more and more significant in many intelligent systems and ap-plications. The diversity and uncertainty in socially inter-active behaviors among a rich variety of agents make this task more challenging than other deterministic computer vision tasks. Researchers have made a lot of efforts to quan-tify the effects of these interactions on future trajectories through different mathematical models and network structures, but this problem has not been well solved. Inspired by marine animals that localize the positions of their com-panions underwater through echoes, we build a new angle-based trainable social interaction representation, named SocialCircle, for continuously reflecting the context of social interactions at different angular orientations relative to the target agent. We validate the effect of the proposed So-ciaiCircle by training it along with several newly released trajectory prediction models, and experiments show that the SocialCircle not only quantitatively improves the prediction performance, but also qualitatively helps better simulate social interactions when forecasting pedestrian trajectories in a way that is consistent with human intuitions. Conghao Wong, Beihao Xia, Ziqian Zou, Xinge You |
CVPR | 5 |
| 2024 | Generalized Sparse Additive Model with Unknown Link FunctionabstractGeneralized additive models (GAMs) have been successfully applied to high dimensional data. However, most existing methods cannot capture the high level feature patterns from complex data. To alleviate this problem, we propose a new sparse additive model, named generalized sparse additive model with unknown link function (GSAMUL), in which the component functions are estimated by B-spline basis and the unknown link function is estimated by a multi-layer perceptron (MLP) network. Furthermore,$\mathscr{l}_{2.1}$-norm regularizer is used for variable selection. The proposed GSAMUL can realize both variable selection and hidden interaction. We integrate this estimation into a bilevel optimization problem, where the data is split into training set and validation set. In theory, we provide the guarantees about the convergence of the approximate procedure. In applications, experimental evaluations on both synthetic and real world data sets consistently validate the effectiveness of GSAMUL. Peipei Yuan, Xinge You, Hong Chen 0004, Qinmu Peng |
ICDM | 2 |
| 2024 | Optimal Kernel Choice for Score Function-based Causal DiscoveryabstractScore-based methods have demonstrated their effectiveness in discovering causal relationships by scoring different causal structures based on their goodness of fit to the data. Recently, Huang et al. proposed a generalized score function that can handle general data distributions and causal relationships by modeling the relations in reproducing kernel Hilbert space (RKHS). The selection of an appropriate kernel within this score function is crucial for accurately characterizing causal relationships and ensuring precise causal discovery. However, the current method involves manual heuristic selection of kernel parameters, making the process tedious and less likely to ensure optimality. In this paper, we propose a kernel selection method within the generalized score function that automatically selects the optimal kernel that best fits the data. Specifically, we model the generative process of the variables involved in each step of the causal graph search procedure as a mixture of independent noise variables. Based on this model, we derive an automatic kernel selection method by maximizing the marginal likelihood of the variables involved in each search step. We conduct experiments on both synthetic data and real-world benchmarks, and the results demonstrate that our proposed method outperforms heuristic kernel selection methods. Biwei Huang, Feng Liu 0003, Xinge You, Tongliang Liu, Kun Zhang 0001, Mingming Gong |
ICML | 4 |
| 2024 | Causal Visual-semantic Correlation for Zero-shot Learning
Shuhuang Chen, Dingjie Fu, Shiming Chen 0002, Shuo Ye, Wenjin Hou, Xinge You |
ACM Multimedia | 6 |
| 2024 | Rethinking attribute localization for zero-shot learning
Shuhuang Chen, Shiming Chen 0002, Guosen Xie, Xiangbo Shu, Xinge You, Xuelong Li 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | Enhancing robustness of person detection: A universal defense filter against adversarial patch attacks
Zimin Mao, Shuiyan Chen, Zhuang Miao, Heng Li 0008, Beihao Xia, Junzhe Cai, Wei Yuan 0001, Xinge You |
Comput. Secur. | 8 |
| 2024 | Semi-supervised anomaly detection with contamination-resilience and incremental training
Liheng Yuan, Fanghua Ye 0001, Heng Li 0008, Cuiying Gao, Chengqing Yu, Wei Yuan 0001, Xinge You |
Eng. Appl. Artif. Intell. | 8 |
| 2024 | Concept drift adaptation with scarce labels: A novel approach based on diffusion and adversarial learning
Liheng Yuan, Fanghua Ye 0001, Wei Yuan 0001, Xinge You |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Hybrid learning based on Fisher linear discriminant
Jiawen Gong, Bin Zou 0002, Chen Xu 0007, Jie Xu 0006, Xinge You |
Inf. Sci. | 5 |
| 2024 | R2-trans: Fine-grained visual categorization with redundancy reduction
Shuo Ye, Shujian Yu, Yu Wang 0106, Xinge You |
Image Vis. Comput. | 4 |
| 2024 | The Image Data and Backbone in Weakly Supervised Fine-Grained Visual Categorization: A Revisit and Further ThinkingabstractWeakly-supervised fine-grained visual categorization (FGVC) aims to achieve subclass classification within the same large class using only label information. Compared to general images, fine-grained images have similar appearances and features, and are often affected by disturbances such as viewpoint, lighting, and occlusion during data collection, resulting in significant intra-class variance and small inter-class variance. To achieve FGVC, carefully designed models are often needed to explore the locally discriminative regions of the image. This paper revisits high-quality FGVC publications based on deep learning and analyzes from two new perspective: fine-grained image data and backbone. We address two ignored but interesting problems in FGVC. First, we argue that the reasons for exacerbating intra-class variance are not the same in data of animal, plant, and commodity types, and it is necessary to consider the effects of posture, covariate shift, and structural changes. Additionally, the “soft boundary” between subclasses intensifies the difficulty of classification. Second, we highlight that convolutional networks and self-attention networks have different receptive fields and shape biases, leading to performance differences when processing different types of fine-grained data. Overall, our analysis provides new insights into recent advances, challenges, and future directions for FGVC based on deep learning, which can help researchers develop more effective models for FGVC. Shuo Ye, Yu Wang 0106, Qinmu Peng, Xinge You, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | EGANS: Evolutionary Generative Adversarial Network Search for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize the novel classes which cannot be collected for training a prediction model. Accordingly, generative models (e.g., generative adversarial network (GAN)) are typically used to synthesize the visual samples conditioned by the class semantic vectors and achieve remarkable progress for ZSL. However, existing GAN-based generative ZSL methods are based on hand-crafted models, which cannot adapt to various datasets/scenarios and fails to model instability. To alleviate these challenges, we propose evolutionary generative adversarial network search (termed EGANS) to automatically design the generative network with good adaptation and stability, enabling reliable visual feature sample synthesis for advancing ZSL. Specifically, we adopt cooperative dual evolution to conduct a neural architecture search for both generator and discriminator under a unified evolutionary adversarial framework. EGANS is learned by two stages: evolution generator architecture search and evolution discriminator architecture search. During the evolution generator architecture search, we adopt a many-to-one adversarial training strategy to evolutionarily search for the optimal generator. Then the optimal generator is further applied to search for the optimal discriminator in the evolution discriminator architecture search with a similar evolution search algorithm. Once the optimal generator and discriminator are searched, we entail them into various generative ZSL baselines for ZSL classification. Extensive experiments show that EGANS consistently improve existing generative ZSL methods on the standard CUB, SUN, AWA2 and FLO datasets. The significant performance gains indicate that the evolutionary neural architecture search explores a virgin field in ZSL. Shiming Chen 0002, Shuhuang Chen, Wenjin Hou, Weiping Ding 0001, Xinge You |
IEEE Trans. Evol. Comput. | 5 |
| 2024 | Few-Shot Learning With Dynamic Graph Structure PreservingabstractIn recent years, few-shot learning has received increasing attention in the Internet of Things areas. Few-shot learning aims to distinguish unseen classes with a few labeled samples from each class. Most recently transductive few-shot studies highly rely on the static geometry distributions generated on the feature space during the label propagation process between unseen class instances. However, these recent methods fail to guarantee that the generated graph structure preserves the true distributions between data properly. In this article, we propose a novel dynamic graph structure preserving (DGSP) model for few-shot learning. Specifically, we formulate the objective function of DGSP by simultaneously considering the data correlations from the feature space and the label space to update the generated graph structure, which can reasonably revise the inappropriate or mistaken local geometry relationships. Then, we design an efficient alternating optimization algorithm to jointly learn the label prediction matrix and the optimal graph structure, the latter of which can be formulated as a linear programming problem. Moreover, our proposed DGSP can be easily combined with any backbone networks during the learning process. We conduct extensive experimental results across different benchmarks, backbones, and task settings, and our method achieves state-of-the-art performance compared with methods based on transductive few-shot learning. Sichao Fu, Qiong Cao, Yunwen Lei, Yibing Zhan, Xinge You |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | ECEA: Extensible Co-Existing Attention for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) identifies objects from extremely few annotated samples. Most existing FSOD methods, recently, apply the two-stage learning paradigm, which transfers the knowledge learned from abundant base classes to assist the few-shot detectors by learning the global features. However, such existing FSOD approaches seldom consider the localization of objects from local to global. Limited by the scarce training data in FSOD, the training samples of novel classes typically capture part of objects, resulting in such FSOD methods being unable to detect the completely unseen object during testing. To tackle this problem, we propose an Extensible Co-Existing Attention (ECEA) module to enable the model to infer the global object according to the local parts. Specifically, we first devise an extensible attention mechanism that starts with a local region and extends attention to co-existing regions that are similar and adjacent to the given local region. We then implement the extensible attention mechanism in different feature scales to progressively discover the full object in various receptive fields. In the training process, the model learns the extensible ability on the base stage with abundant samples and transfers it to the novel stage of continuous extensible learning, which can assist the few-shot model to quickly adapt in extending local regions to co-existing regions. Extensive experiments on the PASCAL VOC and COCO datasets show that our ECEA module can assist the few-shot detector to completely predict the object despite some regions failing to appear in the training samples and achieve the new state-of-the-art compared with existing FSOD methods. Code is released at https://github.com/zhimengXin/ECEA. Zhimeng Xin, Tianxu Wu, Shiming Chen 0002, Yixiong Zou, Ling Shao 0001, Xinge You |
IEEE Trans. Image Process. | 6 |
| 2024 | GNDAN: Graph Navigated Dual Attention Network for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the unseen class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a direct embedding is adopted for associating the visual and semantic domains in ZSL. However, most existing ZSL methods focus on learning the embedding from implicit global features or image regions to the semantic space. Thus, they fail to: 1) exploit the appearance relationship priors between various local regions in a single image, which corresponds to the semantic information and 2) learn cooperative global and local features jointly for discriminative feature representations. In this article, we propose the novel graph navigated dual attention network (GNDAN) for ZSL to address these drawbacks. GNDAN employs a region-guided attention network (RAN) and a region-guided graph attention network (RGAT) to jointly learn a discriminative local embedding and incorporate global context for exploiting explicit global embeddings under the guidance of a graph. Specifically, RAN uses soft spatial attention to discover discriminative regions for generating local embeddings. Meanwhile, RGAT employs an attribute-based attention to obtain attribute-based region features, where each attribute focuses on the most relevant image regions. Motivated by the graph neural network (GNN), which is beneficial for structural relationship representations, RGAT further leverages a graph attention network to exploit the relationships between the attribute-based region features for explicit global embedding representations. Based on the self-calibration mechanism, the joint visual embedding learned is matched with the semantic embedding to form the final prediction. Extensive experiments on three benchmark datasets demonstrate that the proposed GNDAN achieves superior performances to the state-of-the-art methods. Our code and trained models are available at https://github.com/shiming-chen/GNDAN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Qinmu Peng, Xinge You, Weiping Ding 0001, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Discriminative Suprasphere Embedding for Fine-Grained Visual CategorizationabstractDespite the great success of the existing work in fine-grained visual categorization (FGVC), there are still several unsolved challenges, e.g., poor interpretation and vagueness contribution. To circumvent this drawback, motivated by the hypersphere embedding method, we propose a discriminative suprasphere embedding (DSE) framework, which can provide intuitive geometric interpretation and effectively extract discriminative features. Specifically, DSE consists of three modules. The first module is a suprasphere embedding (SE) block, which learns discriminative information by emphasizing weight and phase. The second module is a phase activation map (PAM) used to analyze the contribution of local descriptors to the suprasphere feature representation, which uniformly highlights the object region and exhibits remarkable object localization capability. The last module is a class contribution map (CCM), which quantitatively analyzes the network classification decision and provides insight into the domain knowledge about classified objects. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed method in comparison with state-of-the-art methods. Shuo Ye, Qinmu Peng, Wenju Sun, Jiamiao Xu, Yu Wang 0106, Xinge You, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | MultiCut-MultiMix: a two-level data augmentation method for detecting small and densely distributed objects in large-size images
Zhimeng Xin, Tongwei Lu, Xinge You |
Vis. Comput. | 4 |
| 2023 | Self-Supervised Guided Hypergraph Feature Propagation for Semi-Supervised Classification with Missing Node FeaturesabstractGraph neural networks (GNNs) with missing node features have recently received increasing interest. Such missing node features seriously hurt the performance of the existing GNNs. Some recent methods have been proposed to reconstruct the missing node features by the information propagation among nodes with known and unknown attributes. Although these methods have achieved superior performance, how to exactly exploit the complex data correlations among nodes to reconstruct missing node features is still a great challenge. To solve the above problem, we propose a self-supervised guided hypergraph feature propagation (SGHFP). Specifically, the feature hypergraph is first generated according to the node features with missing information. And then, the reconstructed node features produced by the previous iteration are fed to a two-layer GNNs to construct a pseudo-label hypergraph. Before each iteration, the constructed feature hypergraph and pseudo-label hypergraph are fused effectively, which can better preserve the higher-order data correlations among nodes. After then, we apply the fused hypergraph to the feature propagation for reconstructing missing features. Finally, the reconstructed node features by multi-iteration optimization are applied to the downstream semi-supervised classification task. Extensive experiments demonstrate that the proposed SGHFP outperforms the existing semi-supervised classification with missing node feature methods. Chengxiang Lei, Sichao Fu, Yuetian Wang, Wenhao Qiu, Yachen Hu, Qinmu Peng, Xinge You |
ICASSP | 7 |
| 2023 | Towards Unsupervised Graph Completion Learning on Graphs with Features and Structure MissingabstractIn recent years, graph neural networks (GNN) have achieved significant developments in a variety of graph analytical tasks. Nevertheless, GNN’s superior performance will suffer from serious damage when the collected node features or structure relationships are partially missing owning to numerous unpredictable factors. Recently emerged graph completion learning (GCL) has received increasing attention, which aims to reconstruct the missing node features or structure relationships under the guidance of a specifically supervised task. Although these proposed GCL methods have made great success, they still exist the following problems: the reliance on labels, the bias of the reconstructed node features and structure relationships. Besides, the generalization ability of the existing GCL still faces a huge challenge when both collected node features and structure relationships are partially missing at the same time. To solve the above issues, we propose a more general GCL framework with the aid of self-supervised learning for improving the task performance of the existing GNN variants on graphs with features and structure missing, termed unsupervised GCL (UGCL). Specifically, to avoid the mismatch between missing node features and structure during the message-passing process of GNN, we separate the feature reconstruction and structure reconstruction and design its personalized model in turn. Then, a dual contrastive loss on the structure level and feature level is introduced to maximize the mutual information of node representations from feature reconstructing and structure reconstructing paths for providing more supervision signals. Finally, the reconstructed node features and structure can be applied to the downstream node classification task. Extensive experiments on eight datasets demonstrate the effectiveness of our proposed method. Sichao Fu, Qinmu Peng, Baokun Du, Xinge You |
ICDM | 5 |
| 2023 | Evolving Semantic Prototype Improves Generative Zero-Shot LearningabstractIn zero-shot learning (ZSL), generative methods synthesize class-related sample features based on predefined semantic prototypes. They advance the ZSL performance by synthesizing unseen class sample features for better training the classifier. We observe that each class’s predefined semantic prototype (also referred to as semantic embedding or condition) does not accurately match its real semantic prototype. So the synthesized visual sample features do not faithfully represent the real sample features, limiting the classifier training and existing ZSL performance. In this paper, we formulate this mismatch phenomenon as the visual-semantic domain shift problem. We propose a dynamic semantic prototype evolving (DSP) method to align the empirically predefined semantic prototypes and the real prototypes for class-related feature synthesis. The alignment is learned by refining sample features and semantic prototypes in a unified framework and making the synthesized visual sample features approach real sample features. After alignment, synthesized sample features from unseen classes are closer to the real sample features and benefit DSP to improve existing generative ZSL methods by 8.5%, 8.0%, and 9.7% on the standard CUB, SUN AWA2 datasets, the significant performance improvement indicates that evolving semantic prototype explores a virgin field in ZSL. Shiming Chen 0002, Wenjin Hou, Ziming Hong, Xiaohan Ding, Yibing Song, Xinge You, Tongliang Liu, Kun Zhang 0001 |
ICML | 6 |
| 2023 | Attention-Based Deep Convolutional Network for Speech Recognition Under Multi-scene Noise Environment
Chuanwu Yang, Shuo Ye, Zhishu Lin, Qinmu Peng, Jiamiao Xu, Peipei Yuan, Yuetian Wang, Xinge You |
ICONIP (9) | 8 |
| 2023 | Coping with change: Learning invariant and minimum sufficient representations for fine-grained visual categorization
Shuo Ye, Shujian Yu, Wenjin Hou, Yu Wang 0106, Xinge You |
Comput. Vis. Image Underst. | 5 |
| 2023 | TransZero++: Cross Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the novel class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is typically represented by attribute descriptions shared between different classes, which act as strong priors for localizing object attributes that represent discriminative region features, enabling significant and sufficient visual-semantic interaction for advancing ZSL. Existing attention-based models have struggled to learn inferior region features in a single image by solely using unidirectional attention, which ignore the transferable and discriminative attribute localization of visual features for representing the key semantic knowledge for effective knowledge transfer in ZSL. In this paper, we propose a cross attribute-guided Transformer network, termed TransZero++, to refine visual features and learn accurate attribute localization for key semantic knowledge representations in ZSL. Specifically, TransZero++ employs an attribute → visual Transformer sub-net (AVT) and a visual → attribute Transformer sub-net (VAT) to learn attribute-based visual features and visual-based attribute features, respectively. By further introducing feature-level and prediction-level semantical collaborative losses, the two attribute-guided transformers teach each other to learn semantic-augmented visual embeddings for key semantic knowledge representations via semantical collaborative learning. Finally, the semantic-augmented visual embeddings learned by AVT and VAT are fused to conduct desirable visual-semantic interaction cooperated with class semantic vectors for ZSL classification. Extensive experiments show that TransZero++ achieves the new state-of-the-art results on three golden ZSL benchmarks and on the large-scale ImageNet dataset. The project website is available at: https://shiming-chen.github.io/TransZero-pp/TransZero-pp.html. Shiming Chen 0002, Ziming Hong, Wenjin Hou, Guosen Xie, Yibing Song, Jian Zhao 0006, Xinge You, Shuicheng Yan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Improving the Generalization of MAML in Few-Shot Classification via Bi-Level ConstraintabstractFew-shot classification (FSC), which aims to identify novel classes in the presence of a few labeled samples, has drawn vast attention in recent years. One of the representative few-shot classification methods is model-agnostic meta-learning (MAML), which focuses on learning an initialization that can quickly adapt to novel categories with a few annotated samples. However, due to insufficient samples, MAML can easily fall into the dilemma of overfitting. Most existing MAML-based methods either improve the inner-loop update rule to achieve better generalization or constrain the outer-loop optimization to learn a more desirable initialization, without considering improving the two optimization processes jointly, resulting in unsatisfactory performance. In this paper, we propose a bi-level constrained MAML (BLC-MAML) method for few-shot classification. Specifically, in the inner-loop optimization, we introduce a supervised contrastive loss to constrain the adaptation procedure, which can effectively increase the intra-class aggregation and inter-class separability, thus improving the generalization of the adapted model. In the case of the outer loop, we propose a cross-task metric (CTM) loss to constrain the adapted model to perform well on the different few-shot task. The CTM loss can enforce the adapted model to learn more discriminative and generalized feature representations, further boosting the generalization of the learned initialization. By simultaneously constraining the bi-level optimization procedure, the proposed BLC-MAML can learn an initialization with better generalization. Extensive experiments on several FSC benchmarks show that our method can effectively improve the performance of MAML under both the within-domain and cross-domain settings, and also perform favorably against the state-of-the-art FSC algorithms. Yuanjie Shao, Wenxiao Wu, Xinge You, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learning Performance of Weighted Distributed Learning With Support Vector MachinesabstractThe divide-and-conquer strategy is a very effective method of dealing with big data. Noisy samples in big data usually have a great impact on algorithmic performance. In this article, we introduce Markov sampling and different weights for distributed learning with the classical support vector machine (cSVM). We first estimate the generalization error of weighted distributed cSVM algorithm with uniformly ergodic Markov chain (u.e.M.c.) samples and obtain its optimal convergence rate. As applications, we obtain the generalization bounds of weighted distributed cSVM with strong mixing observations and independent and identically distributed (i.i.d.) samples, respectively. We also propose a novel weighted distributed cSVM based on Markov sampling (DM-cSVM). The numerical studies of benchmark datasets show that the DM-cSVM algorithm not only has better performance but also has less total time of sampling and training compared to other distributed algorithms. Bin Zou 0002, Chen Xu 0007, Jie Xu 0006, Xinge You, Yuan Yan Tang |
IEEE Trans. Cybern. | 5 |
| 2023 | MSN: Multi-Style Network for Trajectory PredictionabstractTrajectory prediction aims to forecast agents’ possible future locations considering their observations along with the video context. It is strongly needed by many autonomous platforms like tracking, detection, robot navigation, and self-driving cars. Whether it is agents’ internal personality factors, interactive behaviors with the neighborhood, or the influence of surroundings, they all impact agents’ future planning. However, many previous methods model and predict agents’ behaviors with the same strategy or feature distribution, making them challenging to make predictions with sufficient style differences. This paper proposes the Multi-Style Network (MSN), which utilizes style proposal and stylized prediction using two sub-networks, to provide multi-style predictions in a novel categorical way adaptively. The proposed network contains a series of style channels, and each channel is bound to a unique and specific behavior style. We use agents’ end-point plannings and their interaction context as the basis for the behavior classification, so as to adaptively learn multiple diverse behavior styles through these channels. Then, we assume that the target agents may plan their future behaviors according to each of these categorized styles, thus utilizing different style channels to make predictions with significant style differences in parallel. Experiments show that the proposed MSN outperforms current state-of-the-art methods up to 10% quantitatively on two widely used datasets, and presents better multi-style characteristics qualitatively. Conghao Wong, Beihao Xia, Qinmu Peng, Wei Yuan 0001, Xinge You |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | A Componentwise Approach to Weakly Supervised Semantic Segmentation Using Dual-Feedback NetworkabstractRecent weakly supervised semantic segmentation methods generate pseudolabels to recover the lost position information in weak labels for training the segmentation network. Unfortunately, those pseudolabels often contain mislabeled regions and inaccurate boundaries due to the incomplete recovery of position information. It turns out that the result of semantic segmentation becomes determinate to a certain degree. In this article, we decompose the position information into two components: high-level semantic information and low-level physical information, and develop a componentwise approach to recover each component independently. Specifically, we propose a simple yet effective pseudolabels updating mechanism to iteratively correct mislabeled regions inside objects to precisely refine high-level semantic information. To reconstruct low-level physical information, we utilize a customized superpixel-based random walk mechanism to trim the boundaries. Finally, we design a novel network architecture, namely, a dual-feedback network (DFN), to integrate the two mechanisms into a unified model. Experiments on benchmark datasets show that DFN outperforms the existing state-of-the-art methods in terms of intersection-over-union (mIoU). Zhengqiang Zhang, Qinmu Peng, Sichao Fu, Yiu-Ming Cheung, Shujian Yu, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Kernelized Similarity Learning and Embedding for Dynamic Texture SynthesisabstractDynamic texture (DT) exhibits statistical stationarity in the spatial domain and stochastic repetitiveness in the temporal dimension, indicating that different frames of DT possess a high similarity correlation that is critical prior knowledge. However, existing methods cannot effectively learn a synthesis model for high-dimensional DT from a small number of training samples. In this article, we propose a novel DT synthesis method, which makes full use of similarity as prior knowledge to address this issue. Our method is based on the proposed kernel similarity embedding, which can not only mitigate the high dimensionality and small sample issues, but also has the advantage of modeling nonlinear feature relationships. Specifically, we first put forward two hypotheses that are essential for the DT model to generate new frames using similarity correlations. Then, we integrate kernel learning and the extreme learning machine into a unified synthesis model to learn kernel similarity embeddings for representing DTs. Extensive experiments on DT videos collected from the Internet and two benchmark datasets, i.e., Gatech Graphcut Textures and Dyntex, demonstrate that the learned kernel similarity embeddings can provide discriminative representations for DTs. Further, our method can preserve the long-term temporal continuity of the synthesized DT sequences with excellent sustainability and generalization. Meanwhile, it effectively generates realistic DT videos with higher speed and lower computation than the current state-of-the-art methods. The code and more synthesis videos are available at our project pagehttps://shiming-chen.github.io/Similarity-page/Similarit.html. Shiming Chen 0002, Peng Zhang 0040, Guosen Xie, Qinmu Peng, Zehong Cao, Wei Yuan 0001, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | TransZero: Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero. Shiming Chen 0002, Ziming Hong, Yang Liu 0069, Guosen Xie, Baigui Sun, Hao Li 0030, Qinmu Peng, Ke Lu 0002, Xinge You |
AAAI | 9 |
| 2022 | MSDN: Mutually Semantic Distillation Network for Zero-Shot LearningabstractThe key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You |
CVPR | 8 |
| 2022 | View Vertically: A Hierarchical Network for Trajectory Prediction via Fourier Spectrums
Conghao Wong, Beihao Xia, Ziming Hong, Qinmu Peng, Wei Yuan 0001, Qiong Cao, Xinge You |
ECCV (22) | 8 |
| 2022 | Leachable Component ClusteringabstractClustering attempts to partition data instances into several distinctive groups, while the similarities among data belonging to the common partition can be principally reserved. Furthermore, incomplete data frequently occurs in many real-world applications, and brings perverse influence on pattern analysis. As a consequence, the specific solutions to data imputation and handling are developed to conduct the missing values of data, and independent stage of knowledge exploitation is absorbed for information understanding. In this work, a novel approach to clustering of incomplete data, termed leachable component clustering, is proposed. Rather than existing methods, the proposed method handles data imputation with Bayes alignment, and collects the lost patterns in theory. Due to the simple numeric computation of equations, the proposed method can learn optimized partitions while the calculation efficiency is held. Experiments on several artificial incomplete data sets demonstrate that, the proposed method is able to present superior performance compared with other state-of-the-art algorithms. Xinge You |
ICPR | 2 |
| 2022 | Semantic Compression Embedding for Generative Zero-Shot LearningabstractGenerative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use visual features extracted by the pre-trained CNN backbone. These visual features lack attribute-level semantic information. Consequently, seen classes are indistinguishable, and the knowledge transfer from seen to unseen classes is limited. To tackle this issue, we propose a novel Semantic Compression Embedding Guided Generation (SC-EGG) model, which cascades a semantic compression embedding network (SCEN) and an embedding guided generative network (EGGN). The SCEN extracts a group of attribute-level local features for each sample and further compresses them into the new low-dimension visual feature. Thus, a dense-semantic visual space is obtained. The EGGN learns a mapping from the class-level semantic space to the dense-semantic visual space, thus improving the discriminability of the synthesized dense-semantic unseen visual features. Extensive experiments on three benchmark datasets, i.e., CUB, SUN and AWA2, demonstrate the significant performance gains of SC-EGG over current state-of-the-art methods and its baselines. Ziming Hong, Shiming Chen 0002, Guosen Xie, Wenhan Yang, Jian Zhao 0006, Yuanjie Shao, Qinmu Peng, Xinge You |
IJCAI | 8 |
| 2022 | Recent Advances in Concept Drift Adaptation Methods for Deep LearningabstractIn the ``Big Data'' age, the amount and distribution of data have increased wildly and changed over time in various time-series-based tasks, e.g weather prediction, network intrusion detection. However, deep learning models may become outdated facing variable input data distribution, which is called concept drift. To address this problem, large number of samples are usually required to update deep learning models, which is impractical in many realistic applications. This challenge drives researchers to explore the effective ways to adapt deep learning models to concept drift. In this paper, we first mathematically describe the categories of concept drift including abrupt drift, gradual drift, recurrent drift, incremental drift. We then divide existing studies into two categories (i.e., model parameter updating and model structure updating), and analyze the pros and cons of representative methods in each category. Finally, we evaluate the performance of these methods, and point out the future directions of concept drift adaptation for deep learning. Liheng Yuan, Heng Li 0008, Beihao Xia, Cuiying Gao, Wei Yuan 0001, Xinge You |
IJCAI | 7 |
| 2022 | DIT-NET: Joint Deformable Network and Intra-class Transfer GAN for Cross-domain 3D Neonatal Brain MRI Segmentation
Xinge You, Qinmu Peng, Chuanwu Yang |
PRCV (2) | 2 |
| 2022 | A Unified B-Spline Framework for Scale-Invariant Keypoint Detection
Qi Zheng 0003, Mingming Gong, Xinge You, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2022 | PSIDP: Unsupervised deep hashing with pretrained semantic information distillation and preservation
Yufeng Shi 0003, Xinge You, Jiamiao Xu, Weihua Ou, Feng Zheng 0001, Qinmu Peng |
Neurocomputing | 2 |
| 2022 | Incremental Fisher linear discriminant based on data denoisingabstractIn this article we consider Incremental Fisher linear discriminant (IFLD) based on data denoising. The data denoising is completed by Markov sampling such that the generated non-noise sample sequence is an uniformly ergodic Markov chain (u.e.M.c.). We first establish the generalization bounds of IFLD with u.e.M.c. samples, and prove that the IFLD algorithm with u.e.M.c. samples is consistent. We also present two new IFLD classification algorithms based on Markov sampling, IFLD based on Markov sampling (IFLD-MS) and improved IFLD based on Markov sampling (IIFLD-MS). Experimental results of benchmark repository suggest that IFLD-MS and IIFLD-MS have better performance than the classical IFLD, the incremental support vector machine (ISVM) and other IFLD algorithms. Ting Liang, Bin Zou 0002, Yaling Cai, Jie Xu 0006, Xinge You |
Knowl. Based Syst. | 6 |
| 2022 | CSCNet: Contextual semantic consistency network for trajectory prediction in crowded spaces
Beihao Xia, Conghao Wong, Qinmu Peng, Wei Yuan 0001, Xinge You |
Pattern Recognit. | 5 |
| 2022 | Deep Adaptively-Enhanced Hashing With Discriminative Similarity Guidance for Unsupervised Cross-Modal RetrievalabstractCross-modal hashing that leverages hash functions to project high-dimensional data from different modalities into the compact common hamming space, has shown immeasurable potential in cross-modal retrieval. To ease labor costs, unsupervised cross-modal hashing methods are proposed. However, existing unsupervised methods still suffer from two factors in the optimization of hash functions: 1) similarity guidance, they barely give a clear definition of whether is similar or not between data points, leading to the residual of the redundant information; 2) optimization strategy, they ignore the fact that the similarity learning abilities of different hash functions are different, which makes the hash function of one modality weaker than the hash function of the other modality. To alleviate such limitations, this paper proposes an unsupervised cross-modal hashing method to train hash functions with discriminative similarity guidance and adaptively-enhanced optimization strategy, termed Deep Adaptively-Enhanced Hashing (DAEH). Specifically, to estimate the similarity relations with discriminability, Information Mixed Similarity Estimation (IMSE) is designed by integrating information from distance distributions and the similarity ratio. Moreover, Adaptive Teacher Guided Enhancement (ATGE) optimization strategy is also designed, which employs information theory to discover the weaker hash function and utilizes an extra teacher network to enhance it. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed DAEH against the state-of-the-arts. Yufeng Shi 0003, Xin Liu 0011, Feng Zheng 0001, Weihua Ou, Xinge You, Qinmu Peng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Deep medical cross-modal attention hashing
Weihua Ou, Yufeng Shi 0003, Jiaxin Deng, Xinge You, Anzhi Wang |
World Wide Web | 5 |
| 2021 | FREE: Feature Refinement for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to over-coming the problems of visual-semantic domain gap and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dataset bias between ImageNet and GZSL benchmarks. Such a bias inevitably results in poor-quality visual features for GZSL tasks, which potentially limits the recognition performance on both seen and unseen classes. In this paper, we propose a simple yet effective GZSL method, termed feature refinement for generalized zero-shot learning (FREE), to tackle the above problem. FREE employs a feature refinement (FR) module that in-corporates semantic→visual mapping into a unified generative model to refine the visual features of seen and unseen class samples. Furthermore, we propose a self-adaptive margin center loss (SAMC-loss) that cooperates with a semantic cycle-consistency loss to guide FR to learn class- and semantically-relevant representations, and concatenate the features in FR to extract the fully refined features. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of FREE over its baseline and current state-of-the-art methods. The code is available at https://github.com/shiming-chen/FREE. Shiming Chen 0002, Beihao Xia, Qinmu Peng, Xinge You, Feng Zheng 0001, Ling Shao 0001 |
ICCV | 5 |
| 2021 | Norm-guided Adaptive Visual Embedding for Zero-Shot Sketch-Based Image RetrievalabstractZero-shot sketch-based image retrieval (ZS-SBIR), which aims to retrieve photos with sketches under the zero-shot scenario, has shown extraordinary talents in real-world applications. Most existing methods leverage language models to generate class-prototypes and use them to arrange the locations of all categories in the common space for photos and sketches. Although great progress has been made, few of them consider whether such pre-defined prototypes are necessary for ZS-SBIR, where locations of unseen class samples in the embedding space are actually determined by visual appearance and a visual embedding actually performs better. To this end, we propose a novel Norm-guided Adaptive Visual Embedding (NAVE) model, for adaptively building the common space based on visual similarity instead of language-based pre-defined prototypes. To further enhance the representation quality of unseen classes for both photo and sketch modality, modality norm discrepancy and noisy label regularizer are jointly employed to measure and repair the modality bias of the learned common embedding. Experiments on two challenging datasets demonstrate the superiority of our NAVE over state-of-the-art competitors. Yufeng Shi 0003, Shiming Chen 0002, Qinmu Peng, Feng Zheng 0001, Xinge You |
IJCAI | 6 |
| 2021 | HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a common (latent) space is adopted for associating the visual and semantic domains in ZSL. However, existing common space learning methods align the semantic and visual domains by merely mitigating distribution disagreement through one-step adaptation. This strategy is usually ineffective due to the heterogeneous nature of the feature representations in the two domains, which intrinsically contain both distribution and structure variations. To address this and advance ZSL, we propose a novel hierarchical semantic-visual adaptation (HSVA) framework. Specifically, HSVA aligns the semantic and visual domains by adopting a hierarchical two-step adaptation, i.e., structure adaptation and distribution adaptation. In the structure adaptation step, we take two task-specific encoders to encode the source data (visual domain) and the target data (semantic domain) into a structure-aligned common space. To this end, a supervised adversarial discrepancy (SAD) module is proposed to adversarially minimize the discrepancy between the predictions of two task-specific classifiers, thus making the visual and semantic feature manifolds more closely aligned. In the distribution adaptation step, we directly minimize the Wasserstein distance between the latent multivariate Gaussian distributions to align the visual and semantic distributions using a common encoder. Finally, the structure and distribution adaptation are derived in a unified framework under two partially-aligned variational autoencoders. Extensive experiments on four benchmark datasets demonstrate that HSVA achieves superior performance on both conventional and generalized ZSL. The code is available at \url{https://github.com/shiming-chen/HSVA}. Shiming Chen 0002, Guosen Xie, Yang Liu 0069, Qinmu Peng, Baigui Sun, Hao Li 0030, Xinge You, Ling Shao 0001 |
NeurIPS | 7 |
| 2021 | Multiset Feature Learning for Highly Imbalanced Data ClassificationabstractWith the expansion of data, increasing imbalanced data has emerged. When the imbalance ratio (IR) of data is high, most existing imbalanced learning methods decline seriously in classification performance. In this paper, we systematically investigate the highly imbalanced data classification problem, and propose an uncorrelated cost-sensitive multiset learning (UCML) approach for it. Specifically, UCML first constructs multiple balanced subsets through random partition, and then employs the multiset feature learning (MFL) to learn discriminant features from the constructed multiset. To enhance the usability of each subset and deal with the non-linearity issue existed in each subset, we further propose a deep metric based UCML (DM-UCML) approach. DM-UCML introduces the generative adversarial network technique into the multiset constructing process, such that each subset can own similar distribution with the original dataset. To cope with the non-linearity issue, DM-UCML integrates deep metric learning with MFL, such that more favorable performance can be achieved. In addition, DM-UCML designs a new discriminant term to enhance the discriminability of learned metrics. Experiments on eight traditional highly class-imbalanced datasets and two large-scale datasets indicate that: the proposed approaches outperform state-of-the-art highly imbalanced learning methods and are more robust to high IR. Xiaoyuan Jing, Xinyu Zhang 0012, Xiaoke Zhu, Fei Wu 0004, Xinge You, Yang Gao 0001, Shiguang Shan, Jing-Yu Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | CDE-GAN: Cooperative Dual Evolution-Based Generative Adversarial NetworkabstractGenerative adversarial networks (GANs) have been a popular deep generative model for real-world applications. Despite many recent efforts on GANs that have been contributed, mode collapse and instability of GANs are still open problems caused by their adversarial optimization difficulties. In this article, motivated by the cooperative co-evolutionary algorithm, we propose a cooperative dual evolution-based GAN (CDE-GAN) to circumvent these drawbacks. In essence, CDE-GAN incorporates dual evolution with respect to the generator(s) and discriminators into a unified evolutionary adversarial framework to conduct effective adversarial multiobjective optimization. Thus, it exploits the complementary properties and injects dual mutation diversity into the training, to steadily diversify the estimated density in capturing multimodes and improve generative performance. Specifically, CDE-GAN decomposes the complex adversarial optimization problem into two subproblems (generation and discrimination), and each subproblem is solved with a separated subpopulation (E-GeneratorsandE-Discriminators), evolved by its own evolutionary algorithm. Additionally, we further propose aSoft Mechanismto balance the tradeoff between E-Generators and E-Discriminators to conduct steady training for CDE-GAN. Extensive experiments on one synthetic dataset and three real-world benchmark image datasets demonstrate that the proposed CDE-GAN achieves a competitive and superior performance in generating good quality and diverse samples over baselines. The code and more generated results are available at our project homepagehttps://shiming-chen.github.io/CDE-GAN-website/CDE-GAN.html. Shiming Chen 0002, Beihao Xia, Xinge You, Qinmu Peng, Zehong Cao, Weiping Ding 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2021 | Hyperspectral Image Classification via Spatial Window-Based Multiview Intact Feature LearningabstractDue to the high dimensionality of hyperspectral images (HSIs), more training samples are needed in general for better classification performance. However, surface materials cannot always provide sufficient training samples in practice. HSI classification with small size training samples is still a challenging problem. Multiview learning is a feasible way to improve the classification accuracy in the case of small training samples by combining information from different views. This article proposes a new spatial window-based multiview intact feature learning method (SWMIFL) for HSI classification. In the proposed SWMIFL, multiple features that reflect different information of the original image are extracted and spatial windows are imposed on training samples to select unlabeled samples. Then, multiview intact feature learning is performed to learn the intact feature of the training and unlabeled samples. Considering that neighboring samples are likely to belong to the same class, labels of spatial neighboring samples are determined by two factors including the labels of training samples that locate in the spatial window and the labels learned from the intact feature. Finally, unlabeled samples that have same labels under these two factors are treated as new training samples. Experimental results demonstrate that the proposed SWMIFL-based classification method outperforms several well-known HSI classification methods on three real-world data sets. Yiu-Ming Cheung, Xinge You, Qinmu Peng, Jiangtao Peng, Peipei Yuan, Yufeng Shi 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Modal-Regression-Based Structured Low-Rank Matrix Recovery for Multiview LearningabstractLow-rank Multiview Subspace Learning (LMvSL) has shown great potential in cross-view classification in recent years. Despite their empirical success, existing LMvSL-based methods are incapable of handling well view discrepancy and discriminancy simultaneously, which, thus, leads to performance degradation when there is a large discrepancy among multiview data. To circumvent this drawback, motivated by the block-diagonal representation learning, we propose structured low-rank matrix recovery (SLMR), a unique method of effectively removing view discrepancy and improving discriminancy through the recovery of the structured low-rank matrix. Furthermore, recent low-rank modeling provides a satisfactory solution to address the data contaminated by the predefined assumptions of noise distribution, such as Gaussian or Laplacian distribution. However, these models are not practical, since complicated noise in practice may violate those assumptions and the distribution is generally unknown in advance. To alleviate such a limitation, modal regression is elegantly incorporated into the framework of SLMR (termed MR-SLMR). Different from previous LMvSL-based methods, our MR-SLMR can handle any zero-mode noise variable that contains a wide range of noise, such as Gaussian noise, random noise, and outliers. The alternating direction method of multipliers (ADMM) framework and half-quadratic theory are used to optimize efficiently MR-SLMR. Experimental results on four public databases demonstrate the superiority of MR-SLMR and its robustness to complicated noise. Jiamiao Xu, Fangzhao Wang, Qinmu Peng, Xinge You, Xiaoyuan Jing, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Emotion Classification Using EEG Brain Signals and the Broad Learning SystemabstractThis article presents a new user-independent emotion classification method that classifies four distinct emotions using electroencephalograph (EEG) signals and the broad learning system (BLS). The public DEAP and MAHNOB-HCI databases are used. Just one EEG electrode channel is selected for the feature extraction process. Continuous wavelet transform (CWT) is then utilized to extract the proposed gray-scale image (GSI) feature which describes the EEG brain activation in both time and frequency domains. Finally, the new BLS is constructed for the emotion classification process, which successfully upgrades the efficiency of emotion classification based on EEG brain signals. The experiment results show that the proposed work produces a robust system with high accuracy of approximately 93.1% and training process time of approximately 0.7 s for the DEAP database, as well as, the high average accuracy of approximately 94.4% and training process time of approximately 0.6 s for MAHNOB-HCI database. Sali Issa, Qinmu Peng, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Adaptive Matching of Kernel MeansabstractAs a promising step, the performance of data analysis and feature learning are able to be improved if certain pattern matching mechanism is available. One of the feasible solutions can refer to the importance estimation of instances, and consequently, kernel mean matching (KMM) has become an important method for knowledge discovery and novelty detection in kernel machines. Furthermore, the existing KMM methods have focused on concrete learning frameworks. In this work, a novel approach to adaptive matching of kernel means is proposed, and selected data with high importance are adopted to achieve calculation efficiency with optimization. In addition, scalable learning can be conducted in proposed method as a generalized solution to matching of appended data. The experimental results on a wide variety of real-world data sets demonstrate the proposed method is able to give outstanding performance compared with several state-of-the-art methods, while calculation efficiency can be preserved. Xinge You |
ICPR | 2 |
| 2020 | Video scene parsing: An overview of deep learning methods and datasets
Xiyu Yan, Huihui Gong, Yong Jiang 0001, Shutao Xia, Feng Zheng 0001, Xinge You, Ling Shao 0001 |
Comput. Vis. Image Underst. | 6 |
| 2020 | Group sparse additive machine with average top-k loss
Peipei Yuan, Xinge You, Hong Chen 0004, Qinmu Peng, Zhou Xu 0003, Xiaoyuan Jing, Zhenyu He 0001 |
Neurocomputing | 2 |
| 2020 | Coarse-to-fine salient object detection with low-rank matrix recovery
Qi Zheng 0003, Shujian Yu, Xinge You |
Neurocomputing | 3 |
| 2020 | Automatic kidney segmentation in ultrasound images using subsequent boundary distance regression and pixelwise classification networks
Qinmu Peng, Zhengqiang Zhang, Xinge You, Katherine Fischer, Susan L. Furth, Gregory Tasian, Yong Fan 0001 |
Medical Image Anal. | 5 |
| 2020 | Multiview Hybrid Embedding: A Divide-and-Conquer ApproachabstractWe present a novel cross-view classification algorithm where the gallery and probe data come from different views. A popular approach to tackle this problem is the multiview subspace learning (MvSL) that aims to learn a latent subspace shared by multiview data. Despite promising results obtained on some applications, the performance of existing methods deteriorates dramatically when the multiview data is sampled from nonlinear manifolds or suffers from heavy outliers. To circumvent this drawback, motivated by the Divide-and-Conquer strategy, we propose multiview hybrid embedding (MvHE), a unique method of dividing the problem of cross-view classification into three subproblems and building one model for each subproblem. Specifically, the first model is designed to remove view discrepancy, whereas the second and third models attempt to discover the intrinsic nonlinear structure and to increase the discriminability in intraview and interview samples, respectively. The kernel extension is conducted to further boost the representation power of MvHE. Extensive experiments are conducted on four benchmark datasets. Our methods demonstrate the overwhelming advantages against the state-of-the-art MvSL-based cross-view classification approaches in terms of classification accuracy and robustness. Jiamiao Xu, Shujian Yu, Xinge You, Mengjun Leng, Xiaoyuan Jing, C. L. Philip Chen |
IEEE Trans. Cybern. | 3 |
| 2019 | Equally-Guided Discriminative Hashing for Cross-modal RetrievalabstractCross-modal hashing intends to project data from two modalities into a common hamming space to perform cross-modal retrieval efficiently. Despite satisfactory performance achieved on real applications, existing methods are incapable of effectively preserving semantic structure to maintain inter-class relationship and improving discriminability to make intra-class samples aggregated simultaneously, which thus limits the higher retrieval performance. To handle this problem, we propose Equally-Guided Discriminative Hashing (EGDH), which jointly takes into consideration semantic structure and discriminability. Specifically, we discover the connection between semantic structure preserving and discriminative methods. Based on it, we directly encode multi-label annotations that act as high-level semantic features to build a common semantic structure preserving classifier. With the common classifier to guide the learning of different modal hash functions equally, hash codes of samples are intra-class aggregated and inter-class relationship preserving. Experimental results on two benchmark datasets demonstrate the superiority of EGDH compared with the state-of-the-arts. Yufeng Shi 0003, Xinge You, Feng Zheng 0001, Qinmu Peng |
IJCAI | 2 |
| 2019 | Common Structured Low-Rank Matrix Recovery for Cross-View Classification
Zihan Long, Jiamiao Xu, Fangzhao Wang, Chuanwu Yang, Xinge You |
PRCV (1) | 5 |
| 2019 | Retrieval by Classification: Discriminative Binary Embedding for Sketch-Based Image Retrieval
Yufeng Shi 0003, Xinge You, Feng Zheng 0001, Qinmu Peng |
PRCV (3) | 2 |
| 2019 | Advances in data representation and learning for pattern analysis
C. L. Philip Chen, Xinge You, Xinbo Gao 0001, Tongliang Liu, Fionn Murtagh, Weifeng Liu 0001 |
Neurocomputing | 2 |
| 2019 | Multi-view common component discriminant analysis for cross-view classification
Xinge You, Jiamiao Xu, Wei Yuan 0001, Xiaoyuan Jing, Dacheng Tao, Taiping Zhang |
Pattern Recognit. | 1 |
| 2019 | Distance learning by mining hard and easy negative samples for person re-identification
Xiaoke Zhu, Xiaoyuan Jing, Fan Zhang 0028, Xinyu Zhang 0012, Xinge You, Xiang Cui |
Pattern Recognit. | 5 |
| 2019 | Robust Visual Tracking Using Multi-Frame Multi-Feature Joint ModelingabstractIt remains a huge challenge to design effective and efficient trackers under complex scenarios, including occlusions, illumination changes and pose variations. To cope with this problem, a promising solution is to integrate the temporal consistency across consecutive frames and multiple feature cues in a unified model. Motivated by this idea, we propose a novel correlation filter-based tracker in this paper, in which the temporal relatedness is reconciled under a multi-task learning framework and the multiple feature cues are modeled using a multi-view learning approach. We demonstrate that the resulting regression model can be efficiently learned by exploiting the structure of blockwise diagonal matrix. A fast blockwise diagonal matrix inversion algorithm is developed thereafter for efficient online tracking. Meanwhile, we incorporate an adaptive scale estimation mechanism to strengthen the stability of scale variation tracking. We implement our tracker using two types of features and test it on two benchmark datasets. The experimental results demonstrate the superiority of our proposed approach when compared with the other state-of-the-art trackers. Peng Zhang 0040, Shujian Yu, Jiamiao Xu, Xinge You, Xiubao Jiang, Xiaoyuan Jing, Dacheng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | New Incremental Learning Algorithm With Support Vector MachinesabstractIncremental learning is one of the most effective methods of learning accumulated data and large-scale data. The newly increased samples of the previously known works on incremental learning are usually independent and identically distributed. To study how dependent sampling methods influence the learning ability of incremental support vector machines (ISVM) algorithm, in this paper we introduce an ISVM based on Markov resampling (MR-ISVM), and give the experimental research on the learning ability of the MR-ISVM algorithm. The experimental results indicate that the MR-ISVM algorithm has not only smaller misclassification rates and sparser of the obtained classifiers, but also less total time of sampling and training compared to ISVM based on randomly independent sampling. We also compare it with other ISVM algorithms. Jie Xu 0006, Chen Xu 0007, Bin Zou 0002, Yuan Yan Tang, Jiangtao Peng, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2018 | Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition
Chaojian Yu, Qi Zheng 0003, Peng Zhang 0040, Xinge You |
ECCV (16) | 5 |
| 2018 | Emotion Assessment Based on EEG Brain Signals
Sali Issa, Qinmu Peng, Xinge You, Wahab Ali Shah |
ISDA (2) | 3 |
| 2018 | Robust Multi-view Subspace Learning Through Structured Low-Rank Matrix Recovery
Jiamiao Xu, Xinge You, Qi Zheng 0003, Fangzhao Wang, Peng Zhang 0040 |
PRCV (3) | 2 |
| 2018 | Multi-view manifold learning with locality alignment
Xinge You, Shujian Yu, Chang Xu 0002, Wei Yuan 0001, Xiaoyuan Jing, Taiping Zhang, Dacheng Tao |
Pattern Recognit. | 2 |
| 2018 | Mixed Noise Removal via Robust Constrained Sparse RepresentationabstractIn recent years, the sparse coding-based techniques have been widely used for image denoising. However, most of the sparse coding-based mixed noise reduction methods fail to take full advantage of the geometric structure of data samples. In other words, they neglect the common information shared by the similar patches in sparse coding. To address this concern, in this paper, we propose a robust constrained sparse representation (RCSR) method to remove mixed noise. By using the center coefficient of similar patches as the guider which is approximated by the coefficient of query patch in sparse coding, the geometric structure of data can be well preserved. Moreover, different from most existing two-stage mixed noise reduction methods that use explicit detectors to restrain impulse noise, the proposed RCSR adaptively adjusts the contribution of each pixel in the loss function to eliminate the influences of outliers. Experiments on the reconstruction of synthetic data and the removal of mixed noise in real images demonstrate the effectiveness of our proposed method. Licheng Liu, C. L. Philip Chen, Xinge You, Yuan Yan Tang, Yushu Zhang 0001, Shutao Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Semi-Supervised Cross-View Projection-Based Dictionary Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-id) has attracted a lot of research interest. When facing dramatic growth in new pedestrian videos, existing video-based person re-id methods usually need large quantities of labeled pedestrian videos to train a discriminative model. In practice, labeling large quantities of pedestrian videos is a costly and time-consuming task, which will limit the application of these methods in the real environment. Therefore, it is valuable and necessary to investigate how to learn a discriminative re-id model by using limited labeled training pedestrian videos. In this paper, we propose a semi-supervised cross-view projection-based dictionary learning (SCPDL) approach for video-based person re-id. Specifically, SCPDL jointly learns a pair of feature projection matrices and a pair of dictionaries by integrating the information contained in labeled and unlabeled pedestrian videos. With the learned feature projection matrices, the influence of variations within each video to the re-id can be reduced. With the learned dictionary pair, pedestrian videos from two different cameras can be converted into coding coefficients in a common representation space, such that the differences between different cameras can be bridged. In the learning process, the labeled pedestrian videos are used to ensure that the learned dictionaries have favorable discriminability; the large quantities of unlabeled pedestrian videos are used to ensure that SCPDL can better capture the variations between pedestrian videos, such that the learned dictionaries can own stronger representative capability. Experiments on two public pedestrian sequence data sets (iLIDS-VID and PRID 2011) demonstrate the effectiveness of the proposed approach. Xiaoke Zhu, Xiaoyuan Jing, Liang Yang 0002, Xinge You, Dan Chen 0001, Guangwei Gao, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection MatrixabstractPerson re-identification plays an important role in video surveillance and forensics applications. In many cases, person re-identification needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of the suspect. We call re-identification in this scenario as image to video person reidentification (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature projection matrix and heterogeneous dictionary pair learning (PHDL) approach for IVPR. Specifically, the PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence caused by the variations within each video on the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted by using the coding coefficients. Furthermore, to ensure that the obtained coding coefficients own favorable discriminability, the PHDL designs a point-to-set coefficient discriminant term. To make better use of the complementary spatial-temporal and visual appearance information contained in pedestrian video data, we further propose a multi-view PHDL approach, which can fuse different video information effectively in the dictionary learning process. Experiments on four publicly available person sequence data sets demonstrate the effectiveness of the proposed approaches. Xiaoke Zhu, Xiaoyuan Jing, Xinge You, Wangmeng Zuo, Shiguang Shan, Wei-Shi Zheng 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Video-Based Person Re-Identification by Simultaneously Learning Intra-Video and Inter-Video Distance MetricsabstractVideo-based person re-identification (re-id) is an important application in practice. Since large variations exist between different pedestrian videos, as well as within each video, it's challenging to conduct re-identification between pedestrian videos. In this paper, we propose a simultaneous intra-video and inter-video distance learning (SI2DL) approach for video-based person re-id. Specifically, SI2DL simultaneously learns an intravideo distance metric and an inter-video distance metric from the training videos. The intra-video distance metric is used to make each video more compact, and the inter-video one is used to ensure that the distance between truly matching videos is smaller than that between wrong matching videos. Considering that the goal of distance learning is to make truly matching video pairs from different persons be well separated with each other, we also propose a pair separation based SI2DL (P-SI2DL). P-SI2DL aims to learn a pair of distance metrics, under which any two truly matching video pairs can be well separated. Experiments on four public pedestrian image sequence datasets show that our approaches achieve the state-of-the-art performance. Xiaoke Zhu, Xiaoyuan Jing, Xinge You, Xinyu Zhang 0012, Taiping Zhang |
IEEE Trans. Image Process. | 3 |
| 2018 | k-Times Markov Sampling for SVMCabstractSupport vector machine (SVM) is one of the most widely used learning algorithms for classification problems. Although SVM has good performance in practical applications, it has high algorithmic complexity as the size of training samples is large. In this paper, we introduce SVM classification (SVMC) algorithm based on -times Markov sampling and present the numerical studies on the learning performance of SVMC with -times Markov sampling for benchmark data sets. The experimental results show that the SVMC algorithm with -times Markov sampling not only have smaller misclassification rates, less time of sampling and training, but also the obtained classifier is more sparse compared with the classical SVMC and the previously known SVMC algorithm based on Markov sampling. We also give some discussions on the performance of SVMC with -times Markov sampling for the case of unbalanced training samples and large-scale training samples. Bin Zou 0002, Chen Xu 0007, Yang Lu 0009, Yuan Yan Tang, Jie Xu 0006, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2017 | Saliency Detection by Compactness Diffusion
Qi Zheng 0003, Peng Zhang 0040, Xinge You |
BMVC | 3 |
| 2017 | Generic Pixel Level Object Tracker Using Bi-Channel Fully Convolutional Network
Zijing Chen, Jun Li 0010, Zhe Chen 0013, Xinge You |
ICONIP (1) | 4 |
| 2017 | Efficient single image dehazing and denoising: An efficient multi-scale correlated wavelet approach
Xin Liu 0011, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
Comput. Vis. Image Underst. | 4 |
| 2017 | Dynamically Modulated Mask Sparse TrackingabstractVisual tracking is a critical task in many computer vision applications such as surveillance and robotics. However, although the robustness to local corruptions has been improved, prevailing trackers are still sensitive to large scale corruptions, such as occlusions and illumination variations. In this paper, we propose a novel robust object tracking technique depends on subspace learning-based appearance model. Our contributions are twofold. First, mask templates produced by frame difference are introduced into our template dictionary. Since the mask templates contain abundant structure information of corruptions, the model could encode information about the corruptions on the object more efficiently. Meanwhile, the robustness of the tracker is further enhanced by adopting system dynamic, which considers the moving tendency of the object. Second, we provide the theoretic guarantee that by adapting the modulated template dictionary system, our new sparse model can be solved by the accelerated proximal gradient algorithm as efficient as in traditional sparse tracking methods. Extensive experimental evaluations demonstrate that our method significantly outperforms 21 other cutting-edge algorithms in both speed and tracking accuracy, especially when there are challenges such as pose variation, occlusion, and illumination changes. Zijing Chen, Xinge You, Boxuan Zhong, Jun Li 0010, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2017 | Robust Object Tracking via Key Patch Sparse RepresentationabstractMany conventional computer vision object tracking methods are sensitive to partial occlusion and background clutter. This is because the partial occlusion or little background information may exist in the bounding box, which tends to cause the drift. To this end, in this paper, we propose a robust tracker based on key patch sparse representation (KPSR) to reduce the disturbance of partial occlusion or unavoidable background information. Specifically, KPSR first uses patch sparse representations to get the patch score of each patch. Second, KPSR proposes a selection criterion of key patch to judge the patches within the bounding box and select the key patch according to its location and occlusion case. Third, KPSR designs the corresponding contribution factor for the sampled patches to emphasize the contribution of the selected key patches. Comparing the KPSR with eight other contemporary tracking methods on 13 benchmark video data sets, the experimental results show that the KPSR tracker outperforms classical or state-of-the-art tracking methods in the presence of partial occlusion, background clutter, and illumination change. Zhenyu He 0001, Shuangyan Yi, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
IEEE Trans. Cybern. | 4 |
| 2017 | Coalition Formation and Spectrum Sharing of Cooperative Spectrum Sensing ParticipantsabstractIn cognitive radio networks, self-interested secondary users (SUs) desire to maximize their own throughput. They compete with each other for transmit time once the absence of primary users (PUs) is detected. To satisfy the requirement of PU protection, on the other hand, they have to form some coalitions and cooperate to conduct spectrum sensing. Such dilemma of SUs between competition and cooperation motivates us to study two interesting issues: 1) how to appropriately form some coalitions for cooperative spectrum sensing (CSS) and 2) how to share transmit time among SUs. We jointly consider these two issues, and propose a noncooperative game model with 2-D strategies. The first dimension determines coalition formation, and the second indicates transmit time allocation. Considering the complexity of solving this game, we decompose the game into two more tractable ones: one deals with the formation of CSS coalitions, and the other focuses on the allocation of transmit time. We characterize the Nash equilibria (NEs) of both games, and show that the combination of these two NEs corresponds to the NE of the original game. We also develop a distributed algorithm to achieve a desirable NE of the original game. When this NE is achieved, the SUs obtain a Dhp-stable coalition structure and a fair transmit time allocation. Numerical results verify our analyses, and demonstrate the effectiveness of our algorithm. Zhensheng Jiang, Wei Yuan 0001, Henry Leung 0001, Xinge You, Qi Zheng 0003 |
IEEE Trans. Cybern. | 4 |
| 2017 | Super-Resolution Person Re-Identification With Semi-Coupled Low-Rank Discriminant Dictionary LearningabstractPerson re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high resolution (HR), while probe images are usually low resolution (LR) in the identification scenarios with large variation of illumination, weather, or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification task. With the HR and LR dictionary pair and mapping matrices learned from the features of HR and LR training images, SLD2L can convert the features of the LR probe images into HR features. To ensure that the converted features have favorable discriminative capability and the learned dictionaries can well characterize intrinsic feature spaces of the HR and LR images, we design a discriminant term and a low-rank regularization term for SLD2L. Moreover, considering that low resolution results in different degrees of loss for different types of visual appearance features, we propose a multi-view SLD2L (MVSLD2L) approach, which can learn the type-specific dictionary pair and mappings for each type of feature. Experimental results on multiple publicly available data sets demonstrate the effectiveness of our proposed approaches for the SR person re-identification task. Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Ruimin Hu, Xinge You, Yunhong Wang 0001, Jing-Yu Yang 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | A Hybrid of Local and Global Saliencies for Detecting Image Salient Region and AppearanceabstractThis paper presents a visual saliency detection approach, which is a hybrid of local feature-based saliency and global feature-based saliency (simply called local saliency and global saliency, respectively, for short). First, we propose an automatic selection of smoothing parameter scheme to make the foreground and background of an input image more homogeneous. Then, we partition the smoothed image into a set of regions and compute the local saliency by measuring the color and texture dissimilarity in the smoothed regions and the original regions, respectively. Furthermore, we utilize the global color distribution model embedded with color coherence, together with the multiple edge saliency, to yield the global saliency. Finally, we combine the local and global saliencies, and utilize the composition information to obtain the final saliency. Experimental results show the efficacy of the proposed method, featuring: 1) the enhanced accuracy of detecting visual salient region and appearance in comparison with the existing counterparts, 2) the robustness against the noise and the low-resolution problem of images, and 3) its applicability to multisaliency detection task. Qinmu Peng, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Single image super-resolution with non-local balanced low-rank matrix restorationabstractSingle image super-resolution (SR) has gained popularity to construct a high-resolution (HR) image from a single low-resolution (LR) version. More recently, non-local self similarity (NSS) has been attracted enormous interests in the field of SR, and the non-local means (NLM)-based methods are classical NSS-based SR methods. However, NLM-based methods neglect the structure information in the patches and structural similarity between patches, so it will be prone to introduce unexpected details into resultant HR images. In this paper, we propose a non-local balanced low rank matrix restoration model (NB-LRM) to improve the performance of SR which will overcome the drawbacks of NLM-based methods and take full advantage of the NSS prior. The proposed algorithm formulates the constrained optimization problem for HR image recovery. First, to take advantage of the local structure in the patch and the structural similarity between the non-local similar patches, we propose a measurement of the similarity based on both Euclidean distance and Pearson distance, then reconstruct the target patch by weighted average the similar patches. Second, to guarantee the structural similarity and linear correlation between the target patch and similar patches, we propose a new low rank regular term. Third, we introduce the iterative low rank regular algorithm to solve our model. Addition, this method doesn't need other image priors and can produce more robust reconstruction of image local structures. Compared with state-of-the-art SR methods, the proposed NB-LRM method achieves highly competitive PSNR and SSIM result, while demonstrating better edge and texture preservation performance. Xinge You, Weiyong Xue, Jiajia Lei, Peng Zhang 0040, Yiu-Ming Cheung, Yuan Yan Tang, Naiding Zhou |
ICPR | 1 |
| 2016 | Big learning in social media analytics
C. L. Philip Chen, Dacheng Tao, Xinge You |
Neurocomputing | 3 |
| 2016 | STFT-like time frequency representations of nonstationary signal with arbitrary sampling schemes
Shujian Yu, Xinge You, Weihua Ou, Xiubao Jiang, Yi Mou |
Neurocomputing | 2 |
| 2016 | Multiscale patch-based contrast measure for small infrared target detection
Yantao Wei, Xinge You, Hong Li 0009 |
Pattern Recognit. | 2 |
| 2016 | Multi-view low-rank dictionary learning for image classification
Fei Wu 0004, Xiaoyuan Jing, Xinge You, Dong Yue 0001, Ruimin Hu, Jing-Yu Yang 0001 |
Pattern Recognit. | 3 |
| 2016 | Sparse discriminative multi-manifold embedding for one-sample face identification
Pengyue Zhang, Xinge You, Weihua Ou, C. L. Philip Chen, Yiu-Ming Cheung |
Pattern Recognit. | 2 |
| 2016 | Dynamic texture modeling and synthesis using multi-kernel Gaussian process dynamic model
Xinge You, Shujian Yu, Jixin Zou, Haiquan Zhao 0001 |
Signal Process. | 2 |
| 2016 | Multiobjective Optimization of Linear Cooperative Spectrum Sensing: Pareto Solutions and RefinementabstractIn linear cooperative spectrum sensing, the weights of secondary users and detection threshold should be optimally chosen to minimize missed detection probability and to maximize secondary network throughput. Since these two objectives are not completely compatible, we study this problem from the viewpoint of multiple-objective optimization. We aim to obtain a set of evenly distributed Pareto solutions. To this end, here, we introduce the normal constraint (NC) method to transform the problem into a set of single-objective optimization (SOO) problems. Each SOO problem usually results in a Pareto solution. However, NC does not provide any solution method to these SOO problems, nor any indication on the optimal number of Pareto solutions. Furthermore, NC has no preference over all Pareto solutions, while a designer may be only interested in some of them. In this paper, we employ a stochastic global optimization algorithm to solve the SOO problems, and then propose a simple method to determine the optimal number of Pareto solutions under a computational complexity constraint. In addition, we extend NC to refine the Pareto solutions and select the ones of interest. Finally, we verify the effectiveness and efficiency of the proposed methods through computer simulations. Wei Yuan 0001, Xinge You, Jing Xu 0005, Henry Leung 0001, Tianhang Zhang, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2016 | Connected Component Model for Multi-Object TrackingabstractIn multi-object tracking, it is critical to explore the data associations by exploiting the temporal information from a sequence of frames rather than the information from the adjacent two frames. Since straightforwardly obtaining data associations from multi-frames is an NP-hard multi-dimensional assignment (MDA) problem, most existing methods solve this MDA problem by either developing complicated approximate algorithms, or simplifying MDA as a 2D assignment problem based upon the information extracted only from adjacent frames. In this paper, we show that the relation between associations of two observations is the equivalence relation in the data association problem, based on the spatial-temporal constraint that the trajectories of different objects must be disjoint. Therefore, the MDA problem can be equivalently divided into independent subproblems by equivalence partitioning. In contrast to existing works for solving the MDA problem, we develop a connected component model (CCM) by exploiting the constraints of the data association and the equivalence relation on the constraints. Based upon CCM, we can efficiently obtain the global solution of the MDA problem for multi-object tracking by optimizing a sequence of independent data association subproblems. Experiments on challenging public data sets demonstrate that our algorithm outperforms the state-of-the-art approaches. Zhenyu He 0001, Xin Li 0034, Xinge You, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2016 | Kernel Learning for Dynamic Texture SynthesisabstractDynamic textures (DTs) that represent moving scenes such as flames, smoke, and waves, exhibit fixed dynamics within a period of time and have been successfully modeled using linear dynamic systems (LDS). In this paper, we show that the widely used LDS model can be approximated using a principal component regression (PCR) model with the main advantage of simplicity. Furthermore, to capture the nonlinearity of training frames, we extend traditional PCR to its kernelized version and introduce kernel principal component regression (KPCR) to model and synthesize DTs. To ensure algorithm stability, we remove the standard state model and directly apply the quantized kernel least mean squares algorithm from signal processing domain to approximate the performance achieved with KPCR. We term this improvement kernel adaptive dynamic texture synthesis (KADTS), which also has the benefits of computational and memory efficiency. These advantages make KADTS ideally suited for real-world applications, since the majority of electronic devices, including cell phones and laptops, suffer from limited memory and real-time constraints. We demonstrate, via both theoretical and experimental analyses, the connections between DT synthesis using KPCR and KADTS with a regularization network theory. We also show the superiority of our proposed algorithms for DT synthesis compared with other dynamic system-based benchmarks. MATLAB code is available from our project homepage http://bmal.hust.edu.cn/project/dts.html. Xinge You, Weigang Guo, Shujian Yu, Kan Li 0002, José C. Príncipe, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | Automatic Ear Landmark Localization, Segmentation, and Pose Classification in Range ImagesabstractMultibiometric systems using face and ear features are increasingly adopted for forensic and civilian applications to address the challenges of facial expressions and occlusions. Although numerous ear and face recognition techniques have been proposed, not much work has been conducted in the field of 3-D fiducial points localization and 3-D ear detection. This paper presents an effective and efficient system of ear landmark localization, ear detection, and pose classification based on 3-D ears captured under large yaw variations. By utilizing the symmetrical property of human heads and classifying the ear with respect to its pose, all three tasks can be fulfilled given either left or right ears, without any prior pose information. A novel ear tree-structured graph (ETG) is proposed to represent the 3-D ear, after which a 3-D flexible mixture model is trained to locate the landmarks automatically. Then, the ear region is segmented based on them and the pose of the ear, i.e., whether it is a left or right ear, is classified based on the detected ETG. To the best of our knowledge, this paper is the first to present automatic landmark localization of 3-D ears extracted from facial scans with significant pose variations. Experiments were conducted at the University of Notre Dame collection F, G and J2, which contain large occlusion and pose variations, validating the effectiveness of the proposed methods. Jiajia Lei, Xinge You, Mohamed Abdel-Mottaleb |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2015 | Super-resolution Person re-identification with semi-coupled low-rank discriminant dictionary learningabstractPerson re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high-resolution (HR) while probe images are usually low-resolution (LR) in the identification scenarios with large variation of illumination, weather or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification. For the given training image set which consists of HR gallery and LR probe images, we aim to convert the features of LR images into discriminating HR features. Specifically, our approach learns a pair of HR and LR dictionaries and a mapping from the features of HR gallery images and LR probe images. To ensure that the converted features using the learned dictionaries and mapping have favorable discriminative capability, we design a discriminant term which requires the converted HR features of LR probe images should be close to the features of HR gallery images from the same person, but far away from the features of HR gallery images from different persons. In addition, we apply low-rank regularization in dictionary learning procedure such that the learned dictionaries can well characterize intrinsic feature space of HR and LR images. Experimental results on public datasets demonstrate the effectiveness of SLD2L. Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Xinge You, Qinglong Liu, Dong Yue 0001, Ruimin Hu, Baowen Xu |
CVPR | 4 |
| 2015 | Webcam-Based Visual Gaze Estimation Under Desktop Environment
Shujian Yu, Weihua Ou, Xinge You, Xiubao Jiang, Yi Mou, Weigang Guo, Yuan Yan Tang, C. L. Philip Chen |
ICONIP (2) | 3 |
| 2015 | Generalized Kernel Normalized Mixed-Norm Algorithm: Analysis and Simulations
Shujian Yu, Xinge You, Xiubao Jiang, Weihua Ou, Yixiao Zhao, C. L. Philip Chen, Yuan Yan Tang |
ICONIP (2) | 2 |
| 2015 | Kernel normalized mixed-norm algorithm for system identificationabstractKernel methods provide an efficient nonparametric model to produce adaptive nonlinear filtering (ANF) algorithms. However, in practical applications, standard squared error based kernel methods suffer from two main issues: (1) a constant step size is used, which degrades the algorithm performance in non-stationary environment, and (2) additive noises are assumed to follow Gaussian distribution, while in practice the noises are generally non-Gaussian and follow other statistical distributions. To address these two issues simultaneously, this paper proposes a novel kernel normalized mixed-norm (KNMN) algorithm. Compared to the standard squared error based kernel methods, the KNMN algorithm extends the linear mixed-norm adaptive filtering algorithms to Reproducing Kernel Hilbert Space (RKHS) and introduces a normalized step size as well as adaptive mixing parameter. We also conduct the mean square convergence analysis and demonstrate the desirable performance of the KNMN algorithm in solving the system identification problem. Shujian Yu, Xinge You, Weihua Ou, Yuan Yan Tang |
IJCNN | 2 |
| 2015 | Dynamic Texture Synthesis via Image ReconstructionabstractThis paper addresses the problem of synthesizing continuous and infinitely varying stream of texture videos by doing operations on finite texture videos. Given an input texture video, such as flame, water, smoke, etc, we can synthesize a longer texture video holding the same texture appearance. Dynamic textures have been modeled as linear dynamic systems (LDS) by unfolding the video frames into column vectors and modeling their dynamic trajectory as time evolves. After the vectors are projected onto a lower dimensional space by Singular Value Decomposition (SVD), dynamic texture synthesis is achieved by driving the system with random noise. However, because of its over-simplified appearance model and under-constrained dynamic model. It is usually hard to synthesize long and visual pleasing texture video sequences. In this paper, we propose a new dynamic texture synthesis framework via creatively fitting the basic LDS with a newly developed patch reconstruction technique to efficiently enhance high quality texture details while maintaining the temporal coherence of the reconstructed texture patches. The patch reconstruction technique is inspired by locally linear embedding (LLE) and based on the assumption that small patches in the low-and high-quality images form manifolds with similar local geometry. The newly synthesized patches are finally stitched together by graph cuts to make up the output texture videos. Experiments on standard dynamic texture databases demonstrate that our method exhibits superior performance on synthesizing dynamic textures. Weigang Guo, Xinge You, Weiyong Xue, Shujian Yu, Xiubao Jiang |
SMC | 2 |
| 2015 | Human Heart Rate Estimation Using Ordinary Cameras under Natural MovementabstractNon-contact face-video based human heart rate (HR) estimation has attracted a lot of attentions in recent years. Almost all the state-of-the-art webcam or smartphone based HR estimation methods comprise three main steps: firstly, a region of interest (ROI) on the human face is detected in each video frame, then, the target signal is obtained by fusing multiple raw traces, which are extracted from the RGB channels across all the video frames, finally, HR is estimated by applying frequency analysis approach to the target signal. However, three major drawbacks impede the applicability of the current methods: (1) the performance of ROI detection is susceptible to head motion and facial expression, (2) there is still a lack of well-accepted method for fusing raw traces to form the target signal, and (3) the adopted frequency analysis approaches always provide estimation results with low resolution and high side lobes. To address these issues, we propose a novel HR estimation method which is applicable to ordinary cameras subject to natural head movement or facial expression. The proposed method features ROI detection via facial feature detection and tracking, target signal extraction via Independent Component Analysis (ICA) in the RGB channels, and HR estimation via real-valued iterative adaptive approach (RIAA). Experimental results validate the superiority of our proposed method. Shujian Yu, Xinge You, Xiubao Jiang, Yi Mou, Weihua Ou, Yuan Yan Tang, C. L. Philip Chen |
SMC | 2 |
| 2015 | A new weighted mean filter with a two-phase detector for removing impulse noise
Licheng Liu, C. L. Philip Chen, Yicong Zhou, Xinge You |
Inf. Sci. | 4 |
| 2015 | One global optimization method in network flow model for multiple object tracking
Zhenyu He 0001, Hongpeng Wang 0002, Xinge You, C. L. Philip Chen |
Knowl. Based Syst. | 4 |
| 2015 | An adaptive hybrid pattern for noise-robust texture analysis
Xinge You, C. L. Philip Chen, Dacheng Tao, Weihua Ou, Xiubao Jiang, Jixing Zou |
Pattern Recognit. | 2 |
| 2015 | Single object tracking via robust combination of particle filter and sparse representation
Shuangyan Yi, Zhenyu He 0001, Xinge You, Yiu-Ming Cheung |
Signal Process. | 3 |
| 2015 | Robust Nonnegative Patch Alignment for Dimensionality ReductionabstractDimensionality reduction is an important method to analyze high-dimensional data and has many applications in pattern recognition and computer vision. In this paper, we propose a robust nonnegative patch alignment for dimensionality reduction, which includes a reconstruction error term and a whole alignment term. We use correntropy-induced metric to measure the reconstruction error, in which the weight is learned adaptively for each entry. For the whole alignment, we propose locality-preserving robust nonnegative patch alignment (LP-RNA) and sparsity-preserviing robust nonnegative patch alignment (SP-RNA), which are unsupervised and supervised, respectively. In the LP-RNA, we propose a locally sparse graph to encode the local geometric structure of the manifold embedded in high-dimensional space. In particular, we select large p -nearest neighbors for each sample, then obtain the sparse representation with respect to these neighbors. The sparse representation is used to build a graph, which simultaneously enjoys locality, sparseness, and robustness. In the SP-RNA, we simultaneously use local geometric structure and discriminative information, in which the sparse reconstruction coefficient is used to characterize the local geometric structure and weighted distance is used to measure the separability of different classes. For the induced nonconvex objective function, we formulate it into a weighted nonnegative matrix factorization based on half-quadratic optimization. We propose a multiplicative update rule to solve this function and show that the objective function converges to a local optimum. Several experimental results on synthetic and real data sets demonstrate that the learned representation is more discriminative and robust than most existing dimensionality reduction methods. Xinge You, Weihua Ou, C. L. Philip Chen, Qiang Li 0024, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | A Noise-Robust Adaptive Hybrid Pattern for Texture ClassificationabstractIn this paper, we focus on developing a novel noise-robust LBP-based texture feature extraction scheme for texture classification. Specifically, two solutions have been proposed to overcome the primary two reasons that cause local binary pattern sensitive to noise. First, a hybrid model is proposed for noise-robust texture description. In this new model, the local primitive micro features are encoded with the texture's global spatial structure to reduce the noise sensitiveness. Second, we design an adaptive quantization algorithm, in which quantization thresholds are choosing adaptively on the basis of the texture's content. Higher noise-tolerance and discriminant power can be obtained in the quantization process. Based on the proposed hybrid texture description model and adaptive quantization algorithm, we develop an adaptive hybrid pattern scheme for noise-robust texture feature extraction. Compared with several state-of-the-art feature extraction schemes, our scheme leads to significant improvement in noisy texture classification. Xinge You, C. L. Philip Chen, Dacheng Tao, Xiubao Jiang, Fanyu You, Jixing Zou |
ICPR | 2 |
| 2014 | A novel joint tracker based on occlusion detection
Xin Li 0034, Zhenyu He 0001, Xinge You, C. L. Philip Chen |
Knowl. Based Syst. | 3 |
| 2014 | Arabic font recognition based on diacritics features
Mohammed Lutf, Xinge You, Yiu-Ming Cheung, C. L. Philip Chen |
Pattern Recognit. | 2 |
| 2014 | Robust face recognition via occlusion dictionary learning
Weihua Ou, Xinge You, Dacheng Tao, Pengyue Zhang, Yuan Yan Tang |
Pattern Recognit. | 2 |
| 2014 | Local Metric Learning for Exemplar-Based Object DetectionabstractObject detection has been widely studied in the computer vision community and it has many real applications, despite its variations, such as scale, pose, lighting, and background. Most classical object detection methods heavily rely on category-based training to handle intra-class variations. In contrast to classical methods that use a rigid category-based representation, exemplar-based methods try to model variations among positives by learning from specific positive samples. However, current existing exemplar-based methods either fail to use any training information or suffer from a significant performance drop when few exemplars are available. In this paper, we design a novel local metric learning approach to well handle exemplar-based object detection task. The main works are two-fold: 1) a novel local metric learning algorithm called exemplar metric learning (EML) is designed and 2) an exemplar-based object detection algorithm based on EML is implemented. We evaluate our method on two generic object detection data sets: UIUC-Car and UMass FDDB. Experiments show that compared with other exemplar-based methods, our approach can effectively enhance object detection performance when few exemplars are available. Xinge You, Qiang Li 0024, Dacheng Tao, Weihua Ou, Mingming Gong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Lip Segmentation under MAP-MRF Framework with Automatic Selection of Local Observation Scale and Number of SegmentsabstractThis paper addresses the problem of segmenting lip region from frontal human face image. Supposing each pixel of the target image has an optimal local scale from the segmentation viewpoint, we treat the lip segmentation problem as a combination of observation scale selection and observed data classification. Accordingly, we propose a hierarchical multiscale Markov random field (MRF) model to represent the membership map of each input pixel to a specific segment and local-scale map simultaneously. Subsequently, lip segmentation can be formulated as an optimal problem in the maximum a posteriori (MAP)-MRF framework. Then, we present a rival-penalized iterative algorithm to implement the segmentation, which is independent of the number of predefined segments. The proposed method mainly features two aspects: 1) its performance is independent of the predefined number of segments, and 2) it takes into account the local optimal observation scale for each pixel. Finally, we conduct the experiments on four benchmark databases, i.e. AR, CVL, GTAV, and VidTIMIT. Experimental results show that the proposed method is robust to the segment number that changes with a speaker's appearance, and can enhance the segmentation accuracy by taking advantage of the local optimal observation scale information. Yiu-Ming Cheung, Meng Li 0015, Xiaochun Cao, Xinge You |
IEEE Trans. Image Process. | 4 |
| 2014 | Group Sparse Multiview Patch Alignment Framework With View Consistency for Image ClassificationabstractNo single feature can satisfactorily characterize the semantic concepts of an image. Multiview learning aims to unify different kinds of features to produce a consensual and efficient representation. This paper redefines part optimization in the patch alignment framework (PAF) and develops a group sparse multiview patch alignment framework (GSM-PAF). The new part optimization considers not only the complementary properties of different views, but also view consistency. In particular, view consistency models the correlations between all possible combinations of any two kinds of view. In contrast to conventional dimensionality reduction algorithms that perform feature extraction and feature selection independently, GSM-PAF enjoys joint feature extraction and feature selection by exploiting l(2,1)-norm on the projection matrix to achieve row sparsity, which leads to the simultaneous selection of relevant features and learning transformation, and thus makes the algorithm more discriminative. Experiments on two real-world image data sets demonstrate the effectiveness of GSM-PAF for image classification. Jie Gui, Dacheng Tao, Zhenan Sun, Yong Luo 0002, Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 5 |
| 2014 | Diverse Expected Gradient Active Learning for Relative AttributesabstractThe use of relative attributes for semantic understanding of images and videos is a promising way to improve communication between humans and machines. However, it is extremely labor- and time-consuming to define multiple attributes for each instance in large amount of data. One option is to incorporate active learning, so that the informative samples can be actively discovered and then labeled. However, most existing active-learning methods select samples one at a time (serial mode), and may therefore lose efficiency when learning multiple attributes. In this paper, we propose a batch-mode active-learning method, called diverse expected gradient active learning. This method integrates an informativeness analysis and a diversity analysis to form a diverse batch of queries. Specifically, the informativeness analysis employs the expected pairwise gradient length as a measure of informativeness, while the diversity analysis forces a constraint on the proposed diverse gradient angle. Since simultaneous optimization of these two parts is intractable, we utilize a two-step procedure to obtain the diverse batch of queries. A heuristic method is also introduced to suppress imbalanced multiclass distributions. Empirical evaluations of three different databases demonstrate the effectiveness and efficiency of the proposed approach. Xinge You, Ruxin Wang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2013 | Detection, localization and pose classification of ear in 3D face profile imagesabstractWe present an efficient and robust system for landmark localization, segmentation and pose classification of ears from 3D profile facial range data. After defining 18 landmarks on the ear, including Triangular Fossa and Incisure Intertragica, a novel Ear Tree-structured Graph (ETG) is proposed to represent the 3D ear. We trained a flexible mixture model to locate these landmarks automatically. Afterwards, the ear region is outlined as the minimum rectangle including all landmarks. Finally, by calculating the turning angle between landmarks on the helix, the ear is classified as either a left or a right ear. To the best of our knowledge, there is no previous work on automatic landmark localization for 3D ear on 3D facial profile depth images. Experiments are conducted on University of Notre Dame Collection F and Collection J2 datasets, containing large occlusion, scale and pose variations. Results demonstrate the effectiveness of the proposed techniques. Jiajia Lei, Jindan Zhou, Mohamed Abdel-Mottaleb, Xinge You |
ICIP | 4 |
| 2013 | Learning a Sparse Representation for Robust Face Recognition
Weihua Ou, Xinge You, Pengyue Zhang, Xiubao Jiang, Duanquan Xu |
ICONIP (3) | 2 |
| 2013 | Generalization performance of magnitude-preserving semi-supervised ranking with graph-based regularization
Zhibin Pan, Xinge You, Hong Chen 0004, Dacheng Tao |
Inf. Sci. | 2 |
| 2012 | Structured sparse coding for image representation based on L1-graph
Weihua Ou, Xinge You, Yiu-Ming Cheung, Qinmu Peng, Mingming Gong, Xiubao Jiang |
ICPR | 2 |
| 2012 | A method using long digital straight segments for fingerprint recognition
Xiubao Jiang, Xinge You, Yuan Yuan 0001, Mingming Gong |
Neurocomputing | 2 |
| 2012 | Fingerprint Enhancement Based on Wavelet and Anisotropic FilteringabstractThe importance of high-fidelity enhancement in low quality fingerprint image cannot be overemphasized. Most of the existing fingerprint enhancement methods are contextual filter-based methods and they often suffer from two shortcomings: (1) there is block effect on the enhanced images; and (2) they blur or destroy ridge structures around singular points. In order to well preserve the ridge structures in singular regions and avoid block effect, we develop a new method for fingerprint enhancement combining nontensor product wavelet filter banks and anisotropic filter. We first decompose the fingerprint image using the nontensor product wavelet filter banks. Then we modify the approximation subimage using anisotropic filtering and adjust the high frequency coefficients of the three other subimages by applying the adaptive approach to reduce the noises according to the geometry feature of images. Finally, the inverse transform is applied to map the result and a final contrast enhancement is done subsequently. Experiments have been conducted on the fingerprint database FVC2004 in our study. The results demonstrate that the proposed approach is capable of overcoming block effect and enhancing low quality fingerprint while preserving the ridge structures around singular points. Jiajia Lei, Qinmu Peng, Xinge You, Hiyam Hatem Jabbar, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | A local region based approach to lip tracking
Yiu-Ming Cheung, Xin Liu 0011, Xinge You |
Pattern Recognit. | 3 |
| 2012 | Shape matching and classification using height functions
Xiang Bai, Xinge You, Wenyu Liu 0001, Longin Jan Latecki |
Pattern Recognit. Lett. | 3 |
| 2012 | Tracking and Pairing Vehicle Headlight in Night ScenesabstractTraffic surveillance is an important topic in computer vision and intelligent transportation systems and has intensively been studied in the past decades. However, most of the state-of-the-art methods concentrate on daytime traffic monitoring. In this paper, we propose a nighttime traffic surveillance system, which consists of headlight detection, headlight tracking and pairing, and camera calibration and vehicle speed estimation. First, a vehicle headlight is detected using a reflection intensity map and a reflection suppressed map based on the analysis of the light attenuation model. Second, the headlight is tracked and paired by utilizing a simple yet effective bidirectional reasoning algorithm. Finally, the trajectories of the vehicle's headlight are employed to calibrate the surveillance camera and estimate the vehicle's speed. Experimental results on typical sequences show that the proposed method can robustly detect, track, and pair the vehicle headlight in night scenes. Extensive quantitative evaluations and related comparisons demonstrate that the proposed method outperforms state-of-the-art methods. Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2011 | Segmentation of retinal blood vessels using the radial projection and semi-supervised approach
Xinge You, Qinmu Peng, Yuan Yuan 0001, Yiu-Ming Cheung, Jiajia Lei |
Pattern Recognit. | 1 |
| 2010 | Extracting corner-cue feature to improve minutiae-matching accuracyabstractThis paper proposes a new feature of fingerprint, called corner-cue. It is based on the curvature of fingerprint ridges. To extract the corner-cue, we first compute the curvature of fingerprint ridges and find the local maximum curvature points. Without regard to the high curvature points near minutiae, corner-cues are obtained. Corner-cues are further utilized in the matching stage to enhance the system's performance. Since high curvature points are important features of a fingerprint, the proposed method can obtain better results than conventional solely minutiae-based methods. Experimental results illustrate its effectiveness. Jiajia Lei, Xinge You, Wu Zeng |
ICIP | 2 |
| 2010 | Wavelet Domain Local Binary Pattern Features For Writer IdentificationabstractThe representation of writing styles is a crucial step of writer identification schemes. However, the large intra-writer variance makes it a challenging task. Thus, a good feature of writing style plays a key role in writer identification. In this paper, we present a simple and effective feature for off-line, text-independent writer identification, namely wavelet domain local binary patterns (WD-LBP). Based on WD-LBP, a writer identification algorithm is developed. WD-LBP is able to capture the essence of characteristics of writer while ignoring the variations intrinsic to every single writer. Unlike other texture framework method, we do not assign any statistical distribution assumption to the proposed method. This prevent us from making any, possibly erroneous, assumptions about the handwritten image feature distributions. The experimental results show that the proposed writer identification method achieves high accuracy of identification and outperforms recent writer identification method such as wavelet-GGD model and Gabor filtering method. Xinge You, Zhifan Gao, Yuan Yan Tang |
ICPR | 2 |
| 2010 | Offline Arabic Handwriting Identification Using Language DiacriticsabstractIn this paper, we present an approach for writer identification using off-line Arabic handwriting. The proposed method introduced Arabic writing in a new form, by presenting Arabic writing in its basic components instead of alphabetic. We split the input document into two parts: one for the letters and the other for the diacritics, we extract all diacritics from the input image and calculate the LBP histogram for each diacritic then concatenate these histograms to use it as handwriting features. We use the IFN/ENIT database in the experiments reported here and our tests involve 287 writers. The results show that our method is very effective and makes the handling of the Arabic handwriting more easily than before. Mohammed Lutf, Xinge You, Hong Li 0009 |
ICPR | 2 |
| 2010 | Retinal Blood Vessels Segmentation Using the Radial Projection and Supervised ClassificationabstractThe low-contrast and narrow blood vessels in retinal images are difficult to be extracted but useful in revealing certain systemic disease. Motivated by the goals of improving detection of such vessels, we propose the radial projection method to locate the vessel centerlines. Then the supervised classification is used for extracting the major structures of vessels. The final segmentation is obtained by the union of the two types of vessels after removal schemes. Our approach is tested on the STARE database, the results demonstrate that our algorithm can yield better segmentation. Qinmu Peng, Xinge You, Yiu-Ming Cheung |
ICPR | 2 |
| 2010 | Shape Classification Using Tree -UnionsabstractIn this paper, we proposed a novel approach to shape classification. A new shape tree based on junction nodes can represent the global structure in a simple way. The statistic distribution of junctions can be learned by merging the shape trees. In the process of learning, context of a junction node is obtained to improve the rate of classification. We illustrate the utility of the proposed method on the problem of 2D shape classification using the new shape tree representation. Bo Wang 0044, Wei Shen 0002, Wenyu Liu 0001, Xinge You, Xiang Bai |
ICPR | 4 |
| 2010 | Rotation invariant iris feature extraction using Gaussian Markov random fields with non-separable wavelet
Jing Huang 0018, Xinge You, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2010 | Image matching using enclosed region detector
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You |
J. Vis. Commun. Image Represent. | 4 |
| 2010 | A Blind Watermarking Scheme Using New Nontensor Product Wavelet Filter BanksabstractAs an effective method for copyright protection of digital products against illegal usage, watermarking in wavelet domain has recently received considerable attention due to the desirable multiresolution property of wavelet transform. In general, images can be represented with different resolutions by the wavelet decomposition, analogous to the human visual system (HVS). Usually, human eyes are insensitive to image singularities revealed by different high frequency subbands of wavelet decomposed images. Hence, adding watermarks into these singularities will improve the imperceptibility that is a desired property of a watermarking scheme. That is, the capability for revealing singularities of images plays a key role in designing wavelet-based watermarking algorithms. Unfortunately, the existing wavelets have a limited ability in revealing singularities in different directions. This motivates us to construct new wavelet filter banks that can reveal singularities in all directions. In this paper, we utilize special symmetric matrices to construct the new nontensor product wavelet filter banks, which can capture the singularities in all directions. Empirical studies will show their advantages of revealing singularities in comparison with the existing wavelets. Based upon these new wavelet filter banks, we, therefore, propose a modified significant difference watermarking algorithm. Experimental results show its promising results. Xinge You, Yiu-Ming Cheung, Qiuhui Chen |
IEEE Trans. Image Process. | 1 |
| 2009 | Writer Identification Using a Hybrid Method Combining Gabor Wavelet and Mesh Fractal Dimension
Zhenyu He 0007, Yiu-Ming Cheung, Xinge You |
IDEAL | 4 |
| 2009 | Editorial
Yiu-Ming Cheung, Yuping Wang 0003, Xinge You, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | Facial Biometrics Using Nontensor Product Wavelet and 2D Discriminant TechniquesabstractA new facial biometric scheme is proposed in this paper. Three steps are included. First, a new nontensor product bivariate wavelet is utilized to get different facial frequency components. Then a modified 2D linear discriminant technique (M2DLD) is applied on these frequency components to enhance the discrimination of the facial features. Finally, support vector machine (SVM) is adopted for classification. Compared with the traditional tensor product wavelet, the new nontensor product wavelet can detect more singular facial features in the high-frequency components. Earlier studies show that the high-frequency components are sensitive to facial expression variations and minor occlusions, while the low-frequency component is sensitive to illumination changes. Therefore, there are two advantages of using the new nontensor product wavelet compared with the traditional tensor product one. First, the low-frequency component is more robust to the expression variations and minor occlusions, which indicates that it is more efficient in facial feature representation. Second, the corresponding high-frequency components are more robust to the illumination changes, subsequently it is more powerful for classification as well. The application of the M2DLD on these wavelet frequency components enhances the discrimination of the facial features while reducing the feature vectors dimension a lot. The experimental results on the AR database and the PIE database verified the efficiency of the proposed method. Dan Zhang 0008, Xinge You, Patrick Shen-Pei Wang, Svetlana N. Yanushkevich, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Texture image retrieval based on non-tensor product wavelet filter banks
Zhenyu He 0007, Xinge You, Yuan Yuan 0001 |
Signal Process. | 2 |
| 2009 | A novel iris segmentation using radial-suppression edge detection
Jing Huang 0018, Xinge You, Yuan Yan Tang, Yuan Yuan 0001 |
Signal Process. | 2 |
| 2008 | Palmprint identification based on non-separable wavelet filter banksabstractCreases, as a special salient feature of palmprint, are large in number and distributed at all directions. It changes slowly in a personpsilas whole life, which qualifies themselves as features in palmprint identification. In this paper, we devised a new algorithm of crease extraction by using non-separable bivariate wavelet filter banks with linear phase. Compared with the traditional wavelet, our research demonstrates that the three high frequency sub-images generated by Non-separable Discrete Wavelet Transform (NDWT) can extract more creases and no longer extensively focus on the three special directions. As a consequence, we proposed a new method by combining NDWT and Support Vector Machines (SVM) for palmprint identification. Tested by our experiment, this method achieves a satisfied identification result and computational efficiency as well. Xinge You, Yuan Yan Tang, Yiu-Ming Cheung |
ICPR | 2 |
| 2008 | An Adaptive Image Watermarking Scheme Using Non-separable Wavelets and Support Vector Regression
Xinge You, Yiu-Ming Cheung |
IDEAL | 2 |
| 2008 | Iris recognition based on non-separable waveletabstractThis paper focuses on the rotation noise of iris recognition. Current iris recognition systems are unable to deal with the rotation noise perfectly. We propose a novel method for iris matching that decompose iris picture into wavelet subband coefficients via 16 non-separable wavelet filters, and use generalized Gaussian density (GGD) modeling of each non-separable orthogonal wavelet coefficients as a means of feature extraction, then compute the Kullback-Leibler distance (KLD) between GGDs and compare the iris code using the Kullback-Leibler distance. Experiments show that the proposed method is rotation invariance, it does not decrease their recognition rate, when the iris image is rotated. Jing Huang 0018, Xinge You, Yuan Yan Tang |
SMC | 2 |
| 2008 | Writer identification using global wavelet-based features
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Neurocomputing | 2 |
| 2008 | Writer identification of Chinese handwriting documents using hidden Markov tree model
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Pattern Recognit. | 2 |
| 2007 | Watermarking technique based on discrete non-separable wavelet filtersabstractThis paper presents a digital watermarking technique based on the Discrete Non-Separable Wavelet Transform (DNWT). In our paper, the discrete non-separable wavelet is constructed based on the standard dilation matrix 2I. Our investigation demonstrates that the constructed discrete nonseparable wavelet can detect more singularities of the host image with reflecting whole orientation while only three orientations being considered by traditional Discrete Separable Transform(DWT). For this reason, more coefficients in the highfrequency sub-bands by DNWT can add the watermark than that by DWT. By using the desirable character of DNWT, the imperceptibility and robustness requirements of watermarks are fulfilled. Experiment results show that the watermarking scheme based on DNWT is robust to some distortions such as noising, JPEG compression, and cropping. It also shows that the decomposing of the host image and the robustness of the watermark are relating to the parameters. Qingyan He, Xinge You, Limin Cui, Zaochao Bao |
SMC | 2 |
| 2007 | Wavelet-Based Approach to Character SkeletonabstractCharacter skeleton plays a significant role in character recognition. The strokes of a character may consist of two regions, i.e., singular and regular regions. The intersections and junctions of the strokes belong to singular region, while the straight and smooth parts of the strokes are categorized to regular region. Therefore, a skeletonization method requires two different processes to treat the skeletons in theses two different regions. All traditional skeletonization algorithms are based on the symmetry analysis technique. The major problems of these methods are as follows. 1) The computation of the primary skeleton in the regular region is indirect, so that its implementation is sophisticated and costly. 2) The extracted skeleton cannot be exactly located on the central line of the stroke. 3) The captured skeleton in the singular region may be distorted by artifacts and branches. To overcome these problems, a novel scheme of extracting the skeleton of character based on wavelet transform is presented in this paper. This scheme consists of two main steps, namely: a) extraction of primary skeleton in the regular region and b) amendment processing of the primary skeletons and connection of them in the singular region. A direct technique is used in the first step, where a new wavelet-based symmetry analysis is developed for finding the central line of the stroke directly. A novel method called smooth interpolation is designed in the second step, where a smooth operation is applied to the primary skeleton, and, thereafter, the interpolation compensation technique is proposed to link the primary skeleton, so that the skeleton in the singular region can be produced. Experiments are conducted and positive results are achieved, which show that the proposed skeletonization scheme is applicable to not only binary image but also gray-level image, and the skeleton is robust against noise and affine transform. Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2006 | Handwriting-based personal identificationabstractHandwriting-based personal identification, which is also called handwriting-based writer identification, is an active research topic in pattern recognition. Despite continuous effort, offline handwriting-based writer identification still remains as a challenging problem because writing features can only be extracted from the handwriting image. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is unavailable for offline writer identification. In this paper, we present a novel wavelet-based Generalized Gaussian Density (GGD) method for offline writer identification. Compared with the 2-D Gabor model, which is currently widely acknowledged as a good method for offline handwriting identification, GGD method not only achieves a better identification accuracy but also greatly reduces the elapsed time on calculation in our experiments. Zhenyu He 0001, Xinge You, Yuan Yan Tang, Bin Fang 0001, Jianwei Du |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2006 | Thinning Character Using Modulus Minima of Wavelet TransformabstractAn essential step in character recognition is to extract the skeleton characteristics of the character. In this paper, an efficient algorithm is proposed to extract visually satisfactory skeleton from printed and handwritten characters, which overcomes fundamental shortcomings of our previous skeletonization technique based on the maximum modulus symmetry of wavelet transform (WT). The proposed method is motivated from some desirable properties of the WT with constructed wavelet functions: namely, the local modulus minima of the WT are scale-independent at different level scales and are located at the medial axis of the symmetrical contours of character stroke. Thus the modulus minima of the WT are computed as the intrinsic skeletons of character strokes. To achieve faster implementation, a multiscale processing technique is employed. Thus major structures of the skeleton are extracted using the coarse scale, while fine structures are extracted using the fine scale. We have tested the algorithm on handwritten and printed character images. Experimental results show that the proposed algorithm is applicable to not only binary image but also gray-level image where it can be impractical to use other skeletonization techniques, such as thinning and distance transforms. Further, it can effectively remove unwanted artifacts and branches from the extracted skeletons at the intersections and junctions of character strokes and is robust against noises while most existing methods perform poorly. Xinge You, Qiuhui Chen, Bin Fang 0001, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2005 | A Novel Method for Off-line Handwriting-based Writer IdentificationabstractHandwriting-based writer identification is a hot research topic in the pattern recognition field. Nowadays, online handwriting-based writer identification is steadily growing toward its maturity. On the contrary, offline handwriting-based writer identification still remains as a challenging problem because writing features only can be extracted from the handwriting image in this situation. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is lost. At present, 2D Gabor filter method is widely acknowledged as a good method for offline handwriting identification, however it still suffers from some inherent disadvantages, such as the high computational cost. In this paper, we present a novel wavelet-based GGD method to replace the traditional 2D Gabor filters. Shown in our experiments, this novel method not only achieves better experiment results but also greatly reduces the elapsed time on calculation. Zhenyu He 0001, Yuan Yan Tang, Bin Fang 0001, Jianwei Du, Xinge You |
ICDAR | 5 |
| 2005 | Similarity Measurement for Off-Line Signature Verification
Xinge You, Bin Fang 0001, Zhenyu He 0001, Yuan Yan Tang |
ICIC (1) | 1 |
| 2005 | Locating Vessel Centerlines in Retinal Images Using Wavelet Transform: A Multilevel Approach
Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001, Jian Huang 0009 |
ICIC (1) | 1 |
| 2005 | A contourlet-based method for writer identificationabstractHandwriting-based writer identification is a hot research topic in the field of pattern recognition. Typically, there are four modes of writer identification: on-line text-dependent, on-line text-independent, off-line text-dependent, off-line text-independent; and off-line text-independent is the most challenging problem among them because many valuable writing features are not available in this case, such as shape features, dynastic writing information and etc. In this paper, we focus on the text-independent writer identification based on off-line Chinese handwriting and present a new contourlet-based GGD (Generalized Gaussian Density) method. This novel method achieves a good experiment result in our experiments. Zhenyu He 0001, Yuan Yan Tang, Xinge You |
SMC | 3 |
| 2005 | Morphological structure reconstruction of retinal vessels in fundus imagesabstractVessels in retinal fundus images are useful in revealing the severity of eye-related diseases. In addition, they can act as landmarks for localizing lesions or the central vision area, and guide laser treatment of neovascularization. In this paper, we propose a two-stage scheme to extract vessels and reconstruct the morphological structure of vessels in retinal images. First, we employ mathematical morphology techniques to highlight large and small vessels with respect to their spatial properties. Different curvature response between vessel and noise patterns allows the use of curvature evaluation to remove enhanced vessel-like noise. A set of linear filters finalize the vessel map. However, the resulting vascular structure is incomplete of some important features in bifurcation points and central reflex. In order to rectify the pitfall, a reconstruction process is performed using dynamic local region growth to recover the morphological structure of vessels. Average performance of our method to extract vessels is 83.7% of TPR(True positive rate) and 3.8% of FPR(False positive rate) for 35 retinal images which include 21 abnormal images. Bin Fang 0001, Xinge You, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | A wavelet-based approach to ridge thinning in fingerprint imagesabstractAs a global feature of fingerprints, the thinning of ridges, extraction of minutiae and computation of orientation field are very important for automatic fingerprint recognition. Many algorithms have been proposed for their computation and estimation, but their results are unsatisfactory, especially for poor quality fingerprint images. In this paper, a robust wavelet-based method to create thinned ridge map of fingerprint for automatic recognition is proposed. Properties of modulus minima based on the spline wavelet function are substantially investigated. Desirable characteristics show that this method is suitable to describe the skeleton of the ridge of the fingerprint image. A multi-scale thinning algorithm based on the modulus minima of wavelet transform is presented. The proposed algorithm is able to improve the skeleton representation of the ridge of the fingerprint without side-effects and limitations of the existing methods. The thinned ridge map can facilitate the extraction of the minutiae for matching in fingerprint recognition. Experiments have been conducted to validate the effectiveness and efficiency of the proposed method. Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | Fingerprint Enhancement Using Wavelet Transform Combined With Gabor FilterabstractThe performance of automatic fingerprint identification system (AFIS) is heavily determined by the quality of the input image, thus an effective method to enhance the fingerprint image is essential in such a system. In this paper, we combine the filter-based method, which is mostly used nowadays with wavelet transform to achieve a more reliable and effective approach to fingerprint enhancement. This novel approach consists of five main steps, namely: (1) normalization, (2) decomposition, (3) wavelet coefficient adjustment, (4) Gabor filtering, and (5) reconstruction. Using this new method, a more clear fingerprint image can be obtained, which can distinctly improve the accuracy of the minutiae extraction module and finally achieve a better performance of the entire system. Experiments have been conducted in our study and positive experimental results have been received, which show that the proposed combined method is more effective and robust than other existing methods such as the filter-based and direct gray-level approaches. Yuan Yan Tang, Xinge You |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2003 | Skeletonization of Character Based on Wavelet Transform
Xinge You, Yuan Yan Tang |
CAIP | 1 |
| 2003 | Skeletonization of Ribbon-Like Shapes Based on a New Wavelet FunctionabstractA wavelet-based scheme to extract skeleton of Ribbon-like shape is proposed in this paper, where a novel wavelet function plays a key role in this scheme, which possesses three significant characteristics, namely, 1) the position of the local maximum moduli of the wavelet transform with respect to the Ribbon-like shape is independent of the gray-levels of the image. 2) When the appropriate scale of the wavelet transform is selected, the local maximum moduli of the wavelet transform of the Ribbon-like shape produce two new parallel contours, which are located symmetrically at two sides of the original one and have the same topological and geometric properties as that of the original shape. 3) The distance between these two parallel contours equals to the scale of the wavelet transform, which is independent of the width of the shape. This new scheme consists of two phases: 1) Generation of wavelet skeleton-based on the desirable properties of the new wavelet function, symmetry analyses of the maximum moduli of the wavelet transform is described. Midpoints of all pairs of contour elements can be connected to generate a skeleton of the shape, which is defined as wavelet skeleton. 2) Modification of the wavelet skeleton. Thereafter, a set of techniques are utilized for modifying the artifacts of the primary wavelet skeleton. The corresponding algorithm is also developed in this paper. Experimental results show that the proposed scheme is capable of extracting exactly the skeleton of the Ribbon-like shape with different width as well as different gray-levels. The skeleton representation is robust against noise and affine transformation. Yuan Yan Tang, Xinge You |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |