Lei Gao 0001

dblp:44/2139-1 · DBLP profile ↗
← Back
34ranked-venue papers
22as first author
16since 2021 · last 2025
0000-0001-5583-713XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 18 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 An Efficient Optimization Criterion for Multi-View Feature Representation Learning
abstract
The training of contemporary machine learning (ML) models, particularly deep neural networks (DNNs), often relies on enormous data sources to properly tune model parameters. As a result, achieving competitive results with limited training data and computational resources has been recognized as a significant bottleneck to advance ML. To address these issues, multi-view representation learning has emerged. However, how to efficiently build multi-view learning models remains a big challenge. In this paper, a novel optimization criterion is proposed to tackle this challenge. Specifically, the proposed criterion ensures speedy and effective parameter selection, reducing the effort to reach optimal design of the model while maintaining performance. To validate the efficiency and generalizability of the presented solution, experiments were conducted on face recognition and few-shot learning for image classification using four databases of different scales. Experimental results demonstrate the superiority of the proposed approach, offering an efficient yet robust solution to the data-scarcity challenge.
Lei Gao 0001, Kai Liu 0032, Kevin Tang, Ling Guan
ISM1
2025 EntroFormer: An entropy-based sparse vision transformer for real-time semantic segmentation
Song Wang 0008, Lin Wu 0001, Deyin Liu, Lei Gao 0001, Lin Qi 0001, Guanghui Wang 0001
Comput. Vis. Image Underst.5
2025 A discriminative multi-modal adaptation neural network model for video action recognition
Lei Gao 0001, Kai Liu 0032, Ling Guan
Neural Networks1
2025 ODMTCNet: An Interpretable Multiview Deep Neural Network Architecture for Feature Representation
abstract
Recently, deep cascade architecture-based algorithms have attracted wide attention and have been applied to numerous application domains successfully. Nevertheless, the black-box structure of such algorithms has always been considered the Achilles' heel by the machine learning community. Moreover, due to its data-driven nature, the deep cascade architecture likely causes over-fitting problems when there is no sufficient data available. In order to solve these pressing issues, this work proposes a novel multiview deep neural network (DNN) model, namely, optimal discriminant multiview tensor convolutional network (ODMTCNet), which integrates statistics-guided optimization (SGO) principles with the DNN architecture. Specifically, a discriminant multiview tensor convolution strategy is proposed and integrated with a deep cascade architecture. Different from the traditional DNN models, the parameters of the convolutional layers in ODMTCNet are determined by solving SGO problems. Based on the SGO principles, the relation between the optimal performance and parameters (e.g., the number of convolutional filters) can be analytically predicted, with each layer generating justified knowledge representations. In addition, information quality (IQ) is adopted to further improve multiview feature representation. Because of its unique design, ODMTCNet is able to handle different types of features (e.g., raw, hand-crafted, prior knowledge-based, and DNN-generated features), forming a general platform for multiview feature representation. To validate the genericness and effectiveness of the ODMTCNet model, we conducted experiments on five datasets of different scales: The Olivetti Research Lab (ORL) database, the Facial Recognition Technology (FERET) database, the ETH-80 database, the Caltech 256 database, and the nanyang technological university (NTU) red green blue-depth (RGB+D) 120 database. Experimental results show the superiority of the presented solution over the state-of-the-art. Implementation codes will be made available in the final version.
Lei Gao 0001, Ling Guan
IEEE Trans. Neural Networks Learn. Syst.1
2025 Mathematics-Inspired Models: A Green and Interpretable Learning Paradigm for Multimedia Computing
abstract
The advances of machine learning (ML), and AI in general, have attracted unprecedented attention in intelligent multimedia computing and many other fields. However, due to the concern for sustainability and black-box nature of ML models, especially deep neural networks (DNNs), green and interpretable learnings have been extensively studied in recent years, despite suspicions on effectiveness, subjectivity of interpretability, and complexity. To address these concerns and suspicions, this article starts with a survey on recent discoveries in green learning and interpretable learning and then presents mathematics-inspired (M-I) learning models. We will demonstrate that the M-I models are green in nature with numerous interpretable properties. Finally, we present several examples in multi-view information computing on both static image-based and dynamic video-based tasks to demonstrate that the M-I methodology promises a plausible and sustainable path for natural evolution of ML, which is worth further investment in.
Lei Gao 0001, Kai Liu 0032, Ling Guan
ACM Trans. Multim. Comput. Commun. Appl.1
2024 An Optimal Edge-weighted Graph Semantic Correlation Framework for Multi-view Feature Representation Learning
abstract
In this article, we present an optimal edge-weighted graph semantic correlation (EWGSC) framework for multi-view feature representation learning. Different from most existing multi-view representation methods, local structural information and global correlation in multi-view feature spaces are exploited jointly in the EWGSC framework, leading to a new and high-quality multi-view feature representation. Specifically, a novel edge-weighted graph model is first conceptualized and developed to preserve local structural information in each of the multi-view feature spaces. Then, the explored structural information is integrated with a semantic correlation algorithm, labeled multiple canonical correlation analysis (LMCCA), to form a powerful platform for effectively exploiting local and global relations across multi-view feature spaces jointly. We then theoretically verified the relation between the upper limit on the number of projected dimensions and the optimal solution to the multi-view feature representation problem. To validate the effectiveness and generality of the proposed framework, we conducted experiments on five datasets of different scales, including visual-based (University of California Irvine (UCI) iris database, Olivetti Research Lab (ORL) face database, and Caltech 256 database), text-image-based (Wiki database), and video-based (Ryerson Multimedia Lab (RML) audio-visual emotion database) examples. The experimental results show the superiority of the proposed framework on multi-view feature representation over state-of-the-art algorithms.
Lei Gao 0001, Ling Guan
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Towards Efficient Multi-view Representation Learning
abstract
The proliferation of deep neural networks (DNNs) has drawn unprecedented interest in the study of various contents such as image, audio, video, to name a few. However, due to the data-driven nature, the high computational requirement and slow running time are considered as Achilles’ heels of DNN-based algorithms, limiting the progress of DNNs in time-sensitive applications. Recently, distinct discriminant canonical correlation analysis network (DDCCANet), a multi-view neural network, has shown great generalizability across multiple application domains, both analytically and experimentally. However, Although the computational requirement and running time by DDCCANet are more manageable than DNN-based algorithms, they can be substantially further improved. This paper proposes two new algorithms for multi-view feature representation learning, namely incremental DDCCANet (IDDCCANet) with substantial save in computational memory and GPU-accelerated DDCCANet (GADDCCANet) with drastically accelerated running time, forming a practically significant platform for multi-view feature representation learning. To validate the power of the proposed algorithms, experiments are conducted on several data sets with different types of inputs (e.g., raw image pixels, classical and DNN-based features). Experimental results clearly show that the proposed algorithms provide promising solutions to address the two longstanding challenges.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan
ISM3
2023 A Discriminant Information Theoretic Learning Framework for Multi-modal Feature Representation
abstract
As sensory and computing technology advances, multi-modal features have been playing a central role in ubiquitously representing patterns and phenomena for effective information analysis and recognition. As a result, multi-modal feature representation is becoming a progressively significant direction of academic research and real applications. Nevertheless, numerous challenges remain ahead, especially in the joint utilization of discriminatory representations and complementary representations from multi-modal features. In this article, a discriminant information theoretic learning (DITL) framework is proposed to address these challenges. By employing this proposed framework, the discrimination and complementation within the given multi-modal features are exploited jointly, resulting in a high-quality feature representation. According to characteristics of the DITL framework, the newly generated feature representation is further optimized, leading to lower computational complexity and improved system performance. To demonstrate the effectiveness and generality of DITL, we conducted experiments on several recognition examples, including both static cases, such as handwritten digit recognition, face recognition, and object recognition, and dynamic cases, such as video-based human emotion recognition and action recognition. The results show that the proposed framework outperforms state-of-the-art algorithms.
Lei Gao 0001, Ling Guan
ACM Trans. Intell. Syst. Technol.1
2022 Interpretable Artificial Intelligence through Locality Guided Neural Networks
Randy Tan, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan
Neural Networks2
2022 A Discriminative Vectorial Framework for Multi-Modal Feature Representation
abstract
Due to the rapid advancements of sensory and computing technology, multi-modal data sources that represent the same pattern or phenomenon have attracted growing attention. As a result, finding means to explore useful information from these multi-modal data sources has quickly become a necessity. In this paper, a discriminative vectorial framework is proposed for multi-modal feature representation in knowledge discovery by employing multi-modal hashing (MH) and discriminative correlation maximization (DCM) analysis. Specifically, the proposed framework is capable of minimizing the semantic similarity among different modalities by MH and exacting intrinsic discriminative representations across multiple data sources by DCM analysis jointly, enabling a novel vectorial framework of multi-modal feature representation. Moreover, the proposed feature representation strategy is analyzed and further optimized based on canonical and non-canonical cases, respectively. Consequently, the generated feature representation leads to effective utilization of the input data sources of high quality, producing improved, sometimes quite impressive, results in various applications. The effectiveness and generality of the proposed framework are demonstrated by utilizing classical features and deep neural network (DNN) based features with applications to image and multimedia analysis and recognition tasks, including data visualization, face recognition, object recognition; cross-modal (text-image) recognition and audio emotion recognition. Experimental results show that the proposed solutions are superior to state-of-the-art statistical machine learning (SML) and DNN algorithms.
Lei Gao 0001, Ling Guan
IEEE Trans. Multim.1
2021 2D-FRFT Based Frequency Shift-Invariant Digital Image Encryption
abstract
In this paper, we study the property of frequency shift in two-dimensional Fractional Fourier Transform (2D-FRFT) domain. Based on the mathematical verification and computer simulations, it is demonstrated that the magnitude of reconstruction from phase information satisfies frequency shift-invariant in 2D-FRFT domain, improving the robustness of image encryption. Experiments are implemented to verify the effectiveness of this property against the frequency shift attack in 2D-FRFT domain, improving the robustness of image encryption.
Lei Gao 0001, Lin Qi 0001, Ling Guan
ICASSP1
2021 A two-stream heterogeneous network for action recognition based on skeleton and RGB modalities
abstract
Recent years, skeleton based action recognition with graph convolutional network (GCN) has achieved great success. However, since skeleton data only includes human body joints coordinates, other key information on actions is missing such as the subtle motion of hands, the objects the human is interacting, leading to an unsatisfactory performance. In this respect, the RGB data can offer help to recognize actions that skeleton-based methods have limitations on. In this work, we propose a novel two-stream heterogeneous network consisting of GCN and CNN networks for action recognition. Specifically, the GCN network takes the skeletal sequence as input to exploit skeleton information. For the RGB video, the CNN model, ResNet (2+1)D, is adapted to exploit RGB information. Afterwards, the discriminant canonical correlation analysis (DCCA) method is utilized to integrate the output feature maps from the skeleton and RGB streams, resulting in improved performance. Experimental results on the large-scale dataset NTU RGB+D show that the proposed model outperforms state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISM2
2021 Integrating vertex and edge features with Graph Convolutional Networks for skeleton-based action recognition
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
Neurocomputing2
2021 A discriminant kernel entropy-based framework for feature representation learning
Lei Gao 0001, Lin Qi 0001, Ling Guan
J. Vis. Commun. Image Represent.1
2021 The Property of Frequency Shift in 2D-FRFT Domain With Application to Image Encryption
abstract
The Fractional Fourier Transform (FRFT) has been playing a unique and increasingly important role in signal and image processing. In this letter, we investigate the property of frequency shift in two-dimensional FRFT (2D-FRFT) domain. It is shown that the magnitude of image reconstruction from phase information is frequency shift-invariant in 2D-FRFT domain, enhancing the robustness of image encryption, an important multimedia security task. Experiments are conducted to demonstrate the effectiveness of this property against the frequency shift attack, improving the robustness of image encryption.
Lei Gao 0001, Lin Qi 0001, Ling Guan
IEEE Signal Process. Lett.1
2021 A Multi-Stream Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action Recognition
abstract
Recently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN with HCRF to retain the human skeleton structure information even during the classification stage. Our model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework which takes the relative coordinate of the joints and bone direction as two static feature streams, and the temporal displacements between two consecutive frames as the dynamic feature stream. Experimental results on three challenging benchmarks (NTU RGB+D, N-UCLA, SYSU) show the superior performance of the proposed model over state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
IEEE Trans. Multim.2
2020 A Distinct Discriminant Canonical Correlation Analysis Network based Deep Information Quality Representation for Image Classification
abstract
In this paper, we present a distinct discriminant canonical correlation analysis network (DDCCANet) based deep information quality representation with application to image classification. Specifically, to explore the sufficient discriminant information between different data sets, the within-class and between-class correlation matrices are employed and optimized jointly. Moreover, different from the existing canonical correlation analysis network (CCANet) and related algorithms, an information theoretic descriptor, information quality (IQ), is adopted to generate the deep-level feature representation for image classification. Benefiting from the explored discriminant information and IQ descriptor, it is potential to gain a more effective deep-level representation from multi-view data sets, leading to improved performance in classification tasks. To demonstrate the effectiveness of the proposed DDCCANet, we conduct experiments on the Olivetti Research Lab (ORL) face database, ETH80 database and CIFAR10 database. Experimental results show the superiority of the proposed solution on image classification.
Lei Gao 0001, Ling Guan
ICPR1
2020 A Vertex-Edge Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
The Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability of exploiting the joint information from the graph structure of the skeleton data. Recently, as a strong and complementary modality for action recognition, the bone information from skeleton data has attracted more attention. However, most existing GCN-based methods extract the bone and joint features with two separate GCN networks, ignoring the dependencies between joints and bones. In this work, a Vertex-Edge Graph Convolutional Network(VE-GCN) is proposed to reveal the information across joints, bones and their relationships simultaneously. In addition, we learn the additional connections among joints and bones for various action samples besides the natural connections of the skeleton. Then we conduct the convolution operation on joints and their neighbors based on these additional connections. Moreover, the conditional random field (CRF) is utilized as the loss function to achieve improved performance. Experimental results on two large-scale datasets NTU RGB+D and NTU RGB+D 120 show that the proposed model outperforms state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISCAS2
2020 A Complete Discriminative Tensor Representation Learning for Two-Dimensional Correlation Analysis
abstract
As an effective tool for two-dimensional data analysis, two-dimensional canonical correlation analysis (2DCCA) is not only capable of preserving the intrinsic structural information of original two-dimensional (2D) data, but also reduces the computational complexity effectively. However, due to the unsupervised nature, 2DCCA is incapable of extracting sufficient discriminatory representations, resulting in an unsatisfying performance. In this letter, we propose a complete discriminative tensor representation learning (CDTRL) method based on linear correlation analysis for analyzing 2D signals (e.g. images). This letter shows that the introduction of the complete discriminatory tensor representation strategy provides an effective vehicle for revealing, and extracting the discriminant representations across the 2D data sets, leading to improved results. Experimental results show that the proposed CDTRL outperforms state-of-the-art methods on the evaluated data sets.
Lei Gao 0001, Ling Guan
IEEE Signal Process. Lett.1
2019 Information Fusion via Multimodal Hashing With Discriminant Correlation Maximization
abstract
Due to low storage cost and fast query speed, hashing has been applied to similarity search in multimedia data widely. In this paper, an effective information fusion algorithm using multimodal hashing with discriminant correlation maximization is presented. The proposed algorithm not only finds the minimum of the semantic similarity across different modalities by multimodal hashing, but also minimizes the between-class correlation and maximizes the within-class correlation simultaneously to extract discriminant representations for information fusion. More importantly, two solutions with canonical case and non-canonical case are presented, and a novel solution to non-canonical case is proposed. Benefiting from the combination of semantic similarity across different modalities from multimodal hashing information and the discriminant representation strategy, the proposed strategy can achieve improved performance. Experimental results show that the proposed approach outperforms the related methods.
Lei Gao 0001, Ling Guan
ICIP1
2019 Information Fusion via Deep Cross-Modal Factor Analysis
abstract
In this paper, we introduce Deep Cross-Modal Factor Analysis (DCFA) to identify complex nonlinear transformations of two variables for information fusion. DCFA is able to represent the coupled patterns between two different sets of variables by minimizing the Frobenius norm distance in the transformed domain. Unlike previous kernel methods, the feature mapping of DCFA is achieved with deep networks (DN) instead of the traditional kernel method. Therefore, the representation of DCFA method is not limited by the fixed kernel. Moreover, DCFA can be considered as a nonlinear extension of the linear Cross-Modal Factor Analysis (CFA), and an alternative to the nonparametric method Kernel Cross-Modal Factor Analysis (KCFA) and the recently proposed Deep Canonical Correlation Analysis (Deep CCA) method. The performance of DCFA is evaluated on MNIST handwritten digit dataset and two audio emotion datasets. Experimental results show that the proposed solution outperforms the methods of KCCA, KCFA, Deep CCA and the deep learning based method-Alexnet, in terms of accuracy.
Lei Gao 0001, Ling Guan
ISCAS1
2019 Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action Recognition
abstract
Recently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN and HCRF to retain the human skeleton structure information during the classification stage. The proposed model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework that takes the relative coordinates of the joints and bone direction as two static feature streams and the temporal displacements as the dynamic feature stream. Experimental results on two challenging benchmarks (NTU RGB+D, N-UCLA) show the superior performance of the proposed model over state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISM2
2019 The Labeled Multiple Canonical Correlation Analysis for Information Fusion
abstract
The objective of multimodal information fusion is to mathematically analyze information carried in different sources and create a new representation that will be more effectively utilized in pattern recognition and other multimedia information processing tasks. In this paper, we introduce a new method for multimodal information fusion and representation based on the Labeled Multiple Canonical Correlation Analysis (LMCCA). By incorporating class label information of the training samples, the proposed LMCCA ensures that the fused features carry discriminative characteristics of the multimodal information representations and are capable of providing superior recognition performance. We implement a prototype of LMCCA to demonstrate its effectiveness on handwritten digit recognition, face recognition, and object recognition utilizing multiple features, bimodal human emotion recognition involving information from both audio and visual domains. The generic nature of LMCCA allows it to take as input features extracted by any means, including those by deep learning (DL) methods. Experimental results show that the proposed method enhanced the performance of both statistical machine learning methods, and methods based on DL.
Lei Gao 0001, Rui Zhang 0010, Lin Qi 0001, Enqing Chen, Ling Guan
IEEE Trans. Multim.1
2018 Discriminative Robust Gaze Estimation Using Kernel-DMCCA Fusion
abstract
The proposed framework employs discriminative analysis for gaze estimation using kernel discriminative multiple canonical correlation analysis (K-DMCCA), which represents different feature vectors that account for variations of head pose, illumination and occlusion. The feature extraction component of the framework includes spatial indexing, statistical and geometrical elements. Gaze estimation is constructed by feature aggregation and transforming features into a higher dimensional space using the RBF kernel γ and spread factor. The output of fused features through K-DMCCA is robust to illumination, occlusion and is calibration free. Our algorithm is validated on MPII, CAVE, ACS and EYEDIAP datasets. The two main contributions of the framework are the following: Enhancing the performance of DMCCA with the kernel and introducing quadtree as an iris region descriptor. Spatial indexing using quadtree is a robust method for detecting which quadrant the iris is situated, detecting the iris boundary and it is inclusive of statistical and geometrical indexing that are calibration free. Our method achieved an accurate gaze estimation of 4.8° using Cave, 4.6° using MPII, 5.1° using ACS and 5.9° using EYEDIAP datasets respectively. The proposed framework provides insight into the methodology of multi-feature fusion for gaze estimation.
Salah Rabba, Matthew J. Kyan, Lei Gao 0001, Azhar Quddus, Ali Shahidi Zandi, Ling Guan
ISM3
2018 Deep Reinforcement Learning with Parameterized Action Space for Object Detection
abstract
Object detection is a fundamental task in computer vision. With the remarkable progress made in big visual data analytics and deep learning, Reinforcement Learning (RL) is becoming a promising framework to model the object detection problem since the detection procedure can be cast as a Markov decision process (MDP). We propose a Reinforcement Learning system with parameterized action space for image object detection. The proposed system uses an active agent exploring in a scene to identify the location of a target object, and learns a policy to refine the geometry of the agent by taking simple actions in parameterized space, which integrates the discrete actions and its corresponding continuous parameters. We then optimize the representation of the generated region proposals with the discriminative multiple canonical correlation analysis (DMCCA) [11] in preparation for classification with Fast R-CNN. Experiments on PASCAL VOC 2007 and 2012 datasets show the effectiveness of the proposed method.
Naimul Mefraz Khan, Lei Gao 0001, Ling Guan
ISM3
2018 Discriminative Multiple Canonical Correlation Analysis for Information Fusion
abstract
In this paper, we propose the discriminative multiple canonical correlation analysis (DMCCA) for multimodal information analysis and fusion. DMCCA is capable of extracting more discriminative characteristics from multimodal information representations. Specifically, it finds the projected directions, which simultaneously maximize the within-class correlation and minimize the between-class correlation, leading to better utilization of the multimodal information. In the process, we analytically demonstrate that the optimally projected dimension by DMCCA can be quite accurately predicted, leading to both superior performance and substantial reduction in computational cost. We further verify that canonical correlation analysis (CCA), multiple canonical correlation analysis (MCCA) and discriminative canonical correlation analysis (DCCA) are special cases of DMCCA, thus establishing a unified framework for canonical correlation analysis. We implement a prototype of DMCCA to demonstrate its performance in handwritten digit recognition and human emotion recognition. Extensive experiments show that DMCCA outperforms the traditional methods of serial fusion, CCA, MCCA, and DCCA.
Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan
IEEE Trans. Image Process.1
2016 Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognition
abstract
As an information fusion tool, Kernel Entropy Component Analysis (KECA) is realized by using descriptor of information entropy and optimized by entropy estimation. However, as an unsuper-vised method, it merely puts the information or features from different channels together without considering their intrinsic structures and relations. In this paper, we introduce an enhanced version of KECA for information fusion, KECA in Discriminative Canonical Correlation Space (DCCS). Not only the intrinsic structures and discriminative representations are considered, but also the natural representations of input data are revealed by entropy estimation, leading to improved recognition accuracy. The effectiveness of the proposed solution is evaluated through experiments on two audio emotion databases. Experimental results show that the proposed solution outperforms the existing methods based on similar principles.
Lei Gao 0001, Lin Qi 0001, Ling Guan
ICASSP1
2016 A Novel Discriminative Framework Integrating Kernel Entropy Component Analysis and Discriminative Multiple Canonical Correlation for Information Fusion
abstract
The effective interpretation and integration of multiple information content are important for the efficacious utilisation of multimedia in a wide variety of application context. The major challenge in information fusion lies in the difficulty of identifying the complementary and discriminatory representations from individual channels or data sources. In this paper, we propose a novel framework integrating kernel entropy-estimation and discriminative multiple canonical correlation (DMCC) to address this challenge. Not only the distribution and complementary representations of input data are revealed by entropy estimation, but also the discriminative representations are considered by DMCC, achieving improved recognition accuracy. The effectiveness of the proposed method is demonstrated on two audio emotion databases. Experimental results show that it outperforms the existing methods based on similar principles.
Lei Gao 0001, Ling Guan, Lin Qi 0001, Enqing Chen
ISM1
2016 Posture Selection Based on Two-Layer AP with Application to Human Action Recognition Using HMM
abstract
In this paper, we propose a posture selection method based on two-layer Affinity propagation (AP) for human action recognition using Hidden markov models (HMMs). A two-layer AP is used as the clustering algorithm instead of K-means in order to avoid the problem of random initialization. After two-layer AP, each cluster center of the skeleton features represents the pose of this activity. Each frame sequence of an action can be labeled by a sequence of these poses, meanwhile, the initial parameters of HMM can be calculated from these sequences. The effectiveness of the proposed method is implemented through MSR Action3D and UTKinect databases.
Mengyan Yuan, Enqing Chen, Lei Gao 0001
ISM3
2015 Sparsity preserving multiple canonical correlation analysis with visual emotion recognition to multi-feature fusion
abstract
Sparsity preserving projections (SPP) aim to preserve the sparse reconstructive relationship among the data and have been successfully applied to face recognition. The projections are invariant to rotations, rescalings, and translations of the data, and more importantly, they contain natural discriminating information even without class labels. Based on the concept of SSP, it presents a new method for multi-feature information fusion based on the Sparsity Preserving Multiple Canonical Correlation Analysis (SPMCCA), which can preserve the sparse reconstructive relationship of the data for recognition from multi-feature information representation. We implement a prototype of SPM-CCA with the application to visual-based human emotion recognition. Experimental results show that the proposed method outperforms the traditional methods of serial fusion, Canonical Correlation Analysis (CCA), Multiple Canonical Correlation Analysis (MCCA) and recently proposed Sparsity Preserving Canonical Correlation Analysis (SPCCA).
Lei Gao 0001, Lin Qi 0001, Ling Guan
ICIP1
2015 Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component Analysis in Canonical Correlation Space
abstract
Kernel Entropy Component Analysis(KECA), an effective information fusion tool, is realized using descriptor of information entropy and optimized by entropy estimation. However, it merely put the information or data from different channels together to achieve the information fusion without considering their intrinsic structures and relations. In this paper, we enhance the performance of KECA by introducing KECA in Canonical Correlation Space (CCS) or KECA+CCS. Not only the intrinsic structures and relations are considered in CCS, but also the nature of input data are revealed by entropy estimation. It improves the recognition accuracy effectively. The effectiveness of the proposed method is evaluated through experimentation on two audio-based emotion databases. The results show that the proposed method outperforms the existing methods based on similar principles.
Lei Gao 0001, Lin Qi 0001, Ling Guan
ISM1
2012 Discriminative Multiple Canonical Correlation Analysis for Multi-feature Information Fusion
abstract
This paper presents a novel approach for multi-feature information fusion. The proposed method is based on the Discriminative Multiple Canonical Correlation Analysis (DMCCA), which can extract more discriminative characteristics for recognition from multi-feature information representation. It represents the different patterns among multiple subsets of features identified by minimizing the Frobenius norm. We will demonstrate that the Canonical Correlation Analysis (CCA), the Multiple Canonical Correlation Analysis (MCCA), and the Discriminative Canonical Correlation Analysis (DCCA) are special cases of the DMCCA. The effectiveness of the DMCCA is demonstrated through experimentation in speaker recognition and speech-based emotion recognition. Experimental results show that the proposed approach outperforms the traditional methods of serial fusion, CCA, MCCA and DCCA.
Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan
ISM1
2012 2D-FRFT Based Rotation Invariant Digital Image Watermarking
abstract
The extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments.
Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan
ISM1
2012 Generalized MMSD feature extraction using QR decomposition
abstract
Multiple Maximum scatter difference (MMSD) discriminant criterion is an effective feature extraction method that computes the discriminant vectors from both the range of the between-class scatter matrix and the null space of the within-class scatter matrix. However, singular value decomposition (SVD) of two times is involved in MMSD, making this method impractical for high dimensional data. In this paper, we propose a novel method for feature extraction and classification based on MMSD criterion, called generalized MMSD (GMMSD), which employs QR decomposition rather than SVD. Unlike MMSD, GMMSD does not require the computation of the whole scatter matrix. Instead, it computes the discriminant vectors from both the range of whitenizated input data matrix and the null space of the within-class scatter matrix. We evaluate the effectiveness of the GMMSD method in terms of classification accuracy in the reduced dimensional space. Our experiments on two facial expression databases demonstrate that the GMMSD method provides favorable performance in terms of both recognition accuracy and computational efficiency.
Ning Zheng 0003, Lin Qi 0001, Lei Gao 0001, Ling Guan
VCIP3