Rui Wang 0050

dblp:06/2293-50 · DBLP profile ↗
← Back
29ranked-venue papers
15as first author
26since 2021 · last 2026
0000-0002-9984-1752ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 12 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Wasserstein-Aligned Hyperbolic Multi-View Clustering
abstract
Multi-view clustering (MVC) aims to uncover the latent structure of multi-view data by learning view-common and view-specific information. Although recent studies have explored hyperbolic representations for better tackling the representation gap between different views, they focus primarily on instance-level alignment and neglect global semantic consistency, rendering them vulnerable to view-specific information (e.g., noise and cross-view discrepancies). To this end, this paper proposes a novel Wasserstein-Aligned Hyperbolic (WAH) framework for multi-view clustering. Specifically, our method exploits a view-specific hyperbolic encoder for each view to embed features into the Lorentz manifold for hierarchical semantic modeling. Whereafter, a global semantic loss based on the hyperbolic sliced-Wasserstein distance is introduced to align manifold distributions across views. This is followed by soft cluster assignments to encourage cross-view semantic consistency. Extensive experiments on multiple benchmarking datasets show that our method can achieve SOTA clustering performance.
Rui Wang 0050, Xiaoqing Luo, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001
AAAI1
2026 Vision Mamba-enhanced Multi-level Context Aggregation Network for polyp segmentation
Jiaye Chen, Tianyang Xu 0001, Rui Wang 0050, Xiaoning Song, Tao Zhou 0002
Pattern Recognit.4
2025 R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration
abstract
Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.
Yang Hua 0002, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Rui Wang 0050, Wenjie Zhang 0009, Xiaojun Wu 0001
AAAI5
2025 Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry
abstract
Covariance matrices have proven highly effective across many scientific fields. Since these matrices lie within the Symmetric Positive Definite (SPD) manifold—a Riemannian space with intrinsic non-Euclidean geometry, the primary challenge in representation learning is to respect this underlying geometric structure. Drawing inspiration from the success of Euclidean deep learning, researchers have developed neural networks on the SPD manifolds for more faithful covariance embedding learning. A notable advancement in this area is the implementation of Riemannian batch normalization (RBN), which has been shown to improve the performance of SPD network models. Nonetheless, the Riemannian metric beneath the existing RBN might fail to effectively deal with the ill-conditioned SPD matrices (ICSM), undermining the effectiveness of RBN. In contrast, the Bures-Wasserstein metric (BWM) demonstrates superior performance for ill-conditioning. In addition, the recently introduced Generalized BWM (GBWM) parameterizes the vanilla BWM via an SPD matrix, allowing for a more nuanced representation of vibrant geometries of the SPD manifold. Therefore, we propose a novel RBN algorithm based on the GBW geometry, incorporating a learnable metric parameter. Moreover, the deformation of GBWM by matrix power is also introduced to further enhance the representational capacity of GBWMbased RBN. Experimental results on different datasets validate the effectiveness of our proposed method. The code is available at https://github.com/jjscc/GBWBN.
Rui Wang 0050, Shaocheng Jin, Ziheng Chen 0001, Xiaoqing Luo, Xiaojun Wu 0001
CVPR1
2025 A Correlation Manifold Self-Attention Network for EEG Decoding
abstract
Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geometrically capture the spatiotemporal dependencies inherent in time series data, e.g., electroencephalography (EEG). Recent studies have highlighted the full-rank correlation matrix as an advantageous alternative to the covariance matrix for data representation, owing to its invariance to the scale of variables. Motivated by these advancements, we propose the Correlation Attention Network (CorAtt) tailored for full-rank correlation matrices and implement it under the permutation-invariant and computationally efficient Off-Log and Log-Scaled geometries, respectively. Extensive evaluations on three benchmarking EEG datasets provide substantial evidence for the effectiveness of our introduced CorAtt. The code and supplementary material can be found at https://github.com/ChenHu-ML/CorAtt.
Rui Wang 0050, Xiaoning Song, Tao Zhou 0002, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001
IJCAI2
2025 SymCL: Riemannian Contrastive Learning on the Symmetric Positive Definite Manifold for Visual Classification
abstract
Symmetric Positive Definite (SPD) matric has been proven to be an effective feature descriptor in the realm of artificial intelligence, as it can encode spatiotemporal statistical information of data on a curved Riemannian manifold, i.e., SPD manifold. Although existing Riemannian neural networks have demonstrated superiority in many scientific fields, the inherent reliance on labels within supervised learning renders them susceptible to label errors. Besides, it is insufficient to depend solely on labels to learn effective feature distributions in some complicated data scenarios. Drawing inspiration from the considerable achievements of contrastive learning (CL) across diverse tasks, we extend the conventional CL paradigm to the context of SPD manifolds, which we denote SymCL, paving the way for a novel approach in SPD matrix-based visual classification. Furthermore, we inject a Riemannian triplet loss-based Riemannian metric learning (RML) into the designed SPD manifold CL framework for the sake of improving the discrimination of the learned geometric representations. Extensive experimental results on four datasets verify the effectiveness of the proposed algorithm.
Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IJCNN2
2025 Learning a Discriminative Grassmannian Neural Network for Visual Classification
abstract
Learning representations on the Grassmannian manifold is popular in quite a few visual classification tasks. With the development of deep learning techniques, several neural networks have recently emerged for processing subspace data. However, the diversely changed appearance of the signal data (video clips and image sets), makes it impossible for the existing Grassmannian networks (GrasNets) that rely on a single cross-entropy loss for end-to-end training to learn effective geometric representations, especially for complicated visual scenarios. To solve this problem, a Riemannian triplet loss-based Riemannian metric learning mechanism is introduced to the original GrasNet, which can explicitly encode and learn the characteristics of the intra- and inter-class data distributions conveyed by the input data during network training. Additionally, given the existence of intra-class diversity and inter-class ambiguity of the input data, we propose a hard sample reward strategy (HSR) to further improve the discriminability of the learned network embedding. Extensive experimental results obtained on four benchmarking datasets demonstrate the effectiveness of the proposed method.
Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Umapada Pal 0001, Josef Kittler
IJCNN2
2025 Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
abstract
Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality.
Xiaoqing Luo, Zhancheng Zhang, Hui Li 0037, Rui Wang 0050, Zhenhua Feng 0001, Xiaoning Song
NeurIPS6
2025 Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds
abstract
Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored to specific geometries, limiting their applicability. On the other hand, recent studies show that several matrix manifolds, such as Symmetric Positive Definite (SPD), Symmetric Positive Semi-Definite (SPSD), and Grassmannian manifolds, admit gyrovector structures, which extend vector addition and scalar product into manifolds. Leveraging these properties, we propose a Gyro Attention (GyroAtt) framework over general gyrovector spaces, applicable to various matrix geometries. Empirically, we manifest GyroAtt on three gyro structures on the SPD manifold, three on the SPSD manifold, and one on the Grassmannian manifold. Extensive experiments on four electroencephalography (EEG) datasets demonstrate the effectiveness of our framework.
Rui Wang 0050, Xiaoning Song, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001
NeurIPS1
2025 Geometry-Aware Self-attention Network with Adaptive Log-Euclidean Metric for EEG Decoding
Zihao Bi, Rui Wang 0050, Tao Zhou 0002, Xiaoning Song, Xiaojun Wu 0001
PRCV (4)2
2025 SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Huan Kang, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Chunyang Cheng, Josef Kittler
Int. J. Comput. Vis.5
2025 Learning a Better SPD Network for Signal Classification: A Riemannian Batch Normalization Method
abstract
Symmetric positive definite (SPD) matrices have been widely used as Riemannian feature descriptors in various scientific fields, due to their capacity to encode effective manifold-valued representations. Inspired by the architectural principles of Euclidean deep learning, the emerging SPD neural networks have achieved more robust signal classification. Among these advancements, Riemannian batch normalization (RBN) based on the affine-invariant Riemannian metric (AIRM) has emerged as a key technique for enhancing the learning capability of SPD-based networks. Nevertheless, the reliance of singular value decomposition (SVD) makes this metric relatively unstable for the computation of SPD matrices, especially for the ill-conditioned case. To address this limitation, we propose a novel RBN algorithm based on the recently introduced log-Cholesky metric (LCM), which leverages Cholesky decomposition. Unlike AIRM, the LCM offers enhanced numerical stability and allows for more efficient computation. Specifically, the LCM-based Riemannian operators such as Fr $\acute {\mathrm {e}}$ chet mean and parallel transport (PT) are much simpler than those of AIRM, and both have closed forms. Besides, since LCM is the pullback metric from the Cholesky manifold via Cholesky decomposition, the LCM-based RBN on the SPD manifold can be computed in the Cholesky manifold, further boosting the efficiency. Extensive experiments conducted on four benchmarking datasets certify the effectiveness of our proposed algorithm. The source code is now available at: https://github.com/jjscc/CBN.git.
Rui Wang 0050, Shaocheng Jin, Zhenyu Cai, Ziheng Chen 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Grassmannian Manifold Self-Attention Network for Signal Classification
Rui Wang 0050, Ziheng Chen 0001, Xiaojun Wu 0001, Xiaoning Song
IJCAI1
2024 A Riemannian Residual Learning Mechanism for SPD Network
abstract
The generalization of Euclidean network paradigm to the Riemannian manifolds has attracted much attention for offering useful geometric representations in processing manifold-valued data in recent years. However, the information degradation during data compression mapping hinders Riemannian networks from going deeper, and there are very few solutions specifically designed for this problem. Given the remarkable success of deep Residual learning in Euclidean networks, a novel Riemannian residual learning mechanism (RRLM) is proposed in the context of Symmetric Positive Definite (SPD) manifolds, enabling the characterization of deep spatiotemporal features while preserving the manifold properties. Based on RRLM, a stack of SPD manifold-constrained residual-like blocks is designed on the tail of the original SPDNet(backbone) for the sake of conducting deep Riemannian residual learning. For simplicity, we refer to the network architecture introduced above as Riemannian residual SPD network (ResSPDNet). The experimental results achieved on three types of visual classification tasks, i.e., facial emotion recognition, drone recognition, and action recognition, demonstrate that our method can achieve improved accuracy with a deepened network structure.
Zhenyu Cai, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IJCNN2
2024 RMLR: Extending Multinomial Logistic Regression into General Geometries
abstract
Riemannian neural networks, which extend deep learning techniques to Riemannian spaces, have gained significant attention in machine learning. To better classify the manifold-valued features, researchers have started extending Euclidean multinomial logistic regression (MLR) into Riemannian manifolds. However, existing approaches suffer from limited applicability due to their strong reliance on specific geometric properties. This paper proposes a framework for designing Riemannian MLR over general geometries, referred to as RMLR. Our framework only requires minimal geometric properties, thus exhibiting broad applicability and enabling its use with a wide range of geometries. Specifically, we showcase our framework on the Symmetric Positive Definite (SPD) manifold and special orthogonal group, i.e., the set of rotation matrices. On the SPD manifold, we develop five families of SPD MLRs under five types of power-deformed metrics. On rotation matrices we propose Lie MLR based on the popular bi-invariant metric. Extensive experiments on different Riemannian backbone networks validate the effectiveness of our framework.
Ziheng Chen 0001, Yue Song 0002, Rui Wang 0050, Xiaojun Wu 0001, Nicu Sebe
NeurIPS3
2024 Manifold-based multi-graph embedding for semi-supervised classification
Jiang-Tao Song, Jia-Sheng Chen, Rui Wang 0050, Xiaojun Wu 0001
Pattern Recognit. Lett.4
2024 Deep Metric Learning on the SPD Manifold for Image Set Classification
abstract
Thanks to the efficacy of Symmetric Positive Definite (SPD) manifold in characterizing video sequences (image sets), image set-based visual classification has made remarkable progress. However, the issue of large intra-class diversity and inter-class similarity is still an open challenge for the research community. Although several recent studies have alleviated the above issue by constructing Riemannian neural networks for SPD matrix nonlinear processing, the degradation of structural information during multi-stage feature transformation impedes them from going deeper. Besides, a single cross-entropy loss is insufficient for discriminative learning as it neglects the peculiarities of data distribution. To this end, this paper develops a novel framework for image set classification. Specifically, we first choose a mainstream neural network built on the SPD manifold (SPDNet)[25]as the backbone with a stacked SPD manifold autoencoder (SSMAE) built on the tail to enrich the structured representations. Due to the associated reconstruction error terms, the embedding mechanism of both SSMAE and each SPD manifold autoencoder (SMAE) forms an approximate identity mapping, simplifying the training of the suggested deeper network. Then, the ReCov layer is introduced with a nonlinear function for the constructed architecture to narrow the discrepancy of the intra-class distributions from the perspective of regularizing the local statistical information of the SPD data. Afterward, two progressive metric learning stages are coupled with the proposed SSMAE to explicitly capture, encode, and analyze the geometric distributions of the generated deep representations during training. In consequence, not only a more powerful Riemannian network embedding but also effective classifiers can be obtained. Finally, a simple maximum voting strategy is applied to the outputs of the learned multiple classifiers for classification. The proposed model is evaluated on three typical visual classification tasks using widely adopted benchmarking datasets. Extensive experiments show its superiority over the state of the arts.
Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.1
2024 SPD Manifold Deep Metric Learning for Image Set Classification
abstract
By characterizing each image set as a nonsingular covariance matrix on the symmetric positive definite (SPD) manifold, the approaches of visual content classification with image sets have made impressive progress. However, the key challenge of unhelpfully large intraclass variability and interclass similarity of representations remains open to date. Although, several recent studies have mitigated the two problems by jointly learning the embedding mapping and the similarity metric on the original SPD manifold, their inherent shallow and linear feature transformation mechanism are not powerful enough to capture useful geometric features, especially in complex scenarios. To this end, this article explores a novel approach, termed SPD manifold deep metric learning (SMDML), for image set classification. Specifically, SMDML first selects a prevailing SPD manifold neural network (SPDNet) as the backbone (encoder) to derive an SPD matrix nonlinear representation. To counteract the degradation of structural information during multistage feature embedding, we construct a Riemannian decoder at the end of the encoder, trained by a reconstruction error term (RT), to induce the generated low-dimensional feature manifold of the hidden layer to capture the pivotal information about the visual data describing the imaged scene. We demonstrate through theory and experiments that it is feasible to replace the Riemannian metric with Euclidean distance in RT. Then, the ReCov layer is introduced into the established Riemannian network to regularize the local statistical information within each input feature matrix, which enhances the effectiveness of the learning process. The theoretical analysis of the activation function used in the ReCov layer in terms of continuity and conditions for generating positive definite matrices is beneficial for network design. Inspired by the fact that the single cross-entropy loss used for training is unable to effectively parse the geometric distribution of the deep representations, we finally endow the suggested model with a novel metric learning regularization term. By explicitly incorporating the encoding and processing of the data variations into the network learning process, this term can not only derive a powerful Riemannian representation but also train an effective classifier. The experimental results show the superiority of the proposed approach on three typical visual classification tasks.
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Josef Kittler
IEEE Trans. Neural Networks Learn. Syst.1
2023 Riemannian Local Mechanism for SPD Neural Networks
abstract
The Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicitly mine the local geometrical information in deep SPD feature representations. Given the great success of local mechanisms in Euclidean methods, we argue that it is of utmost importance to ensure the preservation of local geometric information in the SPD networks. We first analyse the convolution operator commonly used for capturing local information in Euclidean deep networks from the perspective of a higher level of abstraction afforded by category theory. Based on this analysis, we define the local information in the SPD manifold and design a multi-scale submanifold block for mining local geometry. Experiments involving multiple visual tasks validate the effectiveness of our approach.
Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Zhiwu Huang, Josef Kittler
AAAI4
2023 U-SPDNet: An SPD manifold learning-based neural network for visual classification
Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
Neural Networks1
2023 Hybrid Riemannian Graph-Embedding Metric Learning for Image Set Classification
abstract
With the continuously increasing amount of video data, image set classification has recently received widespread attention in the CV&PR community. However, the intra-class diversity and inter-class ambiguity of representations remain an open challenge. To tackle this issue, several methods have been put forward to perform multiple geometry-aware image set modelling and learning. Although the extracted complementary geometric information is beneficial for decision making, the sophisticated computational paradigm (e.g., scatter matrices computation and iterative optimisation) of such algorithms is counterproductive. As a countermeasure, we propose an effective hybrid Riemannian metric learning framework in this paper. Specifically, we design a multiple graph embedding-guided metric learning framework for the sake of fusing these complementary kernel features, obtained via the explicit RKHS embeddings of the Grassmannian manifold, SPD manifold, and Gaussian embedded Riemannian manifold, into a unified subspace for classification. Furthermore, the involved optimisation problem of the developed model can be solved in terms of a series of sub-problems, achieving improved efficiency theoretically and experimentally. Substantial experiments are carried out to evaluate the efficacy of our approach. The experimental results suggest the superiority of it over the state-of-the-art methods.
Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler
IEEE Trans. Big Data4
2022 DreamNet: A Deep Riemannian Manifold Network for SPD Matrix Learning
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler
ACCV (6)1
2022 Learning a discriminative SPD manifold neural network for image set classification
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler
Neural Networks1
2022 Multiple Riemannian Manifold-Valued Descriptors Based Image Set Classification With Multi-Kernel Metric Learning
abstract
The importance of wild video based image set recognition is monotonically increasing due to the large amount of video data being collected by various devices including surveillance cameras, drive recorders, smart phones, and internet. The content of these videos is often complex, and it raises the question of how to perform image set modeling and feature extraction for image set-based classification. In recent years, image set classification methods have advanced considerably by modeling the image set in terms of a covariance matrix, linear subspace, or Gaussian distribution. Moreover, the distinctive geometry spanned by them include Symmetric Positive Definite (SPD) manifold, Grassmannian manifold, and Gaussian embedded Riemannian manifold, respectively. As a matter of fact, most of the approaches just adopt a single geometric model to describe each given image set, which may lose information useful for classification. To tackle this problem, we propose a novel algorithm to model each image set from a multi-geometric perspective. Specifically, the covariance matrix, linear subspace, and Gaussian distribution are applied to set representation simultaneously. In order to fuse these multiple heterogeneous Riemannian manifold-valued features, the well-equipped Riemannian kernel functions are first employed to map them into high dimensional Hilbert spaces. Then, a multi-kernel metric learning framework is devised to embed the learned hybrid kernels into a common lower dimensional subspace to facilitate classification. We conduct experiments on six widely used datasets each representing a different classification task: video-based face recognition, set-based object categorization, video-based emotion recognition, dynamic scene classification, set-based cell identification, and 3D hand pose estimation, to evaluate the classification performance of the proposed algorithm. The extensive experimental results confirm its superiority over the state-of-the-art methods.
Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler
IEEE Trans. Big Data1
2022 SymNet: A Simple Symmetric Positive Definite Manifold Deep Learning Method for Image Set Classification
abstract
By representing each image set as a nonsingular covariance matrix on the symmetric positive definite (SPD) manifold, visual classification with image sets has attracted much attention. Despite the success made so far, the issue of large within-class variability of representations still remains a key challenge. Recently, several SPD matrix learning methods have been proposed to assuage this problem by directly constructing an embedding mapping from the original SPD manifold to a lower dimensional one. The advantage of this type of approach is that it cannot only implement discriminative feature selection but also preserve the Riemannian geometrical structure of the original data manifold. Inspired by this fact, we propose a simple SPD manifold deep learning network (SymNet) for image set classification in this article. Specifically, we first design SPD matrix mapping layers to map the input SPD matrices into new ones with lower dimensionality. Then, rectifying layers are devised to activate the input matrices for the purpose of forming a valid SPD manifold, chiefly to inject nonlinearity for SPD matrix learning with two nonlinear functions. Afterward, we introduce pooling layers to further compress the input SPD matrices, and the log-map layer is finally exploited to embed the resulting SPD matrices into the tangent space via log-Euclidean Riemannian computing, such that the Euclidean learning applies. For SymNet, the (2-D)2principal component analysis (PCA) technique is utilized to learn the multistage connection weights without requiring complicated computations, thus making it be built and trained easier. On the tail of SymNet, the kernel discriminant analysis (KDA) algorithm is coupled with the output vectorized feature representations to perform discriminative subspace learning. Extensive experiments and comparisons with state-of-the-art methods on six typical visual classification tasks demonstrate the feasibility and validity of the proposed SymNet.
Rui Wang 0050, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Neural Networks Learn. Syst.1
2021 Graph Embedding Multi-Kernel Metric Learning for Image Set Classification With Grassmannian Manifold-Valued Features
abstract
In the domain of video-based image set classification, a considerable advance has been made by modeling a sequence of video frames (image set) as a linear subspace, which typically resides on a Grassmannian manifold. As a consequence of the large intra-class variations of the video data, there are two open challenges for the modeling task: how to establish appropriate image set models to encode these variations, and how to effectively measure the similarity between any two image sets. As a possible way to tackle these issues, this paper presents a graph embedding multi-kernel metric learning (GEMKML) algorithm for image set classification. The proposed GEMKML implements set modeling, feature extraction, and classification in two steps. Firstly, the proposed framework constructs a novel cascaded feature learning architecture on Grassmannian manifold with the aim of producing more effective Grassmannian manifold-valued feature representations. To make a better use of these learned features, a graph embedding multi-kernel metric learning scheme is then devised to map them into a lower-dimensional Euclidean space, where the inter-class distances are maximized and the intra-class distances are minimized. We evaluate the proposed GEMKML on five different visual classification tasks using widely adopted datasets. The extensive classification results confirm its superiority over the state-of-the-art methods.
Rui Wang 0050, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Multim.1
2020 GrasNet: A Simple Grassmannian Network for Image Set Classification
Rui Wang 0050, Xiaojun Wu 0001
Neural Process. Lett.1
2018 Riemannian kernel based Nyström method for approximate infinite-dimensional covariance descriptors with application to image set classification
abstract
In the domain of pattern recognition, using the CovDs (Covariance Descriptors) to represent data and taking the metrics of the resulting Riemannian manifold into account have been widely adopted for the task of image set classification. Recently, it has been proven that infinite-dimensional CovDs are more discriminative than their low-dimensional counterparts. However, the form of infinite-dimensional CovDs is implicit and the computational load is high. We propose a novel framework for representing image sets by approximating infinite-dimensional CovDs in the paradigm of the Nyström method based on a Riemannian kernel. We start by modeling the images via CovDs, which lie on the Riemannian manifold spanned by SPD (Symmetric Positive Definite) matrices. We then extend the Nyström method to the SPD manifold and obtain the approximations of CovDs in RKHS (Reproducing Kernel Hilbert Space). Finally, we approximate infinite-dimensional CovDs via these approximations. Empirically, we apply our framework to the task of image set classification. The experimental results obtained on three benchmark datasets show that our proposed approximate infinite-dimensional CovDs outperform the original CovDs.
Kai-Xuan Chen 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler
ICPR3
2018 Multiple Manifolds Metric Learning with Application to Image Set Classification
abstract
In image set classification, a considerable advance has been made by modeling the original image sets by second order statistics or linear subspace, which typically lie on the Riemannian manifold. Specifically, they are Symmetric Positive Definite (SPD) manifold and Grassmann manifold respectively, and some algorithms have been developed on them for classification tasks. Motivated by the inability of existing methods to extract discriminatory features for data on Riemannian manifolds, we propose a novel algorithm which combines multiple manifolds as the features of the original image sets. In order to fuse these manifolds, the well-studied Riemannian kernels have been utilized to map the original Riemannian spaces into high dimensional Hilbert spaces. A metric Learning method has been devised to embed these kernel spaces into a lower dimensional common subspace for classification. The state-of-the-art results achieved on three datasets corresponding to two different classification tasks, namely face recognition and object categorization, demonstrate the effectiveness of the proposed method.
Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler
ICPR1