Xiaobo Shen 0001

dblp:136/1805 · also Xiao-Bo Shen · DBLP profile ↗
← Back
87ranked-venue papers
27as first author
48since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 15 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 18 first-author · 26 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph-Driven Domain Co-Adaptation for Cross-Domain Image Quality Assessment
abstract
As a typical information medium, images are widely utilized across various scenarios. Measuring image quality accurately is meaningful for the subsequent usability of images. However, significant variations exist in image types and distortion types in different scenarios. And, acquiring labeled images for each specific scenario is time-consuming and labor-intensive. Consequently, designing cross-domain image quality assessment (IQA) that generalizes across different scenarios remains a substantial challenge. Existing cross-domain IQA methods primarily focus on content relevance while neglecting distortion differences, leading to limited applicability while distortion fluctuates. To address these limitations, a graph-driven domain co-adaptation framework for cross-domain IQA (GDCIQA) is proposed. Firstly, a graph knowledge sharing (GKS) module that constructs graphs via inter-domain distortion relevance has been proposed. GKS employs graph neural networks to update quality-aware features in the source domain by leveraging target-domain representations. Secondly, the proposed co-adaptation learning (CAL) mechanism can enable joint optimization of different modules, which ensures comprehensive sharing of quality-aware and distortion-related information. Finally, a domain adaptation framework has been designed to train models effectively on labeled source images, yielding target-domain-optimized IQA models. Experimental results demonstrate that GDCIQA achieves higher accuracy and stability in cross-domain scenarios. The proposed GKS and CAL can advance cross-domain IQA research.
Shun Zhu, Xichen Yang, Tianshu Wang 0001, Zhongyuan Mao, Tianyin Li, Zhuoyan Sun, Xiaobo Shen 0001
AAAI8
2025 HUANG: A Robust Diffusion Model-based Targeted Adversarial Attack Against Deep Hashing Retrieval
abstract
Deep hashing models have achieved great success in retrieval tasks due to their powerful representation and strong information compression capabilities. However, they inherit the vulnerability of deep neural networks to adversarial perturbations. Attackers can severely impact the retrieval capability of hashing models by adding subtle, carefully crafted adversarial perturbations to benign images, transforming them into adversarial images. Most existing adversarial attacks target image classification models, with few focusing on retrieval models. We propose HUANG, the first targeted adversarial attack algorithm to leverage a diffusion model for hashing retrieval in black-box scenarios. In our approach, adversarial denoising uses adversarial perturbations and residual image to guide the shift from benign to adversarial distribution. Extensive experiments demonstrate the superiority of HUANG across different datasets, achieving state-of-the-art performance in black-box targeted attacks. Additionally, the dynamic interplay between denoising and adding adversarial perturbations in adversarial denoising endows HUANG with exceptional robustness and transferability.
Chihan Huang, Xiaobo Shen 0001
AAAI2
2025 Adversarial Contrastive Graph Masked AutoEncoder Against Graph Structure and Feature Dual Attacks
abstract
Graph Neural Networks (GNNs) have been shown vulnerable to graph adversarial attacks. Current robust graph representation learning methods mainly defend against graph structure attack, and improves performance of GNNs. However node feature in graph can been easily attacked in reality. The joint defense on graph structure and feature dual attacks remains challenging yet less studied. To fulfill this gap, we propose Adversarial Contrastive Graph Masked AutoEncoder (ACGMAE) to defend against graph structure and feature dual attacks. ACGMAE employs adversarial feature masking for reconstructing node feature to mitigate the influence of feature attack. ACGMAE employs contrastive learning on kNN graph and attacked graph, considers neighbor nodes as positive samples, and further calculates their probabilities being true positive to mitigate the effect of adversarial edges. Extensive experiments on node classification and clustering demonstrate the effectiveness of the proposed ACGMAE especially under graph structure and feature dual attacks.
Weixuan Shen, Xiaobo Shen 0001, Shirui Pan
AAAI2
2025 PoemBERT: A Dynamic Masking Content and Ratio Based Semantic Language Model For Chinese Poem Generation
abstract
Ancient Chinese poetry stands as a crucial treasure in Chinese culture. To address the absence of pre-trained models for ancient poetry, we introduced PoemBERT, a BERT-based model utilizing a corpus of classical Chinese poetry. Recognizing the unique emotional depth and linguistic precision of poetry, we incorporated sentiment and pinyin embeddings into the model, enhancing its sensitivity to emotional information and addressing challenges posed by the phenomenon of multiple pronunciations for the same Chinese character. Additionally, we proposed Character Importance-based masking and dynamic masking strategies, significantly augmenting the model’s capability to extract imagery-related features and handle poetry-specific information. Fine-tuning our PoemBERT model on various downstream tasks, including poem generation and sentiment classification, resulted in state-of-the-art performance in both automatic and manual evaluations. We provided explanations for the selection of the dynamic masking rate strategy and proposed a solution to the issue of a small dataset size.
Chihan Huang, Xiaobo Shen 0001
COLING2
2025 Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
abstract
The low resolution (LR) problem is rather challenging in face analysis. Most existing face hallucination methods assume that LR face images have only one resolution, but multiple resolutions may be available from different sources. To solve this issue, we propose a novel simultaneous facial canonical correlation representation learning method for face hallucination, which seeks latent correlation subspaces for multi-resolution views. Our method jointly solves multiple linear transformations by optimizing a correlation summation criterion of all pairs of resolutions. The neighborhood reconstruction is used to infer the HR facial canonical correlation representation of LR face inputs. Extensive experimental results show the superiority of our proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Yun Li 0010
ICASSP5
2025 Efficient Multi-branch Black-box Semantic-aware Targeted Attack Against Deep Hashing Retrieval
abstract
Deep hashing have achieved exceptional performance in retrieval tasks due to their robust representational capabilities. However, they inherit the vulnerability of deep neural networks to adversarial attacks. These models are susceptible to finely crafted adversarial perturbations that can lead them to return incorrect retrieval results. Although numerous adversarial attack methods have been proposed, there has been a scarcity of research focusing on targeted black-box attacks against deep hashing models. We introduce the Efficient Multi-branch Black-box Semantic-aware Targeted Attack against Deep Hashing Retrieval (EmbSTar), capable of executing targeted black-box attacks on hashing models. Initially, we distill the target model to create a knockoff model. Subsequently, we devised novel Target Fusion and Target Adaptation modules to integrate and enhance the semantic information of the target label and image. Knockoff model is then utilized to align the adversarial image more closely with the target image semantically. With the knockoff model, we can obtain powerful targeted attacks with few queries. Extensive experiments demonstrate that EmbSTar significantly surpasses previous models in its targeted attack capabilities, achieving SOTA performance for targeted black-box attacks.
Chihan Huang, Xiaobo Shen 0001
ICASSP2
2025 Distributed Cascaded Manifold Hashing Network for Compact Image Set Representation
abstract
Conventional image set methods typically learn from image sets stored in a single location. However, in real-world applications, image sets are often distributed across different locations. Learning from such distributed sets using deep neural networks poses challenges for efficient image set classification and retrieval. To address this, we propose Distributed Cascade Manifold Hashing Network (DCMHN) for compact image set representation. DCMHN represents each image set using an SPD manifold and utilizes a manifold hashing network to generate hash codes, enabling efficient classification and retrieval. The network is trained in a cascaded manner, where the bilinear mapping in the BiMap layer is learned first, followed by joint learning of the hash function and classifier in the hash layer. DCMHN enforces local consistency on global variables across neighboring nodes, allowing parallel optimization. Extensive experiments on three benchmark image set datasets demonstrate that the proposed DCMHN achieves competitive accuracies in distributed settings, and outperforms state-of-the-arts in terms of computation and storage efficiency.
Xiaxin Wang, Haoyu Cai, Xiaobo Shen 0001
IJCAI3
2025 Underwater image quality evaluation via deep meta-learning: Dataset and objective method
Tianhai Chen, Xichen Yang, Tianshu Wang 0001, Nengxin Li, Shun Zhu, Xiaobo Shen 0001
Comput. Vis. Image Underst.6
2025 Graph guided local structure propagation for tensorial multi-view subspace clustering
abstract
Multi-view subspace clustering (MSC) is widely studied owing to the ability to capture the diverse and complementary information hidden in multiple views. As a representative model, tensorial MSC can capture global information by leveraging high-order correlations across various perspectives, leading to promising results. However, this approach fails to reveal the local structure in the specific view and ignores the prior information of the self-representation tensor. To address the problems, we propose the novel graph-guided local structure propagation (GGLSP) for tensorial MSC. First, we improve the adaptive graph model to acquire a fused graph similarity matrix for extracting the relationships between samples and propagating the local structure information to the self-representation tensor. Subsequently, we introduce the weighted tensor Schatten p-norm to approximate the tensor rank function by exploiting the contributions of different singular values so that the self-representation tensor can better reveal the global low-rank structure information. Finally, we develop two efficient algorithms to solve the optimization problems. Large numbers of experiments on seven popular datasets confirm the superiority of our proposed GGLSP.
Tao Zhang 0015, Yizhang Wang, Xiaobo Shen 0001, Fan Liu 0003
Intell. Data Anal.4
2025 Semi-Paired Semi-Supervised Deep Hashing for cross-view retrieval
abstract
Abstract Due to its fast computational speed and low storage cost, hashing has been effectively applied to large‐scale multimedia retrieval tasks, such as medical video and security video retrieval. Most existing cross‐view hashing methods require good matching information, however, this exact pairing relationship is difficult to fully realise in practice. The association between views is incomplete, as is the label information. This task of missing paired and labelled information is very challenging, but less explored in research. In this study, a semi‐supervised semi‐paired deep hashing for large‐scale data is proposed, named Semi‐Paired Semi‐Supervised Deep Hashing (SPSDH) to solve this challenging task. SPSDH is a novel end‐to‐end deep neural network model with high‐order affinity. A non‐local higher‐order affinity measure that better considers the multimodal neighbourhood structure is proposed. A common representation to associate different modalities is introduced, which combined with the labelled information greatly maintains the consistency within the modalities. SPSDH is evaluated on three benchmark datasets for large‐scale cross‐view approximate nearest neighbour search and compared with several state‐of‐the‐art hashing methods. Extensive experimental results demonstrate the superior performance of our proposed SPSDH in semi‐supervised semi‐paired retrieval tasks.
Yi Wang 0104, Xiaobo Shen 0001, Zhenmin Tang, Ming Zhang 0033
IET Comput. Vis.2
2025 Meta-learning enhanced global-local feature fusion for image quality assessment
Nengxin Li, Xichen Yang, Tianhai Chen, Shun Zhu, Zhongyuan Mao, Tianshu Wang 0001, Xiaobo Shen 0001
Mach. Vis. Appl.7
2025 Multi-Type Image Quality Assessment Based on Multi-Region Deep Feature Fusion Under Meta-Learning
abstract
Most existing image quality assessment methods need to be retrained when dealing with a new type of task. This approach wastes computing resources and time. Therefore, these methods fail to suit the application scenarios that require processing of multi-type image quality assessment tasks. In the human visual system, the eyes of human tend to pay varying degrees of attention to different regions. Inspired by this system, this paper proposes a multi-type image quality assessment method based on multi-region deep feature fusion under meta-learning (MMQA). First, we utilize the differences in the structural information to screen out salient and non-salient regions. Second, a deep multi-stream network is designed to comprehensively consider and fuse different features related to the quality in salient regions, non-salient regions and the entire image. Third, meta-learning is applied to quickly learn and update the parameters of the model when facing new types of images. By summarizing the prior knowledge in the training of one type of task, the model can be quickly fine-tuned for other types of images. The experimental results demonstrate that the proposed method has advantages over the existing methods in generalization and robustness. Furthermore, the proposed method can adapt well to different distortion types and different image types quickly and accurately.
Shun Zhu, Xichen Yang, Tianshu Wang 0001, Tianhai Chen, Nengxin Li, Xiaobo Shen 0001, Genlin Ji
IEEE Trans. Multim.6
2025 Graph Convolutional Multi-Label Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing encodes different modalities of multimodal data into low-dimensional Hamming space for fast cross-modal retrieval. In multi-label cross-modal retrieval, multimodal data are often annotated with multiple labels, and some labels, e.g., "ocean" and "cloud," often co-occur. However, existing cross-modal hashing methods overlook label dependency that is crucial for improving performance. To fulfill this gap, this article proposes graph convolutional multi-label hashing (GCMLH) for effective multi-label cross-modal retrieval. Specifically, GCMLH first generates word embedding of each label and develops label encoder to learn highly correlated label embedding via graph convolutional network (GCN). In addition, GCMLH develops feature encoder for each modality, and feature fusion module to generate highly semantic feature via GCN. GCMLH uses teacher-student learning scheme to transfer knowledge from the teacher modules, i.e., label encoder and feature fusion module, to the student module, i.e., feature encoder, such that learned hash code can well exploit multi-label dependency and multimodal semantic structure. Extensive empirical results on several benchmarks demonstrate the superiority of the proposed method over existing state-of-the-arts.
Xiaobo Shen 0001, Yinfan Chen, Weiwei Liu 0003, Yuhui Zheng, Quan-Sen Sun, Shirui Pan
IEEE Trans. Neural Networks Learn. Syst.1
2024 Distributed Manifold Hashing for Image Set Classification and Retrieval
abstract
Conventional image set methods typically learn from image sets stored in one location. However, in real-world applications, image sets are often distributed or collected across different positions. Learning from such distributed image sets presents a challenge that has not been studied thus far. Moreover, efficiency is seldom addressed in large-scale image set applications. To fulfill these gaps, this paper proposes Distributed Manifold Hashing (DMH), which models distributed image sets as a connected graph. DMH employs Riemannian manifold to effectively represent each image set and further suggests learning hash code for each image set to achieve efficient computation and storage. DMH is formally formulated as a distributed learning problem with local consistency constraint on global variables among neighbor nodes, and can be optimized in parallel. Extensive experiments on three benchmark datasets demonstrate that DMH achieves highly competitive accuracies in a distributed setting and provides faster classification and retrieval than state-of-the-arts.
Xiaobo Shen 0001, Peizhuo Song, Yun-Hao Yuan 0001, Yuhui Zheng
AAAI1
2024 Learning Spectral Canonical ℱ-Correlation Representation for Face Super-Resolution
abstract
Face super-resolution (FSR) is a powerful technique for restoring high-resolution face images from the captured low-resolution ones with the assistance of prior information. Existing FSR methods based on explicit or implicit covariance matrices are difficult to reveal complex nonlinear relationships between features, as conventional covariance computation is essentially a linear operation process. Besides, the limited number of training samples and noise disturbance lead to the deviation of sample covariance matrices. To solve these issues, we propose a novel FSR method via using spectral canonical ℱ-correlation representation. The proposed method first defines intra-resolution and inter-resolution covariation matrices by considering the nonlinear relationship between different features, and then uses the fractional order idea to rebuild covariation matrices. The qualitative and quantitative results have validated the superiority of the proposed method.
Yun-Hao Yuan 0001, Mingzhi Hao, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICASSP6
2024 Contrastive Transformer Masked Image Hashing for Degraded Image Retrieval
Xiaobo Shen 0001, Haoyu Cai, Xiuwen Gong, Yuhui Zheng
IJCAI1
2024 Contrastive Transformer Cross-Modal Hashing for Video-Text Retrieval
Xiaobo Shen 0001, Qianxin Huang, Long Lan, Yuhui Zheng
IJCAI1
2024 Unsupervised Deep Graph Structure and Embedding Learning
Xiaobo Shen 0001, Xiuwen Gong, Shirui Pan
IJCAI1
2024 Graph Convolutional Semi-Supervised Cross-Modal Hashing
abstract
Cross-modal hashing encodes different modalities of multi-modal data into a low-dimensional Hamming space for fast cross-modal retrieval. Most existing cross-modal hashing methods heavily rely on label semantics to boost retrieval performance; however, semantics are expensive to collect in real applications. To mitigate the heavy reliance on semantics, this work proposes a new semi-supervised deep cross-modal hashing method, namely, Graph Convolutional Semi-Supervised Cross-Modal Hashing (GCSCH), which is trained with limited label supervision. The proposed GCSCH first generates pseudo-multi-labels of the unlabeled samples using the simple yet effective idea of consistency regularization and pseudo-labeling. GCSCH designs a fusion network that merges the two modalities and employs Graph Convolutional Network (GCN) to capture semantic information among ground-truth-labeled and pseudo-labeled multi-modal data. Using the idea of knowledge distillation, GCSCH employs a teacher-student learning scheme that can successfully transfer knowledge from the fusion module to the image and text hashing networks. Empirical studies on three multi-modal benchmark datasets demonstrate the superiority of the proposed GCSCH over state-of-the-art cross-modal hashing methods with limited label supervision.
Xiaobo Shen 0001, Gaoyao Yu, Yinfan Chen, Xichen Yang, Yuhui Zheng
ACM Multimedia1
2024 Similarity Preserving Transformer Cross-Modal Hashing for Video-Text Retrieval
abstract
As social networks grow exponentially, there is an increasing demand for video retrieval using natural language. Cross-modal hashing that encodes multi-modal data using compact hash code has been widely used in large-scale image-text retrieval, primarily due to its computation and storage efficiency. When applied to video-text retrieval, existing unsupervised cross-modal hashing extracts the frame- or word-level features individually, and thus ignores long-term dependencies. In addition, effectively exploiting the multi-modal structure is a remarkable challenge owing to the complex nature of video and text. To address the above issues, we propose Similarity Preserving Transformer Cross-Modal Hashing (SPTCH), a new unsupervised deep cross-modal hashing method for video-text retrieval. SPTCH encodes video and text by bidirectional transformer encoder that exploits their long-term dependencies. SPTCH constructs a multi-modal collaborative graph to model correlations among multi-modal data, and applies semantic aggregation by employing Graph Convolutional Network (GCN) on such graph. SPTCH designs unsupervised multi-modal contrastive loss and neighborhood reconstruction loss to effectively leverage inter- and intra-modal similarity structure among videos and texts. The empirical results on three video benchmark datasets illustrate that the proposed SPTCH generally outperforms state-of-the-arts in video-text retrieval.
Qianxin Huang, Siyao Peng, Xiaobo Shen 0001, Yun-Hao Yuan 0001, Shirui Pan
ACM Multimedia3
2024 Learning adaptive Grassmann neighbors for image-set analysis
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Zhenwen Ren
Expert Syst. Appl.2
2024 Two-step affinity matrix learning for multi-view subspace clustering
Tao Zhang 0015, Yun-Hao Yuan 0001, Xiaobo Shen 0001, Fan Liu 0003
Expert Syst. Appl.3
2024 Linguistic steganalysis via multi-task with crossing generative-natural domain
Huiqing You, Lingyun Xiang, Chunfang Yang, Xiaobo Shen 0001
Neurocomputing4
2024 Query-Guided Prototype Optimization for Few-Shot Classification
abstract
With limited labeled samples, few-shot classification poses a challenge to standard deep models and has attracted a surge of concern. Metric learning based approaches stand out for the minimalist and efficient design, aiming to classify query samples using supervision of support sets in the metric space. Prototypical Network has pioneered the use of mean feature embeddings to represent each support class and leaned on the computed prototypes for query classification. However, inherent bias exists in the mean prototypes generating from scarce support samples versus the actual class prototypes, which induces subsequent inference deviation. In this paper, we propose to diminish the bias by leveraging the semantic information of query samples to guide prototype optimization. Specifically, we exploit the semantic correlation between the local of initial mean prototypes and the global of query samples to generate query-guided masks, thus tailoring optimized prototypes that vary by query samples. This exploration of correlation is first utilized to alleviate the prototype bias problem and shows great brevity compared to existing methods. Extensive experiments are conducted on three few-shot image classification benchmark datasets, and demonstrate the effectiveness of our proposed method.
Yinghui Sun, Xiaobo Shen 0001, Quan-Sen Sun
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multiple Riemannian Kernel Hashing for Large-Scale Image Set Classification and Retrieval
abstract
Conventional image set methods typically learn from small to medium-sized image set datasets. However, when applied to large-scale image set applications such as classification and retrieval, they face two primary challenges: 1) effectively modeling complex image sets; and 2) efficiently performing tasks. To address the above issues, we propose a novel Multiple Riemannian Kernel Hashing (MRKH) method that leverages the powerful capabilities of Riemannian manifold and Hashing on effective and efficient image set representation. MRKH considers multiple heterogeneous Riemannian manifolds to represent each image set. It introduces a multiple kernel learning framework designed to effectively combine statistics from multiple manifolds, and constructs kernels by selecting a small set of anchor points, enabling efficient scalability for large-scale applications. In addition, MRKH further exploits inter- and intra-modal semantic structure to enhance discrimination. Instead of employing continuous feature to represent each image set, MRKH suggests learning hash code for each image set, thereby achieving efficient computation and storage. We present an iterative algorithm with theoretical convergence guarantee to optimize MRKH, and the computational complexity is linear with the size of dataset. Extensive experiments on five image set benchmark datasets including three large-scale ones demonstrate the proposed method outperforms state-of-the-arts in accuracy and efficiency particularly in large-scale image set classification and retrieval.
Xiaobo Shen 0001, Xiaxin Wang, Yuhui Zheng
IEEE Trans. Image Process.1
2023 Graph Convolutional Incomplete Multi-modal Hashing
abstract
Multi-modal hashing (MMH) encodes multi-modal data into latent hash code, and has been widely applied for efficient large-scale multi-modal retrieval. In practice it is common that multi-modal data is often corrupted with missing modalities, e.g., social image often lacks its tags in image-text retrieval. Conventional MMHs can only learn on complete modalities, which however wastes a considerable amount of collected data. To fulfill this gap, this paper proposes Graph Convolutional Incomplete Multi-modal Hashing (GCIMH) to learn hash code on incomplete multi-modal data. GCIMH develops Graph Convolutional Autoencoder to reconstruct incomplete multi-modal data with effective exploit of its semantic structure. GCIMH further develops multi-modal and label networks to encode multiple modalities and label respectively. GCIMH can successfully transfer knowledge of autoencoder and label network to multi-modal hashing network using teacher-student learning framework. GCIMH can handle missing modalities in both offline training and online query stages. Extensive empirical studies on three benchmark datasets demonstrate the superiority of the proposed GCIMH over the state-of-the-arts on both complete and incomplete multi-modal retrieval.
Xiaobo Shen 0001, Yinfan Chen, Shirui Pan, Weiwei Liu 0003, Yuhui Zheng
ACM Multimedia1
2023 Spatiotemporal consistent selection-correction network for deep interactive image segmentation
Tao Wang 0020, Zexuan Ji, Peng Fu 0003, Xiaobo Shen 0001, Quan-Sen Sun
Neural Comput. Appl.5
2023 Compact network embedding for fast node classification
Xiaobo Shen 0001, Yew-Soon Ong, Zheng Mao, Shirui Pan, Weiwei Liu 0003, Yuhui Zheng
Pattern Recognit.1
2023 Efficient Feature Reconstruction via l2,1-Norm Regularization for Few-Shot Classification
abstract
Few-shot classification, which aims to recognize novel classes with the aid of learned knowledge and a few labeled samples, remains challenging and draws emerging concerns in computer vision. Metric-based methods incorporate the idea of metric learning into tackling the few-shot classification problem, i.e., the class membership of query samples can be determined in the latent space. Several recent methods propose to infer such memberships by feature reconstruction, which, however either exploit attention mechanisms or closed-form solutions. In this paper, we revisit the essence of feature reconstruction applied to few-shot classification and introduce a$l_{2,1}$-norm regularization to constraint it. As with the intention of attention mechanisms, we attempt to guide feature reconstruction towards focusing more on semantically rich target regions and diminishing the contribution of profitless features. We thereby obtain more discriminative query reconstructions for each class and perform classification based on these reconstruction errors. We formulate the proposal as a constrained optimization problem and design a simple and efficient iterative method as a solution. Furthermore, we provide the analysis on the convergence of proposed iterative method in theory. We conduct few-shot classification on two fine-grained and two general datasets. Extensive experimental results reveal that our method achieves new state-of-the-art performance.
Xiaobo Shen 0001, Quan-Sen Sun
IEEE Trans. Circuits Syst. Video Technol.2
2023 Discriminative Geometric-Structure-Based Deep Hashing for Large-Scale Image Retrieval
abstract
Deep hashing reaps the benefits of deep learning and hashing technology, and has become the mainstream of large-scale image retrieval. It generally encodes image into hash code with feature similarity preserving, that is, geometric-structure preservation, and achieves promising retrieval results. In this article, we find that existing geometric-structure preservation manner inadequately ensures feature discrimination, while improving feature discrimination of hash code essentially determines hash learning retrieval performance. This fact principally spurs us to propose a discriminative geometric-structure-based deep hashing method (DGDH), which investigates three novel loss terms based on class centers to induce the so-called discriminative geometrical structure. In detail, the margin-aware center loss assembles samples in the same class to the corresponding class centers for intraclass compactness, then a linear classifier based on class center serves to boost interclass separability, and the radius loss further puts different class centers on a hypersphere to tentatively reduce quantization errors. An efficient alternate optimization algorithm with guaranteed desirable convergence is proposed to optimize DGDH. We theoretically analyze the robustness and generalization of the proposed method. The experiments on five popular benchmark datasets demonstrate superior image retrieval performance of the proposed DGDH over several state of the arts.
Guohua Dong, Xiang Zhang 0008, Xiaobo Shen 0001, Long Lan, Zhigang Luo, Xiaomin Ying
IEEE Trans. Cybern.3
2023 Contrastive Transformer Hashing for Compact Video Representation
abstract
Video hashing learns compact representation by mapping video into low-dimensional Hamming space and has achieved promising performance in large-scale video retrieval. It is challenging to effectively exploit temporal and spatial structure in an unsupervised setting. To fulfill this gap, this paper proposes Contrastive Transformer Hashing (CTH) for effective video retrieval. Specifically, CTH develops a bidirectional transformer autoencoder, based on which visual reconstruction loss is proposed. CTH is more powerful to capture bidirectional correlations among frames than conventional unidirectional models. In addition, CTH devises multi-modality contrastive loss to reveal intrinsic structure among videos. CTH constructs inter-modality and intra-modality triplet sets and proposes multi-modality contrastive loss to exploit inter-modality and intra-modality similarities simultaneously. We perform video retrieval tasks on four benchmark datasets, i.e., UCF101, HMDB51, SVW30, FCVID using the learned compact hash representation, and extensive empirical results demonstrate the proposed CTH outperforms several state-of-the-art video hashing methods.
Xiaobo Shen 0001, Yun-Hao Yuan 0001, Xichen Yang, Long Lan, Yuhui Zheng
IEEE Trans. Image Process.1
2023 Hierarchical Hashing Learning for Image Set Classification
abstract
With the development of video network, image set classification (ISC) has received a lot of attention and can be used for various practical applications, such as video based recognition, action recognition, and so on. Although the existing ISC methods have obtained promising performance, they often have extreme high complexity. Due to the superiority in storage space and complexity cost, learning to hash becomes a powerful solution scheme. However, existing hashing methods often ignore complex structural information and hierarchical semantics of the original features. They usually adopt a single-layer hashing strategy to transform high-dimensional data into short-length binary codes in one step. This sudden drop of dimension could result in the loss of advantageous discriminative information. In addition, they do not take full advantage of intrinsic semantic knowledge from whole gallery sets. To tackle these problems, in this paper, we propose a novel Hierarchical Hashing Learning (HHL) for ISC. Specifically, a coarse-to-fine hierarchical hashing scheme is proposed that utilizes a two-layer hash function to gradually refine the beneficial discriminative information in a layer-wise fashion. Besides, to alleviate the effects of redundant and corrupted features, we impose the $\ell _{2,1}$ norm on the layer-wise hash function. Moreover, we adopt a bidirectional semantic representation with the orthogonal constraint to keep intrinsic semantic information of all samples in whole image sets adequately. Comprehensive experiments demonstrate HHL acquires significant improvements in accuracy and running time. We will release the demo code on https://github.com/sunyuan-cs.
Yuan Sun 0016, Xu Wang 0028, Dezhong Peng, Zhenwen Ren, Xiaobo Shen 0001
IEEE Trans. Image Process.5
2023 Sparse Representation Classifier Guided Grassmann Reconstruction Metric Learning With Applications to Image Set Analysis
abstract
Modeling a sequence of video frames as a linear subspace on Grassmann manifold has recently become increasingly attractive in multiple computer vision applications. The success of such algorithms largely depends on a good distance measure, and learning an appropriate metric on Grassmann manifold remains a key challenge. Existing works address this by learning a discriminative mapping from the original Grassmann manifold to Hilbert space or a lower-dimensional, more discriminative Grassmann manifold. However, these approaches always highly rely on nearest neighbor matching of samples on the projected space, which is sensitive to noises and errors. Different from them, this paper proposes a Grassmann Reconstruction Metric Learning (GRML) algorithm guided by sparse representation-based classifier (SRC) for image set classification. SRC selects the coefficients associated with each class to reconstruct training samples, and then we employ it as a criterion to direct the design of a discriminant metric on Grassmann manifold. Specifically, GRML attempts to jointly maximize the inter-class reconstruction residual and minimize the intra-class reconstruction residual in the lower but more discriminative Grassmann manifold. To further explore the intrinsic geometry distance, we present a Grassmann Reconstruction Multiple Kernel Metric Learning (GRMKML) algorithm, which aims to jointly learn a metric and the corresponding kernel from a family of kernels for Grassmann manifold. Extensive experiments on eight benchmark datasets demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods.
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Zhenwen Ren
IEEE Trans. Multim.2
2022 Learning Canonical F-Correlation Projection for Compact Multiview Representation
abstract
Canonical correlation analysis (CCA) matters in multi-view representation learning. But, CCA and its most variants are essentially based on explicit or implicit covariance matrices. It means that they have no ability to model the nonlinear relationship among features due to intrinsic linearity of covariance. In this paper, we address the preceding problem and propose a novel canonical F-correlation framework by exploring and exploiting the nonlinear relationship between different features. The framework projects each feature rather than observation into a certain new space by an arbitrary nonlinear mapping, thus resulting in more flexibility in real applications. With this frame-work as a tool, we propose a correlative covariation projection (CCP) method by using an explicit nonlinear mapping. Moreover, we further propose a multiset version of CCP dubbed MCCP for learning compact representation of more than two views. The proposed MCCP is solved by an iterative method, and we prove the convergence of this iteration. A series of experimental results on six benchmark datasets demonstrate the effectiveness of our proposed CCP and MCCP methods.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Jianping Gou
CVPR6
2022 Online unsupervised cross-view discrete hashing for large-scale retrieval
Yun-Hao Yuan 0001, Shirui Pan, Xiaobo Shen 0001
Appl. Intell.5
2022 Multiview Graph Convolutional Hashing for Multisource Remote Sensing Image Retrieval
abstract
Recently, hashing has been successfully applied for large-scale remote sensing image retrieval (LSRSIR) due to its advantage in terms of computation and storage. In LSRSIR, existing hashing methods mainly focus on single-source remotely sensed data. They cannot effectively fuse multisource remotely sensed data, which has a large potential for LSRSIR. To fulfill this gap, this letter proposes a novel deep hashing method, dubbed Multiview Graph Convolutional Hashing (MGCH) that can successfully fuse multisource remote sensing image. Since graph convolutional network (GCN) has been applied as an effective means that expresses and integrates relationships into features, MGCH applies a GCN to explore inherent structural similarity among multiview data, which will help to generate discriminative hash codes. An asymmetric scheme is developed that optimizes the proposed deep model in an end-to-end manner to improve training efficiency. We evaluate the proposed method by fusing two different kinds of RS images, i.e., multispectral (MUL) image and panchromatic (PAN) image. The experimental results on the dual-source RS image data set (DSRSID) show that the proposed MGCH outperforms state-of-the-art multiview hashing methods.
Xiaobo Shen 0001, Peng Fu 0003, Zexuan Ji, Tao Wang 0020
IEEE Geosci. Remote. Sens. Lett.2
2022 The Emerging Trends of Multi-Label Learning
abstract
Exabytes of data are generated daily by humans, leading to the growing needs for new efforts in dealing with the grand challenges for multi-label learning brought by big data. For example, extreme multi-label classification is an active and rapidly growing research area that deals with classification tasks with extremely large number of classes or labels; utilizing massive data with limited supervision to build a multi-label classification model becomes valuable for practical applications, etc. Besides these, there are tremendous efforts on how to harvest the strong learning capability of deep learning to better capture the label dependencies in multi-label learning, which is the key for deep learning to address real-world classification tasks. However, it is noted that there have been a lack of systemic studies that focus explicitly on analyzing the emerging trends and new challenges of multi-label learning in the era of big data. It is imperative to call for a comprehensive survey to fulfil this mission and delineate future research directions and new applications.
Weiwei Liu 0003, Haobo Wang 0001, Xiaobo Shen 0001, Ivor W. Tsang
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Neighborhood preserving embedding on Grassmann manifold for image-set analysis
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Zhenwen Ren
Pattern Recognit.2
2022 Unsupervised Multiview Distributed Hashing for Large-Scale Retrieval
abstract
Multi-view hashing (MvH) learns compact hash code by efficiently integrating multi-view data, and has achieved promising performance in large-scale retrieval task. In real-world applications, multi-view data is often stored or collected in different locations, and learning hash code in such case is more challenging yet less studied. In addition, unsupervised MvHs hardly achieve impressive retrieval performance due to absence of supervision. To fulfill this gap, this paper introduces a novel unsupervised multi-view distributed hashing (UMvDisH) to learn hash code from multi-view data, which is distributed in different nodes of a network. UMvDisH jointly performs latent factor model and spectral clustering to generate latent hash code and pseudo label respectively in each node. The consistency between hash code and pseudo label improves discrimination of hash code. The proposed distributed learning problem is divided into a set of decentralized subproblems by imposing local consistency among neighbor nodes. As such, the subproblems can be solved in parallel, and training time can be reduced. The communication cost is low due to no exchange of training data. Experimental results on four benchmark image datasets including a very large-scale image dataset show that UMvDisH achieves comparable retrieval performance and trains faster than state-of-the-art unsupervised MvHs in the distributed setting.
Xiaobo Shen 0001, Yunpeng Tang, Yuhui Zheng, Yun-Hao Yuan 0001, Quan-Sen Sun
IEEE Trans. Circuits Syst. Video Technol.1
2022 Scalable Gaussian Process Classification With Additive Noise for Non-Gaussian Likelihoods
abstract
Gaussian process classification (GPC) provides a flexible and powerful statistical framework describing joint distributions over function space. Conventional GPCs, however, suffer from: 1) poor scalability for big data due to the full kernel matrix and 2) intractable inference due to the non-Gaussian likelihoods. Hence, various scalable GPCs have been proposed through: 1) the sparse approximation built upon a small inducing set to reduce the time complexity and 2) the approximate inference to derive analytical evidence lower bound (ELBO). However, these scalable GPCs equipped with analytical ELBO are limited to specific likelihoods or additional assumptions. In this work, we present a unifying framework that accommodates scalable GPCs using various likelihoods. Analogous to GP regression (GPR), we introduce additive noises to augment the probability space for: 1) the GPCs with step, (multinomial) probit, and logit likelihoods via the internal variables and 2) particularly, the GPC using softmax likelihood via the noise variables themselves. This leads to unified scalable GPCs with analytical ELBO by using variational inference. Empirically, our GPCs showcase superiority on extensive binary/multiclass classification tasks with up to two million data points.
Haitao Liu 0002, Yew-Soon Ong, Ziwei Yu, Jianfei Cai 0001, Xiaobo Shen 0001
IEEE Trans. Cybern.5
2022 Hyperspectral Image Few-Shot Classification Network Based on the Earth Mover's Distance
abstract
Deep learning has achieved promising performance in hyperspectral image (HSI) classification. Training deep models usually requires labeling massive HSIs, which however is prohibitively time-consuming and expensive. To fill in the gap, this paper proposes a novel meta-learning method for HSI few-shot classification that conducts HSI classification with a few labeled samples. Specifically, we introduce the Earth Mover’s Distance (EMD) as a metric. The designed EMD metric learning module aims to calculate the similarity of paired embedding features by decomposing embedding features into a set of local representations. The EMD metric aims to find the optimal matching flows between local representations that have the minimum matching cost. Furthermore, we attempt to learn class prototype representation for each hyperspectral class using the EMD metric. The proposed network effectively learns general knowledge from base HSIs and transfers such knowledge to the classification of novel HSIs. We conduct HSI few-shot classification by training on three base HSIs and classification on three novel HSIs. Extensive experimental results on three novel HSI datasets demonstrate that the proposed model outperforms existing state-of-the-art HSI methods, including two HSI few-shot methods.
Xiaobo Shen 0001, Quan-Sen Sun
IEEE Trans. Geosci. Remote. Sens.2
2022 Discrete Metric Learning for Fast Image Set Classification
abstract
In the field of image set classification, most existing works focus on exploiting effective latent discriminative features. However, it remains a research gap to efficiently handle this problem. In this paper, benefiting from the superiority of hashing in terms of its computational complexity and memory costs, we present a novel Discrete Metric Learning (DML) approach based on the Riemannian manifold for fast image set classification. The proposed DML jointly learns a metric in the induced space and a compact Hamming space, where efficient classification is carried out. Specifically, each image set is modeled as a point on Riemannian manifold after which the proposed DML minimizes the Hamming distance between similar Riemannian pairs and maximizes the Hamming distance between dissimilar ones by introducing a discriminative Mahalanobis-like matrix. To overcome the shortcoming of DML that relies on the vectorization of Riemannian representations, we further develop Bilinear Discrete Metric Learning (BDML) to directly manipulate the original Riemannian representations and explore the natural matrix structure for high-dimensional data. Different from conventional Riemannian metric learning methods, which require complicated Riemannian optimizations (e.g., Riemannian conjugate gradient), both DML and BDML can be efficiently optimized by computing the geodesic mean between the similarity matrix and inverse of the dissimilarity matrix. Extensive experiments conducted on different visual recognition tasks (face recognition, object recognition, and action recognition) demonstrate that the proposed methods achieve competitive performance in terms of accuracy and efficiency.
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao
IEEE Trans. Image Process.2
2022 Deep Co-Image-Label Hashing for Multi-Label Image Retrieval
abstract
Deep supervised hashing has greatly improved retrieval performance with the powerful learning capability of deep neural network. In multi-label image retrieval, existing deep hashing simply indicates whether two images are similar by constructing a similarity matrix. However, it ignores the dependency among multiple labels that has been shown important in multi-label application. To fulfill this gap, this paper proposes Deep Co-Image-Label Hashing (DCILH) to discover label dependency. Specifically, DCILH regards image and label as two views, and maps the two views into a common deep Hamming space. DCILH proposes to learn prototype for each label, and preserve similarity among images, labels, and prototypes. To exploit label dependency, DCILH further employs the label-correlation aware loss on the predicted labels, such that predicted output on positive label is enforced to be larger than that on negative label. Extensive experiments on several multi-label benchmarks demonstrate the proposed DCILH outperforms state-of-the-art deep supervised hashing on large-scale multi-label image retrieval.
Xiaobo Shen 0001, Guohua Dong, Yuhui Zheng, Long Lan, Ivor W. Tsang, Quan-Sen Sun
IEEE Trans. Multim.1
2021 Multi-view Fractional Deep Canonical Correlation Analysis for Subspace Clustering
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICONIP (2)6
2021 Composite nonlinear multiset canonical correlation analysis for multiview feature learning and recognition
abstract
Summary In this paper, we propose a composite nonlinear multiset canonical correlation projections (CNMCPs) framework where orthogonal constraints are imposed in each set. This makes CNMCP capable of learning uncorrelated low‐dimensional features with minimum redundancy in Hilbert space. With the CNMCP framework, we further present a particular algorithm called multikernel multiset canonical correlations or mKMCC, which introduces different weights into multiple nonlinear functions in all views. An alternating iterative optimization is designed for computational solution. Numerous experimental results on practical datasets have demonstrated the effectiveness and robustness of mKMCC, in contrast with existing kernel correlation learning approaches.
Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010, Bin Li 0006, Jianping Gou, Jipeng Qiang, Xinfeng Zhang 0003, Quan-Sen Sun
Concurr. Comput. Pract. Exp.2
2021 Adaptive graph guided concept factorization on Grassmann manifold
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Zhenwen Ren
Inf. Sci.2
2021 Label group diffusion for image and image pair segmentation
Tao Wang 0020, Zexuan Ji, Jian Yang 0003, Quan-Sen Sun, Xiaobo Shen 0001, Zhenwen Ren, Qi Ge
Pattern Recognit.5
2021 Unsupervised Discriminative Deep Hashing With Locality and Globality Preservation
abstract
Deep hashing has greatly improved retrieval performance with the powerful learning capability of deep neural network. However, deep unsupervised hashing can hardly achieve impressive performance due to the lack of the semantic supervision. This letter proposes Unsupervised Discriminative Deep Hashing (UD2H) to fulfill this gap. UD2H is formulated to jointly perform hash code learning and clustering, and trained in an asymmetric manner to improve the efficiency. The cluster labels supervise the training of deep model to enable hash code discriminative. Based on the outputs of the deep model, UD2H adaptively constructs a similarity graph that considers the local and global structures. Experiments on three benchmark datasets show that the proposed UD$^2$H outperforms the state-of-the-art unsupervised deep hashing methods.
Zhuyi Ni, Zexuan Ji, Long Lan, Yun-Hao Yuan 0001, Xiaobo Shen 0001
IEEE Signal Process. Lett.5
2020 Supervised Multi-View Distributed Hashing
abstract
Multi-view hashing efficiently integrates multi-view data for learning compact hash codes, and achieves impressive large-scale retrieval performance. In real-world applications, multi-view data are often stored or collected in different locations, where hash code learning is more challenging yet less studied. To fulfill this gap, this paper proposes a novel supervised multi-view distributed hashing (SMvDisH) for hash code learning from multi-view data in a distributed manner. SMvDisH yields the discriminative latent hash codes by joint learning of latent factor model and classifier. With local consistency assumption among neighbor nodes, the distributed learning problem is divided into a set of decentralized sub-problems. The sub-problems can be solved in parallel, and the computational and communication costs are low. Experimental results on three large-scale image datasets demonstrate that SMvDisH achieves competitive retrieval performance and trains faster than state-of-the-art multi-view hashing methods.
Yunpeng Tang, Xiaobo Shen 0001, Zexuan Ji, Tao Wang 0020, Peng Fu 0003, Quan-Sen Sun
ICIP2
2020 Complex-Valued Spatial-Scattering Separated Attention Network for Polsar Image Classification
abstract
Fully polarimetric synthetic aperture radar (PolSAR) images are generally expressed as the complex-valued (CV) matrix, whereas the convolutional neural networks (CNNs) have been successfully utilized for the PolSAR image classification. However, most 2D or 3D CNNs suffer from insufficient exploring the CV features or computationally expensive. To address these issues, this paper presents an efficient complex-valued spatial-scattering separated attention network (CVS3ANet) for PolSAR image classification. The CVS3ANet utilizes CV-3D convolutions to explore the features in both spatial and scattering dimensions of the PolSAR images, and reduces the parameters by factorizing the 3D convolution as a sequential process of 2D spatial convolution followed by 1D scattering convolution. Moreover, a squeeze and fusion attention unit is used to enhance the learning interpretation ability of the network by modeling correlations between channels with respect to attention probability. The experimental results demonstrate that the proposed method can obtain superior results over the state-of-the-art techniques.
Zhaohao Fan, Zexuan Ji, Peng Fu 0003, Tao Wang 0020, Xiaobo Shen 0001, Quan-Sen Sun
IGARSS5
2020 Discriminant sub-dictionary learning with adaptive multiscale superpixel representation for hyperspectral image classification
Xiao Tu, Xiaobo Shen 0001, Peng Fu 0003, Tao Wang 0020, Quan-Sen Sun, Zexuan Ji
Neurocomputing2
2020 Locality-aware group sparse coding on Grassmann manifolds for image set classification
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Wenzhu Yan
Neurocomputing2
2020 Prototype learning and collaborative representation using Grassmann manifolds for image set classification
Dong Wei 0007, Xiaobo Shen 0001, Quan-Sen Sun, Xizhan Gao, Wenzhu Yan
Pattern Recognit.2
2020 Guest Editorial Special Issue on Structured Multi-Output Learning: Modeling, Algorithm, Theory, and Applications
abstract
Structured multioutput learning is a topic in artificial intelligence that considers multiple structured outputs prediction for a given input. The output may involve structured objects in the form of sequence, string, tree, lattice, or graph and has values that are characterized by diverse data types, such as binary, nominal, ordinal, and real-valued variables. Such learning problems arise in a variety of real-world applications, ranging from document classification, computer emulation, sensor network analysis, concept-based information retrieval, and human action/causal induction to video analysis, image annotation/retrieval, gene function prediction, and brain science. As many complex real-world scenarios can be posed as a structured multioutput learning problem, their importance and popularity have been increasing steadily.
Weiwei Liu 0003, Xiaobo Shen 0001, Yew-Soon Ong, Ivor W. Tsang, Chen Gong 0002, Vladimir Pavlovic 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 When Gaussian Process Meets Big Data: A Review of Scalable GPs
abstract
The vast quantity of information brought by big data as well as the evolving computer hardware encourages success stories in the machine learning community. In the meanwhile, it poses challenges for the Gaussian process regression (GPR), a well-known nonparametric, and interpretable Bayesian model, which suffers from cubic complexity to data size. To improve the scalability while retaining desirable prediction quality, a variety of scalable GPs have been presented. However, they have not yet been comprehensively reviewed and analyzed to be well understood by both academia and industry. The review of scalable GPs in the GP community is timely and important due to the explosion of data size. To this end, this article is devoted to reviewing state-of-the-art scalable GPs involving two main categories: global approximations that distillate the entire data and local approximations that divide the data for subspace learning. Particularly, for global approximations, we mainly focus on sparse approximations comprising prior approximations that modify the prior but perform exact inference, posterior approximations that retain exact prior but perform approximate inference, and structured sparse approximations that exploit specific structures in kernel matrix; for local approximations, we highlight the mixture/product of experts that conducts model averaging from multiple local experts to boost predictions. To present a complete review, recent advances for improving the scalability and capability of scalable GPs are reviewed. Finally, the extensions and open issues of scalable GPs in various scenarios are reviewed and discussed to inspire novel ideas for future research avenues.
Haitao Liu 0002, Yew-Soon Ong, Xiaobo Shen 0001, Jianfei Cai 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 Survey on Multi-Output Learning
abstract
The aim of multi-output learning is to simultaneously predict multiple outputs given an input. It is an important learning problem for decision-making since making decisions in the real world often involves multiple complex factors and criteria. In recent times, an increasing number of research studies have focused on ways to predict multiple outputs at once. Such efforts have transpired in different forms according to the particular multi-output learning problem under study. Classic cases of multi-output learning include multi-label learning, multi-dimensional learning, multi-target regression, and others. From our survey of the topic, we were struck by a lack in studies that generalize the different forms of multi-output learning into a common framework. This article fills that gap with a comprehensive review and analysis of the multi-output learning paradigm. In particular, we characterize the four Vs of multi-output learning, i.e., volume, velocity, variety, and veracity, and the ways in which the four Vs both benefit and bring challenges to multi-output learning by taking inspiration from big data. We analyze the life cycle of output labeling, present the main mathematical definitions of multi-output learning, and examine the field's key challenges and corresponding solutions as found in the literature. Several model evaluation metrics and popular data repositories are also discussed. Last but not least, we highlight some emerging challenges with multi-output learning from the perspective of the four Vs as potential research directions worthy of further studies.
Donna Xu, Yaxin Shi, Ivor W. Tsang, Yew-Soon Ong, Chen Gong 0002, Xiaobo Shen 0001
IEEE Trans. Neural Networks Learn. Syst.6
2019 Memetic Evolution Strategy for Reinforcement Learning
abstract
Neuroevolution (i.e., training neural network with Evolution Computation) has successfully unfolded a range of challenging reinforcement learning (RL) tasks. However, existing neuroevolution methods suffer from high sample complexity, as the black-box evaluations (i.e., accumulated rewards of complete Markov Decision Processes (MDPs)) discard bunches of temporal frames (i.e., time-step data instances in MDP). Actually, these temporal frames hold the Markov property of the problem, that benefits the training of neural network as well by temporal difference (TD) learning. In this paper, we propose a memetic reinforcement learning (MRL) framework that optimizes the RL agent by leveraging both black-box evaluations and temporal frames. To this end, an evolution strategy (ES) is associated with Q learning, where ES provides diversified frames to globally train the agent, and Q learning locally exploits the Markov property within frames to refresh the agent. Therefore, MRL conveys a novel memetic framework that allows evaluation free local search by Q learning. Experiments on classical control problem verify the efficiency of the proposed MRL, that achieves significantly faster convergence than canonical ES.
Xinghua Qu, Yew-Soon Ong, Yaqing Hou, Xiaobo Shen 0001
CEC4
2019 Sparse Extreme Multi-label Learning with Oracle Property
abstract
The pioneering work of sparse local embeddings for extreme classification (SLEEC) (Bhatia et al., 2015) has shown great promise in multi-label learning. Unfortunately, the statistical rate of convergence and oracle property of SLEEC are still not well understood. To fill this gap, we present a unified framework for SLEEC with nonconvex penalty. Theoretically, we rigorously prove that our proposed estimator enjoys oracle property (i.e., performs as well as if the underlying model were known beforehand), and obtains a desirable statistical convergence rate. Moreover, we show that under a mild condition on the magnitude of the entries in the underlying model, we are able to obtain an improved convergence rate. Extensive numerical experiments verify our theoretical findings and the superiority of our proposed estimator.
Weiwei Liu 0003, Xiaobo Shen 0001
ICML2
2019 Discrete Multi-graph Hashing for Large-Scale Visual Search
Lingyun Xiang, Xiaobo Shen 0001, Jiaohua Qin, Wei Hao 0002
Neural Process. Lett.2
2019 Hyperspectral Imagery Classification via Stochastic HHSVMs
abstract
Hyperspectral imagery (HSI) has shown promising results in real-world applications. However, the technological evolution of optical sensors poses two main challenges in HSI classification: 1) the spectral band is usually redundant and noisy and 2) HSI with millions of pixels has become increasingly common in real-world applications. Motivated by the recent success of hybrid huberized support vector machines (HHSVMs), which inherit the benefits of both lasso and ridge regression, this paper first investigates the advantages of HHSVM for HSI applications. Unfortunately, the existing HHSVM solvers suffer from prohibitive computational costs on large-scale data sets. To solve this problem, this paper proposes simple and effective stochastic HHSVM algorithms for HSI classification. In the stochastic settings, we show that with a probability of at least , our algorithms find an -accurate solution using iterations. Since the convergence rate of our algorithms does not depend on the size of the training set, our algorithms are suitable for handling large-scale problems. We demonstrate the superiority of our algorithms by conducting experiments on large-scale binary and multiclass classification problems, comparing to the state-of-the-art HHSVM solvers. Finally, we apply our algorithms to real HSI classification and achieve promising results.
Weiwei Liu 0003, Xiaobo Shen 0001, Bo Du 0001, Ivor W. Tsang, Wenjie Zhang 0001, Xuemin Lin 0001
IEEE Trans. Image Process.2
2018 Compact Multi-Label Learning
abstract
Embedding methods have shown promising performance in multi-label prediction, as they can discover the dependency of labels. Most embedding methods cannot well align the input and output, which leads to degradation in prediction performance. Besides, they suffer from expensive prediction computational costs when applied to large-scale datasets. To address the above issues, this paper proposes a Co-Hashing (CoH) method by formulating multi-label learning from the perspective of cross-view learning. CoH first regards the input and output as two views, and then aims to learn a common latent hamming space, where input and output pairs are compressed into compact binary embeddings. CoH enjoys two key benefits: 1) the input and output can be well aligned, and their correlations are explored; 2) the prediction is very efficient using fast cross-view kNN search in the hamming space. Moreover, we provide the generalization error bound for our method. Extensive experiments on eight real-world datasets demonstrate the superiority of the proposed CoH over the state-of-the-art methods in terms of both prediction accuracy and efficiency.
Xiaobo Shen 0001, Weiwei Liu 0003, Ivor W. Tsang, Quan-Sen Sun, Yew-Soon Ong
AAAI1
2018 Low Resolution Face Recognition and Reconstruction Via Deep Canonical Correlation Analysis
abstract
Low-resolution (LR) face identification is always a challenge in computer vision. In this paper, we propose a new LR face recognition and reconstruction method using deep canonical correlation analysis (DCCA). Unlike linear CCA-based methods, our proposed method can learn flexible nonlinear representations by passing LR and high-resolution (HR) image principal component features through multiple stacked layers of nonlinear transformation. As the nonlinear transformation in deep neural networks is implicit, we apply radial basis function based neural network to learn an explicit mapping between principal components and correlational features. In addition, we also design two residual compensation methods for identification and vision enhancement, respectively. The proposed approach is compared with existing LR face recognition and reconstruction algorithms. A number of experimental results on benchmark datasets have demonstrated the effectiveness and robustness of our method.
Zhao Zhang 0018, Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010
ICASSP3
2018 Learning Parallel Canonical Correlations for Scale-Adaptive Low Resolution Face Recognition
abstract
Low resolution is one of the main obstacles in the application of face recognition. Although many methods have been proposed to improve the problem, they assume that low-resolution (LR) face images have a uniform scale. In real scenarios, this prerequisite is very harsh. In this paper, we propose a scale-adaptive LR face recognition approach based on two-dimensional multi-set canonical correlation analysis (2DM-CCA), where face image matrix does not need to be previously transformed into a vector. In the proposed method, training sets with different resolutions are treated as different views, and then projected in parallel into a latent coherent space where the consistency of multi-view face data is maximally enhanced. When a new LR face image with an arbitrary scale is input, we first transform it by using the left and right projection matrices of an appropriate training view, and then reconstruct its high resolution facial feature by neighborhood reconstruction. Experimental results show that our proposed method is more effective and efficient than several existing methods.
Yun-Hao Yuan 0001, Zhao Zhang 0018, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Xiaobo Shen 0001
ICPR6
2018 Deep Discrete Prototype Multilabel Learning
abstract
kNN embedding methods, such as the state-of-the-art LM-kNN, have shown impressive results in multi-label learning. Unfortunately, these approaches suffer expensive computation and memory costs in large-scale settings. To fill this gap, this paper proposes a novel deep prototype compression, i.e., DBPC for fast multi-label prediction. DBPC compresses the database into a small set of short discrete prototypes, and uses the prototypes for prediction. The benefit of DBPC comes from two aspects: 1) The number of distance comparisons are reduced in the prototype; 2) The distance computation cost is significantly decreased in the reduced space. We propose to jointly learn the deep latent subspace and discrete prototypes within one framework. The encoding and decoding neural networks are employed to make deep discrete prototypes well represent the instances and labels. Extensive experiments on several large-scale datasets demonstrate that DBPC achieves several orders of magnitude lower storage and prediction complexity than state-of-the-art multi-label methods, while achieving competitive accuracy.
Xiaobo Shen 0001, Weiwei Liu 0003, Yong Luo 0002, Yew-Soon Ong, Ivor W. Tsang
IJCAI1
2018 Discrete Network Embedding
abstract
Network embedding aims to seek low-dimensional vector representations for network nodes, by preserving the network structure. The network embedding is typically represented in continuous vector, which imposes formidable challenges in storage and computation costs, particularly in large-scale applications. To address the issue, this paper proposes a novel discrete network embedding (DNE) for more compact representations. In particular, DNE learns short binary codes to represent each node. The Hamming similarity between two binary embeddings is then employed to well approximate the ground-truth similarity. A novel discrete multi-class classifier is also developed to expedite classification. Moreover, we propose to jointly learn the discrete embedding and classifier within a unified framework to improve the compactness and discrimination of network embedding. Extensive experiments on node classification consistently demonstrate that DNE exhibits lower storage and computational complexity than state-of-the-art network embedding methods, while obtains competitive classification results.
Xiaobo Shen 0001, Shirui Pan, Weiwei Liu 0003, Yew-Soon Ong, Quan-Sen Sun
IJCAI1
2018 A novel multi-view dimensionality reduction and recognition framework with applications to face recognition
Xiaobo Shen 0001, Yun-Hao Yuan 0001, Fumin Shen, Yang Xu 0006, Quan-Sen Sun
J. Vis. Commun. Image Represent.1
2018 Multiview Discrete Hashing for Scalable Multimedia Search
abstract
Hashing techniques have recently gained increasing research interest in multimedia studies. Most existing hashing methods only employ single features for hash code learning. Multiview data with each view corresponding to a type of feature generally provides more comprehensive information. How to efficiently integrate multiple views for learning compact hash codes still remains challenging. In this article, we propose a novel unsupervised hashing method, dubbed multiview discrete hashing (MvDH), by effectively exploring multiview data. Specifically, MvDH performs matrix factorization to generate the hash codes as the latent representations shared by multiple views, during which spectral clustering is performed simultaneously. The joint learning of hash codes and cluster labels enables that MvDH can generate more discriminative hash codes, which are optimal for classification. An efficient alternating algorithm is developed to solve the proposed optimization problem with guaranteed convergence and low computational complexity. The binary codes are optimized via the discrete cyclic coordinate descent (DCC) method to reduce the quantization errors. Extensive experimental results on three large-scale benchmark datasets demonstrate the superiority of the proposed method over several state-of-the-art methods in terms of both accuracy and scalability.
Xiaobo Shen 0001, Fumin Shen, Li Liu 0004, Yun-Hao Yuan 0001, Weiwei Liu 0003, Quan-Sen Sun
ACM Trans. Intell. Syst. Technol.1
2018 Multilabel Prediction via Cross-View Search
abstract
Embedding methods have shown promising performance in multilabel prediction, as they are able to discover the label dependence. However, most methods ignore the correlations between the input and output, such that their learned embeddings are not well aligned, which leads to degradation in prediction performance. This paper presents a formulation for multilabel learning, from the perspective of cross-view learning, that explores the correlations between the input and the output. The proposed method, called Co-Embedding (CoE), jointly learns a semantic common subspace and view-specific mappings within one framework. The semantic similarity structure among the embeddings is further preserved, ensuring that close embeddings share similar labels. Additionally, CoE conducts multilabel prediction through the cross-view $k$ nearest neighborhood ( $k$ NN) search among the learned embeddings, which significantly reduces computational costs compared with conventional decoding schemes. A hashing-based model, i.e., Co-Hashing (CoH), is further proposed. CoH is based on CoE, and imposes the binary constraint on continuous latent embeddings. CoH aims to generate compact binary representations to improve the prediction efficiency by benefiting from the efficient $k$ NN search of multiple labels in the Hamming space. Extensive experiments on various real-world data sets demonstrate the superiority of the proposed methods over the state of the arts in terms of both prediction accuracy and efficiency.
Xiaobo Shen 0001, Weiwei Liu 0003, Ivor W. Tsang, Quan-Sen Sun, Yew-Soon Ong
IEEE Trans. Neural Networks Learn. Syst.1
2017 Compressed K-Means for Large-Scale Clustering
abstract
Large-scale clustering has been widely used in many applications, and has received much attention. Most existing clustering methods suffer from both expensive computation and memory costs when applied to large-scale datasets. In this paper, we propose a novel clustering method, dubbed compressed k-means (CKM), for fast large-scale clustering. Specifically, high-dimensional data are compressed into short binary codes, which are well suited for fast clustering. CKM enjoys two key benefits: 1) storage can be significantly reduced by representing data points as binary codes; 2) distance computation is very efficient using Hamming metric between binary codes. We propose to jointly learn binary codes and clusters within one framework. Extensive experimental results on four large-scale datasets, including two million-scale datasets demonstrate that CKM outperforms the state-of-the-art large-scale clustering methods in terms of both computation and memory cost, while achieving comparable clustering accuracy.
Xiaobo Shen 0001, Weiwei Liu 0003, Ivor W. Tsang, Fumin Shen, Quan-Sen Sun
AAAI1
2017 Compact Multiple-Instance Learning
abstract
The weakly supervised Multiple-Instance Learning (MIL) problem has been successfully applied in information retrieval tasks. Two related issues might affect the performance of MIL algorithms: how to cope with label ambiguities and how to deal with non-discriminative components, and we propose COmpact MultiPle-Instance LEarning (COMPILE) to consider them simultaneously. To treat label ambiguities, COMPILE seeks ground-truth positive instances in positive bags. By using weakly supervised information to learn data's short binary representations, COMPILE enhances discrimination via strengthening discriminative components and suppressing non-discriminative ones. We adapt block coordinate descent to optimize COMPILE efficiently. Experiments on text categorization empirically show: 1) COMPILE unifies disambiguation and data preprocessing successfully; 2) it generates short binary representations efficiently to enhance discrimination at significantly reduced storage cost.
Jing Chai, Weiwei Liu 0003, Ivor W. Tsang, Xiaobo Shen 0001
CIKM4
2017 Fractional discriminative multiview correlation projection for face feature fusion
abstract
Multiple view data with different feature representations have widely arisen in various practical applications. Due to the information diversity, fusing multiview features is very valuable for classification purpose. In this paper, we propose a new multifeature fusion method called fractional-order discriminative multiview correlation projection (FDMCP), which is based on fractional-order scatter matrices with class label information of the samples. FDMCP first defines supervised covariance matrices in each view. It then constructs fractional supervised scatter matrices. Experimental results on three benchmark face image datasets show that our proposed FDMCP approach outperforms generalized multiview linear discriminant analysis.
Yun-Hao Yuan 0001, Yun Li 0010, Bin Li 0006, Hongkun Ji, Xiaobo Shen 0001
FUSION6
2017 Sparse Embedded k-Means Clustering
abstract
The $k$-means clustering algorithm is a ubiquitous tool in data mining and machine learning that shows promising performance. However, its high computational cost has hindered its applications in broad domains. Researchers have successfully addressed these obstacles with dimensionality reduction methods. Recently, [1] develop a state-of-the-art random projection (RP) method for faster $k$-means clustering. Their method delivers many improvements over other dimensionality reduction methods. For example, compared to the advanced singular value decomposition based feature extraction approach, [1] reduce the running time by a factor of $\min \{n,d\}\epsilon^2 log(d)/k$ for data matrix $X \in \mathbb{R}^{n\times d} $ with $n$ data points and $d$ features, while losing only a factor of one in approximation accuracy. Unfortunately, they still require $\mathcal{O}(\frac{ndk}{\epsilon^2log(d)})$ for matrix multiplication and this cost will be prohibitive for large values of $n$ and $d$. To break this bottleneck, we carefully build a sparse embedded $k$-means clustering algorithm which requires $\mathcal{O}(nnz(X))$ ($nnz(X)$ denotes the number of non-zeros in $X$) for fast matrix multiplication. Moreover, our proposed algorithm improves on [1]'s results for approximation accuracy by a factor of one. Our empirical studies corroborate our theoretical findings, and demonstrate that our approach is able to significantly accelerate $k$-means clustering, while achieving satisfactory clustering performance.
Weiwei Liu 0003, Xiaobo Shen 0001, Ivor W. Tsang
NIPS2
2017 Graph regularized multilayer concept factorization for data representation
Xiaobo Shen 0001, Zhenqiu Shu, Qiaolin Ye, Chunxia Zhao
Neurocomputing2
2017 Laplacian multiset canonical correlations for multiview feature extraction and image recognition
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Quan-Sen Sun, Jinlong Yang 0002
Multim. Tools Appl.3
2017 Semi-Paired Discrete Hashing: Learning Latent Hash Codes for Semi-Paired Cross-View Retrieval
abstract
Due to the significant reduction in computational cost and storage, hashing techniques have gained increasing interests in facilitating large-scale cross-view retrieval tasks. Most cross-view hashing methods are developed by assuming that data from different views are well paired, e.g., text-image pairs. In real-world applications, however, this fully-paired multiview setting may not be practical. The more practical yet challenging semi-paired cross-view retrieval problem, where pairwise correspondences are only partially provided, has less been studied. In this paper, we propose an unsupervised hashing method for semi-paired cross-view retrieval, dubbed semi-paired discrete hashing (SPDH). In specific, SPDH explores the underlying structure of the constructed common latent subspace, where both paired and unpaired samples are well aligned. To effectively preserve the similarities of semi-paired data in the latent subspace, we construct the cross-view similarity graph with the help of anchor data pairs. SPDH jointly learns the latent features and hash codes with a factorization-based coding scheme. For the formulated objective function, we devise an efficient alternating optimization algorithm, where the key binary code learning problem is solved in a bit-by-bit manner with each bit generated with a closed-form solution. The proposed method is extensively evaluated on four benchmark datasets with both fully-paired and semi-paired settings and the results demonstrate the superiority of SPDH over several other state-of-the-art methods in term of both accuracy and scalability.
Xiaobo Shen 0001, Fumin Shen, Quan-Sen Sun, Yang Yang 0002, Yun-Hao Yuan 0001, Heng Tao Shen
IEEE Trans. Cybern.1
2016 Nonnegative matrix factorization with endmember sparse graph learning for hyperspectral unmixing
abstract
Nonnegative matrix factorization (NMF) based hyperspectral unmixing aims at estimating pure spectral signatures and their fractional abundances at each pixel. During the past several years, manifold structures have been introduced as regularization constraints into NMF. However, most methods only consider the constraints on abundance matrix while ignoring the geometric relationship of endmembers. Although such relationship can be described by traditional graph construction approaches based on k-nearest neighbors, its accuracy is questionable. In this paper, we propose a novel hyperspectral unmixing method, namely NMF with endmember sparse graph learning, to tackle the above drawbacks. This method first integrates endmember sparse graph structure into NMF, then simultaneously performs unmixing and graph learning. It is further extended by incorporating abundance smoothness constraint to improve the unmixing performance. Experimental results on both synthetic and real datasets have validated the effectiveness of the proposed method.
Bin Qian 0006, Jun Zhou 0001, Xiaobo Shen 0001, Fan Liu 0003
ICIP4
2016 Semi-discriminative Multiview Canonical Correlation Analysis for Recognition
Yun-Hao Yuan 0001, Yun Li 0010, Hongkun Ji, Chong-Guang Ren, Xiaobo Shen 0001, Quan-Sen Sun
IDEAL5
2016 Fractional-Order Multiview Discriminant Analysis
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Chong-Guang Ren, Chao-Fei Li
IDEAL3
2016 Learning multi-kernel multi-view canonical correlations for image recognition
abstract
canonical correlations (M 2 CCs) framework for subspace learning. In the proposed framework, the input data of each original view are mapped into multiple higher dimensional feature spaces by multiple nonlinear mappings determined by different kernels. This makes M 2 CC can discover multiple kinds of useful information of each original view in the feature spaces. With the framework, we further provide a specific multi-view feature learning method based on direct summation kernel strategy and its regularized version. The experimental results in visual recognition tasks demonstrate the effectiveness and robustness of the proposed method.
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Guoqing Zhang 0002, Quan-Sen Sun
Comput. Vis. Media5
2016 Semi-paired hashing for cross-view retrieval
Xiaobo Shen 0001, Quan-Sen Sun, Yun-Hao Yuan 0001
Neurocomputing1
2016 Robust Cross-view Hashing for Multimedia Retrieval
abstract
Hashing techniques have been widely applied to large-scale cross-view retrieval tasks due to the significant advantage of binary codes in computation and storage efficiency. However, most existing cross-view hashing methods learn binary codes with continuous relaxations, which cause large quantization loss across views. To address this problem, in this letter, we propose a novel cross-view hashing method, where a common Hamming space is learned such that binary codes from different views are consistent and comparable. The quantization loss across views is explicitly reduced by two carefully designed regression terms from original spaces to the Hamming space. In our method, the l2,1-norm regularization is further exploited for discriminative feature selection. To obtain high-quality binary codes, we propose to jointly learn the codes and hash functions, for which an efficient iterative algorithm is presented. We evaluate the proposed method, dubbed Robust Cross-view Hashing (RCH), on two benchmark datasets and the results demonstrate the superiority of RCH over many other state-of-the-art methods in terms of retrieval performance and cross-view consistency.
Xiaobo Shen 0001, Fumin Shen, Quan-Sen Sun, Yun-Hao Yuan 0001, Heng Tao Shen
IEEE Signal Process. Lett.1
2015 Sparse Discrimination based Multiset Canonical Correlation Analysis for Multi-Feature Fusion and Recognition
Hongkun Ji, Xiaobo Shen 0001, Quan-Sen Sun, Zexuan Ji
BMVC2
2015 Multi-view Latent Hashing for Efficient Multimedia Search
abstract
Hashing techniques have attracted broad research interests in recent multimedia studies. However, most of existing hashing methods focus on learning binary codes from data with only one single view, and thus cannot fully utilize the rich information from multiple views of data. In this paper, we propose a novel unsupervised hashing approach, dubbed multi-view latent hashing (MVLH), to effectively incorporate multi-view data into hash code learning. Specifically, the binary codes are learned by the latent factors shared by multiple views from an unified kernel feature space, where the weights of different views are adaptively learned according to the reconstruction error with each view. We then propose to solve the associate optimization problem with an efficient alternating algorithm. To obtain high-quality binary codes, we provide a novel scheme to directly learn the codes without resorting to continuous relaxations, where each bit is efficiently computed in a closed form. We evaluate the proposed method on several large-scale datasets and the results demonstrate the superiority of our method over several other state-of-the-art methods.
Xiaobo Shen 0001, Fumin Shen, Quan-Sen Sun, Yun-Hao Yuan 0001
ACM Multimedia1
2015 A unified multiset canonical correlation analysis framework based on graph embedding for multiple feature extraction
Xiaobo Shen 0001, Quan-Sen Sun, Yun-Hao Yuan 0001
Neurocomputing1
2015 Orthogonal Multiset Canonical Correlation Analysis based on Fractional-Order and Its Application in Multiple Feature Extraction and Recognition
Xiaobo Shen 0001, Quan-Sen Sun
Neural Process. Lett.1
2014 A novel semi-supervised canonical correlation analysis and extensions for multi-view dimensionality reduction
Xiaobo Shen 0001, Quan-Sen Sun
J. Vis. Commun. Image Represent.1
2013 Orthogonal canonical correlation analysis and its application in feature fusion
Xiaobo Shen 0001, Quan-Sen Sun, Yun-Hao Yuan 0001
FUSION1