Xiu-Shen Wei

dblp:160/5936 · DBLP profile ↗
← Back
73ranked-venue papers
21as first author
47since 2021 · last 2026
0000-0002-8200-1845ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 14 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 6 first-author · 24 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Efficient and Effective In-context Demonstration Selection with Coreset
abstract
In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection.
Zihua Wang, Jiarui Wang 0002, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002, Xu Yang 0021, Xiu-Shen Wei, Siya Mi, Yu Zhang 0004
AAAI7
2026 Coarse Labels Matter: Revisiting the Role of Coarse-Grained Supervision in Fine-Grained Learning
abstract
The prohibitive cost of acquiring high-quality fine-grained annotations has spurred significant interest in leveraging readily available coarse labels for fine-grained learning. However, prevailing approaches tend to rely on increasingly sophisticated unsupervised methods to define fine-grained proxy tasks, with coarse labels often playing an auxiliary role. In this paper, we propose CSer, a framework designed to maximize the utility of coarse label information for Coarse-to-Fine learning. Specifically, to reconcile the conflict between preserving fine-grained feature diversity and maintaining strong coarse-grained supervision, our coarse-grained self-distillation strategy fortifies the backbone's discriminative power by distilling knowledge from the final classifier to intermediate layers. Concurrently, we introduce dense supervision on common component features within each coarse class, which are decoupled using Non-negative Matrix Factorization. This enhances responses to distinct components, thereby mitigating the simplicity bias in embeddings that can arise under coarse supervision. Moreover, we leverage relationships among intra-class samples to dynamically adjust the negative sampling strategy in contrastive learning, thereby constructing distinct fine-grained class relationships tailored to different coarse classes. Extensive experiments conducted on multiple benchmark datasets demonstrate the effectiveness of our method, yielding state-of-the-art results surpassing competing methods.
Xin-Yang Zhao, Pengyuan Zhang, Qiyuan Zhuang, Yazhou Yao, Xiu-Shen Wei
IEEE Trans. Image Process.5
2025 Prototype-Based Contrastive Learning with Stage-Wise Progressive Augmentation for Self-Supervised Fine-Grained Learning
Baofeng Tan, Xiu-Shen Wei, Lin Zhao 0003
ICCV2
2025 Object-Level Correlation for Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment objects of novel categories in the query images given only a few annotated support samples. Existing methods primarily build the image-level correlation between the support target object and the entire query image. However, this correlation contains the hard pixel noise, \textit{i.e.}, irrelevant background objects, that is intractable to trace and suppress, leading to the overfitting of the background. To address the limitation of this correlation, we imitate the biological vision process to identify novel objects in the object-level information. Target identification in the general objects is more valid than in the entire image, especially in the low-data regime. Inspired by this, we design an Object-level Correlation Network (OCNet) by establishing the object-level correlation between the support target object and query general objects, which is mainly composed of the General Object Mining Module (GOMM) and Correlation Construction Module (CCM). Specifically, GOMM constructs the query general object feature by learning saliency and high-level similarity cues, where the general objects include the irrelevant background objects and the target foreground object. Then, CCM establishes the object-level correlation by allocating the target prototypes to match the general object feature. The generated object-level correlation can mine the query target feature and suppress the hard pixel noise for the final prediction. Extensive experiments on PASCAL-${5}^{i}$ and COCO-${20}^{i}$ show that our model achieves the state-of-the-art performance.
Chunlin Wen, Yu Zhang 0004, Hongyuan Zhu 0002, Xiu-Shen Wei, Zhiqiang Kou, Shuzhou Sun
ICCV5
2025 Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
abstract
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task definitions (e.g., "image segmentation"), and conjecture it is a key barrier that hampers zero-shot task generalization. Our hypothesis is that without truly understanding previously-seen tasks—due to these terminological definitions—deep models struggle to generalize to novel tasks. To verify this, we introduce Explanatory Instructions, which provide an intuitive way to define CV task objectives through detailed linguistic transformations from input images to outputs. We create a large-scale dataset comprising 12 million "image input $\to$ explanatory instruction $\to$ output" triplets, and train an auto-regressive-based vision-language model (AR-based VLM) that takes both images and explanatory instructions as input. By learning to follow these instructions, the AR-based VLM achieves instruction-level zero-shot capabilities for previously-seen tasks and demonstrates strong zero-shot generalization for unseen CV tasks. Code and dataset will be open-sourced.
Yang Shen 0006, Xiu-Shen Wei, Yifan Sun 0003, YuXin Song 0001, Heyang Xu, Yazhou Yao, Errui Ding
ICML2
2025 Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
abstract
Fine-grained hashing has become a powerful solution for rapid and efficient image retrieval, particularly in scenarios requiring high discrimination between visually similar categories. To enable each hash bit to correspond to specific visual attributes, we propose a novel method that harnesses learnable queries for attribute-aware hash code learning. This method deploys a tailored set of queries to capture and represent nuanced attribute-level information within the hashing process, thereby enhancing both the interpretability and relevance of each hash bit. Building on this query-based optimization framework, we incorporate an auxiliary branch to help alleviate the challenges of complex landscape optimization often encountered with low-bit hash codes. This auxiliary branch models high-order attribute interactions, reinforcing the robustness and specificity of the generated hash codes. Experimental results on benchmark datasets demonstrate that our method generates attribute-aware hash codes and consistently outperforms state-of-the-art techniques in retrieval accuracy and robustness, especially for low-bit hash codes, underscoring its potential in fine-grained image hashing tasks.
Peng Wang 0107, Yong Li 0032, Lin Zhao 0003, Xiu-Shen Wei
ICML4
2025 Equiangular Basis Vectors: A Novel Paradigm for Classification Tasks
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Lingyan Gao
Int. J. Comput. Vis.3
2025 An Empirical Study on Training Paradigms for Deep Supervised Hashing
Yang Shen 0006, Peng Wang 0023, Xiu-Shen Wei, Yazhou Yao
Int. J. Comput. Vis.3
2025 Delving Deep into Simplicity Bias for Long-Tailed Image Recognition
Xiu-Shen Wei, Xuhao Sun, Yang Shen 0006, Peng Wang 0023
Int. J. Comput. Vis.1
2025 Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
abstract
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed, they often overfit to specific datasets and individual features, limiting their cross-domain applicability. To overcome these limitations, we propose a doubly adaptive dropout approach for cross-domain AU detection, which enhances the robustness of convolutional feature maps and spatial tokens against domain shifts. This approach includes a Channel Drop Unit (CD-Unit) and a Token Drop Unit (TD-Unit), which work together to reduce domain-specific noise at both the channel and token levels. The CD-Unit preserves domain-agnostic local patterns in feature maps, while the TD-Unit helps the model identify AU relationships generalizable across domains. An auxiliary domain classifier, integrated at each layer, guides the selective omission of domain-sensitive features. To prevent excessive feature dropout, a progressive training strategy is used, allowing for selective exclusion of sensitive features at any model layer. Our method consistently outperforms existing techniques in cross-domain AU detection, as demonstrated by extensive experimental evaluations. Visualizations of attention maps also highlight clear and meaningful patterns related to both individual and combined AUs, further validating the approach's effectiveness.
Yong Li 0032, Xuesong Niu, Yi Ding 0012, Xiu-Shen Wei, Cuntai Guan
IEEE Trans. Affect. Comput.5
2024 An Asymmetric Augmented Self-Supervised Learning Method for Unsupervised Fine-Grained Image Hashing
abstract
Unsupervised fine-grained image hashing aims to learn compact binary hash codes in unsupervised settings, addressing challenges posed by large-scale datasets and dependence on supervision. In this paper, we first identify a granularity gap between generic and fine-grained datasets for unsupervised hashing methods, highlighting the inadequacy of conventional self-supervised learning for fine-grained visual objects. To bridge this gap, we propose the Asymmetric Augmented Self-Supervised Learning (A2-SSL) method, comprising three modules. The asymmetric augmented SSL module employs suitable augmentation strategies for positive/negative views, preventing fine-grained category confusion inherent in conventional SSL. Part-oriented dense contrastive learning utilizes the Fisher Vector framework to capture and model fine- grained object parts, enhancing unsupervised representations through part-level dense contrastive learning. Self-consistent hash code learning introduces a reconstruction task aligned with the self-consistency principle, guiding the model to emphasize comprehensive features, particularly fine-grained patterns. Experimental results on five benchmark datasets demonstrate the superiority of A2-SSL over existing methods, affirming its efficacy in unsupervised fine-grained image hashing.
Feiran Hu, Chen-Lin Zhang, Jiangliang Guo, Xiu-Shen Wei, Lin Zhao 0003, Lingyan Gao
CVPR4
2024 Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual Reconstruction
abstract
Pre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model parameters, often fall short when confronted with long-tailed distributions, especially in detection tasks. This is largely due to extreme data imbalance and the issue of simplicity bias. In this paper, we introduce a novel pre-training framework for object detection, called Dynamic Rebalancing Contrastive Learning with Dual Reconstruction (2DRCL). Our method builds on a Holistic-Local Contrastive Learning mechanism, which aligns pre-training with object detection by capturing both global contextual semantics and detailed local patterns. To tackle the imbalance inherent in long-tailed data, we design a dynamic rebalancing strategy that adjusts the sampling of underrepresented instances throughout the pre-training process, ensuring better representation of tail classes. Moreover, Dual Reconstruction addresses simplicity bias by enforcing a reconstruction task aligned with the self-consistency principle, specifically benefiting underrepresented tail classes. Experiments on COCO and LVIS v1.0 datasets demonstrate the effectiveness of our method, particularly in improving the mAP/AP scores for tail classes.
Chen-Long Duan, Yong Li 0032, Xiu-Shen Wei, Lin Zhao 0003
NeurIPS3
2024 Negatives Make a Positive: An Embarrassingly Simple Approach to Semi-Supervised Few-Shot Learning
abstract
Semi-Supervised Few-Shot Learning (SSFSL) aims to train a classifier that can adapt to new tasks using limited labeled data and a fixed amount of unlabeled data. Various sophisticated methods have been proposed to tackle the challenges associated with this problem. In this paper, we present a simple but quite effective approach to predict accurate negative pseudo-labels of unlabeled data from an indirect learning perspective. We leverage these pseudo-labels to augment the support set, which is typically limited in few-shot tasks, e.g., 1-shot classification. In such label-constrained scenarios, our approach can offer highly accurate negative pseudo-labels. By iteratively excluding negative pseudo-labels one by one, we ultimately derive a positive pseudo-label for each unlabeled sample in our approach. The integration of negative and positive pseudo-labels complements the limited support set, resulting in significant accuracy improvements for SSFSL. Our approach can be implemented in just few lines of code by only using off-the-shelf operations, yet it outperforms state-of-the-art methods on four benchmark datasets. Furthermore, our approach exhibits good adaptability and generalization capabilities when used as a plug-and-play counterpart alongside existing SSFSL methods and when extended to generalized linear models.
Xiu-Shen Wei, Heyang Xu, Chen-Long Duan, Yuxin Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Contextualizing Meta-Learning via Learning to Decompose
abstract
Meta-learning has emerged as an efficient approach for constructing target models based on support sets. For example, the meta-learned embeddings enable the construction of target nearest-neighbor classifiers for specific tasks by pulling instances closer to their same-class neighbors. However, a single instance can be annotated from various latent attributes, making visually similar instances inside or across support sets have different labels and diverse relationships with others. Consequently, a uniform meta-learned strategy for inferring the target model from the support set fails to capture the instance-wise ambiguous similarity. To this end, we propose Learning to Decompose Network (LeadNet) tocontextualizethe meta-learned “support-to-target” strategy, leveraging the context of instances with one or mixed latent attributes in a support set. In particular, the comparison relationship between instances is decomposed w.r.t. multiple embedding spaces.LeadNetlearns to automatically select the strategy associated with the right attribute via incorporatingthe change of comparison across contextswith polysemous embeddings. We demonstrate the superiority ofLeadNetin various applications, including exploring multiple views of confusing data, out-of-distribution recognition, and few-shot image classification.
Han-Jia Ye, Da-Wei Zhou 0001, Lanqing Hong, Zhenguo Li, Xiu-Shen Wei, De-Chuan Zhan
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 MECOM: A Meta-Completion Network for Fine-Grained Recognition With Incomplete Multi-Modalities
abstract
Our work focuses on tackling the problem of fine-grained recognition with incomplete multi-modal data, which is overlooked by previous work in the literature. It is desirable to not only capture fine-grained patterns of objects but also alleviate the challenges of missing modalities for such a practical problem. In this paper, we propose to leverage a meta-learning strategy to learn model abilities of both fast modal adaptation and more importantly missing modality completion across a variety of incomplete multi-modality learning tasks. Based on that, we develop a meta-completion method, termed as MECOM, to perform multimodal fusion and explicit missing modality completion by our proposals of cross-modal attention and decoupling reconstruction. To further improve fine-grained recognition accuracy, an additional partial stream (as a counterpart of the main stream of MECOM, i.e., holistic) and the part-level features (corresponding to fine-grained objects' parts) selection are designed, which are tailored for fine-grained nature to capture discriminative but subtle part-level patterns. Comprehensive experiments from quantitative and qualitative aspects, as well as various ablation studies, on two fine-grained multimodal datasets and one generic multimodal dataset show our superiority over competing methods. Our code is open-source and available at https://github.com/SEU-VIPGroup/MECOM.
Xiu-Shen Wei, Hong-Tao Yu, Faen Zhang, Yuxin Peng 0001
IEEE Trans. Image Process.1
2023 Equiangular Basis Vectors
abstract
We propose Equiangular Basis Vectors (EBVs) for classification tasks. In deep neural networks, models usually end with a k-way fully connected layer with softmax to handle different classification tasks. The learning objective of these methods can be summarized as mapping the learned feature representations to the samples' label space. While in metric learning approaches, the main objective is to learn a transformation function that maps training data points from the original space to a new space where similar points are closer while dissimilar points become farther apart. Different from previous methods, our EBVs generate normalized vector embeddings as “predefined classifiers” which are required to not only be with the equal status between each other, but also be as orthogonal as possible. By minimizing the spherical distance of the embedding of an input between its categorical EBV in training, the predictions can be obtained by identifying the categorical EBV with the smallest distance during inference. Various experiments on the ImageNet-1K dataset and other downstream tasks demonstrate that our method outperforms the general fully connected classifier while it does not introduce huge additional computation compared with classical metric learning methods. Our EBVs won the first place in the 2022 DIGIX Global AI Challenge, and our code is open-source and available at https://github.com/NJUST-VIPGroup/Equiangular-Basis-Vectors.
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei
CVPR3
2023 Hawkeye: A PyTorch-based Library for Fine-Grained Image Recognition with Deep Learning
abstract
Fine-Grained Image Recognition (FGIR) is a fundamental and challenging task in computer vision and multimedia that plays a crucial role in Intellectual Economy and Industrial Internet applications. However, the absence of a unified open-source software library covering various paradigms in FGIR poses a significant challenge for researchers and practitioners in the field. To address this gap, we present Hawkeye, a PyTorch-based library for FGIR with deep learning. Hawkeye is designed with a modular architecture, emphasizing high-quality code and human-readable configuration, providing a comprehensive solution for FGIR tasks. In Hawkeye, we have implemented 16 state-of-the-art fine-grained methods, covering 6 different paradigms, enabling users to explore various approaches for FGIR. To the best of our knowledge, Hawkeye represents the first open-source PyTorch-based library dedicated to FGIR. It is publicly available at https://github.com/Hawkeye-FineGrained/Hawkeye/, providing researchers and practitioners with a powerful tool to advance their research and development in the field of FGIR.
Yang Shen 0006, Xiu-Shen Wei, Ye Wu 0001
ACM Multimedia3
2023 Hyperbolic Space with Hierarchical Margin Boosts Fine-Grained Learning from Coarse Labels
abstract
Learning fine-grained embeddings from coarse labels is a challenging task due to limited label granularity supervision, i.e., lacking the detailed distinctions required for fine-grained tasks. The task becomes even more demanding when attempting few-shot fine-grained recognition, which holds practical significance in various applications. To address these challenges, we propose a novel method that embeds visual embeddings into a hyperbolic space and enhances their discriminative ability with a hierarchical cosine margins manner. Specifically, the hyperbolic space offers distinct advantages, including the ability to capture hierarchical relationships and increased expressive power, which favors modeling fine-grained objects. Based on the hyperbolic space, we further enforce relatively large/small similarity margins between coarse/fine classes, respectively, yielding the so-called hierarchical cosine margins manner. While enforcing similarity margins in the regular Euclidean space has become popular for deep embedding learning, applying it to the hyperbolic space is non-trivial and validating the benefit for coarse-to-fine generalization is valuable. Extensive experiments conducted on five benchmark datasets showcase the effectiveness of our proposed method, yielding state-of-the-art results surpassing competing methods.
Shu-Lin Xu, Yifan Sun 0003, Faen Zhang, Xiu-Shen Wei, Yi Yang 0001
NeurIPS5
2023 Bridge the gap between supervised and unsupervised learning for fine-grained classification
Jiabao Wang 0001, Yang Li 0015, Xiu-Shen Wei, Hang Li 0008, Zhuang Miao, Rui Zhang 0038
Inf. Sci.3
2023 Learning Graph Convolutional Networks for Multi-Label Recognition and Applications
abstract
The task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important information, we propose graph convolutional networks (GCNs) based models for multi-label image recognition, where directed graphs are constructed over classes and information is propagated between classes to learn inter-dependent class-level representations. Following this idea, we design two particular models that approach multi-label classification from different views. In our first model, the prior knowledge about the class dependencies is integrated into classifier learning. Specifically, we propose Classifier Learning GCN (C-GCN) to map class-level semantic representations (e.g., word embeddings) into classifiers that maintain the inter-class topology. In our second model, we decompose the visual representation of an image into a set of label-aware features and propose prediction learning GCN (P-GCN) to encode such features into inter-dependent image-level prediction scores. Furthermore, we also present an effective correlation matrix construction approach to capture inter-class relationships and consequently guide information propagation among classes. Empirical results on generic multi-label image recognition demonstrate that both of the proposed models can obviously outperform other existing state-of-the-arts. Moreover, the proposed methods also show advantages in some other multi-label classification related applications.
Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval
abstract
Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose attribute-aware hashing networks with self-consistency for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. Our models are also equipped with a feature decorrelation constraint upon these attribute vectors to strengthen their representative abilities. Then, driven by preserving original entities' similarity, the required hash codes can be generated from these attribute-specific vectors and thus become attribute-aware. Furthermore, to combat simplicity bias in deep hashing, we consider the model design from the perspective of the self-consistency principle and propose to further enhance models' self-consistency by equipping an additional image reconstruction path. Comprehensive quantitative experiments under diverse empirical settings on six fine-grained retrieval datasets and two generic retrieval datasets show the superiority of our models over competing methods. Moreover, qualitative results demonstrate that not only the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects, but also our self-consistency designs can effectively overcome simplicity bias in fine-grained hashing.
Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Peng Wang 0023, Yuxin Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Prototype Learning for Automatic Check-Out
abstract
The basic goal of Automatic Check-Out (ACO) task is to accurately predict the categories and quantities of products selected by customers in the check-out images. However, there is a significant domain gap between the single-product exemplars as training data and the check-out images as testing data. To mitigate the domain gap, we propose a novel method termed as Prototype Learning for Automatic Check-Out (PLACO). In PLACO, prototype learning is designed to reach the goal in two ways. Specifically, in the prototype-based classifier learning module, to fully exploit the invariance of category prototypes, the prototypes obtained from the single-product exemplars are employed to generate classifiers for classifying the proposals of check-out image. On the other side, in prototype alignment module, prototypes for both the single-product exemplar and check-out image domains are entered simultaneously to ensure intra-category compactness and inter-category sparsity. Moreover, to further improve the performance of PLACO, we develop a discriminative re-ranking module to both adjust the predicted scores of product proposals for bringing more discriminative ability in classifier learning and provide a reasonable sorting possibility by considering the fine-grained nature. Experiments are conducted on the large-scale RPC dataset for evaluations. Our PLACO obtains the optimal results in both traditional ACO task setting and incremental task setting.
Hao Chen 0052, Xiu-Shen Wei, Liang Xiao 0001
IEEE Trans. Multim.2
2023 Boosting Robust Learning Via Leveraging Reusable Samples in Noisy Web Data
abstract
Webly-supervised fine-grained visual classification (FGVC) has attracted increasing attention in recent years because of the unaffordable cost of obtaining correctly-labeled large-scale fine-grained datasets. However, due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause accumulated errors. Sample selection methods identify clean (“easy”) samples based on the fact that small losses can alleviate the accumulated errors. However, “hard” and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the network. Furthermore, in order to endow our model with the capability to capture richer and more discriminative feature representations, we propose a cross-layer attention-based feature refinement (CLAR) block. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jian Zhang 0002, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2023 QBox: Partial Transfer Learning With Active Querying for Object Detection
abstract
Object detection requires plentiful data annotated with bounding boxes for model training. However, in many applications, it is difficult or even impossible to acquire a large set of labeled examples for the target task due to the privacy concern or lack of reliable annotators. On the other hand, due to the high-quality image search engines, such as Flickr and Google, it is relatively easy to obtain resource-rich unlabeled datasets, whose categories are a superset of those of target data. In this article, to improve the target model with cost-effective supervision from source data, we propose a partial transfer learning approach QBox to actively query labels for bounding boxes of source images. Specifically, we design two criteria, i.e., informativeness and transferability, to measure the potential utility of a bounding box for improving the target model. Based on these criteria, QBox actively queries the labels of the most useful boxes from the source domain and, thus, requires fewer training examples to save the labeling cost. Furthermore, the proposed query strategy allows annotators to simply labeling a specific region, instead of the whole image, and, thus, significantly reduces the labeling difficulty. Extensive experiments are performed on various partial transfer benchmarks and a real COVID-19 detection task. The results validate that QBox improves the detection accuracy with lower labeling cost compared to state-of-the-art query strategies for object detection.
Ying-Peng Tang, Xiu-Shen Wei, Borui Zhao, Sheng-Jun Huang
IEEE Trans. Neural Networks Learn. Syst.2
2022 Dual Attention Networks for Few-Shot Fine-Grained Recognition
abstract
The task of few-shot fine-grained recognition is to classify images belonging to subordinate categories merely depending on few examples. Due to the fine-grained nature, it is desirable to capture subtle but discriminative part-level patterns from limited training data, which makes it a challenging problem. In this paper, to generate fine-grained tailored representations for few-shot recognition, we propose a Dual Attention Network (Dual Att-Net) consisting of two dual branches of both hard- and soft-attentions. Specifically, by producing attention guidance from deep activations of input images, our hard-attention is realized by keeping a few useful deep descriptors and forming them as a bag of multi-instance learning. Since these deep descriptors could correspond to objects' parts, the advantage of modeling as a multi-instance bag is able to exploit inherent correlation of these fine-grained parts. On the other side, a soft attended activation representation can be obtained by applying attention guidance upon original activations, which brings comprehensive attention information as the counterpart of hard-attention. After that, both outputs of dual branches are aggregated as a holistic image embedding w.r.t. input images. By performing meta-learning, we can learn a powerful image embedding in such a metric space to generalize to novel classes. Experiments on three popular fine-grained benchmark datasets show that our Dual Att-Net obviously outperforms other existing state-of-the-art methods.
Shu-Lin Xu, Faen Zhang, Xiu-Shen Wei
AAAI3
2022 Relieving Long-tailed Instance Segmentation via Pairwise Class Balance
abstract
Long-tailed instance segmentation is a challenging task due to the extreme imbalance of training samples among classes. It causes severe biases of the head classes (with majority samples) against the tailed ones. This renders “how to appropriately define and alleviate the bias” one of the most important issues. Prior works mainly use label distribution or mean score information to indicate a coarse-grained bias. In this paper, we explore to excavate the confusion matrix, which carries the fine-grained misclassification details, to relieve the pairwise biases, generalizing the coarse one. To this end, we propose a novel Pairwise Class Balance (PCB) method, built upon a confusion matrix which is updated during training to accumulate the ongoing prediction preferences. PCB generates fightback soft labels for regularization during training. Besides, an iterative learning paradigm is developed to support a progressive and smooth regularization in such debiasing. PCB can be plugged and played to any existing method as a complement. Experimental results on LVIS demonstrate that our method achieves state-of-the-art performance without bells and whistles. Superior results across various architectures show the generalization ability. The code and trained models are available at https://github.com/megvii-research/PCB.
Yin-Yin He, Peizhen Zhang, Xiu-Shen Wei, Xiangyu Zhang 0005, Jian Sun 0001
CVPR3
2022 Automatic Check-Out via Prototype-Based Classifier Learning from Single-Product Exemplars
Hao Chen 0052, Xiu-Shen Wei, Faen Zhang, Yang Shen 0006, Liang Xiao 0001
ECCV (25)2
2022 SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Qing-Yuan Jiang, Jian Yang 0003
ECCV (14)3
2022 A Channel Mix Method for Fine-Grained Cross-Modal Retrieval
abstract
In this paper, we propose a simple but effective method for dealing with the challenging fine-grained cross-modal retrieval task where it aims to enable flexible retrieval among subor-dinate categories across different modalities. Specifically, in order to enhance information interaction in different modalities for fine-grained objects, a channel mix method is developed and performed upon the channels of deep activations across dif-ferent modalities. After that, a 1 x 1 convolution is employed to aggregate the mixed channels into a unified feature vector. Moreover, equipped with a novel fine-grained cross-modal cen-ter loss, our method can further improve the intra-class separa-bility as well as inter-class compactness for multi-modalities. Experiments are conducted on the fine-grained cross-modal benchmark dataset and show our superiority over competing methods. Meanwhile, ablation studies also demonstrate the effectiveness of our proposals.
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Hanxu Hu
ICME3
2022 Webly-Supervised Fine-Grained Recognition with Partial Label Learning
abstract
The task of webly-supervised fine-grained recognition is to boost recognition accuracy of classifying subordinate categories (e.g., different bird species) by utilizing freely available but noisy web data. As the label noises significantly hurt the network training, it is desirable to distinguish and eliminate noisy images. In this paper, we propose two strategies, i.e., open-set noise removal and closed-set noise correction, to both remove such two kinds of web noises w.r.t. fine-grained recognition. Specifically, for open-set noise removal, we utilize a pre-trained deep model to perform deep descriptor transformation to estimate the positive correlation between these web images, and detect the open-set noises based on the correlation values. Regarding closed-set noise correction, we develop a top-k recall optimization loss for firstly assigning a label set towards each web image to reduce the impact of hard label assignment for closed-set noises. Then, we further propose to correct the sample with its label set as the true single label from a partial label learning perspective. Experiments on several webly-supervised fine-grained benchmark datasets show that our method obviously outperforms other existing state-of-the-art methods.
Yu-Yan Xu, Yang Shen 0006, Xiu-Shen Wei, Jian Yang 0003
IJCAI3
2022 An Embarrassingly Simple Approach to Semi-Supervised Few-Shot Learning
abstract
Semi-supervised few-shot learning consists in training a classifier to adapt to new tasks with limited labeled data and a fixed quantity of unlabeled data. Many sophisticated methods have been developed to address the challenges this problem comprises. In this paper, we propose a simple but quite effective approach to predict accurate negative pseudo-labels of unlabeled data from an indirect learning perspective, and then augment the extremely label-constrained support set in few-shot classification tasks. Our approach can be implemented in just few lines of code by only using off-the-shelf operations, yet it is able to outperform state-of-the-art methods on four benchmark datasets.
Xiu-Shen Wei, Heyang Xu, Faen Zhang, Yuxin Peng 0001
NeurIPS1
2022 RPC: a large-scale and fine-grained retail product checkout dataset
Xiu-Shen Wei, Quan Cui, Lei Yang 0048, Peng Wang 0023, Lingqiao Liu, Jian Yang 0003
Sci. China Inf. Sci.1
2022 Prototype-based classifier learning for long-tailed visual recognition
Xiu-Shen Wei, Shu-Lin Xu, Hao Chen 0052, Liang Xiao 0001, Yuxin Peng 0001
Sci. China Inf. Sci.1
2022 Preface
Min-Ling Zhang, Xiu-Shen Wei, Gao Huang 0001
J. Comput. Sci. Technol.2
2022 Fine-Grained Image Analysis With Deep Learning: A Survey
abstract
Fine-grained image analysis (FGIA) is a longstanding and fundamental problem in computer vision and pattern recognition, and underpins a diverse set of real-world applications. The task of FGIA targets analyzing visual objects from subordinate categories, e.g., species of birds or models of cars. The small inter-class and large intra-class variation inherent to fine-grained image analysis makes it a challenging problem. Capitalizing on advances in deep learning, in recent years we have witnessed remarkable progress in deep learning powered FGIA. In this paper we present a systematic survey of these advances, where we attempt to re-define and broaden the field of FGIA by consolidating two fundamental fine-grained research areas - fine-grained image recognition and fine-grained image retrieval. In addition, we also review other key issues of FGIA, such as publicly available benchmark datasets and related domain-specific applications. We conclude by highlighting several research directions and open problems which need further exploration from the community.
Xiu-Shen Wei, Yi-Zhe Song, Oisin Mac Aodha, Jianxin Wu 0001, Yuxin Peng 0001, Jinhui Tang 0001, Jian Yang 0003, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Delving deep into spatial pooling for squeeze-and-excitation networks
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan
Pattern Recognit.3
2022 Self-Supervised Multi-Category Counting Networks for Automatic Check-Out
abstract
The practical task of Automatic Check-Out (ACO) is to accurately predict the presence and count of each product in an arbitrary product combination. Beyond the large-scale and the fine-grained nature of product categories as its main challenges, products are always continuously updated in realistic check-out scenarios, which is also required to be solved in an ACO system. Previous work in this research line almost depends on the supervisions of labor-intensive bounding boxes of products by performing a detection paradigm. While, in this paper, we propose a Self-Supervised Multi-Category Counting (S2MC2) network to leverage the point-level supervisions of products in check-out images to both lower the labeling cost and be able to return ACO predictions in a class incremental setting. Specifically, as a backbone, our S2MC2 is built upon a counting module in a class-agnostic counting fashion. Also, it consists of several crucial components including an attention module for capturing fine-grained patterns and a domain adaptation module for reducing the domain gap between single product images as training and check-out images as test. Furthermore, a self-supervised approach is utilized in S2MC2 to initialize the parameters of its backbone for better performance. By conducting comprehensive experiments on the large-scale automatic check-out dataset RPC, we demonstrate that our proposed S2MC2 achieves superior accuracy in both traditional and incremental settings of ACO tasks over the competing baselines.
Hao Chen 0052, Yangzhun Zhou, Jun Li 0027, Xiu-Shen Wei, Liang Xiao 0001
IEEE Trans. Image Process.4
2022 Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard Examples
abstract
Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop.
Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.4
2022 A Lightweight Encoder-Decoder Path for Deep Residual Networks
abstract
In this article, we present a novel lightweight path for deep residual neural networks. The proposed method integrates a simple plug-and-play module, i.e., a convolutional encoder-decoder (ED), as an augmented path to the original residual building block. Due to the abstract design and ability of the encoding stage, the decoder part tends to generate feature maps where highly semantically relevant responses are activated, while irrelevant responses are restrained. By a simple elementwise addition operation, the learned representations derived from the identity shortcut and original transformation branch are enhanced by our ED path. Furthermore, we exploit lightweight counterparts by removing a portion of channels in the original transformation branch. Fortunately, our lightweight processing does not cause an obvious performance drop but brings a computational economy. By conducting comprehensive experiments on ImageNet, MS-COCO, CUB200-2011, and CIFAR, we demonstrate the consistent accuracy gain obtained by our ED path for various residual architectures, with comparable or even lower model complexity. Concretely, it decreases the top-1 error of ResNet-50 and ResNet-101 by 1.22% and 0.91% on the task of ImageNet classification and increases the mmAP of Faster R-CNN with ResNet-101 by 2.5% on the MS-COCO object detection task. The code is available at https://github.com/Megvii-Nanjing/ED-Net.
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan, Yang Yu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Bag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks
abstract
In recent years, visual recognition on challenging long-tailed distributions, where classes often exhibit extremely imbalanced frequencies, has made great progress mostly based on various complex paradigms (e.g., meta learning). Apart from these complex methods, simple refinements on training procedures also make contributions. These refinements, also called tricks, are minor but effective, such as adjustments in the data distribution or loss functions. However, different tricks might conflict with each other. If users apply these long-tail related tricks inappropriately, it could cause worse recognition accuracy than expected. Unfortunately, there has not been a scientific guideline of these tricks in the literature. In this paper, we first collect existing tricks in long-tailed visual recognition and then perform extensive and systematic experiments, in order to give a detailed experimental guideline and obtain an effective combination of these tricks. Furthermore, we also propose a novel data augmentation approach based on class activation maps for long-tailed recognition, which can be friendly combined with re-sampling methods and shows excellent results. By assembling these tricks scientifically, we can outperform state-of-the-art methods on four long-tailed benchmark datasets, including ImageNet-LT and iNaturalist 2018. Our code is open-source and available at https://github.com/zhangyongshun/BagofTricks-LT.
Xiu-Shen Wei, Boyan Zhou, Jianxin Wu 0001
AAAI2
2021 Contrastive Learning Based Hybrid Networks for Long-Tailed Image Classification
abstract
Learning discriminative image representations plays a vital role in long-tailed image classification because it can ease the classifier learning in imbalanced cases. Given the promising performance contrastive learning has shown recently in representation learning, in this work, we explore effective supervised contrastive learning strategies and tailor them to learn better image representations from imbalanced data in order to boost the classification accuracy thereon. Specifically, we propose a novel hybrid network structure being composed of a supervised contrastive loss to learn image representations and a cross-entropy loss to learn classifiers, where the learning is progressively transited from feature learning to the classifier learning to embody the idea that better features make better classifiers. We explore two variants of contrastive loss for feature learning, which vary in the forms but share a common idea of pulling the samples from the same class together in the normalized embedding space and pushing the samples from different classes apart. One of them is the recently proposed supervised contrastive (SC) loss, which is designed on top of the state-of-the-art unsupervised contrastive loss by incorporating positive samples from the same class. The other is a prototypical supervised contrastive (PSC) learning strategy which addresses the intensive memory consumption in standard SC loss and thus shows more promise under limited memory budget. Extensive experiments on three long-tailed classification datasets demonstrate the advantage of the proposed contrastive learning based hybrid networks in long-tailed classification.
Peng Wang 0023, Kai Han 0001, Xiu-Shen Wei, Lei Zhang 0054, Lei Wang 0001
CVPR3
2021 Distilling Virtual Examples for Long-tailed Recognition
abstract
We tackle the long-tailed visual recognition problem from the knowledge distillation perspective by proposing a Distill the Virtual Examples (DiVE) method. Specifically, by treating the predictions of a teacher model as virtual examples, we prove that distilling from these virtual examples is equivalent to label distribution learning under certain constraints. We show that when the virtual example distribution becomes flatter than the original input distribution, the under-represented tail classes will receive significant improvements, which is crucial in long-tailed recognition. The proposed DiVE method can explicitly tune the virtual example distribution to become flat. Extensive experiments on three benchmark datasets, including the large-scale iNaturalist ones, justify that the proposed DiVE method can significantly outperform state-of-the-art methods. Further-more, additional analyses and experiments verify the virtual example interpretation, and demonstrate the effectiveness of tailored designs in DiVE for long-tailed problems.
Yin-Yin He, Jianxin Wu 0001, Xiu-Shen Wei
ICCV3
2021 Webly Supervised Fine-Grained Recognition: Benchmark Datasets and An Approach
abstract
Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will significantly reduce the labeling costs by leveraging free web data. Despite its significant practical and research value, the webly supervised fine-grained recognition problem is not extensively studied in the computer vision community, largely due to the lack of high-quality datasets. To fill this gap, in this paper we construct two new benchmark webly supervised fine-grained datasets, termed WebFG-496 and WebiNat-5089, respectively. In concretely, WebFG-496 consists of three sub-datasets containing a total of 53,339 web training images with 200 species of birds (Web-bird), 100 types of aircrafts (Web-aircraft), and 196 models of cars (Web-car). For WebiNat-5089, it contains 5089 sub-categories and more than 1.1 million web training images, which is the largest webly supervised fine-grained dataset ever. As a minor contribution, we also propose a novel webly supervised method (termed "Peer-learning") for benchmarking these datasets. Comprehensive experimental results and analyses on two new benchmark datasets demonstrate that the proposed method achieves superior performance over the competing baseline models and states-of-the-art. Our benchmark datasets and the source codes of Peer-learning have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/weblyFG-dataset.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jianxin Wu 0001, Jian Zhang 0002, Heng Tao Shen
ICCV3
2021 MULL'21: First International Workshop on Multimedia Understanding with Less Labeling
abstract
With the advent of deep neural networks, quite a lot of multimedia tasks have been significantly improved. While, however, deep neural networks still lack the ability of learning from less labeling, e.g., with limited exemplars or fast generalizing to new tasks. In order to address the current inefficiency of multimedia, there is pressing need to research methods to drastically reduce requirements for labeled training data. This workshop aims to provide a platform for discussing the challenges and corresponding innovative approaches in multimedia with less labeling. We hope more advanced technologies can be proposed or inspired, and also we invite several domain-specific experts for sharing their insights and research progress on the topic of MULL.
Xiu-Shen Wei, Jufeng Yang, Han-Jia Ye, Jian Yang 0003
ACM Multimedia1
2021 A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval
abstract
Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose an Attribute-Aware hashing Network (A$^2$-Net) for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. A$^2$-Net is also equipped with a feature decorrelation constraint upon these attribute vectors to enhance their representation abilities. Finally, the required hash codes are generated by the attribute vectors driven by preserving original similarities. Qualitative experiments on five benchmark fine-grained datasets show our superiority over competing methods. More importantly, quantitative results demonstrate the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects.
Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Han-Jia Ye, Jian Yang 0003
NeurIPS1
2021 Multi-Instance Learning With Emerging Novel Class
abstract
Diverse applications involving complicated data objects such as proteins and images are solved by applying multi-instance learning (MIL) algorithms. However, few MIL algorithms can deal with problems in an open and dynamic environment, where new categories of samples emerge. In this type of emerging novel class setting, algorithms should be able to not only classify the samples from the observed classes accurately, but also recognize the samples from the novel class. In this paper, we focus on the Multi-Instance learning with Emerging Novel class (MIEN) problem, and formulate MIEN from a metric learning perspective. We extract key instances to form the “super-bag” for each observed class, and non-key instances from all the observed classes to form a “meta super-bag”. Based on these super-bags, we propose the MIEN-metric method to learn discriminative metrics for classifying MIL bags from the observed classes and recognizing bags from the novel class. Experimental results of diverse domains, e.g., biological function annotation, text categorization, and object-centric/scene-centric image classification, show MIEN-metric outperforms other baseline methods significantly when the novel class emerges. Meanwhile, MIEN-metric is comparable with state-of-the-art MIL algorithms for binary classification in the traditional MIL setting.
Xiu-Shen Wei, Han-Jia Ye, Xin Mu, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.1
2021 Disentangling, Embedding and Ranking Label Cues for Multi-Label Image Recognition
abstract
Multi-label image recognition is a fundamental but challenging computer vision and multimedia task. Great progress has been achieved by exploiting label correlations among these multiple labels associated with a single image, which is the most crucial issue for multi-label image recognition. In this paper, to explicitly model label correlations, we propose a unified deep learning framework to Disentangle, Embed and Rank (DER) the corresponding label cues. Specifically, we first obtain class-aware disentangled maps (CADMs) by reforming deep activations in accordance with the class-specific recognition weights. Then, after transforming CADMs into the corresponding label vectors, we propose an embedding operation from a metric learning perspective to pull the relevant label vectors together and push irrelevant label vectors away. Furthermore, a ranking operation is employed, which aims to accurately and robustly measure the similarity/dissimilarity of these label vectors. Our model can be trained in an end-to-end manner with only image-level supervision, during which the proposed embedding and ranking operations can contribute to the CADMs learning through back-propagation. In addition, the obtained CADMs are aggregated and further used as an essential feature stream for the final multi-label classification. We conduct extensive experiments on three commonly used multi-label benchmark datasets. Quantitative results show that our model can significantly and consistently outperform previous competitive methods. Moreover, qualitative analysis of our DER proposal also reveals the effectiveness of our proposed model.
Quan Cui, Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001
IEEE Trans. Multim.3
2020 Exploring Categorical Regularization for Domain Adaptive Object Detection
abstract
In this paper, we tackle the domain adaptive object detection problem, where the main challenge lies in significant domain gaps between source and target domains. Previous work seeks to plainly align image-level and instance-level shifts to eventually minimize the domain discrepancy. However, they still overlook to match crucial image regions and important instances across domains, which will strongly affect domain shift mitigation. In this work, we propose a simple but effective categorical regularization framework for alleviating this issue. It can be applied as a plug-and-play component on a series of Domain Adaptive Faster R-CNN methods which are prominent for dealing with domain adaptive detection. Specifically, by integrating an image-level multi-label classifier upon the detection backbone, we can obtain the sparse but crucial image regions corresponding to categorical information, thanks to the weakly localization ability of the classification manner. Meanwhile, at the instance level, we leverage the categorical consistency between image-level predictions (by the classifier) and instance-level predictions (by the detection head) as a regularization factor to automatically hunt for the hard aligned instances of target domains. Extensive experiments of various domain shift scenarios show that our method obtains a significant performance gain over original Domain Adaptive Faster R-CNN detectors. Furthermore, qualitative visualization and analyses can demonstrate the ability of our method for attending on the key regions/instances targeting on domain adaptation. Our code is open-source and available at https://github.com/Megvii-Nanjing/CR-DA-DET.
Chang-Dong Xu, Xing-Ran Zhao, Xin Jin 0023, Xiu-Shen Wei
CVPR4
2020 BBN: Bilateral-Branch Network With Cumulative Learning for Long-Tailed Visual Recognition
abstract
Our work focuses on tackling the challenging but natural visual recognition task of long-tailed data distribution (i.e., a few classes occupy most of the data, while most classes have rarely few samples). In the literature, class re-balancing strategies (e.g., re-weighting and re-sampling) are the prominent and effective methods proposed to alleviate the extreme imbalance for dealing with long-tailed problems. In this paper, we firstly discover that these re-balancing methods achieving satisfactory recognition accuracy owe to that they could significantly promote the classifier learning of deep networks. However, at the same time, they will unexpectedly damage the representative ability of the learned deep features to some extent. Therefore, we propose a unified Bilateral-Branch Network (BBN) to take care of both representation learning and classifier learning simultaneously, where each branch does perform its own duty separately. In particular, our BBN model is further equipped with a novel cumulative learning strategy, which is designed to first learn the universal patterns and then pay attention to the tail data gradually. Extensive experiments on four benchmark datasets, including the large-scale iNaturalist ones, justify that the proposed BBN can significantly outperform state-of-the-art methods. Furthermore, validation experiments can demonstrate both our preliminary discovery and effectiveness of tailored designs in BBN for long-tailed problems. Our method won the first place in the iNaturalist 2019 large scale species classification competition, and our code is open-source and available at https://github.com/Megvii-Nanjing/BBN.
Boyan Zhou, Quan Cui, Xiu-Shen Wei
CVPR3
2020 Hierarchical Context Embedding for Region-Based Object Detection
Xin Jin 0023, Borui Zhao, Xiu-Shen Wei, Yanwen Guo 0001
ECCV (21)4
2020 ExchNet: A Unified Hashing Network for Large-Scale Fine-Grained Image Retrieval
Quan Cui, Qing-Yuan Jiang, Xiu-Shen Wei, Wu-Jun Li, Osamu Yoshie
ECCV (3)3
2020 PyRetri: A PyTorch-based Library for Unsupervised Image Retrieval by Deep Convolutional Neural Networks
abstract
Despite significant progress of applying deep learning methods to the field of content-based image retrieval, there has not been a software library that covers these methods in a unified manner. In order to fill this gap, we introduce PyRetri, an open source library for deep learning based unsupervised image retrieval. The library encapsulates the retrieval process in several stages and provides functionality that covers various prominent methods for each stage. The idea underlying its design is to provide a unified platform for deep learning based image retrieval research, with high usability and extensibility. The project source code, with usage examples, sample data and pre-trained models are available at https://github.com/PyRetri/.
Benyi Hu, Renjie Song, Xiu-Shen Wei, Yazhou Yao, Xian-Sheng Hua 0001, Yuehu Liu
ACM Multimedia3
2020 CRSSC: Salvage Reusable Samples from Noisy Data for Robust Learning
abstract
Due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause severe accumulated errors. Sample selection methods identify clean ("easy") samples based on the fact that small losses can alleviate the accumulated errors. However, "hard" and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the networks. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Xian-Sheng Hua 0001, Yazhou Yao, Xiu-Shen Wei, Guosheng Hu, Jian Zhang 0002
ACM Multimedia4
2020 Piecewise Hashing: A Deep Hashing Method for Large-Scale Fine-Grained Search
Yimu Wang, Xiu-Shen Wei, Bo Xue 0004, Lijun Zhang 0005
PRCV (2)2
2020 An Adversarial Domain Adaptation Network For Cross-Domain Fine-Grained Recognition
abstract
In this paper, we tackle a valuable yet very challenging visual recognition task, where the instances are within a subordinate category, and the target domain undergoes a shift with the source domain. This task, termed as cross-domain fine-grained recognition, relates closely to many real-life scenarios, e.g., recognizing retail products in storage racks by models trained with images collected in controlled environments. To deal with this problem, we design a new algorithm and propose a corresponding fine-grained domain adaptation dataset. Firstly, we propose a novel end-to-end CNN architecture that integrates two specialized modules: an adversarial module for domain alignment and a self-attention module for fine-grained recognition. The adversarial module is used to handle domain shift by gradually aligning the different domains with domain-level and class-level alignments, and strive to help the classifier learn with domain-invariant features generated by nets. The self-attention module is designed to capture discriminative image regions which are crucial for fine-grained visual recognition. Secondly, we collect a large-scale fine-grained domain adaptation dataset of retail products, which contains 52,011 images of 263 classes from 3 domains. Thirdly, we validate the effectiveness of our method on three datasets, showing that the proposed method can yield significant improvements over baseline methods on fine-grained datasets. Besides, we also evaluate the effectiveness of the self-attention module by performing visualization, which can capture the discriminative image regions in both source and target domains.
Yimu Wang, Renjie Song, Xiu-Shen Wei, Lijun Zhang 0005
WACV3
2020 Adversarial Learning of Structure-Aware Fully Convolutional Networks for Landmark Localization
abstract
Landmark/pose estimation in single monocular images has received much effort in computer vision due to its important applications. It remains a challenging task when input images come with severe occlusions caused by, e.g., adverse camera views. Under such circumstances, biologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geometric constraints of landmark point inter-connectivity. To address the problem, by incorporating priors about the structure of pose components, we propose a novel structure-aware fully convolutional network to implicitly take such priors into account during training of the deep network. Explicit learning of such constraints is typically challenging. Instead, inspired by how human identifies implausible poses, we design discriminators to distinguish the real poses from the fake ones (such as biologically implausible ones). If the pose generator G generates results that the discriminator fails to distinguish from real ones, the network successfully learns the priors. Training of the network follows the strategy of conditional Generative Adversarial Networks (GANs). The effectiveness of the proposed network is evaluated on three pose-related tasks: 2D human pose estimation, 2D facial landmark estimation and 3D human pose estimation. The proposed approach significantly outperforms several state-of-the-art methods and almost always generates plausible pose predictions, demonstrating the usefulness of implicit learning of structures using GANs.
Yu Chen 0037, Chunhua Shen, Hao Chen 0041, Xiu-Shen Wei, Lingqiao Liu, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Structure-aware human pose estimation with graph convolutional networks
Yanrui Bin, Xiu-Shen Wei, Xinya Chen, Changxin Gao, Nong Sang
Pattern Recognit.3
2020 Learning Semantically Enhanced Feature for Fine-Grained Image Classification
abstract
We aim to provide a computationally cheap yet effective approach for fine-grained image classification (FGIC) in this letter. Unlike previous methods that rely on complex part localization modules, our approach learns fine-grained features by enhancing the semantics of sub-features of a global feature. Specifically, we first achieve the sub-feature semantic by arranging feature channels of a CNN into different groups through channel permutation. Meanwhile, to enhance the discriminability of sub-features, the groups are guided to be activated on object parts with strong discriminability by a weighted combination regularization. Our approach is parameter parsimonious and can be easily integrated into the backbone model as a plug-and-play module for end-to-end training with only image-level supervision. Experiments verified the effectiveness of our approach and validated its comparable performance to the state-of-the-art methods. Code is available at https://github.com/cswluo/SEF.
Wei Luo 0006, Hengmin Zhang, Jun Li 0027, Xiu-Shen Wei
IEEE Signal Process. Lett.4
2020 Bi-Modal Progressive Mask Attention for Fine-Grained Recognition
abstract
Traditional fine-grained image recognition is required to distinguish different subordinate categories (e.g., birds species) based on the visual cues beneath raw images. Due to both small inter-class variations and large intra-class variations, it is desirable to capture the subtle differences between these sub-categories, which is crucial but challenging for fine-grained recognition. Recently, language modality aggregation has been proved as a successful technique to improve visual recognition in the experience. In this paper, we introduce an end-to-end trainable Progressive Mask Attention (PMA) model for fine-grained recognition by leveraging both visual and language modalities. Our Bi-Modal PMA model can not only stage-by-stage capture the most discriminative part in the visual modality by our mask-based fashion, but also explore the out-of-visual-domain knowledge from the language modality in an interactional alignment paradigm. Specifically, at each stage, a self-attention module is proposed to attend to the key patch from images or text descriptions. Besides, a query-relational module is designed to seize the key words/phrases of texts and further bridge the connection between two modalities. Later, the learned representations of bi-modality from multiple stages are aggregated as the final features for recognition. Our Bi-Modal PMA model only needs raw images and raw text descriptions, without requiring bounding boxes/part annotations in images or key word annotations in texts. By conducting comprehensive experiments on fine-grained benchmark datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines, on either vision and language bi-modality or single visual modality.
Kaitao Song, Xiu-Shen Wei, Xiangbo Shu, Renjie Song, Jianfeng Lu 0003
IEEE Trans. Image Process.2
2019 Multi-Label Image Recognition With Graph Convolutional Networks
abstract
The task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important dependencies, we propose a multi-label classification model based on Graph Convolutional Network (GCN). The model builds a directed graph over the object labels, where each node (label) is represented by word embeddings of a label, and GCN is learned to map this label graph into a set of inter-dependent object classifiers. These classifiers are applied to the image descriptors extracted by another sub-net, enabling the whole network to be end-to-end trainable. Furthermore, we propose a novel re-weighted scheme to create an effective label correlation matrix to guide information propagation among the nodes in GCN. Experiments on two multi-label image recognition datasets show that our approach obviously outperforms other existing state-of-the-art methods. In addition, visualization analyses reveal that the classifiers learned by our model maintain meaningful semantic topology.
Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001
CVPR2
2019 Multi-Label Image Recognition with Joint Class-Aware Map Disentangling and Label Correlation Embedding
abstract
Multi-label image recognition is a fundamental but challenging computer vision task. Great progress has been achieved by exploring the label correlation among these multiple labels which is the most crucial issue for multi-label recognition. In this paper, we propose a unified deep learning framework to jointly disentangle class-specific maps corresponding to discriminative category-wise information and then evaluate the label co-occurrence of these maps. Specifically, after obtaining the general deep image features and conducting multi-label classification, we employ the classification weights to reform the feature maps into class-aware disentangled maps (CADMs). Then, based on CADMs, we first transfer them into label vectors and then formulate the label correlation dependency from an embedding perspective. The whole model is driven by both the classification loss and the label correlation embedding loss, which is end-to-end trainable with only image-level supervisions. Extensive quantitative results of two benchmark multi-label image datasets show our model consistently outperforms other competing methods by a large margin. Meanwhile, qualitative analyses also demonstrate our model can effectively capture relatively pure class-aware maps and model label correlation dependency as well.
Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001
ICME2
2019 Unsupervised object discovery and co-localization by deep descriptor transformation
Xiu-Shen Wei, Chen-Lin Zhang, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
Pattern Recognit.1
2019 Piecewise Classifier Mappings: Learning Fine-Grained Learners for Novel Categories With Few Examples
abstract
Humans are capable of learning a new fine-grained concept with very little supervision, e.g., few exemplary images for a species of bird, yet our best deep learning systems need hundreds or thousands of labeled examples. In this paper, we try to reduce this gap by studying the fine-grained image recognition problem in a challenging few-shot learning setting, termed few-shot fine-grained recognition (FSFG). The task of FSFG requires the learning systems to build classifiers for the novel fine-grained categories from few examples (only one or less than five). To solve this problem, we propose an end-to-end trainable deep network, which is inspired by the state-of-the-art fine-grained recognition model and is tailored for the FSFG task. Specifically, our network consists of a bilinear feature learning module and a classifier mapping module: while the former encodes the discriminative information of an exemplar image into a feature vector, the latter maps the intermediate feature into the decision boundary of the novel category. The key novelty of our model is a "piecewise mappings" function in the classifier mapping module, which generates the decision boundary via learning a set of more attainable sub-classifiers in a more parameter-economic way. We learn the exemplar-to-classifier mapping based on an auxiliary dataset in a meta-learning fashion, which is expected to be able to generalize to novel categories. By conducting comprehensive experiments on three fine-grained datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines.
Xiu-Shen Wei, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
IEEE Trans. Image Process.1
2018 Coarse-to-Fine: A RNN-Based Hierarchical Attention Model for Vehicle Re-identification
Xiu-Shen Wei, Chen-Lin Zhang, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
ACCV (2)1
2018 Mask-CNN: Localizing parts and selecting descriptors for fine-grained bird species categorization
Xiu-Shen Wei, Chen-Wei Xie, Jianxin Wu 0001, Chunhua Shen
Pattern Recognit.1
2018 Deep Bimodal Regression of Apparent Personality Traits from Short Video Sequences
abstract
Apparent personality analysis (APA) is an important problem of personality computing, and furthermore, automatic APA becomes a hot and challenging topic in computer vision and multimedia. In this paper, we propose a deep learning solution to APA from short video sequences. In order to capture rich information from both the visual and audio modality of videos, we tackle these tasks with our Deep Bimodal Regression (DBR) framework. In DBR, for the visual modality, we modify the traditional convolutional neural networks for exploiting important visual cues. In addition, taking into account the model efficiency, we extract audio representations and build a linear regressor for the audio modality. For combining the complementary information from the two modalities, we ensemble these predicted regression scores by both early fusion and late fusion. Finally, based on the proposed framework, we come up with a solution for the Apparent Personality Analysis competition track in the ChaLearn Looking at People challenge in association with ECCV 2016. Our DBR is the winner (first place) of this challenge with 86 registered participants. Beyond the competition, we further investigate the performance of different loss functions in our visual models, and prove non-convex loss functions for regression are optimal on the human-labeled video data.
Xiu-Shen Wei, Chen-Lin Zhang, Hao Zhang 0038, Jianxin Wu 0001
IEEE Trans. Affect. Comput.1
2017 Adversarial PoseNet: A Structure-Aware Convolutional Network for Human Pose Estimation
abstract
For human pose estimation in monocular images, joint occlusions and overlapping upon human bodies often result in deviated pose predictions. Under these circumstances, biologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geometric constraints of joint inter-connectivity. To address the problem by incorporating priors about the structure of human bodies, we propose a novel structure-aware convolutional network to implicitly take such priors into account during training of the deep network. Explicit learning of such constraints is typically challenging. Instead, we design discriminators to distinguish the real poses from the fake ones (such as biologically implausible ones). If the pose generator (G) generates results that the discriminator fails to distinguish from real ones, the network successfully learns the priors.,,To better capture the structure dependency of human body joints, the generator G is designed in a stacked multi-task manner to predict poses as well as occlusion heatmaps. Then, the pose and occlusion heatmaps are sent to the discriminators to predict the likelihood of the pose being real. Training of the network follows the strategy of conditional Generative Adversarial Networks (GANs). The effectiveness of the proposed network is evaluated on two widely used human pose estimation benchmark datasets. Our approach significantly outperforms the state-of-the-art methods and almost always generates plausible human pose predictions.
Yu Chen 0037, Chunhua Shen, Xiu-Shen Wei, Lingqiao Liu, Jian Yang 0003
ICCV3
2017 Deep Descriptor Transforming for Image Co-Localization
abstract
Reusable model design becomes desirable with the rapid expansion of machine learning applications. In this paper, we focus on the reusability of pre-trained deep convolutional models. Specifically, different from treating pre-trained models as feature extractors, we reveal more treasures beneath convolutional layers, i.e., the convolutional activations could act as a detector for the common object in the image co-localization problem. We propose a simple but effective method, named Deep Descriptor Transforming (DDT), for evaluating the correlations of descriptors and then obtaining the category-consistent regions, which can accurately locate the common object in a set of images. Empirical studies validate the effectiveness of the proposed DDT method. On benchmark image co-localization datasets, DDT consistently outperforms existing state-of-the-art methods by a large margin. Moreover, DDT also demonstrates good generalization ability for unseen categories and robustness for dealing with noisy data.
Xiu-Shen Wei, Chen-Lin Zhang, Yao Li 0003, Chen-Wei Xie, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
IJCAI1
2017 Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
abstract
Deep convolutional neural network models pre-trained for the ImageNet classification task have been successfully adopted to tasks in other domains, such as texture description and object proposal generation, but these tasks require annotations for images in the new domain. In this paper, we focus on a novel and challenging task in the pure unsupervised setting: fine-grained image retrieval. Even with image labels, fine-grained images are difficult to classify, letting alone the unsupervised retrieval task. We propose the selective convolutional descriptor aggregation (SCDA) method. The SCDA first localizes the main object in fine-grained images, a step that discards the noisy background and keeps useful deep descriptors. The selected descriptors are then aggregated and the dimensionality is reduced into a short feature vector using the best practices we found. The SCDA is unsupervised, using no image label or bounding box annotation. Experiments on six fine-grained data sets confirm the effectiveness of the SCDA for fine-grained image retrieval. Besides, visualization of the SCDA features shows that they correspond to visual attributes (even subtle ones), which might explain SCDA's high-mean average precision in fine-grained retrieval. Moreover, on general image retrieval data sets, the SCDA achieves comparable retrieval results with the state-of-the-art general image retrieval approaches.
Xiu-Shen Wei, Jian-Hao Luo, Jianxin Wu 0001, Zhi-Hua Zhou
IEEE Trans. Image Process.1
2017 Scalable Algorithms for Multi-Instance Learning
abstract
Multi-instance learning (MIL) has been widely applied to diverse applications involving complicated data objects, such as images and genes. However, most existing MIL algorithms can only handle small- or moderate-sized data. In order to deal with large-scale MIL problems, we propose MIL based on the vector of locally aggregated descriptors representation (miVLAD) and MIL based on the Fisher vector representation (miFV), two efficient and scalable MIL algorithms. They map the original MIL bags into new vector representations using their corresponding mapping functions. The new feature representations keep essential bag-level information, and at the same time lead to excellent MIL performances even when linear classifiers are used. Thanks to the low computational cost in the mapping step and the scalability of linear classifiers, miVLAD and miFV can handle large-scale MIL data efficiently and effectively. Experiments show that miVLAD and miFV not only achieve comparable accuracy rates with the state-of-the-art MIL algorithms, but also have hundreds of times faster speed. Moreover, we can regard the new miVLAD and miFV representations as multiview data, which improves the accuracy rates in most cases. In addition, our algorithms perform well even when they are used without parameter tuning (i.e., adopting the default parameters), which is convenient for practical MIL applications.
Xiu-Shen Wei, Jianxin Wu 0001, Zhi-Hua Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2016 An empirical study on image bag generators for multi-instance learning
Xiu-Shen Wei, Zhi-Hua Zhou
Mach. Learn.1
2016 Weakly Supervised Fine-Grained Categorization With Part-Based Image Representation
abstract
In this paper, we propose a fine-grained image categorization system with easy deployment. We do not use any object/part annotation (weakly supervised) in the training or in the testing stage, but only class labels for training images. Fine-grained image categorization aims to classify objects with only subtle distinctions (e.g., two breeds of dogs that look alike). Most existing works heavily rely on object/part detectors to build the correspondence between object parts, which require accurate object or object part annotations at least for training images. The need for expensive object annotations prevents the wide usage of these methods. Instead, we propose to generate multi-scale part proposals from object proposals, select useful part proposals, and use them to compute a global image representation for categorization. This is specially designed for the weakly supervised fine-grained categorization task, because useful parts have been shown to play a critical role in existing annotation-dependent works, but accurate part detectors are hard to acquire. With the proposed image representation, we can further detect and visualize the key (most discriminative) parts in objects of different classes. In the experiments, the proposed weakly supervised method achieves comparable or better accuracy than the state-of-the-art weakly supervised methods and most existing annotation-dependent methods on three challenging datasets. Its success suggests that it is not always necessary to learn expensive object/part detectors in fine-grained image categorization.
Yu Zhang 0004, Xiu-Shen Wei, Jianxin Wu 0001, Jianfei Cai 0001, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.2
2014 Scalable Multi-instance Learning
abstract
Multi-instance learning (MIL) has been widely applied to diverse applications involving complicated data objects such as images and genes. However, most existing MIL algorithms can only handle small-or moderate-sized data. In order to deal with the large scale problems in MIL, we propose an efficient and scalable MIL algorithm named miFV. Our algorithm maps the original MIL bags into a new feature vector representation, which can obtain bag-level information, and meanwhile lead to excellent performances even with linear classifiers. In consequence, thanks to the low computational cost in the mapping step and the scalability of linear classifiers, miFV can handle large scale MIL data efficiently and effectively. Experiments show that miFV not only achieves comparable accuracy rates with state-of-the-art MIL algorithms, but has hundreds of times faster speed than other MIL algorithms.
Xiu-Shen Wei, Jianxin Wu 0001, Zhi-Hua Zhou
ICDM1