Xuhao Sun

dblp:305/2467 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-6642-6966ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 42% Learning theory · 31% Learning paradigms · 15%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › image retrieval › content-based image retrieval
fine-grained image retrieval
1.732023
Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2023
SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval · ECCV (14) 2022
A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval · NeurIPS 2021
Information retrieval
image retrieval
1.732023
Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2023
SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval · ECCV (14) 2022
A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval · NeurIPS 2021
Machine learning › Learning theory
classification
0.912025
Equiangular Basis Vectors: A Novel Paradigm for Classification Tasks · Int. J. Comput. Vis. 2025
Machine learning › Learning paradigms
long-tailed recognition
0.912025
Delving Deep into Simplicity Bias for Long-Tailed Image Recognition · Int. J. Comput. Vis. 2025
Machine learning › Learning theory › inductive bias
simplicity bias
0.912025
Delving Deep into Simplicity Bias for Long-Tailed Image Recognition · Int. J. Comput. Vis. 2025
Machine learning › Representation and self-supervised learning › hashing
deep hashing
0.712023
Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.712023
Equiangular Basis Vectors · CVPR 2023
Information retrieval › hashing
hashing-based retrieval
0.712023
Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning
hashing
0.612022
SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval · ECCV (14) 2022
Information retrieval
hashing
0.512021
A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

self-consistency · 1.3encoder-decoder · 1.3attention · 1.3learning to hash · 1.1long-tailed recognition · 0.9equiangular basis vectors · 0.9deep learning · 0.9spherical distance minimization · 0.7softmax classifier · 0.7orthogonal embedding · 0.7similarity preservation · 0.5feature decorrelation · 0.5attention-based encoder-decoder · 0.5
YearPublicationVenuePosition
2025 Equiangular Basis Vectors: A Novel Paradigm for Classification Tasks
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Lingyan Gao
Int. J. Comput. Vis.2
2025 Delving Deep into Simplicity Bias for Long-Tailed Image Recognition
Xiu-Shen Wei, Xuhao Sun, Yang Shen 0006, Peng Wang 0023
Int. J. Comput. Vis.2
2023 Equiangular Basis Vectors
abstract
We propose Equiangular Basis Vectors (EBVs) for classification tasks. In deep neural networks, models usually end with a k-way fully connected layer with softmax to handle different classification tasks. The learning objective of these methods can be summarized as mapping the learned feature representations to the samples' label space. While in metric learning approaches, the main objective is to learn a transformation function that maps training data points from the original space to a new space where similar points are closer while dissimilar points become farther apart. Different from previous methods, our EBVs generate normalized vector embeddings as “predefined classifiers” which are required to not only be with the equal status between each other, but also be as orthogonal as possible. By minimizing the spherical distance of the embedding of an input between its categorical EBV in training, the predictions can be obtained by identifying the categorical EBV with the smallest distance during inference. Various experiments on the ImageNet-1K dataset and other downstream tasks demonstrate that our method outperforms the general fully connected classifier while it does not introduce huge additional computation compared with classical metric learning methods. Our EBVs won the first place in the 2022 DIGIX Global AI Challenge, and our code is open-source and available at https://github.com/NJUST-VIPGroup/Equiangular-Basis-Vectors.
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei
CVPR2
2023 Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval
abstract
Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose attribute-aware hashing networks with self-consistency for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. Our models are also equipped with a feature decorrelation constraint upon these attribute vectors to strengthen their representative abilities. Then, driven by preserving original entities' similarity, the required hash codes can be generated from these attribute-specific vectors and thus become attribute-aware. Furthermore, to combat simplicity bias in deep hashing, we consider the model design from the perspective of the self-consistency principle and propose to further enhance models' self-consistency by equipping an additional image reconstruction path. Comprehensive quantitative experiments under diverse empirical settings on six fine-grained retrieval datasets and two generic retrieval datasets show the superiority of our models over competing methods. Moreover, qualitative results demonstrate that not only the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects, but also our self-consistency designs can effectively overcome simplicity bias in fine-grained hashing.
Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Peng Wang 0023, Yuxin Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Qing-Yuan Jiang, Jian Yang 0003
ECCV (14)2
2022 A Channel Mix Method for Fine-Grained Cross-Modal Retrieval
abstract
In this paper, we propose a simple but effective method for dealing with the challenging fine-grained cross-modal retrieval task where it aims to enable flexible retrieval among subor-dinate categories across different modalities. Specifically, in order to enhance information interaction in different modalities for fine-grained objects, a channel mix method is developed and performed upon the channels of deep activations across dif-ferent modalities. After that, a 1 x 1 convolution is employed to aggregate the mixed channels into a unified feature vector. Moreover, equipped with a novel fine-grained cross-modal cen-ter loss, our method can further improve the intra-class separa-bility as well as inter-class compactness for multi-modalities. Experiments are conducted on the fine-grained cross-modal benchmark dataset and show our superiority over competing methods. Meanwhile, ablation studies also demonstrate the effectiveness of our proposals.
Yang Shen 0006, Xuhao Sun, Xiu-Shen Wei, Hanxu Hu
ICME2
2021 A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval
abstract
Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose an Attribute-Aware hashing Network (A$^2$-Net) for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. A$^2$-Net is also equipped with a feature decorrelation constraint upon these attribute vectors to enhance their representation abilities. Finally, the required hash codes are generated by the attribute vectors driven by preserving original similarities. Qualitative experiments on five benchmark fine-grained datasets show our superiority over competing methods. More importantly, quantitative results demonstrate the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects.
Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Han-Jia Ye, Jian Yang 0003
NeurIPS3