Yuehui Han

dblp:228/4058 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Self-Supervised Vision Graph Neural Networks Based on Contrastive Learning
abstract
In the field of computer vision, Vision Graph Neural Networks (ViG) have demonstrated significant potential in image understanding. By treating the divided image patches as nodes and constructing connection relationships based on neighbor attributes, ViG can efficiently model global dependencies within images with the help of graph attributes. However, most existing ViG methods have the problem of high computational complexity in graph construction and may not be able to effectively and fully explore the graph structure information. Besides, the heavy reliance on manual annotation labels limits the application potential of ViG in practical scenarios. To this end, in this paper, we propose a novel self-supervised vision graph contrastive learning method (S2ViG) based on image mixing strategy for efficient vision graph representation learning. It aims to use self-supervised method to alleviate the dependence on manual annotation and enhance the understanding of the global structure of the graph using two different vision graph construction methods. Specifically, we first employ image mixing strategy to uncover latent semantic relationships among multiple images. Then, we construct dynamic graph structures for image patches from local and global perspectives to obtain augmented contrastive samples. Finally, the multilevel contrastive loss is constructed to optimize the network. Experimental results show that our method achieves excellent performance on multiple datasets such as ImageNet-1K and CIFAR.
Yuehui Han, Jianjun Qian, Jian Yang 0003
ACM Multimedia2
2025 Cross-Modal Contrast with Image Jigsaw for Self-Supervised Representation Learning of 3D Point Clouds
abstract
Contrastive learning has shown impressive progress for the self-supervision based 3D point clouds feature learning. Based on the distance control in the feature space of positive and negative samples, it can obtain effective point cloud feature representations in a self-supervised manner. However, most existing contrast based point cloud learning methods only consider the feature similarity relationship between samples, (e.g., point cloud, voxel or image), which lack of explicit exploration on point cloud structure. Considering that structure is an important property of point clouds, for better feature learning, we propose a effective cross-modal contrast based method with image jigsaw (CrossCon-Jig) to better learn point cloud representations with both semantic and structural information. Specifically, our method includes intra-modal contrast of point cloud, cross-modal contrast between point cloud and rendered image, and point cloud guided image jigsaw. The intra-modal contrast and the contrast of cross-modal focus on the exploring of invariant and consistent feature representations, and image jigsaw guides the model to explore spatial structure information of point clouds. Extensive experimental tests on 3D object classification and 3D object part segmentation tasks have achieved excellent performance, demonstrating the effectiveness of the proposed method.
Yuehui Han, Xinpeng Yu, Can Xu 0006, Qi Liu 0001
QRS1
2025 RagNet3D: Learning distinguishable representation for pooled grids in 3D object detection
Jiaxin Chen 0001, Yuehui Han, Zhiqiang Yan 0001, Jianjun Qian, Jun Li 0027, Jian Yang 0003
Neurocomputing2
2024 Multi-Attribute Interactions Matter for 3D Visual Grounding
abstract
3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm, existing methods mainly focus on embedding object attributes in unimodal feature extraction and multimodal feature fusion, to enhance the discriminability of the proposal feature for accurate grounding. However, most of them ignore the explicit interaction of multiple attributes, causing a bias in unimodal representation and misalignment in multimodal fusion. In this paper, we propose a multi-attribute aware Transformer for 3D visual grounding, learning the multi-attribute interactions to refine the intra-modal and inter-modal grounding cues. Specifically, we first develop an attribute causal analysis module to quantify the causal effect of different attributes for the final prediction, which provides powerful supervision to correct the misleading attributes and adaptively capture other discriminative features. Then, we design an exchanging-based multimodal fusion module, which dynamically replaces tokens with low attribute attention between modalities before directly integrating low-dimensional global features. This ensures an attribute-level multimodal information fusion and helps align the language and vision details more efficiently for fine-grained multimodal features. Extensive experiments show that our method can achieve state-of-the-art performance on ScanRefer and Sr3D/Nr3D datasets. The code is publicly available at https://github.com/volcanoXC/MA2TransVG.
Can Xu 0006, Yuehui Han, Rui Xu 0021, Le Hui, Jin Xie 0001, Jian Yang 0003
CVPR2
2024 Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning
Yuehui Han, Can Xu 0006, Rui Xu 0021, Jianjun Qian, Jin Xie 0001
ECCV (76)1
2024 Learning Local Semantic Region Activations for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) aims to train instance-level locators by exploiting accessible image-level labels. By multiplying channel-wise features with classification weights and then adding them together, most prior works follow the pipeline of the Class Activation Map (CAM) to collect the semantic responses, thereby highlighting regions that contribute to class prediction to achieve WSOL. However, CAM-based methods treat the class contributions of all pixel positions in a channel equally and assign dominant weights for the discriminative channels biasedly. This fails to express the fine-grained pixel-level semantic response of each channel and model the complex contextual relations between channels, resulting in the mixup of the activation value between non-discriminative foreground regions and the background. To alleviate these issues, we present a Local Semantic activation enhancement and Global Spatial correlation mining network (LSGS-Net) for accurate WSOL. Specifically, we first propose a local activation generation module to explicitly learn the semantic response of each pixel position from channels. Then, we design a regularization loss to supervise the consistency between similar local activations, which utilizes the cross-image information to improve the accuracy of local activations. We further propose a K-nearest Neighbors graph module to capture the spatial correlation between different local activations, which can adaptively assign more proper weights when fusing all local activation. In the inference stage, the bounding box will be determined with a foreground threshold. Extensive experiments show that LSGS-Net achieves significant and consistent improvement with various backbones on the CUB, ILSVRC, and OpenImages benchmarks, with a 97.5% and 75.3% GT-Known LOC on CUB and ILSVRC, respectively. For segmentation quality on OpenImages, LSGS-Net already exceeds the SOTA method by 1.2% pIoU and 1.9% PxAP.
Can Xu 0006, Le Hui, Yuehui Han, Haobo Jiang, Jiaxin Chen 0001, Jin Xie 0001, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 Generation-based Multi-view Contrast for Self-supervised Graph Representation Learning
abstract
Graph contrastive learning has made remarkable achievements in the self-supervised representation learning of graph-structured data. By employing perturbation function (i.e., perturbation on the nodes or edges of graph), most graph contrastive learning methods construct contrastive samples on the original graph. However, the perturbation-based data augmentation methods randomly change the inherent information (e.g., attributes or structures) of the graph. Therefore, after nodes embedding on the perturbed graph, we cannot guarantee the validity of the contrastive samples as well as the learned performance of graph contrastive learning. To this end, in this article, we propose a novel generation-based multi-view contrastive learning framework (GMVC) for self-supervised graph representation learning, which generates the contrastive samples based on our generator rather than perturbation function. Specifically, after nodes embedding on the original graph we first employ random walk in the neighborhood to develop multiple relevant node sequences for each anchor node. We then utilize the transformer to generate the representations of relevant contrastive samples of anchor node based on the features and structures of the sampled node sequences. Finally, by maximizing the consistency between the anchor view and the generated views, we force the model to effectively encode graph information into nodes embeddings. We perform extensive experiments of node classification and link prediction tasks on eight benchmark datasets, which verify the effectiveness of our generation-based multi-view graph contrastive learning method.
Yuehui Han
ACM Trans. Knowl. Discov. Data1
2023 Graph Spectral Perturbation for 3D Point Cloud Contrastive Learning
abstract
3D point cloud contrastive learning has attracted increasing attention due to its efficient learning ability. By distinguishing the similarity relationship between positive and negative samples in the feature space, it can learn effective point cloud feature representations without manual annotation. However, most point cloud contrastive learning methods construct contrastive samples by perturbing point clouds in data space or introducing multi-modality/format data, which may be difficult to control the intensity of the perturbation or introduce interference from different modalities/formats. To this end, in this paper, we propose a novel graph spectral perturbation based contrastive learning framework (GSPCon) for efficient and robust self-supervised 3D point cloud representation learning. It aims to perform perturbations in the graph spectral domain to construct contrastive samples of the point cloud. Specifically, we first naturally represent the point cloud as a k-nearest neighbors (KNN) graph, and adaptively transform the coordinates of the points into the graph spectral domain based on the graph Fourier transform (GFT). Then we implement data augmentation in the graph spectral domain by perturbing the spectral representations. Finally, the contrastive samples are generated by employing the inverse graph Fourier transform (IGFT) to transform the augmented spectral representations back to the point clouds. Experimental results show that our method achieves the state-of-the-art performance on various downstream tasks. Source code is available at https://github.com/yh-han/GSPCon.git.
Yuehui Han, Jiaxin Chen 0001, Jianjun Qian, Jin Xie 0001
ACM Multimedia1
2023 Transformer-based Point Cloud Generation Network
abstract
Point cloud generation is an important research topic in 3D computer vision, which can provide high-quality datasets for various downstream tasks. However, efficiently capturing the geometry of point clouds remains a challenging problem due to their irregularities. In this paper, we propose a novel transformer-based 3D point cloud generation network to generate realistic point clouds. Specifically, we first develop a transformer-based interpolation module that utilizes k-nearest neighbors at different scales to learn global and local information about point clouds in the feature space. Based on geometric information, we interpolate new point features to upsample the point cloud features. Then, the upsampled features are used to generate a coarse point cloud with spatial coordinate information. We construct a transformer-based refinement module to enhance the upsampled features in feature space with geometric information in coordinate space. Finally, we use a multi-layer perceptron on the upsampled features to generate the final point cloud. Extensive experiments on ShapeNet and ModelNet demonstrate the effectiveness of our proposed method.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia3
2023 Scene Graph Masked Variational Autoencoders for 3D Scene Generation
abstract
Generating realistic 3D indoor scenes requires a deep understanding of objects and their spatial relationships. However, existing methods often fail to generate realistic 3D scenes due to the limited understanding of object relationships. To tackle this problem, we propose a Scene Graph Masked Variational Auto-Encoder (SG-MVAE) framework that fully captures the relationships between objects to generate more realistic 3D scenes. Specifically, we first introduce a relationship completion module that adaptively learns the missing relationships between objects in the scene graph. To accurately predict the missing relationships, we employ multi-group attention to capture the correlations between the objects with missing relationships and other objects in the scene. After obtaining the complete scene relationships, we mask the relationships between objects and use a decoder to reconstruct the scene. The reconstruction process enhances the model's understanding of relationships, generating more realistic scenes. Extensive experiments on benchmark datasets show that our model outperforms state-of-the-art methods.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia3
2022 Generative Subgraph Contrast for Self-Supervised Graph Representation Learning
Yuehui Han, Le Hui, Haobo Jiang, Jianjun Qian, Jin Xie 0001
ECCV (30)1
2020 Extracting representative user subset of social networks towards user characteristics and topological features
Yuehui Han, An Liu 0002, Zhixu Li, Hongzhi Yin, Wei Chen 0070, Lei Zhao 0001
World Wide Web2
2019 A Combined Approach for the Binarization of Historical Tibetan Document Images
abstract
It is common that historical Tibetan documents belonging to historical collections are poorly preserved and are prone to degradation processes. This causes many challenges that can be addressed by image binarization, the most common of which is stains. A lack of uniform standard datasets makes it difficult to evaluate binarization effects. Motivated by the poor effects and difficulty of evaluating the binarization of historical Tibetan document images, a combined approach is proposed that aims to improve overall performance. The method includes the following parts: first, image generation through standard binarization and the background extraction of color images, which are both used for image synthesis in preparing for an evaluation. Then, preliminary binarization processing is implemented through channel combination in the lab color space and through local binarization. The synthetic images are used to select the coefficient when the channels are combined. Furthermore, Local Binary Pattern (hereafter LBP) and image smoothing is carried out after the combination of the channels to obtain the outline of the text area. Finally, the final binarization image is obtained by combining the preliminary binarization image and the text area contour image. Our method achieved top performance compared to other methods after a large number of synthetic image tests with a variety of background types.
Yuehui Han, Weilan Wang, Huaming Liu
Int. J. Pattern Recognit. Artif. Intell.1
2019 Online Tibetan Handwriting Recognition for Large Character Set on New Databases
abstract
The online handwriting recognition of Tibetan characters is still in its infancy. For further research, an online handwriting database of large Tibetan character set was developed, and a recognition research was carried out on this database as a baseline result. The Northwest Minzu University Online Tibetan Handwriting Database (NMU-OLTHWDB) contains 7240 different types of characters, and the sample number in each type is 5000. The total number of samples is [Formula: see text]. The database covers Tibetan Character Collection, Information Technology Tibetan Coded Character set (Extension Set A), and Information Technology Tibetan Coded Character set (Extension Set B). The characters in the database are composed of 170 types of different components. We studied the online handwritten Tibetan recognition software also, and the character feature extraction, classifier training, and the statistics and analysis of the recognition results on the test set were mainly introduced. The character features included the direction attribute coefficients and spatial combination, and the feature matrix was compressed by Linear Discriminate Analysis (LDA). A quick classifier was designed by a modified quadratic discriminate function (QMQDF), and was trained with 4500 sets of samples. In the large character set, the recognition rates of top 1, top 3, top 5, and top 10 were 75.2%, 89.56%, 93.02%, and 95.96%, respectively. Moreover, an online handwriting recognition system for Tibetan large character set was designed with good performance.
Weilan Wang, Zhengjiang Li, Zhengqi Cai, Xiaobao Lv, Caike Zhaxi, Yuehui Han
Int. J. Pattern Recognit. Artif. Intell.6
2018 Research on the Method of Tibetan Recognition Based on Component Location Information
Yuehui Han, Weilan Wang
PRCV (3)1
2018 Research on Text Line Segmentation of Historical Tibetan Documents Based on the Connected Component Analysis
Weilan Wang, Yuehui Han
PRCV (3)4
2018 A Recognition Method of the Similarity Character for Uchen Script Tibetan Historical Document Based on DNN
Weilan Wang, Yuehui Han, Zhanjun Hao 0001
PRCV (3)5
2018 Extracting Representative User Subset of Social Networks Towards User Characteristics and Topological Features
Yuehui Han, An Liu 0002, Zhixu Li, Hongzhi Yin, Lei Zhao 0001
WISE (1)2