VLDB 2026 Research / reviewers in the wild / expert
Can Xu 0006
dblp:33/965-6
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-Modal Contrast with Image Jigsaw for Self-Supervised Representation Learning of 3D Point CloudsabstractContrastive learning has shown impressive progress for the self-supervision based 3D point clouds feature learning. Based on the distance control in the feature space of positive and negative samples, it can obtain effective point cloud feature representations in a self-supervised manner. However, most existing contrast based point cloud learning methods only consider the feature similarity relationship between samples, (e.g., point cloud, voxel or image), which lack of explicit exploration on point cloud structure. Considering that structure is an important property of point clouds, for better feature learning, we propose a effective cross-modal contrast based method with image jigsaw (CrossCon-Jig) to better learn point cloud representations with both semantic and structural information. Specifically, our method includes intra-modal contrast of point cloud, cross-modal contrast between point cloud and rendered image, and point cloud guided image jigsaw. The intra-modal contrast and the contrast of cross-modal focus on the exploring of invariant and consistent feature representations, and image jigsaw guides the model to explore spatial structure information of point clouds. Extensive experimental tests on 3D object classification and 3D object part segmentation tasks have achieved excellent performance, demonstrating the effectiveness of the proposed method. Yuehui Han, Xinpeng Yu, Can Xu 0006, Qi Liu 0001 |
QRS | 4 |
| 2025 | Graph-in-graph discriminative feature enhancement network for fine-grained visual classification
Yupeng Wang 0004, Can Xu 0006, Yongli Wang 0002, Xiaoli Wang 0003, Weiping Ding 0001 |
Appl. Intell. | 2 |
| 2025 | Strengthen contrastive semantic consistency for fine-grained image classification
Yupeng Wang 0004, Yongli Wang 0002, Qiaolin Ye, Wenxi Lang, Can Xu 0006 |
Pattern Anal. Appl. | 5 |
| 2025 | Weakly Supervised Object Localization With Progressive Activation DiffusionabstractWeakly supervised object localization (WSOL) aims to locate objects with only image-level labels. Previous works mainly follow the framework of class activation map (CAM), which discovers the objects by estimating the contribution of each pixel position to the category prediction. However, most of them overlook the pixel-level spatial and semantic contextual correlation, resulting in: 1) limited activation ranges that only highlight the most discriminative parts rather than the entire object and 2) low activation values for some foreground parts, especially regions near the boundary between foreground and background. To alleviate this issue, we propose an activation diffusion network (ADNet) to progressively refine both the range and value of activations on the localization map. Specifically, a context propagation module is first developed to learn the top-down spatial dependency between adjacent feature maps, which helps back-propagate the activation from the discriminative part to its surroundings for more complete objects. Then, a diffusion probability distillation module (DPDM) is proposed, which transfers the pixel-level semantic correlation emerging in the image generation process to the localization map generation in a teacher-student learning manner. This helps boost the value of the activated foreground region and stimulates the value of neighboring inactivated foreground positions to sharpen the object boundary. Experiments on various datasets and backbones demonstrate the superiority of our ADNet over state-of-the-art (SOTA) methods in object localization and segmentation, yielding 82.2% and 62.2% Top-1 Loc on Caltech-UCSD Birds-200-2011 (CUB) and ImageNet Large-ScaleVisual Recognition Challenge (ILSVRC) datasets and 76.6% pixel average precision (PxAP) on OpenImages dataset. Qualitative results also show that we can achieve a more complete and consistent activation covering the whole object. Can Xu 0006, Le Hui, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multi-Attribute Interactions Matter for 3D Visual Groundingabstract3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm, existing methods mainly focus on embedding object attributes in unimodal feature extraction and multimodal feature fusion, to enhance the discriminability of the proposal feature for accurate grounding. However, most of them ignore the explicit interaction of multiple attributes, causing a bias in unimodal representation and misalignment in multimodal fusion. In this paper, we propose a multi-attribute aware Transformer for 3D visual grounding, learning the multi-attribute interactions to refine the intra-modal and inter-modal grounding cues. Specifically, we first develop an attribute causal analysis module to quantify the causal effect of different attributes for the final prediction, which provides powerful supervision to correct the misleading attributes and adaptively capture other discriminative features. Then, we design an exchanging-based multimodal fusion module, which dynamically replaces tokens with low attribute attention between modalities before directly integrating low-dimensional global features. This ensures an attribute-level multimodal information fusion and helps align the language and vision details more efficiently for fine-grained multimodal features. Extensive experiments show that our method can achieve state-of-the-art performance on ScanRefer and Sr3D/Nr3D datasets. The code is publicly available at https://github.com/volcanoXC/MA2TransVG. Can Xu 0006, Yuehui Han, Rui Xu 0021, Le Hui, Jin Xie 0001, Jian Yang 0003 |
CVPR | 1 |
| 2024 | Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning
Yuehui Han, Can Xu 0006, Rui Xu 0021, Jianjun Qian, Jin Xie 0001 |
ECCV (76) | 2 |
| 2024 | Learning Local Semantic Region Activations for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to train instance-level locators by exploiting accessible image-level labels. By multiplying channel-wise features with classification weights and then adding them together, most prior works follow the pipeline of the Class Activation Map (CAM) to collect the semantic responses, thereby highlighting regions that contribute to class prediction to achieve WSOL. However, CAM-based methods treat the class contributions of all pixel positions in a channel equally and assign dominant weights for the discriminative channels biasedly. This fails to express the fine-grained pixel-level semantic response of each channel and model the complex contextual relations between channels, resulting in the mixup of the activation value between non-discriminative foreground regions and the background. To alleviate these issues, we present a Local Semantic activation enhancement and Global Spatial correlation mining network (LSGS-Net) for accurate WSOL. Specifically, we first propose a local activation generation module to explicitly learn the semantic response of each pixel position from channels. Then, we design a regularization loss to supervise the consistency between similar local activations, which utilizes the cross-image information to improve the accuracy of local activations. We further propose a K-nearest Neighbors graph module to capture the spatial correlation between different local activations, which can adaptively assign more proper weights when fusing all local activation. In the inference stage, the bounding box will be determined with a foreground threshold. Extensive experiments show that LSGS-Net achieves significant and consistent improvement with various backbones on the CUB, ILSVRC, and OpenImages benchmarks, with a 97.5% and 75.3% GT-Known LOC on CUB and ILSVRC, respectively. For segmentation quality on OpenImages, LSGS-Net already exceeds the SOTA method by 1.2% pIoU and 1.9% PxAP. Can Xu 0006, Le Hui, Yuehui Han, Haobo Jiang, Jiaxin Chen 0001, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Computer vision-driven forest wildfire and smoke recognition via IoT drone cameras
Yupeng Wang 0004, Yongli Wang 0002, Can Xu 0006, Xiaoli Wang 0003 |
Wirel. Networks | 3 |
| 2023 | Corrigendum to "Generative detect for occlusion object based on occlusion generation and feature completing" [J. Visual Commun. Image Represent. 78 (2021) 103189]
Can Xu 0006, Wenxi Lang, Kaichen Mao |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Graph-based discriminative features learning for fine-grained image retrieval
Wenxi Lang, Can Xu 0006, Ningzhong Liu, Huiyu Zhou 0001 |
Signal Process. Image Commun. | 3 |
| 2022 | Discriminative feature mining hashing for fine-grained image retrieval
Wenxi Lang, Can Xu 0006, Ningzhong Liu, Huiyu Zhou 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Generative detect for occlusion object based on occlusion generation and feature completing
Can Xu 0006, Peter W. T. Yuen, Wenxi Lang, Kaichen Mao |
J. Vis. Commun. Image Represent. | 1 |