VLDB 2026 Research / reviewers in the wild / expert
Xuekuan Wang
dblp:187/1647
· DBLP profile ↗
15ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-5741-2784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Phase Visual-Language Pretraining and Adaptation for Long-Tailed Multi-Label RecognitionabstractLong-Tailed Multi-Label Recognition (LTML) is a critical yet challenging task due to two core issues: the severe scarcity of training samples for rare "tail" classes, and the complex co-occurrence patterns among labels that often lead to biased models. To address this, we propose DP-VLPA, a novel Dual-Phase Visual-Language Pretraining and Adaptation framework. In the first phase, our Structured Tail-Aware Generation (STAG) module employs a Large Language Model (LLM) to create detailed descriptions that explicitly emphasize tail classes and their contextual relationships, providing a strong and less-biased feature foundation. In the second adaptation phase, we ensure this knowledge is applied effectively. A Dynamic Query Reweighting (DQR) mechanism forces the model to attend to crucial tail-class evidence. Simultaneously, a Co-occurrence-Aware (COA) loss explicitly teaches the model the statistical dependencies between labels, correcting for co-occurrence biases. Extensive experiments on VOC-LT and COCO-LT datasets demonstrate state-of-the-art performance, achieving mAP scores of 90.72% and 74.42% respectively - surpassing previous best methods by 2.84% and 8.23%. Xuekuan Wang, Cairong Zhao |
AAAI | 2 |
| 2026 | DiffPano++: Scalable and Consistent Multi-View Panorama Generation with Spherical Epipolar-Aware Diffusion
Chenhao Ji, Weicai Ye, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Xuekuan Wang, Guofeng Zhang 0001, Song-Hai Zhang, Tong He 0001, Wanli Ouyang, Cairong Zhao |
Int. J. Comput. Vis. | 6 |
| 2026 | SC-DETR: A Text-Guided Small Object Detection via Scale Prompting and Centerpoint Localization
Mingzhu Li, Xuekuan Wang, Cairong Zhao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Fusion4DAL: Offline Multi-modal 3D Object Detection for 4D Auto-labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jincheng Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao |
Int. J. Comput. Vis. | 2 |
| 2025 | Identity aware 3D face reconstruction from in-the-wild images
Ruigang Hu, Xuekuan Wang, Cairong Zhao |
Neurocomputing | 2 |
| 2025 | TGAvatar: Reconstructing 3D Gaussian Avatars With Transformer-Based Tri-PlaneabstractWe introduce TGAvatar, a novel framework for 3D head animation and reconstruction that revolutionizes the use of 3D Gaussian Splatting (3DGS). TGAvatar significantly advances rendering quality by leveraging the intricate properties of 3DGS to achieve detailed and realistic representations of human head geometries and textures. We use an innovative application of linear blending techniques to imitate 3D Morphable Model (3DMM) coefficients within 3DGS, thereby enabling precise and dynamic facial feature and expression modeling. Further enhancing TGAvatar’s capabilities, a transformer based tri-plane module is incorporated to accurately infer spherical harmonics and alpha parameters. This integration is pivotal for the method, as it allows allows us to efficiently and precisely represent the visual characteristics of gaussians, tailored specifically to the intricate details of the head’s components. Our exhaustive evaluations show that TGAvatar not only elevates the fidelity and realism of 3D head reconstructions but also sets a new standard by surpassing existing methods in rendering quality and computational efficiency. Please see our project page athttps://hrg0417.github.io/TGAvatar/ Ruigang Hu, Xuekuan Wang, Yichao Yan, Cairong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual LossabstractText Super-Resolution (SR) technology aims to recover lost information in low-resolution text images. With the proposal of TextZoom, which is the first dataset aiming at text super-resolution in real scenes, more and more scene text super-resolution models have been presented on the basis of it. Although these methods have achieved excellent performance, they do not consider how to make full and efficient use of semantic information. Out of this consideration, a Semantic-aware Trident Network (STNet) for Scene Text Image Super-Resolution is proposed. Specifically, pre-trained text recognition model ASTER (Attentional Scene Text Recognizer) is utilized to assist this process in two ways. Firstly, a novel basic block named Semantic-aware Trident Block (STB) is designed to build the STNet, which incorporates an added branch for semantic distillation to learn semantic information of pre-trained recognition model. Secondly, we expand our model in an adversarial training manner and propose new text perceptual loss based on ASTER to further enhance semantic information in SR images. Extensive experiments on TextZoom dataset show that compared with directly recognizing bicubic images, the proposed STNet boosts the recognition accuracy of ASTER, MORAN (Multi-Object Rectified Attention Network), and CRNN (Convolutional Recurrent Neural Network) by 17.4%, 18.2%, and 24.3%, respectively, which is higher than the performance of several existing state-of-the-art (SOTA) SR network models. Besides, experiments in real scenes (on ICDAR 2015 dataset) and in restricted scenarios (defense against adversarial attacks) validate that addition of semantic information enables the proposed method to achieve promising cross-dataset performance. Since the proposed method is trained on cropped images, when applied to real-world scenarios, locations of text in natural images are firstly localized through scene text detection methods, and then cropped text images are obtained based on detected text positions. Cairong Zhao, Shuyang Feng, Xuekuan Wang |
IEEE Trans. Multim. | 5 |
| 2024 | Interactive 3D Object Detection with Prompts
Rui Zhang 0003, Xiangru Lin, Wei Zhang 0197, Jincheng Lu, Xuekuan Wang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Guanbin Li |
ECCV (17) | 5 |
| 2024 | Uni4DAL: A Unified Baseline for Multi-dataset 4D Auto-Labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jinchen Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao |
ICPR (30) | 2 |
| 2020 | Similarity learning with joint transfer constraints for person re-identification
Cairong Zhao, Xuekuan Wang, Wangmeng Zuo, Fumin Shen, Ling Shao 0001, Duoqian Miao 0001 |
Pattern Recognit. | 2 |
| 2018 | Kernelized random KISS metric learning for person re-identification
Cairong Zhao, Yipeng Chen, Xuekuan Wang, Wai Keung Wong, Duoqian Miao 0001, Jingsheng Lei |
Neurocomputing | 3 |
| 2018 | Maximal granularity structure and generalized multi-view discriminant analysis for person re-identification
Cairong Zhao, Xuekuan Wang, Duoqian Miao 0001, Hanli Wang, Wei-Shi Zheng 0001, Yong Xu 0001, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2017 | Multiple metric learning based on bar-shape descriptor for person re-identification
Cairong Zhao, Xuekuan Wang, Wai Keung Wong, Wei-Shi Zheng 0001, Jian Yang 0003, Duoqian Miao 0001 |
Pattern Recognit. | 2 |
| 2016 | Mutli-channel micro-structure difference descriptor for image retrievalabstractThis paper presents a novel image feature representation method, called multi-channel micro-structure difference descriptor (MCMSDD) for image retrieval. With the local feature extraction from a micro-structure and MAX operator, MCMSDD integrates the advantages of multi-channel local binary encoding and color difference histogram , which are the fusion of color, texture and spatial distribution information. Although it extracts feature from full color image, the dimension of the feature vector is relatively low without learning and segmentation. To improve the performance of retrieval, a simple re-ranking algorithm is employed. Finally, the proposed MCMSDD is extensively tested on Corel-2K and Washington datasets, and the experimental results show that the proposed MCMSDD is more effective than the state-of-the-art. Xuekuan Wang, Cairong Zhao, Duoqian Miao 0001, Cuijun Liu, Yipeng Chen, Zhihui Lai 0001 |
ICPR | 1 |
| 2016 | Fusion of multiple channel features for person re-identification
Xuekuan Wang, Cairong Zhao, Duoqian Miao 0001, Zhihua Wei 0001, Renxian Zhang, Tingfei Ye |
Neurocomputing | 1 |