VLDB 2026 Research / reviewers in the wild / expert
Xiaoxuan He
dblp:188/4696
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal Formalization
Yi Yang 0001, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu 0001, Wei Chen 0001 |
ICCV | 2 |
| 2025 | LTB-Solver: Long-tailed Bias Solver for image synthesis of diffusion modelsabstractThough diffusion models have shown the merits of generating high-quality visual data while preserving better diversity in recent studies, they do not generalize well on long-tailed datasets due to the minority classes lacking of diversity and semantic information . To overcome the aforementioned challenges, we first take a closer look at the collapse of tail category patterns under long-tail distributed data and propose an alternative but easy-to-use and effective solution, a L ong- T ailed B ias Solver in diffusion model image synthesis ( LTB-Solver ), which thereby enhances the overall diversity and quality of synthetic samples building upon the properties of the long-tailed distribution training data. Especially, we extract rich generative distribution knowledge of ‘head’ categories within proxy model and transfer the head-tail consistency distance to ‘tail’ categories, enabling the target diffusion model to learn diverse generation preserving inter-sample variation during the diffusion training process. Moreover, we incorporate the minority guidance loss function that better aligns training objectives with sampling behaviors and adjust the loss values for different classes by multiplying them with different weights. Extensive experiments are conducted on various datasets and several state-of-the-art diffusion model frameworks to verify the effectiveness of the proposed method. The results show that our method significantly improves the performance of diffusion models on long-tailed datasets by a large margin. Siming Fu, Xiaoxuan He, Haoji Hu |
Neurocomputing | 2 |
| 2025 | SemiGMMPoint: Semi-supervised point cloud segmentation based on Gaussian mixture models
Xianwei Zhuang, Hualiang Wang, Xiaoxuan He, Siming Fu, Haoji Hu |
Pattern Recognit. | 3 |
| 2024 | Robustness-Guided Image Synthesis for Data-Free QuantizationabstractQuantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks. Jianhong Bai, Huanpeng Chu, Hualiang Wang, Zuozhu Liu, Ruizhe Chen, Xiaoxuan He, Lianrui Mu, Chengfei Cai, Haoji Hu |
AAAI | 7 |
| 2024 | Unified Medical Image Pre-training in Language-Guided Common Semantic Space
Xiaoxuan He, Yifan Yang 0004, Xinyang Jiang, Xufang Luo, Haoji Hu, Siyun Zhao, Dongsheng Li 0002, Yuqing Yang 0001, Lili Qiu |
ECCV (81) | 1 |
| 2023 | Video Surveillance on Mobile Edge Networks: Exploiting Multi-Exit NetworkabstractVideo surveillance systems are playing increasingly important roles in our everyday lives. To get meaningful surveillance information in a timely and accurate manner, it is vital to optimally allocate computation and communication resources for image classification tasks. In this paper, taking face recognition as an example, we propose a novel end-to-edge collaborative computing system based on a multi-exit network to dynamically allocate computation at the front end (the camera sensor) and back end (the mobile edge computing server). With the ∊-greedy algorithm for reinforcement learning, the decision module decides whether to obtain recognition results from earlier exits at the front end or transmit the feature maps to the back end to obtain more accurate results. The module balances recognition accuracy and time overhead under different channel conditions. Experimental results show that the proposed system can significantly save inference time and maintain competitive accuracy in various communication channel conditions. Yuchen Cao 0005, Siming Fu, Xiaoxuan He, Haoji Hu, Hangguan Shan, Lu Yu 0003 |
ICC | 3 |
| 2023 | Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionabstractRecently, large-scale pre-trained vision-language models have presented benefits for alleviating class imbalance in long-tailed recognition. However, the long-tailed data distribution can corrupt the representation space, where the distance between head and tail categories is much larger than the distance between two tail categories. This uneven feature space distribution causes the model to exhibit unclear and inseparable decision boundaries on the uniformly distributed test set, which lowers its performance. To address these challenges, we propose the uniformly category prototype-guided vision-language framework to effectively mitigate feature space bias caused by data imbalance. Especially, we generate a set of category prototypes uniformly distributed on a hypersphere. Category prototype-guided mechanism for image-text matching makes the features of different classes converge to these distinct and uniformly distributed category prototypes, which maintain a uniform distribution in the feature space, and improve class boundaries. Additionally, our proposed irrelevant text filtering and attribute enhancement module allows the model to ignore irrelevant noisy text and focus more on key attribute information, thereby enhancing the robustness of our framework. In the image recognition fine-tuning stage, to address the positive bias problem of the learnable classifier, we design the class feature prototype-guided classifier, which compensates for the performance of tail classes while maintaining the performance of head classes. Our method outperforms previous vision-language methods for long-tailed learning work by a large margin and achieves state-of-the-art performance. Xiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 0005, Hualiang Wang |
ACM Multimedia | 1 |
| 2023 | Class semantic enhancement network for semantic segmentation
Siming Fu, Hualiang Wang, Haoji Hu, Xiaoxuan He, Yongwen Long, Jianhong Bai, Yangtao Ou, Yuanjia Huang, Mengqiu Zhou |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Hierarchical Self-Supervised Learning for 3D Tooth Segmentation in Intra-Oral Mesh ScansabstractAccurately delineating individual teeth and the gingiva in the three-dimension (3D) intraoral scanned (IOS) mesh data plays a pivotal role in many digital dental applications, e.g., orthodontics. Recent research shows that deep learning based methods can achieve promising results for 3D tooth segmentation, however, most of them rely on high-quality labeled dataset which is usually of small scales as annotating IOS meshes requires intensive human efforts. In this paper, we propose a novel self-supervised learning framework, named STSNet, to boost the performance of 3D tooth segmentation leveraging on large-scale unlabeled IOS data. The framework follows two-stage training, i.e., pre-training and fine-tuning. In pre-training, three hierarchical-level, i.e., point-level, region-level, cross-level, contrastive losses are proposed for unsupervised representation learning on a set of predefined matched points from different augmented views. The pretrained segmentation backbone is further fine-tuned in a supervised manner with a small number of labeled IOS meshes. With the same amount of annotated samples, our method can achieve an mIoU of 89.88%, significantly outperforming the supervised counterparts. The performance gain becomes more remarkable when only a small amount of labeled samples are available. Furthermore, STSNet can achieve better performance with only 40% of the annotated samples as compared to the fully supervised baselines. To the best of our knowledge, we present the first attempt of unsupervised pre-training for 3D tooth segmentation, demonstrating its strong potential in reducing human efforts for annotation and verification. Zuozhu Liu, Xiaoxuan He, Hualiang Wang, Huimin Xiong, Yan Zhang 0004, Gaoang Wang, Jin Hao, Yang Feng 0011, Fudong Zhu, Haoji Hu |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Meta-prototype Decoupled Training for Long-Tailed Learning
Siming Fu, Huanpeng Chu, Xiaoxuan He, Hualiang Wang, Haoji Hu |
ACCV (6) | 3 |
| 2022 | Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-Tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He, Hangxiang Fang, Zuozhu Liu, Haoji Hu |
ECCV (24) | 3 |