VLDB 2026 Research / reviewers in the wild / expert
Yuelei Xu
dblp:34/4741
· DBLP profile ↗
15ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0002-9868-7693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation
Xuanbin Wang, Yuelei Xu |
ICCV | 5 |
| 2025 | Unlocking Instance Semantic Awareness for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to improve model generalization across domains. While existing UDA methods leverage labels (source domain) and pseudo-labels (target domain) to learn domain-invariant features, they often treat each object as a singular entity, failing to capture the hierarchical understanding of instance semantics, and thus overfitting on certain samples. To address this, we propose a unified Inter- and Intra-instance Self-supervised learning framework for Domain Adaptive semantic segmentation, called I2SDA, which leverages diffusion models to unlock the awareness of instance semantics by capturing both inter-instance correlations and intra-instance structures. Building on the impressive compositional generalization abilities of diffusion models, this framework enables more effective domain-invariant feature learning. For inter-instance correlation, we leverage the similarity metric of diffusion feature across samples to provide additional supervision for the segmentation model, explicitly encouraging learning of the contextual dependencies between instances for more reliable and precise category predictions. For intra-instance structure, we generate diverse yet structurally consistent class-specific instances on target domain samples via diffusion models, guided by source domain labels, enabling the model to develop fine-grained understanding of intrinsic instance structures. Extensive experiments on SYNTHIA → Cityscape and GTA5 → Cityscape benchmarks demonstrate state-of-the-art performance, validating the effectiveness of our method. Zhaoxiang Zhang 0002, Chengming Xu 0001, Yuelei Xu |
ICME | 6 |
| 2025 | Better to Teach than to Give: Domain Generalized Semantic Segmentation via Agent Queries with Diffusion Model GuidanceabstractDomain Generalized Semantic Segmentation (DGSS) trains a model on a labeled source domain to generalize to unseen target domains with consistent contextual distribution and varying visual appearance.
Most existing methods rely on domain randomization or data generation but struggle to capture the underlying scene distribution, resulting in the loss of useful semantic information.
Inspired by the diffusion model's capability to generate diverse variations within a given scene context, we consider harnessing its rich prior knowledge of scene distribution to tackle the challenging DGSS task.
In this paper, we propose a novel agent \textbf{Query}-driven learning framework based on \textbf{Diff}usion model guidance for DGSS, named QueryDiff.
Our recipe comprises three key ingredients: (1) generating agent queries from segmentation features to aggregate semantic information about instances within the scene;
(2) learning the inherent semantic distribution of the scene through agent queries guided by diffusion features;
(3) refining segmentation features using optimized agent queries for robust mask predictions.
Extensive experiments across various settings demonstrate that our method significantly outperforms previous state-of-the-art methods.
Notably, it enhances the model's ability to generalize effectively to extreme domains, such as cubist art styles. Code is available at https://github.com/FanLiHub/QueryDiff. Zhaoxiang Zhang 0002, Yuelei Xu |
ICML | 5 |
| 2025 | No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion ModelsabstractEnhancing the cross-domain generalization of 3D semantic segmentation is a pivotal task in computer vision that has recently gained increasing attention. Most existing methods, whether using consistency regularization or cross-modal feature fusion, focus solely on individual objects while overlooking implicit semantic dependencies among them, resulting in the loss of useful semantic information. Inspired by the diffusion model's ability to flexibly compose diverse objects into high-quality images across varying domains, we seek to harness its capacity for capturing underlying contextual distributions and spatial arrangements among objects to address the challenging task of cross-domain 3D semantic segmentation. In this paper, we propose a novel cross-modal learning framework based on diffusion models to enhance the generalization of 3D semantic segmentation, named XDiff3D. XDiff3D comprises three key ingredients: (1) constructing object agent queries from diffusion features to aggregate instance semantic information; (2) decoupling fine-grained local details from object agent queries to prevent interference with 3D semantic representation; (3) leveraging object agent queries as an interface to enhance the modeling of object semantic dependencies in 3D representations. Extensive experiments validate the effectiveness of our method, achieving state-of-the-art performance across multiple benchmarks in different task settings. Code is available at \url{https://github.com/FanLiHub/XDiff3D}. Xuanbin Wang, Yuelei Xu |
NeurIPS | 5 |
| 2023 | Truncated Quantile Critics Algorithm for Cryptocurrency Portfolio OptimizationabstractThis paper investigates portfolio management algorithm for the cryptocurrency market by using the TQC (Truncated Quantile Critics) algorithm. The study is based on the daily prices of cryptocurrencies. TQC is a deep reinforcement learning algorithm with the Actor-Critic architecture. It alleviates the overestimation problem of traditional value learning algorithm. In this paper, the data of cryptocurrencies are first processed as input to the networks. The inputs to the networks include not only the closing prices of cryptocurrencies, but also the relative strength index, moving average line, and moving average convergence divergence. Various metrics measuring algorithm returns and algorithm stability are used as evaluation criteria in this paper. In this paper, common deep reinforcement learning algorithms are compared. The experimental results show that the TQC algorithm has a highest return of 33.9 % during the test period, which is 3 %, 3 % and 15.6 % higher than A2C, PPO and DDPG respectively. And, the TQC algorithm has the highest stability of return, which is an important evaluation metric for portfolio management algorithms. Despite the high volatility of the cryptocurrency market, the performance of the TQC algorithm has remained relatively stable. This illustrates the positive effects of the TOC algorithm. Leibing Xiao, Xinchao Wei, Yuelei Xu, Kun Gong |
SMC | 3 |
| 2023 | Water-surface infrared small object detection based on spatial feature weighting and class balancing methodabstractAbstract Infrared imaging is widely used due to its penetration capability to operate under many weather or lighting condition. However, due to the far distance of aerial view, feature blur, and the scarcity of aerial infrared data, the detection of small infrared targets on the water surface remains a challenging problem. In response to the problem of unclear features, we propose the spatial feature weighting method based on 2D Gaussian distribution. This method increases the weight of the target area by adaptively adjusting the feature activation. Secondly, for the problem of rare aerial perspective infrared data, we propose the cross‐spectral data migration method. By introducing the domain difference loss function to optimize the pseudo‐label selection process, the range of target domain distribution is expanded, and the adaptability of the detector is improved. Finally, in response to the problem of underfitting caused by category imbalance in transfer learning, we propose the class balancing method that effectively reduces the false detection. Extensive experiments were conducted on both benchmark datasets and the self‐built dataset to evaluate the effectiveness and robustness of our method. The proposed method was evaluated with different models and various scenarios, and the results demonstrated the effectiveness. Tian Hui 0001, Yuelei Xu, Huafeng Li 0003, Rasol Jarhinbek |
IET Image Process. | 2 |
| 2022 | Unsupervised SAR and Optical Image Matching Using Siamese Domain AdaptationabstractDue to its highly complementary information about remote sensing, synthetic aperture radar (SAR) and optical imagery matching have drawn much attention in recent years. Compared with traditional methods, deep learning-based SAR-optical image matching models largely rely on supervision with ground truths, where the matching accuracy suffers because of unseen image domains. To mitigate loads in burdensome labeling tasks, transferring deep learning models trained with annotated source domains to nonannotated target domains has attracted great concern. Due to the domain gap, the difference between the source and target domains is likely to deteriorate the matching accuracy on target data if the training process is directly conducted without proper domain adaptation (DA). In this research, a Siamese DA (SDA) approach with a combined loss function is developed in the context of multimodality image matching. Then, a novel rotation/scale-invariant transformation module with regression modules is designed to extract rotation/scale-equivariant features. Finally, the causal inference-based self-learning method and the multiresolution histogram matching approach are employed to enhance the unsupervised matching performance. Experimental results on the RadarSat/Planet dataset and the Sentinel-1/2 dataset demonstrate that the developed model can achieve competitive matching performance with a low overlap ratio between domains and little data labeling. By alleviating the domain discrepancy, the developed model drastically reduces the average L2 score of the unsupervised matching from 9.576 to 0.658, while the less-than-one-pixel matching error rate is enhanced from 66.3% to 90.6%. Zhaoxiang Zhang 0002, Yuelei Xu, Qing Zhou 0001, Linhua Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Cross-View Images Matching and Registration Technology Based on Deep Learning
Qing Zhou 0001, Ronggang Zhu, Yuelei Xu, Zhaoxiang Zhang 0002 |
ICIG (1) | 3 |
| 2021 | Remote Sensing Image Jitter Restoration Based on Deep Generative Adversarial NetworkabstractHigh stability of the observation satellite platform is increasingly demanded in recent years. However, attitude jitter of observation satellites is a problem that degenerates the development of imaging quality and resolution. In order to reduce the geo-positioning errors and improve the geometric accuracy of remote sensing images, satellite jitter have been studied in recent years. In this work, a generative adversarial network (GAN) architecture is proposed to automatically learn and correct the deformed scene features from a single remote sensing image. In the proposed GAN, a convolutional neural network (CNN) is designed to discriminate the inputs and another CNN is used to generate so-called fake inputs. In order to explore the usefulness and effectiveness of GAN for jitter detection, the proposed GAN are trained on part of PatternNet dataset and tested on three popular remote sensing datasets. Several experiments show that the proposed models provide competitive results compared to other methods. the proposed GAN reveals the huge potential of GAN-based methods for the analysis of attitude jitter from remote sensing images. Zhaoxiang Zhang 0002, Qing Zhou 0001, Yuelei Xu, Linhua Ma, Akira Iwasaki |
IGARSS | 3 |
| 2021 | Detail texture detection based on Yolov4-tiny combined with attention mechanism and bicubic interpolationabstractAbstract Aero‐engine blades crack detection is one of the important tasks in daily ground maintenance, crack is a kind of texture feature, due to the random distribution, irregular shape and vague characteristics, which is still a challenging task to realize automatic detection in working environment. A detection model based on the Yolov4‐tiny is proposed that is universal and focuses more on the characteristics of cracks, and it is implemented in embedded device. First, in order to distinguish the cracks and noises, an improved attention module is introduced into the backbone of Yolov4‐tiny to enhance the model's capability to focus on crack areas; second, in order to improve the effect of multi‐scale feature fusion, the bicubic interpolation is implemented in upsampling module; finally, in order to solve the redundant detection results of bounding‐boxes in crack areas, the optimized non‐maximum suppression method is proposed to make the detection results better corresponding to the groundTruth. The robustness of proposed detection model was demonstrated by evaluating varying lighting and noise images. The average precision on integrated datasets is 81.6%, which outperforms the original Yolov4‐tiny by an increase of 12.3%. Tian Hui 0001, Yuelei Xu, Rasol Jarhinbek |
IET Image Process. | 2 |
| 2021 | A hierarchical sampling based triplet network for fine-grained image classification
Guiqing He, Qiyao Wang, Zongwen Bai, Yuelei Xu |
Pattern Recognit. | 5 |
| 2020 | A concept ontology triplet network for learning discriminative representations of fine-grained classes
Guiqing He, Haixi Zhang, Yuelei Xu, Jianping Fan 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Remote Sensing Airport Detection Based on End-to-End Deep Transferable Convolutional Neural NetworksabstractRapid intelligent detection of airports from remote sensing images is required to accomplish autonomous intelligent landing of unmanned aerial vehicles (UAVs) and other tasks. To address the insufficiency of traditional models in detecting airports under complicated backgrounds from remote sensing images, we propose an end-to-end remote sensing airport hierarchical expression and detection model based on deep transferable convolutional neural networks. Based on transfer learning, we solve the fundamental problem of overfitting due to the inadequate number of labeled remote sensing images by transferring the network model from natural image source domain to remote sensing image target domain. In addition, we introduce a cascade region proposal network with soft-decision nonmaximal suppression to improve the network structure and the performance of our method under complex backgrounds. Moreover, we use skip-layer feature fusion and hard example mining methods to improve the object expression ability and the training efficiency. Finally, the experimental results demonstrate that the method established in this letter can quickly and effectively detect different types of airports over complex backgrounds and obtain better detection performance than the other detection methods. Shuai Li 0009, Yuelei Xu, Mingming Zhu, Shiping Ma |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Edge Detection Based on Primary Visual Pathway
Yuelei Xu, Xulei Zhang, Shiping Ma, Shuai Li 0009, Peng Xin, Mingning Zhu, Hongqiang Ma |
ICIG (2) | 2 |
| 2005 | Analytical Solution for Dynamic of Neuronal Populations
Licheng Jiao, Shiping Ma, Yuelei Xu |
ICANN (1) | 4 |