VLDB 2026 Research / reviewers in the wild / expert
Xuelu Li
dblp:52/8508
· DBLP profile ↗
9ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-0217-9350ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PIXELS: Progressive Image Xemplar-based Editing with Latent SurgeryabstractRecent advancements in language-guided diffusion models for image editing are often bottle-necked by cumbersome prompt engineering to precisely articulate desired changes. An intuitive alternative calls on guidance from in-the-wild image exemplars to help users bring their imagined edits to life. Contemporary exemplar-based editing methods shy away from leveraging the rich latent space learnt by pre-existing large text-to-image (TTI) models and fall back on training with curated objective functions to achieve the task. Though somewhat effective, this demands significant computational resources and lacks compatibility with diverse base models and arbitrary exemplar count. On further investigation, we also find that these techniques restrict user control to only applying uniform global changes over the entire edited region. In this paper, we introduce a novel framework for progressive exemplar-driven editing with off-the-shelf diffusion models, dubbed PIXELS, to enable customization by providing granular control over edits, allowing adjustments at the pixel or region level. Our method operates solely during inference to facilitate imitative editing, enabling users to draw inspiration from a dynamic number of reference images, or multimodal prompts, and progressively incorporate all the desired changes without retraining or fine-tuning existing TTI models. This capability of fine-grained control opens up a range of new possibilities, including selective modification of individual objects and specifying gradual spatial changes. We demonstrate that PIXELS delivers high-quality edits efficiently, leading to a notable improvement in quantitative metrics as well as human evaluation. By making high-quality image editing more accessible, PIXELS has the potential to enable professional-grade edits to a wider audience with the ease of using any open-source image generation model. Shristi Das Biswas, Matthew Shreve, Xuelu Li, Prateek Singhal |
AAAI | 3 |
| 2023 | Maturity-Aware Active Learning for Semantic Segmentation with Hierarchically-Adaptive Sample Assessment
Amirsaeed Yazdani, Xuelu Li, Vishal Monga |
BMVC | 2 |
| 2023 | PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain AdaptationabstractTraditional Unsupervised Domain Adaptation (UDA) leverages the labeled source domain to tackle the learning tasks on the unlabeled target domain. It can be more challenging when a large domain gap exists between the source and the target domain. A more practical setting is to utilize a large-scale pre-trained model to fill the domain gap. For example, CLIP shows promising zero-shot generalizability to bridge the gap. However, after applying traditional fine-tuning to specifically adjust CLIP on a target domain, CLIP suffers from catastrophic forgetting issues where the new domain knowledge can quickly override CLIP’s pre-trained knowledge and decreases the accuracy by half. We propose Catastrophic Forgetting Measurement (CFM) to adjust the learning rate to avoid excessive training (thus mitigating the catastrophic forgetting issue). We then utilize CLIP’s zero-shot prediction to formulate a Pseudo-labeling setting with Adaptive Debiasing in CLIP (PADCLIP) by adjusting causal inference with our momentum and CFM. Our PADCLIP allows end-to-end training on source and target domains without extra overhead. We achieved the best results on four public datasets, with a significant improvement (+18.5% accuracy) on DomainNet. Zhengfeng Lai, Noranart Vesdapunt, Cong Phuoc Huynh, Xuelu Li, Kah Kuen Fu, Chen-Nee Chuah |
ICCV | 6 |
| 2022 | Structural Prior Models for 3-D Deep Vessel SegmentationabstractWe address the problem of 3-D blood vessel segmentation with a deep learning method that incorporates domain information via priors and regularizers on vessel structure and morphology. Inspired by the observation that 3-D vessel structures project onto 2-D image slices with distinctive edges that can aid 3-D vessel segmentation, we propose a novel multi-task learning architecture comprising a shared encoder and two decoders that respectively predict vessel segmentation maps and edge profiles. 3-D features from the two branches are concatenated to facilitate edge-guidance when learning segmentation maps. We introduce new regularization terms that encourage local homogeneity of 3-D blood vessel volumes brought about by biomarkers, as well as sparsity of edge pixels. Experiments on benchmark datasets demonstrate superior performance of our method over the state-of-the-art, especially when training data is limited. Xuelu Li, Raja Bala, Vishal Monga |
ICASSP | 1 |
| 2022 | Robust Deep 3D Blood Vessel Segmentation Using Structural PriorsabstractDeep learning has enabled significant improvements in the accuracy of 3D blood vessel segmentation. Open challenges remain in scenarios where labeled 3D segmentation maps for training are severely limited, as is often the case in practice, and in ensuring robustness to noise. Inspired by the observation that 3D vessel structures project onto 2D image slices with informative and unique edge profiles, we propose a novel deep 3D vessel segmentation network guided by edge profiles. Our network architecture comprises a shared encoder and two decoders that learn segmentation maps and edge profiles jointly. 3D context is mined in both the segmentation and edge prediction branches by employing bidirectional convolutional long-short term memory (BCLSTM) modules. 3D features from the two branches are concatenated to facilitate learning of the segmentation map. As a key contribution, we introduce new regularization terms that: a) capture the local homogeneity of 3D blood vessel volumes in the presence of biomarkers; and b) ensure performance robustness to domain-specific noise by suppressing false positive responses. Experiments on benchmark datasets with ground truth labels reveal that the proposed approach outperforms state-of-the-art techniques on standard measures such as DICE overlap and mean Intersection-over-Union. The performance gains of our method are even more pronounced when training is limited. Furthermore, the computational cost of our network inference is among the lowest compared with state-of-the-art. Xuelu Li, Raja Bala, Vishal Monga |
IEEE Trans. Image Process. | 1 |
| 2020 | Multiview Automatic Target Recognition for Infrared Imagery Using Collaborative Sparse PriorsabstractThe low resolution of infrared (IR) images makes feature extraction for classification of a challenging work. Learning-based methods, therefore, are preferred to be used on such raw imagery. In this article, in order to avoid difficulties in feature extraction, a novel multitask extension of the widely used sparse-representation-classification (SRC) method is proposed in both single and multiview set-ups. That is, the test sample could be a single IR image or images from different views. In both single-view and multiview scenarios, we try to employ collaborative spike and slab priors. This is because the traditional sparsity-inducing measures such as the l0-row pseudonorm makes it hard to capture the sparse structure of the coefficient matrix when expanded in terms of a training dictionary, and the priors are proved to be able to capture fairly general sparse structures. Furthermore, a joint prior and sparse coefficient estimation method (JPCEM) is proposed for the first time in this article in order to alleviate the need to handpick prior parameters required before classification. Multiple experiments are conducted on a synthetic Comanche Forward Looking IR (FLIR) Automatic Target Recognition (ATR) database collected by Army Research Lab and a challenging mid-wave IR (MWIR) image ATR database made available by the U.S. Army Night Vision and Electronic Sensors Directorate. The final results substantiate the merits of the proposed JPCEM through comparisons with other state-of-the-art methods, including both the ones based on SRC and the ones constructed using deep learning frameworks. Xuelu Li, Vishal Monga, Abhijit Mahalanobis |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Group Based Deep Shared Feature Learning for Fine-grained Image Classification
Xuelu Li, Vishal Monga |
BMVC | 1 |
| 2018 | Collaborative Sparse Priors for Infrared Image Multi-View ATRabstractFeature extraction from infrared (IR) images remains a challenging task. Learning based methods that can work on raw imagery/patches have therefore assumed significance. We propose a novel multi-task extension of the widely used sparse-representation-classification (SRC) method in both single and multi-view set-ups. That is, the test sample could be a single IR image or images from different views. When expanded in terms of a training dictionary, the coefficient matrix in a multi-view scenario admits a sparse structure that is not easily captured by traditional sparsity-inducing measures such as the l0-row pseudo norm. To that end, we employ collaborative spike and slab priors on the coefficient matrix, which can capture fairly general sparse structures. Our work involves joint parameter and sparse coefficient estimation (JPCEM) which alleviates the need to handpick prior parameters before classification. The experimental merits of JPCEM are substantiated through comparisons with other state-of-art methods on a challenging mid-wave IR image (MWIR) ATR database made available by the US Army Night Vision and Electronic Sensors Directorate. Xuelu Li, Vishal Monga |
IGARSS | 1 |
| 2018 | 3-D Imaging Based on Combination of the ISAR Technique and a MIMO Radar SystemabstractIn this paper, a novel 3-D imaging method achieved by combining the inverse synthetic aperture radar (ISAR) technique and a multiple-input-multiple-output (MIMO) radar system is presented. The high-resolution image of a target can be obtained within a limited imaging time interval, and the computational load can be reduced simultaneously because the time samples obtained by the ISAR technique can be used to play the role of space samples needed in the MIMO radar system. The adopted process of time selection method can help to realize the exact combination of the signals received by different antenna elements without producing nonequivalence between the time samples and the space samples which need to be substituted. Besides, the 2-D smoothed (2-D SL0) method, which is especially designed for 2-D signals in the compressed sensing (CS) technique, is used to realize the high-resolution imaging when there are gaps in the global observation angle without increasing the computation load or introducing the false estimated data. Furthermore, the algorithms for range alignment and velocity estimation under the 3-D imaging radar configuration are described explicitly. Finally, the simulation results are provided to prove the effectiveness of the algorithms proposed in this paper. Yong Wang 0017, Xuelu Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |