Xuesong Li 0001

dblp:76/5858-1 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-9295-7333ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 CoAttn-PEFT: Complementary Attention in Parameter-Efficient Fine-Tuning for Improved Adaptation and Generalization
Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li 0001, Hongdong Li
ICPR (2)4
2026 MFGS: Mask-free Gaussian separation for 3D object reconstruction
abstract
Accurate 3D reconstruction from multi-view images is a fundamental problem in computer vision. A common acquisition strategy involves placing an object on a rotating turntable while moving the camera to capture it from various viewpoints. In such scenarios, object moves relative to the background, many existing reconstruction methods rely on object masks to separate the foreground from the background. The quality of these masks significantly affects the final reconstruction, yet obtaining high-quality and consistent masks is a challenging and laborious process, especially when controlled environments like green screens are unavailable. To address this limitation, we introduce Mask-free Gaussian Separation (MFGS), a novel method that performs simultaneous object reconstruction and segmentation without requiring any input masks. Our approach builds on Gaussian Splatting and automatically disentangles the scene by extending each Gaussian primitive with a learnable parameter that represents its probability of belonging to the dynamic foreground object. This separation is optimized in a self-supervised manner, optimized by the object and camera transformation constraints. We evaluated MFGS on new synthetic and real-world datasets designed to reflect this challenging capture scenario. Experimental results demonstrate that our mask-free approach significantly outperforms existing methods. Notably, MFGS surpasses the performance of the state-of-the-art method(2DGS) that relies on high-quality segmentation masks, achieving a 27% improvement in novel view synthesis and a 7% improvement in geometry reconstruction.
Jinguang Tong, Xuesong Li 0001, Sundaram Muthu, Fahira A. Maken, Lars Petersson, Hongdong Li
Pattern Recognit.2
2025 GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
abstract
3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the advantage of high speed and detailed real-time rendering, but extracting surfaces from the Gaussians can be noisy due to the lack of geometric constraints. To bridge the gap between these approaches, we propose a novel reconstruction method called GS-2DGS for reflective objects based on 2D Gaussian Splatting (2DGS). Our approach combines the rapid rendering capabilities of Gaussian Splatting with additional geometric information from foundation models. Experimental results on synthetic and real datasets demonstrate that our method significantly outperforms Gaussian-based techniques in terms of reconstruction and relighting and achieves performance comparable to SDF-based methods while being an order of magnitude faster. Code is available at https://github.com/hirotong/GS2DGS
Jinguang Tong, Xuesong Li 0001, Fahira A. Maken, Sundaram Muthu, Lars Petersson, Hongdong Li
CVPR2
2025 SSTD: Stripe-Like Space Target Detection Using Single-Point Weak Supervision
abstract
Stripe-like space target detection (SSTD) plays a key role in enhancing space situational awareness, but it faces three challenges: the lack of public datasets, interference from space noise, and the variability of stripe-like targets, making manual labeling both inaccurate and labor-intensive. In response, we introduce ‘AstroStripeSet’, a pioneering dataset for SSTD, aiming to bridge the gap in academic resources. Furthermore, we propose a novel teacher-student label evolution framework with single-point weak supervision, providing a new solution to the challenges of manual labeling. It starts with generating initial pseudo-labels using the zero-shot capabilities of the Segment Anything Model (SAM) in a single-point setting. Then, the fine-tuned StripeSAM serves as the teacher and the newly developed StripeNet as the student, improving segmentation performance through label evolution, which iteratively refines these labels. We also introduce ‘GeoDice’, a new loss function tailored for the linear characteristics of stripe-like targets. Extensive experiments show that our method matches fully supervised approaches, exhibits strong zero-shot generalization for diverse real-world images, and sets a new state-of-the-art benchmark. Our dataset and code are available at https://github.com/BenZae/SSTD.
Ali Zia, Xuesong Li 0001, Bingbing Dan, Yuebo Ma, Enhai Liu, Rujin Zhao
ICME3
2025 Enhancing Features in Long-tailed Data Using Large Vision Model
abstract
Language-based foundation models, such as large language models (LLMs) or large vision-language models (LVLMs), have been widely studied in long-tailed recognition. However, the need for linguistic data is not applicable to all practical tasks. In this study, we aim to explore using large vision models (LVMs) or visual foundation models (VFMs) to enhance long-tailed data features without any language information. Specifically, we extract features from the LVM and fuse them with features in the baseline network's map and latent space to obtain the augmented features. Moreover, we design several prototype-based losses in the latent space to further exploit the potential of the augmented features. In the experimental section, we validate our approach on two benchmark datasets: ImageNet-LT and iNaturalist2018.
Pengxiao Han, Changkun Ye, Jinguang Tong, Xuesong Li 0001
IJCNN7
2025 DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction
abstract
Dynamic scene reconstruction from monocular video is essential for real-world applications. We introduce DGNS, a hybrid framework integrating Deformable Gaussian Splatting and Dynamic Neural Surfaces, effectively addressing dynamic novel-view synthesis and 3D geometry reconstruction simultaneously. During training, depth maps generated by the deformable Gaussian splatting module guide the ray sampling for faster processing and provide depth supervision within the dynamic neural surface module to improve geometry reconstruction. Conversely, the dynamic neural surface directs the distribution of Gaussian primitives around the surface, enhancing rendering quality. In addition, we propose a depth-filtering approach to further refine depth supervision. Extensive experiments conducted on public datasets demonstrate that DGNS achieves state-of-the-art performance in 3D reconstruction, along with competitive results in novel-view synthesis.
Xuesong Li 0001, Jinguang Tong, Vivien Rolland, Lars Petersson
ACM Multimedia1
2025 BioNet and NeFF: Crop Biomass Prediction from Point Clouds to Drone Imagery
abstract
Crop biomass offers crucial insights into plant health and yield, making it essential for crop science, farming systems, and agricultural research. However, current measurement methods, which are labor-intensive, destructive, and imprecise, hinder large-scale quantification of this trait. To address this limitation, we present a biomass prediction network (BioNet), designed for adaptation across different data modalities, including point clouds and drone imagery. Our BioNet, utilizing a sparse 3D convolutional neural network (CNN) and a transformer-based prediction module, processes point clouds and other 3D data representations to predict biomass. To further extend BioNet for drone imagery, we integrate a neural feature field (NeFF) module, enabling 3D structure reconstruction and the transformation of 2D semantic features from vision foundation models into the corresponding 3D surfaces. For the point cloud modality, BioNet demonstrates superior performance on two public datasets, with an approximate 6.1% relative improvement (RI) over the state-of-the-art. In the RGB image modality, the combination of BioNet and NeFF achieves a 7.9% RI. Additionally, the NeFF-based approach utilizes inexpensive, portable drone-mounted cameras, providing a scalable solution for large field applications.
Xuesong Li 0001, Zeeshan Hayder, Ali Zia, Connor Cassidy, Shiming Liu, Warwick Stiller, Eric A. Stone, Warren Conaty, Lars Petersson, Vivien Rolland
WACV1
2024 Backpropagation-free Network for 3D Test-time Adaptation
abstract
Real-world systems often encounter new data over time, which leads to experiencing target domain shifts. Existing Test- Time Adaptation (TTA) methods tend to apply computationally heavy and memory-intensive backpropagation-based approaches to handle this. Here, we propose a novel method that uses a backpropagation-free approach for TTA for the specific case of 3D data. Our model uses a two-stream architecture to maintain knowledge about the source domain as well as complementary target-domain-specific information. The backpropagation-free property of our model helps address the well-known forgetting prob-lem and mitigates the error accumulation issue. The pro-posed method also eliminates the need for the usually noisy process of pseudo-labeling and reliance on costly self-supervised training. Moreover, our method leverages sub-space learning, effectively reducing the distribution vari-ance between the two domains. Furthermore, the source-domain-specific and the target-domain-specific streams are aligned using a novel entropy-based adaptive fusion strat-egy. Extensive experiments on popular benchmarks demon-strate the effectiveness of our method. The code will be available at https://github.com/abie-e/BFTT3D.
Yanshuo Wang, Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, David Ahmedt-Aristizabal, Xuesong Li 0001, Lars Petersson, Mehrtash Harandi
CVPR8
2024 Spatial Transcriptomics Analysis of Zero-Shot Gene Expression Prediction
Yan Yang 0011, Xuesong Li 0001, Shafin Rahman, Eric A. Stone
MICCAI (4)3
2023 Cross-Modal and Cross-Domain Knowledge Transfer for Label-Free 3D Segmentation
Huitong Yang, Daijie Wu, Jacky W. Keung, Xuesong Li 0001, Xinge Zhu, Yuexin Ma
PRCV (3)5
2023 Efficient and Accurate Object Detection With Simultaneous Classification and Tracking Under Limited Computing Power
abstract
Interacting with the environment, such as object detection and tracking, is a crucial ability of mobile robots. Besides high accuracy, efficiency in terms of processing effort and energy consumption are also desirable. To satisfy both requirements, we propose a detection framework based on simultaneous classification and tracking in the point stream. In this framework, a tracker performs data association in sequences of the point cloud, guiding the detector to avoid redundant processing (i.e. classifying already-known objects). For objects whose classification is not sufficiently certain, a fusion model is designed to fuse selected key observations that provide different perspectives across the tracking span. Therefore, performance (accuracy and efficiency of detection) can be enhanced. This method is particularly suitable for detecting and tracking moving objects, a process which would require expensive computations if solved using conventional procedures. Experiments were conducted on the benchmark dataset, and the results showed that the proposed method outperforms original tracking-by-detection approaches in both efficiency and accuracy.
Xuesong Li 0001, José E. Guivant
IEEE Trans. Intell. Transp. Syst.1
2018 Non-linear Estimation with Generalised Compressed Kalman Filter
abstract
The optimal estimation of dynamic random fields is a relevant problem in diverse areas of robotics application. The associated estimation process in these problems implicitly requires dealing with high dimensional multi-variate Probability Density Functions (PDFs) with unaffordable processing cost. The Generalised Compressed Kalman Filter (GCKF) with subsystem switching and proper information exchange architecture is capable of solving such problems with comparable performance to the optimal full Gaussian estimators but at a remarkably lower cost. In this paper, an explicit algorithm is proposed for replacing the Kalman Filter core with a suitable Gaussian Filter core to solve non-linear estimation problems. The computational advantages of GCKF are highlighted, where the computational complexities of different Gaussian Filters are compared against their compressed counterpart. The performance of the algorithm has been verified through its application in solving linear Stochastic Partial Differential Equations (SPDEs) with unknown parameters.
Karan Narula, José E. Guivant, Xuesong Li 0001
FUSION3
2017 Research on joint segment optimisation and stereo matching
abstract
Image segments are often used as a constraint in stereo matching. However, both over‐segmentation and under‐segmentation can lead to disparity degradation in some regions. To obtain an accurate disparity map, a modified semi‐global matching (SGM) algorithm is proposed which is based on adaptive window models. Introducing the object notion and build a new global energy function to optimise segments and estimate a disparity map jointly. The effective and efficient block coordinate descent approach is used to optimise the global energy function by merging small segments. The authors’ demonstrate the performance of the proposed algorithm on the KITTI and Middlebury benchmarks. The results show that the authors’ algorithm outperforms many state‐of‐the‐art methods and confirm the effectiveness of approach.
Xuesong Li 0001, Shengheng Liu
IET Commun.3
2016 Efficient methods using slanted support windows for slanted surfaces
abstract
The frontal‐parallel assumption is made by many matching algorithms, but this assumption fails for slanted surfaces. This study proposes a matching algorithm intended to improve the matching results for slanted surfaces. First, a mathematical model is constructed to prove that slanted surfaces in the environment have corresponding slanted disparity surfaces in the disparity space image, and the model is to help find the proper plane parameters of slanted support windows, then improved cost aggregation and post‐processing methods are proposed. The algorithm is tested using the Middlebury and Karlsruhe Institute of Technology and Toyota Technical Institute at Chicago (KITTI) benchmarks. The results demonstrate that the algorithm exhibits good performance and is efficient for slanted surfaces.
Xuesong Li 0001, Heng Fu
IET Comput. Vis.1