Xiaomao Li

dblp:72/3469 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Beyond Vision: Vision-Language Distillation and Edge-Aware Mix Diffusion in Semi-Supervised Semantic Segmentation
abstract
In semi-supervised semantic segmentation (SSSS), segmentation performance is heavily constrained by the quality of pseudo labels. However, prevalent pseudo-label optimization approaches rely on the model’s internal self-correction. When the model fails to recognize or adequately represent certain classes, this self-enhancement mechanism amplifies initial mistakes, ultimately leading to poor semantic or spatial consistency. To address this limitation, we propose ViLaDiff to enhance pseudo-label quality. Specifically, ViLaDiff first employs a prompt-guided image captioning task to generate descriptive text for each input image, providing high-level semantic context. To our knowledge, this is the first attempt to introduce vision-language modeling into SSSS. We design a vision-language fusion module to enhance feature semantics and discriminative capability. It integrates cross-modal interactions with dual-path knowledge to ensure semantic consistency. Additionally, while language provides high-level semantic guidance, it is inherently limited in expressing fine-grained spatial structures. Therefore, we propose an edge-aware mixed-noise diffusion process. It simulates feature-level uncertainty through Gaussian perturbations and introduces class-flipping noise into the masks to model misclassification errors. To enhance boundary refinement, we apply a higher flipping probability along mask edges, enabling edge-aware modeling during denoising. Extensive experiments on public benchmarks validate that our method significantly improves pseudo-label quality and segmentation performance.
Yuehua Liu, Xiaomao Li, Shaorong Xie
AAAI4
2025 PolarNeXt: Rethink Instance Segmentation with Polar Representation
abstract
One of the roadblocks for instance segmentation today is heavy computational overhead and model parameters. Previous methods based on Polar Representation made the initial mark to address this challenge by formulating instance segmentation as polygon detection, but failed to align with mainstream methods in performance. In this paper, we highlight that Representation Errors, arising from the limited capacity of polygons to capture boundary details, have long been overlooked, which results in severe performance degradation. Observing that optimal starting point selection effectively alleviates this issue, we propose an Adaptive Polygonal Sample Decision strategy to dynamically capture the positional variation of representation errors across samples. Additionally, we design a Union-aligned Rasterization Module to incorporate these errors into polygonal assessment, further advancing the proposed strategy. With these components, our framework PolarNeXt achieves a remarkable performance boost of over 4.8% AP compared to other polar-based methods. PolarNeXt is markedly more lightweight and efficient than state-of-the-art instance segmentation methods, while achieving comparable segmentation accuracy. We expect this work will open up a new direction for instance segmentation in high-resolution images and resource-limited scenarios. Codes can be found at https://github.com/Sun15194/PolarNeXt.
Xinghong Zhou, Yiqiang Wu, Jiaxuan Lu, Xiaomao Li
CVPR7
2025 Decoupled and Interactive Regression Modeling for High-performance One-stage 3D Object Detection
abstract
Inadequate bounding box modeling in regression tasks constrains the performance of one-stage 3D object detection. We pinpoint two key factors behind this limitation: (1) Restricted center-offset prediction severely impairs bounding box localization, as many peak response positions deviate significantly from object centers. (2) Low-quality samples ignored in regression tasks notably affect bounding box prediction, leading to unreliable IoU-based quality adjustments. To tackle these problems, we propose Decoupled and Interactive Regression Modeling (DIRM) for one-stage detection. Decoupled Attribute Regression (DAR) is implemented to facilitate long regression range modeling for the center attribute through an adaptive multi-sample assignment strategy. Additionally, to enhance the reliability of IoU predictions for low-quality results, Interactive Quality Prediction (IQP) integrates the classification task, proficient in modeling negative samples, with quality prediction for joint optimization. Extensive experiments on Waymo and ONCE datasets demonstrate that DIRM achieves state-of-the-art performance, balancing accuracy and computational cost for real-time autonomous driving.
Weiping Xiao, Yiqiang Wu, Chenghai Mao, Xiaomao Li
ICME6
2025 Enhanced soft domain adaptation for object detection in the dark
Xiaomao Li
J. Vis. Commun. Image Represent.4
2025 Contrastive-Domain Mean Teacher for Domain Adaptive Object Detection
abstract
Most semi-supervised domain object detection (SDAOD) methods are based on the mean-teacher framework. This framework primarily utilizes object-level features provided by pseudo-labels. However, the pseudo-labels generated by the Teacher model often contain notable noise, which limits the detector’s performance. Unlike pseudo-labels, domain labels are more precise and can offer accurate domain-level features. Motivated by this, we incorporate domain-level features into contrastive learning by designing different label assignment strategies and thus propose Contrastive-Domain Mean Teacher (CDMT) for SDAOD. Specifically, domain-level features include both inter-domain and intra-domain features. For inter-domain features, our strategy regards samples with the same domain label as positive pairs, enabling contrastive learning to extract global feature representations. While, intra-domain features from the same image are treated as positive pairs, which helps contrastive learning to extract fine-grained feature representations. Thorough experiments demonstrate that CDMT achieves state-of-the-art performance on Foggy Cityscapes and Clipart combined with recent Mean Teacher framework methods. Notably, for more challenging foggiest images (’0.02’ split) based on the Probabilistic Teacher (PT) baseline, CDMT outperforms the previously best CMT by 4.1% on mAP, which shows its priority on cross-domain detection tasks.
Yiqiang Wu, Xiaomao Li
IEEE Trans. Circuits Syst. Video Technol.4
2025 SampleDet3D: Sample Enhanced 3D Object Detection
abstract
Center-based 3D object detection has underperformed recently compared to advanced techniques. We experimentally find that the root lies in two weaknesses of the basic sample mechanism: (1) Unreasonable assignment that close-range and high-frequency objects dominate the network optimization since samples are equally assigned to each object. (2) Ambiguous encoding that samples exhibit suboptimal object discrimination ability, as the encoding process is restricted to a limited receptive field. To realize a reasonable assignment, Dynamic Multi-Quality Assignment (DMQA) is proposed, which dynamically assigns and supervises samples through fine-grained control. Concretely, initial samples are defined based on prior attributes (category and distance) per object, and dynamically adjusted upon the learning effect (classification and localization confidence). Besides, multi-scale auxiliary losses are introduced, ensuring precise sample learning. As for ambiguous encoding, Interactive Enhancement (IE) is introduced to improve sample representation through cross-task and cross-sample interaction. Cross-task interaction first aggregates neighborhood context from another task map. Parallel attention further performs cross-sample interaction on both local and global levels. Based on DMQA and IE, we propose a novel 3D detector named Sample Enhanced 3D Object Detection (SampleDet3D). Comprehensive experiments demonstrate that SampleDet3D effectively enhances center-based detection and achieves state-of-the-art performance on both Waymo and ONCE datasets.
Yiqiang Wu, Chang Liu 0082, Chenghai Mao, Xiaomao Li
IEEE Trans. Circuits Syst. Video Technol.7
2024 Misalignment-resistant domain adaptive learning for one-stage object detection
Chang Liu 0082, Xiaomao Li
Knowl. Based Syst.4
2024 Enhancing low-light images via dehazing principles: Essence and method
Fei Li 0034, Caiju Wang, Xiaomao Li
Pattern Recognit. Lett.3
2024 Spatial-Aware Learning in Feature Embedding and Classification for One-Stage 3-D Object Detection
abstract
One-stage 3D object detection, known for its simplicity and high-speed inference, is attracting increasing attention in autonomous driving scenarios. However, current one-stage detectors tend to perform sub-optimally compared to two-stage competitors. Our experimental findings suggest that one-stage detectors underperform due to the underutilization of spatial information in feature embedding and classification. Concretely, the spatial context is severely lost during feature propagation, inducing distorted spatial awareness. On the other hand, category recognition relies on the full utilization of spatial information, which is neglected by current detectors. This inadequate spatial awareness of the classification branch can exacerbate misclassification. To address these issues, we propose Spatial-aware Learning in Feature Embedding and Classification for One-stage 3D Object Detection (SLDet). Specifically, to restore the distorted spatial awareness, Category-wise Spatial Augmentation (CSA) is proposed to adaptively bring the network with pre-encoding multi-scale spatial contexts. As for misclassification, Spatial Guiding Classification (SGC) is introduced to guide the classification using explicit scale information. It employs the natural scale divergences among categories to rectify misclassification. Comprehensive experiments demonstrate that SLDet efficiently utilizes spatial information and achieves newly state-of-the-art performance on both the Waymo Open Dataset and the ONCE Dataset. Furthermore, additional experiments demonstrate the excellent generalization capacity of SLDet.
Yiqiang Wu, Weiping Xiao, Jiantao Gao, Chang Liu 0082, Yan Peng 0001, Xiaomao Li
IEEE Trans. Geosci. Remote. Sens.7
2024 CCDet: Confidence-Consistent Learning for Dense Object Detection
abstract
Modern detectors commonly employ classification scores to reflect the localization quality of detection results. However, there exists an inconsistency between them, misguiding the selection of high-quality predictions and providing unreliable results for downstream applications. In this paper, we find that the root of this confidence inconsistency lies in the inaccurate IoU estimation and the spatial misalignment of the learned features between the classification and localization tasks. Therefore, a Confidence-Consistent Detector (CCDet) which includes the Distribution-based IoU Prediction (DIP) and Consistency-aware label assignment (CLA), is proposed. DIP provides more stable and accurate IoU estimation by learning the probability distribution over the IoU range and employing the expectation as the predicted IoU. CLA adopts both the prediction performance and consistency degree of samples as assignment metrics to select positives, which guides the classification and localization tasks to promote similar feature distribution. Comprehensive experiments demonstrate that CCDet can effectively mitigate the confidence inconsistency between classification and localization, and achieve stable improvement across different baselines. On the test-dev set of MS COCO, CCDet acquires a single-model single-scale AP of 50.1%, surpassing most of the existing object detectors.
Chang Liu 0082, Xiaomao Li, Weiping Xiao, Shaorong Xie
IEEE Trans. Image Process.2
2023 Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection
abstract
With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since classification scores cannot properly represent the localization quality. (2) Assignment ambiguity that samples are matched with improper labels in pseudo-label assignment, as the strategy is misguided by missed objects and inaccurate pseudo boxes. To tackle these problems, we propose a Ambiguity-Resistant Semi-supervised Learning (ARSL) for one-stage detectors. Specifically, to alleviate the selection ambiguity, Joint-Confidence Estimation (JCE) is proposed to jointly quantifies the classification and localization quality of pseudo labels. As for the assignment ambiguity, Task-Separation Assignment (TSA) is introduced to assign labels based on pixel-level predictions rather than unreliable pseudo boxes. It employs a ‘divide-and-conquer’ strategy and separately exploits positives for the classification and localization task, which is more robust to the assignment ambiguity. Comprehensive experiments demonstrate that ARSL effectively mitigates the ambiguities and achieves state-of-the-art SSOD performance on MS COCO and PASCAL VOC. Codes can be found at https://github.com/PaddlePaddle/PaddleDetection.
Chang Liu 0082, Weiming Zhang 0006, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Xiaomao Li, Errui Ding, Jingdong Wang 0001
CVPR7
2023 Mitigate the classification ambiguity via localization-classification sequence in object detection
Chang Liu 0082, Shaorong Xie, Xiaomao Li, Jiantao Gao, Weiping Xiao, Baojie Fan, Yan Peng 0001
Pattern Recognit.3
2023 Balanced Sample Assignment and Objective for Single-Model Multi-Class 3D Object Detection
abstract
Accurately detecting multi-class objects in a single pass is critical but challenging for real-world autonomous driving scenarios. Several single-class anchor-based methods have recently achieved the state-of-the-art performance in the car category, but when extending to multi-class detection tasks, their performance on small objects (i.e., pedestrians and cyclists) is limited. We find that the core problem that causes this phenomenon lies in the unbalanced sample quality and the classification objective. To address this problem, we proposed a single-model multi-class 3D object detector with balanced sample assignment and objective, named BSAODet. Specifically, the quality-balanced sample assignment (QBSA) is introduced to dynamically collect stable high-quality samples for each class according to the predicted sample performance and geometric constraints. In conjunction with the QBSA, the class-balanced classification objective (CBCO) performs instance-wise label normalization and weighting on positive samples, preventing the model from biasing toward objects with more samples. Extensive experiments on the popular KITTI dataset, the latest large-scale ONCE dataset, and the challenging Waymo Open Dataset show that our method steadily improves the performance of current state-of-the-art detectors by 2–7 mAP in pedestrians and cyclists while maintaining competitiveness in cars. Moreover, our best model achieves 66.31 mAP on three classes, outperforming all published LiDAR-only detectors on the KITTI benchmark.
Weiping Xiao, Yan Peng 0001, Chang Liu 0082, Jiantao Gao, Yiqiang Wu, Xiaomao Li
IEEE Trans. Circuits Syst. Video Technol.6
2022 Policy Learning for Optimal Individualized Dose Intervals
abstract
We study the problem of learning individualized dose intervals using observational data. There are very few previous works for policy learning with continuous treatment, and all of them focused on recommending an optimal dose rather than an optimal dose interval. In this paper, we propose a new method to estimate such an optimal dose interval, named probability dose interval (PDI). The potential outcomes for doses in the PDI are guaranteed better than a pre-specified threshold with a given probability (e.g., $50%$). The associated nonconvex optimization problem can be efficiently solved by the Difference-of-Convex functions (DC) algorithm. We prove that our estimated policy is consistent, and its risk converges to that of the best-in-class policy at a root-n rate. Numerical simulations show the advantage of the proposed method over outcome modeling based benchmarks. We further demonstrate the performance of our method in determining individualized Hemoglobin A1c (HbA1c) control intervals for elderly patients with diabetes.
Xiaomao Li, Menggang Yu
AISTATS2
2022 RE-Det3D: RoI-enhanced 3D object detector
Yiqiang Wu, Weiping Xiao, Chang Liu 0082, Jiantao Gao, Guozhu Tan, Xiaomao Li
Image Vis. Comput.7
2022 3D-VDNet: Exploiting the vertical distribution characteristics of point clouds for 3D object detection and augmentation
Weiping Xiao, Xiaomao Li, Chang Liu 0082, Jiantao Gao, Jun Luo 0006, Yan Peng 0001
Image Vis. Comput.2
2020 Diverse receptive field network with context aggregation for fast object detection
Shaorong Xie, Chang Liu 0082, Jiantao Gao, Xiaomao Li, Jun Luo 0006, Baojie Fan, Jiahong Chen, Huayan Pu, Yan Peng 0001
J. Vis. Commun. Image Represent.4
2020 Data driven hybrid edge computing-based hierarchical task guidance for efficient maritime escorting with multiple unmanned surface vehicles
Jiajia Xie, Jun Luo 0006, Yan Peng 0001, Shaorong Xie, Huayan Pu, Xiaomao Li, Zhou Su 0001, Yuan Liu 0025
Peer-to-Peer Netw. Appl.6
2018 Structured and weighted multi-task low rank tracker
Baojie Fan, Xiaomao Li, Yang Cong, Yandong Tang
Pattern Recognit.2
2017 Layered Multitask Tracker via Spatial-Temporal Laplacian Graph
abstract
Most multitask trackers define the trace of each candidate as one task, and assume all tasks are equally related. Multitask learning is only evaluated on the current frame. In fact, these assumptions are limited, and ignore the multitask relationship in consecutive frames. In this letter, we propose a discriminative layered multitask tracker via spatial-temporal Laplacian graphs, which defines the layered tasks from a novel view, and naturally incorporates the global and local target information into reverse multitask tracking process. The spatial-temporal Laplacian graphs not only exploit the sequential consistent information of the target, but also make full use of the geometric structure corresponding to the tasks among the adjacent frames. Besides, l0norm constraint and labeling information are used to improve the tracking robustness. Encouraging experimental results on challenging sequences justify that the proposed method performs well both in accuracy and robustness against some related trackers.
Baojie Fan, Xiaomao Li, Yang Cong
IEEE Signal Process. Lett.2
2016 Structured low rank tracker with smoothed regularization
abstract
In this paper, we propose a structured low rank learning algorithm with smoothed regularization for robust object tracking, under particle filter framework. Specifically, the relationships among the particles are exploited with structured low rank regularization term, and simultaneously handle the outlier using a group sparsity regularization. The label information from training data is incorporated into the tracking objective function as the classification error term and idea coding regularization term respectively. By the smoothed regularization, the developed structured low rank learning based tracker can be efficiently solved by iterative reweighed least squares algorithm(IRLS), and avoids svd operation. Moreover, the collaborate normalized metric is developed to find the best candidate. Compared with some state-of-the-art tracking methods on 50 challenging sequences, the proposed algorithms perform well in terms of accuracy, robustness.
Baojie Fan, Yang Cong, Xiaomao Li, Yandong Tang
VCIP3
2008 Lunar terrain reconstruction using PDEs
abstract
Based on the geometry features of lunar terrain, this paper treats lunar terrain reconstruction as a surface reconstruction problem. We define an energy functional model consisting of local energy term and smooth energy term for lunar terrain reconstruction. The solution to minimize the functional (by partial differential equations) is defined as the optimal surface. In the smooth energy term, we design a vector field of depth discontinuousness likelihood (VFDDL) to control the direction and degree of smoothing. Experiments indicate that accurate VFDDL can lead to an exact reconstructed surface. Thus, VFDDL transfers 3D terrain reconstruction into a 2D image processing problem. An innovative method is proposed to estimate VFDDL, using image local and statistical features. Experiments verify our method and show a good performance in terrain reconstruction.
Ji Liu 0002, Yang Cong, Xiaomao Li, Yuechao Wang, Yandong Tang, Chuan Zhou 0010
ICIP3