Xiao Lu 0002

dblp:37/4424-2 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0003-0880-0160ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 PortraitFormer: Global Illumination Helps Portrait Shadow Removal
Xuze Jiao, Jiangjian Yu, Xiao Lu 0002, Chunxia Xiao
CGI (2)4
2025 DL2G: Degradation-guided Local-to-Global Restoration for Eyeglass Reflection Removal
abstract
Eyeglass reflection removal can restore the texture information in the reflection destructed eye area, which is meaningful for various tasks on the facial images. It is still challenging to correctly eliminate reflections, reasonably restore the lost contents, and guarantee that the final result has a consistent color and illumination with the input image. In this paper, we introduce a Degradation-guided Local-to-Global (DL2G) restoration framework to address this problem. We first propose a multiplicative reflection degradation model, which is used to alleviate reflection degradation to obtain a preliminary result. Then, in the local details restoration stage, we propose a local structure-aware diffusion model to learn the true distribution of texture details in the eye area. This helps in recovering lost contents in the regions of heavy degradation where the background is invisible. Finally, in the global consistency refinement stage, we utilize the input image as a reference image to generate the final result that is consistent with the input image in color and illumination. Extensive experiments demonstrate that our method can improve the effect of reflection removal and generate results with more reasonable semantics, exquisite details, and harmonious illumination.
Zhilv Yi, Xiao Lu 0002, Jingbo Hu, Chunxia Xiao
CVPR2
2025 Object-Preserving Counterfactual Diffusion Augmentation for Single-Domain Generalized Object Detection
abstract
Recent latent diffusion models (LDMs) have been explored to generate diverse domain-specific images based on source domain data, showing promising performance in domain generalization tasks. However, although the generated images present counterfactual augmentation, such as the background and style changes, the distortion of object details disrupts the causal factors, such as texture and shape. This leads to negative outcomes when directly applying LDM to domain generalization in object detection. To address the problems mentioned above, we propose Object-Preserving Counterfactual Diffusion augmentation method (OPCD) to explore the diffusion model to generate diverse domain-specific images without disrupting the object details. First, we construct a region-aware image generation framework, which leverages labeled source domain data to guide LDM in generating region-constrained images that preserve the semantic consistency of the original source images. Second, we propose object-preserving counterfactual augmentation, which retains the object region of the generated image and fuses diversified global information. This ensures that object details are not distorted and that the generated information is maintained. Third, to reduce the resource burden of generating a large number of images in LDM, we design a random insertion strategy. It mixes generated and source domain images, turning limited diversity samples into abundant training data. Experimental results on several benchmark datasets show that OPCD outperforms existing methods in single-domain generalized object detection. Codes can be found at https://github.com/qinhongda8/OPCD.
Hongda Qin, Xiao Lu 0002, Ningjiang Chen
ACM Multimedia2
2025 CoCNet: A Chain-of-Clues framework for zero-shot referring expression comprehension
Xuanyu Zhou, Zengcan Xue, Xiao Lu 0002, Tianxing Xiao, Lianhua Wu, Xuan Li 0019
Expert Syst. Appl.4
2025 Facial Highlight Removal With Cross-Context Attention and Texture Enhancement
abstract
Facial highlight removal aims to identify and remove the specular highlight components in the facial image, ensuring that the generated image has a consistent facial tone and high-fidelity texture detail. Existing methods struggle to remove the highlight and recover the details in disturbed areas simultaneously, often resulting in specular residues or distorted local details (i.e. texture, illumination, and color). To rectify these issues, this work proposes a novel two-stage facial highlight removal network (FHR-Net), which mainly consists of a Cross-Context Attention Module (CCAM) and a Texture Enhancement Module (TEM). In the first stage, according to the detected highlight mask, the CCAM explicitly integrates cross-context information to obtain coarse highlight removal results consistent with the surrounding facial context. Building upon the coarse result, the TEM in the second stage utilizes patch-wise attention to refine the texture details in the highlight areas, thereby producing a high-fidelity facial image. To improve coherence between the removed highlight areas and non-highlight areas, this work introduces a face feature loss that makes the processed highlight-disturbed areas align well with the surrounding facial architecture. Additionally, to address the lack of high-quality datasets in the research community and satisfy the training demands for data-driven facial highlight removal, this work builds a real-world Paired Facial Specular-Diffuse (PFSD) dataset through cross-polarization. Experimental results on PFSD and other datasets demonstrate that FHR-Net can effectively remove the facial highlight and recover original color and texture details.
Hongsheng Zheng, Wenju Xu, Xiao Lu 0002, Chunxia Xiao
IEEE Trans. Circuits Syst. Video Technol.4
2024 Low-Light Salient Object Detection by Learning to Highlight the Foreground Objects
abstract
Previous methods in salient object detection (SOD) mainly focused on favorable illumination circumstances while neglecting the performance in low-light condition, which significantly impedes the development of related down-stream tasks. In this work, considering that it is impractical to annotate the large-scale labels for this task, we present a framework (HDNet) to detect the salient objects in low-light images with the synthetic images. Our HDNet consists of a foreground highlight sub-network (HNet) and an appearance-aware detection sub-network (DNet), both of which can be learned jointly in an end-to-end manner. Specifically, to highlight the foreground objects, we design the HNet to estimate the parameters to adjust the dynamic range for each pixel adaptively, which can be trained via the weak supervision signals of the salient object labels. In addition, we design a simple detection network (DNet) with a contextual feature fusion module and a multi-scale feature refine module for detailed feature fusion and refinement. Furthermore, we contribute the first annotated dataset for salient object detection in low-light images (SOD-LL), including 6,000 labeled synthetic images (SOD-LLS) and 2,000 labeled real images (SOD-LLR). Experimental results on SOD-LL and other low-light videos in the wild demonstrate the effectiveness and generalization ability of our method. Our dataset and code are available at https://github.com/Ylinyuan/HDNet.
Xiao Lu 0002, Yulin Yuan, Lucai Wang, Xuanyu Zhou, Yimin Yang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Eyeglass Reflection Removal With Joint Learning of Reflection Elimination and Content Inpainting
abstract
Eyeglass reflection removal is of great importance to the portrait image processing. However, it remains a challenge to eliminate the reflections on the glass and restore the textual contents of eyes without introducing visual artifacts. Addressing this problem, in this paper, we propose an Eyeglass Reflection Removal Network (ER2Net) by learning reflection elimination and content inpainting jointly. The reflection elimination branch is effective in weak reflection regions, and the content inpainting branch is dedicated to content reasoning in strong reflection regions. We then propose a result fusion module (RFM), which adaptively fuses the elimination result and the inpainting result according to the reflection intensity of each pixel, to produce high-quality result. We also design a memory module for improving the content inpainting result, and propose an eye-symmetry loss to avoid visual artifacts. Additionally, we construct the first Real-world eyeglass Reflection (ReyeR) dataset for eyeglass reflection removal. Extensive quantitative and qualitative experiments demonstrate the superiority of the ER2Net over state-of-the-art methods for eyeglass reflection removal.
Wentao Zou, Xiao Lu 0002, Zhilv Yi, Ling Zhang 0017, Gang Fu 0003, Ping Li 0016, Chunxia Xiao
IEEE Trans. Circuits Syst. Video Technol.2
2023 Adaptive Refining-Aggregation-Separation Framework for Unsupervised Domain Adaptation Semantic Segmentation
abstract
Unsupervised domain adaptation has attracted widespread attention as a promising method to solve the labeling difficulties of semantic segmentation tasks. It trains a segmentation network for unlabeled real target images using easily available labeled virtual source images. To improve performance, clustering is used to obtain domain-invariant feature representations. However, most clustering-based methods indiscriminately cluster all features mapped by category from both domains, causing the centroid shift and affecting the generation of discriminative features. We propose a novel clustering-based method that uses an adaptive refining-aggregation-separation framework, which learns the discriminative features by designing different adaptive schemes for different domains and features. The clustering does not require any tunable thresholds. To estimate more accurate domain-invariant centroids, we design different ways to guide the adaptive refinement of different domain features. A critic is proposed to directly evaluate the confidence of target features to solve the absence of target labels. We introduce a domain-balanced aggregation loss and two adaptive separation losses for distance and similarity respectively, which can discriminate clustering features by combining the refinement strategy to improve segmentation performance. Experimental results on GTA$5\rightarrow $Cityscapes and SYNTHIA$\rightarrow $Cityscapes benchmarks show that our method outperforms existing state-of-the-art methods.
Yihong Cao, Hui Zhang 0023, Xiao Lu 0002, Yurong Chen 0003, Yaonan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training
abstract
It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Spatio-Temporal Interpolation Consistency Training (STICT) framework to rationally feed the unlabeled video frames together with the labeled images into an image shadow detection network training. Specifically, we propose the Spatial and Temporal ICT, in which we define two new interpolation schemes, i.e., the spatial interpolation and the temporal interpolation. We then derive the spatial and temporal interpolation consistency constraints accordingly for enhancing generalization in the pixel-wise classification task and for encouraging temporal consistent predictions, respectively. In addition, we design a Scale- Aware Network for multi-scale shadow knowledge learning in images, and propose a scale-consistency constraint to minimize the discrepancy among the predictions at different scales. Our proposed approach is extensively validated on the ViSha dataset and a self-annotated dataset. Experimental results show that, even without video labels, our approach is better than most state of the art supervised, semi-supervised or unsupervised image/video shadow detection methods and other methods in related tasks. Code and dataset are available at https://github.com/yihong-97/STICT.
Xiao Lu 0002, Yihong Cao, Chengjiang Long, Zipei Chen, Xuanyu Zhou, Yimin Yang 0001, Chunxia Xiao
CVPR1
2022 Semi-supervised Video Shadow Detection via Image-assisted Pseudo-label Generation
abstract
Although learning-based methods have shown their potential for image shadow detection, video shadow detection is still a challenging problem. It is due to the absence of large-scale, temporally consistent annotated video shadow detection dataset. To this end, we propose a semi-supervised video shadow detection method by seeking the assistance of the existing labeled image dataset to generate pseudo-labels as the additional supervision signals. Specifically, we first introduce a novel image-assisted video pseudo-label generator with a spatio-temporally aligned network (STANet). It generates high-quality and temporally consistent pseudo-labels. Then, with these pseudo-labels, we propose an uncertainty-guided semi-supervised learning strategy to reduce the impact of noise from them. Moreover, we also design a memory propagated long-term network (MPLNet), which produces video shadow detection results with long-term consistency in a light-weight way by using the memory mechanism. Extensive experiments on ViSha and our collected real-world video shadow detection dataset RVSD show that our approach not only achieves superior performance in the benchmark dataset but also generalizes well in more practical applications, which demonstrates the effectiveness of our method.
Zipei Chen, Xiao Lu 0002, Ling Zhang 0017, Chunxia Xiao
ACM Multimedia2
2021 Real-time stage-wise object tracking in traffic scenes: an online tracker selection method via deep reinforcement learning
Xiao Lu 0002, Yihong Cao, Xuanyu Zhou, Yimin Yang 0001
Neural Comput. Appl.1
2020 A Surface Defect Detection Framework for Glass Bottle Bottom Using Visual Attention Model and Wavelet Transform
abstract
Glass bottles must be thoroughly inspected before they are used for packaging. However, the vision inspection of bottle bottoms for defects remains a challenging task in quality control due to inaccurate localization, the difficulty in detecting defects in the texture region, and the intrinsically nonuniform brightness across the central panel. To overcome these problems, we propose a surface defect detection framework, which is composed of three main parts. First, a new localization method named entropy rate superpixel circle detection (ERSCD), which combines least-squares circle detection and entropy rate superpixel (ERS) with an improved randomized circle detection, is proposed to accurately obtain the region of interest (ROI) of the bottle bottom. Then, according to the structure-property, the ROI is divided into two measurement regions: central panel region and annular texture region. For the former, a defect detection method named frequency-tuned anisotropic diffusion super-pixel segmentation (FTADSP) that integrates frequency-tuned salient region detection (FT), anisotropic diffusion, and an improved superpixel segmentation is proposed to precisely detect the regions and boundaries of defects. For the latter, a defect detection strategy called wavelet transform multiscale filtering (WTMF) based on a wavelet transform and a multiscale filtering algorithm is proposed to reduce the influence of texture and to improve the robustness to localization error. The proposed framework is tested on four data sets obtained by our designed vision system. The experimental results demonstrate that our framework achieves the best performance compared with many traditional methods.
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Changyan Xiao, Xiao Lu 0002, Hui Zhang 0023
IEEE Trans. Ind. Informatics6
2019 SSG: superpixel segmentation and GrabCut-based salient object segmentation
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Changyan Xiao, Xiao Lu 0002
Vis. Comput.5
2018 Optimal Transmission Estimation via Fog Density Perception for Efficient Single Image Defogging
abstract
Single image defogging algorithms based on prior assumptions or constraints have captured much attention because of their simplicity and practicality. However, they still have some challenges to deal with foggy images captured under weather conditions where these assumptions or constraints may not be effective or efficient enough. In this paper, we aim to develop a novel image defogging algorithm by directly predicting the fog density of recovered images rather than adopting prior assumptions or constraints. In order to achieve this goal, two specific steps are introduced. First, we adopt three fog-relevant statistical features derived from foggy images, and further develop a simple fog density evaluator (SFDE) by creating a linear combination of these fog-relevant features. This proposed evaluator can efficiently perceive the fog density of a single image without reference to a corresponding fog-free image and has a low computational load compared with an existing method. Second, a physics-based mathematical relationship between the transmission and the fog density score of the recovered image is developed via SFDE, thus image defogging can be posed as a minimization problem on the fog density score of the recovered image. As a result, two optimal transmission models, called an optimal transmission model via SFDE (OTSFDE) and a simpler optimal transmission models via SFDE (SOTSFDE), are present to determine the key transmission map for efficient fog removal. Compared to OTSFDE, SOTSFDE has low computational complexity with slight performance degradation. Experimental results demonstrate that the proposed algorithms can effectively remove fog and are not confined by any assumptions or constraints, both quantitatively and qualitatively, compared with some existing algorithms.
Zhigang Ling, Jianwei Gong, Guoliang Fan 0001, Xiao Lu 0002
IEEE Trans. Multim.4
2017 Perception oriented transmission estimation for high quality image dehazing
Zhigang Ling, Guoliang Fan 0001, Jianwei Gong, Yaonan Wang 0001, Xiao Lu 0002
Neurocomputing5
2017 Traffic Sign Recognition via Multi-Modal Tree-Structure Embedded Multi-Task Learning
abstract
Traffic sign recognition is a rather challenging task for intelligent transportation systems since signs in different subsets, e.g., speed limit signs, prohibition signs, and mandatory signs, are very different from each other in color or shape, whereas they share some similarities to the ones in the same subset. Therefore, it is important to integrate different modalities of visual features, such as color and shape, and select discriminative features for better sign description; in addition, it benefits to explore the correlations between the classes of traffic signs to learn the classifiers jointly to improve the generalization performance. In this paper, we propose Multi- Modal tree-structure embedded Multi-Task Learning called M2- tMTL to select discriminative visual features both between and within modalities, as well as the correlated features shared by similar classification tasks. Our method simultaneously introduces two structured sparsity-induced norms into a least squares regression. One of the norms can be used not only to select modality of features but also to conduct within-modality feature selection. Moreover, the hierarchical correlations among the classification tasks are well represented by a tree structure, and therefore, the tree-structure sparsity-induced norm is used for learning the regression coefficients jointly to boost the performance of multi-class traffic sign recognition. Alternating direction method of multipliers (ADMM) is used to efficiently solve the proposed model with guaranteed convergence. Extensive experiments on public benchmark data sets demonstrate that the proposed algorithm leads to a quite interpretable model, and it has better or competitive performance with several state-of-the-art methods but with less computational and memory cost.
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhenjun Zhang, Zhigang Ling
IEEE Trans. Intell. Transp. Syst.1
2016 Learning deep transmission network for single image dehazing
abstract
State-of-the-art single image dehazing algorithms have some challenges to deal with images captured under complex weather conditions because their assumptions usually do not hold in those situations. In this paper, we develop a deep transmission network for robust single image dehazing. This deep transmission network simultaneously copes with three color channels and local patch information to automatically explore and exploit haze-relevant features in a learning framework. We further explore different network structures and parameter settings to achieve tradeoffs between performance and speed, which shows that color channels information is the most useful haze-relevant feature rather than local information. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world datasets.
Zhigang Ling, Guoliang Fan 0001, Yaonan Wang 0001, Xiao Lu 0002
ICIP4
2016 A Method for Metric Learning with Multiple-Kernel Embedding
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhigang Ling
Neural Process. Lett.1
2016 Erratum to: A Method for Metric Learning with Multiple-Kernel Embedding
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhigang Ling
Neural Process. Lett.1
2016 Adaptive transmission compensation via human visual system for efficient single image dehazing
Zhigang Ling, Shutao Li 0001, Yaonan Wang 0001, He Shen 0001, Xiao Lu 0002
Vis. Comput.5
2015 Adaptive extended piecewise histogram equalisation for dark image enhancement
abstract
Histogram equalisation has been widely used for image enhancement because of its simple implementation and satisfactory performance. However, traditional histogram equalisation uniformly redistributes an entire histogram or multiple piecewise histograms with the same equalisation strategy, which may produce unnatural artefacts, over‐enhancement or under‐enhancement in wide dynamic range dark image enhancement. This study proposes an adaptive extended piecewise histogram equalisation algorithm (AEPHE) for dark image enhancement. First, an original histogram is divided into a group of extended piecewise histograms. Then, an adaptive histogram equalisation, which balances intensity preservation and contrast boosting, is further developed and respectively applied to these extended piecewise histograms. The final histogram for image enhancement is produced by a weighted fusion of these equalised histograms. The experimental results indicate that AEPHE is superior to multiple state‐of‐the‐art algorithms.
Zhigang Ling, Yan Liang 0001, Yaonan Wang 0001, He Shen 0001, Xiao Lu 0002
IET Image Process.5
2015 A Method to Calibrate Vehicle-Mounted Cameras Under Urban Traffic Scenes
abstract
We address the problem of vehicle-mounted camera calibration under urban traffic scenes regarding the fact that the traditional calibration methods are practically restricted, since the internal parameters should be calibrated in the laboratory and it is impossible for recalibration that resulted from the parameters drifting or re-focusing when driving on roads. In this paper, we propose to utilize the manual lines lying in Manhattan directions in the scenes to compute their corresponding vanishing points for camera calibration, as the urban traffic scenes are usually man-made and the important lines and signs for driving are typically lying in the Manhattan directions. For “Manhattan world” scenes, where there are plenty of lines lying in Manhattan directions, the lines in the scene are detected automatically, and the clusters corresponding to Manhattan directions are obtained using RANSAC-like methods. For the more general “quasi-Manhattan world” scenes, where only the lines in two directions can be found naturally, while the lines in the other direction are usually detected trivially or even can be hardly detected, we propose a method to estimate the lines in the third direction to improve the vanishing point estimation accuracy. The method proposed is tested on both two types of scenes, and the accuracy and practicability of this method are demonstrated. Furthermore, calibration experiments on both one image and multiple images are conducted, which show that the results can be more accurate when more images are used.
Yaonan Wang 0001, Xiao Lu 0002, Zhigang Ling, Yimin Yang 0001, Zhenjun Zhang, Kena Wang
IEEE Trans. Intell. Transp. Syst.2
2013 A new method for camera stratified self-calibration under circular motion
Xiao Lu 0002, Yaonan Wang 0001, Hai-Xia Xu 0001, Xuanyu Zhou
Vis. Comput.1