Xuan Shao

dblp:208/4733 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-4096-9428ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation
Cairong Yan, Jinyi Han, Jin Ju, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)6
2025 KG-TS: Knowledge Graph-Driven Thompson Sampling for Online Recommendation
Cairong Yan, Hualu Xu, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)5
2025 A Virtual Camera Assisted Surround-View SLAM System for Robust Parking
Xuan Shao, Feiyang Lu, Cairong Yan
ICIC (14)1
2025 Revisit Data Association in Semantic SLAM Systems for Autonomous Parking
Xuan Shao, Leming Huang
MMM (3)1
2025 Towards a Robust Visual-Inertial-Surround-View SLAM System for Autonomous Indoor Parking
abstract
An autonomous parking system is a low-speed unmanned driving system applied in indoor parking environments. Real-time and high-precision vehicle localization and map construction of the environment are two core functional modules of the system. Camera and IMU (Inertial Measurement Unit) sensors provide complementary data to create a Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) system. However, existing SLAM systems face challenges in complex parking environments. Moreover, limitations inherent in VI-SLAM systems further compromise their perception accuracy, affecting both localization and optimization. This article addresses the shortcomings of current VI-SLAM systems by proposing the RVIS SLAM system. This robust semantic SLAM system integrates data from three sensors: a front-view camera, an IMU, and a surround-view system. To ensure localization accuracy, the system utilizes metric information from common semantic objects on the ground. These objects include parking-slots, speed bumps, and parking-slot numbers captured in surround-view images to build scale-aware constraints. These constraints refine the initial scale of the SLAM system, which is often compromised under low IMU excitation conditions. Additionally, in optimization, SLAM systems ideally assume that the front-end produces optimization graphs without data association outliers. However, in real-world indoor parking environments, sensor noise and vehicle vibrations make this assumption unrealistic. To mitigate the adverse effects of outliers in SLAM systems, this article proposes a robust surround-view semantic data association strategy. This strategy quantifies the uncertainty of surround-view semantic landmarks for the first time, ensuring reliable localization and mapping in challenging environments. Extensive experiments in typical indoor parking environments validate the effectiveness and efficiency of the proposed RVIS SLAM system.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 LiDUT-Depth: A Lightweight Self-supervised Depth Estimation Model Featuring Dynamic Upsampling and Triplet Loss Optimization
Hao Jiang 0014, Zhijun Fang 0001, Xuan Shao, Jenq-Neng Hwang
ICPR (16)3
2023 Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
Dandan Zhu 0001, Xuan Shao, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
Appl. Intell.2
2023 SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic Framework
abstract
For the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS .
Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.1
2023 A Novel Lightweight Audio-visual Saliency Model for Videos
abstract
Audio information has not been considered an important factor in visual attention models regardless of many psychological studies that have shown the importance of audio information in the human visual perception system. Since existing visual attention models only utilize visual information, their performance is limited but also requires high-computational complexity due to the limited information available. To overcome these problems, we propose a lightweight audio-visual saliency (LAVS) model for video sequences. To the best of our knowledge, this article is the first trial to utilize audio cues for an efficient deep-learning model for the video saliency estimation. First, spatial-temporal visual features are extracted by the lightweight receptive field block (RFB) with the bidirectional ConvLSTM units. Then, audio features are extracted by using an improved lightweight environment sound classification model. Subsequently, deep canonical correlation analysis (DCCA) aims at capturing the correspondence between audio and spatial-temporal visual features, thus obtaining a spatial-temporal auditory saliency. Lastly, the spatial-temporal visual and auditory saliency are fused to obtain the audio-visual saliency map. Extensive comparative experiments and ablation studies validate the performance of the LAVS model in terms of effectiveness and complexity.
Dandan Zhu 0001, Xuan Shao, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2022 MOFISSLAM: A Multi-Object Semantic SLAM System With Front-View, Inertial, and Surround-View Sensors for Indoor Parking
abstract
The semantic SLAM (Simultaneous Localization And Mapping) system is a crucial module for autonomous indoor parking. Visual cameras (monocular/binocular) and IMU (Inertial Measurement Unit) constitute the basic configuration to build such a system. The performance of existing SLAM systems typically deteriorates in the presence of dynamically movable objects or objects with little texture. By contrast, semantic objects on the ground embody the most salient and stable features in the indoor parking environment. Due to their inabilities to perceive such features on the ground, existing SLAM systems are prone to tracking inconsistency during navigation. In this paper, we present MOFISSLAM, a novel tightly-coupled${M}$ulti-${O}$bject semantic SLAM system integrating${F}$ront-view,${I}$nertial, and${S}$urround-view sensors for autonomous indoor parking. The proposed system moves beyond existing semantic SLAM systems by complementing the sensor configuration with a surround-view system capturing images from a top-down viewpoint. In MOFISSLAM, apart from low-level visual features and inertial motion data, typical semantic objects (parking-slots, parking-slot IDs and speed bumps) detected in surround-views are also incorporated in optimization, forming robust surround-view constraints. Specifically, each surround-view feature imposes a surround-view constraint that can be split into a contact term and a registration term. The former pre-defines the position of each individual surround-view feature subject to whether it has semantic contact with other surround-view features. Three contact modes, defined ascomplementary,adjacentandcoincident, are identified to guarantee a unified form of all contact terms. The latter further constrains by registering each surround-view observation and its position in the world coordinate system. In parallel, to objectively evaluate SLAM studies for autonomous indoor parking, a large-scale dataset with groundtruth trajectories is collected, which is the first of its kind. Its groundtruth trajectories, commonly unavailable, are obtained by tracking artificial features scattered in the indoor parking environment, whose 3D coordinates are measured with an ETS (Electronic Total Station). The collected dataset has been made publicly available athttps://shaoxuan92.github.io/MOFIS.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2021 ROECS: A Robust Semi-direct Pipeline Towards Online Extrinsics Correction of the Surround-view System
abstract
Generally, a surround-view system (SVS), which is an indispensable component of advanced driving assistant systems (ADAS), consists of four to six wide-angle fisheye cameras. As long as both intrinsics and extrinsics of all cameras have been calibrated, a top-down surround-view with the real scale can be synthesized at runtime from fisheye images captured by these cameras. However, when the vehicle is driving on the road, relative poses between cameras in the SVS may change from the initial calibrated states due to bumps or collisions. In case that extrinsics' representations are not adjusted accordingly, on the surround-view, obvious geometric misalignment will appear. Currently, the researches on correcting the extrinsics of the SVS in an online manner are quite sporadic, and a mature and robust pipeline is still lacking. As an attempt to fill this research gap to some extent, in this work, we present a novel extrinsics correction pipeline designed specially for the SVS, namely ROECS (Robust Online Extrinsics Correction of the Surround-view system). Specifically, a "refined bi-camera error" model is firstly designed. Then, by minimizing the overall "bi-camera error" within a sparse and semi-direct framework, the SVS's extrinsics can be iteratively optimized and become accurate eventually. Besides, an innovative three-step pixel selection strategy is also proposed. The superior robustness and the generalization capability of ROECS are validated by both quantitative and qualitative experimental results. To make the results reproducible, the collected data and the source code have been released at https://cslinzhang.github.io/ROECS/.
Tianjun Zhang, Brian Nlong Zhao, Ying Shen 0005, Xuan Shao, Lin Zhang 0014, Yicong Zhou
ACM Multimedia4
2020 A Tightly-coupled Semantic SLAM System with Visual, Inertial and Surround-view Sensors for Autonomous Indoor Parking
abstract
The semantic SLAM (simultaneous localization and mapping) system is an indispensable module for autonomous indoor parking. Monocular and binocular visual cameras constitute the basic configuration to build such a system. Features used in existing SLAM systems are often dynamically movable, blurred and repetitively textured. By contrast, semantic features on the ground are more stable and consistent in the indoor parking environment. Due to their inabilities to perceive salient features on the ground, existing SLAM systems are prone to tracking loss during navigation. Therefore, a surround-view camera system capturing images from a top-down viewpoint is necessarily called for. To this end, this paper proposes a novel tightly-coupled semantic SLAM system by integrating Visual, Inertial, and Surround-view sensors, VIS SLAM for short, for autonomous indoor parking. In VIS SLAM, apart from low-level visual features and IMU (inertial measurement unit) motion data, parking-slots in surround-view images are also detected and geometrically associated, forming semantic constraints. Specifically, each parking-slot can impose a surround-view constraint that can be split into an adjacency term and a registration term. The former pre-defines the position of each individual parking-slot subject to whether it has an adjacent neighbor. The latter further constrains by registering between each observed parking-slot and its position in the world coordinate system. To validate the effectiveness and efficiency of VIS SLAM, a large-scale dataset composed of synchronous multi-sensor data collected from typical indoor parking sites is established, which is the first of its kind. The collected dataset has been made publicly available at https://cslinzhang.github.io/VISSLAM/.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Hongyu Li 0001, Yicong Zhou
ACM Multimedia1
2019 Revisit Surround-view Camera System Calibration
abstract
The surround-view system is an essential component of an advanced driver assistance system especially when the vehicle runs in tight parking space or on a narrow road. To ensure successful maneuvering, a panoramic bird's-eye image with no blind spots is necessarily called for. Hence, a typical surround-view system consists of several cameras mounted around the vehicle capturing images from a top-down viewpoint, and an accurate extrinsic calibration for such system is prerequisite for providing a seamless surround-view image. To achieve this goal, this paper presents a novel extrinsic calibration pipeline which is both easy-to-use and reliable to operate on multiple cameras. Instead of taking the vehicle to a fixed position in a specific calibration site, a single chessboard is the only demand. We adopt a novel refinement procedure that jointly optimizes camera poses in a closed-loop manner. The effectiveness and efficiency of the proposed pipeline to calibrate a surround-view camera system has been corroborated by experiments.
Xuan Shao, Xiao Liu 0030, Lin Zhang 0014, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang
ICME1
2018 Salient object detection via a local and global method based on deep residual network
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Qiangqiang Zhou, Laurent Itti
J. Vis. Commun. Image Represent.4
2017 Saliency prediction based on new deep multi-layer convolution neural network
abstract
Recent advances in saliency detection have utilized deep learning to obtain high-level features to detect salient regions. These advances have demonstrated superior results over previous works that utilize hand-crafted low-level features for saliency detection. In this paper, we propose a new multilayer Convolutional Neural Network (CNN) model to learn high-level features for saliency detection. Compared to other methods, our method presents two merits. First, when performing features extraction, apart from the convolution and pooling step in our method, we add Restricted Boltzmann Machine (RBM) into the CNN framework to obtain more accurate features in intermediate step. Second, in order to deal with case of non-linear classification, we add the Deep Belief Network (DBN) classifier at the end of this model to classify the salient and non-salient regions. Quantitative and qualitative experiments on three benchmark datasets demonstrate that our method performs favorably against the state-of-the-art methods.
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti
ICIP3
2017 Scanpath Prediction Based on High-Level Features and Memory Bias
Xuan Shao, Ye Luo 0004, Dandan Zhu 0001, Laurent Itti
ICONIP (3)1
2017 Deep Salient Object Detection via Hierarchical Network Learning
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti
ICONIP (3)4