EDBT 2026 Demo / reviewers in the wild / expert
Bo Yan 0007
dblp:63/6796-7
· DBLP profile ↗
20ranked-venue papers
0as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised Monocular Visual Odometry Based on Multi-View Spatio-Temporal Feature Fusion
Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 3 |
| 2026 | SwinFVO: Self-Supervised Visual Odometry With Enhanced Global Spatiotemporal PerceptionabstractPose estimation using visual sensors has become a fundamental component in robotic navigation and autonomous driving systems. Learning-based monocular visual odometry (VO) has attracted substantial attention due to its resilience to camera parameter variations and dynamic environments. Given that camera movement manifests as pixel-level motion across the entire image in optical flow data, capturing both global contextual information and local feature details is crucial for accurate pose estimation. To address this challenge, we propose SwinFVO, a novel self-supervised visual odometry framework that incorporates enhanced motion perception to achieve global spatial dependency modeling with temporal continuity. Leveraging quadrant-based motion characteristics, we perform cross-regional feature interaction through a refined Swin Transformer architecture. Two robust spatiotemporal feature extractors are designed to extend the single-frame-based Swin Transformer to a temporally-aware framework for sequential understanding. Through the exploration of long-range spatial correlations and preservation of temporal consistency, SwinFVO delivers accurate and consistent pose estimation. Extensive experiments across multiple datasets demonstrate the superior performance and generalization capability of SwinFVO in both pose and depth estimation tasks. It achieves competitive results against classical algorithms and outperforms related state-of-the-art (SOTA) methods by up to 20.6% and 72.4% on average translational and rotational evaluations, respectively. Rujun Song, Ruoqi Li, Zhuoling Xiao, Bo Yan 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | ZSCIL: Zero-Shot Class Incremental Learning Method for Signal RecognitionabstractA significant challenge in signal recognition tasks is identifying classes not present in the dataset. Zero-shot learning-based signal recognition addresses the challenge by identifying previously unseen classes in a mixed signal space without supervision. However, most existing methodologies are limited to one-time recognition processes. We propose a zero-shot class incremental learning (ZSCIL) method to achieve continuous unseen classes identification. Our model employs an encoder-decoder architecture and incorporates a triplet loss function to train the classifier, thereby enhancing the model’s ability to recognize mixed signals through a metric learning paradigm. Additionally, we utilize class incremental learning, where the identified unseen signals are stored in a fixed-size buffer with a maximum diversity data replay mechanism. These signals are then used for incremental training. The framework’s effectiveness and generality of our method are demonstrated through a series of experiments on two datasets. For instance, we achieved a significant 14.4% accuracy improvement for seen classes and that of the unseen classes by 4.2% on the DeepSig 2016.04C dataset. To the best of our knowledge, ZSCIL is the first method to implement sustainable identification for unseen classes in the mixed signal space. Rujun Song, Sidi Liang, Di He 0002, Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 6 |
| 2025 | LoSeVO: Local Sequence Constraints for Deep Visual OdometryabstractMany current visual odometry (VO) methods that utilize deep learning primarily concentrate on the constraints of motion relationships between adjacent frames, neglecting the modeling of temporal correlations within sequence data. Consequently, this paper introduces LoSeVO to effectively capture and model the temporal correlation features present in images. We design a Joint Feature Extraction component that not only performs joint feature extraction on adjacent frames but also extracts features from cross-frame images, which are called feature-guided maps. Then, we apply a Local Consistency Constraint component to the joint features between adjacent frames. It can adaptively constrain adjacent frames at different temporal positions within a sequence using different feature-guided maps. Extensive experiments based on the KITTI and Malaga datasets have shown that, compared to our previous DeepAVO model, LoSeVO can improve pose estimation performance by up to 23% and 9% in translation and rotation estimation, respectively. Rujun Song, Di He 0002, Tingyong Yang, Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 7 |
| 2025 | IPLC+: SAM-Guided Iterative Pseudo Label Correction for Source-Free Domain Adaptation in Medical Image SegmentationabstractDomain Adaptation (DA) is important for a segmentation model to deal with domain shift in a new target domain. Due to the privacy concern of medical data and the expensive annotation process, Source-Free Domain Adaptation (SFDA) is appealing without access to source data and labels of target domain images for the adaptation. However, existing SFDA methods have limited performance due to insufficient supervision and unreliable pseudo labels. In this paper, we propose an enhanced Iterative Pseudo Label Correction (IPLC+) SFDA framework guided by Segment Anything Model (SAM) for medical image segmentation. Specifically, with a pre-trained source model and SAM, we propose a Reliable SAM Pseudo-label Generator (RSPG) to obtain high-quality and reliable pseudo labels in the target domain based on multiple prompts randomly sampled from the model's prediction. To provide more efficient constraints during adaptation, we introduce self-training pseudo labels weighted by the uncertainty, and propose regularization using mean curvature minimization based on shape-prior knowledge for smoother segmentation. We also propose an Iterative Correction Learning (ICL) strategy to iteratively refine pseudo labels using SAM with updated prompts and combine supervisions to optimize the model sufficiently. Experiments on two public multi-site datasets for prostate and heart segmentation show that our method effectively outperformed ten state-of-the-art SFDA methods, improved the quality of pseudo labels, and even achieved better results than fully supervised learning in the target domain in some cases. Guoning Zhang 0002, Xiaoran Qi, Jianghao Wu 0001, Bo Yan 0007, Guotai Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | GraSS: Graph Neural Networks for Loop Closure Detection with Semantic and Spatial AssistanceabstractLoop Closure Detection (LCD) is an essential part of minimizing drift due to the accumulation of previously pose errors in Simultaneous Localization and Mapping (SLAM). The existing loop detection methods are limited by the changes of external conditions such as illumination, viewpoint and appearance. Previous work has mainly focused on the feature descriptor matching methods, which usually only consider the keypoints themselves. Here, we propose a fusion method GraSS, which uses the Graph Neural Network (GNN) based on visual features, and introduces semantics and depth, so as to enhance the spatial characteristics of the keypoints and the information correlation between them in the graph. Furthermore, a learnable parameter is added when two keypoints share the same semantic labels, their matching scores are increased, mitigating to some extent the issue of mismatch caused by significant differences in external conditions between two keypoints that should ideally be paired. Our findings show that GraSS has better performance than other state-of-the-art LCD methods when facing obvious illumination, appearance changes and slight viewpoint changes. Shihang Lu, Zhuolin Peng, Zhuoling Xiao, Bo Yan 0007, Shuisheng Lin, Di He 0002 |
ISCAS | 4 |
| 2024 | CLFusion: 3D Semantic Segmentation Based on Camera and Lidar FusionabstractIn the field of autonomous driving, semantic segmentation is crucial for scene understanding. Currently, there are two main methods: camera-based and Lidar-based approaches. To address the issues of Lidar segmentation lacking texture features and image segmentation lacking distance information, this paper proposes a fusion of camera and Lidar to achieve 3D semantic segmentation. The method utilizes a dual-stream encoder-decoder network to process camera images and Lidar point cloud and incorporates a specially designed attention mechanism module for feature fusion. To avoid expensive manual annotation of 3D point clouds, the study also introduces a cross-dataset and cross-modal self-supervised training approach. Experimental results show a 2.4% improvement compared to the Lidar-only mode baseline results on the SemanticKITTI dataset and a 6% improvement on the nuScenes dataset. Tianyue Wang, Rujun Song, Zhuoling Xiao, Bo Yan 0007, Haojie Qin, Di He 0002 |
ISCAS | 4 |
| 2024 | IPLC: Iterative Pseudo Label Correction Guided by SAM for Source-Free Domain Adaptation in Medical Image Segmentation
Guoning Zhang 0002, Xiaoran Qi, Bo Yan 0007, Guotai Wang |
MICCAI (11) | 3 |
| 2024 | GraphAVO: Self-Supervised Visual Odometry Based on Graph-Assisted Geometric ConsistencyabstractLearning-based monocular visual odometry (VO) has recently attracted considerable attention for its robustness to camera parameters and environmental variations. Despite traditional pose graph optimization enhancing pose accuracy, its integration with deep learning may lead to error accumulation due to insufficient motion information exchange. Our method, GraphAVO, concentrates simultaneously on the adjacent and interval co-visibility correspondence to establish a feature and pose graph optimization for pose consistency. We design a graph-assisted Windowed Feature Graph Refinement (WFGR) component to operationalize pose graph optimization for deep feature refinement. The geometric consistency is further constrained by a Cycle Consistency Loss. Additionally, the Cascade Dilated Convolution Fusion (CDCF) component is incorporated to handle different degrees of pixel movement, facilitating the joint detection of slight and distinct motion cues for subsequent feature enhancement. Extensive experiments on the KITTI, Malaga, RobotCar, and self-collected outdoor datasets have demonstrated the promising performance and generalization ability of GraphAVO. It achieves competitive results against classical algorithms and outperforms related state-of-the-art methods by up to 24.4% and 40.1% on average translational and rotational evaluation, respectively. Rujun Song, Zhuoling Xiao, Bo Yan 0007 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Adaptive Semantic Fusion Framework for Unsupervised Monocular Depth EstimationabstractUnsupervised monocular depth estimation plays an important role in autonomous driving, and has been received considerable research attention in recent years. Nevertheless, numerous existing methods relying on photometric consistency are excessively susceptible to variations in illumination and suffer in the regions with strong reflection. To overcome this limitation, we propose a novel unsupervised depth estimation framework named ColorDepth, which forces the model to explore object semantic to infer depth. Specifically, we extract pixel-level semantic prior clues of objects using the semantic segmentation network. These priors and the original image are then adaptively fused into color data by a learnable parameter for depth estimation. The incorporation of semantics endows our model with the ability to perceive scene structure information. The fused data effectively alleviates the depth ambiguity within the same semantic block, leading to improved consistency and robustness in challenging scenarios. Extensive experiments on the KITTI and Make3D datasets show that our method surpasses the previous state-of-the-art methods even those supervised by additional constraints, and brings significant performance improvement particularly in the regions of high reflection. Ruoqi Li, Kaiyang Du, Zhuoling Xiao, Bo Yan 0007, Zhengxi Yuan |
ICASSP | 5 |
| 2023 | Self-supervised Visual Odometry Based on Geometric ConsistencyabstractLearning-based monocular visual odometry (VO) has lately drawn significant attention for its robustness to camera parameters and environmental variations. Unlike most self-supervised learning-based methods, our approach simultaneously focuses on the adjacent and interval co-visibility correspondence to improve the pose estimation. To handle different pixel displacements, we apply the Multi-scale Feature Fusion component for the full exploration of latent motion features. Besides, the Interval Feature Guided Refinement component is incorporated to adaptively exploit the continuity of camera motions and steer the network for retaining pose consistency in the time domain. Extensive experiments on the KITTI and Malaga datasets have demonstrated the promising performance of our approaches. The proposed method produces competitive results against classic algorithms and outperform state-of-the-art methods by up to 23.9 % and 15.4 % on average translational and rotational evaluation. Rujun Song, Kaisheng Liao, Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 5 |
| 2023 | GlobalDepth: Global-Aware Attention Model for Unsupervised Monocular Depth EstimationabstractMonocular depth estimation is a significant task in computer vision, which can be widely used in Simultaneous Localization and Mapping (SLAM) and navigation. However, the current unsupervised approaches have limitations in global information perception, especially at distant objects and the boundaries of the objects. To overcome this weakness, we propose a global-aware attention model called GlobalDepth for depth estimation, which includes two essential modules: Global Feature Extraction (GFE) and Selective Feature Fusion (SFF). GFE considers the correlation among multiple channels and refines the encoder feature by extending the receptive field of the network. Furthermore, we restructure the skip connection by employing SFF between the low-level and the high-level features in element wise, rather than simply concatenation or addition at the feature level. Our model excavates the key information and enhances the ability of global perception to predict details of the scene. Extensive experimental results demonstrate that our method reduces the absolute relative error by 10.32% compared with other state-of-the-art models on KITTI datasets. Ruoqi Li, Zhuoling Xiao, Bo Yan 0007 |
ISCAS | 4 |
| 2023 | Adaptive map matching based on dynamic word embeddings for indoor positioning
Xinyue Lan, Lijia Zhang, Zhuoling Xiao, Bo Yan 0007 |
Neurocomputing | 4 |
| 2023 | ContextAVO: Local context guided and refining poses for deep visual odometry
Rujun Song, Zhuoling Xiao, Bo Yan 0007 |
Neurocomputing | 4 |
| 2022 | DeepAVO: Efficient pose refining with feature distilling for deep Visual Odometry
Rujun Song, Bo Yan 0007, Zhuoling Xiao |
Neurocomputing | 5 |
| 2021 | Radar Update (RUPT): A Pedestrian Navigation System with Enhanced Trajectory PerformanceabstractTo alleviate the dependence on sensor quality and to reduce the accumulated error in traditional inertial navigation systems, this paper proposes RUPT: a millimeter-wave radar aided pedestrian dead reckoning system with dual foot-mounted inertial measurement units (IMU). RUPT in this paper is a comprehensive data processing procedure which pre-processes both inertial data and millimeter-wave data and fuses them in a complementary way. Extensive experiments have demonstrated that the accuracy of RUPT has been improved by up to 65% over the conventional dual-foot mounted pedestrian tracking system. Yuquan Dai, Kemeng Li, Jin Chai, Zhuoling Xiao, Bo Yan 0007, Wenyu Peng |
ISCAS | 5 |
| 2021 | Adaptive Real-Time Loop Closure Detection Based on Image Feature ConcatenationabstractSimultaneous Localization and Mapping (SLAM) is used to solve the problem of autonomous localization and navigation of mobile robots in unknown environments. Loop closure detection is a key part of SLAM, which largely determines accuracy and stability of SLAM. In recent years, some experiments have proved that the loop closure detection system based on neural network is superior to the traditional loop closure detection in both accuracy and real-time performance. In this paper, we propose an adaptive real-time loop closure detection (AR-Loop) method based on monocular vision. A pre-trained convolutional neural network (CNN) is used to extract image features. Then features of different layers are concatenated as image descriptors. In addition, the adaptive candidate matching range algorithm and image-to- sequence calibration algorithm are proposed to improve the performance of the algorithm. Extensive experiments have been conducted on several open datasets to validate the performance of AR-Loop. It has been demonstrated that the recall rate is increased by over 18% compared with other state-of-the-art algorithms when the precision is 100%. Xiaorui Lin, Zhuoling Xiao, Bo Yan 0007, Shuisheng Lin |
ISCAS | 6 |
| 2019 | LightVO: Lightweight Inertial-Assisted Monocular Visual Odometry with Dense Neural NetworksabstractMonocular visual odometry (VO) is one of the most practical ways in vehicle autonomous positioning, through which a vehicle can automatically locate itself in a completely unknown environment. Although some existing VO algorithms have proved the superiority, they usually need another precise adjustment to operate well when using a different camera or in different environments. The existing VO methods based on deep learning require few manual calibration, but most of them occupy a tremendous amount of computing resources and cannot realize real-time VO. We propose a highly real-time VO system based on the optical flow and DenseNet structure accompanied with the inertial measurement unit (IMU). It cascade the optical flow network and DenseNet structure to calculate the translation and rotation, using the calculated information and IMU for construction and self- correction of the map. We have verified its computational complexity and performance on the KITTI dataset. The experiments have shown that the proposed system only requires less than 50% computation power than the main stream deep learning VO. It can also achieve 30% higher translation accuracy as well. Zibin Guo, Ninghao Chen, Zhuoling Xiao, Bo Yan 0007, Shuisheng Lin |
GLOBECOM | 5 |
| 2019 | AZUPT: Adaptive Zero Velocity Update Based on Neural Networks for Pedestrian TrackingabstractZero Velocity Update (ZUPT) has played a key role in Pedestrian Dead Reckoning (PDR) with inertial measurement units (IMU). However, it is both crucial and difficult to determine ZUPT conditions given complex and varying motion types such as walking, fast walking or running, and different walking habits of distinct people, which have direct and significant impact on the tracking accuracy. In this research we proposed a model based on deep neural networks to determine moments when the ZUPT should be conducted. The proposed model ensures nearly identical performance regardless of different motion types. It has been demonstrated by extensive experiments conducted in three different scenarios that our model can work equally well with different pedestrians and walking patterns, enabling the wide use of PDR in real-world applications. Xinguo Yu, Xinyue Lan, Zhuoling Xiao, Shuisheng Lin, Bo Yan 0007 |
GLOBECOM | 6 |
| 2019 | The Research of Stance-Phase Detection to Improve ZUPT-Aided Pedestrian Navigation SystemabstractInertial navigation is a fundamental method for pervasive indoor tacking and navigation. Although PDR based on inertial navigation can achieve robust indoors and outdoors positioning, the positioning accuracy does not meet the accuracy we need, due to the error divergence of the system. We present ZUPT with Kalman filter, a precise, robust technique tracks well even when presented with very noisy sensor data. Key to our ZUPT is zero velocity detection, the step to determine if the person's foot is in stance phase during walking. We used three different methods to detect zero velocity moments and compare their accuracy. Finally, we found that ZUPT using asymptotic zero velocity detection greatly improved the accuracy of inertial navigation. We believe that such a convergent and high precision approach will improve the application of inertial navigation in indoor positioning. Jianbo Liang, Zhuoling Xiao, Bo Yan 0007, Shuisheng Lin, Xinchun Liu |
ISCAS | 4 |