Rujun Song

dblp:52/7943 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SwinFVO: Self-Supervised Visual Odometry With Enhanced Global Spatiotemporal Perception
abstract
Pose estimation using visual sensors has become a fundamental component in robotic navigation and autonomous driving systems. Learning-based monocular visual odometry (VO) has attracted substantial attention due to its resilience to camera parameter variations and dynamic environments. Given that camera movement manifests as pixel-level motion across the entire image in optical flow data, capturing both global contextual information and local feature details is crucial for accurate pose estimation. To address this challenge, we propose SwinFVO, a novel self-supervised visual odometry framework that incorporates enhanced motion perception to achieve global spatial dependency modeling with temporal continuity. Leveraging quadrant-based motion characteristics, we perform cross-regional feature interaction through a refined Swin Transformer architecture. Two robust spatiotemporal feature extractors are designed to extend the single-frame-based Swin Transformer to a temporally-aware framework for sequential understanding. Through the exploration of long-range spatial correlations and preservation of temporal consistency, SwinFVO delivers accurate and consistent pose estimation. Extensive experiments across multiple datasets demonstrate the superior performance and generalization capability of SwinFVO in both pose and depth estimation tasks. It achieves competitive results against classical algorithms and outperforms related state-of-the-art (SOTA) methods by up to 20.6% and 72.4% on average translational and rotational evaluations, respectively.
Rujun Song, Ruoqi Li, Zhuoling Xiao, Bo Yan 0007
IEEE Trans. Circuits Syst. Video Technol.1
2025 ZSCIL: Zero-Shot Class Incremental Learning Method for Signal Recognition
abstract
A significant challenge in signal recognition tasks is identifying classes not present in the dataset. Zero-shot learning-based signal recognition addresses the challenge by identifying previously unseen classes in a mixed signal space without supervision. However, most existing methodologies are limited to one-time recognition processes. We propose a zero-shot class incremental learning (ZSCIL) method to achieve continuous unseen classes identification. Our model employs an encoder-decoder architecture and incorporates a triplet loss function to train the classifier, thereby enhancing the model’s ability to recognize mixed signals through a metric learning paradigm. Additionally, we utilize class incremental learning, where the identified unseen signals are stored in a fixed-size buffer with a maximum diversity data replay mechanism. These signals are then used for incremental training. The framework’s effectiveness and generality of our method are demonstrated through a series of experiments on two datasets. For instance, we achieved a significant 14.4% accuracy improvement for seen classes and that of the unseen classes by 4.2% on the DeepSig 2016.04C dataset. To the best of our knowledge, ZSCIL is the first method to implement sustainable identification for unseen classes in the mixed signal space.
Rujun Song, Sidi Liang, Di He 0002, Zhuoling Xiao, Bo Yan 0007
ISCAS2
2025 LoSeVO: Local Sequence Constraints for Deep Visual Odometry
abstract
Many current visual odometry (VO) methods that utilize deep learning primarily concentrate on the constraints of motion relationships between adjacent frames, neglecting the modeling of temporal correlations within sequence data. Consequently, this paper introduces LoSeVO to effectively capture and model the temporal correlation features present in images. We design a Joint Feature Extraction component that not only performs joint feature extraction on adjacent frames but also extracts features from cross-frame images, which are called feature-guided maps. Then, we apply a Local Consistency Constraint component to the joint features between adjacent frames. It can adaptively constrain adjacent frames at different temporal positions within a sequence using different feature-guided maps. Extensive experiments based on the KITTI and Malaga datasets have shown that, compared to our previous DeepAVO model, LoSeVO can improve pose estimation performance by up to 23% and 9% in translation and rotation estimation, respectively.
Rujun Song, Di He 0002, Tingyong Yang, Zhuoling Xiao, Bo Yan 0007
ISCAS3
2024 CLFusion: 3D Semantic Segmentation Based on Camera and Lidar Fusion
abstract
In the field of autonomous driving, semantic segmentation is crucial for scene understanding. Currently, there are two main methods: camera-based and Lidar-based approaches. To address the issues of Lidar segmentation lacking texture features and image segmentation lacking distance information, this paper proposes a fusion of camera and Lidar to achieve 3D semantic segmentation. The method utilizes a dual-stream encoder-decoder network to process camera images and Lidar point cloud and incorporates a specially designed attention mechanism module for feature fusion. To avoid expensive manual annotation of 3D point clouds, the study also introduces a cross-dataset and cross-modal self-supervised training approach. Experimental results show a 2.4% improvement compared to the Lidar-only mode baseline results on the SemanticKITTI dataset and a 6% improvement on the nuScenes dataset.
Tianyue Wang, Rujun Song, Zhuoling Xiao, Bo Yan 0007, Haojie Qin, Di He 0002
ISCAS2
2024 GraphAVO: Self-Supervised Visual Odometry Based on Graph-Assisted Geometric Consistency
abstract
Learning-based monocular visual odometry (VO) has recently attracted considerable attention for its robustness to camera parameters and environmental variations. Despite traditional pose graph optimization enhancing pose accuracy, its integration with deep learning may lead to error accumulation due to insufficient motion information exchange. Our method, GraphAVO, concentrates simultaneously on the adjacent and interval co-visibility correspondence to establish a feature and pose graph optimization for pose consistency. We design a graph-assisted Windowed Feature Graph Refinement (WFGR) component to operationalize pose graph optimization for deep feature refinement. The geometric consistency is further constrained by a Cycle Consistency Loss. Additionally, the Cascade Dilated Convolution Fusion (CDCF) component is incorporated to handle different degrees of pixel movement, facilitating the joint detection of slight and distinct motion cues for subsequent feature enhancement. Extensive experiments on the KITTI, Malaga, RobotCar, and self-collected outdoor datasets have demonstrated the promising performance and generalization ability of GraphAVO. It achieves competitive results against classical algorithms and outperforms related state-of-the-art methods by up to 24.4% and 40.1% on average translational and rotational evaluation, respectively.
Rujun Song, Zhuoling Xiao, Bo Yan 0007
IEEE Trans. Intell. Transp. Syst.1
2023 Self-supervised Visual Odometry Based on Geometric Consistency
abstract
Learning-based monocular visual odometry (VO) has lately drawn significant attention for its robustness to camera parameters and environmental variations. Unlike most self-supervised learning-based methods, our approach simultaneously focuses on the adjacent and interval co-visibility correspondence to improve the pose estimation. To handle different pixel displacements, we apply the Multi-scale Feature Fusion component for the full exploration of latent motion features. Besides, the Interval Feature Guided Refinement component is incorporated to adaptively exploit the continuity of camera motions and steer the network for retaining pose consistency in the time domain. Extensive experiments on the KITTI and Malaga datasets have demonstrated the promising performance of our approaches. The proposed method produces competitive results against classic algorithms and outperform state-of-the-art methods by up to 23.9 % and 15.4 % on average translational and rotational evaluation.
Rujun Song, Kaisheng Liao, Zhuoling Xiao, Bo Yan 0007
ISCAS1
2023 ContextAVO: Local context guided and refining poses for deep visual odometry
Rujun Song, Zhuoling Xiao, Bo Yan 0007
Neurocomputing1
2022 DeepAVO: Efficient pose refining with feature distilling for deep Visual Odometry
Rujun Song, Bo Yan 0007, Zhuoling Xiao
Neurocomputing4