Shaojun Cai

dblp:164/8455 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
2 papers
Human-robot interaction · 72% Accessibility and assistive technology · 28%
Artificial intelligence
4 papers
3D vision · 46% Robot navigation and mapping · 44% Planning, search and constraint satisfaction · 10%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Human-robot interaction
assistive robotics
1.012026
Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental Interactions · HRI 2026
Human-robot interaction
human-robot collaboration
1.012026
Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental Interactions · HRI 2026
Accessibility and assistive technology › assistive navigation
navigation assistance for visually impaired
0.812024
Navigating Real-World Challenges: A Quadruped Robot Guiding System for Visually Impaired People in Diverse Environments · CHI 2024
Robotics › Robot navigation and mapping
visual navigation
0.512021
Differentiable SLAM-Net: Learning Particle SLAM for Visual Navigation · CVPR 2021
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.512021
Differentiable SLAM-Net: Learning Particle SLAM for Visual Navigation · CVPR 2021
Computer vision › 3D vision
camera pose estimation
0.412020
Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020
Computer vision › 3D vision › visual localization
camera relocalization
0.412020
Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020
Computer vision › 3D vision
multi-view geometry
0.412020
Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › multi-agent planning
intention-aware planning
0.212015
Intention-aware online POMDP planning for autonomous driving in a crowd · ICRA 2015
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
POMDP planning
0.112015
Intention-aware online POMDP planning for autonomous driving in a crowd · ICRA 2015

Methods — techniques the papers use, named apart from their topics

voice feedback · 1.5force feedback · 1.5lead mode and adaptation mode · 1.0collaborative human-robot approach · 1.0particle filter · 0.5differentiable computation graph · 0.5backpropagation · 0.5graph neural network · 0.4convolutional neural network · 0.4online planning · 0.2POMDP · 0.2
YearPublicationVenuePosition
2026 Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental Interactions
abstract
Robotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of navigation—environmental interactions that require manipulating objects to enable movement. These interactions are challenging for a human–robot pair because they demand (i) precise localization and manipulation of interaction targets (e.g., pressing elevator buttons) and (ii) dynamic coordination between the user’s and robot’s movements (e.g., pulling out a chair to sit). We present a collaborative human–robot approach that combines our robotic guide dog’s precise sensing and localization capabilities with the user’s ability to perform physical manipulation. The system alternates between two modes: lead mode, where the robot detects and guides the user to the target, and adaptation mode, where the robot adjusts its motion as the user interacts with the environment (e.g., opening a door). Evaluation results show that our system enables navigation that is safer, smoother, and more efficient than both a traditional white cane and a non-adaptive guiding system, with the performance gap widening as tasks demand higher precision in locating interaction targets. These findings highlight the promise of human–robot collaboration in advancing assistive technologies toward more generalizable and realistic navigation support.
Shaojun Cai, Nuwan Janaka, Ashwin Ram 0002, Janidu Shehan, Yingjia Wan, Kotaro Hara, David Hsu
HRI1
2026 Multimodal Feature Interaction and High-Quality Pseudolabel Generation With Self-Training for Cognitive State Detection
abstract
Cognitive state detection holds significant research value in the field of human–computer interaction and neural engineering. However, existing works are insufficient in modeling the temporal dynamics of multimodal physiological signals, which leads to heterogeneous distribution differences in cross-modal feature interactions. In addition, domain shift issues under cross-subject and few-sample conditions restrict the model generalization performance. To cope with these problems, this work proposes a cognitive state detection framework that integrates Transformer-based multimodal feature interaction and self-training of pseudolabel optimization. First, the multihead attention mechanism is introduced to model the temporal evolution patterns across modalities, dynamically harmonizing cross-modal contributions to extract cognitive state-related shared features. Then, a dual-model cross-validation strategy is designed to filter high-quality pseudolabeled samples from the target domain for subsequent self-training, effectively avoiding the dependency on auxiliary modules in domain adaptation. Finally, Extensive experiments show that the proposed work significantly improves the recognition accuracy, and the designed pseudolabel optimization mechanism can be transferred to related tasks without increasing model complexity.
Kevin W. Tong, Xuefeng Men, Haoran Duan 0001, Shaojun Cai, Changyu Li, Ping Li 0044, Guangyu Zhu 0001, Qi Wu 0003, Limin Zhu 0001
IEEE Trans. Ind. Informatics6
2024 Navigating Real-World Challenges: A Quadruped Robot Guiding System for Visually Impaired People in Diverse Environments
abstract
Blind and Visually Impaired (BVI) people find challenges in navigating unfamiliar environments, even using assistive tools such as white canes or smart devices. Increasingly affordable quadruped robots offer us opportunities to design autonomous guides that could improve how BVI people find ways around unfamiliar environments and maneuver therein. In this work, we designed RDog, a quadruped robot guiding system that supports BVI individuals’ navigation and obstacle avoidance in indoor and outdoor environments. RDog combines an advanced mapping and navigation system to guide users with force feedback and preemptive voice feedback. Using this robot as an evaluation apparatus, we conducted experiments to investigate the difference in BVI people’s ambulatory behaviors using a white cane, a smart cane, and RDog. Results illustrated the benefits of RDog-based ambulation, including faster and smoother navigation with fewer collisions and limitations, and reduced cognitive load. We discuss the implications of our work for multi-terrain assistive guidance systems.
Shaojun Cai, Ashwin Ram 0002, Zhengtai Gou, Mohd Alqama Wasim Shaikh, Yu-An Chen, Yingjia Wan, Kotaro Hara, Shengdong Zhao 0001, David Hsu
CHI1
2021 Differentiable SLAM-Net: Learning Particle SLAM for Visual Navigation
abstract
Simultaneous localization and mapping (SLAM) remains challenging for a number of downstream applications, such as visual robot navigation, because of rapid turns, featureless walls, and poor camera quality. We introduce the Differentiable SLAM Network (SLAM-net) along with a navigation architecture to enable planar robot navigation in previously unseen indoor environments. SLAM-net encodes a particle filter based SLAM algorithm in a differentiable computation graph, and learns task-oriented neural network components by backpropagating through the SLAM algorithm. Because it can optimize all model components jointly for the end-objective, SLAM-net learns to be robust in challenging conditions. We run experiments in the Habitat platform with different real-world RGB and RGB-D datasets. SLAM-net significantly outperforms the widely adapted ORB-SLAM in noisy conditions. Our navigation architecture with SLAMnet improves the state-of-the-art for the Habitat Challenge 2020 PointNav task by a large margin (37% to 64% success). Project website: http://sites.google.com/view/slamnet
Péter Karkus, Shaojun Cai, David Hsu
CVPR2
2020 A Self-Learned Arbitration Between Model-Based and Model-Free Navigation Strategies in Autonomous Driving
Shaojun Cai, Yingjia Wan
CogSci1
2020 Learning Multi-View Camera Relocalization With Graph Neural Networks
abstract
We propose to construct a view graph to excavate the information of the whole given sequence for absolute camera pose estimation. Specifically, we harness GNNs to model the graph, allowing even non-consecutive frames to exchange information with each other. Rather than adopting the regular GNNs directly, we redefine the nodes, edges, and embedded functions to fit the relocalization task. Redesigned GNNs cooperate with CNNs in guiding knowledge propagation and feature extraction respectively to process multi-view high-dimension image features iteratively at different levels. Besides, a general graph-based loss function beyond constraints between consecutive views is employed for training the network in an end-to-end fashion. Extensive experiments conducted on both indoor and outdoor datasets demonstrate that our method outperforms previous approaches especially in large-scale and challenging scenarios.
Shaojun Cai, Junqiu Wang
CVPR3
2019 Localizing Discriminative Visual Landmarks for Place Recognition
abstract
We address the problem of visual place recognition with perceptual changes. The fundamental problem of visual place recognition is generating robust image representations which are not only insensitive to environmental changes but also distinguishable to different places. Taking advantage of the feature extraction ability of Convolutional Neural Networks (CNNs), we further investigate how to localize discriminative visual landmarks that positively contribute to the similarity measurement, such as buildings and vegetations. In particular, a Landmark Localization Network (LLN) is designed to indicate which regions of an image are used for discrimination. Detailed experiments are conducted on open source datasets with varied appearance and viewpoint changes. The proposed approach achieves superior performance against state-of-the-art methods.
Zhe Xin, Yinghao Cai, Tao Lu 0006, Xiaoxia Xing, Shaojun Cai, Jixiang Zhang 0001
ICRA5
2018 3DTNet: Learning Local Features Using 2D and 3D Cues
abstract
We present an approach to learn 3D local descriptor by combining both 2D texture and 3D geometric information, which can be used to register partial 3D data for a variety of vision applications. Unlike previous approaches which simply concatenate features learned from multiple sources into one feature descriptor, we learn 2D and 3D feature representations jointly. We design a network, 3DTNet with an architecture particularly designed for learning robust local feature representation leveraging both texture and geometric information. Two types of information are interacted with each other which results in more robust and stable feature representation. Finally, feature representations of multi-scale neighborhoods are aggregated to further improve the performance of feature matching. Extensive experimental results show that our method outperforms state-of-art 2D or 3D descriptors in terms of both accuracy and efficiency.
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Shaojun Cai, Dayong Wen
3DV4
2018 CubemapSLAM: A Piecewise-Pinhole Monocular Fisheye SLAM System
Shaojun Cai, Shijie Li 0006, Yun Liu 0011, Yangyan Guo, Tao Li 0022, Ming-Ming Cheng
ACCV (6)2
2018 Visual Localization in Changing Environments using Place Recognition Techniques
abstract
This paper proposes a visual localization system combining Convolutional Neural Networks (CNNs) and sparse point features to estimate the 6-DOF pose of the robot. The challenges of visual localization across time lie in that the same place captured across time appears dramatically different due to different illumination and weather conditions, viewpoint variations and dynamic objects. In this paper, a novel CNN-based place recognition approach is proposed, which requires no time-consuming feature generation process and no task-specific training. Moreover, we demonstrate that the rich semantic context information obtained from place recognition can greatly improve the subsequent feature matching process for pose estimation. The semantic constraint performs much better than traditional Bag-of-Words based methods for establishing correspondences between the query image and the map. To evaluate the robustness of the algorithm, the proposed system is integrated into ORB-SLAM2 and verified on the data collected over various illumination and weather conditions. Extensive experimental results show that even with weak ORB descriptors, the proposed system can significantly improve the success rate of localization under severe appearance changes.
Zhe Xin, Yinghao Cai, Shaojun Cai, Jixiang Zhang 0001
ICPR3
2015 Intention-aware online POMDP planning for autonomous driving in a crowd
abstract
This paper presents an intention-aware online planning approach for autonomous driving amid many pedestrians. To drive near pedestrians safely, efficiently, and smoothly, autonomous vehicles must estimate unknown pedestrian intentions and hedge against the uncertainty in intention estimates in order to choose actions that are effective and robust. A key feature of our approach is to use the partially observable Markov decision process (POMDP) for systematic, robust decision making under uncertainty. Although there are concerns about the potentially high computational complexity of POMDP planning, experiments show that our POMDP-based planner runs in near real time, at 3 Hz, on a robot golf cart in a complex, dynamic environment. This indicates that POMDP planning is improving fast in computational efficiency and becoming increasingly practical as a tool for robot planning under uncertainty.
Haoyu Bai, Shaojun Cai, David Hsu, Wee Sun Lee
ICRA2