EDBT 2026 Demo / reviewers in the wild / expert
Shaojun Cai
dblp:164/8455
· DBLP profile ↗
11ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
2 papers |
Human-robot interaction · 72% Accessibility and assistive technology · 28% | |
| Artificial intelligence
4 papers |
3D vision · 46% Robot navigation and mapping · 44% Planning, search and constraint satisfaction · 10% |
Topics — the 10 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Human-robot interaction
assistive robotics |
1.0 | 1 | 2026 | Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental Interactions · HRI 2026 |
Human-robot interaction
human-robot collaboration |
1.0 | 1 | 2026 | Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental Interactions · HRI 2026 |
Accessibility and assistive technology › assistive navigation
navigation assistance for visually impaired |
0.8 | 1 | 2024 | Navigating Real-World Challenges: A Quadruped Robot Guiding System for Visually Impaired People in Diverse Environments · CHI 2024 |
Robotics › Robot navigation and mapping
visual navigation |
0.5 | 1 | 2021 | Differentiable SLAM-Net: Learning Particle SLAM for Visual Navigation · CVPR 2021 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.5 | 1 | 2021 | Differentiable SLAM-Net: Learning Particle SLAM for Visual Navigation · CVPR 2021 |
Computer vision › 3D vision
camera pose estimation |
0.4 | 1 | 2020 | Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020 |
Computer vision › 3D vision › visual localization
camera relocalization |
0.4 | 1 | 2020 | Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020 |
Computer vision › 3D vision
multi-view geometry |
0.4 | 1 | 2020 | Learning Multi-View Camera Relocalization With Graph Neural Networks · CVPR 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › multi-agent planning
intention-aware planning |
0.2 | 1 | 2015 | Intention-aware online POMDP planning for autonomous driving in a crowd · ICRA 2015 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
POMDP planning |
0.1 | 1 | 2015 | Intention-aware online POMDP planning for autonomous driving in a crowd · ICRA 2015 |
Methods — techniques the papers use, named apart from their topics
voice feedback · 1.5force feedback · 1.5lead mode and adaptation mode · 1.0collaborative human-robot approach · 1.0particle filter · 0.5differentiable computation graph · 0.5backpropagation · 0.5graph neural network · 0.4convolutional neural network · 0.4online planning · 0.2POMDP · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Navigation beyond Wayfinding: Robots Collaborating with Visually Impaired Users for Environmental InteractionsabstractRobotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of navigation—environmental interactions that require manipulating objects to enable movement. These interactions are challenging for a human–robot pair because they demand (i) precise localization and manipulation of interaction targets (e.g., pressing elevator buttons) and (ii) dynamic coordination between the user’s and robot’s movements (e.g., pulling out a chair to sit). We present a collaborative human–robot approach that combines our robotic guide dog’s precise sensing and localization capabilities with the user’s ability to perform physical manipulation. The system alternates between two modes: lead mode, where the robot detects and guides the user to the target, and adaptation mode, where the robot adjusts its motion as the user interacts with the environment (e.g., opening a door). Evaluation results show that our system enables navigation that is safer, smoother, and more efficient than both a traditional white cane and a non-adaptive guiding system, with the performance gap widening as tasks demand higher precision in locating interaction targets. These findings highlight the promise of human–robot collaboration in advancing assistive technologies toward more generalizable and realistic navigation support. Shaojun Cai, Nuwan Janaka, Ashwin Ram 0002, Janidu Shehan, Yingjia Wan, Kotaro Hara, David Hsu |
HRI | 1 |
| 2026 | Multimodal Feature Interaction and High-Quality Pseudolabel Generation With Self-Training for Cognitive State DetectionabstractCognitive state detection holds significant research value in the field of human–computer interaction and neural engineering. However, existing works are insufficient in modeling the temporal dynamics of multimodal physiological signals, which leads to heterogeneous distribution differences in cross-modal feature interactions. In addition, domain shift issues under cross-subject and few-sample conditions restrict the model generalization performance. To cope with these problems, this work proposes a cognitive state detection framework that integrates Transformer-based multimodal feature interaction and self-training of pseudolabel optimization. First, the multihead attention mechanism is introduced to model the temporal evolution patterns across modalities, dynamically harmonizing cross-modal contributions to extract cognitive state-related shared features. Then, a dual-model cross-validation strategy is designed to filter high-quality pseudolabeled samples from the target domain for subsequent self-training, effectively avoiding the dependency on auxiliary modules in domain adaptation. Finally, Extensive experiments show that the proposed work significantly improves the recognition accuracy, and the designed pseudolabel optimization mechanism can be transferred to related tasks without increasing model complexity. Kevin W. Tong, Xuefeng Men, Haoran Duan 0001, Shaojun Cai, Changyu Li, Ping Li 0044, Guangyu Zhu 0001, Qi Wu 0003, Limin Zhu 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Navigating Real-World Challenges: A Quadruped Robot Guiding System for Visually Impaired People in Diverse EnvironmentsabstractBlind and Visually Impaired (BVI) people find challenges in navigating unfamiliar environments, even using assistive tools such as white canes or smart devices. Increasingly affordable quadruped robots offer us opportunities to design autonomous guides that could improve how BVI people find ways around unfamiliar environments and maneuver therein. In this work, we designed RDog, a quadruped robot guiding system that supports BVI individuals’ navigation and obstacle avoidance in indoor and outdoor environments. RDog combines an advanced mapping and navigation system to guide users with force feedback and preemptive voice feedback. Using this robot as an evaluation apparatus, we conducted experiments to investigate the difference in BVI people’s ambulatory behaviors using a white cane, a smart cane, and RDog. Results illustrated the benefits of RDog-based ambulation, including faster and smoother navigation with fewer collisions and limitations, and reduced cognitive load. We discuss the implications of our work for multi-terrain assistive guidance systems. Shaojun Cai, Ashwin Ram 0002, Zhengtai Gou, Mohd Alqama Wasim Shaikh, Yu-An Chen, Yingjia Wan, Kotaro Hara, Shengdong Zhao 0001, David Hsu |
CHI | 1 |
| 2021 | Differentiable SLAM-Net: Learning Particle SLAM for Visual NavigationabstractSimultaneous localization and mapping (SLAM) remains challenging for a number of downstream applications, such as visual robot navigation, because of rapid turns, featureless walls, and poor camera quality. We introduce the Differentiable SLAM Network (SLAM-net) along with a navigation architecture to enable planar robot navigation in previously unseen indoor environments. SLAM-net encodes a particle filter based SLAM algorithm in a differentiable computation graph, and learns task-oriented neural network components by backpropagating through the SLAM algorithm. Because it can optimize all model components jointly for the end-objective, SLAM-net learns to be robust in challenging conditions. We run experiments in the Habitat platform with different real-world RGB and RGB-D datasets. SLAM-net significantly outperforms the widely adapted ORB-SLAM in noisy conditions. Our navigation architecture with SLAMnet improves the state-of-the-art for the Habitat Challenge 2020 PointNav task by a large margin (37% to 64% success). Project website: http://sites.google.com/view/slamnet Péter Karkus, Shaojun Cai, David Hsu |
CVPR | 2 |
| 2020 | A Self-Learned Arbitration Between Model-Based and Model-Free Navigation Strategies in Autonomous Driving
Shaojun Cai, Yingjia Wan |
CogSci | 1 |
| 2020 | Learning Multi-View Camera Relocalization With Graph Neural NetworksabstractWe propose to construct a view graph to excavate the information of the whole given sequence for absolute camera pose estimation. Specifically, we harness GNNs to model the graph, allowing even non-consecutive frames to exchange information with each other. Rather than adopting the regular GNNs directly, we redefine the nodes, edges, and embedded functions to fit the relocalization task. Redesigned GNNs cooperate with CNNs in guiding knowledge propagation and feature extraction respectively to process multi-view high-dimension image features iteratively at different levels. Besides, a general graph-based loss function beyond constraints between consecutive views is employed for training the network in an end-to-end fashion. Extensive experiments conducted on both indoor and outdoor datasets demonstrate that our method outperforms previous approaches especially in large-scale and challenging scenarios. Shaojun Cai, Junqiu Wang |
CVPR | 3 |
| 2019 | Localizing Discriminative Visual Landmarks for Place RecognitionabstractWe address the problem of visual place recognition with perceptual changes. The fundamental problem of visual place recognition is generating robust image representations which are not only insensitive to environmental changes but also distinguishable to different places. Taking advantage of the feature extraction ability of Convolutional Neural Networks (CNNs), we further investigate how to localize discriminative visual landmarks that positively contribute to the similarity measurement, such as buildings and vegetations. In particular, a Landmark Localization Network (LLN) is designed to indicate which regions of an image are used for discrimination. Detailed experiments are conducted on open source datasets with varied appearance and viewpoint changes. The proposed approach achieves superior performance against state-of-the-art methods. Zhe Xin, Yinghao Cai, Tao Lu 0006, Xiaoxia Xing, Shaojun Cai, Jixiang Zhang 0001 |
ICRA | 5 |
| 2018 | 3DTNet: Learning Local Features Using 2D and 3D CuesabstractWe present an approach to learn 3D local descriptor by combining both 2D texture and 3D geometric information, which can be used to register partial 3D data for a variety of vision applications. Unlike previous approaches which simply concatenate features learned from multiple sources into one feature descriptor, we learn 2D and 3D feature representations jointly. We design a network, 3DTNet with an architecture particularly designed for learning robust local feature representation leveraging both texture and geometric information. Two types of information are interacted with each other which results in more robust and stable feature representation. Finally, feature representations of multi-scale neighborhoods are aggregated to further improve the performance of feature matching. Extensive experimental results show that our method outperforms state-of-art 2D or 3D descriptors in terms of both accuracy and efficiency. Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Shaojun Cai, Dayong Wen |
3DV | 4 |
| 2018 | CubemapSLAM: A Piecewise-Pinhole Monocular Fisheye SLAM System
Shaojun Cai, Shijie Li 0006, Yun Liu 0011, Yangyan Guo, Tao Li 0022, Ming-Ming Cheng |
ACCV (6) | 2 |
| 2018 | Visual Localization in Changing Environments using Place Recognition TechniquesabstractThis paper proposes a visual localization system combining Convolutional Neural Networks (CNNs) and sparse point features to estimate the 6-DOF pose of the robot. The challenges of visual localization across time lie in that the same place captured across time appears dramatically different due to different illumination and weather conditions, viewpoint variations and dynamic objects. In this paper, a novel CNN-based place recognition approach is proposed, which requires no time-consuming feature generation process and no task-specific training. Moreover, we demonstrate that the rich semantic context information obtained from place recognition can greatly improve the subsequent feature matching process for pose estimation. The semantic constraint performs much better than traditional Bag-of-Words based methods for establishing correspondences between the query image and the map. To evaluate the robustness of the algorithm, the proposed system is integrated into ORB-SLAM2 and verified on the data collected over various illumination and weather conditions. Extensive experimental results show that even with weak ORB descriptors, the proposed system can significantly improve the success rate of localization under severe appearance changes. Zhe Xin, Yinghao Cai, Shaojun Cai, Jixiang Zhang 0001 |
ICPR | 3 |
| 2015 | Intention-aware online POMDP planning for autonomous driving in a crowdabstractThis paper presents an intention-aware online planning approach for autonomous driving amid many pedestrians. To drive near pedestrians safely, efficiently, and smoothly, autonomous vehicles must estimate unknown pedestrian intentions and hedge against the uncertainty in intention estimates in order to choose actions that are effective and robust. A key feature of our approach is to use the partially observable Markov decision process (POMDP) for systematic, robust decision making under uncertainty. Although there are concerns about the potentially high computational complexity of POMDP planning, experiments show that our POMDP-based planner runs in near real time, at 3 Hz, on a robot golf cart in a complex, dynamic environment. This indicates that POMDP planning is improving fast in computational efficiency and becoming increasingly practical as a tool for robot planning under uncertainty. Haoyu Bai, Shaojun Cai, David Hsu, Wee Sun Lee |
ICRA | 2 |