Shaoan Wang

dblp:326/4452 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-8175-1567ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Perceptive Locomotion and Navigation for Quadruped Robots via Depth-Based Representation
abstract
Enabling quadruped robots to navigate complex, unstructured 3-D environments using onboard sensors remains a significant challenge, particularly in the absence of maps. This difficulty stems from coupling long-horizon decision-making with terrain-aware control under partial observability and limited onboard perception. To address this, we propose a depth-based distillation learning framework featuring a hierarchical architecture that decouples high-level navigation from low-level locomotion. The navigation policy predicts velocity commands from raw depth and goal observations, while the locomotion controller executes motor actions based on proprioception and compact depth features distilled from privileged geometric scans of terrain. To improve representation quality and training efficiency, we incorporate contrastive forward prediction and inverse dynamics modeling into the reinforcement learning loop. A GPU-parallel Warp-based depth renderer is integrated into Isaac Gym to accelerate visual simulation and support large-scale training. To ensure robustness in cluttered environments, we employ an adaptive curriculum learning strategy that progressively increases terrain complexity during training. Our system enables robust, map-free navigation across both structured and unstructured terrain, and demonstrates successful zero-shot deployment on a real quadruped robot.
Aocheng Luo, Shaoan Wang, Shihan Kong, Kaiwei Zhu, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.3
2025 An Evidence-Based Tri-Branch Cross-Pseudo Supervision Method for Semi-Supervised Medical Image Segmentation
abstract
The semi-supervised medical image segmentation with a few annotated data can provide significant help in robot-assisted surgery. This step plays a pivotal role in identification of pathological regions, more appropriate planning of surgical procedures, and so on. In this work, we develop an evidence-based tri-branch cross-pseudo supervision model, which integrates evidence-based uncertainty estimation and multi-branch cross supervision to bolster the effectiveness of semi-supervised learning. The overall framework consists of a vanilla network and an evidential dual-branch network. Two evidential branches EPB and ERB are proposed to complement each other and improve the quality of pseudo-labels. The EPB places more focus on classification accuracy at the pixel level and the ERB emphasizes the similarity and overall integrity of the segmented regions. Then, a novel cross-pseudo supervision strategy among the three branches is designed, to guarantee that valuable and diverse unlabeled knowledge is explored and transferred for segmentation improvement. The effectiveness of the proposed method was verified on the ACDC dataset, achieving outstanding performance compared with other state-of-the-art methods. In addition, we conducted ablation study to validate the effectiveness of the evidential branches (EPB and ERB) and tri-branch cross-supervision strategy, respectively.
Aocheng Luo, Shaoan Wang, Yaoqing Hu, Jie Pan 0008, Junzhi Yu 0001
IROS3
2025 Adjusting Distributed Cameras for Robust Moving Object Pose Estimation
abstract
Robust moving object pose estimation is crucial in fine manipulation tasks, such as surgical instrument tracking. This paper presents a distributed-camera system with robotic adjustments to maintain consistent tracking of moving objects, thus avoiding tracking failures. An integrated framework for camera adjustment and pose estimation is developed for this distributed-camera system. In each detection cycle, the camera exhibiting the largest deviation with the object is adjusted by a visual servoing technique. After adjustment, the camera extrinsics are re-calibrated in the following detection cycles. For the unadjusted cameras, an online extrinsic optimization method based on multi-frame detection results is proposed to refine the camera extrinsics. Based on the refined camera extrinsics and detection results from multiple cameras, the pose of moving objects relative to the principal camera can be robustly estimated. We test the performance of this system in both simulation environments and real-world scenarios. The results indicate that our system achieves higher pose estimation accuracy and exhibits strong resistance to limited field-of-view (FoV) compared to conventional equivalent fixed multi-camera systems.
Yaoqing Hu, Shaoan Wang, Xingyu Chen 0002, Mingzhu Zhu, Zhanhua Xin, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.2
2025 A Lightweight Integrated Positioning System With Occlusion-Aware Region-Based Pose Tracking for Oral and Maxillofacial Surgery
abstract
The development of an accurate and robust positioning system for oral and maxillofacial surgery (OMS) is a challenging task, primarily due to the oral space limitations and line-of-sight occlusions. This paper presents a novel lightweight integrated positioning system for OMS, which can provide practical guidance utilizing only a micro camera installed on the end of the surgical instrument. An efficient region-based pose tracking method for texture-less teeth is proposed, which can use search lines around the object contour and simple local region partitioning strategy to improve pose accuracy. Besides, to deal with the possible partial occlusions of target during surgery, an occlusion-aware weight function is presented and utilized seamlessly in the pose optimization pipeline. This function calculates the pixel-wise occlusion probability using object contour and distance constraint, helping to improve the tracking robustness. Pivot calibration evaluation reveals that the tracking accuracy of the proposed camera-based handpiece is higher than the marker-based handpieces. Comparative experiments demonstrate that proposed pose tracking method has higher accuracy than existing state-of-the-art methods and ablation study confirms the effectiveness of the occlusion handling strategy. The overall positioning experiment indicates that the proposed system has satisfactory static poses stability and positioning accuracy. Furthermore, the main advantage of our system is that it is lighter and more integrated than other systems, which can reduce the system complexity, decrease the risk of line-of-sight occlusion, and lower the surgery cost.Note to Practitioners—This paper is motivated by the problem of restricted oral space constraints and partial occlusions during positioning for OMS. Compared with traditional OMS navigation systems, the designed system is more lightweight and more integrated without other external cameras and additional fiducial markers. Our system can provide practical guidance utilizing only a micro camera installed on the end of the surgical instrument. In addition, an efficient region-based pose tracking method for texture-less teeth is proposed to increase pose accuracy. Since the target can partially be occluded during the procedure, we present a novel occlusion-aware strategy to improve the tracking performance of partial occlusions. Our proposed system achieves a decent balance between positioning accuracy and hardware cost, and can easily be integrated into various dental surgical tools, thus it has tremendous potential for commercialization.
Yaoqing Hu, Shaoan Wang, Mingzhu Zhu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.3
2025 Accurate and Automatic Dental Crown Components Segmentation With Multi-Scale Attention Based U-Net and Hybrid Level Set Models
abstract
This paper presents a two-step method to automatically and accurately segment the dental crown components from CT images. Firstly, a multi-scale attention based U-Net model is proposed for pulp segmentation, which is embedded with global and local attention modules. The constructed attention modules can automatically aggregate pixel-wise contextual information and focus on catching the real dental pulp region. Secondly, two efficient level set models are proposed: one is the shape constraint-based level set model for enamel and dentin segmentation, the other is the region mutual exclusion-based level set model for neighboring teeth segmentation. The proposed shape constraint term can better handle topology changes of teeth and the region mutual exclusion term can more effectively avoid intersecting segmentation. Besides, a starting slice initialization method is introduced to achieve automatic segmentation, and an accurate contour propagation strategy is developed for slice-by-slice segmentation. We set up a series of comparative experiments for evaluation. Experimental results verify that the proposed method obtains promising performance for each crown component segmentation, and outperforms state-of-the-art tooth segmentation methods in terms of accuracy. This suggests that the proposed method can be used to accurately segment the crown components for precise tooth preparation treatment.Note to Practitioners—The motivation of this work is to reduce the burden on dentists during tooth preparation treatment, which requires accurate segmentation of crown components (i.e., enamel, dentin, and pulp) from dental CT images. Existing methods only focused on the segmentation of teeth or alveolar bone. Therefore, we present a novel automatic segmentation model for the dental crown components with high accuracy. A key strength of this study is the combination of a data-driven method (deep learning) and model-driven methods (level-set), which can provide good accuracy under limited training samples. This ability is highly desirable for practitioners by saving labor-intensive, costly labeling efforts. Furthermore, our proposed method will provide tools to help reduce subjectivity and human errors, as well as streamline and expedite the clinical workflow. This will significantly facilitate tooth preparation automation.
Mingzhu Zhu, Shaoan Wang, Yaoqing Hu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.3
2024 A Novel Lightweight Navigation System for Oral and Maxillofacial Surgery Using an External Curved Self-Identifying Checkerboard
abstract
This paper presents a novel lightweight navigation system for oral and maxillofacial surgery (OMS). An external curved checkerboard with self-identifying markers is set as the reference object around the surgical scene. A customized oral clip with a micro camera is designed for oral localization by tracking the external checkerboard. Similarly, the dental handpiece is also equipped with a micro camera, which can be localized like the clip. The spatial model of the markers is provided by binocular stereoscopic reconstruction. A front surface mirror is taken up for the registration between the oral cavity and the camera on the clip. The pivot calibration of the dental handpiece is accomplished by our proposed calibration method. We set up an experimental group and a control group for evaluation. The surgical tools of the experimental group were approximately 30% lighter, 35% less bulky, and 90% cheaper than those of the control group. Our system yielded the comprehensive navigation accuracy of 0.92 mm whereas the accuracy of the control group was 0.87 mm. Results revealed that our system can achieve similar accuracy compared with a prevailing system at a lighter weight, a more compact volume, and a lower cost. Note to Practitioners—The motivation of this work is to reduce the burden on patients and surgeons during OMS. Current commercial navigation systems are still limited by the burden of extra cumbersome fiducial markers and high hardware costs. Their high accuracy benefits from the large size of fiducial markers. To take full advantage of the camera’s localization effect, we propose the concept of “marker-camera inverse projection”, i.e., reversing the roles of the camera and the markers. In this way, cameras on surgical tools detect more points with a more uniform distribution. Our proposed system achieves a decent balance between navigation accuracy and hardware cost, which facilitates the development of surgical tools to be lighter and more economical, and involves tremendous potential for commercialization.
Yaoqing Hu, Mingzhu Zhu, Shaoan Wang, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.3
2024 CylinderTag: An Accurate and Flexible Marker for Cylinder-Shape Objects Pose Estimation Based on Projective Invariants
abstract
High-precision pose estimation based on visual markers has been a thriving research topic in the field of computer vision. However, the suitability of traditional flat markers on curved objects is limited due to the diverse shapes of curved surfaces, which hinders the development of high-precision pose estimation for curved objects. Therefore, this paper proposes a novel visual marker called CylinderTag, which is designed for developable curved surfaces such as cylindrical surfaces. CylinderTag is a cyclic marker that can be firmly attached to objects with a cylindrical shape. Leveraging the manifold assumption, the cross-ratio in projective invariance is utilized for encoding in the direction of zero curvature on the surface. Additionally, to facilitate the usage of CylinderTag, we propose a heuristic search-based marker generator and a high-performance recognizer as well. Moreover, an all-encompassing evaluation of CylinderTag properties is conducted by means of extensive experimentation, covering detection rate, detection speed, dictionary size, localization jitter, and pose estimation accuracy. CylinderTag showcases superior detection performance from varying view angles in comparison to traditional visual markers, accompanied by higher localization accuracy. Furthermore, CylinderTag boasts real-time detection capability and an extensive marker dictionary, offering enhanced versatility and practicality in a wide range of applications. Experimental results demonstrate that the CylinderTag is a highly promising visual marker for use on cylindrical-like surfaces, thus offering important guidance for future research on high-precision visual localization of cylinder-shaped objects.
Shaoan Wang, Mingzhu Zhu, Yaoqing Hu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans. Vis. Comput. Graph.1