Boshu Lei

dblp:319/3035 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0000-1153-1702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 43% Robot navigation and mapping · 27% Robot manipulation · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
tactile sensing
1.322026
TensorTouch: Calibration of Tactile Sensors for High Resolution Stress Tensor and Deformation for Dexterous Manipulation · IEEE Trans. Robotics 2026
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Robotics › Robot navigation and mapping › active perception
active mapping
0.912025
Multimodal LLM Guided Exploration and Active Mapping Using Fisher Information · ICCV 2025
Robotics › Robot navigation and mapping
active perception
0.912025
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Computer vision › 3D vision › depth estimation
depth uncertainty
0.912025
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Machine learning › Reinforcement learning
exploration
0.912025
Multimodal LLM Guided Exploration and Active Mapping Using Fisher Information · ICCV 2025
Robotics › Robot navigation and mapping › view planning
next-best-view planning
0.912025
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Robotics › Robot navigation and mapping › active vision
active view selection
0.812024
FisherRF: Active View Selection and Mapping with Radiance Fields Using Fisher Information · ECCV (13) 2024
Computer vision › 3D vision
camera pose estimation
0.812024
Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization · ICRA 2024
Computer vision › Image recognition and object detection
image retrieval
0.812024
Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization · ICRA 2024
Computer vision › 3D vision › novel view synthesis
radiance field
0.812024
FisherRF: Active View Selection and Mapping with Radiance Fields Using Fisher Information · ECCV (13) 2024
Computer vision › 3D vision
visual localization
0.812024
Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization · ICRA 2024
Computer vision › 3D vision › local feature descriptor
descriptor learning
0.712023
GRM: Gradient Rectification Module for Visual Place Retrieval · ICRA 2023
Robotics › Robot manipulation
dexterous manipulation
0.312026
TensorTouch: Calibration of Tactile Sensors for High Resolution Stress Tensor and Deformation for Dexterous Manipulation · IEEE Trans. Robotics 2026
Robotics › Robot manipulation
grasping
0.312026
TensorTouch: Calibration of Tactile Sensors for High Resolution Stress Tensor and Deformation for Dexterous Manipulation · IEEE Trans. Robotics 2026
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312025
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Computer vision › 3D vision
3d reconstruction
0.312025
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting · ICRA 2025
Computer vision › 3D vision
structure from motion
0.212024
Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization · ICRA 2024
Computer vision › Image recognition and object detection
image classification
0.212023
GRM: Gradient Rectification Module for Visual Place Retrieval · ICRA 2023

Methods — techniques the papers use, named apart from their topics

fisher information · 1.6gradient rectification · 1.3learning-based calibration · 1.0finite element analysis · 1.0semantic depth alignment · 0.9segment anything model 2 · 0.9multimodal large language model · 0.9FisherRF · 0.9pose estimation · 0.8local feature aggregation · 0.8prototype learning · 0.7
YearPublicationVenuePosition
2026 TensorTouch: Calibration of Tactile Sensors for High Resolution Stress Tensor and Deformation for Dexterous Manipulation
abstract
Advanced dexterous manipulation requires identifying and controlling multiple simultaneous contacts, including interactions with compliant objects where deformations are large. Raw optical tactile images are information rich but lack calibrated physical meaning, limiting cross-sensor use and real-world deployment. We present TensorTouch, a calibration framework that combines finite element analysis and learning to infer dense deformation and stress/force fields from a single tactile image. Across real sensors, TensorTouch achieves contact localization errors under 1.29 mm and mean force errors under 0.139 N per axis. In a multi-object selective grasp task with two simultaneously contacted objects (including identical cables), the system achieves up to 90.0% success. We further demonstrate robustness under repeated loading, yielding 38.295 dB PSNR between the initial tactile image and the image after 20,000 contacts. The learned model runs in real time at 95 Hz on an RTX 5090, proving to be suitable for contact-rich, dexterous manipulation.
Won Kyung Do, Matthew Strong, Aiden Swann, Boshu Lei, Monroe Kennedy III
IEEE Trans. Robotics4
2025 Multimodal LLM Guided Exploration and Active Mapping Using Fisher Information
Wen Jiang 0008, Boshu Lei, Katrina Ashton, Kostas Daniilidis
ICCV2
2025 Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting
abstract
We propose a framework for active next best view and touch selection for robotic manipulators using 3D Gaussian Splatting (3DGS). 3DGS is emerging as a useful explicit 3D scene representation for robotics, as it has the ability to represent scenes in a both photorealistic and geometrically accurate manner. However, in real-world, online robotic scenes where the number of views is limited given efficiency requirements, random view selection for 3DGS becomes impractical as views are often overlapping and redundant. We address this issue by proposing an end-to-end online training and active view selection pipeline, which enhances the performance of 3DGS in few-view robotics settings. We first elevate the performance of few-shot 3DGS with a novel semantic depth alignment method using Segment Anything Model 2 (SAM2) that we supplement with Pearson depth and surface normal loss to improve color and depth reconstruction of real-world scenes. We then extend FisherRF, a next-best-view selection method for 3DGS, to select views and touch poses based on depth uncertainty. We perform online view selection on a real robot system during live 3DGS training. We motivate our improvements to few-shot GS scenes, and extend depth-based FisherRF to them, where we demonstrate both qualitative and quantitative improvements on challenging robot scenes. For more information, please see our project page at arm.stanford.edu/next-best-sense.
Matthew Strong, Boshu Lei, Aiden Swann, Wen Jiang 0008, Kostas Daniilidis, Monroe Kennedy III
ICRA2
2024 FisherRF: Active View Selection and Mapping with Radiance Fields Using Fisher Information
Wen Jiang 0008, Boshu Lei, Kostas Daniilidis
ECCV (13)2
2024 Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization
abstract
Visual localization is an important sub-task in SfM and visual SLAM that involves estimating a 6-DoF camera pose for an input query image relative to a given 3D model of the environment. The most accurate approach is a hierarchical one that splits the task into two stages: image retrieval and camera pose estimation. Each stage requires different image features, with global features compactly encoding holistic image information for the first stage and local features encoding the appearance around salient image points for the second stage. While existing methods use independent networks to extract these features, one for global and one for local, this strategy is suboptimal in terms of computational efficiency. In this paper, we propose a novel approach that achieves state-of-the-art inference accuracy with significantly improved efficiency. Our approach’s core component is SuperGF, a network that aggregates local features optimized for camera pose estimation to create a global feature that enables precise image retrieval. Through extensive experiments on the standard benchmark tests, we demonstrate that the method offers a better trade-off between accuracy and computational cost.
Wenzheng Song, Boshu Lei, Takayuki Okatani
ICRA3
2023 GRM: Gradient Rectification Module for Visual Place Retrieval
abstract
Visual place retrieval aims to search images in the database that depict similar places as the query image. However, global descriptors encoded by the network usually fall into a low dimensional principal space, which is harmful to the retrieval performance. We first analyze the cause of this phenomenon, pointing out that it is due to degraded distribution of the gradients of descriptors. Then, we propose Gradient Rectification Module (GRM) to alleviate this issue. GRM is appended after the final pooling layer and can rectify gradients to the complementary space of the principal space. With GRM, the network is encouraged to generate descriptors more uniformly in the whole space. At last, we conduct experiments on multiple datasets and generalize our method to classification task under prototype learning framework.
Boshu Lei, Limeng Qiao, Xi Qiu
ICRA1