EDBT 2026 Demo / reviewers in the wild / expert
Benshuang Chen
dblp:392/1987
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
3D vision · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › pose estimation
multi-modal pose estimation |
0.9 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Computer vision › 3D vision
pose estimation |
0.9 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Computer vision › 3D vision › camera pose estimation
relative pose regression |
0.9 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Medical and health informatics › surgical navigation
endoscopic navigation |
0.9 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Medical and health informatics
surgical robotics |
0.9 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Medical and health informatics › medical visualization
surgical visualization |
0.3 | 1 | 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning · ICRA 2025 |
Methods — techniques the papers use, named apart from their topics
optical flow · 1.7multimodal feature learning · 1.7attention mechanism · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NEPose: A novel benchmark dataset with an improved framework for vision-based nasal endoscope pose estimation
Liangjing Shao, Benshuang Chen, Xinrong Chen |
Pattern Recognit. | 2 |
| 2026 | EndoLoc: Relative Pose Regression Framework With Transformation and Correlation Features for Visual Localization of EndoscopeabstractReal-time localization of endoscope is significant for the navigation and automation of endoscopic diagnosis and minimally invasive surgery. However, traditional localization based on optical tracking or magnetic tracking is easily influenced by occlusion or electromagnetic interference, while the implementation is complicated. Meanwhile, transformation and correlation information in image pairs are still ignored in existing visual localization methods for endoscopy. In this work, a novel relative pose regression framework is proposed for relative pose estimation and absolute pose tracking of endoscope based on endoscopic videos. Firstly, scene features and transformation features are respectively extracted from endoscopic observations and the corresponding optical flow by the proposed feature encoder based on gated convolution, which can prevent gradient vanishing when training the encoder from scratch on endoscopic data. Furthermore, a novel correlation module based on cross-attention is proposed to extract correlation features from two input images, which can capture more key features in endoscopic frames with more limited vision from local to global. Moreover, a novel pose decoder with upsampling and downsampling on the channel dimension is utilized to extract richer representation from the concatenated feature map for relative transformation vector prediction. The proposed method outperforms the state-of-the-art methods on the datasets from nasal endoscopy and colonoscopy, with an average localization error of less than 5%. The further experiments also demonstrate the efficiency of the proposed method. The demo videos of visual localization can be found on https://endoloc.netlify.app/ Liangjing Shao, Benshuang Chen, Shuting Zhao, Fuming Yang, Xinrong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature LearningabstractReal-time ego-motion tracking for endoscope is a significant task for efficient navigation and robotic automation of endoscopy. In this paper, a novel framework is proposed to perform real-time ego-motion tracking for endoscope. Firstly, a multi-modal visual feature learning network is proposed to perform relative pose prediction, in which the motion feature from the optical flow, the scene features and the joint feature from two adjacent observations are all extracted for prediction. Due to more correlation information in the channel dimension of the concatenated image, a novel feature extractor is designed based on an attention mechanism to integrate multi-dimensional information from the concatenation of two continuous frames. To extract more complete feature representation from the fused features, a novel pose decoder is proposed to predict the pose transformation from the concatenated feature map at the end of the framework. At last, the absolute pose of endoscope is calculated based on relative poses. The experiment is conducted on three datasets of various endoscopic scenes and the results demonstrate that the proposed method outperforms state-of-the-art methods. Besides, the inference speed of the proposed method is over 30 frames per second, which meets the real-time requirement. The project page is here: remote-bmxs.netlify.app Liangjing Shao, Benshuang Chen, Shuting Zhao, Xinrong Chen |
ICRA | 2 |
| 2025 | EndoMODE: A Multimodal Visual Feature-Based Ego-Motion Estimation Framework for Monocular Odometry and Depth Estimation in Various Endoscopic ScenesabstractEgo-motion estimation is a critical task for both monocular odometry and depth estimation, which are significant for navigation and scene perception in endoscopy. Most of existing ego-motion estimation frameworks only utilize a single modality of visual feature. In this work, a novel multimodal visual feature-based framework is proposed to perform real-time vision-based ego-motion estimation for visual odometry and depth estimation in endoscopic scenes. In the framework, to extract correlation information of two adjacent endoscopic frames, a channel attention-based module is proposed to integrate multidimensional features from the concatenated image and a feature interaction module is proposed to fuse scene features from two observations. Moreover, a novel pose decoder based on depthwise separable convolution is proposed to extract multiscale feature representation. Furthermore, a fully supervised pipeline and a self-supervised pipeline with the proposed framework are designed for monocular odometry and depth estimation, respectively. In the experiments, the proposed framework is compared with several advanced methods on five different datasets of various endoscopic scenes. Both quantitative and qualitative results demonstrate the pipeline with the proposed framework provides the most accurate monocular odometry and depth estimation in different kinds of endoscopic scenes. Liangjing Shao, Benshuang Chen, Shuting Zhao, Xinrong Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | 3D Clothed Human Reconstruction From One In-the-Wild RGB ImageabstractIn recent years, much achievement have been made in the field of 3D clothed human reconstruction. However, most of researches performed not well for reconstruction from in-the-wild images due to the domain gap between the synthetic images of training datasets and the in-the-wild images. In this study, a modular model, including clothes encoder, body encoder and cloth generator, is proposed to perform 3D clothed human reconstruction from one single-view in-the-wild RGB image. In particular, we introduce the adaptive aggregation of convolution and multi-head attention into the cloth encoder and apply the adjustment of the segmentation at the preprocessing stage. According to experiments on MSCOCO and 3DPW datasets, the proposed method achieves state-of-the-art performance on 3D clothed human reconstruction from in-the-wild images compared with previous works. Liangjing Shao, Benshuang Chen, Xinrong Chen |
ICIP | 2 |
| 2024 | NETrack: A Lightweight Attention-Based Network for Real-Time Pose Tracking of Nasal Endoscope Based on Endoscopic Image
Liangjing Shao, Benshuang Chen, Xinrong Chen |
PRICAI (5) | 2 |