Fangxun Zhong

dblp:178/2943 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-1151-1995ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 6-DoF Shape Servoing of Deformable Objects in Co-Rotated Space of Modal Graph
abstract
Shape control of deformable objects under both rotational and translational deformations is important for versatile robotic applications. However, deformation control with full 6-degree-of-freedom (DoF) manipulation is an open problem, since modeling and describing rotational deformations lead to significant challenges. To tackle the problem, this paper proposes a novel method by introducing a co-rotated space for the modal graph representation of objects with unknown physical and geometric models. In this space, we design new deformation features that can encode local rotations while preserving a compact and low-frequency shape representation. Moreover, these features can be mapped analytically to the robot manipulation, enabling the design of adaptive control laws with guaranteed stability for unmodeled objects. Experiments on complex volumetric objects demonstrate the effectiveness and advantage of our method with raw, noisy, and unregistered point clouds. The results highlight the importance of integrating co-rotated features to address rotational deformations.
Bohan Yang 0005, Fangxun Zhong, Yun-Hui Liu 0001
ICRA3
2025 Smooth Surface-to-Surface Contact Control for Rope-Base Soft-Tip Manipulator
abstract
A new control pipeline has been proposed for the Rope-Base Soft-tip Manipulator (RBSM) to execute the surface contact task to prevent the jamming and slipping problems. The control pipeline enables smooth surface-to-surface contact for the RBSM using only force sensors, eliminating the dependence on additional pose measurement of the window surface plane and soft-tip deformation information. The pipeline consists of three steps: free contact step implemented by an exponential force shape controller to avoid force overshoot to the window surface; orientation refinement step implemented by a force and torque combined controller to make the RBSM cleaning head surface stable adapt to the smooth window surface; and finally, a release normal force step to reduce head jamming and region covering with a pre-defined vibration-less cleaning trajectory for smooth cleaning on the slippery window surface. The proposed pipeline has been validated in a Rope base Cleaning Manipulator prototype to clean a common window surface. The force and velocity curves during the cleaning experiment show that the proposed method achieves smooth scraping and cleaning under unknown initial significant errors in surface orientation.
Guangli Sun, Fangxun Zhong, Peng Li 0019, Linzhu Yue, Zhi Chen 0020, Xiang Li 0009, Yun-Hui Liu 0001
IEEE Trans Autom. Sci. Eng.2
2025 Absolute Monocular Depth Estimation on Robotic Visual and Kinematics Data via Self-Supervised Learning
abstract
Accurate estimation of absolute depth from a monocular endoscope is a fundamental task for automatic navigation systems in robotic surgery. Previous works solely rely on uni-modal data (i.e., monocular images), which can only estimate depth values arbitrarily scaled with the real world. In this paper, we present a novel framework, SADER, which explores vision and robot kinematics to estimate the high-quality absolute depth for monocular surgical scenes. To jointly learn the multi-modal data, we introduce a self-distillation based two-stage training policy in the framework. In the first stage, a boosting depth module based on vision transformer is proposed to improve the relative depth estimation network that is trained in a self-supervised method. Then, we develop an algorithm to automatically compute the scale from robot kinematics. By coupling the scale and relative depth data, pseudo absolute depth labels for all images are yielded. In the second stage, we re-train the network with 3D loss supervised by pseudo labels. To make our method generalize to different endoscopes, the learning of endoscopic intrinsics is integrated into the network. In addition, we did cadaver experiments to collect new surgical depth estimation data about robotic laparoscopy for evaluation. Experimental results on public SCARED and cadaver data demonstrate that the SADER outperforms previous state-of-art even stereo-based methods with an accuracy error under 1.90 mm, proving the feasibility of our approach to recover the absolute depth with monocular inputs. Note to Practitioners—This paper aims to solve the problem of absolute monocular depth estimation in automatic surgical navigation by leveraging the multi-modal data from the robot-based endoscopic system. Accurate depth perception with real scales of the monocular scene is essential for the control of surgical robots in automatic navigation. However, current methods can only predict the relative depth of the surgical scene using monocular images. In this article, we propose a self-supervised learning-based method to achieve high-quality absolute depth estimation of monocular endoscopic images. It neither needs manual data annotation, nor other imaging modalities. The experiments extensively validate the feasibility and high performance of our framework for absolute depth estimation on monocular endoscopes. This absolute depth perception framework can be potentially encapsulated into the automatic navigation system in the near future.
Ruofeng Wei, Bin Li 0082, Fangxun Zhong, Hangjie Mo, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001
IEEE Trans Autom. Sci. Eng.3
2025 Efficient Underwater Object Detection With Enhanced Feature Extraction and Fusion
abstract
Underwater object detection is critical for applications, such as environmental monitoring, resource exploration, and the navigation of autonomous underwater vehicles. However, accurately detecting small objects in underwater environments remains challenging due to noisy imaging conditions, variable illumination, and complex backgrounds. To address these challenges, we propose the adaptive residual attention network (ARAN), an optimized deep learning framework designed to enhance the detection and precise identification of diminutive targets in complex aquatic settings. ARAN incorporates the proposed Fusion path aggregation network (PANet), which refines spatial features by effectively distinguishing objects from their backgrounds. The framework integrates three novel modules: first, multiscale feature attention, which enhances low-level feature extraction; second, high–low feature residual learning, which rearranges channel and batch dimensions to capture pixel-level relationships through cross-dimensional interactions; and third, multilevel feature dynamic aggregation, which dynamically adjusts fusion weights to facilitate progressive multilevel feature fusion and mitigate conflicts in multiscale integration, ensuring that small objects are not overshadowed. Extensive experiments on four benchmark datasets demonstrate that ARAN significantly outperforms mainstream models, achieving state-of-the-art performance. Notably, on the CSIRO dataset, ARAN attains a mean average precision at 50% of 98%, precision of 94.7%, F2-score of 94.6%, and recall of 94.7%. These results confirm our model's superior accuracy, robustness, and efficiency in underwater object detection, highlighting its potential for practical deployment in challenging aquatic environments. We will release the code on GitHub upon acceptance of the article.
Ziyi Wang 0006, Rong Dai, Yaqing Wang 0005, Fangxun Zhong, Yun-Hui Liu 0001
IEEE Trans. Ind. Informatics5
2023 Model-Free 3-D Shape Control of Deformable Objects Using Novel Features Based on Modal Analysis
abstract
Shape control of deformable objects is a challenging and important robotic problem. This article proposes a model-free controller using novel 3-D global deformation features based on modal analysis. Unlike most existing controllers using geometric features, our controller employs physically based deformation features designed by decoupling global deformation into low-frequency modes. Although modal analysis is widely adopted in computer vision and simulation, its usage in robotic deformation control is still an open topic. We develop a new model-free framework for the modal-based deformation control. Physical interpretation of the modes enables us to formulate an analytical deformation Jacobian matrix mapping the robot manipulation onto changes of the modal features. In the Jacobian matrix, unknown geometric and physical models of the object are treated as low-dimensional modal parameters, which can be used to linearly parameterize the closed-loop system. Thus, an adaptive controller with proven stability can be designed to deform the object while online estimating the modal parameters. Simulations and experiments are conducted using linear, planar, and volumetric objects under different settings. The results not only confirm the superior performance of our controller, but also demonstrate its advantages over the baseline method.
Bohan Yang 0005, Bo Lu 0001, Wei Chen 0068, Fangxun Zhong, Yun-Hui Liu 0001
IEEE Trans. Robotics4
2023 Robot-Camera Calibration in Tightly Constrained Environment Using Interactive Perception
abstract
Manipulation in tight environment is challenging but increasingly common in vision-guided robotic applications. The significantly reduced amount of available feedback (limited visual cues, field of view, robot motion space, etc.) hinders solving the hand-eye relationship accurately. In this article, we propose a new generic approach for online robot–camera calibration that could deal with the least feedback input available in tight environment: an arbitrarily restricted motion space and a single feature point with unknown position for the robot end-effector. We introduce the interactive perception to generate prescribed but tunable robot motions to reveal high-dimensional sensory feedback, which is not obtainable from static images. We then define the interactive feature plane (IFP), whose spatial property corresponds to the robot-actuating trajectories. A depth-free adaptive controller is proposed based on image feedback, where the converged orientation of IFP directly harvests the data for solving the hand–eye relationship. Our algorithm requires neither external calibration sensors/objects nor large-scale data acquisition process. Simulations demonstrate the validity of our method to accurately calibrate different types of robot under various system set-ups. In experiments, we show good results of our algorithm in terms of accuracy and consistency under tight motion space compared to existing approaches using external objects and/or optimization.
Fangxun Zhong, Bin Li 0082, Wei Chen 0068, Yun-Hui Liu 0001
IEEE Trans. Robotics1
2022 A Visual Navigation Perspective for Category-Level Object Pose Estimation
Fangxun Zhong, Rong Xiong, Yun-Hui Liu 0001, Yue Wang 0020, Yiyi Liao
ECCV (6)2
2022 Distilled Visual and Robot Kinematics Embeddings for Metric Depth Estimation in Monocular Scene Reconstruction
abstract
Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth information, which is difficult to transfer to the soft robotics-based surgical systems due to the use of monocular endoscopy. In this paper, we present a novel framework that combines robot kinematics and monocular endoscope images with deep unsupervised learning into a single network for metric depth estimation and then achieve 3D reconstruction of complex anatomy. Specifically, we first obtain the relative depth maps of surgical scenes by leveraging a brightness-aware monocular depth estimation method. Then, the corresponding endoscope poses are computed based on non-linear optimization of geo-metric and photometric reprojection residuals. Afterwards, we develop a Depth-driven Sliding Optimization (DDSO) algorithm to extract the scaling coefficient from kinematics and calculated poses offline. By coupling the metric scale and relative depth data, we form a robust ensemble that represents the metric and consistent depth. Next, we treat the ensemble as supervisory labels to train a metric depth estimation network for surgeries (i.e., MetricDepthS-Net) that distills the embeddings from the robot kinematics, endoscopic videos, and poses. With accurate metric depth estimation, we utilize a dense visual reconstruction method to recover the 3D structure of the whole surgical site. We have extensively evaluated the proposed framework on public SCARED and achieved comparable performance with stereo-based depth estimation methods. Our results demon-strate the feasibility of the proposed approach to recover the metric depth and 3D structure with monocular inputs.
Ruofeng Wei, Bin Li 0082, Hangjie Mo, Fangxun Zhong, Yonghao Long 0001, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001
IROS4
2022 AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Surgical Automation in Laparoscopic Hysterectomy
Ziyi Wang 0006, Bo Lu 0001, Yonghao Long 0001, Fangxun Zhong, Tak Hong Cheung, Qi Dou 0001, Yun-Hui Liu 0001
MICCAI (8)4
2016 Robust image-based computation of the 3D position of RCM instruments and its application to image-guided manipulation
abstract
In this paper, we address the 3D position control of RCM-constrained instruments with monocular cameras. To compute the instrument's position from a single 2D image, we develop an innovative gradient descent algorithm which rotates and translates a line segment (over the plane spanned by the imaged instrument and the optical centre) until it best aligns with the manipulated tool. In contrast with other approaches in the literature, our algorithm only requires to simultaneously observe two feature points; the proposed iterative algorithm is not based on the exact solution, therefore it can still work with noisy image measurements. We derive a kinematic controller that uses the proposed position estimator to guide the 3D motion of a robotic instrument with a monocular camera. We evaluate the performance of our approach with numerical simulations and experiments.
David Navarro-Alarcon, Zerui Wang, Hiu Man Yip, Yun-Hui Liu 0001, Fangxun Zhong, Tianxue Zhang, Jiadong Shi, Hesheng Wang 0001
ICRA5
2016 Adaptive 3D pose computation of suturing needle using constraints from static monocular image feedback
abstract
In this paper, we address the problem of the image-based 3D pose computation of a semi-circle suturing needle using monocular image feedback for laparoscopy. We propose a constrained two-degree-of-freedom (2-DOF) geometry-based modelling method to parametrise the needle's 6-DOF pose, including depth information. The modelling solely relies on the simultaneous observation of the needle's apparent tip and junction. No external markers are needed for extra constraints. An adaptive controller combining gradient descent and vector-flow method is introduced to iteratively guide the needle's initial guessing pose to its real pose by minimizing image-based position errors. Experiments have been conducted using both numerical simulations and simulated laparoscopic scenarios to evaluate the performance of the algorithm.
Fangxun Zhong, David Navarro-Alarcon, Zerui Wang, Yun-Hui Liu 0001, Tianxue Zhang, Hiu Man Yip, Hesheng Wang 0001
IROS1
2016 Automatic 3-D Manipulation of Soft Objects by Robotic Arms With an Adaptive Deformation Model
abstract
In this paper, we present a new feedback method to automatically servo-control the 3-D shape of soft objects with robotic manipulators. The soft object manipulation problem has recently received a great deal of attention from robotics researchers because of its potential applications in, e.g., food industry, home robots, medical robotics, and manufacturing. A major complication to automatically control the shape of an object is the estimation of its deformation properties, which determines how the manipulator's motion actively transforms into deformations. Note that these properties are rarely known beforehand, and its offline parametric identification is difficult and/or impractical to conduct in many applications. To cope with this issue, we developed a new algorithm that computes in real time the unknown deformation parameters of a soft object; this algorithm provides a valuable adaptive behavior to the deformation controller, something we cannot achieve with traditional fixed-model approaches. In contrast with most controllers in the literature, our new method can explicitly servo-control 3-D deformations (and not just 2-D image projections) in an entirely model-free way. To validate the proposed adaptive controller, we present a detailed experimental study with robotic manipulators.
David Navarro-Alarcon, Hiu Man Yip, Zerui Wang, Yun-Hui Liu 0001, Fangxun Zhong, Tianxue Zhang, Peng Li 0019
IEEE Trans. Robotics5