Dianye Huang

dblp:221/1712 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-7719-6505ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improving Probe Localization for Freehand 3D Ultrasound Using Lightweight Cameras
abstract
Ultrasound (US) probe localization relative to the examined subject is essential for freehand 3D US imaging, which offers significant clinical value due to its affordability and unrestricted field of view. However, existing methods often rely on expensive tracking systems or bulky probes, while recent US image-based deep learning methods suffer from accumulated errors during probe maneuvering. To address these challenges, this study proposes a versatile, cost-effective probe pose localization method for freehand 3D US imaging, utilizing two lightweight cameras. To eliminate accumulated errors during US scans, we introduce PoseNet, which directly predicts the probe's 6 D pose relative to a preset world coordinate system based on camera observations. We first jointly train pose and camera image encoders based on pairs of 6 D pose and camera observations densely sampled in simulation. This will encourage each pair of probe pose and its corresponding camera observation to share the same representation in latent space. To ensure the two encoders handle unseen images and poses effectively, we incorporate a triplet loss that enforces smaller differences in latent features between nearby poses compared to distant ones. Then, the pose decoder uses the latent representation of the camera images to predict the probe's 6 D pose. To bridge the sim-to-real gap, in the real world, we use the trained image encoder and pose decoder for initial predictions, followed by an additional MLP layer to refine the estimated pose, improving accuracy. The results obtained from an arm phantom demonstrate the effectiveness of the proposed method, which notably surpasses state-of-the-art techniques, achieving average positional and rotational errors of 2.03 mm and 0.37°, respectively. Code:https://github.com/dianyeHuang/FreehandUS_Pose_Estimation
Dianye Huang, Nassir Navab, Zhongliang Jiang
ICRA1
2025 Vibration-Based Energy Metric for Restoring Needle Alignment in Autonomous Robotic Ultrasound
abstract
Precise needle alignment is essential for percutaneous needle insertion in robotic ultrasound-guided procedures. However, inherent challenges such as speckle noise, needle-like artifacts, and low image resolution complicate robust needle detection, which is essential for alignment in ultrasound images. These issues become particularly problematic when visibility is reduced or lost, diminishing the effectiveness of visual-based needle alignment methods. In this paper, we propose a method to restore effectively when the ultrasound imaging plane and the needle insertion plane are misaligned. Unlike many existing approaches that rely heavily on needle visibility in ultrasound images, our method uses a more robust feature by periodically vibrating the needle using a mechanical system. Specifically, we propose a new vibration-based energy metric that remains effective even when the needle is fully out of plane. Using this metric, we develop an elegant control strategy to reposition the ultrasound probe in response to misalignments between the imaging plane and the needle insertion plane in both translation and rotation. Experiments conducted on ex-vivo porcine tissue samples using a dual-arm robotic ultrasound-guided needle insertion system demonstrate the effectiveness of the proposed approach. The experimental results show the translational error of 0.41±0.27 mm and the rotational error of 0.51±0.19 degrees.
Chenyang Li 0004, Dianye Huang, Zhongliang Jiang, Stefanie Speidel, Xiangyu Chu, K. W. Samuel Au
IROS4
2025 Tactile-Guided Robotic Ultrasound: Mapping Preplanned Scan Paths for Intercostal Imaging
abstract
Medical ultrasound (US) imaging is widely used in clinical examinations due to its portability, real-time capability, and radiation-free nature. To address inter- and intra-operator variability, robotic ultrasound systems have gained increasing attention. However, their application in challenging intercostal imaging remains limited due to the lack of an effective scan path generation method within the constrained acoustic window. To overcome this challenge, we explore the potential of tactile cues for characterizing subcutaneous rib structures as an alternative signal for ultrasound segmentation-free bone surface point cloud extraction. Compared to 2D US images, 1D tactile-related signals offer higher processing efficiency and are less susceptible to acoustic noise and artifacts. By leveraging robotic tracking data, a sparse tactile point cloud is generated through a few scans along the rib, mimicking human palpation. To robustly map the scanning trajectory into the intercostal space, the sparse tactile bone location point cloud is first interpolated to form a denser representation. This refined point cloud is then registered to an image-based dense bone surface point cloud, enabling accurate scan path mapping for individual patients. Additionally, to ensure full coverage of the object of interest, we introduce an automated tilt angle adjustment method to visualize structures beneath the bone. To validate the proposed method, we conducted comprehensive experiments on four distinct phantoms. The final scanning waypoint mapping achieved Mean Nearest Neighbor Distance (MNND) and Hausdorff distance (HD) errors of 3.41 mm and 3.65 mm, respectively, while the reconstructed object beneath the bone had errors of 0.69 mm and 2.2 mm compared to the CT ground truth.
Dianye Huang, Nassir Navab, Zhongliang Jiang
IROS2
2025 Semantic Scene Graph for Ultrasound Image Explanation and Scanning Guidance
Dianye Huang, Nassir Navab, Zhongliang Jiang
MICCAI (9)2
2025 Robot-Assisted Deep Venous Thrombosis Ultrasound Examination Using Virtual Fixture
abstract
Deep Venous Thrombosis (DVT) is a common vascular disease with blood clots inside deep veins, which may block blood flow or even cause a life-threatening pulmonary embolism. A typical exam for DVT using ultrasound (US) imaging is by pressing the target vein until its lumen is fully compressed. However, the compression exam is highly operator-dependent. To alleviate intra-and inter-variations, we present a robotic US system with a novel hybrid force motion control scheme ensuring position and force tracking accuracy, and soft landing of the probe onto the target surface. In addition, a path-based virtual fixture is proposed to realize easy human-robot interaction for repeat compression operation at the lesion location. To ensure the biometric measurements obtained in different examinations are comparable, the 6D scanning path is determined in a coarse-to-fine manner using both an external RGBD camera and US images. The RGBD camera is first used to extract a rough scanning path on the object. Then, the segmented vascular lumen from US images are used to optimize the scanning path to ensure the visibility of the target object. To generate a continuous scan path for developing virtual fixtures, an arc-length based path fitting model considering both position and orientation is proposed. Finally, the whole system is evaluated on a human-like arm phantom with an uneven surface. The code (https://github.com/dianyeHuang/RobDVTUS) and intuitive demonstration video (https://www.youtube.com/ watch?v=3xFyqU1rV8c) can be publicly accessed.Note to Practitioners—Robotic ultrasound (US) systems have attracted attention for various applications in the past decades. However, the existing studies are not mature and intelligent enough for some challenging applications, such as DVT exam, which requires rich contact interaction between patients and clinicians. To tackle with this challenge, this study presents a novel human-centric robotic DVT exam program using the technique of virtual fixture. The coarse-to-fine path planning module ensures the repeatability of US acquisitions carried out at different times. During DVT exam, the proposed continuous 6D path virtual fixture can guide clinicians to freely move the probe along the scan path while limiting the probe motion in other directions. In order to perform the compress-release exam, a decoupled position/force controller is developed to precisely generate the contact force conveyed by clinicians and to restrict the probe motion along the probe centerline. We believe such a robot-assisted system is a promising solution to take both advantages of robots about the accuracy and repeatability and human operators about the advanced physiological knowledge.
Dianye Huang, Chenguang Yang 0001, Mingchuan Zhou, Angelos Karlas, Nassir Navab, Zhongliang Jiang
IEEE Trans Autom. Sci. Eng.1
2025 VibNet: Vibration-Boosted Needle Detection in Ultrasound Images
abstract
Precise percutaneous needle detection is crucial for ultrasound (US)-guided interventions. However, inherent limitations such as speckles, needle-like artifacts, and low resolution make it challenging to robustly detect needles, especially when their visibility is reduced or imperceptible. To address this challenge, we propose VibNet, a learning-based framework designed to enhance the robustness and accuracy of needle detection in US images by leveraging periodic vibration applied externally to the needle shafts. VibNet integrates neural Short-Time Fourier Transform and Hough Transform modules to achieve successive sub-goals, including motion feature extraction in the spatiotemporal space, frequency feature aggregation, and needle detection in the Hough space. Due to the periodic subtle vibration, the features are more robust in the frequency domain than in the image intensity domain, making VibNet more effective than traditional intensity-based methods. To demonstrate the effectiveness of VibNet, we conducted experiments on distinct ex vivo porcine and bovine tissue samples. The results obtained on porcine samples demonstrate that VibNet effectively detects needles even when their visibility is severely reduced, with a tip error of ${1}.{61}\pm {1}.{56}~\textit {mm}$ compared to ${8}.{15}\pm {9}.{98}~\textit {mm}$ for UNet and ${6}.{63}\pm {7}.{58}~\textit {mm}$ for WNet, and a needle direction error of ${1}.{64}\pm {1}.{86}^{\circ }$ compared to ${9}.{29}~\pm ~{15}.{30}^{\circ }$ for UNet and ${8}.{54}~\pm ~{17}.{92}^{\circ }$ for WNet. Code: https://github.com/marslicy/VibNet.
Dianye Huang, Chenyang Li 0004, Angelos Karlas, Xiangyu Chu, K. W. Samuel Au, Nassir Navab, Zhongliang Jiang
IEEE Trans. Medical Imaging1
2025 Improving Robustness to Out-of-Distribution States in Imitation Learning via Deep Koopman-Boosted Diffusion Policy
abstract
Integrating generative models with action chunking has shown significant promise in imitation learning for robotic manipulation. However, the existing diffusion-based paradigm often struggles to capture strong temporal dependencies across multiple steps, particularly when incorporating proprioceptive input. This limitation can lead to task failures, where the policy overfits to proprioceptive cues at the expense of capturing the visually derived features of the task. To overcome this challenge, we propose the Deep Koopman-boosted Dual-branch Diffusion Policy (D3P) algorithm. D3P introduces a dual-branch architecture to decouple the roles of different sensory modality combinations. The visual branch encodes the visual observations to indicate task progression, while the fused branch integrates both visual and proprioceptive inputs for precise manipulation. Within this architecture, when the robot fails to accomplish intermediate goals, such as grasping a drawer handle, the policy can dynamically switch to execute action chunks generated by the visual branch, allowing recovery to previously observed states and facilitating retrial of the task. To further enhance visual representation learning, we incorporate a Deep Koopman Operator module that captures structured temporal dynamics from visual inputs. During inference, we use the test-time loss of the generative model as a confidence signal to guide the aggregation of the temporally overlapping predicted action chunks, thereby enhancing the reliability of policy execution. In simulation experiments across six RLBench tabletop tasks, D3P outperforms the state-of-the-art diffusion policy by an average of 14.6%. On three real-world robotic manipulation tasks, it achieves a 15.0% improvement. Code:https://github.com/dianyeHuang/D3P.
Dianye Huang, Nassir Navab, Zhongliang Jiang
IEEE Trans. Robotics1
2024 SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
abstract
Object rearrangement is pivotal in robotic-environment interactions, representing a significant capability in embodied AI. In this paper, we present SG-Bot, a novel rearrangement framework that utilizes a coarse-to-fine scheme with a scene graph as the scene representation. Unlike previous methods that rely on either known goal priors or zero-shot large models, SG-Bot exemplifies lightweight, real-time, and user-controllable characteristics, seamlessly blending the consideration of commonsense knowledge with automatic generation capabilities. SG-Bot employs a three-fold procedure– observation, imagination, and execution–to adeptly address the task. Initially, objects are discerned and extracted from a cluttered scene during the observation. These objects are first coarsely organized and depicted within a scene graph, guided by either commonsense or user-defined criteria. Then, this scene graph subsequently informs a generative model, which forms a fine-grained goal scene considering the shape information from the initial scene and object semantics. Finally, for execution, the initial and envisioned goal scenes are matched to formulate robotic action policies. Experimental results demonstrate that SG-Bot outperforms competitors by a large margin.
Guangyao Zhai, Xiaoni Cai, Dianye Huang, Yan Di, Fabian Manhardt, Federico Tombari, Nassir Navab, Benjamin Busam
ICRA3
2024 Safe Multiagent Learning With Soft Constrained Policy Optimization in Real Robot Control
abstract
Due to a lack of safety considerations, a wide range of multiagent reinforcement learning (MARL) applications are limited in real-world environments. Thus, ensuring MARL safety is essential and urgent in the domain. However, merely a few studies consider the safe MARL problem, and the investigation of real-world applications using safe MARL algorithms still needs to be improved. To fill this gap, we provide a framework with soft constrained policy optimization, in which we develop practical algorithms to address the problem in a cooperative game setting. First, the problem formulation of safe MARL is introduced. Second, the safe policy optimization of safe MARL algorithms based on soft constrained optimization is analyzed, and we further propose a safe learning framework for safe MARL. The framework can be plugged into MARL algorithms without manually fine-tuning safety bounds. Third, we investigate the sim-to-real problems, and conduct simulation and real-world experiments to evaluate the effectiveness of our algorithms. Finally, the comprehensive experimental results indicate that our method has significant benefits regarding the balance between reward and safety performance and outperforms several strong baselines.
Shangding Gu, Dianye Huang, Muning Wen, Guang Chen 0001, Alois C. Knoll
IEEE Trans. Ind. Informatics2
2023 MonoGraspNet: 6-DoF Grasping with a Single RGB Image
abstract
6-DoF robotic grasping is a long-lasting but un-solved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but performing unsatisfactorily on photometrically challenging objects, e.g., objects in transparent or reflective materials. The bottleneck lies in that the surface of these objects can not reflect accurate depth due to the absorption or refraction of light. In this paper, in contrast to exploiting the inaccurate depth data, we propose the first RGB-only 6-DoF grasping pipeline called MonoGraspNet that utilizes stable 2D features to simultaneously handle arbitrary object grasping and overcome the problems induced by photometrically challenging objects. MonoGraspNet leverages a keypoint heatmap and a normal map to recover the 6-DoF grasping poses represented by our novel representation parameterized with 2D keypoints with corresponding depth, grasping direction, grasping width, and angle. Extensive experiments in real scenes demonstrate that our method can achieve competitive results in grasping common objects and surpass the depth-based competitor by a large margin in grasping photometrically challenging objects. To further stimulate robotic manipulation research, we annotate and open-source a multi-view grasping dataset in the real world containing 44 sequence collections of mixed photometric complexity with nearly 20M accurate grasping labels.
Guangyao Zhai, Dianye Huang, Yan Di, Fabian Manhardt, Federico Tombari, Nassir Navab, Benjamin Busam
ICRA2
2023 Motion Magnification in Robotic Sonography: Enabling Pulsation-Aware Artery Segmentation
abstract
Ultrasound (US) imaging is widely used for diagnosing and monitoring arterial diseases, mainly due to the advantages of being non-invasive, radiation-free, and real-time. In order to provide additional information to assist clinicians in diagnosis, the tubular structures are often segmented from US images. To improve the artery segmentation accuracy and stability during scans, this work presents a novel pulsation-assisted segmentation neural network (PAS-NN) by explicitly taking advantage of the cardiac-induced motions. Motion magnification techniques are employed to amplify the subtle motion within the frequency band of interest to extract the pulsation signals from sequential US images. The extracted real-time pulsation information can help to locate the arteries on cross-section US images; therefore, we explicitly integrated the pulsation into the proposed PAS-NN as attention guidance. Notably, a robotic arm is necessary to provide stable movement during US imaging since magnifying the target motions from the US images captured along a scan path is not manually feasible due to the hand tremor. To validate the proposed robotic US system for imaging arteries, experiments are carried out on volunteers' carotid and radial arteries. The results demonstrated that the PAS-NN could achieve comparable results as state-of-the-art on carotid and can effectively improve the segmentation performance for small vessels (radial artery). The code11Code: https://qithub.com/dianveHuanq/RobPMEPASNN and demonstration video22Video: https://youtu.belc9AM042_lUQ can be publicly accessed.
Dianye Huang, Yuan Bi, Nassir Navab, Zhongliang Jiang
IROS1
2021 Composite Learning Enhanced Neural Control for Robot Manipulator With Output Error Constraints
abstract
This article presents a control scheme for robot manipulators with the consideration of output error constraints, unknown dynamics, and bounded disturbances. A modified virtual input variable in the second stage design of the dynamic surface control scheme is proposed, which can enhance the robustness of the controller. Bounded disturbances due to the situations that the base is not well fixed if the robot manipulator is mounted at a mobile platform are considered and suppressed. Besides, the detailed implementation process of the composite learning laws adopted for enhancing the radial basis function neural network is presented. Lyapunov stability analysis verifies that the proposed control scheme ensures the trajectory tracking errors stay within predefined boundaries and parameter estimate errors converge without a stringent condition termed persistent excitation. Experimental results show the superiority of the proposed controller regarding parameter estimation and tracking capabilities.
Dianye Huang, Chenguang Yang 0001, Yongping Pan 0001, Long Cheng 0001
IEEE Trans. Ind. Informatics1
2021 Neural Control of Robot Manipulators With Trajectory Tracking Constraints and Input Saturation
abstract
This article presents a control scheme for the robot manipulator's trajectory tracking task considering output error constraints and control input saturation. We provide an alternative way to remove the feasibility condition that most BLF-based controllers should meet and design a control scheme on the premise that constraint violation possibly happens due to the control input saturation. A bounded barrier Lyapunov function is proposed and adopted to handle the output error constraints. Besides, to suppress the input saturation effect, an auxiliary system is designed and emerged into the control scheme. Moreover, a simplified RBFNN structure is adopted to approximate the lumped uncertainties. Simulation and experimental results demonstrate the effectiveness of the proposed control scheme.
Chenguang Yang 0001, Dianye Huang, Wei He 0001, Long Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.2