Haichong K. Zhang

dblp:150/3109 · also Haichong (Kai) Zhang · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-1314-8456ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2025 Deep Loss Convexification for Learning Iterative Models
abstract
Iterative methods such as iterative closest point (ICP) for point cloud registration often suffer from bad local optimality (e.g. saddle points), due to the nature of nonconvex optimization. To address this fundamental challenge, in this paper we propose learning to form the loss landscape of a deep iterative method w.r.t. predictions at test time into a convex- like shape locally around each ground truth given data, namely Deep Loss Convexification (DLC), thanks to the overparametrization in neural networks. To this end, we formulate our learning objective based on adversarial training by manipulating the ground-truth predictions, rather than input data. In particular, we propose using star-convexity, a family of structured nonconvex functions that are unimodal on all lines that pass through a global minimizer, as our geometric constraint for reshaping loss landscapes, leading to (1) extra novel hinge losses appended to the original loss and (2) near-optimal predictions. We demonstrate the state-of-the-art performance using DLC with existing network architectures for the tasks of training recurrent neural networks (RNNs), 3D point cloud registration, and multimodel image alignment.
Yuping Shao, Yiqing Zhang 0003, Fangzhou Lin, Haichong K. Zhang, Elke A. Rundensteiner
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Guiding the Last Centimeter: Novel Anatomy-Aware Probe Servoing for Standardized Imaging Plane Navigation in Robotic Lung Ultrasound
abstract
Navigating the ultrasound (US) probe to the standardized imaging plane (SIP) for image acquisition is a critical but operator-dependent task in conventional freehand diagnostic US. Robotic US systems (RUSS) offer the potential to enhance imaging consistency by leveraging real-time US image feedback to optimize the probe pose, thereby reducing reliance on operator expertise. However, determining the proper approach to extracting generalizable features from the US images for probe pose adjustment remains challenging. In this work, we propose a SIP navigation framework for RUSS, exemplified in the context of robotic lung ultrasound (LUS). This framework facilitates automatic probe adjustment when in proximity to the SIP. This is achieved by explicitly extracting multiple anatomical features presented in real-time LUS images and performing non-patient-specific template matching to generate probe motion towards the SIP using image-based visual servoing (IBVS). The framework is further integrated with the active-sensing end-effector (A-SEE), a customized robot end-effector that leverages patient external body geometry to maintain optimal probe alignment with the contact surface, thus preserving US signal quality throughout the navigation. The proposed approach ensures procedural interpretability and inter-patient adaptability. Validation is conducted through anatomy-mimicking phantom and in-vivo evaluations involving five human subjects. The results show the framework’s high navigating precision with the probe correctly located at the SIP for all cases, exhibiting positioning error of under 2 mm in translation and under 2 degrees in rotation. These results demonstrate the navigation process’s capability to accommodate anatomical variations among patients. Note to Practitioners—Compared with traditional freehand ultrasound (US) imaging, robotic ultrasound systems (RUSS) have the potential to largely standardize the US diagnosis outcome caused by varying operator expertise if an inter-patient consistent, automatic standardized imaging plane (SIP) navigation process is available. This paper presents a SIP navigation framework for lung US (LUS) examination, which recognizes anatomical landmarks from the US images and fine-tunes the pose of the US probe so that the landmarks are positioned in accordance with a non-patient-specific template image. The special end-effector, active-sensing end-effector (A-SEE), maintains the probe at an optimal orientation with respect to the body, allowing consistent-quality US images to be acquired throughout the navigation. Unlike previous works, our approach can navigate to complicated SIP containing multiple anatomies with interpretable robot arm motion. We verified our framework’s ability to navigate the probe to the SIP with millimeter-level accuracy under phantom and human experiment settings. While preliminary results demonstrate the framework’s efficacy in guiding the robotic LUS procedure, the performance of the system on other examinations (e.g., liver and thyroid US) involving soft tissues requires further validation. In the future, the framework can be applied in various US examinations by implementing specific anatomical feature detection modules.
Xihan Ma, Mingjie Zeng, Jeffrey C. Hill, Beatrice Hoffmann, Haichong K. Zhang
IEEE Trans Autom. Sci. Eng.6
2024 Forearm Ultrasound Based Gesture Recognition on Edge
abstract
Ultrasound imaging of the forearm has demon-strated significant potential for accurate hand gesture classification. Despite this progress, there has been limited focus on developing a stand -alone end-to-end gesture recognition system which makes it mobile, real-time and more user friendly. To bridge this gap, this paper explores the deployment of deep neural networks for forearm ultrasound-based hand gesture recognition on edge devices. Utilizing quantization techniques, we achieve substantial reductions in model size while maintaining high accuracy and low latency. Our best model, with Float16 quantization, achieves a test accuracy of 92% and an inference time of 0.31 seconds on a Raspberry Pi. These results demonstrate the feasibility of efficient, real-time gesture recognition on resource-limited edge devices, paving the way for wearable ultrasound-based systems.
Keshav Bimbraw, Haichong K. Zhang, Bashima Islam
BSN2
2024 Loss Distillation via Gradient Matching for Point Cloud Completion with Weighted Chamfer Distance
abstract
3D point clouds enhanced the robot’s ability to perceive the geometrical information of the environments, making it possible for many downstream tasks such as grasp pose detection and scene understanding. The performance of these tasks, though, heavily relies on the quality of data input, as incomplete can lead to poor results and failure cases. Recent training loss functions designed for deep learning-based point cloud completion, such as Chamfer distance (CD) and its variants (e.g. HyperCD [1]), imply a good gradient weighting scheme can significantly boost performance. However, these CD-based loss functions usually require data-related parameter tuning, which can be time-consuming for data-extensive tasks. To address this issue, we aim to find a family of weighted training losses (weighted CD) that requires no parameter tuning. To this end, we propose a search scheme, Loss Distillation via Gradient Matching, to find good candidate loss functions by mimicking the learning behavior in backpropagation between HyperCD and weighted CD. Once this is done, we propose a novel bilevel optimization formula to train the backbone network based on the weighted CD loss. We observe that: (1) with proper weighted functions, the weighted CD can always achieve similar performance to HyperCD, and (2) the Landau weighted CD, namely Landau CD, can outperform HyperCD for point cloud completion and lead to new state-of-the-art results on several benchmark datasets. Our demo code is available at https://github.com/Zhang-VISLab/IROS2024-LossDistillationWeightedCD.
Fangzhou Lin, Haoying Zhou, Songlin Hou, Kazunori D. Yamada, Gregory S. Fischer, Haichong K. Zhang
IROS8
2024 Thermal Ablation Therapy Control with Tissue Necrosis-driven Temperature Feedback Enabled by Neural State Space Model with Extended Kalman Filter
abstract
Thermal ablation therapy is a major minimally invasive treatment. One of the challenges is that the targeted region and therapeutic progression are often invisible to clinicians, requiring feedback provided in numerical information or imaging. Several emerging imaging modalities offer visualization of the ablation-induced necrosis formation; however, relying solely on necrosis monitoring can result in tissue overheating and endangering patients. Some of the necrosis monitoring modalities are known for their capabilities in temperature sensing, but the principles on which they are based have several limitations, such as sensitivity to the tissue motion and their environment. In this study, we propose a necrosis progression-based temperature estimation technique as an added safety feature for avoiding overheating. This model-based method does not require additional sensing hardware. It is designed to work as an independent estimator or a complimentary estimation component with other thermometers for improved robustness. For this objective, the Neural State Space model is used to approximate the ablation therapy, whose theoretical models involve nonlinear partial differential equations. Then, the Extended Kalman Filter is designed based on the model. The simulation study shows the estimation module robustly estimates the tissue temperature under several types of noise. The maximum estimation error observed before terminating ablation was around 1 °C, and the desired safety feature was successfully demonstrated. The estimator is expected to be used in a variety of necrosis monitoring modalities to guarantee more precise and safer treatment. More ambitiously, the architecture with the Neural State Space model and Extended Kalman Filter is generalizable to other medical/biological procedures involving nonlinear and patient/environment-specific physics and even to procedures having no reliable theoretical models.
Ryo Murakami, Satoshi Mori, Haichong K. Zhang
IROS3
2023 Leveraging Ultrasound Sensing for Virtual Object Manipulation in Immersive Environments
abstract
Hand gesture recognition is a fundamental component of intuitive and immersive user interfaces in virtual reality (VR) applications. This paper presents a data-driven approach utilizing ultrasound data and deep learning techniques for hand gesture recognition in VR interfaces. The proposed methodology involves acquiring data from a subject, training a model using the acquired data, and evaluating the model’s performance on both the training data and during real-time inference. The evaluation metrics primarily focus on accuracy percentage: measuring the classifier’s performance in correctly classifying hand gestures. 4 hand gestures were primarily considered for the study and demonstration. For offline evaluation with a 20% test-train split, an accuracy percentage of 91% was observed. For online evaluation, an accuracy percentage of 92% was achieved. Results on the classification of 7 hand gestures were also analyzed for both online and online evaluation due to the promising results from the 4 gesture classification. The latency of the pipeline, from ultrasound data acquisition using screenshots to sending commands for VR object manipulation, was measured to be 59.48 milliseconds. The results demonstrate the effectiveness of the approach in accurately recognizing hand gestures, both during training and in real-time inference. We supplement our results with a video of the forearm ultrasound data being used to control a custom-designed VR game in a low-latency fashion. This research provides valuable insights into the performance and applicability of ultrasound-based hand gesture recognition techniques in VR interfaces. By employing deep learning and leveraging real-time data acquisition, this approach paves the way for intuitive and immersive interactions in various VR applications. The study contributes to the field of body sensor networks, highlighting the potential of forearm ultrasound based data-driven techniques for enhancing user interaction and immersion in VR environments.
Keshav Bimbraw, Jack Rothenberg, Haichong K. Zhang
BSN3
2022 Prediction of Metacarpophalangeal Joint Angles and Classification of Hand Configurations Based on Ultrasound Imaging of the Forearm
abstract
With the advancement in computing and robotics, it is necessary to develop fluent and intuitive methods for inter-acting with digital systems, augmented/virtual reality (AR/VR) interfaces, and physical robotic systems. Hand movement recognition is widely used to enable such interaction. Hand configuration classification and metacarpophalangeal (MCP) joint angle detection are important for a comprehensive reconstruction of hand motion. Surface electromyography (sEMG) and other technologies have been used for the detection of hand motions. Ultrasound images of the forearm offer a way to visualize the internal physiology of the hand from a musculoskeletal perspective. Recent works have shown that these images can be classified using machine learning to predict various hand configurations. In this paper, we propose a Convolutional Neu-ral Network (CNN) based deep learning pipeline for predicting the MCP joint angles. We supplement our results by using a Support Vector Classifier (SVC) to classify the ultrasound information into several predefined hand configurations based on activities of daily living (ADL). Ultrasound data from the forearm were obtained from six subjects who were instructed to move their hands according to predefined hand configurations relevant to ADLs. Motion capture data was acquired as the ground truth for hand movements at three speeds (0.5 Hz, 1 Hz, and 2 Hz) for the index, middle, ring, and pinky fingers. We demonstrated the perfect prediction of hand configurations through SVC classification and a correspondence between the predicted MCP joint angles and the actual MCP joint angles for the fingers, with an average root mean square error of 7.35 degrees. A low latency (6.25 – 9.10 Hz) pipeline was implemented for the prediction of both MCP joint angles and hand configuration estimation aimed for real-time implementation.
Keshav Bimbraw, Christopher J. Nycz, Matthew J. Schueler, Haichong K. Zhang
ICRA5
2021 Autonomous Scanning Target Localization for Robotic Lung Ultrasound Imaging
abstract
Under the ceaseless global COVID-19 pandemic, lung ultrasound (LUS) is the emerging way for effective diagnosis and severeness evaluation of respiratory diseases. However, close physical contact is unavoidable in conventional clinical ultrasound, increasing the infection risk for health-care workers. Hence, a scanning approach involving minimal physical contact between an operator and a patient is vital to maximize the safety of clinical ultrasound procedures. A robotic ultrasound platform can satisfy this need by remotely manipulating the ultrasound probe with a robotic arm. This paper proposes a robotic LUS system that incorporates the automatic identification and execution of the ultrasound probe placement pose without manual input. An RGB-D camera is utilized to recognize the scanning targets on the patient through a learning-based human pose estimation algorithm and solve for the landing pose to attach the probe vertically to the tissue surface; A position/force controller is designed to handle intraoperative probe pose adjustment for maintaining the contact force. We evaluated the scanning area localization accuracy, motion execution accuracy, and ultrasound image acquisition capability using an upper torso mannequin and a realistic lung ultrasound phantom with healthy and COVID-19-infected lung anatomy. Results demonstrated the overall scanning target localization accuracy of 19.67 ± 4.92 mm and the probe landing pose estimation accuracy of 6.92 ± 2.75 mm in translation, 10.35 ± 2.97 deg in rotation. The contact force-controlled robotic scanning allowed the successful ultrasound image collection, capturing pathological landmarks.
Xihan Ma, Haichong K. Zhang
IROS3
2021 SBO-RNN: Reformulating Recurrent Neural Networks via Stochastic Bilevel Optimization
abstract
In this paper we consider the training stability of recurrent neural networks (RNNs) and propose a family of RNNs, namely SBO-RNN, that can be formulated using stochastic bilevel optimization (SBO). With the help of stochastic gradient descent (SGD), we manage to convert the SBO problem into an RNN where the feedforward and backpropagation solve the lower and upper-level optimization for learning hidden states and their hyperparameters, respectively. We prove that under mild conditions there is no vanishing or exploding gradient in training SBO-RNN. Empirically we demonstrate our approach with superior performance on several benchmark datasets, with fewer parameters, less training data, and much faster convergence. Code is available at https://zhang-vislab.github.io.
Yun Yue, Guojun Wu, Haichong K. Zhang
NeurIPS5
2018 Towards a Fast and Safe LED-Based Photoacoustic Imaging Using Deep Convolutional Neural Network
Emran Mohammad Abu Anas, Haichong K. Zhang, Jin Kang, Emad Boctor
MICCAI (4)2
2016 Photoacoustic Imaging Paradigm Shift: Towards Using Vendor-Independent Ultrasound Scanners
Haichong K. Zhang, Behnoosh Tavakoli, Emad Boctor
MICCAI (1)1
2014 Active Echo: A New Paradigm for Ultrasound Calibration
Alexis Cheng, Haichong K. Zhang, Hyun Jae Kang 0002, Ralph Etienne-Cummings, Emad Boctor
MICCAI (2)3