Mingzhu Zhu

dblp:34/7496 · DBLP profile ↗
← Back
23ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 IFPs-Free Marker Design and Threshold-Free Detection for Deformable Surface Tracking
abstract
Visual localization is essential in many vision-driven interaction domains, including automation, AR (augmented reality), and surgical navigation. This paper presents a visual marker specifically designed for deformable surface tracking. It has three advantages: 1) IFPs (inner false positives) detection is avoided through grayscale integration design, leading to more robust results. 2) Without relying on thresholds, our detection algorithm enhances the method’s reliability in deformable surface with complex shading. 3) We use a position-sensing marker design with higher information density than the self-identifing marker to ensure the supply of features. In the experiments, our marker achieved superior localization accuracy and zero IFPs. Additionally, we present two compelling case studies showcasing the marker’s practical applications in augmented reality and surgical instrument tracking. Our work offers a significant advancement in visual localization, especially in challenging scenarios involving deformable surfaces, providing valuable solutions for researchers and developers across various application domains.
Xianrui Gu, Mingzhu Zhu, Bingwei He
IEEE Trans. Circuits Syst. Video Technol.2
2025 Adjusting Distributed Cameras for Robust Moving Object Pose Estimation
abstract
Robust moving object pose estimation is crucial in fine manipulation tasks, such as surgical instrument tracking. This paper presents a distributed-camera system with robotic adjustments to maintain consistent tracking of moving objects, thus avoiding tracking failures. An integrated framework for camera adjustment and pose estimation is developed for this distributed-camera system. In each detection cycle, the camera exhibiting the largest deviation with the object is adjusted by a visual servoing technique. After adjustment, the camera extrinsics are re-calibrated in the following detection cycles. For the unadjusted cameras, an online extrinsic optimization method based on multi-frame detection results is proposed to refine the camera extrinsics. Based on the refined camera extrinsics and detection results from multiple cameras, the pose of moving objects relative to the principal camera can be robustly estimated. We test the performance of this system in both simulation environments and real-world scenarios. The results indicate that our system achieves higher pose estimation accuracy and exhibits strong resistance to limited field-of-view (FoV) compared to conventional equivalent fixed multi-camera systems.
Yaoqing Hu, Shaoan Wang, Xingyu Chen 0002, Mingzhu Zhu, Zhanhua Xin, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.5
2025 A Lightweight Integrated Positioning System With Occlusion-Aware Region-Based Pose Tracking for Oral and Maxillofacial Surgery
abstract
The development of an accurate and robust positioning system for oral and maxillofacial surgery (OMS) is a challenging task, primarily due to the oral space limitations and line-of-sight occlusions. This paper presents a novel lightweight integrated positioning system for OMS, which can provide practical guidance utilizing only a micro camera installed on the end of the surgical instrument. An efficient region-based pose tracking method for texture-less teeth is proposed, which can use search lines around the object contour and simple local region partitioning strategy to improve pose accuracy. Besides, to deal with the possible partial occlusions of target during surgery, an occlusion-aware weight function is presented and utilized seamlessly in the pose optimization pipeline. This function calculates the pixel-wise occlusion probability using object contour and distance constraint, helping to improve the tracking robustness. Pivot calibration evaluation reveals that the tracking accuracy of the proposed camera-based handpiece is higher than the marker-based handpieces. Comparative experiments demonstrate that proposed pose tracking method has higher accuracy than existing state-of-the-art methods and ablation study confirms the effectiveness of the occlusion handling strategy. The overall positioning experiment indicates that the proposed system has satisfactory static poses stability and positioning accuracy. Furthermore, the main advantage of our system is that it is lighter and more integrated than other systems, which can reduce the system complexity, decrease the risk of line-of-sight occlusion, and lower the surgery cost.Note to Practitioners—This paper is motivated by the problem of restricted oral space constraints and partial occlusions during positioning for OMS. Compared with traditional OMS navigation systems, the designed system is more lightweight and more integrated without other external cameras and additional fiducial markers. Our system can provide practical guidance utilizing only a micro camera installed on the end of the surgical instrument. In addition, an efficient region-based pose tracking method for texture-less teeth is proposed to increase pose accuracy. Since the target can partially be occluded during the procedure, we present a novel occlusion-aware strategy to improve the tracking performance of partial occlusions. Our proposed system achieves a decent balance between positioning accuracy and hardware cost, and can easily be integrated into various dental surgical tools, thus it has tremendous potential for commercialization.
Yaoqing Hu, Shaoan Wang, Mingzhu Zhu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.4
2025 Accurate and Automatic Dental Crown Components Segmentation With Multi-Scale Attention Based U-Net and Hybrid Level Set Models
abstract
This paper presents a two-step method to automatically and accurately segment the dental crown components from CT images. Firstly, a multi-scale attention based U-Net model is proposed for pulp segmentation, which is embedded with global and local attention modules. The constructed attention modules can automatically aggregate pixel-wise contextual information and focus on catching the real dental pulp region. Secondly, two efficient level set models are proposed: one is the shape constraint-based level set model for enamel and dentin segmentation, the other is the region mutual exclusion-based level set model for neighboring teeth segmentation. The proposed shape constraint term can better handle topology changes of teeth and the region mutual exclusion term can more effectively avoid intersecting segmentation. Besides, a starting slice initialization method is introduced to achieve automatic segmentation, and an accurate contour propagation strategy is developed for slice-by-slice segmentation. We set up a series of comparative experiments for evaluation. Experimental results verify that the proposed method obtains promising performance for each crown component segmentation, and outperforms state-of-the-art tooth segmentation methods in terms of accuracy. This suggests that the proposed method can be used to accurately segment the crown components for precise tooth preparation treatment.Note to Practitioners—The motivation of this work is to reduce the burden on dentists during tooth preparation treatment, which requires accurate segmentation of crown components (i.e., enamel, dentin, and pulp) from dental CT images. Existing methods only focused on the segmentation of teeth or alveolar bone. Therefore, we present a novel automatic segmentation model for the dental crown components with high accuracy. A key strength of this study is the combination of a data-driven method (deep learning) and model-driven methods (level-set), which can provide good accuracy under limited training samples. This ability is highly desirable for practitioners by saving labor-intensive, costly labeling efforts. Furthermore, our proposed method will provide tools to help reduce subjectivity and human errors, as well as streamline and expedite the clinical workflow. This will significantly facilitate tooth preparation automation.
Mingzhu Zhu, Shaoan Wang, Yaoqing Hu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.2
2024 Robust Federated Semi-Supervised Learning for Medical Image Classification via Pseudo-Label Filtering
abstract
Federated learning (FL) enables collaborative model training across multiple medical institutions to ensure data security. However, due to the variations in medical imaging equipment and regions at different medical institutions, FL methods usually suffer from insufficient data annotations and irrelevant noise within private datasets. To address these issues, a robust federated semi-supervised learning method via pseudo-label filtering (PFRFed) is introduced to utilize unlabeled data while mitigating the impact of noise data. Compared with existing federated semi-supervised learning methods, we propose a pseudo-label filtering mechanism with double dynamic thresholds, which allows the model to adopt more unlabeled data by adjusting the confidence and entropy thresholds at each stage of model training. Moreover, to reduce the degradation caused by noise data in private datasets from different clients, a noise-tolerant loss function and a grouping aggregation method based on the local model similarity are employed. The comparative experiments demonstrate the effectiveness of PFRFed, which has achieved the best classification accuracy of 95.20% and 88.72% on two public medical datasets. Also, PFRFed exhibits heightened resilience to variations in noisy data ratio and labeled data ratio, reaffirming its versatility and robustness.
Shuyu Guo, Mingzhu Zhu, Tian Bai 0002
BIBM3
2024 2D-3D Feature Co-Embedding Network with Sparse Annotation for 3D Medical Image Segmentation
abstract
Supervised methods on 3D medical image segmentation need large amounts of annotated data, but annotating is time-consuming. Also, existing 3D segmentation methods capture more global structural information but overlook local detailed features, which negatively impacts the segmentation of small tissues. In this paper, we propose a novel weakly-supervised 2D-3D Feature Co-Embedding Network (2D-3D CoENet) that includes 2D and 3D encoding layers, simultaneously extracting 2D local detailed and 3D global structural features. To reduce annotation costs, we use fewer labeled slices as ground truth and pseudo-labels are generated by 2D-3D CoENet for other slices. Additionally, multi-view learning is introduced to capture more 2D local detailed information, and a Multi-view Semantic Consistency loss (MSC loss) is proposed to constrain features from multiple perspectives. To further enhance the local detailed texture features, we propose an Edge Enhancement Module (EEM) in the 3D segmentation network to enhance the edge detail features. Our experimental results on the SKI10 dataset and OAI ZIB dataset demonstrate that our method outperforms the SOTA weakly-supervised segmentation methods. Moreover, our approach achieves results that are comparable to the fully-supervised upper bound results.
Mingzhu Zhu, Shuyu Guo, Jianhang Jiao, Tian Bai 0002
BIBM1
2024 A Novel Lightweight Navigation System for Oral and Maxillofacial Surgery Using an External Curved Self-Identifying Checkerboard
abstract
This paper presents a novel lightweight navigation system for oral and maxillofacial surgery (OMS). An external curved checkerboard with self-identifying markers is set as the reference object around the surgical scene. A customized oral clip with a micro camera is designed for oral localization by tracking the external checkerboard. Similarly, the dental handpiece is also equipped with a micro camera, which can be localized like the clip. The spatial model of the markers is provided by binocular stereoscopic reconstruction. A front surface mirror is taken up for the registration between the oral cavity and the camera on the clip. The pivot calibration of the dental handpiece is accomplished by our proposed calibration method. We set up an experimental group and a control group for evaluation. The surgical tools of the experimental group were approximately 30% lighter, 35% less bulky, and 90% cheaper than those of the control group. Our system yielded the comprehensive navigation accuracy of 0.92 mm whereas the accuracy of the control group was 0.87 mm. Results revealed that our system can achieve similar accuracy compared with a prevailing system at a lighter weight, a more compact volume, and a lower cost. Note to Practitioners—The motivation of this work is to reduce the burden on patients and surgeons during OMS. Current commercial navigation systems are still limited by the burden of extra cumbersome fiducial markers and high hardware costs. Their high accuracy benefits from the large size of fiducial markers. To take full advantage of the camera’s localization effect, we propose the concept of “marker-camera inverse projection”, i.e., reversing the roles of the camera and the markers. In this way, cameras on surgical tools detect more points with a more uniform distribution. Our proposed system achieves a decent balance between navigation accuracy and hardware cost, which facilitates the development of surgical tools to be lighter and more economical, and involves tremendous potential for commercialization.
Yaoqing Hu, Mingzhu Zhu, Shaoan Wang, Fusong Yuan, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.2
2024 Robust Visual Feedback Control for Precise In-Hand Manipulation Using Parallel Soft Actuators
abstract
Soft robotic hands are reliable for grasping objects of various shapes. However, they perform poorly in the high-precision manipulation of grasped objects because the modeling and sensing of soft actuator deformation are complex. To overcome this problem, in a previous study, we proposed a robust visual feedback control method for precise in-hand manipulation using parallel soft actuators. This method enables precise in-hand manipulation without measuring the soft actuator deformations. Generally, in the feedback control of a parallel drive system, the actuator force is converted into the force/torque applied to the grasped object. The conversion matrix constantly changes depending on the position/orientation of the object and the contact points of the soft actuators with the object. Consequently, accurately measuring the conversion matrix is difficult. Therefore, we estimated it as a constant matrix in our previous studies, and its robustness was confirmed experimentally. However, its theoretical robustness was not analyzed sufficiently. Therefore, in this article, we discuss the robustness of the estimated constant matrix through a mathematical stability proof. Furthermore, we investigate the characteristics of the estimated drive matrix. Then, we perform a numerical robustness analysis. Finally, the robustness of the proposed method is studied via the above investigation and verification experiments.
Yoshiki Mori, Mingzhu Zhu, Sadao Kawamura
IEEE Trans. Robotics2
2024 CylinderTag: An Accurate and Flexible Marker for Cylinder-Shape Objects Pose Estimation Based on Projective Invariants
abstract
High-precision pose estimation based on visual markers has been a thriving research topic in the field of computer vision. However, the suitability of traditional flat markers on curved objects is limited due to the diverse shapes of curved surfaces, which hinders the development of high-precision pose estimation for curved objects. Therefore, this paper proposes a novel visual marker called CylinderTag, which is designed for developable curved surfaces such as cylindrical surfaces. CylinderTag is a cyclic marker that can be firmly attached to objects with a cylindrical shape. Leveraging the manifold assumption, the cross-ratio in projective invariance is utilized for encoding in the direction of zero curvature on the surface. Additionally, to facilitate the usage of CylinderTag, we propose a heuristic search-based marker generator and a high-performance recognizer as well. Moreover, an all-encompassing evaluation of CylinderTag properties is conducted by means of extensive experimentation, covering detection rate, detection speed, dictionary size, localization jitter, and pose estimation accuracy. CylinderTag showcases superior detection performance from varying view angles in comparison to traditional visual markers, accompanied by higher localization accuracy. Furthermore, CylinderTag boasts real-time detection capability and an extensive marker dictionary, offering enhanced versatility and practicality in a wide range of applications. Experimental results demonstrate that the CylinderTag is a highly promising visual marker for use on cylindrical-like surfaces, thus offering important guidance for future research on high-precision visual localization of cylinder-shaped objects.
Shaoan Wang, Mingzhu Zhu, Yaoqing Hu, Fusong Yuan, Junzhi Yu 0001
IEEE Trans. Vis. Comput. Graph.2
2023 HydraMarker: Efficient, Flexible, and Multifold Marker Field Generation
abstract
An n-order marker field is a special binary matrix whose n×n subregions are all distinct from each other in four orientations. It is commonly used to guide the composing process of position-sensing markers, which can be detected and identified in a camera image with very limited scope or severe visibility problems. Despite the advantages, position-sensing markers are rare and overlooked because generating marker fields is difficult. In this article, we broaden the definition of marker field, making it more powerful and flexible. Then, we propose bWFC (binary wave function collapse) and its high-speed version, fast-bWFC, to solve the generation problem. The methods are packaged into an open-sourced toolkit named HydraMarker, with which, users not only can generate marker fields on laptops within a short period of time, but also can highly customize them: preset values; fields and subregions in any shape; multifold local uniqueness. Comparative results indicate that the proposed method has superior efficiency, quality, and capability. It makes marker field generation accessible to common marker designers, opening up more possibilities for fiducial markers.
Mingzhu Zhu, Bingwei He, Junzhi Yu 0001, Fusong Yuan
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Encode-decode network with fully connected CRF for dynamic objects detection and static maps reconstruction
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001
Signal Process. Image Commun.3
2021 A Multi-Modal Edge Consistency Metric Based on Regression Robustness of Truncated SVD
abstract
In this paper, we propose a novel edge consistency metric for multi-modal correspondence. It is based on a novel observation on image truncated SVD (singular value decomposition) termed regression robustness, which describes the fact that, a good approximation from image truncated SVD can be inherited even if the eigen-images change due to expansion and channel-dependent offsets. Compared to state-of-the-arts, multi-modal edge consistency metric can simultaneously handle multiple images with complex modality changes, including local variation, gradient reverse, intensity order change, and texture loss. Its complexity is almost linear to pixel number. Remarkable accuracies have been achieved in experiments.
Mingzhu Zhu, Junzhi Yu 0001, Zhang Gao, Bingwei He
IEEE Signal Process. Lett.1
2020 ALRe: Outlier Detection for Guided Refinement
Mingzhu Zhu, Zhang Gao, Junzhi Yu 0001, Bingwei He
ECCV (7)1
2020 Boosting dark channel dehazing via weighted local constant assumption
Mingzhu Zhu, Bingwei He, Junzhi Yu 0001
Signal Process.1
2019 Scene flow estimation by depth map upsampling and layer assignment for camera-LiDAR system
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001
J. Vis. Commun. Image Represent.3
2019 Learning motion field of LiDAR point cloud with convolutional networks
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001
Pattern Recognit. Lett.3
2019 Dark Channel: The Devil is in the Details
abstract
Dark channel prior is a highly valued theory mostly used in the area of haze removal. It reveals the wide existence of pixels with very small channel. However, the definition is in a decorated way that relies on the concept of dark channel. The dark channel works like an image domain wherein non-sky regions of clear outdoor images are close to zero. In this letter we show that, such a concept is unnecessary and misleading, and can be improved by using a more basic concept named dark pixel. A novel transmission estimation method based on finding dark pixels is proposed. The method is concise and easy to understand. In comparisons with state-of-the-art methods, our method shows significant improvements on both transmission estimation and haze removal results.
Mingzhu Zhu, Bingwei He
IEEE Signal Process. Lett.1
2018 Development of a Pneumatically Driven Flexible Finger with Feedback Control of a Polyurethane Bend Sensor
abstract
A pneumatically-driven flexible finger equipped with a flexible sensor is realized for improving the performance of the soft robotic hand. First, we propose a flexible angle estimation sensor. This sensor measures the change in the amount of light passing through polyurethane material and estimates the angle with high repeatability. Next, we design a flexible finger that makes this sensor easy to incorporate. The flexible fingers are produced with a multi-material 3D printer that can use flexible material. The flexible finger can accommodate the proposed flexible sensor within it. It is possible to place the sensor's signal line in the air pressure pipeline. Because the flexible finger is produced with a 3D printer, variations in each model's characteristics are small as compared with manufacturing through molding. In this paper, we show an improvement of positional accuracy in the proposed flexible finger using angle feedback control from the proposed sensor. The effectiveness of this sensor is also shown to solve the problem of vibration problems for the flexible finger during high speed motion.
Yoshiki Mori, Mingzhu Zhu, Hye-Jong Kim, Akira Wada, Masahiko Mitsuzuka, Yoshiro Tajitsu, Sadao Kawamura
IROS2
2018 Single Image Dehazing Based on Dark Channel Prior and Energy Minimization
abstract
Hazy images have limited visibility and low contrast. The degradation is expressed by transmission map, which is one of the most important estimates of single image dehazing. Transmission map estimation is an underconstraint problem, and lots of priors have been proposed. Among them, the dark channel prior is widely recognized. However, traditional methods have not fully exploited its power due to improper assumptions or operations, which cause unwanted artifacts. The postrefinement algorithms employed to remove these artifacts in turn undermine the merits of the prior. In this letter, a novel method for estimating transmission map by energy minimization is proposed to solve this problem. The energy function combines the dark channel prior with piecewise smoothness. The method is compared to the state-of-the-art methods and shows outstanding performance.
Mingzhu Zhu, Bingwei He, Qiang Wu 0001
IEEE Signal Process. Lett.1
2017 Atmospheric light estimation in hazy images based on color-plane model
Mingzhu Zhu, Bingwei He, Li-Wei Zhang
Comput. Vis. Image Underst.1
2014 Learning to Rank with Only Positive Examples
abstract
Search By Multiple Examples (SBME) is a new search paradigm that allows users to specify their information needs as a set of relevant documents rather than as a set of keywords. In this study, we propose a Transductive Positive Unlabeled learning (TPU learning) based framework for SBME. The framework consists of two steps: 1) identifying potential relevant documents for searching space reduction, and 2) adopting TPU learning methods to re-rank the documents in the new searching space. Using MAP and p@k, we evaluate two state-of-the-art PU learning algorithms and the Rocchio classifier (Rc) for document ranking in the proposed framework. We then adopt the idea of ensemble learning to combine Rc with the two state-of-the-art PU learning algorithms respectively. Experiments conducted on a benchmark dataset show that the ensemble learning based methods lead to a significant improvement in effectiveness.
Mingzhu Zhu, Yi-fang Brook Wu
ICMLA1
2014 Search by multiple examples
abstract
It is often difficult for users to adopt keywords to express their information needs. Search-By-Multiple-Examples (SBME), a promising method for overcoming this problem, allows users to specify their information needs as a set of relevant documents rather than as a set of keywords. Most of the studies on SBME adopt the Positive Unlabeled learning (PU learning) techniques by treating the users' provided examples (denote as query examples) as positive set and the entire data collection as unlabeled set. However, it is inefficient to treat the entire data collection as unlabeled set, as its size can be huge. In addition, the query examples are treated as being relevant to a single topic, but it is often the case that they can be relevant to multiple topics. As the query examples are much fewer than the unlabeled data, the system performance may downgrade dramatically because of the class imbalance problem. What's more, the experiments conducted in these studies have not taken into account the settings in online search, which are very different from the controlled experiments scenario. This proposed research seeks to explore how to improve SBME by exploring: (1) how to predict user' information needs by modeling the content of the documents using probabilistic topic models; (2) how to deal with the class imbalance problem by reducing the size of the unlabeled data and adopting machine learning techniques. We will also conduct extensive experiments to better evaluate SBME using different sizes of query examples to simulate users' information needs.
Mingzhu Zhu, Yi-fang Brook Wu
WSDM1
2013 Predicting gene regulatory networks of soybean nodulation from RNA-Seq transcriptome data
abstract
BACKGROUND: High-throughput RNA sequencing (RNA-Seq) is a revolutionary technique to study the transcriptome of a cell under various conditions at a systems level. Despite the wide application of RNA-Seq techniques to generate experimental data in the last few years, few computational methods are available to analyze this huge amount of transcription data. The computational methods for constructing gene regulatory networks from RNA-Seq expression data of hundreds or even thousands of genes are particularly lacking and urgently needed. RESULTS: We developed an automated bioinformatics method to predict gene regulatory networks from the quantitative expression values of differentially expressed genes based on RNA-Seq transcriptome data of a cell in different stages and conditions, integrating transcriptional, genomic and gene function data. We applied the method to the RNA-Seq transcriptome data generated for soybean root hair cells in three different development stages of nodulation after rhizobium infection. The method predicted a soybean nodulation-related gene regulatory network consisting of 10 regulatory modules common for all three stages, and 24, 49 and 70 modules separately for the first, second and third stage, each containing both a group of co-expressed genes and several transcription factors collaboratively controlling their expression under different conditions. 8 of 10 common regulatory modules were validated by at least two kinds of validations, such as independent DNA binding motif analysis, gene function enrichment test, and previous experimental data in the literature. CONCLUSIONS: We developed a computational method to reliably reconstruct gene regulatory networks from RNA-Seq transcriptome data. The method can generate valuable hypotheses for interpreting biological data and designing biological experiments such as ChIP-Seq, RNA interference, and yeast two hybrid experiments.
Mingzhu Zhu, Jeremy L. Dahmen, Gary Stacey, Jianlin Cheng
BMC Bioinform.1