Zhiming Cui 0001

dblp:21/6934-1 · DBLP profile ↗
← Back
70ranked-venue papers
4as first author
66since 2021 · last 2026
0000-0002-3798-4504ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 53 · 3 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 30 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Multi-structure segmentation in CBCT volumes: The ToothFairy2 challenge
abstract
Cone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT.
Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana
Medical Image Anal.22
2026 UniSurf: Universal lifespan cortical surface reconstruction
Zifeng Lian, Jiameng Liu, Xiaoye Li, Han Zhang 0002, Zhiming Cui 0001, Feng Shi 0001, Dinggang Shen
Medical Image Anal.7
2026 3D vessel reconstruction from sparse-view dynamic DSA images via vessel probability guided attenuation learning
Huangxuan Zhao, Wenhui Qin, Zhenghong Zhou, Xinggang Wang, Wenping Wang 0001, Xiaochun Lai, Dinggang Shen, Zhiming Cui 0001
Medical Image Anal.9
2026 HALO: High-frequency enhanced dose-aware diffusion model for arbitrary low-dose PET reconstruction
Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Dinggang Shen
Medical Image Anal.4
2026 3D craniofacial generative model for surgical planning in mandibular reconstruction
Chenfan Xu, Haoshen Wang, Jiepeng Wang 0001, Wenbo Du 0001, Xin Peng 0001, Zhiming Cui 0001
Medical Image Anal.9
2026 Automated Dental Landmarks Recognition in Special Models via Neural Network
abstract
Background: Accurate dental landmark recognition is essential in orthodontics and prosthodontics applications but can be challenged by malocclusion and tooth wear. This study expends automated landmark recognition to include such complex cases, aiming to improve anatomical analysis and establish a treatment evaluation framework for Chinese patients. Methods: The study incorporated a dataset of 919 pairs of pretreatment intraoral scan (IOS) models. Initially, a subset of 100 IOS models was segmented and labeled using MeshLab software to generate 3-D dental grid data. Subsequently, 500 IOS models were meticulously labeled by three researchers to train and develop a neural network termed multiscale automated landmark recognition. This network reformulates the task of dental landmark recognition as a geodesic distance field problem on the tooth surface, effectively capturing both local and global geometric features essential for precise detection. The remaining 319 IOS models were then subjected to both manual labeling and automated recognition using the trained network, and the consistency between the two methods was evaluated. Finally, arch width measurements were compared between the two recognition approaches. Results: The neural network shows high concordance with manual labeling, with a mean Euclidean distance discrepancy of 0.375 mm, and no significant difference in arch width measurements between the two methods (P > 0.05). Conclusion: Based on the 0.500 mm clinically acceptable difference standard defined by the ABO objective grading system, the neural network developed in this study demonstrates superior performance in the automated recognition of IOS models exhibiting malocclusion and tooth wear.
Yue Lai, Zhiming Cui 0001, Minhui Tan, Tianmin Xu, Guangying Song
IEEE Trans. Comput. Soc. Syst.3
2026 ToothAxis: Generalizable Tooth Axis Estimation Network From CBCT or IOS Models
abstract
Tooth axes, indicating the orientation of teeth, are crucial in orthodontics and dental implants. The precise and automated estimation of tooth axes in 3D dental models is of significant importance. In clinical settings, Cone-beam computed tomography (CBCT) images and intraoral scanning (IOS) models are the two primary forms of digital data, providing 3D volumetric and surface information of the oral cavity, respectively. However, the detection of tooth axes remains largely manual annotation due to the complexities associated with geometric definitions and the variations among different tooth types and individuals. In this paper, we propose a novel two-stage network, named ToothAxis, for tooth axis estimation using either CBCT or IOS models. Given that IOS models only capture the tooth crown surface and lack information about the tooth roots, we initially employ an implicit-function tooth completion module for 3D tooth completion in the first stage. Subsequently, with the 3D tooth models segmented from CBCT images or completed from IOS models, a point-wise offset-based module is proposed in the second stage to accurately estimate the tooth axes. This design aims to encode tooth orientation into a dense representation, which is better suited for sparse information regression tasks, such as tooth axis estimation. Additionally, we incorporate a class-specific feature attention module to integrate global context representation, thereby enhancing robustness in managing diverse tooth shapes. We evaluated ToothAxis on a dataset obtained from real-world dental clinics, comprising 529 tooth models with corresponding CBCT images and paired IOS models. Finally, the ToothAxis achieves angle errors of LA ($2.921^{\circ }$), PSA ($4.801^{\circ }$), and LSA ($5.074^{\circ }$) on tooth models extracted from CBCT images, and LA ($5.326^{\circ }$), PSA ($6.360^{\circ }$), and LSA ($6.520^{\circ }$) on partial crowns extracted from IOS models. Extensive evaluations, ablation studies, and comparative analyses demonstrate that our method achieves accurate tooth axis estimations and surpasses state-of-the-art approaches.
Qingyao Luo, Zhiming Cui 0001, Yue Zhao 0012
IEEE J. Biomed. Health Informatics4
2026 Progressive Orthodontic Motion Planning Based on Hierarchical Diffusion Transformer
abstract
Orthodontic motion planning plays a crucial role in digital orthodontics by predicting tooth motion sequences to assist dentists in formulating treatment plans efficiently. Most prior work generates the entire intermediate tooth motion sequence given the initial and target tooth alignments. In practice, only the initial alignment of the patient is obtained. However, no existing method can predict the complete motion sequence using only the initial tooth alignment. To address this gap, we propose OrthoDiff, a novel target-free framework that uses only initial tooth alignment through a progressive generation strategy. This strategy generates tooth motion sequences by decomposing the entire motion sequence into multi-level motions, progressively constraining the inference space and reducing the complexity of target-free planning from coarse to fine. Moreover, we design a hierarchical diffusion transformer as the backbone of OrthoDiff, which treats tooth alignment as a sequence of tooth tokens and fully leverages the topological prior knowledge of the dental model. Through extensive evaluations, we demonstrate that our method significantly outperforms state-of-the-art techniques in target-free tooth motion generation. Ablation studies further confirm the efficacy of key components in our network design. Meanwhile, we also achieve state-of-the-art results in tooth target alignment prediction, benefiting from our framework. The code and data will be publicly available at https://github.com/Intelligent-Orthodontics/OrthoDiff.github.io.
Yeying Fan, Yuanfeng Zhou, Guangshun Wei, Zhiming Cui 0001, Yiran Shen 0001, Yong-Jin Liu 0001, Wenping Wang 0001
IEEE Trans. Medical Imaging5
2026 Structure-Preserving Two-Stage Diffusion Model for CBCT Metal Artifact Reduction
abstract
Cone-beam computed tomography (CBCT) plays a crucial role in dental clinical applications, but metal implants often cause severe artifacts, challenging accurate diagnosis. Most deep learning-based methods attempt to achieve metal artifact reduction (MAR) by training neural networks on paired simulated data. However, they often struggle to preserve anatomical structures around metal implants, and fail to bridge the domain gap between real-world and simulated data, leading to suboptimal performance in practice. To address these issues, we propose a two-stage diffusion framework with a strong emphasis on structure preservation and domain generalization. In Stage I, a structure-aware diffusion model is trained to extract artifact-free clean edge maps from artifact-affected CBCT images. This training is supervised by the tooth contours derived from the fusion of intraoral scan (IOS) data and CBCT images to improve generalization to real-world data. In Stage II, these extracted clean edge maps serve as structural priors to guide the MAR process. Additionally, we introduce a segmentation-guided sampling (SGS) strategy in this stage to further enhance structure preservation during inference. Experiments on both simulated and real-world data demonstrate that our method achieves superior artifact reduction and better preservation of dental structures compared to competing approaches.
Haoshen Wang, Minhui Tan, Zhiming Cui 0001
IEEE Trans. Medical Imaging5
2026 Semi-Supervised Landmark Tracking in Echocardiography Video via Spatial-Temporal Co-Training and Perception-Aware Attention
abstract
Precise landmark annotation in cardiac ultrasound images is fundamental for quantitative cardiac health assessment. However, the time-intensive nature of manual annotation typically constrains clinicians to annotate only selected key frames, limiting comprehensive temporal analysis capabilities. While recent automated landmark detection methods have demonstrated success for key-frame analysis, they fail to effectively utilize the intrinsic temporal information across cardiac sequence. To bridge this gap, we present SemiEchoTracker, a novel semi-supervised framework that enables comprehensive landmark tracking throughout echocardiography sequences while requiring supervision only on key frames. Our framework introduces three key innovative strategies: 1) a co-training mechanism that enforces mutual consistency between spatial detection and temporal tracking, enabling accurate intermediate frame detection without additional annotations, 2) a guided DINOv2 pretraining strategy that is specially tailored for extracting fine-grained echocardiography-specific spatial features, and 3) a perception-aware spatial-temporal (PAST) attention module that efficiently captures inter- and intra-frame relationships in echocardiography videos. Extensive validation on three datasets across multiple cardiac views demonstrates that our method not only achieves state-of-the-art detection performance on the keyframes but also yields accurate frame-by-frame prediction, which is important for dynamic cardiac analysis in clinicians.
Han Wu 0007, Zhiming Cui 0001, Dinggang Shen
IEEE Trans. Medical Imaging5
2026 Dual Cross-Image Semantic Consistency With Self-Aware Pseudo Labeling for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning has proven highly effective in tackling the challenge of limited labeled training data in medical image segmentation. In general, current approaches, which rely on intra-image pixel-wise consistency training via pseudo-labeling, overlook the consistency at more comprehensive semantic levels (e.g., object region) and suffer from severe discrepancy of extracted features resulting from an imbalanced number of labeled and unlabeled data. To overcome these limitations, we present a new Dual Cross-image Semantic Consistency (DuCiSC) learning framework, for semi-supervised medical image segmentation. Concretely, beyond enforcing pixel-wise semantic consistency, DuCiSC proposes dual paradigms to encourage region-level semantic consistency across: 1) labeled and unlabeled images; and 2) labeled and fused images, by explicitly aligning their prototypes. Relying on the dual paradigms, DuCiSC can effectively establish consistent cross-image semantics via prototype representations, thereby addressing the feature discrepancy issue. Moreover, we devise a novel self-aware confidence estimation strategy to accurately select reliable pseudo labels, allowing for exploiting the training dynamics of unlabeled data. Our DuCiSC method is extensively validated on four datasets, including two popular binary benchmarks in segmenting the left atrium and pancreas, a multi-class Automatic Cardiac Diagnosis Challenge dataset, and a challenging scenario of segmenting the inferior alveolar nerve that features complicated anatomical structures, showing superior segmentation results over previous state-of-the-art approaches. Our code is publicly available at https://github.com/ShanghaiTech-IMPACT/DuCiSC.
Han Wu 0007, Chong Wang 0012, Zhiming Cui 0001
IEEE Trans. Medical Imaging3
2025 ConsMatch: A Semi-Supervised Segmentation Approach for Dental CBCT by Leveraging Geometric Information to Refine Pseudo-Labels
abstract
Precise tooth instance segmentation from dental CBCT images is essential for accurate diagnosis, yet the scarcity of labeled data and the complex geometric variations of teeth make this task challenging. To address these issues, we propose ConsMatch, a semi-supervised framework that explicitly integrates geometric information into the learning process. It establishes task-level consistency between instance segmentation and boundary extraction, guiding the model to capture finegrained geometric structures. Furthermore, two geometry-aware strategies-Threshold Adjustment Strategy (TAS) and Weight Adjustment Strategy (WAS)-dynamically refine pseudo-label generation by adapting class-specific thresholds and supervision weights based on geometric consistency. This enables the model to focus on high-confidence, structure-consistent pseudolabels, enhancing training stability and segmentation accuracy. Experimental results on dental CBCT data show that ConsMatch achieves superior performance across Dice, Jaccard, and HD95 metrics, consistently outperforming existing semisupervised methods.
Shuyi Lu, Zhiming Cui 0001, Chuanxiang Yang, Guangshun Wei, Yuanfeng Zhou
BIBM3
2025 Tree-Diffusion: Octree-Based Conditional Diffusion Model for Small Bowel Skeleton Generation with Geometric Direction Modeling
abstract
Accurate 3D reconstruction of the small bowel skeleton is vital for understanding intestinal morphology, de-tecting structural abnormalities, and supporting diagnosis, yet limited resolution, organ adhesion, complex anatomy, and scarce annotations make continuous skeleton extraction from masks challenging. Voxel-based methods often struggle with the sparse topology and geometric directionality inherent in the small bowel skeleton, leading to inefficiency and high memory cost. To address these limitations, we propose a novel octree-based conditional diffusion model (i.e., Tree-Diffusion) that generates anatomically consistent small bowel skeletons guided by 3D segmentation masks. Specifically, we introduce two modules that captures structural priors from masks and topology characteristics from skeletons, ensuring cross-domain alignment and high-quality skeleton generation. Besides, we design a synthesis strategy to generate anatomically plausible skeleton-mask pairs, serving as topological priors to guide the diffusion model toward realis-tic structure predictions. To efficiently represent the elongated skeleton, we adopt an octree- based spatial encoding of hierarchical geometric features. Compared with baselines, our model achieves superior performance in anatomical fidelity, directional consistency, and inference efficiency. The code is available at: https://github.com/Small-Bowel-Skeleton-GenerationlCode
Zhichao Liang, Dengqiang Jia, Yaofei Duan, Xinyu Xie, Kaicong Sun, Zhiming Cui 0001, Tao Tan 0002, Dinggang Shen
BIBM7
2025 Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
abstract
Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Very recent works address this problem by applying a video diffusion model to generate video depth conditioned on the input video, which is training-expensive and can only produce scale-invariant depth values without camera poses. In this paper, we propose a novel video-depth estimation method called Align3R to estimate temporally consistent depth maps for a dynamic video. Our key idea is to utilize the recent DUSt3R model to align estimated monocular depth maps of different timesteps. First, we fine-tune the DUSt3R model with additional estimated monocular depth as inputs for the dynamic scenes. Then, we apply optimization to reconstruct both depth maps and camera poses. Extensive experiments demonstrate that Align3R estimates consistent video depth and camera poses for a monocular video with superior performance than baseline methods.
Jiahao Lu 0001, Zhiyang Dou, Cheng Lin 0001, Zhiming Cui 0001, Zhen Dong 0005, Sai-Kit Yeung, Wenping Wang 0001, Yuan Liu 0025
CVPR6
2025 VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch Generation
Zhiming Cui 0001, Changjian Li 0001
ICCV2
2025 Adapting Foundation Model for Dental Caries Detection with Dual-View Co-training
Tao Luo 0010, Han Wu 0007, Dinggang Shen, Zhiming Cui 0001
MICCAI (16)5
2025 A New Paradigm for Low-Dose PET/CT Reconstruction with Mamba-Powered Progressive Network and Physics-Informed Consistency
Caiwen Jiang, Zhiming Cui 0001, Dinggang Shen
MICCAI (11)3
2025 Location-Guided Automated Lesion Captioning in Whole-Body PET/CT Images
Mingyang Yu 0009, Yaozong Gao, Yiran Shu, Yanbo Chen 0003, Jingyu Liu 0002, Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Shaonan Zhong, Xinlu Wang, Meixin Zhao, Dinggang Shen
MICCAI (5)8
2025 Mask2Surface: Motion Correction and Super-Resolution for Cardiac Surface Reconstruction Using Latent Diffusion
Zeng Zhang, Zhiming Cui 0001
MICCAI (2)4
2025 Diff-OSGN: Diffusion-Based Occlusal Surface Generation Network with Geometric Constraints
abstract
Designing a functional occlusal surface for denture crowns is a complex and important task in prosthodontics. Manual design is time-consuming and heavily relies on the dentist's experience, as it requires careful consideration of occlusal function. Due to the limitations of manual design, the field has turned to data-driven methods for occlusal surface design. However, many of these methods neglect critical geometric details, such as normals and curvature, impacting the quality of the occlusal surface. In this paper, we introduce Diff-OSGN, a novel denture crown occlusal surface generation network based on a denoising diffusion model, which focuses on generating the detailed geometric structure of denture crowns. We model the occlusal surface as a geometry map based on the occlusal plane, incorporating height and normal maps rasterized from intra-oral crown scanning. Both maps represent occlusal surface geometry, and their combination further enhances these details. Considering the crucial occlusal information, we extract features from the geometry maps of adjacent and occlusal teeth, using them as conditions in the reverse diffusion process to train our network for optimal occlusal function. Additionally, we define three geometric operators and corresponding loss functions as constraints to better extract geometric features of the target occlusal surface, such as ridges and grooves, for adequate supervision. Our results demonstrate that Diff-OSGN provides quantitatively and qualitatively superior performance than competing baselines and state-of-the-art methods.
Chen Wang 0054, Guangshun Wei, James Kit Hon Tsoi, Zhiming Cui 0001, Shuyi Lu, Zhenpeng Liu, Yuanfeng Zhou
Comput. Vis. Media4
2025 CLIK-Diffusion: Clinical Knowledge-informed Diffusion Model for Tooth Alignment
Yulong Dou, Han Wu 0007, Changjian Li 0001, Chen Wang 0054, Dinggang Shen, Zhiming Cui 0001
Medical Image Anal.8
2025 Clinical knowledge-guided hybrid classification network for automatic periodontal disease diagnosis in X-ray image
Lanzhuju Mei, Zhiming Cui 0001, Yu Fang 0008, Yuan Liu 0025, Hongchang Lai, Maurizio Tonetti, Dinggang Shen
Medical Image Anal.3
2025 Segmenting the Inferior Alveolar Canal in CBCTs Volumes: The ToothFairy Challenge
abstract
In recent years, several algorithms have been developed for the segmentation of the Inferior Alveolar Canal (IAC) in Cone-Beam Computed Tomography (CBCT) scans. However, the availability of public datasets in this domain is limited, resulting in a lack of comparative evaluation studies on a common benchmark. To address this scientific gap and encourage deep learning research in the field, the ToothFairy challenge was organized within the MICCAI 2023 conference. In this context, a public dataset was released to also serve as a benchmark for future research. The dataset comprises 443 CBCT scans, with voxel-level annotations of the IAC available for 153 of them, making it the largest publicly available dataset of its kind. The participants of the challenge were tasked with developing an algorithm to accurately identify the IAC using the 2D and 3D-annotated scans. This paper presents the details of the challenge and the contributions made by the most promising methods proposed by the participants. It represents the first comprehensive comparative evaluation of IAC segmentation methods on a common benchmark dataset, providing insights into the current state-of-the-art algorithms and outlining future research directions. Furthermore, to ensure reproducibility and promote future developments, an open-source repository that collects the implementations of the best submissions was released.
Federico Bolelli, Luca Lumetti, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Arrigo Pellacani, Kevin Marchesini, Niels van Nistelrooij, Pieter van Lierop, Tong Xi 0001, Yusheng Liu 0001, Rui Xin 0003, Tao Yang 0037, Lisheng Wang, Haoshen Wang, Chenfan Xu, Zhiming Cui 0001, Marek Wodzinski, Henning Müller, Yannick Kirchhoff, Maximilian Rokuss, Klaus H. Maier-Hein, Jae-Hwan Han, Wan Kim, Hong-Gi Ahn, Tomasz Szczepanski, Michal K. Grzeszczyk, Przemyslaw Korzeniowski, Vicent Caselles, Xavier Paolo Burgos-Artizzu, Ferran Prados, Stefaan Bergé, Bram van Ginneken, Alexandre Anesi, Costantino Grana
IEEE Trans. Medical Imaging16
2025 Geometry-Aware Attenuation Learning for Sparse-View CBCT Reconstruction
abstract
Cone Beam Computed Tomography (CBCT) plays a vital role in clinical imaging. Traditional methods typically require hundreds of 2D X-ray projections to reconstruct a high-quality 3D CBCT image, leading to considerable radiation exposure. This has led to a growing interest in sparse-view CBCT reconstruction to reduce radiation doses. While recent advances, including deep learning and neural rendering algorithms, have made strides in this area, these methods either produce unsatisfactory results or suffer from time inefficiency of individual optimization. In this paper, we introduce a novel geometry-aware encoder-decoder framework to solve this problem. Our framework starts by encoding multi-view 2D features from various 2D X-ray projections with a 2D CNN encoder. Leveraging the geometry of CBCT scanning, it then back-projects the multi-view 2D features into the 3D space to formulate a comprehensive volumetric feature map, followed by a 3D CNN decoder to recover 3D CBCT image. Importantly, our approach respects the geometric relationship between 3D CBCT image and its 2D X-ray projections during feature back projection stage, and enjoys the prior knowledge learned from the data population. This ensures its adaptability in dealing with extremely sparse view inputs without individual training, such as scenarios with only 5 or 10 X-ray projections. Extensive evaluations on two simulated datasets and one real-world dataset demonstrate exceptional reconstruction quality and time efficiency of our method.
Yu Fang 0008, Changjian Li 0001, Han Wu 0007, Yuan Liu 0025, Dinggang Shen, Zhiming Cui 0001
IEEE Trans. Medical Imaging7
2025 3D MedDiffusion: A 3D Medical Latent Diffusion Model for Controllable and High-Quality Medical Image Generation
abstract
The generation of medical images presents significant challenges due to their high-resolution and three-dimensional nature. Existing methods often yield suboptimal performance in generating high-quality 3D medical images, and there is currently no universal generative framework for medical imaging. In this paper, we introduce a 3D Medical Latent Diffusion (3D MedDiffusion) model for controllable, high-quality 3D medical image generation. 3D MedDiffusion incorporates a novel, highly efficient Patch-Volume Autoencoder that compresses medical images into latent space through patch-wise encoding and recovers back into image space through volume-wise decoding. Additionally, we design a new noise estimator to capture both local details and global structural information during diffusion denoising process. 3D MedDiffusion can generate fine-detailed, high-resolution images (up to ${512}\times {512}\times {512}$ ) and effectively adapt to various downstream tasks as it is trained on large-scale datasets covering CT and MRI modalities and different anatomical regions (from head to leg). Experimental results demonstrate that 3D MedDiffusion surpasses state-of-the-art methods in generative quality and exhibits strong generalizability across tasks such as sparse-view CT reconstruction, fast MRI reconstruction, and data augmentation for segmentationand classification. Source code and checkpoints are available at https://github.com/ShanghaiTech-IMPACT/3D-MedDiffusion.
Haoshen Wang, Kaicong Sun, Dinggang Shen, Zhiming Cui 0001
IEEE Trans. Medical Imaging6
2025 Integrating Eye Tracking With Grouped Fusion Networks for Semantic Segmentation on Mammogram Images
abstract
Medical image segmentation has seen great progress in recent years, largely due to the development of deep neural networks. However, unlike in computer vision, high-quality clinical data is relatively scarce, and the annotation process is often a burden for clinicians. As a result, the scarcity of medical data limits the performance of existing medical image segmentation models. In this paper, we propose a novel framework that integrates eye tracking information from experienced radiologists during the screening process to improve the performance of deep neural networks with limited data. Our approach, a grouped hierarchical network, guides the network to learn from its faults by using gaze information as weak supervision. We demonstrate the effectiveness of our framework on mammogram images, particularly for handling segmentation classes with large scale differences. We evaluate the impact of gaze information on medical image segmentation tasks and show that our method achieves better segmentation performance compared to state-of-the-art models. A robustness study is conducted to investigate the influence of distraction or inaccuracies in gaze collection. We also develop a convenient system for collecting gaze data without interrupting the normal clinical workflow. Our work offers novel insights into the potential benefits of integrating gaze information into medical image segmentation tasks.
Jiaming Xie, Zhiming Cui 0001, Chong Ma 0004, Wenping Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging3
2025 Coupled Diffusion Models for Metal Artifact Reduction of Clinical Dental CBCT Images
abstract
Metal dental implants may introduce metal artifacts (MA) during the CBCT imaging process, causing significant interference in subsequent diagnosis. In recent years, many deep learning methods for metal artifact reduction (MAR) have been proposed. Due to the huge difference between synthetic and clinical MA, supervised learning MAR methods may perform poorly in clinical settings. Many existing unsupervised MAR methods trained on clinical data often suffer from incorrect dental morphology. To alleviate the above problems, in this paper, we propose a new MAR method of Coupled Diffusion Models (CDM) for clinical dental CBCT images. Specifically, we separately train two diffusion models on clinical MA-degraded images and clinical clean images to obtain prior information, respectively. During the denoising process, the variances of noise levels are calculated from MA images and the prior of diffusion models. Then we develop a noise transformation module between the two diffusion models to transform the MA noise image into a new initial value for the denoising process. Our designs effectively exploit the inherent transformation between the misaligned MA-degraded images and clean images. Additionally, we introduce an MA-adaptive inference technique to better accommodate the MA degradation in different areas of an MA-degraded image. Experiments on our clinical dataset demonstrate that our CDM outperforms the comparison methods on both objective metrics and visual quality, especially for severe MA degradation. We will publicly release our code.
Zhouzhuo Zhang, Juncheng Yan, Zhiming Cui 0001, Jun Xu 0019, Dinggang Shen
IEEE Trans. Medical Imaging4
2025 Robust Hybrid Learning for Automatic Teeth Segmentation and Labeling on 3D Dental Models
abstract
Automatic teeth segmentation and labeling on dental models are basic tasks in computer-aided dentistry. Many existing works can achieve promising results in teeth segmentation, but they heavily rely on aligned input dental models, which leads to additional manual intervention. Moreover, tooth labeling is an essential task in digital dentistry for treatment planning (e.g., orthodontic), and is usually ignored in these methods. In this article, we propose an AlignNet for aligning dental models of arbitrary sizes and orientations automatically. Meanwhile, a multi-task hybrid learning network is designed that effectively plays the advantages of semantic segmentation and instance segmentation, and synergistically improves the performance of teeth point clouds segmentation and labeling. Particularly, for the teeth-gingival boundaries with large segmentation errors, we utilize the filtered curvature information as a constrained feature to detect the weak boundary more accurately. At last, we propose a DiffLoss and postprocessing step based on the dental arch to address the teeth classification problem. Through extensive evaluations of oral scanning models, our method is robust to handle dental model point clouds with arbitrary size and orientation, and outperforms state-of-the-art teeth segmentation and labeling methods, demonstrating its full automation and robustness in clinical practice.
Shaojie Zhuang 0001, Guangshun Wei, Zhiming Cui 0001, Yuanfeng Zhou
IEEE Trans. Multim.3
2025 A Novel Hierarchical Cross-Stream Aggregation Neural Network for Semantic Segmentation of 3-D Dental Surface Models
abstract
Accurate teeth delineation on 3-D dental models is essential for individualized orthodontic treatment planning. Pioneering works like PointNet suggest a promising direction to conduct efficient and accurate 3-D dental model analyses in end-to-end learnable fashions. Recent studies further imply that multistream architectures to concurrently learn geometric representations from different inputs/views (e.g., coordinates and normals) are beneficial for segmenting teeth with varying conditions. However, such multistream networks typically adopt simple late-fusion strategies to combine features captured from raw inputs that encode complementary but fundamentally different geometric information, potentially hampering their accuracy in end-to-end semantic segmentation. This article presents a hierarchical cross-stream aggregation (HiCA) network to learn more discriminative point/cell-wise representations from multiview inputs for fine-grained 3-D semantic segmentation. Specifically, based upon our multistream backbone with input-tailored feature extractors, we first design a contextual cross-steam aggregation (CA) module conditioned on interstream consistency to boost each view's contextual representation learning jointly. Then, before the late fusion of different streams' outputs for segmentation, we further deploy a discriminative cross-stream aggregation (DA) module to concurrently update all views' discriminative representation learning by leveraging a specific graph attention strategy induced by multiview prototype learning. On both public and in-house datasets of real-patient dental models, our method significantly outperformed state-of-the-art (SOTA) deep learning methods for teeth semantic segmentation. In addition, extended experimental results suggest the applicability of HiCA to other general 3-D shape segmentation tasks. The code is available at https://github.com/ladderlab-xjtu/HiCA.
Kehan Li 0009, Jihua Zhu, Zhiming Cui 0001, Xinning Chen, Yang Liu 0157, Fan Wang 0038, Yue Zhao 0012
IEEE Trans. Neural Networks Learn. Syst.3
2024 A Prior-information-guided Residual Diffusion Model for Multi-modal PET Synthesis from MRI
Zaixin Ou, Caiwen Jiang, Yongsheng Pan, Yuanwang Zhang, Zhiming Cui 0001, Dinggang Shen
IJCAI5
2024 A Graph-Embedded Latent Space Learning and Clustering Framework for Incomplete Multimodal Multiclass Alzheimer's Disease Diagnosis
Zaixin Ou, Caiwen Jiang, Yuanwang Zhang, Zhiming Cui 0001, Dinggang Shen
MICCAI (7)5
2024 HF-ResDiff: High-Frequency-Guided Residual Diffusion for Multi-dose PET Reconstruction
Caiwen Jiang, Zhiming Cui 0001, Dinggang Shen
MICCAI (7)3
2024 Cephalometric Landmark Detection Across Ages with Prototypical Network
Han Wu 0007, Chong Wang 0012, Lanzhuju Mei, Dinggang Shen, Zhiming Cui 0001
MICCAI (5)7
2024 TeethDreamer: 3D Teeth Reconstruction from Five Intra-Oral Photographs
Chenfan Xu, Yuan Liu 0025, Yulong Dou, Jiepeng Wang 0001, Minjiao Wang, Dinggang Shen, Zhiming Cui 0001
MICCAI (7)9
2024 DTR-Net: Dual-Space 3D Tooth Model Reconstruction From Panoramic X-Ray Images
abstract
In digital dentistry, cone-beam computed tomography (CBCT) can provide complete 3D tooth models, yet suffers from a long concern of requiring excessive radiation dose and higher expense. Therefore, 3D tooth model reconstruction from 2D panoramic X-ray image is more cost-effective, and has attracted great interest in clinical applications. In this paper, we propose a novel dual-space framework, namely DTR-Net, to reconstruct 3D tooth model from 2D panoramic X-ray images in both image and geometric spaces. Specifically, in the image space, we apply a 2D-to-3D generative model to recover intensities of CBCT image, guided by a task-oriented tooth segmentation network in a collaborative training manner. Meanwhile, in the geometric space, we benefit from an implicit function network in the continuous space, learning using points to capture complicated tooth shapes with geometric properties. Experimental results demonstrate that our proposed DTR-Net achieves state-of-the-art performance both quantitatively and qualitatively in 3D tooth model reconstruction, indicating its potential application in dental practice.
Lanzhuju Mei, Yu Fang 0008, Yue Zhao 0012, Xiang Sean Zhou, Zhiming Cui 0001, Dinggang Shen
IEEE Trans. Medical Imaging6
2024 ChatCAD+: Toward a Universal and Reliable Interactive CAD Using LLMs
abstract
The integration of Computer-Aided Diagnosis (CAD) with Large Language Models (LLMs) presents a promising frontier in clinical applications, notably in automating diagnostic processes akin to those performed by radiologists and providing consultations similar to a virtual family doctor. Despite the promising potential of this integration, current works face at least two limitations: (1) From the perspective of a radiologist, existing studies typically have a restricted scope of applicable imaging domains, failing to meet the diagnostic needs of different patients. Also, the insufficient diagnostic capability of LLMs further undermine the quality and reliability of the generated medical reports. (2) Current LLMs lack the requisite depth in medical expertise, rendering them less effective as virtual family doctors due to the potential unreliability of the advice provided during patient consultations. To address these limitations, we introduce ChatCAD+, to be universal and reliable. Specifically, it is featured by two main modules: (1) Reliable Report Generation and (2) Reliable Interaction. The Reliable Report Generation module is capable of interpreting medical images from diverse domains and generate high-quality medical reports via our proposed hierarchical in-context learning. Concurrently, the interaction module leverages up-to-date information from reputable medical websites to provide reliable medical advice. Together, these designed modules synergize to closely align with the expertise of human medical professionals, offering enhanced consistency and reliability for interpretation and advice. The source code is available at GitHub.
Zihao Zhao 0002, Sheng Wang 0014, Jinchen Gu, Yitao Zhu, Lanzhuju Mei, Zixu Zhuang, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging7
2024 Hierarchical Organ-Aware Total-Body Standard-Dose PET Reconstruction From Low-Dose PET and CT Images
abstract
Positron emission tomography (PET) is an important functional imaging technology in early disease diagnosis. Generally, the gamma ray emitted by standard-dose tracer inevitably increases the exposure risk to patients. To reduce dosage, a lower dose tracer is often used and injected into patients. However, this often leads to low-quality PET images. In this article, we propose a learning-based method to reconstruct total-body standard-dose PET (SPET) images from low-dose PET (LPET) images and corresponding total-body computed tomography (CT) images. Different from previous works focusing only on a certain part of human body, our framework can hierarchically reconstruct total-body SPET images, considering varying shapes and intensity distributions of different body parts. Specifically, we first use one global total-body network to coarsely reconstruct total-body SPET images. Then, four local networks are designed to finely reconstruct head-neck, thorax, abdomen-pelvic, and leg parts of human body. Moreover, to enhance each local network learning for the respective local body part, we design an organ-aware network with a residual organ-aware dynamic convolution (RO-DC) module by dynamically adapting organ masks as additional inputs. Extensive experiments on 65 samples collected from uEXPLORER PET/CT system demonstrate that our hierarchical framework can consistently improve the performance of all body parts, especially for total-body PET images with PSNR of 30.6 dB, outperforming the state-of-the-art methods in SPET image reconstruction.
Zhiming Cui 0001, Caiwen Jiang, Fei Gao 0010, Dinggang Shen
IEEE Trans. Neural Networks Learn. Syst.2
2023 3D Structure-guided Network for Tooth Alignment in 2D Photograph
Yulong Dou, Lanzhuju Mei, Dinggang Shen, Zhiming Cui 0001
BMVC4
2023 HC-Net: Hybrid Classification Network for Automatic Periodontal Disease Diagnosis
Lanzhuju Mei, Yu Fang 0008, Zhiming Cui 0001, Nizhuan Wang 0001, Xuming He 0001, Yiqiang Zhan, Xiang Sean Zhou, Maurizio Tonetti, Dinggang Shen
MICCAI (6)3
2023 Multi-view Vertebra Localization and Identification from CT Images
Han Wu 0007, Yu Fang 0008, Nizhuan Wang 0001, Zhiming Cui 0001, Dinggang Shen
MICCAI (5)6
2023 Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002
MICCAI (3)9
2023 TS-DSANN: Texture and shape focused dual-stream attention neural network for benign-malignant diagnosis of thyroid nodules in ultrasound images
Lu Tang 0001, Chuangeng Tian, Zhiming Cui 0001, Dinggang Shen
Medical Image Anal.4
2023 Semi-Supervised Standard-Dose PET Image Generation via Region-Adaptive Normalization and Structural Consistency Constraint
abstract
Positron Emission Tomography (PET) is an important nuclear medical imaging technique, and has been widely used in clinical applications, e.g., tumor detection and brain disease diagnosis. As PET imaging could put patients at risk of radiation, the acquisition of high-quality PET images with standard-dose tracers should be cautious. However, if dose is reduced in PET acquisition, the imaging quality could become worse and thus may not meet clinical requirement. To safely reduce the tracer dose and also maintain high quality of PET imaging, we propose a novel and effective approach to estimate high-quality Standard-dose PET (SPET) images from Low-dose PET (LPET) images. Specifically, to fully utilize both the rare paired and the abundant unpaired LPET and SPET images, we propose a semi-supervised framework for network training. Meanwhile, based on this framework, we further design a Region-adaptive Normalization (RN) and a structural consistency constraint to track the task-specific challenges. RN performs region-specific normalization in different regions of each PET image to suppress negative impact of large intensity variation across different regions, while the structural consistency constraint maintains structural details during the generation of SPET images from LPET images. Experiments on real human chest-abdomen PET images demonstrate that our proposed approach achieves state-of-the-art performance quantitatively and qualitatively.
Caiwen Jiang, Yongsheng Pan, Zhiming Cui 0001, Dong Nie, Dinggang Shen
IEEE Trans. Medical Imaging3
2023 BowelNet: Joint Semantic-Geometric Ensemble Learning for Bowel Segmentation From Both Partially and Fully Labeled CT Images
abstract
Accurate bowel segmentation is essential for diagnosis and treatment of bowel cancers. Unfortunately, segmenting the entire bowel in CT images is quite challenging due to unclear boundary, large shape, size, and appearance variations, as well as diverse filling status within the bowel. In this paper, we present a novel two-stage framework, named BowelNet, to handle the challenging task of bowel segmentation in CT images, with two stages of 1) jointly localizing all types of the bowel, and 2) finely segmenting each type of the bowel. Specifically, in the first stage, we learn a unified localization network from both partially- and fully-labeled CT images to robustly detect all types of the bowel. To better capture unclear bowel boundary and learn complex bowel shapes, in the second stage, we propose to jointly learn semantic information (i.e., bowel segmentation mask) and geometric representations (i.e., bowel boundary and bowel skeleton) for fine bowel segmentation in a multi-task learning scheme. Moreover, we further propose to learn a meta segmentation network via pseudo labels to improve segmentation accuracy. By evaluating on a large abdominal CT dataset, our proposed BowelNet method can achieve Dice scores of 0.764, 0.848, 0.835, 0.774, and 0.824 in segmenting the duodenum, jejunum-ileum, colon, sigmoid, and rectum, respectively. These results demonstrate the effectiveness of our proposed BowelNet framework in segmenting the entire bowel from CT images.
Chong Wang 0012, Zhiming Cui 0001, Miaofei Han, Gustavo Carneiro 0001, Dinggang Shen
IEEE Trans. Medical Imaging2
2023 Breast Fibroglandular Tissue Segmentation for Automated BPE Quantification With Iterative Cycle-Consistent Semi-Supervised Learning
abstract
Background Parenchymal Enhancement (BPE) quantification in Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) plays a pivotal role in clinical breast cancer diagnosis and prognosis. However, the emerging deep learning-based breast fibroglandular tissue segmentation, a crucial step in automated BPE quantification, often suffers from limited training samples with accurate annotations. To address this challenge, we propose a novel iterative cycle-consistent semi-supervised framework to leverage segmentation performance by using a large amount of paired pre-/post-contrast images without annotations. Specifically, we design the reconstruction network, cascaded with the segmentation network, to learn a mapping from the pre-contrast images and segmentation predictions to the post-contrast images. Thus, we can implicitly use the reconstruction task to explore the inter-relationship between these two-phase images, which in return guides the segmentation task. Moreover, the reconstructed post-contrast images across multiple auto-context modeling-based iterations can be viewed as new augmentations, facilitating cycle-consistent constraints across each segmentation output. Extensive experiments on two datasets with various data distributions show great segmentation and BPE quantification accuracy compared with other state-of-the-art semi-supervised methods. Importantly, our method achieves 11.80 times of quantification accuracy improvement along with 10 times faster, compared with clinical physicians, demonstrating its potential for automated BPE quantification. The code is available at https://github.com/ZhangJD-ong/Iterative-Cycle-consistent-Semi-supervised-Learning-for-fibroglandular-tissue-segmentation.
Zhiming Cui 0001, Luping Zhou, Yiqun Sun, Zhenhui Li, Zaiyi Liu, Dinggang Shen
IEEE Trans. Medical Imaging2
2022 Curvature-Enhanced Implicit Function Network for High-quality Tooth Model Generation from CBCT Images
Yu Fang 0008, Zhiming Cui 0001, Lei Ma 0006, Lanzhuju Mei, Yue Zhao 0012, Zhihao Jiang 0001, Yiqiang Zhan, Yongsheng Pan, Dinggang Shen
MICCAI (5)2
2022 Deep Learning-Based Head and Neck Radiotherapy Planning Dose Prediction via Beam-Wise Dose Decomposition
Bin Wang 0068, Lanzhuju Mei, Zhiming Cui 0001, Xuanang Xu, Qianjin Feng 0003, Dinggang Shen
MICCAI (8)4
2022 Mapping in Cycles: Dual-Domain PET-CT Synthesis Framework with Cycle-Consistent Constraints
Zhiming Cui 0001, Caiwen Jiang, Jingyang Zhang, Fei Gao 0010, Dinggang Shen
MICCAI (6)2
2022 Learning Towards Synchronous Network Memorizability and Generalizability for Continual Segmentation Across Multiple Sites
Jingyang Zhang, Peng Xue 0005, Ran Gu, Yuning Gu, Mianxin Liu, Yongsheng Pan, Zhiming Cui 0001, Lei Ma 0006, Dinggang Shen
MICCAI (5)7
2022 Dense representative tooth landmark/axis detection network on 3D model
Guangshun Wei, Zhiming Cui 0001, Lei Yang 0048, Yuanfeng Zhou, Pradeep Singh 0003, Min Gu 0003, Wenping Wang 0001
Comput. Aided Geom. Des.2
2022 TAD-Net: tooth axis detection network based on rotation transformation encoding
Yeying Fan, Guangshun Wei, Zhiming Cui 0001, Yuanfeng Zhou, Wenping Wang 0001
Graph. Model.4
2022 Semi-supervised anatomical landmark detection via shape-regulated self-training
Runnan Chen, Yuexin Ma, Lingjie Liu, Nenglun Chen, Zhiming Cui 0001, Guodong Wei, Wenping Wang 0001
Neurocomputing5
2022 Grayscale self-adjusting network with weak feature enhancement for 3D lumbar anatomy segmentation
Jinhua Liu 0003, Zhiming Cui 0001, Christian Desrosiers, Shuyi Lu, Yuanfeng Zhou
Medical Image Anal.2
2022 GAN-Guided Deformable Attention Network for Identifying Thyroid Nodules in Ultrasound Images
abstract
Early detection and identification of malignant thyroid nodules, a vital precursory to the treatment, is a difficult task even for experienced clinicians. Many Computer-Aided Diagnose (CAD) systems have been developed to assist clinicians in performing this task on ultrasonic images. Learning-based CAD systems for thyroid nodules generally accommodate both nodule detection/ segmentation and fine-grained classification for its malignancy, and prior researches often treat aforementioned tasks in separate stages, leading to additional computational costs. In this paper, we utilize an online class activation mapping (CAM) mechanism to guide the network to learn discriminative features for identifying thyroid nodules in ultrasound images, called CAM attention network. It takes nodule masks as localization cues for direct spatial attention of the classification module, thereby avoiding isolated training for classification. Meanwhile, we propose a deformable convolution module to add offsets to the regular grid sampling locations in the standard convolution, guiding the network to capture more discriminative features of nodule areas. Furthermore, we use a generative adversarial network (GAN)to ensure reliable deformations of nodules from the deformable convolution module. Our proposed CAM attention network has already achieved the 2nd place in the classification task of TN-SCUI 2020, a MICCAI 2020 Challenge with the largest set of thyroid nodule ultrasound images according to our knowledge. The further inclusion of our proposed GAN-guided deformable module allows for capturing more fine-grained features between benign and malignant nodules, and further improves the classification accuracy to a new state-of-the-art level.
Jintao Lu, Xi Ouyang, Xueda Shen, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics5
2022 Structure-Aware Long Short-Term Memory Network for 3D Cephalometric Landmark Detection
abstract
Detecting 3D landmarks on cone-beam computed tomography (CBCT) is crucial to assessing and quantifying the anatomical abnormalities in 3D cephalometric analysis. However, the current methods are time-consuming and suffer from large biases in landmark localization, leading to unreliable diagnosis results. In this work, we propose a novel Structure-Aware Long Short-Term Memory framework (SA-LSTM) for efficient and accurate 3D landmark detection. To reduce the computational burden, SA-LSTM is designed in two stages. It first locates the coarse landmarks via heatmap regression on a down-sampled CBCT volume and then progressively refines landmarks by attentive offset regression using multi-resolution cropped patches. To boost accuracy, SA-LSTM captures global-local dependence among the cropping patches via self-attention. Specifically, a novel graph attention module implicitly encodes the landmark's global structure to rationalize the predicted position. Moreover, a novel attention-gated module recursively filters irrelevant local features and maintains high-confident local predictions for aggregating the final result. Experiments conducted on an in-house dataset and a public dataset show that our method outperforms state-of-the-art methods, achieving 1.64 mm and 2.37 mm average errors, respectively. Furthermore, our method is very efficient, taking only 0.5 seconds for inferring the whole CBCT volume of resolution 768×768×576 .
Runnan Chen, Yuexin Ma, Nenglun Chen, Lingjie Liu, Zhiming Cui 0001, Yanhong Lin, Wenping Wang 0001
IEEE Trans. Medical Imaging5
2022 Semantic Graph Attention With Explicit Anatomical Association Modeling for Tooth Segmentation From CBCT Images
abstract
Accurate tooth identification and delineation in dental CBCT images are essential in clinical oral diagnosis and treatment. Teeth are positioned in the alveolar bone in a particular order, featuring similar appearances across adjacent and bilaterally symmetric teeth. However, existing tooth segmentation methods ignored such specific anatomical topology, which hampers the segmentation accuracy. Here we propose a semantic graph-based method to explicitly model the spatial associations between different anatomical targets (i.e., teeth) for their precise delineation in a coarse-to-fine fashion. First, to efficiently control the bilaterally symmetric confusion in segmentation, we employ a lightweight network to roughly separate teeth as four quadrants. Then, designing a semantic graph attention mechanism to explicitly model the anatomical topology of the teeth in each quadrant, based on which voxel-wise discriminative feature embeddings are learned for the accurate delineation of teeth boundaries. Extensive experiments on a clinical dental CBCT dataset demonstrate the superior performance of the proposed method compared with other state-of-the-art approaches.
Pengcheng Li 0017, Yang Liu 0157, Zhiming Cui 0001, Feng Yang 0015, Yue Zhao 0012, Chunfeng Lian, Chenqiang Gao
IEEE Trans. Medical Imaging3
2022 Two-Stream Graph Convolutional Network for Intra-Oral Scanner Image Segmentation
abstract
Precise segmentation of teeth from intra-oral scanner images is an essential task in computer-aided orthodontic surgical planning. The state-of-the-art deep learning-based methods often simply concatenate the raw geometric attributes (i.e., coordinates and normal vectors) of mesh cells to train a single-stream network for automatic intra-oral scanner image segmentation. However, since different raw attributes reveal completely different geometric information, the naive concatenation of different raw attributes at the (low-level) input stage may bring unnecessary confusion in describing and differentiating between mesh cells, thus hampering the learning of high-level geometric representations for the segmentation task. To address this issue, we design a two-stream graph convolutional network (i.e., TSGCN), which can effectively handle inter-view confusion between different raw attributes to more effectively fuse their complementary information and learn discriminative multi-view geometric representations. Specifically, our TSGCN adopts two input-specific graph-learning streams to extract complementary high-level geometric representations from coordinates and normal vectors, respectively. Then, these single-view representations are further fused by a self-attention module to adaptively balance the contributions of different views in learning more discriminative multi-view representations for accurate and fully automatic tooth segmentation. We have evaluated our TSGCN on a real-patient dataset of dental (mesh) models acquired by 3D intraoral scanners. Experimental results show that our TSGCN significantly outperforms state-of-the-art methods in 3D tooth (surface) segmentation.
Yue Zhao 0012, Yang Liu 0157, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen
IEEE Trans. Medical Imaging5
2021 TSGCNet: Discriminative Geometric Feature Learning With Two-Stream Graph Convolutional Network for 3D Dental Model Segmentation
abstract
The ability to segment teeth precisely from digitized 3D dental models is an essential task in computer-aided orthodontic surgical planning. To date, deep learning based methods have been popularly used to handle this task. State-of-the-art methods directly concatenate the raw attributes of 3D inputs, namely coordinates and normal vectors of mesh cells, to train a single-stream network for fully-automated tooth segmentation. This, however, has the drawback of ignoring the different geometric meanings provided by those raw attributes. This issue might possibly confuse the network in learning discriminative geometric features and result in many isolated false predictions on the dental model. Against this issue, we propose a two-stream graph convolutional network (TSGCNet) to learn multi-view geometric information from different geometric attributes. Our TSGCNet adopts two graph-learning streams, designed in an input-aware fashion, to extract more discriminative high-level geometric representations from coordinates and normal vectors, respectively. These feature representations learned from the designed two different streams are further fused to integrate the multi-view complementary information for the cell-wise dense prediction task. We evaluate our proposed TSGCNet on a real-patient dataset of dental models acquired by 3D intraoral scanners, and experimental results demonstrate that our method significantly outperforms state-of-the-art methods for 3D shape segmentation.
Yue Zhao 0012, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen
CVPR4
2021 VertNet: Accurate Vertebra Localization and Identification Network from CT Images
Zhiming Cui 0001, Changjian Li 0001, Lei Yang 0048, Chunfeng Lian, Feng Shi 0001, Wenping Wang 0001, Dijia Wu, Dinggang Shen
MICCAI (5)1
2021 Domain Generalization for Mammography Detection via Multi-style and Multi-view Contrastive Learning
Zheren Li, Zhiming Cui 0001, Sheng Wang 0014, Yuji Qi, Xi Ouyang, Qitian Chen, Yuezhi Yang, Zhong Xue, Dinggang Shen, Jie-Zhi Cheng
MICCAI (7)2
2021 Motion Correction for Liver DCE-MRI with Time-Intensity Curve Constraint
Dongming Wei, Zhiming Cui 0001, Yujia Zhou 0001, Caiwen Jiang, Jiameng Liu, Qianjin Feng 0003, Dinggang Shen
MICCAI (7)3
2021 Consistent Segmentation of Longitudinal Brain MR Images with Spatio-Temporal Constrained Networks
Feng Shi 0001, Zhiming Cui 0001, Yongsheng Pan, Yong Xia 0001, Dinggang Shen
MICCAI (1)3
2021 Confidence-Aware Cascaded Network for Fetal Brain Segmentation on MR Images
Xukun Zhang, Zhiming Cui 0001, Changan Chen, Jingjiao Lou, Wenxin Hu, He Zhang 0023, Tao Zhou 0002, Feng Shi 0001, Dinggang Shen
MICCAI (3)2
2021 Self-attention implicit function networks for 3D dental data completion
Yuhan Ping, Guodong Wei, Lei Yang 0048, Zhiming Cui 0001, Wenping Wang 0001
Comput. Aided Geom. Des.4
2021 TSegNet: An efficient and accurate tooth segmentation network on 3D dental model
Zhiming Cui 0001, Changjian Li 0001, Nenglun Chen, Guodong Wei, Runnan Chen, Yuanfeng Zhou, Dinggang Shen, Wenping Wang 0001
Medical Image Anal.1
2021 Structure-Driven Unsupervised Domain Adaptation for Cross-Modality Cardiac Segmentation
abstract
Performance degradation due to domain shift remains a major challenge in medical image analysis. Unsupervised domain adaptation that transfers knowledge learned from the source domain with ground truth labels to the target domain without any annotation is the mainstream solution to resolve this issue. In this paper, we present a novel unsupervised domain adaptation framework for cross-modality cardiac segmentation, by explicitly capturing a common cardiac structure embedded across different modalities to guide cardiac segmentation. In particular, we first extract a set of 3D landmarks, in a self-supervised manner, to represent the cardiac structure of different modalities. The high-level structure information is then combined with another complementary feature, the Canny edges, to produce accurate cardiac segmentation results both in the source and target domains. We extensively evaluate our method on the MICCAI 2017 MM-WHS dataset for cardiac segmentation. The evaluation, comparison and comprehensive ablation studies demonstrate that our approach achieves satisfactory segmentation results and outperforms state-of-the-art unsupervised domain adaptation methods by a significant margin.
Zhiming Cui 0001, Changjian Li 0001, Zhixu Du, Nenglun Chen, Guodong Wei, Runnan Chen, Lei Yang 0048, Dinggang Shen, Wenping Wang 0001
IEEE Trans. Medical Imaging1
2020 Unsupervised Learning of Intrinsic Structural Representation Points
abstract
Learning structures of 3D shapes is a fundamental problem in the field of computer graphics and geometry processing. We present a simple yet interpretable unsupervised method for learning a new structural representation in the form of 3D structure points. The 3D structure points produced by our method encode the shape structure intrinsically and exhibit semantic consistency across all the shape instances with similar structures. This is a challenging goal that has not fully been achieved by other methods. Specifically, our method takes a 3D point cloud as input and encodes it as a set of local features. The local features are then passed through a novel point integration module to produce a set of 3D structure points. The chamfer distance is used as reconstruction loss to ensure the structure points lie close to the input point cloud. Extensive experiments have shown that our method outperforms the state-of-the-art on the semantic shape correspondence task and achieves comparable performance with the state-of-the-art on the segmentation label transfer task. Moreover, the PCA based shape embedding built upon consistent structure points demonstrates good performance in preserving the shape structures. Code is available at https://github.com/NolenChen/3DStructurePoints.
Nenglun Chen, Lingjie Liu, Zhiming Cui 0001, Runnan Chen, Duygu Ceylan, Changhe Tu, Wenping Wang 0001
CVPR3
2020 TANet: Towards Fully Automatic Tooth Arrangement
Guodong Wei, Zhiming Cui 0001, Nenglun Chen, Runnan Chen, Guiqing Li, Wenping Wang 0001
ECCV (15)2
2020 Mapping in a Cycle: Sinkhorn Regularized Unsupervised Learning for Point Cloud Shapes
Lei Yang 0048, Wenxi Liu, Zhiming Cui 0001, Nenglun Chen, Wenping Wang 0001
ECCV (10)3
2019 ToothNet: Automatic Tooth Instance Segmentation and Identification From Cone Beam CT Images
abstract
This paper proposes a method that uses deep convolutional neural networks to achieve automatic and accurate tooth instance segmentation and identification from CBCT (cone beam CT) images for digital dentistry. The core of our method is a two-stage network. In the first stage, an edge map is extracted from the input CBCT image to enhance image contrast along shape boundaries. Then this edge map and the input images are passed to the second stage. In the second stage, we build our network upon the 3D region proposal network (RPN) with a novel learned-similarity matrix to help efficiently remove redundant proposals, speed up training and save GPU memory. To resolve the ambiguity in the identification task, we encode teeth spatial relationships as an additional feature input in the identification task, which helps to remarkably improve the identification accuracy. Our evaluation, comparison and comprehensive ablation studies demonstrate that our method produces accurate instance segmentation and identification results automatically and outperforms the state-of-the-art approaches. To the best of our knowledge, our method is the first to use neural networks to achieve automatic tooth segmentation and identification from CBCT images.
Zhiming Cui 0001, Changjian Li 0001, Wenping Wang 0001
CVPR1