Danni Ai

dblp:78/8613 · DBLP profile ↗
← Back
67ranked-venue papers
4as first author
44since 2021 · last 2026
0000-0002-2285-0570ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MOMA: multi-expert framework with missing pattern awareness for rectal cancer neoadjuvant therapy
Yucong Lin, Kailun Fei, Bowen Liu 0011, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Hong Song 0003, Jian Yang 0003
Inf. Sci.6
2026 Multimodal hybrid mamba classification model for tumor pathological grade prediction using magnetic resonance images
Langtao Zhou, Tianyu Fu 0003, Xiaoxia Qu, Jiaoyang Wu, Yangrui Huang, Hong Song 0003, Jingfan Fan, Danni Ai, Deqiang Xiao, Junfang Xian, Jian Yang 0003
Neural Networks8
2026 Sculpting Margin Penalty: Intra-Task Adapter Merging and Classifier Calibration for Few-Shot Class-Incremental Learning
Liang Bai 0006, Hong Song 0003, Jinfu Li 0004, Yucong Lin, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.7
2026 Hierarchical Heterogeneous Aggregation Network for Multi-Shape Coronary Stenosis Detection in X-Ray Angiography Sequences
abstract
Accurate detection of multi-shape coronary artery stenoses from X-ray angiography (XRA) sequences plays a crucial role in diagnosing and planning interventions for coronary artery disease. However, vessel overlap, background noise, and nonlinear cardiac motion introduce significant challenges. These factors often result in missed detections, intra-frame class conflict, and temporal category drift, particularly for subtle and morphologically complex stenoses such as focal and bifurcation stenoses. To address these challenges, we propose a Hierarchical Heterogeneous Aggregation Network that effectively integrates both spatial and temporal cues across XRA sequences. The proposed framework incorporates a Channel Importance-guided Fusion module, which aims to enhance the representation of small-stenosis features by dynamically selecting high-importance channels across scales. Furthermore, we introduce a Hierarchical Heterogeneous Aggregator designed to reduce spatial redundancy and explicitly generate discriminative features across frames based on heterogeneous relationships, thereby improving temporal consistency and classification robustness. Existing experiments conducted on two clinical datasets indicate that our method outperforms existing detectors and stenosis methods in terms of detection accuracy and generalization.
Sigeng Chen, Jingfan Fan, Yujie Xie, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Wenyuan Yu, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.4
2026 Double-Decomposition Motion Tracking of Intraoperative 3D Structures via Cross-Spatio-Temporal Semantics Alignment
abstract
3D motion tracking in X-ray image-guided operations using pre- and intra-operative image registration has recently gained attention. However, due to pre- and intra-operative acquisitions exist spatio-temporal misalignment (i.e., limited 3D prior versus continuous 2D images) and distinct respiratory phase difference, recent methods still struggle to accurately estimate 3D dynamic structures from X-ray images. To overcome these issues, we propose a novel double-decomposition tracking (DD-Track) framework that aligns with multi-organ motion characteristics via two alignment pipes: 1) Temporal alignment aims to compensate in-plane respiratory phases difference between the projection of static 3D prior and continuous X-ray images. A dual-excitation mechanism in the image and frequency domains is proposed to extract discriminate motion features while suppressing irrelevant background information. 2) Spatial alignment subsequently integrates the extracted 2D motion features into the cross-modal registration process to accurately warp the 3D prior. Further, we decompose the motion tracking into the common trajectory and organ-specific deformation to align with the multi-organ motion nature, avoiding excessive organ stretching for sliding compensation. Comprehensive quantitative and qualitative experiments on simulated and clinical multi-organ datasets demonstrate that DD-Track outperforms state-of-the-art methods, and we also validate its generalization for tracking intra-organ lesions on simulated data.
Haixiao Geng, Jingfan Fan, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009
IEEE Trans. Medical Imaging4
2026 SG-3DGS: Sequential Growing 3D Gaussian Splatting for Scene Reconstruction of Monocular Endoscope Video
abstract
The reconstruction of monocular endoscope video scenes is essential for enhancing the application and analysis of surgical endoscopic images. However, restricted by the narrow space of endoscopic movement and the obstruction of vision within cavities, it is difficult for most conventional methods to perform high-quality reconstruction. To address these challenges, a novel dynamic growing 3D Gaussian splatting architecture is proposed to construct the 3D model of endoscopic scene without precomputed camera poses or Structure from Motion. Firstly, to establish spatial feature associations between interframes, a 2D-3D displacement fields are designed by utilizing dense feature matches and depth prediction. On this basis, a novel displacement field variational optimization is developed to obtain relative poses by minimizing the energy functional associated with field transformation. Secondly, to address the constraint of the endoscopic view, by Gaussian sequential transformation and differential gradient field optimization, a novel Sequential Gaussian Growing Module is proposed to grow the local Gaussian model sequentially. Finally, a novel Forward-Reconstruction&Backward-Optimization architecture is proposed to generate the global Gaussian model. The evaluation is conducted on two public endoscopic datasets: Scared and C3VD. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods in both quantitative metrics (PSNR, SSIM, LPIPS, ATE, RMSE, MAE) and qualitative comparisons. The project page is https://iheckzza.github.io/ DG-3DGS/.
Hong Song 0003, Jingfan Fan, Long Shao, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Yucong Lin, Jian Yang 0009
IEEE Trans. Medical Imaging6
2026 Pelvic Fracture Reduction Planning via Joint Shape-Intensity Reference
abstract
Pelvic fracture reduction planning is clinically critical yet technically demanding due to the complex anatomical structure of pelvis and the topological discontinuities introduced by fractures. Existing computer-assisted planning approaches dominantly rely on shape-based models, overlooking the rich CT intensity information that is essential for accurate and patient-specific planning. To address this limitation, we propose SIRDiff, a novel framework that incorporates anatomical shape and CT intensity information to generate biomechanically plausible reference models for pelvic fracture reduction planning. SIRDiff comprises three key components: 1) the structure-aware diffusion model to reconstruct the global anatomical structure, 2) the topology-adaptive structural conditioning strategy that maps fracture landmarks into a healthy anatomical graph domain for robust structure guidance, and 3) the detail-preserved autoencoder to ensure the fine-grained image reconstruction from latent representations. Additionally, SIRDiff adopts a multi-task learning approach to jointly predict the reference CT image and corresponding bone segmentation map, which enhances its potential for clinical application and ensures better anatomical consistency. Despite being trained exclusively on synthetic fracture data, SIRDiff shows the strong generalizability to real clinical cases and consistently outperforms existing methods across multiple clinically relevant evaluation metrics, demonstrating its potential as a robust and deployable solution for pelvic fracture reduction planning.
Xirui Zhao, Deqiang Xiao, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Junqiang Wang, Jian Yang 0009
IEEE Trans. Medical Imaging5
2026 Anatomy-Aware Sketch-Guided Latent Diffusion Model for Orbital Tumor Multi-Parametric MRI Missing Modalities Synthesis
abstract
Synthesizing missing modalities in multi-parametric MRI (mpMRI) is vital for accurate tumor diagnosis, yet remains challenging due to incomplete acquisitions and modality heterogeneity. Diffusion models have shown strong generative capability, but conventional approaches typically operate in the image domain with high memory costs and often rely solely on noise-space supervision, which limits anatomical fidelity. Latent diffusion models (LDMs) improve efficiency by performing denoising in latent space, but standard LDMs lack explicit structural priors and struggle to integrate multiple modalities effectively. To address these limitations, we propose the anatomy-aware sketch-guided latent diffusion model (ASLDM), a novel LDM-based framework designed for flexible and structure-preserving MRI synthesis. ASLDM incorporates an anatomy-aware feature fusion module, which encodes tumor region masks and edge-based anatomical sketches via cross-attention to guide the denoising process with explicit structure priors. A modality synergistic reconstruction strategy enables the joint modeling of available and missing modalities, enhancing cross-modal consistency and supporting arbitrary missing scenarios. Additionally, we introduce image-level losses for pixel-space supervision using L1 and SSIM losses, overcoming the limitations of pure noise-based loss training and improving the anatomical accuracy of synthesized outputs. Extensive experiments on a five-modality orbital tumor mpMRI private dataset and a four-modality public BraTS2024 dataset demonstrate that ASLDM outperforms state-of-the-art methods in both synthesis quality and structural consistency, showing strong potential for clinically reliable multi-modal MRI completion. Our code is publicly available at: https://github.com/zltshadow/ASLDM.git.
Langtao Zhou, Xiaoxia Qu, Tianyu Fu 0003, Jiaoyang Wu, Hong Song 0003, Jingfan Fan, Danni Ai, Deqiang Xiao, Junfang Xian, Jian Yang 0009
IEEE Trans. Medical Imaging7
2026 Enhanced CT-CBCT image registration for orthopedic surgery: Integrating rigid-elastic motion models
abstract
Computed tomography (CT) and cone-beam computed tomography (CBCT) image registration play pivotal roles in computer-assisted navigation for orthopedic surgery. Traditional methods often apply uniform deformation models, neglecting the biomechanical differences between rigid structures and soft tissues, which compromises registration accuracy, especially during significant bone displacements. To address this issue, we introduce RE-Reg, a rigid-elastic CT-CBCT image registration framework that jointly learns rigid bone motion and soft tissue deformation. RE-Reg incorporates a rigid alignment (RA) module to estimate global bone motion and an elastic deformation (ED) module to model soft tissue deformation, preserving bony structures through bone shape preservation (BSP) loss. Our comprehensive evaluation on publicly available datasets demonstrates that RE-Reg significantly outperforms existing methods in terms of registration accuracy and rigid bone structure preservation, achieving a 1.3% improvement in Dice similarity coefficient (DSC) and a 23% reduction in rigid bone deformation ( % Δ vol ) compared with the best baseline. This framework not only enhances anatomical fidelity but also ensures biomechanical plausibility and provides a valuable tool for image-guided orthopedic surgery. This code is available at https://github.com/Zq-Huang/RE-Reg.
Deqiang Xiao, Hongxun Liu, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Jian Yang 0009
Virtual Real. Intell. Hardw.5
2026 Augmented reality surgical navigation: Clinical applications, key technologies, and future directions
abstract
Surgical navigation has evolved significantly through advances in augmented reality, virtual reality, and mixed reality, improving precision and safety across many clinical applications, including neurosurgery, maxillofacial, spinal, and arthroplasty procedures. By integrating preoperative imaging with real-time intraoperative data, these systems provide dynamic guidance, reduce radiation exposure, and minimize tissue damage. Key challenges persist, including intraoperative registration accuracy, flexible tissue deformation, respiratory compensation, and real-time imaging quality. Emerging solutions include artificial intelligence-driven segmentation, deformation-field modeling, and hybrid registration techniques. Future developments will include lightweight, portable systems, improved non-rigid registration algorithms, and greater clinical adoption. Despite advances in rigid-tissue applications, soft-tissue navigation requires additional innovation to address motion variability and registration reliability, ultimately advancing minimally invasive surgery and precision medicine.
Jingfan Fan, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Long Shao, Tao Chen 0022, Hong Song 0003, Yongtian Wang, Jian Yang 0009
Virtual Real. Intell. Hardw.5
2025 Reducing Redundancy in Small Lesion Features for Multi-Shape Stenosis Detection in Coronary X-ray Angiography
abstract
The automatic detection of multi-shape coronary artery stenosis through X-ray angiography has important clinical implications for the diagnosis and treatment of coronary artery disease. However, there are several challenges in accurately detecting multi-shape stenoses, including the small size of stenoses, large size variations among multi-shape stenoses, unclear boundaries, and deformation caused by cardiac motion, such as stretching and contraction. To this end, we propose a framework for multi-shape stenosis detection at the sequence level by optimizing the representation of small lesions. Specifically, we propose a contribution-guided redundancy reduction module to suppress the feature redundancy of the background region while dynamically optimizing stenosis representations. Furthermore, to tackle the challenges of classification caused by the similarity in lesion appearance and unclear boundaries, we leverage the relationship between task-specific features and cross-enhancement to improve classification performance. Finally, a sequence relaxation strategy to extend the representation of lesions from the single-frame level to the sequence level. Experimental results indicate that the proposed method demonstrates a significant advantage over comparative methods in the task of multi-shape stenosis detection.
Sigeng Chen, Jingfan Fan, Danni Ai, Jian Yang 0009
IJCNN4
2025 DetectDiffuse: Aggregation- and Attention-Driven Universal Lesion Detection with Multi-scale Diffusion Model
Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Deqiang Xiao, Jian Yang 0009
MICCAI (5)2
2025 Temporal Modulated Multi-scale Deformation Fusion via Knowledge Distillation for 4D Medical Image Interpolation
Jiaju Zhang, Danni Ai, Zhikun Gan, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Deqiang Xiao
MICCAI (8)2
2025 Incremental energy-based recurrent transformer-KAN for time series deformation simulation of soft tissue
Jiaxi Jiang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Danni Ai, Deqiang Xiao, Yongtian Wang, Jian Yang 0009
Expert Syst. Appl.6
2025 SFCLI-Net: Spatial-frequency collaborative learning interpolation network for Computed Tomography slice synthesis
Hong Song 0003, Danni Ai, Jieliang Shi, Jingfan Fan, Deqiang Xiao, Tianyu Fu 0003, Yucong Lin, Wencan Wu, Jian Yang 0009
Expert Syst. Appl.3
2025 Collective Migration-Inspired Large-Deformation Compensation for Nonrigid Image Registration
Dingkun Liu, Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Yongtian Wang, Jian Yang 0009
Int. J. Comput. Vis.2
2025 SEMNet: a simple and efficient MLP-based network for 3D Face point clouds landmarks localization
Mingyang Lei, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Danni Ai, Jingfan Fan, Jian Yang 0009
Multim. Syst.5
2025 FCF-CSM: A Fuzzy Clustering Framework Based on Chromaticity Statistical Model for Automatic Segmentation of Port Wine Stains
abstract
Extracting lesions with accurate boundaries from clinical images is crucial for the clinical diagnosis, progression monitoring, and efficacy evaluation of port wine stains (PWS). However, accurately delineating lesion boundaries remains challenging due to complex boundary structures and unreliable annotations. In this paper, we propose a fuzzy clustering framework, FCF-CSM, for automated segmentation of PWS lesions. It combines prior knowledge of PWS color distribution with superpixels’ boundary characterization capability. Firstly, a chromaticity statistical model (CSM) is established based on 1000 collected PWS images, providing prior probabilities of PWS lesions to guide the improvement of the fuzzy clustering framework. Secondly, a superpixel method incorporating CSM is applied to PWS images, generating superpixels with improved boundary characterization. Thirdly, these superpixels are finely clustered using color features related to erythema index, color statistics, and color volume, improving PWS lesion distinction from complex backgrounds. Finally, a CSM-based automatic decision method distinguishes lesions from the background, achieving fully automated PWS segmentation within a fuzzy clustering framework. In addition, a boundary local fitting (BLF) metric is proposed to evaluate the segmentation precision of the PWS boundaries. Comparative experiments are conducted to verify the superiority of FCF-CSM. It achieves comparable overall segmentation performance with Jaccard and Dice metrics of 84.14% and 91.11%, respectively, compared to state-of-the-art methods. In terms of boundary segmentation, FCF-CSM outperforms other methods with an 81.52% BLF metric. FCF-CSM has proven to be effective for PWS segmentation and is promising to improve boundary delineation. The code is available athttps://github.com/JinrongMu/FCF-CSM. Note to Practitioners—The motivation of this study was to construct a statistical model of port wine stain (PWS) color to quantify prior knowledge of PWS lesion color in RGB images. Existing methods for automatic segmentation of PWS rely heavily on annotated data, but the lack of publicly available datasets hinders the development of such algorithms due to the privacy of clinical data. This paper proposes a fuzzy clustering framework based on the chromaticity statistical model, which can achieve high-precision and fine-grained delineation of the boundaries of PWS lesions without annotating data. In this study, we describe the prior of PWS color distribution based on colorimetry theory and apply this knowledge to the PWS automatic segmentation task, thereby realizing knowledge sharing while protecting patient privacy from being leaked. Comprehensive comparative experiments demonstrate the effectiveness and reliability of the chromaticity statistical model. However, the dataset used to build this model only includes populations with yellow skin tones. In future research, we will address the automatic identification of PWS lesions applicable to other skin tones.
Jinrong Mu, Hong Song 0003, Xianqi Meng, Jingfan Fan, Danni Ai, Defu Chen, Haixia Qiu, Jian Yang 0009
IEEE Trans Autom. Sci. Eng.7
2025 Multidomain Dependency-Aware Guided Unified-Stage Coronary Artery Branch Recognition Network
abstract
Clinical scoring in X-ray coronary angiography image sequences is widely used for revascularization decision-making in cases of coronary artery disease. Accurately recognizing coronary artery branches is a fundamental step in assessing the severity of quantitative stenosis. Existing methods employ a multistage process that includes view separation, skeletonization, graph building, and classification using topological features. However, the graph often suffers from skeleton errors, leading to incorrect topological connections during the classification stage, which requires manual correction. To address these issues, we propose a unified-stage coronary artery branch recognition network (UniCABR) that integrates the segmentation, skeletonization, and graph-building stages. Specifically, we design a dependency-aware module to build dependency graphs in both semantic and spatial domains, avoiding the use of rigid inter-branch topological connections and thus eliminating the need for manual correction of misconnections resulting from skeleton errors. Furthermore, to suppress nontarget branches according to clinical criteria and enhance the performance of side branches, we introduce a small feature supplementation module coupled with an adaptive merged binary supervision method at the pixel level. Extensive experiments on two datasets and a generalization study demonstrate the superiority of UniCABR in performance and generalization ability for coronary artery branch recognition tasks.
Sigeng Chen, Jingfan Fan, Danni Ai, Deqiang Xiao, Yucong Lin, Hong Song 0003, Wenyuan Yu, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.3
2025 Double-Shot 3D Shape Measurement With a Dual-Branch Network for Structured Light Projection Profilometry
abstract
The structured light (SL)-based three-dimensional (3D) measurement techniques with deep learning have been widely studied to improve measurement efficiency, among which fringe projection profilometry (FPP) and speckle projection profilometry (SPP) are two popular methods. However, they generally use a single projection pattern for reconstruction, resulting in fringe order ambiguity or poor reconstruction accuracy. To alleviate these problems, we propose a parallel dual-branch Convolutional Neural Network (CNN)-Transformer network (PDCNet), to take advantage of convolutional operations and self-attention mechanisms for processing different SL modalities. Within PDCNet, a Transformer branch is used to capture global perception in the fringe images, while a CNN branch is designed to collect local details in the speckle images. To fully integrate complementary features, we design a double-stream attention aggregation module (DAAM) that consists of a parallel attention subnetwork for aggregating multi-scale spatial structure information. This module can dynamically retain local and global representations to the maximum extent. Moreover, an adaptive mixture density head with bimodal Gaussian distribution is proposed for learning a representation that is precise near discontinuities. Compared to the standard disparity regression strategy, this adaptive mixture head can effectively improve performance at object boundaries. Extensive experiments demonstrate that our method can reduce fringe order ambiguity while producing high-accuracy results on self-made datasets.
Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.6
2025 Structured Light Image Planar-Topography Feature Decomposition for Generalizable 3D Shape Measurement
abstract
The application of structured light (SL) techniques has achieved remarkable success in three-dimensional (3D) measurements. Traditional methods generally calculate SL information pixel by pixel to obtain the measurement results. Recently, the rise of deep learning (DL) has led to significant developments in this task. However, existing DL-based methods generally learn all features within the image in an end-to-end manner, ignoring the distinction between SL and non-SL information. Therefore, these methods may encounter difficulties in focusing on subtle variations in SL patterns across different scenes, thereby degrading measurement precision. To overcome this challenge, we propose a novel SL Image Planar-Topography Feature Decomposition Network (SIDNet). To fully utilize the information from different SL modality images (fringe and speckle), we decompose different modalities into topography features (modality-specific) and planar features (modality-shared). A physics-driven decomposition loss is proposed to make the topography/planar features dissimilar/similar, which guides the network to distinguish between SL and non-SL information. Moreover, to obtain modality-fused features with global overview and local detail information, we propose a wrapped phase-driven feature fusion module. Specifically, a novel Tri-modality Mamba block is designed to integrate different sources with the guidance of the wrapped phase features. Extensive experiments demonstrate the superiority of our SIDNet in multiple simulated 3D measurement scenes. Moreover, our method shows better generalization ability than other DL models and can be directly applicable to unseen real-world scenes.
Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.6
2025 Landmark and Pose Prediction in Occluded Facial Point Cloud via Explicit Joint Feature Fusion Network
abstract
Facial point clouds collected in practical applications often suffer from pose variations and occlusion. Existing studies typically focus on either pose estimation or landmarks localization, neglecting to fully utilize the effective information from various facial features, thus limiting the improvement of prediction accuracy. Therefore, we propose an innovative 3D facial multi-task prediction network. The proposed network embeds the output of related tasks into feature extraction from the point level to the global level based on the physical dependencies between tasks. This facilitates explicit multi-task knowledge transfer, enabling the simultaneous prediction of facial landmarks, occlusion, and head pose. We introduce a training strategy based on posterior knowledge correction to iteratively refine and improve multi-task prediction results. Moreover, no single dataset provides annotations for all these tasks at once, so we synthesized a 3D landmarks, occlusion and pose (3D-LOP) dataset, which includes annotations for landmarks coordinates, occlusion probability, and head pose. The proposed method was compared with state-of-the-art methods on two public datasets and 3D-LOP. The landmarks localization accuracy improved by 7.1% on the two public datasets, and the pose estimation accuracy and stability on 3D-LOP improved by 28.5% and 32.7%, respectively. The performance on wild data also shows its potential in practical applications.
Jingfan Fan, Long Shao, Mingyang Lei, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Hong Song 0003, Yucong Lin, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.6
2025 Hepatic Vessel Roadmap Prediction Using Adaptive Tracking and Bending Energy Modeling in X-Ray Fluoroscopy
abstract
Dynamic visualization of the hepatic vessel is crucial in X-ray image-guided transjugular intrahepatic portosystemic shunt (TIPS) procedures. However, intraoperative breathing and the presence of guidewires complicate the prediction of the vessel position and posture without contrast agents. The respiration compensation technique aims to utilize the intraoperative respiration modeling to deform the initial vessel roadmap, thereby achieving the dynamic vessel prediction in the X-ray image sequence for the interventional guidance. Therefore, we propose a novel respiration compensation framework utilizing the adaptive tracking and bending energy modeling to achieve the stable vessel roadmap prediction under free breathing. First, we introduce the inter-frame rigid displacement compensation module based on the domain adaptation and adaptive centroid tracking. This module fits the respiratory curve from the X-ray images, providing the temporal motion priors for aligning roadmaps across frames. Second, we propose the novel deformation compensation module based on the bending energy modeling to correct the respiratory motion, wherein we utilize the energy features of the guidewires to drive the non-rigid registration. The control points sampled by the bending energy guide the local image to form the deformation field, facilitating the dynamic overlap of the vessel roadmaps in X-ray images. Experimental results on simulated and clinical datasets show an average tracking error of 0.95 $\pm$ 0.26 mm and 1.49 $\pm$ 0.40 mm, respectively. The effective and fast (mean 57 ms per frame) compensation achieved by our framework has the potential for improving the outcome of liver intervention and reducing the reliance on contrast agents.
Deqiang Xiao, Haixiao Geng, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Jian Yang 0009
IEEE J. Biomed. Health Informatics4
2024 CLIP and image integrative prompt for anterior mediastinal lesion segmentation in CT image
abstract
The automatic segmentation of anterior mediastinal lesions in enhanced CT imaging is of significant importance in clinical diagnostics. Anterior mediastinal lesions are characterized by various types and blurred boundary, which increases the difficulty of anterior mediastinal lesions segmentation. This study leverages the robust zero-shot classification capability and semantic expression of CLIP to formulate CLIP-prompt that express the semantic correlation between images and text, so that CLIP-prompt guided cross-attention has been proposed. By integrating the CLIP-prompt into the image features through cross-attention, the network can focus more intently on the lesion areas. Additionally, to better capture the unknown categorical features of the images, this paper introduces a learnable image prompt that works in conjunction with an attention module integrated with textual information, thereby enhancing the constraints on the segmentation targets. Finally, to address the blurred boundary of the anterior mediastinal lesion, this study proposes a boundary-enhanced loss. By augmenting the weights of difficult-to-segment edge points, the network is enabled to focus on these challenging boundary areas, consequently improving the segmentation accuracy of these points. Compared to existing state-of-the-art methods, our approach has achieved an overall Dice coefficient of 89.43% and has achieved good performance in terms of ASSD metric for segmentation edges.
Su Huang, Danni Ai, Guolin Ma, Jian Yang 0009
BIBM3
2024 PrixMatch: Semi-supervised Network for Multi-modal Medical Image Segmentation with Cross-modal Data Augmentation and Adaptive Prior Knowledge Thresholding
abstract
Semi-supervised medical image segmentation has made significant strides, yet most existing methods are confined to single-modality data, limiting both the volume of data and the generalizability of the models. Multi-modal data can provide richer information, expand the dataset and enhance model robustness. However, integrating multi-modal learning into semi-supervised medical image segmentation presents challenges, primarily in how to deal with the scarcity of labels and alignment across different modalities simultaneously. In this paper, we propose PrixMatch, a multi-modal semi-supervised model with a teacher-student strategy for medical image segmentation. Initially, we propose a cross-modal data augmentation strategy, which randomly exchanges image blocks of the same location between different modalities, to guide the student model to learn cross-modal consistency without the need for additional network modules. Secondly, we design a cross-modal adaptive pseudo-label threshold setting strategy, which can align the prior anatomical knowledge of different modalities, and combine the modal-aligned prior knowledge and model learning state to filter the pseudo-labels at the pixel-level, flexibly alleviating the confirmation bias that occurs during semi-supervised training. Experiments demonstrate that PrixMatch achieves a Dice Similarity Coefficient (DSC) of 87.2% on the BTCV (CT) and CHAOS (MR) multi-modal datasets with only 10% labeling ratio, bringing nearly 5.5% improvement over the latest state-of-the-art method.
Hong Song 0003, Yucong Lin, Long Shao, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009
BIBM7
2024 Deformation Correction in Laparoscopic Liver Surgical Navigation Using Point Cloud Completion and Biomechanical Model
abstract
In the minimally invasive liver resection surgery, deformation estimation of liver is required to correct the preoperative virtual model to match the intraoperative scenarios, in which the liver deforms due to respiration and surgical operations. Existing methods on liver deformation estimation often struggle to achieve high accuracy when the intraoperative liver surface is limited in size. To overcome the challenge of sparse intraoperative point cloud data and improve the accuracy of liver deformation predictions, this paper introduces an innovative method for estimating liver deformation. This method comprises two main components: intraoperative point cloud completion and liver deformation estimation. Intraoperative point cloud completion uses registration techniques to integrate preoperative topological structures into the intraoperative phase. Liver deformation estimation combines optimization control with biomechanical modeling to accurately align the preoperative liver model with its intraoperative counterpart. Comparative and ablation experiments, as well as investigations into the impact of different completion ratios, were conducted. The results demonstrate that this method effectively utilizes preoperative liver geometric features to enhance intraoperative visualization, even with limited intraoperative data. Additionally, the opti-mization control method provides reliable deformation estimates with acceptable accuracy. This study offers new insights and methodologies for the development of augmented reality surgical navigation systems, contributing to the computer assisted liver surgey.
Deqiang Xiao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Jian Yang 0009
BIBM3
2024 N-Gram Swin Transformer for CT Image Super-Resolution
Zhenghao Gao, Danni Ai, Hong Song 0003, Jian Yang 0009
ICXR2
2024 STQD-Det: Spatio-Temporal Quantum Diffusion Model for Real-Time Coronary Stenosis Detection in X-Ray Angiography
abstract
Detecting coronary stenosis accurately in X-ray angiography (XRA) is important for diagnosing and treating coronary artery disease (CAD). However, challenges arise from factors like breathing and heart motion, poor imaging quality, and the complex vascular structures, making it difficult to identify stenosis fast and precisely. In this study, we proposed a Quantum Diffusion Model with Spatio-Temporal Feature Sharing to Real-time detect Stenosis (STQD-Det). Our framework consists of two modules: Sequential Quantum Noise Boxes module and spatio-temporal feature module. To evaluate the effectiveness of the method, we conducted a 4-fold cross-validation using a dataset consisting of 233 XRA sequences. Our approach achieved the F1 score of 92.39% with a real-time processing speed of 25.08 frames per second. These results outperform 17 state-of-the-art methods. The experimental results show that the proposed method can accomplish the stenosis detection quickly and accurately.
Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Jian Yang 0009
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Segmentation of 3D Anatomically Diffused Tissues in Magnetic Resonance Images Through Edge-Preserving Constrained Center-Free Fuzzy $C$-Means
abstract
Anatomically diffused tissues (ADTs) refer to soft tissues containing many anatomical regions that are spatially dispersed and structurally irregular. In magnetic resonance images, ADTs exhibit blurred morphology and heterogeneous texture, making the accurate extraction of their 3D anatomy challenging. Center-free fuzzy C-means (FCM) can effectively partition nonlinear or nonspherical clusters, providing a promising scheme for ADT segmentation. It solves the uncertainty arising from unreliable center estimation by introducing a similarity criterion. However, the similarity criterion is sensitive to the number of target objects and their adjacent members in the images. Moreover, memberships of the existing algorithms are susceptible to losing real ADT details. To handle these issues, we propose an edge-preserving constrained center-free FCM algorithm for segmenting 3D ADTs in magnetic resonance images. To overcome the sensitivity of the similarity criterion, a novel object-to-cluster similarity measure is first proposed to utilize refined member-toobject adjacency. Specifically, the similarity measure focuses on members in the feature space, which share approximately homogeneous characteristics with each target object. Gradient-domain edge-preserving filtering is then combined with the improved similarity criterion to construct the novel objective function of center-free FCM. With the assistance of the designed imagedriven edge-preserving regularization, the gradient information of clusters is constrained, eventually approaching that of ADTs in the guidance image. Experiments are conducted on two public brain datasets and one local intrahepatic vein dataset. The results demonstrate that the proposed algorithm is more effective for ADT segmentation than the state-of-the-art peers, exhibiting superior generalization capability.
Qing Guo 0008, Hong Song 0003, Cong Wang 0033, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009
IEEE Trans. Fuzzy Syst.5
2024 Bi-Fusion of Structure and Deformation at Multi-Scale for Joint Segmentation and Registration
abstract
Medical image segmentation and registration are two fundamental and highly related tasks. However, current works focus on the mutual promotion between the two at the loss function level, ignoring the feature information generated by the encoder-decoder network during the task-specific feature mapping process and the potential inter-task feature relationship. This paper proposes a unified multi-task joint learning framework based on bi-fusion of structure and deformation at multi-scale, called BFM-Net, which simultaneously achieves the segmentation results and deformation field in a single-step estimation. BFM-Net consists of a segmentation subnetwork (SegNet), a registration subnetwork (RegNet), and the multi-task connection module (MTC). The MTC module is used to transfer the latent feature representation between segmentation and registration at multi-scale and link different tasks at the network architecture level, including the spatial attention fusion module (SAF), the multi-scale spatial attention fusion module (MSAF) and the velocity field fusion module (VFF). Extensive experiments on MR, CT and ultrasound images demonstrate the effectiveness of our approach. The MTC module can increase the Dice scores of segmentation and registration by 3.2%, 1.6%, 2.2%, and 6.2%, 4.5%, 3.0%, respectively. Compared with six state-of-the-art algorithms for segmentation and registration, BFM-Net can achieve superior performance in various modal images, fully demonstrating its effectiveness and generalization.
Jiaju Zhang, Tianyu Fu 0003, Deqiang Xiao, Jingfan Fan, Hong Song 0003, Danni Ai, Jian Yang 0009
IEEE Trans. Image Process.6
2024 Cross-Anatomy Transfer Learning via Shape-Aware Adaptive Fine-Tuning for 3D Vessel Segmentation
abstract
Deep learning methods have recently achieved remarkable performance in vessel segmentation applications, yet require numerous labor-intensive labeled data. To alleviate the requirement of manual annotation, transfer learning methods can potentially be used to acquire the related knowledge of tubular structures from public large-scale labeled vessel datasets for target vessel segmentation in other anatomic sites of the human body. However, the cross-anatomy domain shift is a challenging task due to the formidable discrepancy among various vessel structures in different anatomies, resulting in the limited performance of transfer learning. Therefore, we propose a cross-anatomy transfer learning framework for 3D vessel segmentation, which first generates a pre-trained model on a public hepatic vessel dataset and then adaptively fine-tunes our target segmentation network initialized from the model for segmentation of other anatomic vessels. In the framework, the adaptive fine-tuning strategy is presented to dynamically decide on the frozen or fine-tuned filters of the target network for each input sample with a proxy network. Moreover, we develop a Gaussian-based signed distance map that explicitly encodes vessel-specific shape context. The prediction of the map is added as an auxiliary task in the segmentation network to capture geometry-aware knowledge in the fine-tuning. We demonstrate the effectiveness of our method through extensive experiments on two small-scale datasets of coronary artery and brain vessel. The results indicate the proposed method effectively overcomes the discrepancy of cross-anatomy domain shift to achieve accurate vessel segmentation for these two datasets.
Danni Ai, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009
IEEE J. Biomed. Health Informatics2
2024 Local Contractive Registration With Biomechanical Model: Assessing Microwave Ablation After Compensation for Tissue Shrinkage
abstract
Microwave ablation (MWA) is a minimally invasive procedure for the treatment of liver tumor. Accumulating clinical evidence has considered the minimal ablative margin (MAM) as a significant predictor of local tumor progression (LTP). In clinical practice, MAM assessment is typically carried out through image registration of pre- and post-MWA images. However, this process faces two main challenges: non-homologous match between tumor and coagulation with inconsistent image appearance, and tissue shrinkage caused by thermal dehydration. These challenges result in low precision when using traditional registration methods for MAM assessment. In this paper, we present a local contractive nonrigid registration method using a biomechanical model (LC-BM) to address these challenges and precisely assess the MAM. The LC-BM contains two consecutive parts: (1) local contractive decomposition (LC-part), which reduces the incorrect match between the tumor and coagulation and quantifies the shrinkage in the external coagulation region, and (2) biomechanical model constraint (BM-part), which compensates for the shrinkage in the internal coagulation region. After quantifying and compensating for tissue shrinkage, the warped tumor is overlaid on the coagulation, and then the MAM is assessed. We evaluated the method using prospectively collected data from 36 patients with 47 liver tumors, comparing LC-BM with 11 state-of-the-art methods. LTP was diagnosed through contrast-enhanced MR follow-up images, serving as the ground truth for tumor recurrence. LC-BM achieved the highest accuracy (97.9%) in predicting LTP, outperforming other methods. Therefore, our proposed method holds significant potential to improve MAM assessment in MWA surgeries.
Dingkun Liu, Danni Ai, Tianyu Fu 0003, Yuanjin Gao, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009
IEEE J. Biomed. Health Informatics2
2024 DSC-Recon: Dual-Stage Complementary 4-D Organ Reconstruction From X-Ray Image Sequence for Intraoperative Fusion
abstract
Accurately reconstructing 4D critical organs contributes to the visual guidance in X-ray image-guided interventional operation. Current methods estimate intraoperative dynamic meshes by refining a static initial organ mesh from the semantic information in the single-frame X-ray images. However, these methods fall short of reconstructing an accurate and smooth organ sequence due to the distinct respiratory patterns between the initial mesh and X-ray image. To overcome this limitation, we propose a novel dual-stage complementary 4D organ reconstruction (DSC-Recon) model for recovering dynamic organ meshes by utilizing the preoperative and intraoperative data with different respiratory patterns. DSC-Recon is structured as a dual-stage framework: 1) The first stage focuses on addressing a flexible interpolation network applicable to multiple respiratory patterns, which could generate dynamic shape sequences between any pair of preoperative 3D meshes segmented from CT scans. 2) In the second stage, we present a deformation network to take the generated dynamic shape sequence as the initial prior and explore the discriminate feature (i.e., target organ areas and meaningful motion information) in the intraoperative X-ray images, predicting the deformed mesh by introducing a designed feature mapping pipeline integrated into the initialized shape refinement process. Experiments on simulated and clinical datasets demonstrate the superiority of our method over state-of-the-art methods in both quantitative and qualitative aspects.
Haixiao Geng, Jingfan Fan, Sigeng Chen, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009
IEEE Trans. Medical Imaging6
2023 IDAA-NET: An Image Domain Adaptive Alignment Network for Unsupervised Liver Vessel Segmentation from CTA Images
abstract
Accurate segmentation of liver vessel from CTA image is important for the diagnosis and treatment of liver diseases. The quality of labeled data directly affects the prediction results of the segmentation model. Compared with CTA image, MRA image has clearer 3D vasculature. Therefore, in order to reduce the reliance of the labeled CTA image which may contain ambiguous vessel contours, we propose a novel unsupervised liver vessel segmentation method based on image domain adaptive alignment network (IDAA-Net) by using labeled MRA and unlabeled CTA images. The IDAA-Net mainly contains three modules: 1) A spatial alignment module (SAM) is introduced to convert MRA image slice to synthetic CTA image slice for achieving spatial alignment of the different modality data in the feature and image levels; 2) An artifact removal module (ARM) is designed to eliminate background artifacts of synthetic CTA from SAM by using the liver label in MRA; 3) An adversarial segmentation module (ASM) is proposed to obtain the optimal segmentation by jointly adversarial learning and supervised learning between the predicted segmentation and the ground-truth label of MRA image. Experiments on the public and private datasets show that our method achieves comparable performance with state-of-the-art supervised method and outperforms the existing unsupervised segmentation methods.
Haixiao Geng, Danni Ai, Jingfan Fan, Feng Duan 0001, Yujia Yuan, Jian Yang 0009
BIBM2
2023 CPSS-Net: A Cross-pseudo Semi-supervised Network for Liver Vessel Segmentation from CTA Images
abstract
Accurate segmentation of liver vessel from CTA images is often challenging due to the limited availability of labeled data. In the field of medical image segmentation, semi-supervised learning has garnered significant attention as it utilizes unlabeled data to enhance the training of segmentation models. In this paper, we propose a novel cross-pseudo semi-supervised network (CPSS-Net) based on nnU-Net. The CPSS-Net contains three innovative components: 1) An probability prediction (PP) module is designed to generate probability maps for both labeled and unlabeled datasets, capturing model uncertainty through parallel nnU-Net; 2) A double pseudo-label (DPL) module is used to convert the predicted probability maps into double soft pseudo-labels using an adaptive sharpening function; 3) A cross pseudo-supervised (CPS) module is introduced to learn the mutual consistency of double pseudo-labels. Test experiments on both public and private datasets show that our method achieved a Dice score of 0.67 and a sensitivity score of 0.69, surpassing the segmentation accuracy of existing related methods.
Danni Ai, Deqiang Xiao, Feng Duan 0001, Yujia Yuan, Jian Yang 0009
BIBM2
2023 Densely Connected U-Net With Criss-Cross Attention for Automatic Liver Tumor Segmentation in CT Images
abstract
Automatic liver tumor segmentation plays a key role in radiation therapy of hepatocellular carcinoma. In this paper, we propose a novel densely connected U-Net model with criss-cross attention (CC-DenseUNet) to segment liver tumors in computed tomography (CT) images. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intra-slice features of liver tumors. Moreover, the criss-cross attention is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver tumors. We evaluated the proposed CC-DenseUNet on the LiTS dataset and the 3DIRCADb dataset. Experimental results show that the proposed method reaches the state-of-the-art performance for liver tumor segmentation. We further experimentally demonstrate the robustness of the proposed method on a clinical dataset comprising 20 CT volumes.
Qiang Li 0049, Hong Song 0003, Zenghui Wei, Fengbo Yang, Jingfan Fan, Danni Ai, Yucong Lin, Xiaoling Yu, Jian Yang 0009
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 M-CSAFN: Multi-Color Space Adaptive Fusion Network for Automated Port-Wine Stains Segmentation
abstract
Automatic segmentation of port-wine stains (PWS) from clinical images is critical for accurate diagnosis and objective assessment of PWS. However, this is a challenging task due to the color heterogeneity, low contrast, and indistinguishable appearance of PWS lesions. To address such challenges, we propose a novel multi-color space adaptive fusion network (M-CSAFN) for PWS segmentation. First, a multi-branch detection model is constructed based on six typical color spaces, which utilizes rich color texture information to highlight the difference between lesions and surrounding tissues. Second, an adaptive fusion strategy is used to fuse complementary predictions, which address the significant differences within the lesions caused by color heterogeneity. Third, a structural similarity loss with color information is proposed to measure the detail error between predicted lesions and truth lesions. Additionally, a PWS clinical dataset consisting of 1413 image pairs was established for the development and evaluation of PWS segmentation algorithms. To verify the effectiveness and superiority of the proposed method, we compared it with other state-of-the-art methods on our collected dataset and four publicly available skin lesion datasets (ISIC 2016, ISIC 2017, ISIC 2018, and PH2). The experimental results show that our method achieves remarkable performance in comparison with other state-of-the-art methods on our collected dataset, achieving 92.29% and 86.14% on Dice and Jaccard metrics, respectively. Comparative experiments on other datasets also confirmed the reliability and potential capability of M-CSAFN in skin lesion segmentation.
Jinrong Mu, Yucong Lin, Xianqi Meng, Jingfan Fan, Danni Ai, Defu Chen, Haixia Qiu, Jian Yang 0009
IEEE J. Biomed. Health Informatics5
2022 Volume-awareness and outlier-suppression co-training for weakly-supervised MRI breast mass segmentation with partial annotations
Xianqi Meng, Jingfan Fan, Jinrong Mu, Zongyu Li, Aocai Yang, Kuan Lv, Danni Ai, Yucong Lin, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009
Knowl. Based Syst.9
2022 Portal Vein and Hepatic Vein Segmentation in Multi-Phase MR Images Using Flow-Guided Change Detection
abstract
Segmenting portal vein (PV) and hepatic vein (HV) from magnetic resonance imaging (MRI) scans is important for hepatic tumor surgery. Compared with single phase-based methods, multiple phases-based methods have better scalability in distinguishing HV and PV by exploiting multi-phase information. However, these methods just coarsely extract HV and PV from different phase images. In this paper, we propose a unified framework to automatically and robustly segment 3D HV and PV from multi-phase MR images, which considers both the change and appearance caused by the vascular flow event to improve segmentation performance. Firstly, inspired by change detection, flow-guided change detection (FGCD) is designed to detect the changed voxels related to hepatic venous flow by generating hepatic venous phase map and clustering the map. The FGCD uniformly deals with HV and PV clustering by the proposed shared clustering, thus making the appearance correlated with portal venous flow robustly delineate without increasing framework complexity. Then, to refine vascular segmentation results produced by both HV and PV clustering, interclass decision making (IDM) is proposed by combining the overlapping region discrimination and neighborhood direction consistency. Finally, our framework is evaluated on multi-phase clinical MR images of the public dataset (TCGA) and local hospital dataset. The quantitative and qualitative evaluations show that our framework outperforms the existing methods.
Qing Guo 0008, Hong Song 0003, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009
IEEE Trans. Image Process.4
2021 CC-DenseUNet: Densely Connected U-Net with Criss-Cross Attention for Liver and Tumor Segmentation in CT Volumes
abstract
The automatic segmentation of liver and tumor is important for hepatic tumor surgery. In this paper, we propose a novel densely connected U-Net (CC-DenseUNet), which integrates criss-cross attention (CCA) module, to segment the liver and tumor in computed tomography (CT) volumes. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intraslice features of liver and tumors. Moreover, the CCA module is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver or tumors. We evaluated the proposed CCDenseUNet on the Liver Tumor Segmentation Challenge and 3DIRCADb datasets. Experimental results show that our method outperformed the state-of-the-art methods in liver tumor segmentation and achieved a highly competitive performance in liver segmentation.
Qiang Li 0049, Hong Song 0003, Jingfan Fan, Danni Ai, Yucong Lin, Jian Yang 0009
BIBM5
2021 Cross-Domain Transfer Learning for Vessel Segmentation in Computed Tomographic Coronary Angiographic Images
Ruirui An, Danni Ai, Yongtian Wang, Jian Yang 0009
ICIG (2)4
2021 Novel Augmented Reality System for Oral and Maxillofacial Surgery
Lele Ding, Long Shao, Zehua Zhao, Tao Zhang 0152, Danni Ai, Jian Yang 0009, Yongtian Wang
ICIG (2)5
2021 Divergence-Free Fitting-Based Incompressible Deformation Quantification of Liver
abstract
Liver is an incompressible organ that maintains its volume during the respiration-induced deformation. Quantifying this deformation with the incompressible constraint is significant for liver tracking. The constraint can be accomplished with retaining the divergence-free field obtained by the deformation decomposition. However, the decomposition process is time-consuming, and the removal of non-divergence-free field weakens the deformation. In this study, a divergence-free fitting-based registration method is proposed to quantify the incompressible deformation rapidly and accurately. First, the deformation to be estimated is mapped to the velocity in a diffeomorphic space. Then, this velocity is decomposed by a fast Fourier-based Hodge-Helmholtz decomposition to obtain the divergence-free, curl-free, and harmonic fields. The curl-free field is replaced and fitted by the obtained harmonic field with a translation field to generate a new divergence-free velocity. By optimizing this velocity, the final incompressible deformation is obtained. Moreover, a deep learning framework (DLF) is constructed to accelerate the incompressible deformation quantification. An incompressible respiratory motion model is built for the DLF by using the proposed registration method and is then used to augment the training data. An encoder-decoder network is introduced to learn appearance-velocity correlation at patch scale. In the experiment, we compare the proposed registration with three state-of-the-art methods. The results show that the proposed method can accurately achieve the incompressible registration of liver with a mean liver overlap ratio of 95.33%. Moreover, the time consumed by DLF is nearly 15 times shorter than that by other methods.
Tianyu Fu 0003, Jingfan Fan, Dingkun Liu, Hong Song 0003, Chaoyi Zhang, Danni Ai, Zhigang Cheng, Jian Yang 0009
IEEE J. Biomed. Health Informatics6
2021 Topological distance-constrained feature descriptor learning model for vessel matching in coronary angiographies
abstract
Feature matching technology is vital to establish the association between virtual and real objects in virtual reality and augmented reality systems. Specifically, it provides them with the ability to match a dynamic scene. Many image matching methods, of which most are deep learning-based, have been proposed over the past few decades. However, vessel fracture, stenosis, artifacts, high background noise, and uneven vessel gray-scale make vessel matching in coronary angiography extremely difficult. Traditional matching methods perform poorly in this regard. In this study, a topological distance-constrained feature descriptor learning model is proposed. This model regards the topology of the vasculature as the connection relationship of the centerline. The topological distance combines the geodesic distance between the input patches and constrains the descriptor network by maximizing the feature difference between connected and unconnected patches to obtain more useful potential feature relationships. Matching patches of different sequences of angiographic images are generated for the experiments. The matching accuracy and stability of the proposed method is superior to those of the existing models. The proposed method solves the problem of matching coronary angiographies by generating a topological distance-constrained feature descriptor.
Xiaojiao Song, Jingfan Fan, Danni Ai, Jian Yang 0009
Virtual Real. Intell. Hardw.4
2020 A General Endoscopic Image Enhancement Method Based on Pre-trained Generative Adversarial Networks
abstract
Endoscopic images frequently have image quality problems due to the limitations of surgical instruments and the impact of surgical operations, such as uneven illumination, smogginess and color deviation. For deep learning based on enhancement methods, independent training lacks sufficient defect images and generalization capability, and combined training with mixture of data cannot identify diverse specific tasks. To address these issues, we propose a general method based on pre-trained generative adversarial network with a specified transfer learning strategy to obtain high-quality images. Initially, we independently train a standard network based on a universal task, e.g., uneven illumination, where a pre-trained model is extracted as a backbone with partially shared generator. Then, we transfer the backbone to more potential image enhancement tasks. Experiments on uneven illumination, smogginess, and color deviation indicate that the model successfully shares common features of high-quality images and responds specifically to different defects as well.
Jingfan Fan, Danni Ai, Hong Song 0003, Yongtian Wang, Jian Yang 0009
BIBM3
2020 Local Contractive Registration for Quantification of Tissue Shrinkage in Assessment of Microwave Ablation
Dingkun Liu, Tianyu Fu 0003, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009
MICCAI (3)3
2020 Prior information constrained alternating direction method of multipliers for longitudinal compressive sensing MR imaging
Ruirui Kang, Danni Ai, Gangrong Qu, Qingbo Li, Yurong Jiang, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
Neurocomputing2
2020 Groupwise registration with global-local graph shrinkage in atlas construction
Tianyu Fu 0003, Jian Yang 0009, Danni Ai, Hong Song 0003, Yurong Jiang, Yongtian Wang, Alejandro F. Frangi
Medical Image Anal.4
2020 Spatial probabilistic distribution map-based two-channel 3D U-net for visual pathway segmentation
Danni Ai, Zhiqi Zhao, Jingfan Fan, Hong Song 0003, Xiaoxia Qu, Junfang Xian, Jian Yang 0009
Pattern Recognit. Lett.1
2020 Topology Optimization Using Multiple-Possibility Fusion for Vasculature Extraction
abstract
Vascular centerline extraction from angiography images plays an important role in computer-aided diagnosis of vascular disease. To solve the common problems related to noise and inconsistent vasculatures from uneven perfusion, this paper proposes an automatic framework for accurate vascular centerline extraction from angiograms that uses multi-probability fusion-based topology optimization. In this framework, vascular region is first segmented using a learning-based method. Then, initial centerlines are obtained by applying iterative filtering operation and multi-direction indexed non-maximum suppression. Topology optimization is achieved by gap filling. A connection probability map is constructed utilizing the information of initial centerlines, texture, and orientation of vasculatures. Shortest path tracking is employed to search for optimal connections around gaps in the initial centerlines. The proposed framework is evaluated using simulative and clinical coronary angiographies. The experimental results demonstrate that the proposed method can extract centerlines with F1 score of 97.28% ± 1.2% for vasculatures in 12 clinical angiographic images. It is evident that the proposed method can extract complete and accurate vascular centerlines from angiograms and can be used to repair gaps in other filamentary structures, such as roads and retinal blood vessels. This endows our method a great potential in the analysis of filamentary structures.
Huihui Fang, Danni Ai, Weijian Cong, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.2
2020 Greedy Soft Matching for Vascular Tracking of Coronary Angiographic Image Sequences
abstract
Vascular tracking of coronary angiographic image sequences is one of the most clinically important tasks in the diagnostic assessment and interventional guidance of cardiac disease. It is difficult to automate this application because the vascular structure is complex; moreover, unsatisfactory angiography image quality may exacerbate the difficulty of vasculature extraction. This paper converts vascular tracking into branch matching and proposes a novel and automatic greedy soft match algorithm. Our method is based on a graph framework. A graph model building module is proposed to represent the vascular structure. Then, a greedy branch searching method is adopted to acquire all possible paths in the graph that may match the reference vessel. Finally, a soft batch matching method that combines branch descriptor and dynamic time warping is presented to select the best matching branch. The solution to the problem takes advantage of both spatial and temporal continuity between successive frames. The experimental results demonstrate that the proposed algorithm is effective and robust for vascular tracking. The F1 score of a single branch dataset, which contains 12 angiographic image sequences with 77 angiograms of contrast agent-filled vessels, is 0.89 ± 0.06 and of a vessel tree dataset which contains nine sequences with 58 angiograms is 0.88 ± 0.05. Extensive experimental results well demonstrate the superior performance of the algorithm. In addition, it provides a universal solution to address the problem of filamentary structure tracking.
Huihui Fang, Danni Ai, Yong Huang 0002, Yurong Jiang, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.3
2020 Spatio-Temporal Constrained Online Layer Separation for Vascular Enhancement in X-Ray Angiographic Image Sequence
abstract
Automatic vascular enhancement is crucial to vascular structure identification in X-ray angiographic (XRA) image sequences. In this work, we propose a novel spatio-temporal constrained online layer separation (STOLS) method to achieve vascular enhancement in XRA image sequences. The proposed method integrates the motion consistency of structures into the temporal-constrained online robust principal component analysis (ORPCA) to remove quasi-static structures (e.g., bones) from the enhanced vascular images. Furthermore, smoothing technique is integrated into the spatial-constrained ORPCA to reduce motion artifacts and the noise introduced by non-uniform illumination. To make the proposed method more adaptive to various vascular structures, the spatial-constrained ORPCA is adjusted by an adaptive weight using the proportion of the vessel region in the previous frame. The performance of the proposed method is compared with five state-of-the-art subtraction methods with respect to local and global revised contrast-to-noise ratios (rCNRs) and reconstruction errors. For the proposed method, the local and global rCNRs of the final vessel layer reached 2.54 and 1.24, respectively, while the error between the original and reconstructed images from the respiratory, background, and vessel layer reached 0.0354. The proposed STOLS can enhance the angiograms in a real-time and online manner without fine-tuning parameters, and can thus be used for intra-operation diagnosis and interventional procedures of coronary artery diseases.
Shuang Song 0005, Chenbing Du, Danni Ai, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.3
2019 Monte Carlo Tree Search for 3D/2D Registration of Vessel Graphs
abstract
3D/2D registration techniques can compensate for the deficiencies of X-ray angiography-based navigation in vascular interventional surgery, such as the lack of depth information and excessive use of contrast agents. In this study, we propose a novel Monte Carlo tree search-based 3D/2D vessel graph registration method. The registration problem is transferred to a tree search problem according to the topology of vessel centerlines. Then, the Monte Carlo tree search method is applied to find the optimal vessel matching associated with highest registration score. Experiments on uninitialized vessel data demonstrate that the proposed method can achieve the highest accuracy among four state-of-the-art methods. An average accuracy of 1.91 mm on clinical coronary artery data is obtained. For the independence of initial pose and robustness to noise, the proposed method can align 3D and 2D vessels without prior initialization in vascular interventional surgery.
Shuang Song 0005, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009
BIBM4
2019 Spatial Probabilistic Distribution Map Based 3D FCN for Visual Pathway Segmentation
Zhiqi Zhao, Danni Ai, Jingfan Fan, Hong Song 0003, Yongtian Wang, Jian Yang 0009
ICIG (2)2
2019 Patch-Based Adaptive Background Subtraction for Vascular Enhancement in X-Ray Cineangiograms
abstract
OBJECTIVE: Automatic vascular enhancement in X-ray cineangiography is of crucial interest, for instance, for better visualizing and quantifying coronary arteries in diagnostic and interventional procedures. METHODS: A novel patch-based adaptive background subtraction method (PABSM) is proposed automatically enhancing vessels in coronary X-ray cineangiography. First, pixels in the cineangiogram are described by the vesselness and Gabor features. Second, a classifier is utilized to separate the cineangiogram into the rough vascular and non-vascular region. Dilation is applied to the classified binary image to include more vascular region. Third, a patch-based background synthesis is utilized to fill the removed vascular region. RESULTS: A database containing 320 cineangiograms of 175 patients was collected, and then an interventional cardiologist annotated all vascular structures. The performance of PABSM is compared with six state-of-the-art vascular enhancement methods regarding the precision-recall curve and C-value. The area under the precision-recall curve is 0.7133, and the C-value is 0.9659. CONCLUSION: PABSM can automatically enhance the coronary artery in the cineangiograms. It preserves the integrity of vascular topological structures, particularly in complex vascular regions, and removes noise caused by the non-uniform gray-level distribution in the cineangiogram. SIGNIFICANCE: PABSM can avoid the motion artifacts and it eases the subsequent vascular segmentation, which is crucial for the diagnosis and interventional procedures of coronary artery diseases.
Shuang Song 0005, Alejandro F. Frangi, Jian Yang 0009, Danni Ai, Chenbing Du, Yong Huang 0002, Hong Song 0003, Luosha Zhang, Yechen Han, Yongtian Wang
IEEE J. Biomed. Health Informatics4
2018 Inter/Intra-Constraints Optimization for Fast Vessel Enhancement in X-ray Angiographic Image Sequence
Chenbing Du, Shuang Song 0005, Danni Ai, Hong Song 0003, Yong Huang 0002, Yongtian Wang, Jian Yang 0009
BIBM3
2018 Local statistical deformation models for deformable image registration
Songyuan Tang, Weijian Cong, Jian Yang 0009, Tianyu Fu 0003, Hong Song 0003, Danni Ai, Yongtian Wang
Neurocomputing6
2017 Registration and fusion quantification of augmented reality based nasal endoscopic surgery
Yakui Chu, Jian Yang 0009, Shaodong Ma, Danni Ai, Hong Song 0003, Duanduan Chen, Lei Chen 0073, Yongtian Wang
Medical Image Anal.4
2017 Convex Hull Aided Registration Method (CHARM)
abstract
Non-rigid registration finds many applications such as photogrammetry, motion tracking, model retrieval, and object recognition. In this paper we propose a novel convex hull aided registration method (CHARM) to match two point sets subject to a non-rigid transformation. First, two convex hulls are extracted from the source and target respectively. Then, all points of the point sets are projected onto the reference plane through each triangular facet of the hulls. From these projections, invariant features are extracted and matched optimally. The matched feature point pairs are mapped back onto the triangular facets of the convex hulls to remove outliers that are outside any relevant triangular facet. The rigid transformation from the source to the target is robustly estimated by the random sample consensus (RANSAC) scheme through minimizing the distance between the matched feature point pairs. Finally, these feature points are utilized as the control points to achieve non-rigid deformation in the form of thin-plate spline of the entire source point set towards the target one. The experimental results based on both synthetic and real data show that the proposed algorithm outperforms several state-of-the-art ones with respect to sampling, rotational angle, and data noise. In addition, the proposed CHARM algorithm also shows higher computational efficiency compared to these methods.
Jingfan Fan, Jian Yang 0009, Yitian Zhao, Danni Ai, Yonghuai Liu, Ge Wang 0001, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.4
2016 3-Points Convex Hull Matching (3PCHM) for fast and robust point set registration
Jingfan Fan, Jian Yang 0009, Feng Lu 0005, Danni Ai, Yitian Zhao, Yongtian Wang
Neurocomputing4
2016 Shape context and projection geometry constrained vasculature matching for 3D reconstruction of coronary artery
Ruoxiu Xiao, Jian Yang 0009, Jingfan Fan, Danni Ai, Guangzhi Wang, Yongtian Wang
Neurocomputing4
2016 Local statistics and non-local mean filter for speckle noise reduction in medical ultrasound image
Jian Yang 0009, Jingfan Fan, Danni Ai, Xuehu Wang, Yongchang Zheng, Songyuan Tang, Yongtian Wang
Neurocomputing3
2016 Convex hull indexed Gaussian mixture model (CH-GMM) for 3D point set registration
Jingfan Fan, Jian Yang 0009, Danni Ai, Likun Xia, Yitian Zhao, Yongtian Wang
Pattern Recognit.3
2013 Generalized N-dimensional independent component analysis and its application to multiple feature selection and fusion for image classification
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001
Neurocomputing1
2012 Multiple feature selection and fusion based on generalized N-dimensional independent component analysis
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001
ICPR1
2011 Analysis of cypriot icon faces using ICA-enhanced active shape model representation
abstract
Religious iconography is an integral component of the cultural heritage of Cyprus, which was once a part of the great Byzantine empire. On one hand, icons exhibit strict adherence to conventional symbols, poses and apparel. On the other hand, there is a great variety in the style of depiction that can be attributed to different schools and periods. This paper proposes an active shape model (ASM) based technique for icon face representation that can be used for style comparison and attribution. For centuries-old icons suffering from loss of paint, cracks and added noise from digitization artifacts, we apply an independent component analysis (ICA) technique to enhance the paintings' original work. The experimental results show that our method can effectively characterize Cypriot icons.
Guifang Duan, Neela Sawant, James Z. Wang 0001, Dean R. Snow, Danni Ai, Yen-Wei Chen 0001
ACM Multimedia5
2010 Adaptive Color Independent Components Based SIFT Descriptors for Image Classification
abstract
This paper proposes an adaptive color independent components based SIFT descriptor (termed CIC-SIFT) for image classification. Our motivation is to seek an adaptive and efficient color space for color SIFT feature extraction. Our work has two key contributions. First, based on independent component analysis (ICA), an adaptive and efficient color space is proposed for color image representation. Second, in this ICA-based color space, a discriminative CIC-SIFT descriptor is calculated for image classification. The experiment results indicate that (1) contrast between objects and background can be enhanced on the ICA-based color space and (2) the CIC-SIFT descriptor outperforms other conventional color SIFT descriptors on image classification.
Danni Ai, Xianhua Han, Xiang Ruan, Yen-Wei Chen 0001
ICPR1