EDBT 2026 Demo / reviewers in the wild / expert
Hong Song 0003
dblp:09/4495-3
· DBLP profile ↗
78ranked-venue papers
4as first author
56since 2021 · last 2026
0000-0002-3171-2604ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 16 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedKit: Multi-level feature distillation with knowledge injection for radiology report generation
Zhaoli Su, Hong Song 0003, Yucong Lin, Xutao Weng, Zhongxuan Mao, Bowen Liu 0011, Hongxia Yin, Jian Yang 0009 |
Expert Syst. Appl. | 2 |
| 2026 | MOMA: multi-expert framework with missing pattern awareness for rectal cancer neoadjuvant therapy
Yucong Lin, Kailun Fei, Bowen Liu 0011, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Hong Song 0003, Jian Yang 0003 |
Inf. Sci. | 11 |
| 2026 | Motion data segmentation using robust subspace clustering with noise suppression
Qian Wang 0001, Hong Song 0003, Yungang Hao, Yunzhi Luo, Jingfan Fan, Jian Yang 0009 |
Knowl. Based Syst. | 2 |
| 2026 | KAB: A knowledge-aligned benchmark for reproducible evaluation of distantly supervised relation extraction
Bowen Liu 0011, Junhang Hu, Yucong Lin, Hong Song 0003, Yaqing Nie, Hongmin Xiao, Zichao Lin, Xutao Weng, Zhaoli Su, Jinfu Li 0004, Jian Yang 0003 |
Neural Networks | 4 |
| 2026 | Multimodal hybrid mamba classification model for tumor pathological grade prediction using magnetic resonance images
Langtao Zhou, Tianyu Fu 0003, Xiaoxia Qu, Jiaoyang Wu, Yangrui Huang, Hong Song 0003, Jingfan Fan, Danni Ai, Deqiang Xiao, Junfang Xian, Jian Yang 0003 |
Neural Networks | 6 |
| 2026 | Sculpting Margin Penalty: Intra-Task Adapter Merging and Classifier Calibration for Few-Shot Class-Incremental Learning
Liang Bai 0006, Hong Song 0003, Jinfu Li 0004, Yucong Lin, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Non-Linear Motion Estimation Network for Frame Interpolation in Coronary Angiographic SequencesabstractVideo frame interpolation (VFI) is significant for generating a high frame rate coronary sequence without additional radiation exposure. Due to the coronary reciprocating pattern alternating between systolic and diastolic phases, the linear assumption-based existing methods fail to capture the complex motion especially during the transitions between the two phases. Different from the linear methods, a Non-linear Motion Estimation Network (NLME-Net) is proposed to effectively capture the periodic reciprocating motion pattern by accurately estimating both bidirectional flows and long-distance motion. Specifically, the specialized motion estimation decoder is guided not only by target frame reconstruction loss but also by direct supervision through a self-supervised flow loss. This enhanced modeling of reciprocating motion enables accurate intermediate flow estimation in scenarios involving variable directional movement, thereby improving the accuracy and robustness of frame interpolation. Additionally, the interpolation decoder fully exploits the inherent mutual dependency between intermediate flow and target frame features to refine the final interpolation result. According to the experiment results of the proposed and twelve state-of-the-art methods using the coronary dataset with 6486 groups of angiographic images from 399 sequences, the proposed method improves the PSNR score by an average 0.59dB. Tianyu Fu 0003, Hong Song 0003, Deqiang Xiao, Jingfan Fan, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Hierarchical Heterogeneous Aggregation Network for Multi-Shape Coronary Stenosis Detection in X-Ray Angiography SequencesabstractAccurate detection of multi-shape coronary artery stenoses from X-ray angiography (XRA) sequences plays a crucial role in diagnosing and planning interventions for coronary artery disease. However, vessel overlap, background noise, and nonlinear cardiac motion introduce significant challenges. These factors often result in missed detections, intra-frame class conflict, and temporal category drift, particularly for subtle and morphologically complex stenoses such as focal and bifurcation stenoses. To address these challenges, we propose a Hierarchical Heterogeneous Aggregation Network that effectively integrates both spatial and temporal cues across XRA sequences. The proposed framework incorporates a Channel Importance-guided Fusion module, which aims to enhance the representation of small-stenosis features by dynamically selecting high-importance channels across scales. Furthermore, we introduce a Hierarchical Heterogeneous Aggregator designed to reduce spatial redundancy and explicitly generate discriminative features across frames based on heterogeneous relationships, thereby improving temporal consistency and classification robustness. Existing experiments conducted on two clinical datasets indicate that our method outperforms existing detectors and stenosis methods in terms of detection accuracy and generalization. Sigeng Chen, Jingfan Fan, Yujie Xie, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Wenyuan Yu, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | PLPFusion: Plane-Line-Pixel Fully Sparse Fusion for Robust Multi-Modal 3D Object DetectionabstractFully sparse fusion makes an excellent balance between efficiency and accuracy in multi-modal 3D object detection. However, most existing methods focus on foreground objects while overlooking background context. This oversight compromises detection robustness, especially for occluded or small-sized objects, leading to suboptimal detection performance. To address this limitation, we propose a novel fully sparse fusion framework (PLPFusion), which introduces a hierarchical Plane-Line-Pixel representation to progressively model the object-context relationships. PLPFusion comprises three key modules: the Plane Enhancement Module (PEM), the Line Alignment Module (LAM) and the Pixel-Level Aggregation Module (PLAM). Firstly, PEM utilizes geometric cues from LiDAR feature planes to generate spatially-aware object queries. Secondly, LAM further refines these queries with geometric priors for semantic awareness. Lastly, PLAM aggregates pixel-level context to enhance discriminative completeness by leveraging the semantically-aware object queries. On the nuScenes benchmark, PLPFusion achieves 71.9% mAP and 74.0% NDS, outperforming the baseline method FUTR3D by +2.5% mAP and +1.9% NDS, respectively. On the KITTI benchmark, it achieves 72.68% BEV mAP and 67.39% 3D mAP. These results confirm its robustness and effectiveness in diverse multi-modal 3D scenarios. The code of PLPFusion is available on the https://github.com/Text357/PLPFusion. Jingfu Hou, Hong Song 0003, Jinfu Li 0004, Yucong Lin, Jugang He, Xiuwei He, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Double-Decomposition Motion Tracking of Intraoperative 3D Structures via Cross-Spatio-Temporal Semantics Alignmentabstract3D motion tracking in X-ray image-guided operations using pre- and intra-operative image registration has recently gained attention. However, due to pre- and intra-operative acquisitions exist spatio-temporal misalignment (i.e., limited 3D prior versus continuous 2D images) and distinct respiratory phase difference, recent methods still struggle to accurately estimate 3D dynamic structures from X-ray images. To overcome these issues, we propose a novel double-decomposition tracking (DD-Track) framework that aligns with multi-organ motion characteristics via two alignment pipes: 1) Temporal alignment aims to compensate in-plane respiratory phases difference between the projection of static 3D prior and continuous X-ray images. A dual-excitation mechanism in the image and frequency domains is proposed to extract discriminate motion features while suppressing irrelevant background information. 2) Spatial alignment subsequently integrates the extracted 2D motion features into the cross-modal registration process to accurately warp the 3D prior. Further, we decompose the motion tracking into the common trajectory and organ-specific deformation to align with the multi-organ motion nature, avoiding excessive organ stretching for sliding compensation. Comprehensive quantitative and qualitative experiments on simulated and clinical multi-organ datasets demonstrate that DD-Track outperforms state-of-the-art methods, and we also validate its generalization for tracking intra-organ lesions on simulated data. Haixiao Geng, Jingfan Fan, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | SG-3DGS: Sequential Growing 3D Gaussian Splatting for Scene Reconstruction of Monocular Endoscope VideoabstractThe reconstruction of monocular endoscope video scenes is essential for enhancing the application and analysis of surgical endoscopic images. However, restricted by the narrow space of endoscopic movement and the obstruction of vision within cavities, it is difficult for most conventional methods to perform high-quality reconstruction. To address these challenges, a novel dynamic growing 3D Gaussian splatting architecture is proposed to construct the 3D model of endoscopic scene without precomputed camera poses or Structure from Motion. Firstly, to establish spatial feature associations between interframes, a 2D-3D displacement fields are designed by utilizing dense feature matches and depth prediction. On this basis, a novel displacement field variational optimization is developed to obtain relative poses by minimizing the energy functional associated with field transformation. Secondly, to address the constraint of the endoscopic view, by Gaussian sequential transformation and differential gradient field optimization, a novel Sequential Gaussian Growing Module is proposed to grow the local Gaussian model sequentially. Finally, a novel Forward-Reconstruction&Backward-Optimization architecture is proposed to generate the global Gaussian model. The evaluation is conducted on two public endoscopic datasets: Scared and C3VD. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods in both quantitative metrics (PSNR, SSIM, LPIPS, ATE, RMSE, MAE) and qualitative comparisons. The project page is https://iheckzza.github.io/ DG-3DGS/. Hong Song 0003, Jingfan Fan, Long Shao, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 2 |
| 2026 | Pelvic Fracture Reduction Planning via Joint Shape-Intensity ReferenceabstractPelvic fracture reduction planning is clinically critical yet technically demanding due to the complex anatomical structure of pelvis and the topological discontinuities introduced by fractures. Existing computer-assisted planning approaches dominantly rely on shape-based models, overlooking the rich CT intensity information that is essential for accurate and patient-specific planning. To address this limitation, we propose SIRDiff, a novel framework that incorporates anatomical shape and CT intensity information to generate biomechanically plausible reference models for pelvic fracture reduction planning. SIRDiff comprises three key components: 1) the structure-aware diffusion model to reconstruct the global anatomical structure, 2) the topology-adaptive structural conditioning strategy that maps fracture landmarks into a healthy anatomical graph domain for robust structure guidance, and 3) the detail-preserved autoencoder to ensure the fine-grained image reconstruction from latent representations. Additionally, SIRDiff adopts a multi-task learning approach to jointly predict the reference CT image and corresponding bone segmentation map, which enhances its potential for clinical application and ensures better anatomical consistency. Despite being trained exclusively on synthetic fracture data, SIRDiff shows the strong generalizability to real clinical cases and consistently outperforms existing methods across multiple clinically relevant evaluation metrics, demonstrating its potential as a robust and deployable solution for pelvic fracture reduction planning. Xirui Zhao, Deqiang Xiao, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Junqiang Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 9 |
| 2026 | Anatomy-Aware Sketch-Guided Latent Diffusion Model for Orbital Tumor Multi-Parametric MRI Missing Modalities SynthesisabstractSynthesizing missing modalities in multi-parametric MRI (mpMRI) is vital for accurate tumor diagnosis, yet remains challenging due to incomplete acquisitions and modality heterogeneity. Diffusion models have shown strong generative capability, but conventional approaches typically operate in the image domain with high memory costs and often rely solely on noise-space supervision, which limits anatomical fidelity. Latent diffusion models (LDMs) improve efficiency by performing denoising in latent space, but standard LDMs lack explicit structural priors and struggle to integrate multiple modalities effectively. To address these limitations, we propose the anatomy-aware sketch-guided latent diffusion model (ASLDM), a novel LDM-based framework designed for flexible and structure-preserving MRI synthesis. ASLDM incorporates an anatomy-aware feature fusion module, which encodes tumor region masks and edge-based anatomical sketches via cross-attention to guide the denoising process with explicit structure priors. A modality synergistic reconstruction strategy enables the joint modeling of available and missing modalities, enhancing cross-modal consistency and supporting arbitrary missing scenarios. Additionally, we introduce image-level losses for pixel-space supervision using L1 and SSIM losses, overcoming the limitations of pure noise-based loss training and improving the anatomical accuracy of synthesized outputs. Extensive experiments on a five-modality orbital tumor mpMRI private dataset and a four-modality public BraTS2024 dataset demonstrate that ASLDM outperforms state-of-the-art methods in both synthesis quality and structural consistency, showing strong potential for clinically reliable multi-modal MRI completion. Our code is publicly available at: https://github.com/zltshadow/ASLDM.git. Langtao Zhou, Xiaoxia Qu, Tianyu Fu 0003, Jiaoyang Wu, Hong Song 0003, Jingfan Fan, Danni Ai, Deqiang Xiao, Junfang Xian, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Enhanced CT-CBCT image registration for orthopedic surgery: Integrating rigid-elastic motion modelsabstractComputed tomography (CT) and cone-beam computed tomography (CBCT) image registration play pivotal roles in computer-assisted navigation for orthopedic surgery. Traditional methods often apply uniform deformation models, neglecting the biomechanical differences between rigid structures and soft tissues, which compromises registration accuracy, especially during significant bone displacements. To address this issue, we introduce RE-Reg, a rigid-elastic CT-CBCT image registration framework that jointly learns rigid bone motion and soft tissue deformation. RE-Reg incorporates a rigid alignment (RA) module to estimate global bone motion and an elastic deformation (ED) module to model soft tissue deformation, preserving bony structures through bone shape preservation (BSP) loss. Our comprehensive evaluation on publicly available datasets demonstrates that RE-Reg significantly outperforms existing methods in terms of registration accuracy and rigid bone structure preservation, achieving a 1.3% improvement in Dice similarity coefficient (DSC) and a 23% reduction in rigid bone deformation ( % Δ vol ) compared with the best baseline. This framework not only enhances anatomical fidelity but also ensures biomechanical plausibility and provides a valuable tool for image-guided orthopedic surgery. This code is available at https://github.com/Zq-Huang/RE-Reg. Deqiang Xiao, Hongxun Liu, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 9 |
| 2026 | Augmented reality surgical navigation: Clinical applications, key technologies, and future directionsabstractSurgical navigation has evolved significantly through advances in augmented reality, virtual reality, and mixed reality, improving precision and safety across many clinical applications, including neurosurgery, maxillofacial, spinal, and arthroplasty procedures. By integrating preoperative imaging with real-time intraoperative data, these systems provide dynamic guidance, reduce radiation exposure, and minimize tissue damage. Key challenges persist, including intraoperative registration accuracy, flexible tissue deformation, respiratory compensation, and real-time imaging quality. Emerging solutions include artificial intelligence-driven segmentation, deformation-field modeling, and hybrid registration techniques. Future developments will include lightweight, portable systems, improved non-rigid registration algorithms, and greater clinical adoption. Despite advances in rigid-tissue applications, soft-tissue navigation requires additional innovation to address motion variability and registration reliability, ultimately advancing minimally invasive surgery and precision medicine. Jingfan Fan, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Long Shao, Tao Chen 0022, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 10 |
| 2025 | SPPReg: Structure-Aware Partial-to-Complete Point Cloud Registration in Computer-Assisted Orthopedic SurgeryabstractAccurate alignment between partial intraoperative and complete preoperative bone surfaces is essential for navigation in computer-assisted orthopedic surgery. However, this task remains challenging due to low surface overlap, significant initial pose discrepancies, and noise inherent in intraoperative data, which often compromise the effectiveness of existing registration methods. To address these challenges, we propose a structure-aware partial-to-complete point cloud registration framework, named SPPReg, for accurate intraoperative-to-preoperative alignment, featuring a two-stage coarse-to-fine design. In the coarse alignment stage, a point completion network reconstructs missing structures in partial scans and leverages global geometric features to facilitate initial alignment under large pose variations. For the fine registration stage, we adopt a self-attention-based feature matching strategy that constructs a feature similarity matrix to establish accurate point correspondences. To reduce uncertainty interference, we design an overlap estimation block that learns point-wise overlap scores to select representative and reliable correspondences within overlapping regions, thereby improving the accuracy of fine registration. Comparative and ablation studies on a public bone point cloud dataset demonstrate that our method outperforms existing approaches in both accuracy and robustness, highlighting its effectiveness and potential for clinical application. Deqiang Xiao, Jingyi Bian, Long Shao, Hong Song 0003, Jian Yang 0009 |
BIBM | 5 |
| 2025 | DetectDiffuse: Aggregation- and Attention-Driven Universal Lesion Detection with Multi-scale Diffusion Model
Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
MICCAI (5) | 5 |
| 2025 | Temporal Modulated Multi-scale Deformation Fusion via Knowledge Distillation for 4D Medical Image Interpolation
Jiaju Zhang, Danni Ai, Zhikun Gan, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Deqiang Xiao |
MICCAI (8) | 6 |
| 2025 | Incremental energy-based recurrent transformer-KAN for time series deformation simulation of soft tissue
Jiaxi Jiang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Danni Ai, Deqiang Xiao, Yongtian Wang, Jian Yang 0009 |
Expert Syst. Appl. | 5 |
| 2025 | SFCLI-Net: Spatial-frequency collaborative learning interpolation network for Computed Tomography slice synthesis
Hong Song 0003, Danni Ai, Jieliang Shi, Jingfan Fan, Deqiang Xiao, Tianyu Fu 0003, Yucong Lin, Wencan Wu, Jian Yang 0009 |
Expert Syst. Appl. | 2 |
| 2025 | MixFuse: An iterative mix-attention transformer for multi-modal image fusion
Jinfu Li 0004, Hong Song 0003, Lei Liu 0070, Jianghan Xia, Jingfan Fan, Yucong Lin, Jian Yang 0009 |
Expert Syst. Appl. | 2 |
| 2025 | Collective Migration-Inspired Large-Deformation Compensation for Nonrigid Image Registration
Dingkun Liu, Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Yongtian Wang, Jian Yang 0009 |
Int. J. Comput. Vis. | 3 |
| 2025 | SEMNet: a simple and efficient MLP-based network for 3D Face point clouds landmarks localization
Mingyang Lei, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Danni Ai, Jingfan Fan, Jian Yang 0009 |
Multim. Syst. | 2 |
| 2025 | FCF-CSM: A Fuzzy Clustering Framework Based on Chromaticity Statistical Model for Automatic Segmentation of Port Wine StainsabstractExtracting lesions with accurate boundaries from clinical images is crucial for the clinical diagnosis, progression monitoring, and efficacy evaluation of port wine stains (PWS). However, accurately delineating lesion boundaries remains challenging due to complex boundary structures and unreliable annotations. In this paper, we propose a fuzzy clustering framework, FCF-CSM, for automated segmentation of PWS lesions. It combines prior knowledge of PWS color distribution with superpixels’ boundary characterization capability. Firstly, a chromaticity statistical model (CSM) is established based on 1000 collected PWS images, providing prior probabilities of PWS lesions to guide the improvement of the fuzzy clustering framework. Secondly, a superpixel method incorporating CSM is applied to PWS images, generating superpixels with improved boundary characterization. Thirdly, these superpixels are finely clustered using color features related to erythema index, color statistics, and color volume, improving PWS lesion distinction from complex backgrounds. Finally, a CSM-based automatic decision method distinguishes lesions from the background, achieving fully automated PWS segmentation within a fuzzy clustering framework. In addition, a boundary local fitting (BLF) metric is proposed to evaluate the segmentation precision of the PWS boundaries. Comparative experiments are conducted to verify the superiority of FCF-CSM. It achieves comparable overall segmentation performance with Jaccard and Dice metrics of 84.14% and 91.11%, respectively, compared to state-of-the-art methods. In terms of boundary segmentation, FCF-CSM outperforms other methods with an 81.52% BLF metric. FCF-CSM has proven to be effective for PWS segmentation and is promising to improve boundary delineation. The code is available athttps://github.com/JinrongMu/FCF-CSM. Note to Practitioners—The motivation of this study was to construct a statistical model of port wine stain (PWS) color to quantify prior knowledge of PWS lesion color in RGB images. Existing methods for automatic segmentation of PWS rely heavily on annotated data, but the lack of publicly available datasets hinders the development of such algorithms due to the privacy of clinical data. This paper proposes a fuzzy clustering framework based on the chromaticity statistical model, which can achieve high-precision and fine-grained delineation of the boundaries of PWS lesions without annotating data. In this study, we describe the prior of PWS color distribution based on colorimetry theory and apply this knowledge to the PWS automatic segmentation task, thereby realizing knowledge sharing while protecting patient privacy from being leaked. Comprehensive comparative experiments demonstrate the effectiveness and reliability of the chromaticity statistical model. However, the dataset used to build this model only includes populations with yellow skin tones. In future research, we will address the automatic identification of PWS lesions applicable to other skin tones. Jinrong Mu, Hong Song 0003, Xianqi Meng, Jingfan Fan, Danni Ai, Defu Chen, Haixia Qiu, Jian Yang 0009 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Multidomain Dependency-Aware Guided Unified-Stage Coronary Artery Branch Recognition NetworkabstractClinical scoring in X-ray coronary angiography image sequences is widely used for revascularization decision-making in cases of coronary artery disease. Accurately recognizing coronary artery branches is a fundamental step in assessing the severity of quantitative stenosis. Existing methods employ a multistage process that includes view separation, skeletonization, graph building, and classification using topological features. However, the graph often suffers from skeleton errors, leading to incorrect topological connections during the classification stage, which requires manual correction. To address these issues, we propose a unified-stage coronary artery branch recognition network (UniCABR) that integrates the segmentation, skeletonization, and graph-building stages. Specifically, we design a dependency-aware module to build dependency graphs in both semantic and spatial domains, avoiding the use of rigid inter-branch topological connections and thus eliminating the need for manual correction of misconnections resulting from skeleton errors. Furthermore, to suppress nontarget branches according to clinical criteria and enhance the performance of side branches, we introduce a small feature supplementation module coupled with an adaptive merged binary supervision method at the pixel level. Extensive experiments on two datasets and a generalization study demonstrate the superiority of UniCABR in performance and generalization ability for coronary artery branch recognition tasks. Sigeng Chen, Jingfan Fan, Danni Ai, Deqiang Xiao, Yucong Lin, Hong Song 0003, Wenyuan Yu, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Double-Shot 3D Shape Measurement With a Dual-Branch Network for Structured Light Projection ProfilometryabstractThe structured light (SL)-based three-dimensional (3D) measurement techniques with deep learning have been widely studied to improve measurement efficiency, among which fringe projection profilometry (FPP) and speckle projection profilometry (SPP) are two popular methods. However, they generally use a single projection pattern for reconstruction, resulting in fringe order ambiguity or poor reconstruction accuracy. To alleviate these problems, we propose a parallel dual-branch Convolutional Neural Network (CNN)-Transformer network (PDCNet), to take advantage of convolutional operations and self-attention mechanisms for processing different SL modalities. Within PDCNet, a Transformer branch is used to capture global perception in the fringe images, while a CNN branch is designed to collect local details in the speckle images. To fully integrate complementary features, we design a double-stream attention aggregation module (DAAM) that consists of a parallel attention subnetwork for aggregating multi-scale spatial structure information. This module can dynamically retain local and global representations to the maximum extent. Moreover, an adaptive mixture density head with bimodal Gaussian distribution is proposed for learning a representation that is precise near discontinuities. Compared to the standard disparity regression strategy, this adaptive mixture head can effectively improve performance at object boundaries. Extensive experiments demonstrate that our method can reduce fringe order ambiguity while producing high-accuracy results on self-made datasets. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Structured Light Image Planar-Topography Feature Decomposition for Generalizable 3D Shape MeasurementabstractThe application of structured light (SL) techniques has achieved remarkable success in three-dimensional (3D) measurements. Traditional methods generally calculate SL information pixel by pixel to obtain the measurement results. Recently, the rise of deep learning (DL) has led to significant developments in this task. However, existing DL-based methods generally learn all features within the image in an end-to-end manner, ignoring the distinction between SL and non-SL information. Therefore, these methods may encounter difficulties in focusing on subtle variations in SL patterns across different scenes, thereby degrading measurement precision. To overcome this challenge, we propose a novel SL Image Planar-Topography Feature Decomposition Network (SIDNet). To fully utilize the information from different SL modality images (fringe and speckle), we decompose different modalities into topography features (modality-specific) and planar features (modality-shared). A physics-driven decomposition loss is proposed to make the topography/planar features dissimilar/similar, which guides the network to distinguish between SL and non-SL information. Moreover, to obtain modality-fused features with global overview and local detail information, we propose a wrapped phase-driven feature fusion module. Specifically, a novel Tri-modality Mamba block is designed to integrate different sources with the guidance of the wrapped phase features. Extensive experiments demonstrate the superiority of our SIDNet in multiple simulated 3D measurement scenes. Moreover, our method shows better generalization ability than other DL models and can be directly applicable to unseen real-world scenes. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Landmark and Pose Prediction in Occluded Facial Point Cloud via Explicit Joint Feature Fusion NetworkabstractFacial point clouds collected in practical applications often suffer from pose variations and occlusion. Existing studies typically focus on either pose estimation or landmarks localization, neglecting to fully utilize the effective information from various facial features, thus limiting the improvement of prediction accuracy. Therefore, we propose an innovative 3D facial multi-task prediction network. The proposed network embeds the output of related tasks into feature extraction from the point level to the global level based on the physical dependencies between tasks. This facilitates explicit multi-task knowledge transfer, enabling the simultaneous prediction of facial landmarks, occlusion, and head pose. We introduce a training strategy based on posterior knowledge correction to iteratively refine and improve multi-task prediction results. Moreover, no single dataset provides annotations for all these tasks at once, so we synthesized a 3D landmarks, occlusion and pose (3D-LOP) dataset, which includes annotations for landmarks coordinates, occlusion probability, and head pose. The proposed method was compared with state-of-the-art methods on two public datasets and 3D-LOP. The landmarks localization accuracy improved by 7.1% on the two public datasets, and the pose estimation accuracy and stability on 3D-LOP improved by 28.5% and 32.7%, respectively. The performance on wild data also shows its potential in practical applications. Jingfan Fan, Long Shao, Mingyang Lei, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Hong Song 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | VTAG: Visual-Textual Association Guided Radiology Reports GenerationabstractRadiology report generation, which automatically generates diagnostic textual reports from medical images, plays a crucial role in improving clinical efficiency and diagnostic accuracy. However, existing radiology report generation models face numerous challenges, such as lack of interpretability as well as description inaccuracy. To address these issues, we propose an integrated framework that enhances radiology report generation by combining target detection with contextual alignment of relevant region descriptions. Target detection focuses on clinically significant areas within medical images, while contextual alignment ensures that the generated text is directly linked to visual findings. Additionally, we introduce a full-spectrum feature fusion method that combines both high- and low-frequency features from the images. This approach captures details and broader structures, allowing the model to gain a more comprehensive and hierarchical understanding of the images. We validated the effectiveness of our method on the public dataset MIMIC-CXR. The results indicate that our method outperforms previous approaches on multiple evaluation metrics. Notably, in terms of the average of the six traditional metrics, our method (VTAG) achieved a significant improvement of 14.3%, compared to the state-of-the-art model MLRG. Zhaoli Su, Yucong Lin, Hong Song 0003, Ruoyi Jian, Bowen Liu 0011, Jian Yang 0009 |
IEEE Trans. Image Process. | 3 |
| 2025 | Hepatic Vessel Roadmap Prediction Using Adaptive Tracking and Bending Energy Modeling in X-Ray FluoroscopyabstractDynamic visualization of the hepatic vessel is crucial in X-ray image-guided transjugular intrahepatic portosystemic shunt (TIPS) procedures. However, intraoperative breathing and the presence of guidewires complicate the prediction of the vessel position and posture without contrast agents. The respiration compensation technique aims to utilize the intraoperative respiration modeling to deform the initial vessel roadmap, thereby achieving the dynamic vessel prediction in the X-ray image sequence for the interventional guidance. Therefore, we propose a novel respiration compensation framework utilizing the adaptive tracking and bending energy modeling to achieve the stable vessel roadmap prediction under free breathing. First, we introduce the inter-frame rigid displacement compensation module based on the domain adaptation and adaptive centroid tracking. This module fits the respiratory curve from the X-ray images, providing the temporal motion priors for aligning roadmaps across frames. Second, we propose the novel deformation compensation module based on the bending energy modeling to correct the respiratory motion, wherein we utilize the energy features of the guidewires to drive the non-rigid registration. The control points sampled by the bending energy guide the local image to form the deformation field, facilitating the dynamic overlap of the vessel roadmaps in X-ray images. Experimental results on simulated and clinical datasets show an average tracking error of 0.95 $\pm$ 0.26 mm and 1.49 $\pm$ 0.40 mm, respectively. The effective and fast (mean 57 ms per frame) compensation achieved by our framework has the potential for improving the outcome of liver intervention and reducing the reliance on contrast agents. Deqiang Xiao, Haixiao Geng, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | PrixMatch: Semi-supervised Network for Multi-modal Medical Image Segmentation with Cross-modal Data Augmentation and Adaptive Prior Knowledge ThresholdingabstractSemi-supervised medical image segmentation has made significant strides, yet most existing methods are confined to single-modality data, limiting both the volume of data and the generalizability of the models. Multi-modal data can provide richer information, expand the dataset and enhance model robustness. However, integrating multi-modal learning into semi-supervised medical image segmentation presents challenges, primarily in how to deal with the scarcity of labels and alignment across different modalities simultaneously. In this paper, we propose PrixMatch, a multi-modal semi-supervised model with a teacher-student strategy for medical image segmentation. Initially, we propose a cross-modal data augmentation strategy, which randomly exchanges image blocks of the same location between different modalities, to guide the student model to learn cross-modal consistency without the need for additional network modules. Secondly, we design a cross-modal adaptive pseudo-label threshold setting strategy, which can align the prior anatomical knowledge of different modalities, and combine the modal-aligned prior knowledge and model learning state to filter the pseudo-labels at the pixel-level, flexibly alleviating the confirmation bias that occurs during semi-supervised training. Experiments demonstrate that PrixMatch achieves a Dice Similarity Coefficient (DSC) of 87.2% on the BTCV (CT) and CHAOS (MR) multi-modal datasets with only 10% labeling ratio, bringing nearly 5.5% improvement over the latest state-of-the-art method. Hong Song 0003, Yucong Lin, Long Shao, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009 |
BIBM | 2 |
| 2024 | Deformation Correction in Laparoscopic Liver Surgical Navigation Using Point Cloud Completion and Biomechanical ModelabstractIn the minimally invasive liver resection surgery, deformation estimation of liver is required to correct the preoperative virtual model to match the intraoperative scenarios, in which the liver deforms due to respiration and surgical operations. Existing methods on liver deformation estimation often struggle to achieve high accuracy when the intraoperative liver surface is limited in size. To overcome the challenge of sparse intraoperative point cloud data and improve the accuracy of liver deformation predictions, this paper introduces an innovative method for estimating liver deformation. This method comprises two main components: intraoperative point cloud completion and liver deformation estimation. Intraoperative point cloud completion uses registration techniques to integrate preoperative topological structures into the intraoperative phase. Liver deformation estimation combines optimization control with biomechanical modeling to accurately align the preoperative liver model with its intraoperative counterpart. Comparative and ablation experiments, as well as investigations into the impact of different completion ratios, were conducted. The results demonstrate that this method effectively utilizes preoperative liver geometric features to enhance intraoperative visualization, even with limited intraoperative data. Additionally, the opti-mization control method provides reliable deformation estimates with acceptable accuracy. This study offers new insights and methodologies for the development of augmented reality surgical navigation systems, contributing to the computer assisted liver surgey. Deqiang Xiao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Jian Yang 0009 |
BIBM | 7 |
| 2024 | N-Gram Swin Transformer for CT Image Super-Resolution
Zhenghao Gao, Danni Ai, Hong Song 0003, Jian Yang 0009 |
ICXR | 4 |
| 2024 | Revisiting class-incremental object detection: An efficient approach via intrinsic characteristics alignment and task decoupling
Liang Bai 0006, Hong Song 0003, Tao Feng 0014, Tianyu Fu 0003, Qingzhe Yu, Jian Yang 0009 |
Expert Syst. Appl. | 2 |
| 2024 | A Novel Long Short-Term Memory Learning Strategy for Object TrackingabstractIn this paper, a novel integrated long short‐term memory (LSTM) network and dynamic update model are proposed for long‐term object tracking in video images. The LSTM network tracking method is introduced to improve the effect of tracking failure caused by target occlusion. Stable tracking of the target is achieved using the LSTM method to predict the motion trajectory of the target when it is occluded and dynamically updating the tracking template. First, in target tracking, global average peak‐to‐correlation energy (GAPCE) is used to determine whether the tracking target is blocked or temporarily disappearing such that the follow‐up response tracking strategy can be adjusted accordingly. Second, the data with target motion characteristics are utilized to train the designed LSTM model to obtain an offline model, which effectively predicts the motion trajectory during the period when the target is occluded or has disappeared. Therefore, it can be captured again when the target reappears. Finally, in the dynamic template adjustment stage, the historical information of the target movement is combined, and the corresponding value of the current target is compared with the historical response value to realize the dynamic adjustment of the target tracking template. Compared with the current mainstream efficient convolution operators, namely, the E.T.Track, ToMP, KeepTrack, and RTS algorithms, on the OTB100 and LaSOT datasets, the proposed algorithm increases the distance precision by 9.9% when the distance threshold is 5 pixels, increases the overlap success rate by 0.94% when the overlap threshold is 0.75, and decreases the center location error by 18.9%. The proposed method has higher tracking accuracy and robustness and is more suitable for long‐term tracking of targets in actual scenarios than are the main approaches. Qian Wang 0001, Jian Yang 0009, Hong Song 0003 |
Int. J. Intell. Syst. | 3 |
| 2024 | MSLR: A Self-supervised Representation Learning Method for Tabular Data Based on Multi-scale Ladder Reconstruction
Xutao Weng, Hong Song 0003, Yucong Lin, Bowen Liu 0011, Jian Yang 0009 |
Inf. Sci. | 2 |
| 2024 | Domain base dynamic convolution and distance map guidance for anterior mediastinal lesion segmentation
Su Huang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009 |
Knowl. Based Syst. | 5 |
| 2024 | STQD-Det: Spatio-Temporal Quantum Diffusion Model for Real-Time Coronary Stenosis Detection in X-Ray AngiographyabstractDetecting coronary stenosis accurately in X-ray angiography (XRA) is important for diagnosing and treating coronary artery disease (CAD). However, challenges arise from factors like breathing and heart motion, poor imaging quality, and the complex vascular structures, making it difficult to identify stenosis fast and precisely. In this study, we proposed a Quantum Diffusion Model with Spatio-Temporal Feature Sharing to Real-time detect Stenosis (STQD-Det). Our framework consists of two modules: Sequential Quantum Noise Boxes module and spatio-temporal feature module. To evaluate the effectiveness of the method, we conducted a 4-fold cross-validation using a dataset consisting of 233 XRA sequences. Our approach achieved the F1 score of 92.39% with a real-time processing speed of 25.08 frames per second. These results outperform 17 state-of-the-art methods. The experimental results show that the proposed method can accomplish the stenosis detection quickly and accurately. Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Fusion-competition framework of local topology and global texture for head pose estimation
Tianyu Fu 0003, Kaibin Cao, Jingfan Fan, Deqiang Xiao, Hong Song 0003, Jian Yang 0009 |
Pattern Recognit. | 7 |
| 2024 | Self-supervised local rotation-stable descriptors for 3D ultrasound registration using translation equivariant FCN
Yifan Wang 0037, Tianyu Fu 0003, Jingfan Fan, Deqiang Xiao, Hong Song 0003, Jian Yang 0009 |
Pattern Recognit. | 6 |
| 2024 | Segmentation of 3D Anatomically Diffused Tissues in Magnetic Resonance Images Through Edge-Preserving Constrained Center-Free Fuzzy $C$-MeansabstractAnatomically diffused tissues (ADTs) refer to soft tissues containing many anatomical regions that are spatially dispersed and structurally irregular. In magnetic resonance images, ADTs exhibit blurred morphology and heterogeneous texture, making the accurate extraction of their 3D anatomy challenging. Center-free fuzzy C-means (FCM) can effectively partition nonlinear or nonspherical clusters, providing a promising scheme for ADT segmentation. It solves the uncertainty arising from unreliable center estimation by introducing a similarity criterion. However, the similarity criterion is sensitive to the number of target objects and their adjacent members in the images. Moreover, memberships of the existing algorithms are susceptible to losing real ADT details. To handle these issues, we propose an edge-preserving constrained center-free FCM algorithm for segmenting 3D ADTs in magnetic resonance images. To overcome the sensitivity of the similarity criterion, a novel object-to-cluster similarity measure is first proposed to utilize refined member-toobject adjacency. Specifically, the similarity measure focuses on members in the feature space, which share approximately homogeneous characteristics with each target object. Gradient-domain edge-preserving filtering is then combined with the improved similarity criterion to construct the novel objective function of center-free FCM. With the assistance of the designed imagedriven edge-preserving regularization, the gradient information of clusters is constrained, eventually approaching that of ADTs in the guidance image. Experiments are conducted on two public brain datasets and one local intrahepatic vein dataset. The results demonstrate that the proposed algorithm is more effective for ADT segmentation than the state-of-the-art peers, exhibiting superior generalization capability. Qing Guo 0008, Hong Song 0003, Cong Wang 0033, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Bi-Fusion of Structure and Deformation at Multi-Scale for Joint Segmentation and RegistrationabstractMedical image segmentation and registration are two fundamental and highly related tasks. However, current works focus on the mutual promotion between the two at the loss function level, ignoring the feature information generated by the encoder-decoder network during the task-specific feature mapping process and the potential inter-task feature relationship. This paper proposes a unified multi-task joint learning framework based on bi-fusion of structure and deformation at multi-scale, called BFM-Net, which simultaneously achieves the segmentation results and deformation field in a single-step estimation. BFM-Net consists of a segmentation subnetwork (SegNet), a registration subnetwork (RegNet), and the multi-task connection module (MTC). The MTC module is used to transfer the latent feature representation between segmentation and registration at multi-scale and link different tasks at the network architecture level, including the spatial attention fusion module (SAF), the multi-scale spatial attention fusion module (MSAF) and the velocity field fusion module (VFF). Extensive experiments on MR, CT and ultrasound images demonstrate the effectiveness of our approach. The MTC module can increase the Dice scores of segmentation and registration by 3.2%, 1.6%, 2.2%, and 6.2%, 4.5%, 3.0%, respectively. Compared with six state-of-the-art algorithms for segmentation and registration, BFM-Net can achieve superior performance in various modal images, fully demonstrating its effectiveness and generalization. Jiaju Zhang, Tianyu Fu 0003, Deqiang Xiao, Jingfan Fan, Hong Song 0003, Danni Ai, Jian Yang 0009 |
IEEE Trans. Image Process. | 5 |
| 2024 | Cross-Anatomy Transfer Learning via Shape-Aware Adaptive Fine-Tuning for 3D Vessel SegmentationabstractDeep learning methods have recently achieved remarkable performance in vessel segmentation applications, yet require numerous labor-intensive labeled data. To alleviate the requirement of manual annotation, transfer learning methods can potentially be used to acquire the related knowledge of tubular structures from public large-scale labeled vessel datasets for target vessel segmentation in other anatomic sites of the human body. However, the cross-anatomy domain shift is a challenging task due to the formidable discrepancy among various vessel structures in different anatomies, resulting in the limited performance of transfer learning. Therefore, we propose a cross-anatomy transfer learning framework for 3D vessel segmentation, which first generates a pre-trained model on a public hepatic vessel dataset and then adaptively fine-tunes our target segmentation network initialized from the model for segmentation of other anatomic vessels. In the framework, the adaptive fine-tuning strategy is presented to dynamically decide on the frozen or fine-tuned filters of the target network for each input sample with a proxy network. Moreover, we develop a Gaussian-based signed distance map that explicitly encodes vessel-specific shape context. The prediction of the map is added as an auxiliary task in the segmentation network to capture geometry-aware knowledge in the fine-tuning. We demonstrate the effectiveness of our method through extensive experiments on two small-scale datasets of coronary artery and brain vessel. The results indicate the proposed method effectively overcomes the discrepancy of cross-anatomy domain shift to achieve accurate vessel segmentation for these two datasets. Danni Ai, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Embedding-Alignment Fusion-Based Graph Convolution Network With Mixed Learning Strategy for 4D Medical Image ReconstructionabstractIn recent years, 4D medical image involving structural and motion information of tissue has attracted increasing attention. The key to the 4D image reconstruction is to stack the 2D slices based on matching the aligned motion states. In this study, the distribution of the 2D slices with the different motion states is modeled as a manifold graph, and the reconstruction is turned to be the graph alignment. An embedding-alignment fusion-based graph convolution network (GCN) with a mixed-learning strategy is proposed to align the graphs. Herein, the embedding and alignment processes of graphs interact with each other to realize a precise alignment with retaining the manifold distribution. The mixed strategy of self- and semi-supervised learning makes the alignment sparse to avoid the mismatching caused by outliers in the graph. In the experiment, the proposed 4D reconstruction approach is validated on the different modalities including Computed Tomography (CT), Magnetic Resonance Imaging (MRI), and Ultrasound (US). We evaluate the reconstruction accuracy and compare it with those of state-of-the-art methods. The experiment results demonstrate that our approach can reconstruct a more accurate 4D image. Tianyu Fu 0003, Hong Song 0003, Jingfan Fan, Deqiang Xiao, Yucong Lin, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Local Contractive Registration With Biomechanical Model: Assessing Microwave Ablation After Compensation for Tissue ShrinkageabstractMicrowave ablation (MWA) is a minimally invasive procedure for the treatment of liver tumor. Accumulating clinical evidence has considered the minimal ablative margin (MAM) as a significant predictor of local tumor progression (LTP). In clinical practice, MAM assessment is typically carried out through image registration of pre- and post-MWA images. However, this process faces two main challenges: non-homologous match between tumor and coagulation with inconsistent image appearance, and tissue shrinkage caused by thermal dehydration. These challenges result in low precision when using traditional registration methods for MAM assessment. In this paper, we present a local contractive nonrigid registration method using a biomechanical model (LC-BM) to address these challenges and precisely assess the MAM. The LC-BM contains two consecutive parts: (1) local contractive decomposition (LC-part), which reduces the incorrect match between the tumor and coagulation and quantifies the shrinkage in the external coagulation region, and (2) biomechanical model constraint (BM-part), which compensates for the shrinkage in the internal coagulation region. After quantifying and compensating for tissue shrinkage, the warped tumor is overlaid on the coagulation, and then the MAM is assessed. We evaluated the method using prospectively collected data from 36 patients with 47 liver tumors, comparing LC-BM with 11 state-of-the-art methods. LTP was diagnosed through contrast-enhanced MR follow-up images, serving as the ground truth for tumor recurrence. LC-BM achieved the highest accuracy (97.9%) in predicting LTP, outperforming other methods. Therefore, our proposed method holds significant potential to improve MAM assessment in MWA surgeries. Dingkun Liu, Danni Ai, Tianyu Fu 0003, Yuanjin Gao, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | DSC-Recon: Dual-Stage Complementary 4-D Organ Reconstruction From X-Ray Image Sequence for Intraoperative FusionabstractAccurately reconstructing 4D critical organs contributes to the visual guidance in X-ray image-guided interventional operation. Current methods estimate intraoperative dynamic meshes by refining a static initial organ mesh from the semantic information in the single-frame X-ray images. However, these methods fall short of reconstructing an accurate and smooth organ sequence due to the distinct respiratory patterns between the initial mesh and X-ray image. To overcome this limitation, we propose a novel dual-stage complementary 4D organ reconstruction (DSC-Recon) model for recovering dynamic organ meshes by utilizing the preoperative and intraoperative data with different respiratory patterns. DSC-Recon is structured as a dual-stage framework: 1) The first stage focuses on addressing a flexible interpolation network applicable to multiple respiratory patterns, which could generate dynamic shape sequences between any pair of preoperative 3D meshes segmented from CT scans. 2) In the second stage, we present a deformation network to take the generated dynamic shape sequence as the initial prior and explore the discriminate feature (i.e., target organ areas and meaningful motion information) in the intraoperative X-ray images, predicting the deformed mesh by introducing a designed feature mapping pipeline integrated into the initialized shape refinement process. Experiments on simulated and clinical datasets demonstrate the superiority of our method over state-of-the-art methods in both quantitative and qualitative aspects. Haixiao Geng, Jingfan Fan, Sigeng Chen, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Robust multi-view low-rank embedding clustering
Jian Dai 0002, Hong Song 0003, Yunzhi Luo, Zhenwen Ren, Jian Yang 0009 |
Neural Comput. Appl. | 2 |
| 2023 | Densely Connected U-Net With Criss-Cross Attention for Automatic Liver Tumor Segmentation in CT ImagesabstractAutomatic liver tumor segmentation plays a key role in radiation therapy of hepatocellular carcinoma. In this paper, we propose a novel densely connected U-Net model with criss-cross attention (CC-DenseUNet) to segment liver tumors in computed tomography (CT) images. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intra-slice features of liver tumors. Moreover, the criss-cross attention is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver tumors. We evaluated the proposed CC-DenseUNet on the LiTS dataset and the 3DIRCADb dataset. Experimental results show that the proposed method reaches the state-of-the-art performance for liver tumor segmentation. We further experimentally demonstrate the robustness of the proposed method on a clinical dataset comprising 20 CT volumes. Qiang Li 0049, Hong Song 0003, Zenghui Wei, Fengbo Yang, Jingfan Fan, Danni Ai, Yucong Lin, Xiaoling Yu, Jian Yang 0009 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | DSANet: Dual-Branch Shape-Aware Network for Echocardiography Segmentation in Apical ViewsabstractEchocardiography is an essential examination for cardiac disease diagnosis, from which anatomical structures segmentation is the key to assessing various cardiac functions. However, the obscure boundaries and large shape deformations due to cardiac motion make it challenging to accurately identify the anatomical structures in echocardiography, especially for automatic segmentation. In this study, we propose a dual-branch shape-aware network (DSANet) to segment the left ventricle, left atrium, and myocardium from the echocardiography. Specifically, the elaborate dual-branch architecture integrating shape-aware modules boosts the corresponding feature representation and segmentation performance, which guides the model to explore shape priors and anatomical dependence using an anisotropic strip attention mechanism and cross-branch skip connections. Moreover, we develop a boundary-aware rectification module together with a boundary loss to regulate boundary consistency, adaptively rectifying the estimation errors nearby the ambiguous pixels. We evaluate our proposed method on the publicly available and in-house echocardiography dataset. Comparative experiments with other state-of-the-art methods demonstrate the superiority of DSANet, which suggests its potential in advancing echocardiography segmentation. Guangquan Zhou, Wen-Bo Zhang, Zhong-Qing Shi, Zhan-Ru Qi, Kai-Ni Wang, Hong Song 0003, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | 3D ARCNN: An Asymmetric Residual CNN for Decreasing False Positive Rate of Lung Nodules DetectionabstractLung cancer is with the highest morbidity and mortality, and early detection of cancerous changes is essential to reduce the risk of death. To achieve this, it is necessary to reduce the false positive rate of detection. In this paper, we propose a novel asymmetric residual network, called 3D ARCNN, to reduce false positive rate of lung nodules detection. 3D ARCNN consists of asymmetric convolutional and multilayer cascaded residual network structures. To solve the problem of deep neural network with large amounts of parameters and poor reproduction ability, the proposed model uses asymmetric convolution to reduce model parameters and enhance the generalization ability of the model. In addition, the model uses an internally cascaded multi-stage residual to prevent the gradient vanishing and exploding problems of deep networks. Experiments are performed on the public dataset LUNA16. Our method achieved high detection sensitivity of 91.6%, 92.7%, 93.2% and 95.8% at 1, 2, 4 and 8 false positives per scan, respectively, which got an average CPM index of 0.912. Experimental results show that the proposed 3D ARCNN is very useful for reducing the false positive rate of lung nodules in the clinic. Bowen Liu 0011, Hong Song 0003, Qiang Li 0049, Yucong Lin, Jian Yang 0009 |
BIBM | 2 |
| 2022 | Denoising of MR and CT images using cascaded multi-supervision convolutional neural networks with progressive training
Hong Song 0003, Lei Chen 0073, Yutao Cui, Qiang Li 0049, Jingfan Fan, Jian Yang 0009, Le Zhang 0004 |
Neurocomputing | 1 |
| 2022 | Volume-awareness and outlier-suppression co-training for weakly-supervised MRI breast mass segmentation with partial annotations
Xianqi Meng, Jingfan Fan, Jinrong Mu, Zongyu Li, Aocai Yang, Kuan Lv, Danni Ai, Yucong Lin, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009 |
Knowl. Based Syst. | 11 |
| 2022 | Portal Vein and Hepatic Vein Segmentation in Multi-Phase MR Images Using Flow-Guided Change DetectionabstractSegmenting portal vein (PV) and hepatic vein (HV) from magnetic resonance imaging (MRI) scans is important for hepatic tumor surgery. Compared with single phase-based methods, multiple phases-based methods have better scalability in distinguishing HV and PV by exploiting multi-phase information. However, these methods just coarsely extract HV and PV from different phase images. In this paper, we propose a unified framework to automatically and robustly segment 3D HV and PV from multi-phase MR images, which considers both the change and appearance caused by the vascular flow event to improve segmentation performance. Firstly, inspired by change detection, flow-guided change detection (FGCD) is designed to detect the changed voxels related to hepatic venous flow by generating hepatic venous phase map and clustering the map. The FGCD uniformly deals with HV and PV clustering by the proposed shared clustering, thus making the appearance correlated with portal venous flow robustly delineate without increasing framework complexity. Then, to refine vascular segmentation results produced by both HV and PV clustering, interclass decision making (IDM) is proposed by combining the overlapping region discrimination and neighborhood direction consistency. Finally, our framework is evaluated on multi-phase clinical MR images of the public dataset (TCGA) and local hospital dataset. The quantitative and qualitative evaluations show that our framework outperforms the existing methods. Qing Guo 0008, Hong Song 0003, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009 |
IEEE Trans. Image Process. | 2 |
| 2021 | CC-DenseUNet: Densely Connected U-Net with Criss-Cross Attention for Liver and Tumor Segmentation in CT VolumesabstractThe automatic segmentation of liver and tumor is important for hepatic tumor surgery. In this paper, we propose a novel densely connected U-Net (CC-DenseUNet), which integrates criss-cross attention (CCA) module, to segment the liver and tumor in computed tomography (CT) volumes. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intraslice features of liver and tumors. Moreover, the CCA module is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver or tumors. We evaluated the proposed CCDenseUNet on the Liver Tumor Segmentation Challenge and 3DIRCADb datasets. Experimental results show that our method outperformed the state-of-the-art methods in liver tumor segmentation and achieved a highly competitive performance in liver segmentation. Qiang Li 0049, Hong Song 0003, Jingfan Fan, Danni Ai, Yucong Lin, Jian Yang 0009 |
BIBM | 2 |
| 2021 | An automatic framework for endoscopic image restoration and enhancement
Muhammad Asif 0016, Lei Chen 0073, Hong Song 0003, Jian Yang 0009, Alejandro F. Frangi |
Appl. Intell. | 3 |
| 2021 | Divergence-Free Fitting-Based Incompressible Deformation Quantification of LiverabstractLiver is an incompressible organ that maintains its volume during the respiration-induced deformation. Quantifying this deformation with the incompressible constraint is significant for liver tracking. The constraint can be accomplished with retaining the divergence-free field obtained by the deformation decomposition. However, the decomposition process is time-consuming, and the removal of non-divergence-free field weakens the deformation. In this study, a divergence-free fitting-based registration method is proposed to quantify the incompressible deformation rapidly and accurately. First, the deformation to be estimated is mapped to the velocity in a diffeomorphic space. Then, this velocity is decomposed by a fast Fourier-based Hodge-Helmholtz decomposition to obtain the divergence-free, curl-free, and harmonic fields. The curl-free field is replaced and fitted by the obtained harmonic field with a translation field to generate a new divergence-free velocity. By optimizing this velocity, the final incompressible deformation is obtained. Moreover, a deep learning framework (DLF) is constructed to accelerate the incompressible deformation quantification. An incompressible respiratory motion model is built for the DLF by using the proposed registration method and is then used to augment the training data. An encoder-decoder network is introduced to learn appearance-velocity correlation at patch scale. In the experiment, we compare the proposed registration with three state-of-the-art methods. The results show that the proposed method can accurately achieve the incompressible registration of liver with a mean liver overlap ratio of 95.33%. Moreover, the time consumed by DLF is nearly 15 times shorter than that by other methods. Tianyu Fu 0003, Jingfan Fan, Dingkun Liu, Hong Song 0003, Chaoyi Zhang, Danni Ai, Zhigang Cheng, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | A General Endoscopic Image Enhancement Method Based on Pre-trained Generative Adversarial NetworksabstractEndoscopic images frequently have image quality problems due to the limitations of surgical instruments and the impact of surgical operations, such as uneven illumination, smogginess and color deviation. For deep learning based on enhancement methods, independent training lacks sufficient defect images and generalization capability, and combined training with mixture of data cannot identify diverse specific tasks. To address these issues, we propose a general method based on pre-trained generative adversarial network with a specified transfer learning strategy to obtain high-quality images. Initially, we independently train a standard network based on a universal task, e.g., uneven illumination, where a pre-trained model is extracted as a backbone with partially shared generator. Then, we transfer the backbone to more potential image enhancement tasks. Experiments on uneven illumination, smogginess, and color deviation indicate that the model successfully shares common features of high-quality images and responds specifically to different defects as well. Jingfan Fan, Danni Ai, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
BIBM | 4 |
| 2020 | Local Contractive Registration for Quantification of Tissue Shrinkage in Assessment of Microwave Ablation
Dingkun Liu, Tianyu Fu 0003, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009 |
MICCAI (3) | 5 |
| 2020 | Prior information constrained alternating direction method of multipliers for longitudinal compressive sensing MR imaging
Ruirui Kang, Danni Ai, Gangrong Qu, Qingbo Li, Yurong Jiang, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
Neurocomputing | 8 |
| 2020 | Groupwise registration with global-local graph shrinkage in atlas construction
Tianyu Fu 0003, Jian Yang 0009, Danni Ai, Hong Song 0003, Yurong Jiang, Yongtian Wang, Alejandro F. Frangi |
Medical Image Anal. | 5 |
| 2020 | Spatial probabilistic distribution map-based two-channel 3D U-net for visual pathway segmentation
Danni Ai, Zhiqi Zhao, Jingfan Fan, Hong Song 0003, Xiaoxia Qu, Junfang Xian, Jian Yang 0009 |
Pattern Recognit. Lett. | 4 |
| 2020 | Topology Optimization Using Multiple-Possibility Fusion for Vasculature ExtractionabstractVascular centerline extraction from angiography images plays an important role in computer-aided diagnosis of vascular disease. To solve the common problems related to noise and inconsistent vasculatures from uneven perfusion, this paper proposes an automatic framework for accurate vascular centerline extraction from angiograms that uses multi-probability fusion-based topology optimization. In this framework, vascular region is first segmented using a learning-based method. Then, initial centerlines are obtained by applying iterative filtering operation and multi-direction indexed non-maximum suppression. Topology optimization is achieved by gap filling. A connection probability map is constructed utilizing the information of initial centerlines, texture, and orientation of vasculatures. Shortest path tracking is employed to search for optimal connections around gaps in the initial centerlines. The proposed framework is evaluated using simulative and clinical coronary angiographies. The experimental results demonstrate that the proposed method can extract centerlines with F1 score of 97.28% ± 1.2% for vasculatures in 12 clinical angiographic images. It is evident that the proposed method can extract complete and accurate vascular centerlines from angiograms and can be used to repair gaps in other filamentary structures, such as roads and retinal blood vessels. This endows our method a great potential in the analysis of filamentary structures. Huihui Fang, Danni Ai, Weijian Cong, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Greedy Soft Matching for Vascular Tracking of Coronary Angiographic Image SequencesabstractVascular tracking of coronary angiographic image sequences is one of the most clinically important tasks in the diagnostic assessment and interventional guidance of cardiac disease. It is difficult to automate this application because the vascular structure is complex; moreover, unsatisfactory angiography image quality may exacerbate the difficulty of vasculature extraction. This paper converts vascular tracking into branch matching and proposes a novel and automatic greedy soft match algorithm. Our method is based on a graph framework. A graph model building module is proposed to represent the vascular structure. Then, a greedy branch searching method is adopted to acquire all possible paths in the graph that may match the reference vessel. Finally, a soft batch matching method that combines branch descriptor and dynamic time warping is presented to select the best matching branch. The solution to the problem takes advantage of both spatial and temporal continuity between successive frames. The experimental results demonstrate that the proposed algorithm is effective and robust for vascular tracking. The F1 score of a single branch dataset, which contains 12 angiographic image sequences with 77 angiograms of contrast agent-filled vessels, is 0.89 ± 0.06 and of a vessel tree dataset which contains nine sequences with 58 angiograms is 0.88 ± 0.05. Extensive experimental results well demonstrate the superior performance of the algorithm. In addition, it provides a universal solution to address the problem of filamentary structure tracking. Huihui Fang, Danni Ai, Yong Huang 0002, Yurong Jiang, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Spatio-Temporal Constrained Online Layer Separation for Vascular Enhancement in X-Ray Angiographic Image SequenceabstractAutomatic vascular enhancement is crucial to vascular structure identification in X-ray angiographic (XRA) image sequences. In this work, we propose a novel spatio-temporal constrained online layer separation (STOLS) method to achieve vascular enhancement in XRA image sequences. The proposed method integrates the motion consistency of structures into the temporal-constrained online robust principal component analysis (ORPCA) to remove quasi-static structures (e.g., bones) from the enhanced vascular images. Furthermore, smoothing technique is integrated into the spatial-constrained ORPCA to reduce motion artifacts and the noise introduced by non-uniform illumination. To make the proposed method more adaptive to various vascular structures, the spatial-constrained ORPCA is adjusted by an adaptive weight using the proportion of the vessel region in the previous frame. The performance of the proposed method is compared with five state-of-the-art subtraction methods with respect to local and global revised contrast-to-noise ratios (rCNRs) and reconstruction errors. For the proposed method, the local and global rCNRs of the final vessel layer reached 2.54 and 1.24, respectively, while the error between the original and reconstructed images from the respiratory, background, and vessel layer reached 0.0354. The proposed STOLS can enhance the angiograms in a real-time and online manner without fine-tuning parameters, and can thus be used for intra-operation diagnosis and interventional procedures of coronary artery diseases. Shuang Song 0005, Chenbing Du, Danni Ai, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Securing smart vehicles from relay attacks using machine learning
Hong Song 0003, Awais Bilal, Mamoun Alazab, Alireza Jolfaei |
J. Supercomput. | 2 |
| 2019 | Liver Segmentation in CT Images Using a Non-Local Fully Convolutional Neural NetworkabstractLiver segmentation is a critical step in diagnosing various kinds of hepatic diseases. Based on the segmentation results, physicians can make further assessments more accurately. Although deep learning methods have achieved excellent performance in liver segmentation tasks, the traditional convolution encoder-decoder architecture may easily loss the spatial information due to the stacked convolution and pooling layers. In this paper, we present a non-local spatial feature based neural network (referred as NL-Net) to learn more spatial features of liver for more accurate segmentation. The NL-Net consists of an encoder block, a non-local spatial feature learning block and a decoder block. We utilized the pretrained ResNet model with transfer learning as the encoder. The non-local block can learn long range dependencies of the liver pixel position by computing the response at a position as a weighted sum of the responses at all positions, which can help the network learn more robust features. We applied the proposed model to ISBI 2019 CHAOs liver Segmentation Challenge task and evaluated it on the testing set. Experimental results show that the proposed NL-Net achieved an average dice of 0.972, RAVD of 1.593, ASSD of 1.926 and MSSD of 110.658 on the segmentation results. Lei Chen 0073, Hong Song 0003, Qiang Li 0049, Yutao Cui, Jian Yang 0009, Xiaohua Hu 0001 |
BIBM | 2 |
| 2019 | Monte Carlo Tree Search for 3D/2D Registration of Vessel Graphsabstract3D/2D registration techniques can compensate for the deficiencies of X-ray angiography-based navigation in vascular interventional surgery, such as the lack of depth information and excessive use of contrast agents. In this study, we propose a novel Monte Carlo tree search-based 3D/2D vessel graph registration method. The registration problem is transferred to a tree search problem according to the topology of vessel centerlines. Then, the Monte Carlo tree search method is applied to find the optimal vessel matching associated with highest registration score. Experiments on uninitialized vessel data demonstrate that the proposed method can achieve the highest accuracy among four state-of-the-art methods. An average accuracy of 1.91 mm on clinical coronary artery data is obtained. For the independence of initial pose and robustness to noise, the proposed method can align 3D and 2D vessels without prior initialization in vascular interventional surgery. Shuang Song 0005, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009 |
BIBM | 6 |
| 2019 | Spatial Probabilistic Distribution Map Based 3D FCN for Visual Pathway Segmentation
Zhiqi Zhao, Danni Ai, Jingfan Fan, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
ICIG (2) | 5 |
| 2019 | Liver tumor segmentation in CT volumes using an adversarial densely connected networkabstractBACKGROUND: Malignant liver tumor is one of the main causes of human death. In order to help physician better diagnose and make personalized treatment schemes, in clinical practice, it is often necessary to segment and visualize the liver tumor from abdominal computed tomography images. Due to the large number of slices in computed tomography sequence, developing an automatic and reliable segmentation method is very favored by physicians. However, because of the noise existed in the scan sequence and the similar pixel intensity of liver tumors with their surrounding tissues, besides, the size, position and shape of tumors also vary from one patient to another, automatic liver tumor segmentation is still a difficult task. RESULTS: We perform the proposed algorithm to the Liver Tumor Segmentation Challenge dataset and evaluate the segmentation results. Experimental results reveal that the proposed method achieved an average Dice score of 68.4% for tumor segmentation by using the designed network, and ASD, MSD, VOE and RVD improved from 27.8 to 21, 147 to 124, 0.52 to 0.46 and 0.69 to 0.73, respectively after performing adversarial training strategy, which proved the effectiveness of the proposed method. CONCLUSIONS: The testing results show that the proposed method achieves improved performance, which corroborated the adversarial training based strategy can achieve more accurate and robustness results on liver tumor segmentation task. Lei Chen 0073, Hong Song 0003, Chi Wang 0004, Yutao Cui, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004 |
BMC Bioinform. | 2 |
| 2019 | Patch-Based Adaptive Background Subtraction for Vascular Enhancement in X-Ray CineangiogramsabstractOBJECTIVE: Automatic vascular enhancement in X-ray cineangiography is of crucial interest, for instance, for better visualizing and quantifying coronary arteries in diagnostic and interventional procedures. METHODS: A novel patch-based adaptive background subtraction method (PABSM) is proposed automatically enhancing vessels in coronary X-ray cineangiography. First, pixels in the cineangiogram are described by the vesselness and Gabor features. Second, a classifier is utilized to separate the cineangiogram into the rough vascular and non-vascular region. Dilation is applied to the classified binary image to include more vascular region. Third, a patch-based background synthesis is utilized to fill the removed vascular region. RESULTS: A database containing 320 cineangiograms of 175 patients was collected, and then an interventional cardiologist annotated all vascular structures. The performance of PABSM is compared with six state-of-the-art vascular enhancement methods regarding the precision-recall curve and C-value. The area under the precision-recall curve is 0.7133, and the C-value is 0.9659. CONCLUSION: PABSM can automatically enhance the coronary artery in the cineangiograms. It preserves the integrity of vascular topological structures, particularly in complex vascular regions, and removes noise caused by the non-uniform gray-level distribution in the cineangiogram. SIGNIFICANCE: PABSM can avoid the motion artifacts and it eases the subsequent vascular segmentation, which is crucial for the diagnosis and interventional procedures of coronary artery diseases. Shuang Song 0005, Alejandro F. Frangi, Jian Yang 0009, Danni Ai, Chenbing Du, Yong Huang 0002, Hong Song 0003, Luosha Zhang, Yechen Han, Yongtian Wang |
IEEE J. Biomed. Health Informatics | 7 |
| 2018 | Inter/Intra-Constraints Optimization for Fast Vessel Enhancement in X-ray Angiographic Image Sequence
Chenbing Du, Shuang Song 0005, Danni Ai, Hong Song 0003, Yong Huang 0002, Yongtian Wang, Jian Yang 0009 |
BIBM | 4 |
| 2018 | Automatic Liver Segmentation Using Multi-plane Integrated Fully Convolutional Neural Networks
Chi Wang 0004, Hong Song 0003, Lei Chen 0073, Qiang Li 0049, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004 |
BIBM | 2 |
| 2018 | Local statistical deformation models for deformable image registration
Songyuan Tang, Weijian Cong, Jian Yang 0009, Tianyu Fu 0003, Hong Song 0003, Danni Ai, Yongtian Wang |
Neurocomputing | 5 |
| 2017 | Dorsal hand vein recognition based on convolutional neural networksabstractIn this paper, we proposed a dorsal hand vein recognition method based on Convolutional Neural Network (CNN), compared the recognition rate of different depth CNN models and analyzed the influence of dataset size on dorsal hand vein recognition rate. Firstly, the region of interest (ROI) of dorsal hand vein images was extracted, and contrast limited adaptive histogram equalization (CLAHE) and Gaussian smoothing filter algorithm were used to preprocess the images. Then Reference-CaffeNet AlexNet and VGG depth CNN were trained to extract image feature. Finally, logistic regression was applied for identification. The experimental results on two different size of dataset shown that the depth of network and size of data set size have different degree effect on recognition rate, the dorsal hand vein recognition rate based on VGG-19 reaches 99.7%. In this paper, we also explored the feasibility of ensemble learning on SqueezeNet. The recognition rate declined slightly with 99.52%, but the model size has been decreased sharply. Haipeng Wan, Lei Chen 0073, Hong Song 0003, Jian Yang 0009 |
BIBM | 3 |
| 2017 | Registration and fusion quantification of augmented reality based nasal endoscopic surgery
Yakui Chu, Jian Yang 0009, Shaodong Ma, Danni Ai, Hong Song 0003, Duanduan Chen, Lei Chen 0073, Yongtian Wang |
Medical Image Anal. | 6 |
| 2016 | Automatic schizophrenia discrimination on fNIRS by using PCA and SVMabstractA method is proposed to distinguish patients with schizophrenia from healthy controls based on data measured by functional near-infrared spectroscopy (fNIRS) during a cognitive task, which combines principal component analysis (PCA) and support vector machine (SVM). Firstly, a data reduction technique is applied prior to PCA, and then PCA is used to extract features on oxygenated hemoglobin (oxy-Hb) signals from 52-channel fNIRS data of schizophrenia and healthy subjects. Secondly, a classifier based on SVM is designed to discriminate schizophrenia from healthy controls. We recruited a large sample of 52 schizophrenia patients and 38 healthy controls. The hemoglobin response was measured in the prefrontal cortex during the one-back memory task using a 52-channel fNIRS system. The experimental results indicate that the proposed method can achieve a satisfactory classification with the accuracy of 93.33%, 100% for schizophrenia samples and 84.62% for healthy controls. Also, our results suggested that fNIRS has the potential capacity to be an effective objective biomarker for the diagnosis of schizophrenia. Hong Song 0003, Iordachescu Ilie Mihaita Bogdan, Shuliang Wang 0001, Wentian Dong, Wenxiang Quan, Weimin Dang |
BIBM | 1 |
| 2014 | Liver segmentation based on SKFCM and improved GrowCut for CT imagesabstractAccurate liver segmentation is an essential and crucial step for computer-aided liver disease diagnosis and surgical planning. In this paper, a new coarse-to-fine method is proposed to segment liver for abdominal computed tomography (CT) images. This hierarchical framework consists of rough segmentation and refined segmentation. The rough segmentation is implemented based on a kernel fuzzy C-means algorithm with spatial information (SKFCM) algorithm and the refined segmentation is performed based on the proposed improved GrowCut (IGC) algorithm. The SKFCM algorithm introduces a kernel function and spatial constraint based on fuzzy c-means clustering (FCM) algorithm, which can reduce the effect of noise and improve the clustering ability. The IGC algorithm makes good use of the continuity of CT series in space which can automatically generate the seed labels and improve the efficiency of segmentation. The proposed method was applied to segment the liver for the whole dataset of abdominal CT images. The performance evaluation of segmentation results shows that the proposed liver segmentation method is accurate and efficient. Experimental results have been shown visually and achieve reasonable consistency. Hong Song 0003, Shuliang Wang 0001 |
BIBM | 1 |
| 2014 | Splitting touching cells based on concave-point and improved watershed algorithms
Hong Song 0003, Qingjie Zhao, Yinghong Liu |
Frontiers Comput. Sci. | 1 |