EDBT 2026 Demo / reviewers in the wild / expert
Jian Yang 0009
dblp:y/JianYang9
· DBLP profile ↗
119ranked-venue papers
2as first author
69since 2021 · last 2026
0000-0003-1250-6319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 50 · 30 since 2021Artificial intelligence and machine learning · 34 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedKit: Multi-level feature distillation with knowledge injection for radiology report generation
Zhaoli Su, Hong Song 0003, Yucong Lin, Xutao Weng, Zhongxuan Mao, Bowen Liu 0011, Hongxia Yin, Jian Yang 0009 |
Expert Syst. Appl. | 9 |
| 2026 | Motion data segmentation using robust subspace clustering with noise suppression
Qian Wang 0001, Hong Song 0003, Yungang Hao, Yunzhi Luo, Jingfan Fan, Jian Yang 0009 |
Knowl. Based Syst. | 6 |
| 2026 | Generative data-engine foundation model for universal few-shot 2D vascular image segmentationabstractThe segmentation of 2D vascular structures via deep learning holds significant clinical value but is hindered by the scarcity of annotated data, severely limiting its widespread application. Developing a universal few-shot vascular segmentation model is highly desirable, yet remains challenging due to the need for extensive training and the inherent complexities of vascular imaging. In this work, we propose UniVG (Generative Data-engine Foundation Model for Universal Few-shot 2D Vascular Image Segmentation), a novel approach that learns the compositionality of vascular images and constructing a generative foundation model for robust vascular segmentation. UniVG enables the synthesis and learning of diverse and realistic vascular images through two key innovations: 1) Compositional learning for flexible and diverse vascular synthesis: It decomposes and recombines vascular structures with varying morphological features and diverse foreground-background configurations to generate richly diverse synthetic image-label pairs. 2) Few-shot generative adaptation for transferable segmentation: It fine-tunes pre-trained models with minimal annotated data to bridge the gap between synthetic and real vascular domains, synthesizing authentic and diverse vessel images for downstream few-shot vascular segmentation learning. To support our approach, we develop UniVG-58K, a large dataset comprising 58,689 vascular images across five imaging modalities, facilitating robust large-scale generative pre-training. Extensive experiments on 11 vessel segmentation tasks cross 5 modalties (only with 5 labeled images on each task) demonstrate that UniVG achieves performance comparable to fully supervised models, significantly reducing data collection and annotation costs. All code and datasets will be made publicly available at https://github.com/XinAloha/UniVG. Rongjun Ge, Yuxing Liu, Chengliang Liu 0003, Pinzheng Zhang, Jiong Zhang 0004, Jian Yang 0009, Jean-Louis Dillenseger, Yuting He 0001, Yang Chen 0008 |
Medical Image Anal. | 7 |
| 2026 | FDA-Recon: Feature and data alignment reconstruction for sparse-view CBCT
Yikun Zhang 0001, Dianlin Hu, Tianling Lyu, Yan Xi, Jian Yang 0009, Yang Chen 0008 |
Medical Image Anal. | 8 |
| 2026 | Sculpting Margin Penalty: Intra-Task Adapter Merging and Classifier Calibration for Few-Shot Class-Incremental Learning
Liang Bai 0006, Hong Song 0003, Jinfu Li 0004, Yucong Lin, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Non-Linear Motion Estimation Network for Frame Interpolation in Coronary Angiographic SequencesabstractVideo frame interpolation (VFI) is significant for generating a high frame rate coronary sequence without additional radiation exposure. Due to the coronary reciprocating pattern alternating between systolic and diastolic phases, the linear assumption-based existing methods fail to capture the complex motion especially during the transitions between the two phases. Different from the linear methods, a Non-linear Motion Estimation Network (NLME-Net) is proposed to effectively capture the periodic reciprocating motion pattern by accurately estimating both bidirectional flows and long-distance motion. Specifically, the specialized motion estimation decoder is guided not only by target frame reconstruction loss but also by direct supervision through a self-supervised flow loss. This enhanced modeling of reciprocating motion enables accurate intermediate flow estimation in scenarios involving variable directional movement, thereby improving the accuracy and robustness of frame interpolation. Additionally, the interpolation decoder fully exploits the inherent mutual dependency between intermediate flow and target frame features to refine the final interpolation result. According to the experiment results of the proposed and twelve state-of-the-art methods using the coronary dataset with 6486 groups of angiographic images from 399 sequences, the proposed method improves the PSNR score by an average 0.59dB. Tianyu Fu 0003, Hong Song 0003, Deqiang Xiao, Jingfan Fan, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Hierarchical Heterogeneous Aggregation Network for Multi-Shape Coronary Stenosis Detection in X-Ray Angiography SequencesabstractAccurate detection of multi-shape coronary artery stenoses from X-ray angiography (XRA) sequences plays a crucial role in diagnosing and planning interventions for coronary artery disease. However, vessel overlap, background noise, and nonlinear cardiac motion introduce significant challenges. These factors often result in missed detections, intra-frame class conflict, and temporal category drift, particularly for subtle and morphologically complex stenoses such as focal and bifurcation stenoses. To address these challenges, we propose a Hierarchical Heterogeneous Aggregation Network that effectively integrates both spatial and temporal cues across XRA sequences. The proposed framework incorporates a Channel Importance-guided Fusion module, which aims to enhance the representation of small-stenosis features by dynamically selecting high-importance channels across scales. Furthermore, we introduce a Hierarchical Heterogeneous Aggregator designed to reduce spatial redundancy and explicitly generate discriminative features across frames based on heterogeneous relationships, thereby improving temporal consistency and classification robustness. Existing experiments conducted on two clinical datasets indicate that our method outperforms existing detectors and stenosis methods in terms of detection accuracy and generalization. Sigeng Chen, Jingfan Fan, Yujie Xie, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Wenyuan Yu, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2026 | PLPFusion: Plane-Line-Pixel Fully Sparse Fusion for Robust Multi-Modal 3D Object DetectionabstractFully sparse fusion makes an excellent balance between efficiency and accuracy in multi-modal 3D object detection. However, most existing methods focus on foreground objects while overlooking background context. This oversight compromises detection robustness, especially for occluded or small-sized objects, leading to suboptimal detection performance. To address this limitation, we propose a novel fully sparse fusion framework (PLPFusion), which introduces a hierarchical Plane-Line-Pixel representation to progressively model the object-context relationships. PLPFusion comprises three key modules: the Plane Enhancement Module (PEM), the Line Alignment Module (LAM) and the Pixel-Level Aggregation Module (PLAM). Firstly, PEM utilizes geometric cues from LiDAR feature planes to generate spatially-aware object queries. Secondly, LAM further refines these queries with geometric priors for semantic awareness. Lastly, PLAM aggregates pixel-level context to enhance discriminative completeness by leveraging the semantically-aware object queries. On the nuScenes benchmark, PLPFusion achieves 71.9% mAP and 74.0% NDS, outperforming the baseline method FUTR3D by +2.5% mAP and +1.9% NDS, respectively. On the KITTI benchmark, it achieves 72.68% BEV mAP and 67.39% 3D mAP. These results confirm its robustness and effectiveness in diverse multi-modal 3D scenarios. The code of PLPFusion is available on the https://github.com/Text357/PLPFusion. Jingfu Hou, Hong Song 0003, Jinfu Li 0004, Yucong Lin, Jugang He, Xiuwei He, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Double-Decomposition Motion Tracking of Intraoperative 3D Structures via Cross-Spatio-Temporal Semantics Alignmentabstract3D motion tracking in X-ray image-guided operations using pre- and intra-operative image registration has recently gained attention. However, due to pre- and intra-operative acquisitions exist spatio-temporal misalignment (i.e., limited 3D prior versus continuous 2D images) and distinct respiratory phase difference, recent methods still struggle to accurately estimate 3D dynamic structures from X-ray images. To overcome these issues, we propose a novel double-decomposition tracking (DD-Track) framework that aligns with multi-organ motion characteristics via two alignment pipes: 1) Temporal alignment aims to compensate in-plane respiratory phases difference between the projection of static 3D prior and continuous X-ray images. A dual-excitation mechanism in the image and frequency domains is proposed to extract discriminate motion features while suppressing irrelevant background information. 2) Spatial alignment subsequently integrates the extracted 2D motion features into the cross-modal registration process to accurately warp the 3D prior. Further, we decompose the motion tracking into the common trajectory and organ-specific deformation to align with the multi-organ motion nature, avoiding excessive organ stretching for sliding compensation. Comprehensive quantitative and qualitative experiments on simulated and clinical multi-organ datasets demonstrate that DD-Track outperforms state-of-the-art methods, and we also validate its generalization for tracking intra-organ lesions on simulated data. Haixiao Geng, Jingfan Fan, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 11 |
| 2026 | SG-3DGS: Sequential Growing 3D Gaussian Splatting for Scene Reconstruction of Monocular Endoscope VideoabstractThe reconstruction of monocular endoscope video scenes is essential for enhancing the application and analysis of surgical endoscopic images. However, restricted by the narrow space of endoscopic movement and the obstruction of vision within cavities, it is difficult for most conventional methods to perform high-quality reconstruction. To address these challenges, a novel dynamic growing 3D Gaussian splatting architecture is proposed to construct the 3D model of endoscopic scene without precomputed camera poses or Structure from Motion. Firstly, to establish spatial feature associations between interframes, a 2D-3D displacement fields are designed by utilizing dense feature matches and depth prediction. On this basis, a novel displacement field variational optimization is developed to obtain relative poses by minimizing the energy functional associated with field transformation. Secondly, to address the constraint of the endoscopic view, by Gaussian sequential transformation and differential gradient field optimization, a novel Sequential Gaussian Growing Module is proposed to grow the local Gaussian model sequentially. Finally, a novel Forward-Reconstruction&Backward-Optimization architecture is proposed to generate the global Gaussian model. The evaluation is conducted on two public endoscopic datasets: Scared and C3VD. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods in both quantitative metrics (PSNR, SSIM, LPIPS, ATE, RMSE, MAE) and qualitative comparisons. The project page is https://iheckzza.github.io/ DG-3DGS/. Hong Song 0003, Jingfan Fan, Long Shao, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Pelvic Fracture Reduction Planning via Joint Shape-Intensity ReferenceabstractPelvic fracture reduction planning is clinically critical yet technically demanding due to the complex anatomical structure of pelvis and the topological discontinuities introduced by fractures. Existing computer-assisted planning approaches dominantly rely on shape-based models, overlooking the rich CT intensity information that is essential for accurate and patient-specific planning. To address this limitation, we propose SIRDiff, a novel framework that incorporates anatomical shape and CT intensity information to generate biomechanically plausible reference models for pelvic fracture reduction planning. SIRDiff comprises three key components: 1) the structure-aware diffusion model to reconstruct the global anatomical structure, 2) the topology-adaptive structural conditioning strategy that maps fracture landmarks into a healthy anatomical graph domain for robust structure guidance, and 3) the detail-preserved autoencoder to ensure the fine-grained image reconstruction from latent representations. Additionally, SIRDiff adopts a multi-task learning approach to jointly predict the reference CT image and corresponding bone segmentation map, which enhances its potential for clinical application and ensures better anatomical consistency. Despite being trained exclusively on synthetic fracture data, SIRDiff shows the strong generalizability to real clinical cases and consistently outperforms existing methods across multiple clinically relevant evaluation metrics, demonstrating its potential as a robust and deployable solution for pelvic fracture reduction planning. Xirui Zhao, Deqiang Xiao, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Junqiang Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 11 |
| 2026 | Anatomy-Aware Sketch-Guided Latent Diffusion Model for Orbital Tumor Multi-Parametric MRI Missing Modalities SynthesisabstractSynthesizing missing modalities in multi-parametric MRI (mpMRI) is vital for accurate tumor diagnosis, yet remains challenging due to incomplete acquisitions and modality heterogeneity. Diffusion models have shown strong generative capability, but conventional approaches typically operate in the image domain with high memory costs and often rely solely on noise-space supervision, which limits anatomical fidelity. Latent diffusion models (LDMs) improve efficiency by performing denoising in latent space, but standard LDMs lack explicit structural priors and struggle to integrate multiple modalities effectively. To address these limitations, we propose the anatomy-aware sketch-guided latent diffusion model (ASLDM), a novel LDM-based framework designed for flexible and structure-preserving MRI synthesis. ASLDM incorporates an anatomy-aware feature fusion module, which encodes tumor region masks and edge-based anatomical sketches via cross-attention to guide the denoising process with explicit structure priors. A modality synergistic reconstruction strategy enables the joint modeling of available and missing modalities, enhancing cross-modal consistency and supporting arbitrary missing scenarios. Additionally, we introduce image-level losses for pixel-space supervision using L1 and SSIM losses, overcoming the limitations of pure noise-based loss training and improving the anatomical accuracy of synthesized outputs. Extensive experiments on a five-modality orbital tumor mpMRI private dataset and a four-modality public BraTS2024 dataset demonstrate that ASLDM outperforms state-of-the-art methods in both synthesis quality and structural consistency, showing strong potential for clinically reliable multi-modal MRI completion. Our code is publicly available at: https://github.com/zltshadow/ASLDM.git. Langtao Zhou, Xiaoxia Qu, Tianyu Fu 0003, Jiaoyang Wu, Hong Song 0003, Jingfan Fan, Danni Ai, Deqiang Xiao, Junfang Xian, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Enhanced CT-CBCT image registration for orthopedic surgery: Integrating rigid-elastic motion modelsabstractComputed tomography (CT) and cone-beam computed tomography (CBCT) image registration play pivotal roles in computer-assisted navigation for orthopedic surgery. Traditional methods often apply uniform deformation models, neglecting the biomechanical differences between rigid structures and soft tissues, which compromises registration accuracy, especially during significant bone displacements. To address this issue, we introduce RE-Reg, a rigid-elastic CT-CBCT image registration framework that jointly learns rigid bone motion and soft tissue deformation. RE-Reg incorporates a rigid alignment (RA) module to estimate global bone motion and an elastic deformation (ED) module to model soft tissue deformation, preserving bony structures through bone shape preservation (BSP) loss. Our comprehensive evaluation on publicly available datasets demonstrates that RE-Reg significantly outperforms existing methods in terms of registration accuracy and rigid bone structure preservation, achieving a 1.3% improvement in Dice similarity coefficient (DSC) and a 23% reduction in rigid bone deformation ( % Δ vol ) compared with the best baseline. This framework not only enhances anatomical fidelity but also ensures biomechanical plausibility and provides a valuable tool for image-guided orthopedic surgery. This code is available at https://github.com/Zq-Huang/RE-Reg. Deqiang Xiao, Hongxun Liu, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 10 |
| 2026 | Augmented reality surgical navigation: Clinical applications, key technologies, and future directionsabstractSurgical navigation has evolved significantly through advances in augmented reality, virtual reality, and mixed reality, improving precision and safety across many clinical applications, including neurosurgery, maxillofacial, spinal, and arthroplasty procedures. By integrating preoperative imaging with real-time intraoperative data, these systems provide dynamic guidance, reduce radiation exposure, and minimize tissue damage. Key challenges persist, including intraoperative registration accuracy, flexible tissue deformation, respiratory compensation, and real-time imaging quality. Emerging solutions include artificial intelligence-driven segmentation, deformation-field modeling, and hybrid registration techniques. Future developments will include lightweight, portable systems, improved non-rigid registration algorithms, and greater clinical adoption. Despite advances in rigid-tissue applications, soft-tissue navigation requires additional innovation to address motion variability and registration reliability, ultimately advancing minimally invasive surgery and precision medicine. Jingfan Fan, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Long Shao, Tao Chen 0022, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 12 |
| 2025 | Multi-Task Learning for Optical Flow-Guided Self-Supervised Depth Estimation and Semantic Segmentation in Endoscopic SurgeryabstractSemantic segmentation in surgical scenes requires precise differentiation of organs and tissues, as well as strong generalization capabilities in complex environments, while the training of the model demands a large number of labeled images. To address these challenges and improve efficiency, we propose a joint-learning framework for self-supervised depth estimation and semantic segmentation, enhanced by optical flow. Our approach effectively leverages high-dimensional input features, using depth estimation to guide and improve semantic segmentation performance at a lower labeling cost. Meanwhile, an optical flow module is introduced and jointly trained with the segmentation network, with its outputs fused with RGB images as multimodal inputs to enhance dynamic feature perception and exploit semantic cues from motion. We evaluate our method on the CholecSeg8k and the experimental results demonstrate the effectiveness and robustness of our proposed approach. Xuexin Jiang, Sifan Cao, Long Shao, Jingfan Fan, Jian Yang 0009 |
BIBM | 6 |
| 2025 | SPPReg: Structure-Aware Partial-to-Complete Point Cloud Registration in Computer-Assisted Orthopedic SurgeryabstractAccurate alignment between partial intraoperative and complete preoperative bone surfaces is essential for navigation in computer-assisted orthopedic surgery. However, this task remains challenging due to low surface overlap, significant initial pose discrepancies, and noise inherent in intraoperative data, which often compromise the effectiveness of existing registration methods. To address these challenges, we propose a structure-aware partial-to-complete point cloud registration framework, named SPPReg, for accurate intraoperative-to-preoperative alignment, featuring a two-stage coarse-to-fine design. In the coarse alignment stage, a point completion network reconstructs missing structures in partial scans and leverages global geometric features to facilitate initial alignment under large pose variations. For the fine registration stage, we adopt a self-attention-based feature matching strategy that constructs a feature similarity matrix to establish accurate point correspondences. To reduce uncertainty interference, we design an overlap estimation block that learns point-wise overlap scores to select representative and reliable correspondences within overlapping regions, thereby improving the accuracy of fine registration. Comparative and ablation studies on a public bone point cloud dataset demonstrate that our method outperforms existing approaches in both accuracy and robustness, highlighting its effectiveness and potential for clinical application. Deqiang Xiao, Jingyi Bian, Long Shao, Hong Song 0003, Jian Yang 0009 |
BIBM | 6 |
| 2025 | Reducing Redundancy in Small Lesion Features for Multi-Shape Stenosis Detection in Coronary X-ray AngiographyabstractThe automatic detection of multi-shape coronary artery stenosis through X-ray angiography has important clinical implications for the diagnosis and treatment of coronary artery disease. However, there are several challenges in accurately detecting multi-shape stenoses, including the small size of stenoses, large size variations among multi-shape stenoses, unclear boundaries, and deformation caused by cardiac motion, such as stretching and contraction. To this end, we propose a framework for multi-shape stenosis detection at the sequence level by optimizing the representation of small lesions. Specifically, we propose a contribution-guided redundancy reduction module to suppress the feature redundancy of the background region while dynamically optimizing stenosis representations. Furthermore, to tackle the challenges of classification caused by the similarity in lesion appearance and unclear boundaries, we leverage the relationship between task-specific features and cross-enhancement to improve classification performance. Finally, a sequence relaxation strategy to extend the representation of lesions from the single-frame level to the sequence level. Experimental results indicate that the proposed method demonstrates a significant advantage over comparative methods in the task of multi-shape stenosis detection. Sigeng Chen, Jingfan Fan, Danni Ai, Jian Yang 0009 |
IJCNN | 6 |
| 2025 | DetectDiffuse: Aggregation- and Attention-Driven Universal Lesion Detection with Multi-scale Diffusion Model
Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
MICCAI (5) | 7 |
| 2025 | Incremental energy-based recurrent transformer-KAN for time series deformation simulation of soft tissue
Jiaxi Jiang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Danni Ai, Deqiang Xiao, Yongtian Wang, Jian Yang 0009 |
Expert Syst. Appl. | 9 |
| 2025 | SFCLI-Net: Spatial-frequency collaborative learning interpolation network for Computed Tomography slice synthesis
Hong Song 0003, Danni Ai, Jieliang Shi, Jingfan Fan, Deqiang Xiao, Tianyu Fu 0003, Yucong Lin, Wencan Wu, Jian Yang 0009 |
Expert Syst. Appl. | 10 |
| 2025 | MixFuse: An iterative mix-attention transformer for multi-modal image fusion
Jinfu Li 0004, Hong Song 0003, Lei Liu 0070, Jianghan Xia, Jingfan Fan, Yucong Lin, Jian Yang 0009 |
Expert Syst. Appl. | 9 |
| 2025 | Collective Migration-Inspired Large-Deformation Compensation for Nonrigid Image Registration
Dingkun Liu, Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Yongtian Wang, Jian Yang 0009 |
Int. J. Comput. Vis. | 9 |
| 2025 | SEMNet: a simple and efficient MLP-based network for 3D Face point clouds landmarks localization
Mingyang Lei, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Danni Ai, Jingfan Fan, Jian Yang 0009 |
Multim. Syst. | 9 |
| 2025 | FCF-CSM: A Fuzzy Clustering Framework Based on Chromaticity Statistical Model for Automatic Segmentation of Port Wine StainsabstractExtracting lesions with accurate boundaries from clinical images is crucial for the clinical diagnosis, progression monitoring, and efficacy evaluation of port wine stains (PWS). However, accurately delineating lesion boundaries remains challenging due to complex boundary structures and unreliable annotations. In this paper, we propose a fuzzy clustering framework, FCF-CSM, for automated segmentation of PWS lesions. It combines prior knowledge of PWS color distribution with superpixels’ boundary characterization capability. Firstly, a chromaticity statistical model (CSM) is established based on 1000 collected PWS images, providing prior probabilities of PWS lesions to guide the improvement of the fuzzy clustering framework. Secondly, a superpixel method incorporating CSM is applied to PWS images, generating superpixels with improved boundary characterization. Thirdly, these superpixels are finely clustered using color features related to erythema index, color statistics, and color volume, improving PWS lesion distinction from complex backgrounds. Finally, a CSM-based automatic decision method distinguishes lesions from the background, achieving fully automated PWS segmentation within a fuzzy clustering framework. In addition, a boundary local fitting (BLF) metric is proposed to evaluate the segmentation precision of the PWS boundaries. Comparative experiments are conducted to verify the superiority of FCF-CSM. It achieves comparable overall segmentation performance with Jaccard and Dice metrics of 84.14% and 91.11%, respectively, compared to state-of-the-art methods. In terms of boundary segmentation, FCF-CSM outperforms other methods with an 81.52% BLF metric. FCF-CSM has proven to be effective for PWS segmentation and is promising to improve boundary delineation. The code is available athttps://github.com/JinrongMu/FCF-CSM. Note to Practitioners—The motivation of this study was to construct a statistical model of port wine stain (PWS) color to quantify prior knowledge of PWS lesion color in RGB images. Existing methods for automatic segmentation of PWS rely heavily on annotated data, but the lack of publicly available datasets hinders the development of such algorithms due to the privacy of clinical data. This paper proposes a fuzzy clustering framework based on the chromaticity statistical model, which can achieve high-precision and fine-grained delineation of the boundaries of PWS lesions without annotating data. In this study, we describe the prior of PWS color distribution based on colorimetry theory and apply this knowledge to the PWS automatic segmentation task, thereby realizing knowledge sharing while protecting patient privacy from being leaked. Comprehensive comparative experiments demonstrate the effectiveness and reliability of the chromaticity statistical model. However, the dataset used to build this model only includes populations with yellow skin tones. In future research, we will address the automatic identification of PWS lesions applicable to other skin tones. Jinrong Mu, Hong Song 0003, Xianqi Meng, Jingfan Fan, Danni Ai, Defu Chen, Haixia Qiu, Jian Yang 0009 |
IEEE Trans Autom. Sci. Eng. | 10 |
| 2025 | Multidomain Dependency-Aware Guided Unified-Stage Coronary Artery Branch Recognition NetworkabstractClinical scoring in X-ray coronary angiography image sequences is widely used for revascularization decision-making in cases of coronary artery disease. Accurately recognizing coronary artery branches is a fundamental step in assessing the severity of quantitative stenosis. Existing methods employ a multistage process that includes view separation, skeletonization, graph building, and classification using topological features. However, the graph often suffers from skeleton errors, leading to incorrect topological connections during the classification stage, which requires manual correction. To address these issues, we propose a unified-stage coronary artery branch recognition network (UniCABR) that integrates the segmentation, skeletonization, and graph-building stages. Specifically, we design a dependency-aware module to build dependency graphs in both semantic and spatial domains, avoiding the use of rigid inter-branch topological connections and thus eliminating the need for manual correction of misconnections resulting from skeleton errors. Furthermore, to suppress nontarget branches according to clinical criteria and enhance the performance of side branches, we introduce a small feature supplementation module coupled with an adaptive merged binary supervision method at the pixel level. Extensive experiments on two datasets and a generalization study demonstrate the superiority of UniCABR in performance and generalization ability for coronary artery branch recognition tasks. Sigeng Chen, Jingfan Fan, Danni Ai, Deqiang Xiao, Yucong Lin, Hong Song 0003, Wenyuan Yu, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | Double-Shot 3D Shape Measurement With a Dual-Branch Network for Structured Light Projection ProfilometryabstractThe structured light (SL)-based three-dimensional (3D) measurement techniques with deep learning have been widely studied to improve measurement efficiency, among which fringe projection profilometry (FPP) and speckle projection profilometry (SPP) are two popular methods. However, they generally use a single projection pattern for reconstruction, resulting in fringe order ambiguity or poor reconstruction accuracy. To alleviate these problems, we propose a parallel dual-branch Convolutional Neural Network (CNN)-Transformer network (PDCNet), to take advantage of convolutional operations and self-attention mechanisms for processing different SL modalities. Within PDCNet, a Transformer branch is used to capture global perception in the fringe images, while a CNN branch is designed to collect local details in the speckle images. To fully integrate complementary features, we design a double-stream attention aggregation module (DAAM) that consists of a parallel attention subnetwork for aggregating multi-scale spatial structure information. This module can dynamically retain local and global representations to the maximum extent. Moreover, an adaptive mixture density head with bimodal Gaussian distribution is proposed for learning a representation that is precise near discontinuities. Compared to the standard disparity regression strategy, this adaptive mixture head can effectively improve performance at object boundaries. Extensive experiments demonstrate that our method can reduce fringe order ambiguity while producing high-accuracy results on self-made datasets. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | Structured Light Image Planar-Topography Feature Decomposition for Generalizable 3D Shape MeasurementabstractThe application of structured light (SL) techniques has achieved remarkable success in three-dimensional (3D) measurements. Traditional methods generally calculate SL information pixel by pixel to obtain the measurement results. Recently, the rise of deep learning (DL) has led to significant developments in this task. However, existing DL-based methods generally learn all features within the image in an end-to-end manner, ignoring the distinction between SL and non-SL information. Therefore, these methods may encounter difficulties in focusing on subtle variations in SL patterns across different scenes, thereby degrading measurement precision. To overcome this challenge, we propose a novel SL Image Planar-Topography Feature Decomposition Network (SIDNet). To fully utilize the information from different SL modality images (fringe and speckle), we decompose different modalities into topography features (modality-specific) and planar features (modality-shared). A physics-driven decomposition loss is proposed to make the topography/planar features dissimilar/similar, which guides the network to distinguish between SL and non-SL information. Moreover, to obtain modality-fused features with global overview and local detail information, we propose a wrapped phase-driven feature fusion module. Specifically, a novel Tri-modality Mamba block is designed to integrate different sources with the guidance of the wrapped phase features. Extensive experiments demonstrate the superiority of our SIDNet in multiple simulated 3D measurement scenes. Moreover, our method shows better generalization ability than other DL models and can be directly applicable to unseen real-world scenes. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | Landmark and Pose Prediction in Occluded Facial Point Cloud via Explicit Joint Feature Fusion NetworkabstractFacial point clouds collected in practical applications often suffer from pose variations and occlusion. Existing studies typically focus on either pose estimation or landmarks localization, neglecting to fully utilize the effective information from various facial features, thus limiting the improvement of prediction accuracy. Therefore, we propose an innovative 3D facial multi-task prediction network. The proposed network embeds the output of related tasks into feature extraction from the point level to the global level based on the physical dependencies between tasks. This facilitates explicit multi-task knowledge transfer, enabling the simultaneous prediction of facial landmarks, occlusion, and head pose. We introduce a training strategy based on posterior knowledge correction to iteratively refine and improve multi-task prediction results. Moreover, no single dataset provides annotations for all these tasks at once, so we synthesized a 3D landmarks, occlusion and pose (3D-LOP) dataset, which includes annotations for landmarks coordinates, occlusion probability, and head pose. The proposed method was compared with state-of-the-art methods on two public datasets and 3D-LOP. The landmarks localization accuracy improved by 7.1% on the two public datasets, and the pose estimation accuracy and stability on 3D-LOP improved by 28.5% and 32.7%, respectively. The performance on wild data also shows its potential in practical applications. Jingfan Fan, Long Shao, Mingyang Lei, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Hong Song 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | VTAG: Visual-Textual Association Guided Radiology Reports GenerationabstractRadiology report generation, which automatically generates diagnostic textual reports from medical images, plays a crucial role in improving clinical efficiency and diagnostic accuracy. However, existing radiology report generation models face numerous challenges, such as lack of interpretability as well as description inaccuracy. To address these issues, we propose an integrated framework that enhances radiology report generation by combining target detection with contextual alignment of relevant region descriptions. Target detection focuses on clinically significant areas within medical images, while contextual alignment ensures that the generated text is directly linked to visual findings. Additionally, we introduce a full-spectrum feature fusion method that combines both high- and low-frequency features from the images. This approach captures details and broader structures, allowing the model to gain a more comprehensive and hierarchical understanding of the images. We validated the effectiveness of our method on the public dataset MIMIC-CXR. The results indicate that our method outperforms previous approaches on multiple evaluation metrics. Notably, in terms of the average of the six traditional metrics, our method (VTAG) achieved a significant improvement of 14.3%, compared to the state-of-the-art model MLRG. Zhaoli Su, Yucong Lin, Hong Song 0003, Ruoyi Jian, Bowen Liu 0011, Jian Yang 0009 |
IEEE Trans. Image Process. | 6 |
| 2025 | Hepatic Vessel Roadmap Prediction Using Adaptive Tracking and Bending Energy Modeling in X-Ray FluoroscopyabstractDynamic visualization of the hepatic vessel is crucial in X-ray image-guided transjugular intrahepatic portosystemic shunt (TIPS) procedures. However, intraoperative breathing and the presence of guidewires complicate the prediction of the vessel position and posture without contrast agents. The respiration compensation technique aims to utilize the intraoperative respiration modeling to deform the initial vessel roadmap, thereby achieving the dynamic vessel prediction in the X-ray image sequence for the interventional guidance. Therefore, we propose a novel respiration compensation framework utilizing the adaptive tracking and bending energy modeling to achieve the stable vessel roadmap prediction under free breathing. First, we introduce the inter-frame rigid displacement compensation module based on the domain adaptation and adaptive centroid tracking. This module fits the respiratory curve from the X-ray images, providing the temporal motion priors for aligning roadmaps across frames. Second, we propose the novel deformation compensation module based on the bending energy modeling to correct the respiratory motion, wherein we utilize the energy features of the guidewires to drive the non-rigid registration. The control points sampled by the bending energy guide the local image to form the deformation field, facilitating the dynamic overlap of the vessel roadmaps in X-ray images. Experimental results on simulated and clinical datasets show an average tracking error of 0.95 $\pm$ 0.26 mm and 1.49 $\pm$ 0.40 mm, respectively. The effective and fast (mean 57 ms per frame) compensation achieved by our framework has the potential for improving the outcome of liver intervention and reducing the reliance on contrast agents. Deqiang Xiao, Haixiao Geng, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Multimodal Entity Linking With Dynamic Modality Selection and Interactive Prompt LearningabstractRecent advances in Multimodal Entity Linking leverage multimodal information to link target mentions to corresponding entities. However, existing methods uniformly adopt a “one-size-fits-all” approach, which overlooks the unique requirements of individual samples and fails to adequately balance modality-assisted disambiguation and modality-induced noise. Also, the commonly used separate large-scale visual and text pretrained models for feature extraction do not address inter-modal heterogeneity and the high computational cost of fine-tuning. To resolve these two issues, we introduce a novel approach named Multimodal Entity Linking with Dynamic Modality Selection and Interactive Prompt Learning (DSMIP). First, we design three expert networks that utilize different subsets of modalities tailored to the task and train them individually. Specifically, for the multimodal expert network, we enhance entity and mention feature extraction by updating multimodal prompts and setting up a coupling function to realize the interaction of prompts between modalities. Subsequently, to select the best-suited expert network for each specific sample, we devise a Modality Selection Gating Network to gain the optimal one-hot selection vector by applying a specialized reparameterization technique and a two-stage training process. Experimental results on three public benchmark datasets demonstrate that the proposed DSMIP outperforms all state-of-the-art baselines. The code is released on https://github.com/mayy-seu/DSMIP-code. Yingyao Ma, Jiasong Wu, Lotfi Senhadji, Huazhong Shu, Jian Yang 0009 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | GDP-Net: Global Dependency-Enhanced Dual-Domain Parallel Network for Ring Artifact RemovalabstractIn Computed Tomography (CT) imaging, the ring artifacts caused by the inconsistent detector response can significantly degrade the reconstructed images, having negative impacts on the subsequent applications. The new generation of CT systems based on photon-counting detectors are affected by ring artifacts more severely. The flexibility and variety of detector responses make it difficult to build a well-defined model to characterize the ring artifacts. In this context, this study proposes the global dependency-enhanced dual-domain parallel neural network for Ring Artifact Removal (RAR). First, based on the fact that the features of ring artifacts are different in Cartesian and Polar coordinates, the parallel architecture is adopted to construct the deep neural network so that it can extract and exploit the latent features from different domains to improve the performance of ring artifact removal. Besides, the ring artifacts are globally relevant whether in Cartesian or Polar coordinate systems, but convolutional neural networks show inherent shortcomings in modeling long-range dependency. To tackle this problem, this study introduces the novel Mamba mechanism to achieve a global receptive field without incurring high computational complexity. It enables effective capture of the long-range dependency, thereby enhancing the model performance in image restoration and artifact reduction. The experiments on the simulated data validate the effectiveness of the dual-domain parallel neural network and the Mamba mechanism, and the results on two unseen real datasets demonstrate the promising performance of the proposed RAR algorithm in eliminating ring artifacts and recovering image details. Yikun Zhang 0001, Guannan Liu 0002, Shipeng Xie, Jiabing Gu, Zujian Huang, Tianling Lyu, Yan Xi, Shouping Zhu, Jian Yang 0009, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 11 |
| 2024 | CLIP and image integrative prompt for anterior mediastinal lesion segmentation in CT imageabstractThe automatic segmentation of anterior mediastinal lesions in enhanced CT imaging is of significant importance in clinical diagnostics. Anterior mediastinal lesions are characterized by various types and blurred boundary, which increases the difficulty of anterior mediastinal lesions segmentation. This study leverages the robust zero-shot classification capability and semantic expression of CLIP to formulate CLIP-prompt that express the semantic correlation between images and text, so that CLIP-prompt guided cross-attention has been proposed. By integrating the CLIP-prompt into the image features through cross-attention, the network can focus more intently on the lesion areas. Additionally, to better capture the unknown categorical features of the images, this paper introduces a learnable image prompt that works in conjunction with an attention module integrated with textual information, thereby enhancing the constraints on the segmentation targets. Finally, to address the blurred boundary of the anterior mediastinal lesion, this study proposes a boundary-enhanced loss. By augmenting the weights of difficult-to-segment edge points, the network is enabled to focus on these challenging boundary areas, consequently improving the segmentation accuracy of these points. Compared to existing state-of-the-art methods, our approach has achieved an overall Dice coefficient of 89.43% and has achieved good performance in terms of ASSD metric for segmentation edges. Su Huang, Danni Ai, Guolin Ma, Jian Yang 0009 |
BIBM | 5 |
| 2024 | PrixMatch: Semi-supervised Network for Multi-modal Medical Image Segmentation with Cross-modal Data Augmentation and Adaptive Prior Knowledge ThresholdingabstractSemi-supervised medical image segmentation has made significant strides, yet most existing methods are confined to single-modality data, limiting both the volume of data and the generalizability of the models. Multi-modal data can provide richer information, expand the dataset and enhance model robustness. However, integrating multi-modal learning into semi-supervised medical image segmentation presents challenges, primarily in how to deal with the scarcity of labels and alignment across different modalities simultaneously. In this paper, we propose PrixMatch, a multi-modal semi-supervised model with a teacher-student strategy for medical image segmentation. Initially, we propose a cross-modal data augmentation strategy, which randomly exchanges image blocks of the same location between different modalities, to guide the student model to learn cross-modal consistency without the need for additional network modules. Secondly, we design a cross-modal adaptive pseudo-label threshold setting strategy, which can align the prior anatomical knowledge of different modalities, and combine the modal-aligned prior knowledge and model learning state to filter the pseudo-labels at the pixel-level, flexibly alleviating the confirmation bias that occurs during semi-supervised training. Experiments demonstrate that PrixMatch achieves a Dice Similarity Coefficient (DSC) of 87.2% on the BTCV (CT) and CHAOS (MR) multi-modal datasets with only 10% labeling ratio, bringing nearly 5.5% improvement over the latest state-of-the-art method. Hong Song 0003, Yucong Lin, Long Shao, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009 |
BIBM | 9 |
| 2024 | Deformation Correction in Laparoscopic Liver Surgical Navigation Using Point Cloud Completion and Biomechanical ModelabstractIn the minimally invasive liver resection surgery, deformation estimation of liver is required to correct the preoperative virtual model to match the intraoperative scenarios, in which the liver deforms due to respiration and surgical operations. Existing methods on liver deformation estimation often struggle to achieve high accuracy when the intraoperative liver surface is limited in size. To overcome the challenge of sparse intraoperative point cloud data and improve the accuracy of liver deformation predictions, this paper introduces an innovative method for estimating liver deformation. This method comprises two main components: intraoperative point cloud completion and liver deformation estimation. Intraoperative point cloud completion uses registration techniques to integrate preoperative topological structures into the intraoperative phase. Liver deformation estimation combines optimization control with biomechanical modeling to accurately align the preoperative liver model with its intraoperative counterpart. Comparative and ablation experiments, as well as investigations into the impact of different completion ratios, were conducted. The results demonstrate that this method effectively utilizes preoperative liver geometric features to enhance intraoperative visualization, even with limited intraoperative data. Additionally, the opti-mization control method provides reliable deformation estimates with acceptable accuracy. This study offers new insights and methodologies for the development of augmented reality surgical navigation systems, contributing to the computer assisted liver surgey. Deqiang Xiao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Hong Song 0003, Jian Yang 0009 |
BIBM | 8 |
| 2024 | N-Gram Swin Transformer for CT Image Super-Resolution
Zhenghao Gao, Danni Ai, Hong Song 0003, Jian Yang 0009 |
ICXR | 5 |
| 2024 | Revisiting class-incremental object detection: An efficient approach via intrinsic characteristics alignment and task decoupling
Liang Bai 0006, Hong Song 0003, Tao Feng 0014, Tianyu Fu 0003, Qingzhe Yu, Jian Yang 0009 |
Expert Syst. Appl. | 6 |
| 2024 | A Novel Long Short-Term Memory Learning Strategy for Object TrackingabstractIn this paper, a novel integrated long short‐term memory (LSTM) network and dynamic update model are proposed for long‐term object tracking in video images. The LSTM network tracking method is introduced to improve the effect of tracking failure caused by target occlusion. Stable tracking of the target is achieved using the LSTM method to predict the motion trajectory of the target when it is occluded and dynamically updating the tracking template. First, in target tracking, global average peak‐to‐correlation energy (GAPCE) is used to determine whether the tracking target is blocked or temporarily disappearing such that the follow‐up response tracking strategy can be adjusted accordingly. Second, the data with target motion characteristics are utilized to train the designed LSTM model to obtain an offline model, which effectively predicts the motion trajectory during the period when the target is occluded or has disappeared. Therefore, it can be captured again when the target reappears. Finally, in the dynamic template adjustment stage, the historical information of the target movement is combined, and the corresponding value of the current target is compared with the historical response value to realize the dynamic adjustment of the target tracking template. Compared with the current mainstream efficient convolution operators, namely, the E.T.Track, ToMP, KeepTrack, and RTS algorithms, on the OTB100 and LaSOT datasets, the proposed algorithm increases the distance precision by 9.9% when the distance threshold is 5 pixels, increases the overlap success rate by 0.94% when the overlap threshold is 0.75, and decreases the center location error by 18.9%. The proposed method has higher tracking accuracy and robustness and is more suitable for long‐term tracking of targets in actual scenarios than are the main approaches. Qian Wang 0001, Jian Yang 0009, Hong Song 0003 |
Int. J. Intell. Syst. | 2 |
| 2024 | MSLR: A Self-supervised Representation Learning Method for Tabular Data Based on Multi-scale Ladder Reconstruction
Xutao Weng, Hong Song 0003, Yucong Lin, Bowen Liu 0011, Jian Yang 0009 |
Inf. Sci. | 7 |
| 2024 | Domain base dynamic convolution and distance map guidance for anterior mediastinal lesion segmentation
Su Huang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009 |
Knowl. Based Syst. | 8 |
| 2024 | STQD-Det: Spatio-Temporal Quantum Diffusion Model for Real-Time Coronary Stenosis Detection in X-Ray AngiographyabstractDetecting coronary stenosis accurately in X-ray angiography (XRA) is important for diagnosing and treating coronary artery disease (CAD). However, challenges arise from factors like breathing and heart motion, poor imaging quality, and the complex vascular structures, making it difficult to identify stenosis fast and precisely. In this study, we proposed a Quantum Diffusion Model with Spatio-Temporal Feature Sharing to Real-time detect Stenosis (STQD-Det). Our framework consists of two modules: Sequential Quantum Noise Boxes module and spatio-temporal feature module. To evaluate the effectiveness of the method, we conducted a 4-fold cross-validation using a dataset consisting of 233 XRA sequences. Our approach achieved the F1 score of 92.39% with a real-time processing speed of 25.08 frames per second. These results outperform 17 state-of-the-art methods. The experimental results show that the proposed method can accomplish the stenosis detection quickly and accurately. Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Fusion-competition framework of local topology and global texture for head pose estimation
Tianyu Fu 0003, Kaibin Cao, Jingfan Fan, Deqiang Xiao, Hong Song 0003, Jian Yang 0009 |
Pattern Recognit. | 9 |
| 2024 | Self-supervised local rotation-stable descriptors for 3D ultrasound registration using translation equivariant FCN
Yifan Wang 0037, Tianyu Fu 0003, Jingfan Fan, Deqiang Xiao, Hong Song 0003, Jian Yang 0009 |
Pattern Recognit. | 8 |
| 2024 | Segmentation of 3D Anatomically Diffused Tissues in Magnetic Resonance Images Through Edge-Preserving Constrained Center-Free Fuzzy $C$-MeansabstractAnatomically diffused tissues (ADTs) refer to soft tissues containing many anatomical regions that are spatially dispersed and structurally irregular. In magnetic resonance images, ADTs exhibit blurred morphology and heterogeneous texture, making the accurate extraction of their 3D anatomy challenging. Center-free fuzzy C-means (FCM) can effectively partition nonlinear or nonspherical clusters, providing a promising scheme for ADT segmentation. It solves the uncertainty arising from unreliable center estimation by introducing a similarity criterion. However, the similarity criterion is sensitive to the number of target objects and their adjacent members in the images. Moreover, memberships of the existing algorithms are susceptible to losing real ADT details. To handle these issues, we propose an edge-preserving constrained center-free FCM algorithm for segmenting 3D ADTs in magnetic resonance images. To overcome the sensitivity of the similarity criterion, a novel object-to-cluster similarity measure is first proposed to utilize refined member-toobject adjacency. Specifically, the similarity measure focuses on members in the feature space, which share approximately homogeneous characteristics with each target object. Gradient-domain edge-preserving filtering is then combined with the improved similarity criterion to construct the novel objective function of center-free FCM. With the assistance of the designed imagedriven edge-preserving regularization, the gradient information of clusters is constrained, eventually approaching that of ADTs in the guidance image. Experiments are conducted on two public brain datasets and one local intrahepatic vein dataset. The results demonstrate that the proposed algorithm is more effective for ADT segmentation than the state-of-the-art peers, exhibiting superior generalization capability. Qing Guo 0008, Hong Song 0003, Cong Wang 0033, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009 |
IEEE Trans. Fuzzy Syst. | 8 |
| 2024 | Bi-Fusion of Structure and Deformation at Multi-Scale for Joint Segmentation and RegistrationabstractMedical image segmentation and registration are two fundamental and highly related tasks. However, current works focus on the mutual promotion between the two at the loss function level, ignoring the feature information generated by the encoder-decoder network during the task-specific feature mapping process and the potential inter-task feature relationship. This paper proposes a unified multi-task joint learning framework based on bi-fusion of structure and deformation at multi-scale, called BFM-Net, which simultaneously achieves the segmentation results and deformation field in a single-step estimation. BFM-Net consists of a segmentation subnetwork (SegNet), a registration subnetwork (RegNet), and the multi-task connection module (MTC). The MTC module is used to transfer the latent feature representation between segmentation and registration at multi-scale and link different tasks at the network architecture level, including the spatial attention fusion module (SAF), the multi-scale spatial attention fusion module (MSAF) and the velocity field fusion module (VFF). Extensive experiments on MR, CT and ultrasound images demonstrate the effectiveness of our approach. The MTC module can increase the Dice scores of segmentation and registration by 3.2%, 1.6%, 2.2%, and 6.2%, 4.5%, 3.0%, respectively. Compared with six state-of-the-art algorithms for segmentation and registration, BFM-Net can achieve superior performance in various modal images, fully demonstrating its effectiveness and generalization. Jiaju Zhang, Tianyu Fu 0003, Deqiang Xiao, Jingfan Fan, Hong Song 0003, Danni Ai, Jian Yang 0009 |
IEEE Trans. Image Process. | 7 |
| 2024 | Cross-Anatomy Transfer Learning via Shape-Aware Adaptive Fine-Tuning for 3D Vessel SegmentationabstractDeep learning methods have recently achieved remarkable performance in vessel segmentation applications, yet require numerous labor-intensive labeled data. To alleviate the requirement of manual annotation, transfer learning methods can potentially be used to acquire the related knowledge of tubular structures from public large-scale labeled vessel datasets for target vessel segmentation in other anatomic sites of the human body. However, the cross-anatomy domain shift is a challenging task due to the formidable discrepancy among various vessel structures in different anatomies, resulting in the limited performance of transfer learning. Therefore, we propose a cross-anatomy transfer learning framework for 3D vessel segmentation, which first generates a pre-trained model on a public hepatic vessel dataset and then adaptively fine-tunes our target segmentation network initialized from the model for segmentation of other anatomic vessels. In the framework, the adaptive fine-tuning strategy is presented to dynamically decide on the frozen or fine-tuned filters of the target network for each input sample with a proxy network. Moreover, we develop a Gaussian-based signed distance map that explicitly encodes vessel-specific shape context. The prediction of the map is added as an auxiliary task in the segmentation network to capture geometry-aware knowledge in the fine-tuning. We demonstrate the effectiveness of our method through extensive experiments on two small-scale datasets of coronary artery and brain vessel. The results indicate the proposed method effectively overcomes the discrepancy of cross-anatomy domain shift to achieve accurate vessel segmentation for these two datasets. Danni Ai, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Embedding-Alignment Fusion-Based Graph Convolution Network With Mixed Learning Strategy for 4D Medical Image ReconstructionabstractIn recent years, 4D medical image involving structural and motion information of tissue has attracted increasing attention. The key to the 4D image reconstruction is to stack the 2D slices based on matching the aligned motion states. In this study, the distribution of the 2D slices with the different motion states is modeled as a manifold graph, and the reconstruction is turned to be the graph alignment. An embedding-alignment fusion-based graph convolution network (GCN) with a mixed-learning strategy is proposed to align the graphs. Herein, the embedding and alignment processes of graphs interact with each other to realize a precise alignment with retaining the manifold distribution. The mixed strategy of self- and semi-supervised learning makes the alignment sparse to avoid the mismatching caused by outliers in the graph. In the experiment, the proposed 4D reconstruction approach is validated on the different modalities including Computed Tomography (CT), Magnetic Resonance Imaging (MRI), and Ultrasound (US). We evaluate the reconstruction accuracy and compare it with those of state-of-the-art methods. The experiment results demonstrate that our approach can reconstruct a more accurate 4D image. Tianyu Fu 0003, Hong Song 0003, Jingfan Fan, Deqiang Xiao, Yucong Lin, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Local Contractive Registration With Biomechanical Model: Assessing Microwave Ablation After Compensation for Tissue ShrinkageabstractMicrowave ablation (MWA) is a minimally invasive procedure for the treatment of liver tumor. Accumulating clinical evidence has considered the minimal ablative margin (MAM) as a significant predictor of local tumor progression (LTP). In clinical practice, MAM assessment is typically carried out through image registration of pre- and post-MWA images. However, this process faces two main challenges: non-homologous match between tumor and coagulation with inconsistent image appearance, and tissue shrinkage caused by thermal dehydration. These challenges result in low precision when using traditional registration methods for MAM assessment. In this paper, we present a local contractive nonrigid registration method using a biomechanical model (LC-BM) to address these challenges and precisely assess the MAM. The LC-BM contains two consecutive parts: (1) local contractive decomposition (LC-part), which reduces the incorrect match between the tumor and coagulation and quantifies the shrinkage in the external coagulation region, and (2) biomechanical model constraint (BM-part), which compensates for the shrinkage in the internal coagulation region. After quantifying and compensating for tissue shrinkage, the warped tumor is overlaid on the coagulation, and then the MAM is assessed. We evaluated the method using prospectively collected data from 36 patients with 47 liver tumors, comparing LC-BM with 11 state-of-the-art methods. LTP was diagnosed through contrast-enhanced MR follow-up images, serving as the ground truth for tumor recurrence. LC-BM achieved the highest accuracy (97.9%) in predicting LTP, outperforming other methods. Therefore, our proposed method holds significant potential to improve MAM assessment in MWA surgeries. Dingkun Liu, Danni Ai, Tianyu Fu 0003, Yuanjin Gao, Jingfan Fan, Hong Song 0003, Deqiang Xiao, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | A New Multi-Atlas Based Deep Learning Segmentation Framework With Differentiable Atlas Feature WarpingabstractDeep learning based multi-atlas segmentation (DL-MA) has achieved the state-of-the-art performance in many medical image segmentation tasks, e.g., brain parcellation. In DL-MA methods, atlas-target correspondence is the key for accurate segmentation. In most existing DL-MA methods, such correspondence is usually established using traditional or deep learning based registration methods at image level with no further feature level adaption. This could cause possible atlas-target feature inconsistency. As a result, the information from atlases often has limited positive and even counteractive impact on the final segmentation results. To tackle this issue, in this paper, we propose a new DL-MA framework, where a novel differentiable atlas feature warping module with a new smooth regularization term is presented to establish feature level atlas-target correspondence. Comparing with the existing DL-MA methods, in our framework, atlas features containing anatomical prior knowledge are more relevant to the target image feature, leading the final segmentation results to a high accuracy level. We evaluate our framework in the context of brain parcellation using two public MR brain image datasets: LPBA40 and NIREP-NA0. The experimental results demonstrate that our framework outperforms both traditional multi-atlas segmentation (MAS) and state-of-the-art DL-MA methods with statistical significance. Further ablation studies confirm the effectiveness of the proposed differentiable atlas feature warping module. Huabing Liu, Dong Nie, Jian Yang 0009, Jinda Wang, Zhenyu Tang 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | DSC-Recon: Dual-Stage Complementary 4-D Organ Reconstruction From X-Ray Image Sequence for Intraoperative FusionabstractAccurately reconstructing 4D critical organs contributes to the visual guidance in X-ray image-guided interventional operation. Current methods estimate intraoperative dynamic meshes by refining a static initial organ mesh from the semantic information in the single-frame X-ray images. However, these methods fall short of reconstructing an accurate and smooth organ sequence due to the distinct respiratory patterns between the initial mesh and X-ray image. To overcome this limitation, we propose a novel dual-stage complementary 4D organ reconstruction (DSC-Recon) model for recovering dynamic organ meshes by utilizing the preoperative and intraoperative data with different respiratory patterns. DSC-Recon is structured as a dual-stage framework: 1) The first stage focuses on addressing a flexible interpolation network applicable to multiple respiratory patterns, which could generate dynamic shape sequences between any pair of preoperative 3D meshes segmented from CT scans. 2) In the second stage, we present a deformation network to take the generated dynamic shape sequence as the initial prior and explore the discriminate feature (i.e., target organ areas and meaningful motion information) in the intraoperative X-ray images, predicting the deformed mesh by introducing a designed feature mapping pipeline integrated into the initialized shape refinement process. Experiments on simulated and clinical datasets demonstrate the superiority of our method over state-of-the-art methods in both quantitative and qualitative aspects. Haixiao Geng, Jingfan Fan, Sigeng Chen, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 12 |
| 2023 | IDAA-NET: An Image Domain Adaptive Alignment Network for Unsupervised Liver Vessel Segmentation from CTA ImagesabstractAccurate segmentation of liver vessel from CTA image is important for the diagnosis and treatment of liver diseases. The quality of labeled data directly affects the prediction results of the segmentation model. Compared with CTA image, MRA image has clearer 3D vasculature. Therefore, in order to reduce the reliance of the labeled CTA image which may contain ambiguous vessel contours, we propose a novel unsupervised liver vessel segmentation method based on image domain adaptive alignment network (IDAA-Net) by using labeled MRA and unlabeled CTA images. The IDAA-Net mainly contains three modules: 1) A spatial alignment module (SAM) is introduced to convert MRA image slice to synthetic CTA image slice for achieving spatial alignment of the different modality data in the feature and image levels; 2) An artifact removal module (ARM) is designed to eliminate background artifacts of synthetic CTA from SAM by using the liver label in MRA; 3) An adversarial segmentation module (ASM) is proposed to obtain the optimal segmentation by jointly adversarial learning and supervised learning between the predicted segmentation and the ground-truth label of MRA image. Experiments on the public and private datasets show that our method achieves comparable performance with state-of-the-art supervised method and outperforms the existing unsupervised segmentation methods. Haixiao Geng, Danni Ai, Jingfan Fan, Feng Duan 0001, Yujia Yuan, Jian Yang 0009 |
BIBM | 6 |
| 2023 | CPSS-Net: A Cross-pseudo Semi-supervised Network for Liver Vessel Segmentation from CTA ImagesabstractAccurate segmentation of liver vessel from CTA images is often challenging due to the limited availability of labeled data. In the field of medical image segmentation, semi-supervised learning has garnered significant attention as it utilizes unlabeled data to enhance the training of segmentation models. In this paper, we propose a novel cross-pseudo semi-supervised network (CPSS-Net) based on nnU-Net. The CPSS-Net contains three innovative components: 1) An probability prediction (PP) module is designed to generate probability maps for both labeled and unlabeled datasets, capturing model uncertainty through parallel nnU-Net; 2) A double pseudo-label (DPL) module is used to convert the predicted probability maps into double soft pseudo-labels using an adaptive sharpening function; 3) A cross pseudo-supervised (CPS) module is introduced to learn the mutual consistency of double pseudo-labels. Test experiments on both public and private datasets show that our method achieved a Dice score of 0.67 and a sensitivity score of 0.69, surpassing the segmentation accuracy of existing related methods. Danni Ai, Deqiang Xiao, Feng Duan 0001, Yujia Yuan, Jian Yang 0009 |
BIBM | 6 |
| 2023 | Robust multi-view low-rank embedding clustering
Jian Dai 0002, Hong Song 0003, Yunzhi Luo, Zhenwen Ren, Jian Yang 0009 |
Neural Comput. Appl. | 5 |
| 2023 | Densely Connected U-Net With Criss-Cross Attention for Automatic Liver Tumor Segmentation in CT ImagesabstractAutomatic liver tumor segmentation plays a key role in radiation therapy of hepatocellular carcinoma. In this paper, we propose a novel densely connected U-Net model with criss-cross attention (CC-DenseUNet) to segment liver tumors in computed tomography (CT) images. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intra-slice features of liver tumors. Moreover, the criss-cross attention is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver tumors. We evaluated the proposed CC-DenseUNet on the LiTS dataset and the 3DIRCADb dataset. Experimental results show that the proposed method reaches the state-of-the-art performance for liver tumor segmentation. We further experimentally demonstrate the robustness of the proposed method on a clinical dataset comprising 20 CT volumes. Qiang Li 0049, Hong Song 0003, Zenghui Wei, Fengbo Yang, Jingfan Fan, Danni Ai, Yucong Lin, Xiaoling Yu, Jian Yang 0009 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2023 | M-CSAFN: Multi-Color Space Adaptive Fusion Network for Automated Port-Wine Stains SegmentationabstractAutomatic segmentation of port-wine stains (PWS) from clinical images is critical for accurate diagnosis and objective assessment of PWS. However, this is a challenging task due to the color heterogeneity, low contrast, and indistinguishable appearance of PWS lesions. To address such challenges, we propose a novel multi-color space adaptive fusion network (M-CSAFN) for PWS segmentation. First, a multi-branch detection model is constructed based on six typical color spaces, which utilizes rich color texture information to highlight the difference between lesions and surrounding tissues. Second, an adaptive fusion strategy is used to fuse complementary predictions, which address the significant differences within the lesions caused by color heterogeneity. Third, a structural similarity loss with color information is proposed to measure the detail error between predicted lesions and truth lesions. Additionally, a PWS clinical dataset consisting of 1413 image pairs was established for the development and evaluation of PWS segmentation algorithms. To verify the effectiveness and superiority of the proposed method, we compared it with other state-of-the-art methods on our collected dataset and four publicly available skin lesion datasets (ISIC 2016, ISIC 2017, ISIC 2018, and PH2). The experimental results show that our method achieves remarkable performance in comparison with other state-of-the-art methods on our collected dataset, achieving 92.29% and 86.14% on Dice and Jaccard metrics, respectively. Comparative experiments on other datasets also confirmed the reliability and potential capability of M-CSAFN in skin lesion segmentation. Jinrong Mu, Yucong Lin, Xianqi Meng, Jingfan Fan, Danni Ai, Defu Chen, Haixia Qiu, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 8 |
| 2022 | 3D ARCNN: An Asymmetric Residual CNN for Decreasing False Positive Rate of Lung Nodules DetectionabstractLung cancer is with the highest morbidity and mortality, and early detection of cancerous changes is essential to reduce the risk of death. To achieve this, it is necessary to reduce the false positive rate of detection. In this paper, we propose a novel asymmetric residual network, called 3D ARCNN, to reduce false positive rate of lung nodules detection. 3D ARCNN consists of asymmetric convolutional and multilayer cascaded residual network structures. To solve the problem of deep neural network with large amounts of parameters and poor reproduction ability, the proposed model uses asymmetric convolution to reduce model parameters and enhance the generalization ability of the model. In addition, the model uses an internally cascaded multi-stage residual to prevent the gradient vanishing and exploding problems of deep networks. Experiments are performed on the public dataset LUNA16. Our method achieved high detection sensitivity of 91.6%, 92.7%, 93.2% and 95.8% at 1, 2, 4 and 8 false positives per scan, respectively, which got an average CPM index of 0.912. Experimental results show that the proposed 3D ARCNN is very useful for reducing the false positive rate of lung nodules in the clinic. Bowen Liu 0011, Hong Song 0003, Qiang Li 0049, Yucong Lin, Jian Yang 0009 |
BIBM | 5 |
| 2022 | Denoising of MR and CT images using cascaded multi-supervision convolutional neural networks with progressive training
Hong Song 0003, Lei Chen 0073, Yutao Cui, Qiang Li 0049, Jingfan Fan, Jian Yang 0009, Le Zhang 0004 |
Neurocomputing | 7 |
| 2022 | Volume-awareness and outlier-suppression co-training for weakly-supervised MRI breast mass segmentation with partial annotations
Xianqi Meng, Jingfan Fan, Jinrong Mu, Zongyu Li, Aocai Yang, Kuan Lv, Danni Ai, Yucong Lin, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009 |
Knowl. Based Syst. | 15 |
| 2022 | Portal Vein and Hepatic Vein Segmentation in Multi-Phase MR Images Using Flow-Guided Change DetectionabstractSegmenting portal vein (PV) and hepatic vein (HV) from magnetic resonance imaging (MRI) scans is important for hepatic tumor surgery. Compared with single phase-based methods, multiple phases-based methods have better scalability in distinguishing HV and PV by exploiting multi-phase information. However, these methods just coarsely extract HV and PV from different phase images. In this paper, we propose a unified framework to automatically and robustly segment 3D HV and PV from multi-phase MR images, which considers both the change and appearance caused by the vascular flow event to improve segmentation performance. Firstly, inspired by change detection, flow-guided change detection (FGCD) is designed to detect the changed voxels related to hepatic venous flow by generating hepatic venous phase map and clustering the map. The FGCD uniformly deals with HV and PV clustering by the proposed shared clustering, thus making the appearance correlated with portal venous flow robustly delineate without increasing framework complexity. Then, to refine vascular segmentation results produced by both HV and PV clustering, interclass decision making (IDM) is proposed by combining the overlapping region discrimination and neighborhood direction consistency. Finally, our framework is evaluated on multi-phase clinical MR images of the public dataset (TCGA) and local hospital dataset. The quantitative and qualitative evaluations show that our framework outperforms the existing methods. Qing Guo 0008, Hong Song 0003, Jingfan Fan, Danni Ai, Yuanjin Gao, Xiaoling Yu, Jian Yang 0009 |
IEEE Trans. Image Process. | 7 |
| 2022 | Few-Shot Learning for Deformable Medical Image Registration With Perception-Correspondence Decoupling and Reverse TeachingabstractDeformable medical image registration estimates corresponding deformation to align the regions of interest (ROIs) of two images to a same spatial coordinate system. However, recent unsupervised registration models only have correspondence ability without perception, making misalignment on blurred anatomies and distortion on task-unconcerned backgrounds. Label-constrained (LC) registration models embed the perception ability via labels, but the lack of texture constraints in labels and the expensive labeling costs causes distortion internal ROIs and overfitted perception. We propose the first few-shot deformable medical image registration framework, Perception-Correspondence Registration (PC-Reg), which embeds perception ability to registration models only with few labels, thus greatly improving registration accuracy and reducing distortion. 1) We propose the Perception-Correspondence Decoupling which decouples the perception and correspondence actions of registration to two CNNs. Therefore, independent optimizations and feature representations are available avoiding interference of the correspondence due to the lack of texture constraints. 2) For few-shot learning, we propose Reverse Teaching which aligns labeled and unlabeled images to each other to provide supervision information to the structure and style knowledge in unlabeled images, thus generating additional training data. Therefore, these data will reversely teach our perception CNN more style and structure knowledge, improving its generalization ability. Our experiments on three datasets with only five labels demonstrate that our PC-Reg has competitive registration accuracy and effective distortion-reducing ability. Compared with LC-VoxelMorph( λ = 1), we achieve the 12.5%, 6.3% and 1.0% Reg-DSC improvements on three datasets, revealing our framework with great potential in clinical application. Yuting He 0001, Rongjun Ge, Jian Yang 0009, Youyong Kong, Huazhong Shu, Guanyu Yang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | MVSGAN: Spatial-Aware Multi-View CMR Fusion for Accurate 3D Left Ventricular Myocardium SegmentationabstractThe accurate 3D left ventricular (LV) myocardium segmentation in short-axis (SAX) view of cardiac magnetic resonance (CMR) is challenged by the sparse spatial structure of CMR. The strategy of multi-view CMR fusion can provide fine-grained spatial structure for accurate segmentation. However, the large information misalignment and lack of dense 3D CMR as fusion target in multi-view CMR fusion, and the different spatial resolution between the fusion result and the ground truth in segmentation limit the strategy. In this study, we propose a multi-view spatial-aware adversarial network (MVSGAN). It studies the perception of fine-grained cardiac structure for accurate segmentation by the spatialaware multi-view CMR fusion. It consists of three modules: (1) A residual adversarial fusion (RAF) module takes inter-slices deep correlation and anatomical prior to refine the spatial structures by residual supplement and adversarial optimization. (2) A structural perception-aggregation (SPA) module establishes the spatial correlation between the dense cardiac model and sparse label for accurate CMR LV myocardium segmentation. (3) A joint training strategy utilizes the dense SAX volume as explicit and implicit goals to jointly optimize the framework. The experiments are applied on a public dataset and a clinical dataset to evaluate the performance of MVSGAN. The average Dice and Jaccard score of LV myocardium segmentation obtained by MVSGAN are highest among seven existing state-of-the-art methods, which are up to 0.92 and 0.75. It is concluded that the spatial-aware multi-view CMR fusion can provide meaningful spatial correlation for accurate LV myocardium segmentation. Xiaoming Qi, Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Jian Yang 0009, Wangyag Liu, Yinsu Zhu, Yi Xu 0001, Huazhong Shu, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | CC-DenseUNet: Densely Connected U-Net with Criss-Cross Attention for Liver and Tumor Segmentation in CT VolumesabstractThe automatic segmentation of liver and tumor is important for hepatic tumor surgery. In this paper, we propose a novel densely connected U-Net (CC-DenseUNet), which integrates criss-cross attention (CCA) module, to segment the liver and tumor in computed tomography (CT) volumes. The dense interconnections in CC-DenseUNet ensure the maximum information flow between encoder layers when extracting intraslice features of liver and tumors. Moreover, the CCA module is used in CC-DenseUNet to efficiently capture only the necessary and meaningful non-local contextual information of CT images containing liver or tumors. We evaluated the proposed CCDenseUNet on the Liver Tumor Segmentation Challenge and 3DIRCADb datasets. Experimental results show that our method outperformed the state-of-the-art methods in liver tumor segmentation and achieved a highly competitive performance in liver segmentation. Qiang Li 0049, Hong Song 0003, Jingfan Fan, Danni Ai, Yucong Lin, Jian Yang 0009 |
BIBM | 7 |
| 2021 | Cross-Domain Transfer Learning for Vessel Segmentation in Computed Tomographic Coronary Angiographic Images
Ruirui An, Danni Ai, Yongtian Wang, Jian Yang 0009 |
ICIG (2) | 6 |
| 2021 | Novel Augmented Reality System for Oral and Maxillofacial Surgery
Lele Ding, Long Shao, Zehua Zhao, Tao Zhang 0152, Danni Ai, Jian Yang 0009, Yongtian Wang |
ICIG (2) | 6 |
| 2021 | An automatic framework for endoscopic image restoration and enhancement
Muhammad Asif 0016, Lei Chen 0073, Hong Song 0003, Jian Yang 0009, Alejandro F. Frangi |
Appl. Intell. | 4 |
| 2021 | GSCFN: A graph self-construction and fusion network for semi-supervised brain tissue segmentation in MRI
Yan Zhang 0094, Youyong Kong, Jiasong Wu, Jian Yang 0009, Huazhong Shu, Gouenou Coatrieux |
Neurocomputing | 5 |
| 2021 | Meta grayscale adaptive network for 3D integrated renal structures segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Rongjun Ge, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2021 | Divergence-Free Fitting-Based Incompressible Deformation Quantification of LiverabstractLiver is an incompressible organ that maintains its volume during the respiration-induced deformation. Quantifying this deformation with the incompressible constraint is significant for liver tracking. The constraint can be accomplished with retaining the divergence-free field obtained by the deformation decomposition. However, the decomposition process is time-consuming, and the removal of non-divergence-free field weakens the deformation. In this study, a divergence-free fitting-based registration method is proposed to quantify the incompressible deformation rapidly and accurately. First, the deformation to be estimated is mapped to the velocity in a diffeomorphic space. Then, this velocity is decomposed by a fast Fourier-based Hodge-Helmholtz decomposition to obtain the divergence-free, curl-free, and harmonic fields. The curl-free field is replaced and fitted by the obtained harmonic field with a translation field to generate a new divergence-free velocity. By optimizing this velocity, the final incompressible deformation is obtained. Moreover, a deep learning framework (DLF) is constructed to accelerate the incompressible deformation quantification. An incompressible respiratory motion model is built for the DLF by using the proposed registration method and is then used to augment the training data. An encoder-decoder network is introduced to learn appearance-velocity correlation at patch scale. In the experiment, we compare the proposed registration with three state-of-the-art methods. The results show that the proposed method can accurately achieve the incompressible registration of liver with a mean liver overlap ratio of 95.33%. Moreover, the time consumed by DLF is nearly 15 times shorter than that by other methods. Tianyu Fu 0003, Jingfan Fan, Dingkun Liu, Hong Song 0003, Chaoyi Zhang, Danni Ai, Zhigang Cheng, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | Topological distance-constrained feature descriptor learning model for vessel matching in coronary angiographiesabstractFeature matching technology is vital to establish the association between virtual and real objects in virtual reality and augmented reality systems. Specifically, it provides them with the ability to match a dynamic scene. Many image matching methods, of which most are deep learning-based, have been proposed over the past few decades. However, vessel fracture, stenosis, artifacts, high background noise, and uneven vessel gray-scale make vessel matching in coronary angiography extremely difficult. Traditional matching methods perform poorly in this regard. In this study, a topological distance-constrained feature descriptor learning model is proposed. This model regards the topology of the vasculature as the connection relationship of the centerline. The topological distance combines the geodesic distance between the input patches and constrains the descriptor network by maximizing the feature difference between connected and unconnected patches to obtain more useful potential feature relationships. Matching patches of different sequences of angiographic images are generated for the experiments. The matching accuracy and stability of the proposed method is superior to those of the existing models. The proposed method solves the problem of matching coronary angiographies by generating a topological distance-constrained feature descriptor. Xiaojiao Song, Jingfan Fan, Danni Ai, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 5 |
| 2020 | A General Endoscopic Image Enhancement Method Based on Pre-trained Generative Adversarial NetworksabstractEndoscopic images frequently have image quality problems due to the limitations of surgical instruments and the impact of surgical operations, such as uneven illumination, smogginess and color deviation. For deep learning based on enhancement methods, independent training lacks sufficient defect images and generalization capability, and combined training with mixture of data cannot identify diverse specific tasks. To address these issues, we propose a general method based on pre-trained generative adversarial network with a specified transfer learning strategy to obtain high-quality images. Initially, we independently train a standard network based on a universal task, e.g., uneven illumination, where a pre-trained model is extracted as a backbone with partially shared generator. Then, we transfer the backbone to more potential image enhancement tasks. Experiments on uneven illumination, smogginess, and color deviation indicate that the model successfully shares common features of high-quality images and responds specifically to different defects as well. Jingfan Fan, Danni Ai, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
BIBM | 6 |
| 2020 | Local Contractive Registration for Quantification of Tissue Shrinkage in Assessment of Microwave Ablation
Dingkun Liu, Tianyu Fu 0003, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009 |
MICCAI (3) | 6 |
| 2020 | Classification of Retinal Vessels into Artery-Vein in OCT Angiography Guided by Fundus Images
Jianyang Xie, Yonghuai Liu, Yalin Zheng, Pan Su 0001, Jian Yang 0009, Jiang Liu 0001, Yitian Zhao |
MICCAI (6) | 6 |
| 2020 | Coarse-to-fine classification for diabetic retinopathy grading using convolutional neural network
Zhan Wu, Gonglei Shi, Yang Chen 0008, Xinjian Chen 0001, Gouenou Coatrieux, Jian Yang 0009, Limin Luo 0001, Shuo Li 0001 |
Artif. Intell. Medicine | 7 |
| 2020 | Vessel Structure Extraction using Constrained Minimal Path Propagation
Guanyu Yang 0001, Tianling Lv, Yunpeng Shen, Shuo Li 0001, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux |
Artif. Intell. Medicine | 5 |
| 2020 | Prior information constrained alternating direction method of multipliers for longitudinal compressive sensing MR imaging
Ruirui Kang, Danni Ai, Gangrong Qu, Qingbo Li, Yurong Jiang, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
Neurocomputing | 10 |
| 2020 | Cerebrovascular segmentation from TOF-MRA using model- and data-driven method via sparse labelsabstractCerebrovascular segmentation from time-of-flight magnetic resonance angiography (TOF-MRA) data is of great importance in blood supply structure analysis, diagnosis, and treatment of cerebrovascular pathologies. However, complete and accurate segmentation is still a challenge due to the complex image context and vascular morphology. The existing model-driven methods are often difficult to obtain prominent accuracy and robustness. Deep-learning based technology has achieved unimaginable success, but always faces the problem of insufficient labeled data. In this paper, a novel strategy is proposed to automate cerebrovascular segmentation, which integrates model- and data-driven methods. Firstly, the TOF-MRA data are sparsely labeled by three radiologists. Secondly, a semi-supervised mixture probability model is proposed to fit the cerebrovascular intensity distribution precisely, which starts from the sparse annotations and generates massive labeled points. Thirdly, mislabeled points are corrected by a Clean-Mechanism model, to acquire a well-labeled point-set of good quality. Finally, we construct and train a dilated dense convolution network (DD-CNN) by the resultant labeled point-set. The proposed method is validated on 109 clinical TOR-MRA data from a public dataset. Compared with the other state-of-the-art segmentation methods, our method segments cerebrovascular structure with better completeness and sensibility, especially for slender vascularity. The experimental results show that our method reaches an average dice score of 93.20%, which also indicates that the DD-CNN is very competent for cerebrovascular segmentation from TOF-MRA volume. Baochang Zhang 0003, Shoujun Zhou, Jian Yang 0009, Na Li 0048, Zonghan Wu, Jun Xia 0002 |
Neurocomputing | 4 |
| 2020 | Groupwise registration with global-local graph shrinkage in atlas construction
Tianyu Fu 0003, Jian Yang 0009, Danni Ai, Hong Song 0003, Yurong Jiang, Yongtian Wang, Alejandro F. Frangi |
Medical Image Anal. | 2 |
| 2020 | Dense biased networks with deep priori anatomy and hard region adaptation: Semi-supervised learning for fine renal artery segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2020 | Discriminative feature representation for Noisy image quality assessment
Yunbo Gu, Tianling Lv, Yang Chen 0008, Lu Zhang 0037, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux |
Multim. Tools Appl. | 7 |
| 2020 | Spatial probabilistic distribution map-based two-channel 3D U-net for visual pathway segmentation
Danni Ai, Zhiqi Zhao, Jingfan Fan, Hong Song 0003, Xiaoxia Qu, Junfang Xian, Jian Yang 0009 |
Pattern Recognit. Lett. | 7 |
| 2020 | Topology Optimization Using Multiple-Possibility Fusion for Vasculature ExtractionabstractVascular centerline extraction from angiography images plays an important role in computer-aided diagnosis of vascular disease. To solve the common problems related to noise and inconsistent vasculatures from uneven perfusion, this paper proposes an automatic framework for accurate vascular centerline extraction from angiograms that uses multi-probability fusion-based topology optimization. In this framework, vascular region is first segmented using a learning-based method. Then, initial centerlines are obtained by applying iterative filtering operation and multi-direction indexed non-maximum suppression. Topology optimization is achieved by gap filling. A connection probability map is constructed utilizing the information of initial centerlines, texture, and orientation of vasculatures. Shortest path tracking is employed to search for optimal connections around gaps in the initial centerlines. The proposed framework is evaluated using simulative and clinical coronary angiographies. The experimental results demonstrate that the proposed method can extract centerlines with F1 score of 97.28% ± 1.2% for vasculatures in 12 clinical angiographic images. It is evident that the proposed method can extract complete and accurate vascular centerlines from angiograms and can be used to repair gaps in other filamentary structures, such as roads and retinal blood vessels. This endows our method a great potential in the analysis of filamentary structures. Huihui Fang, Danni Ai, Weijian Cong, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2020 | Greedy Soft Matching for Vascular Tracking of Coronary Angiographic Image SequencesabstractVascular tracking of coronary angiographic image sequences is one of the most clinically important tasks in the diagnostic assessment and interventional guidance of cardiac disease. It is difficult to automate this application because the vascular structure is complex; moreover, unsatisfactory angiography image quality may exacerbate the difficulty of vasculature extraction. This paper converts vascular tracking into branch matching and proposes a novel and automatic greedy soft match algorithm. Our method is based on a graph framework. A graph model building module is proposed to represent the vascular structure. Then, a greedy branch searching method is adopted to acquire all possible paths in the graph that may match the reference vessel. Finally, a soft batch matching method that combines branch descriptor and dynamic time warping is presented to select the best matching branch. The solution to the problem takes advantage of both spatial and temporal continuity between successive frames. The experimental results demonstrate that the proposed algorithm is effective and robust for vascular tracking. The F1 score of a single branch dataset, which contains 12 angiographic image sequences with 77 angiograms of contrast agent-filled vessels, is 0.89 ± 0.06 and of a vessel tree dataset which contains nine sequences with 58 angiograms is 0.88 ± 0.05. Extensive experimental results well demonstrate the superior performance of the algorithm. In addition, it provides a universal solution to address the problem of filamentary structure tracking. Huihui Fang, Danni Ai, Yong Huang 0002, Yurong Jiang, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2020 | Spatio-Temporal Constrained Online Layer Separation for Vascular Enhancement in X-Ray Angiographic Image SequenceabstractAutomatic vascular enhancement is crucial to vascular structure identification in X-ray angiographic (XRA) image sequences. In this work, we propose a novel spatio-temporal constrained online layer separation (STOLS) method to achieve vascular enhancement in XRA image sequences. The proposed method integrates the motion consistency of structures into the temporal-constrained online robust principal component analysis (ORPCA) to remove quasi-static structures (e.g., bones) from the enhanced vascular images. Furthermore, smoothing technique is integrated into the spatial-constrained ORPCA to reduce motion artifacts and the noise introduced by non-uniform illumination. To make the proposed method more adaptive to various vascular structures, the spatial-constrained ORPCA is adjusted by an adaptive weight using the proportion of the vessel region in the previous frame. The performance of the proposed method is compared with five state-of-the-art subtraction methods with respect to local and global revised contrast-to-noise ratios (rCNRs) and reconstruction errors. For the proposed method, the local and global rCNRs of the final vessel layer reached 2.54 and 1.24, respectively, while the error between the original and reconstructed images from the respiratory, background, and vessel layer reached 0.0354. The proposed STOLS can enhance the angiograms in a real-time and online manner without fine-tuning parameters, and can thus be used for intra-operation diagnosis and interventional procedures of coronary artery diseases. Shuang Song 0005, Chenbing Du, Danni Ai, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2019 | Liver Segmentation in CT Images Using a Non-Local Fully Convolutional Neural NetworkabstractLiver segmentation is a critical step in diagnosing various kinds of hepatic diseases. Based on the segmentation results, physicians can make further assessments more accurately. Although deep learning methods have achieved excellent performance in liver segmentation tasks, the traditional convolution encoder-decoder architecture may easily loss the spatial information due to the stacked convolution and pooling layers. In this paper, we present a non-local spatial feature based neural network (referred as NL-Net) to learn more spatial features of liver for more accurate segmentation. The NL-Net consists of an encoder block, a non-local spatial feature learning block and a decoder block. We utilized the pretrained ResNet model with transfer learning as the encoder. The non-local block can learn long range dependencies of the liver pixel position by computing the response at a position as a weighted sum of the responses at all positions, which can help the network learn more robust features. We applied the proposed model to ISBI 2019 CHAOs liver Segmentation Challenge task and evaluated it on the testing set. Experimental results show that the proposed NL-Net achieved an average dice of 0.972, RAVD of 1.593, ASSD of 1.926 and MSSD of 110.658 on the segmentation results. Lei Chen 0073, Hong Song 0003, Qiang Li 0049, Yutao Cui, Jian Yang 0009, Xiaohua Hu 0001 |
BIBM | 5 |
| 2019 | Monte Carlo Tree Search for 3D/2D Registration of Vessel Graphsabstract3D/2D registration techniques can compensate for the deficiencies of X-ray angiography-based navigation in vascular interventional surgery, such as the lack of depth information and excessive use of contrast agents. In this study, we propose a novel Monte Carlo tree search-based 3D/2D vessel graph registration method. The registration problem is transferred to a tree search problem according to the topology of vessel centerlines. Then, the Monte Carlo tree search method is applied to find the optimal vessel matching associated with highest registration score. Experiments on uninitialized vessel data demonstrate that the proposed method can achieve the highest accuracy among four state-of-the-art methods. An average accuracy of 1.91 mm on clinical coronary artery data is obtained. For the independence of initial pose and robustness to noise, the proposed method can align 3D and 2D vessels without prior initialization in vascular interventional surgery. Shuang Song 0005, Danni Ai, Jingfan Fan, Hong Song 0003, Jian Yang 0009 |
BIBM | 8 |
| 2019 | Spatial Probabilistic Distribution Map Based 3D FCN for Visual Pathway Segmentation
Zhiqi Zhao, Danni Ai, Jingfan Fan, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
ICIG (2) | 7 |
| 2019 | 3D Convolutional Two-Stream Network for Action Recognition in VideosabstractIn recent years, action recognition based on two-stream networks has developed rapidly. However, most existing methods describe incomplete and distorted video content due to cropped and warped frame or clip-level feature extraction. This paper proposed an approach based on deep learning that preserves the complete contextual relation of temporal human actions in videos. The proposed architecture follows the two-stream network with a novel 3D Convolutional Network (ConvNets) and pyramid pooling layer, to design an end-to-end behavioral feature learning method. The 3D ConvNets extract video-level, spatial-temporal features from two input streams, the RGB images and the corresponding optical flow. The multi-scale pyramid pooling layer fixed the generated feature maps into a unified size regardless of input video size. The final predictions are resulted from a fused softmax scores of two streams, and subject to the weighting factor of each stream. Our experimental results suggest spatial stream slightly higher than the temporal stream, and the performance of the trained model is conditionally optimized. The proposed method is experimented on two challenging video action datasets UCF101 and HMDB51, in which our method achieves the most advanced performance above 96.1% on UCF101 dataset. Yuezhu Qi, Jian Yang 0009, Junxing Ren, Hong Du |
ICTAI | 3 |
| 2019 | Fusing RFID and Computer Vision for Occlusion-Aware Object Identifying and Tracking
Jian Yang 0009, Hong Du |
WASA | 4 |
| 2019 | Liver tumor segmentation in CT volumes using an adversarial densely connected networkabstractBACKGROUND: Malignant liver tumor is one of the main causes of human death. In order to help physician better diagnose and make personalized treatment schemes, in clinical practice, it is often necessary to segment and visualize the liver tumor from abdominal computed tomography images. Due to the large number of slices in computed tomography sequence, developing an automatic and reliable segmentation method is very favored by physicians. However, because of the noise existed in the scan sequence and the similar pixel intensity of liver tumors with their surrounding tissues, besides, the size, position and shape of tumors also vary from one patient to another, automatic liver tumor segmentation is still a difficult task. RESULTS: We perform the proposed algorithm to the Liver Tumor Segmentation Challenge dataset and evaluate the segmentation results. Experimental results reveal that the proposed method achieved an average Dice score of 68.4% for tumor segmentation by using the designed network, and ASD, MSD, VOE and RVD improved from 27.8 to 21, 147 to 124, 0.52 to 0.46 and 0.69 to 0.73, respectively after performing adversarial training strategy, which proved the effectiveness of the proposed method. CONCLUSIONS: The testing results show that the proposed method achieves improved performance, which corroborated the adversarial training based strategy can achieve more accurate and robustness results on liver tumor segmentation task. Lei Chen 0073, Hong Song 0003, Chi Wang 0004, Yutao Cui, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004 |
BMC Bioinform. | 5 |
| 2019 | Vessel segmentation using centerline constrained level set method
Tianling Lv, Guanyu Yang 0001, Yudong Zhang 0001, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001 |
Multim. Tools Appl. | 4 |
| 2019 | A mobilized automatic human body measure system using neural network
Likun Xia, Jian Yang 0009, Huiming Xu, Yitian Zhao, Yongtian Wang |
Multim. Tools Appl. | 2 |
| 2019 | Patch-Based Adaptive Background Subtraction for Vascular Enhancement in X-Ray CineangiogramsabstractOBJECTIVE: Automatic vascular enhancement in X-ray cineangiography is of crucial interest, for instance, for better visualizing and quantifying coronary arteries in diagnostic and interventional procedures. METHODS: A novel patch-based adaptive background subtraction method (PABSM) is proposed automatically enhancing vessels in coronary X-ray cineangiography. First, pixels in the cineangiogram are described by the vesselness and Gabor features. Second, a classifier is utilized to separate the cineangiogram into the rough vascular and non-vascular region. Dilation is applied to the classified binary image to include more vascular region. Third, a patch-based background synthesis is utilized to fill the removed vascular region. RESULTS: A database containing 320 cineangiograms of 175 patients was collected, and then an interventional cardiologist annotated all vascular structures. The performance of PABSM is compared with six state-of-the-art vascular enhancement methods regarding the precision-recall curve and C-value. The area under the precision-recall curve is 0.7133, and the C-value is 0.9659. CONCLUSION: PABSM can automatically enhance the coronary artery in the cineangiograms. It preserves the integrity of vascular topological structures, particularly in complex vascular regions, and removes noise caused by the non-uniform gray-level distribution in the cineangiogram. SIGNIFICANCE: PABSM can avoid the motion artifacts and it eases the subsequent vascular segmentation, which is crucial for the diagnosis and interventional procedures of coronary artery diseases. Shuang Song 0005, Alejandro F. Frangi, Jian Yang 0009, Danni Ai, Chenbing Du, Yong Huang 0002, Hong Song 0003, Luosha Zhang, Yechen Han, Yongtian Wang |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Domain Progressive 3D Residual Convolution Network to Improve Low-Dose CT ImagingabstractThe wide applications of X-ray computed tomography (CT) bring low-dose CT (LDCT) into a clinical prerequisite, but reducing the radiation exposure in CT often leads to significantly increased noise and artifacts, which might lower the judgment accuracy of radiologists. In this paper, we put forward a domain progressive 3D residual convolution network (DP-ResNet) for the LDCT imaging procedure that contains three stages: sinogram domain network (SD-net), filtered back projection (FBP), and image domain network (ID-net). Though both are based on the residual network structure, the SD-net and ID-net provide complementary effect on improving the final LDCT quality. The experimental results with both simulated and real projection data show that this domain progressive deep-learning network achieves significantly improved performance by combing the network processing in the two domains. Xiangrui Yin, Jean-Louis Coatrieux, Qianlong Zhao, Jin Liu 0019, Wei Yang 0006, Jian Yang 0009, Guotao Quan, Yang Chen 0008, Huazhong Shu, Limin Luo 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Inter/Intra-Constraints Optimization for Fast Vessel Enhancement in X-ray Angiographic Image Sequence
Chenbing Du, Shuang Song 0005, Danni Ai, Hong Song 0003, Yong Huang 0002, Yongtian Wang, Jian Yang 0009 |
BIBM | 7 |
| 2018 | Automatic Liver Segmentation Using Multi-plane Integrated Fully Convolutional Neural Networks
Chi Wang 0004, Hong Song 0003, Lei Chen 0073, Qiang Li 0049, Jian Yang 0009, Xiaohua Hu 0001, Le Zhang 0004 |
BIBM | 5 |
| 2018 | K-mer Counting: memory-efficient strategy, parallel computing and field of application for Bioinformatics
Ming Xiao 0002, Song Hong, Yongtao Yang, Jianxin Wang 0001, Jian Yang 0009, Wenbiao Ding, Le Zhang 0004 |
BIBM | 7 |
| 2018 | Local statistical deformation models for deformable image registration
Songyuan Tang, Weijian Cong, Jian Yang 0009, Tianyu Fu 0003, Hong Song 0003, Danni Ai, Yongtian Wang |
Neurocomputing | 3 |
| 2018 | Structure-Adaptive Fuzzy Estimation for Random-Valued Impulse Noise SuppressionabstractNoise detection accuracy is crucial in suppressing random-valued impulse noise. Both false and miss detections determine the final estimation performance. Deterministic detection methods, which distinctly classify pixels into noisy or uncorrupted pixels, tend to increase the estimation error because some uncorrupted edge points are hard to discriminate from the random-valued impulse noise points. This paper proposes an iterative structure-adaptive fuzzy estimation (SAFE) for random-valued impulse noise suppression. This SAFE method is developed in the framework of Gaussian maximum likelihood estimation. The structure-adaptive fuzziness is reflected by two structure-adaptive metrics based on pixel reliability and patch similarity, respectively. The reliability metric for each pixel (as noise free) is estimated via a novel-minimal-path-based structure propagation to give full consideration of the spatially varying image structures. A robust iteration stopping strategy is also proposed by evaluating the reestimation error of the uncorrupted intensity information. The comparative experimental results show that the proposed structure-adaptive fuzziness can lead to effective restoration. An efficient implementation of this SAFE method is also realized via graphics-processing-unit-based parallelization. Yang Chen 0008, Yudong Zhang 0001, Huazhong Shu, Jian Yang 0009, Limin Luo 0001, Jean-Louis Coatrieux, Qianjin Feng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | 3D Feature Constrained Reconstruction for Low-Dose CT ImagingabstractLow-dose computed tomography (LDCT) images are often highly degraded by amplified mottle noise and streak artifacts. Maintaining image quality under low-dose scan protocols is a well-known challenge. Recently, sparse representation-based techniques have been shown to be efficient in improving such CT images. In this paper, we propose a 3D feature constrained reconstruction (3D-FCR) algorithm for LDCT image reconstruction. The feature information used in the 3D-FCR algorithm relies on a 3D feature dictionary constructed from available high quality standard-dose CT sample. The CT voxels and the sparse coefficients are sequentially updated using an alternating minimization scheme. The performance of the 3D-FCR algorithm was assessed through experiments conducted on phantom simulation data and clinical data. A comparison with previously reported solutions was also performed. Qualitative and quantitative results show that the proposed method can lead to a promising improvement of LDCT image quality. Jin Liu 0019, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001, Qianjing Feng, Zhiguo Gui, Gouenou Coatrieux |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Robust Stereoscopic Crosstalk PredictionabstractWe propose a new metric to predict perceived crosstalk using the original images rather than both the original and ghosted images. The proposed metrics are based on color information. First, we extract a disparity map, a color difference map, and a color contrast map from original image pairs. Then, we use those maps to construct two new metrics (Vdispc and Vdlogc). Metric Vdispc considers the effect of the disparity map and the color difference map, while Vdlogc addresses the influence of the color contrast map. The prediction performance is evaluated using various types of stereoscopic crosstalk images. By incorporating Vdispc and Vdlogc, the new metric Vpdlc is proposed to achieve a higher correlation with the perceived subject crosstalk scores. Experimental results show that the new metrics achieve better performance than previous methods, which indicate that color information is one key factor for crosstalk visible prediction. Furthermore, we construct a new data set to evaluate our new metrics. Jianbing Shen, Yan Zhang 0094, Zhiyuan Liang, Chang Liu 0071, Hanqiu Sun, Xiaopeng Hao, Jianhong Liu, Jian Yang 0009, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2018 | Video Saliency Detection Using Object ProposalsabstractIn this paper, we introduce a novel approach to identify salient object regions in videos via object proposals. The core idea is to solve the saliency detection problem by ranking and selecting the salient proposals based on object-level saliency cues. Object proposals offer a more complete and high-level representation, which naturally caters to the needs of salient object detection. As well as introducing this novel solution for video salient object detection, we reorganize various discriminative saliency cues and traditional saliency assumptions on object proposals. With object candidates, a proposal ranking and voting scheme, based on various object-level saliency cues, is designed to screen out nonsalient parts, select salient object regions, and to infer an initial saliency estimate. Then a saliency optimization process that considers temporal consistency and appearance differences between salient and nonsalient regions is used to refine the initial saliency estimates. Our experiments on public datasets (SegTrackV2, Freiburg-Berkeley Motion Segmentation Dataset, and Densely Annotated Video Segmentation) validate the effectiveness, and the proposed method produces significant improvements over state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Ling Shao 0001, Jian Yang 0009, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Cybern. | 5 |
| 2018 | Multimodal Fusion With Reference: Searching for Joint Neuromarkers of Working Memory Deficits in SchizophreniaabstractBy exploiting cross-information among multiple imaging data, multimodal fusion has often been used to better understand brain diseases. However, most current fusion approaches are blind, without adopting any prior information. There is increasing interest to uncover the neurocognitive mapping of specific clinical measurements on enriched brain imaging data; hence, a supervised, goal-directed model that employs prior information as a reference to guide multimodal data fusion is much needed and becomes a natural option. Here, we proposed a fusion with reference model called "multi-site canonical correlation analysis with reference + joint-independent component analysis" (MCCAR+jICA), which can precisely identify co-varying multimodal imaging patterns closely related to the reference, such as cognitive scores. In a three-way fusion simulation, the proposed method was compared with its alternatives on multiple facets; MCCAR+jICA outperforms others with higher estimation precision and high accuracy on identifying a target component with the right correspondence. In human imaging data, working memory performance was utilized as a reference to investigate the co-varying working memory-associated brain patterns among three modalities and how they are impaired in schizophrenia. Two independent cohorts (294 and 83 subjects respectively) were used. Similar brain maps were identified between the two cohorts along with substantial overlaps in the central executive network in fMRI, salience network in sMRI, and major white matter tracts in dMRI. These regions have been linked with working memory deficits in schizophrenia in multiple reports and MCCAR+jICA further verified them in a repeatable, joint manner, demonstrating the ability of the proposed method to identify potential neuromarkers for mental disorders. Shile Qi, Vince D. Calhoun, Theo G. M. van Erp, Juan R. Bustillo, Eswar Damaraju, Jessica A. Turner, Yuhui Du, Jian Yang 0009, Jiayu Chen 0003, Qingbao Yu, Daniel H. Mathalon, Judith M. Ford, James Voyvodic, Bryon A. Mueller, Aysenil Belger, Sarah C. McEwen, Steven G. Potkin, Adrian Preda, Tianzi Jiang, Jing Sui |
IEEE Trans. Medical Imaging | 8 |
| 2017 | Dorsal hand vein recognition based on convolutional neural networksabstractIn this paper, we proposed a dorsal hand vein recognition method based on Convolutional Neural Network (CNN), compared the recognition rate of different depth CNN models and analyzed the influence of dataset size on dorsal hand vein recognition rate. Firstly, the region of interest (ROI) of dorsal hand vein images was extracted, and contrast limited adaptive histogram equalization (CLAHE) and Gaussian smoothing filter algorithm were used to preprocess the images. Then Reference-CaffeNet AlexNet and VGG depth CNN were trained to extract image feature. Finally, logistic regression was applied for identification. The experimental results on two different size of dataset shown that the depth of network and size of data set size have different degree effect on recognition rate, the dorsal hand vein recognition rate based on VGG-19 reaches 99.7%. In this paper, we also explored the feasibility of ensemble learning on SqueezeNet. The recognition rate declined slightly with 99.52%, but the model size has been decreased sharply. Haipeng Wan, Lei Chen 0073, Hong Song 0003, Jian Yang 0009 |
BIBM | 4 |
| 2017 | Sparse-view X-ray CT reconstruction with Gamma regularization
Jian Yang 0009, Yang Chen 0008, Jean-Louis Coatrieux, Limin Luo 0001 |
Neurocomputing | 3 |
| 2017 | Saliency driven vasculature segmentation with infinite perimeter active contour model
Yitian Zhao, Jingliang Zhao, Jian Yang 0009, Yonghuai Liu, Yifan Zhao 0001, Yalin Zheng, Likun Xia, Yongtian Wang |
Neurocomputing | 3 |
| 2017 | Registration and fusion quantification of augmented reality based nasal endoscopic surgery
Yakui Chu, Jian Yang 0009, Shaodong Ma, Danni Ai, Hong Song 0003, Duanduan Chen, Lei Chen 0073, Yongtian Wang |
Medical Image Anal. | 2 |
| 2017 | Discriminative Feature Representation to Improve Projection Data Inconsistency for Low Dose CT ImagingabstractIn low dose computed tomography (LDCT) imaging, the data inconsistency of measured noisy projections can significantly deteriorate reconstruction images. To deal with this problem, we propose here a new sinogram restoration approach, the sinogram- discriminative feature representation (S-DFR) method. Different from other sinogram restoration methods, the proposed method works through a 3-D representation-based feature decomposition of the projected attenuation component and the noise component using a well-designed composite dictionary containing atoms with discriminative features. This method can be easily implemented with good robustness in parameter setting. Its comparison to other competing methods through experiments on simulated and real data demonstrated that the S-DFR method offers a sound alternative in LDCT. Jin Liu 0019, Jianhua Ma 0001, Yi Zhang 0018, Yang Chen 0008, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux, Wei Yang 0006, Qianjin Feng 0004, Wufan Chen |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Intensity and Compactness Enabled Saliency Estimation for Leakage Detection in Diabetic and Malarial RetinopathyabstractLeakage in retinal angiography currently is a key feature for confirming the activities of lesions in the management of a wide range of retinal diseases, such as diabetic maculopathy and paediatric malarial retinopathy. This paper proposes a new saliency-based method for the detection of leakage in fluorescein angiography. A superpixel approach is firstly employed to divide the image into meaningful patches (or superpixels) at different levels. Two saliency cues, intensity and compactness, are then proposed for the estimation of the saliency map of each individual superpixel at each level. The saliency maps at different levels over the same cues are fused using an averaging operator. The two saliency maps over different cues are fused using a pixel-wise multiplication operator. Leaking regions are finally detected by thresholding the saliency map followed by a graph-cut segmentation. The proposed method has been validated using the only two publicly available datasets: one for malarial retinopathy and the other for diabetic retinopathy. The experimental results show that it outperforms one of the latest competitors and performs as well as a human expert for leakage detection and outperforms several state-of-the-art methods for saliency detection. Yitian Zhao, Yalin Zheng, Yonghuai Liu, Jian Yang 0009, Yifan Zhao 0001, Duanduan Chen, Yongtian Wang |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Convex Hull Aided Registration Method (CHARM)abstractNon-rigid registration finds many applications such as photogrammetry, motion tracking, model retrieval, and object recognition. In this paper we propose a novel convex hull aided registration method (CHARM) to match two point sets subject to a non-rigid transformation. First, two convex hulls are extracted from the source and target respectively. Then, all points of the point sets are projected onto the reference plane through each triangular facet of the hulls. From these projections, invariant features are extracted and matched optimally. The matched feature point pairs are mapped back onto the triangular facets of the convex hulls to remove outliers that are outside any relevant triangular facet. The rigid transformation from the source to the target is robustly estimated by the random sample consensus (RANSAC) scheme through minimizing the distance between the matched feature point pairs. Finally, these feature points are utilized as the control points to achieve non-rigid deformation in the form of thin-plate spline of the entire source point set towards the target one. The experimental results based on both synthetic and real data show that the proposed algorithm outperforms several state-of-the-art ones with respect to sampling, rotational angle, and data noise. In addition, the proposed CHARM algorithm also shows higher computational efficiency compared to these methods. Jingfan Fan, Jian Yang 0009, Yitian Zhao, Danni Ai, Yonghuai Liu, Ge Wang 0001, Yongtian Wang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | 3-Points Convex Hull Matching (3PCHM) for fast and robust point set registration
Jingfan Fan, Jian Yang 0009, Feng Lu 0005, Danni Ai, Yitian Zhao, Yongtian Wang |
Neurocomputing | 2 |
| 2016 | Shape context and projection geometry constrained vasculature matching for 3D reconstruction of coronary artery
Ruoxiu Xiao, Jian Yang 0009, Jingfan Fan, Danni Ai, Guangzhi Wang, Yongtian Wang |
Neurocomputing | 2 |
| 2016 | Local statistics and non-local mean filter for speckle noise reduction in medical ultrasound image
Jian Yang 0009, Jingfan Fan, Danni Ai, Xuehu Wang, Yongchang Zheng, Songyuan Tang, Yongtian Wang |
Neurocomputing | 1 |
| 2016 | Region-based saliency estimation for 3D shape analysis and understanding
Yitian Zhao, Yonghuai Liu, Baogang Wei, Jian Yang 0009, Yifan Zhao 0001, Yongtian Wang |
Neurocomputing | 5 |
| 2016 | Convex hull indexed Gaussian mixture model (CH-GMM) for 3D point set registration
Jingfan Fan, Jian Yang 0009, Danni Ai, Likun Xia, Yitian Zhao, Yongtian Wang |
Pattern Recognit. | 2 |
| 2016 | Curve-Like Structure Extraction Using Minimal Path Propagation With BacktrackingabstractMinimal path techniques can efficiently extract geometrically curve-like structures by finding the path with minimal accumulated cost between two given endpoints. Though having found wide practical applications (e.g., line identification, crack detection, and vascular centerline extraction), minimal path techniques suffer from some notable problems. The first one is that they require setting two endpoints for each line to be extracted (endpoint problem). The second one is that the connection might fail when the geodesic distance between the two points is much shorter than the desirable minimal path (shortcut problem). In addition, when connecting two distant points, the minimal path connection might become inefficient as the accumulated cost increases over the propagation and results in leakage into some non-feature regions near the starting point (accumulation problem). To address these problems, this paper proposes an approach termed minimal path propagation with backtracking. We found that the information in the process of backtracking from reached points can be well utilized to overcome the above problems and improve the extraction performance. The whole algorithm is robust to parameter setting and allows a coarse setting of the starting point. Extensive experiments with both simulated and realistic data are performed to validate the performance of the proposed method. Yang Chen 0008, Yudong Zhang 0001, Jian Yang 0009, Guanyu Yang 0001, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux, Qianjing Feng |
IEEE Trans. Image Process. | 3 |
| 2014 | Artifact Suppressed Dictionary Learning for Low-Dose CT Image ProcessingabstractLow-dose computed tomography (LDCT) images are often severely degraded by amplified mottle noise and streak artifacts. These artifacts are often hard to suppress without introducing tissue blurring effects. In this paper, we propose to process LDCT images using a novel image-domain algorithm called "artifact suppressed dictionary learning (ASDL)." In this ASDL method, orientation and scale information on artifacts is exploited to train artifact atoms, which are then combined with tissue feature atoms to build three discriminative dictionaries. The streak artifacts are cancelled via a discriminative sparse representation operation based on these dictionaries. Then, a general dictionary learning processing is applied to further reduce the noise and residual artifacts. Qualitative and quantitative evaluations on a large set of abdominal and mediastinum CT images are carried out and the results show that the proposed method can be efficiently applied in most current CT systems. Yang Chen 0008, Luyao Shi, Qianjing Feng, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux, Wufan Chen |
IEEE Trans. Medical Imaging | 4 |
| 2013 | Rigid registration of 3-D medical image using convex hull matchingabstractIn this paper, a robust approach called convex hull matching (CHM) technique is proposed for registration of medical images that differ from each other with Euclidean transformation. Firstly, point sets on the surface of the medical image are extracted, and then the 3-D convex hull is constructed from the point sets and triangle patches on the surface of convex hulls are specified by predefining their normal vectors. Secondly, each edge of the referenced triangle is compared with all the edges of the triangle in other point set to find the congruent pair set and also to obtain the scaling factor. Thereafter, the transformation parameters of each triangle pairs including rotation and translation are optimized by minimizing the Euclidian distance between the corresponding vertex pairs. Hence, rigid transformation of the two point sets is obtained by iteratively enumerating and evaluating similarity measures of the triangle patches chosen. Global optimization is achieved through RANSAC optimization by removing the correspondence pairs that may lead to large matching errors of the whole point sets. The experiments evaluate the performance of the proposed algorithm on simulated data with the presence of outliers and noise. The results show the efficiency of CHM by quantitative analysis and comparative study with existing approaches like EM-ICP, LM-ICP and 4PCS. Finally, the real clinical data experiments confirm the proposed algorithm is a strong performer in medical image registration. Jingfan Fan, Jian Yang 0009, Mahima Goyal, Yongtian Wang |
BIBM | 2 |
| 2011 | Segmentation Based on Routing Image AlgorithmsabstractThis paper presents a novel image segmentation method in which energy function is based on global region information while not only on edge information. Image segmentation can be viewed as a routing problem. In order to obtain the optimal segmentation, the Shortest Path Faster Algorithm (SPFA) is used to optimize the discrete grid energy function. As the commonly used Live-Wire algorithm is easy to obtain mistake segmentation when the strong edges and the weak edges are close to each other, the interactive segmentation method is proposed for the precise boundaries estimation. The developed method has been tested on both clinical medical images and natural scene images. It can be seen that the developed method is very fast and effective, and can obtain good segmentation results. Hongzhe Yang, Jian Yang 0009, Yongtian Wang, Yue Liu 0005 |
ICIG | 2 |
| 2009 | Novel Approach for 3-D Reconstruction of Coronary Arteries From Two Uncalibrated Angiographic ImagesabstractThree-dimensional reconstruction of vessels from digital X-ray angiographic images is a powerful technique that compensates for limitations in angiography. It can provide physicians with the ability to accurately inspect the complex arterial network and to quantitatively assess disease induced vascular alterations in three dimensions. In this paper, both the projection principle of single view angiography and mathematical modeling of two view angiographies are studied in detail. The movement of the table, which commonly occurs during clinical practice, complicates the reconstruction process. On the basis of the pinhole camera model and existing optimization methods, an algorithm is developed for 3-D reconstruction of coronary arteries from two uncalibrated monoplane angiographic images. A simple and effective perspective projection model is proposed for the 3-D reconstruction of coronary arteries. A nonlinear optimization method is employed for refinement of the 3-D structure of the vessel skeletons, which takes the influence of table movement into consideration. An accurate model is suggested for the calculation of contour points of the vascular surface, which fully utilizes the information in the two projections. In our experiments with phantom and patient angiograms, the vessel centerlines are reconstructed in 3-D space with a mean positional accuracy of 0.665 mm and with a mean back projection error of 0.259 mm. This shows that the algorithm put forward in this paper is very effective and robust. Jian Yang 0009, Yongtian Wang, Yue Liu 0005, Songyuan Tang, Wufan Chen |
IEEE Trans. Image Process. | 1 |