VLDB 2026 Research / reviewers in the wild / expert
Yang Chen 0008
dblp:48/4792-8
· DBLP profile ↗
136ranked-venue papers
4as first author
101since 2021 · last 2026
0000-0002-5660-6349ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 82 · 2 first-author · 68 since 2021Artificial intelligence and machine learning · 34 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 20 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multisource space-frequency joint learning: A novel paradigm for ultrasound image quality assessment
Tuo Liu, Xuejuan Wang, Yang Chen 0008, Rongjun Ge, Faqin Lv, Guangquan Zhou |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | GraphMorph: Equilibrium adjustment regularized dual-stream GCN for 4D-CT lung imaging with sliding motion
Fei Lyu 0004, Yudong Zhang 0001, Zhan Wu, Jianmin Dong 0003, Tianling Lyu, Wei Zhao 0029, Jean-Louis Coatrieux, Yang Chen 0008 |
Neurocomputing | 11 |
| 2026 | Generative data-engine foundation model for universal few-shot 2D vascular image segmentationabstractThe segmentation of 2D vascular structures via deep learning holds significant clinical value but is hindered by the scarcity of annotated data, severely limiting its widespread application. Developing a universal few-shot vascular segmentation model is highly desirable, yet remains challenging due to the need for extensive training and the inherent complexities of vascular imaging. In this work, we propose UniVG (Generative Data-engine Foundation Model for Universal Few-shot 2D Vascular Image Segmentation), a novel approach that learns the compositionality of vascular images and constructing a generative foundation model for robust vascular segmentation. UniVG enables the synthesis and learning of diverse and realistic vascular images through two key innovations: 1) Compositional learning for flexible and diverse vascular synthesis: It decomposes and recombines vascular structures with varying morphological features and diverse foreground-background configurations to generate richly diverse synthetic image-label pairs. 2) Few-shot generative adaptation for transferable segmentation: It fine-tunes pre-trained models with minimal annotated data to bridge the gap between synthetic and real vascular domains, synthesizing authentic and diverse vessel images for downstream few-shot vascular segmentation learning. To support our approach, we develop UniVG-58K, a large dataset comprising 58,689 vascular images across five imaging modalities, facilitating robust large-scale generative pre-training. Extensive experiments on 11 vessel segmentation tasks cross 5 modalties (only with 5 labeled images on each task) demonstrate that UniVG achieves performance comparable to fully supervised models, significantly reducing data collection and annotation costs. All code and datasets will be made publicly available at https://github.com/XinAloha/UniVG. Rongjun Ge, Yuxing Liu, Chengliang Liu 0003, Pinzheng Zhang, Jiong Zhang 0004, Jian Yang 0009, Jean-Louis Dillenseger, Yuting He 0001, Yang Chen 0008 |
Medical Image Anal. | 11 |
| 2026 | VCC-DSA: A novel vascular consistency constrained DSA imaging model for motion artifact suppression
Rongjun Ge, Weilong Mao, Guanyu Yang 0001, Yang Chen 0008, Shuo Li 0001 |
Medical Image Anal. | 11 |
| 2026 | Causality-inspired representation learning with spatiotemporal memory for polyp detection in endoscopic videos
Changjin Sun, Xiaopu He, Cheng Xue 0003, Guangquan Zhou, Yang Chen 0008 |
Medical Image Anal. | 7 |
| 2026 | FDA-Recon: Feature and data alignment reconstruction for sparse-view CBCT
Yikun Zhang 0001, Dianlin Hu, Tianling Lyu, Yan Xi, Jian Yang 0009, Yang Chen 0008 |
Medical Image Anal. | 9 |
| 2026 | CHAP: Channel-spatial hierarchical adversarial perturbation for semi-supervised medical image segmentation
Siping Zhou, Zhi-Fang Gong, Kai-Ni Wang, Yang Chen 0008, Guangquan Zhou |
Medical Image Anal. | 5 |
| 2026 | Contourlet-informed prior controllable adaptation of ultrasound foundation model for abdominal trauma assessment
Tuo Liu, Xiuzhu Ma, Xuejuan Wang, Rongjun Ge, Faqin Lv, Yang Chen 0008, Guangquan Zhou |
Pattern Recognit. | 8 |
| 2026 | RecHCA: Hierarchical Context Awareness for one-step sensorless freehand 3D ultrasound reconstruction
Qing-Han Yang, Jing-Yang Zhang, Xing-Yang Liu, Yan Xi, Yang Chen 0008, Guangquan Zhou |
Pattern Recognit. | 7 |
| 2026 | Adaptation Follow Human Attention: Gaze-Assisted Medical Segment Anything ModelabstractSegment Anything Model (SAM) has demonstrated state-of-the-art performance in most segmentation tasks. However, due to insufficient training in the medical domain, SAM’s ability to generalize to medical images is limited. Although preliminary efforts have fine-tuned SAM for the medical domain, the fine-tuned model still struggles with variability in medical tasks. Some recent studies have explored weakly supervised learning to mitigate SAM’s performance degradation in the medical domain. However, the effectiveness of weakly supervised learning is heavily dependent on the quality of weakly supervised information, with performance significantly dropping as the quality declines. Doctors’ attention is closely related to the target area during diagnosis. Integrating gaze information into SAM’s adaptation process for medical image segmentation enhances efficiency and significantly improves performance in medical tasks. In this paper, we first propose a Gaze-assisted medical segment Anything Model (GAM), which utilizes gaze information to enable the adaptation of SAM in medical images following doctor’s attention. It has two innovations: 1) Feature-level adaptation: Gaze Alignment (GA) learning makes the feature-level adaptation follow the doctor’s attention which mines the human guidance from gaze heatmaps and guides model to extract general features for downstream tasks. 2) Output-level adaptation: Gaze-Balance (GB) learning makes the output-level adaptation follow the doctor’s attention which utilizes gaze heatmaps to enhance the human-focused area and solve the problem of over/under segmentation from the output-level. Our promising results on 7 tasks with 12 targets have demonstrated the powerful adaptation ability of our GAM in the medical domain. Our GAM demonstrates significant potential for low-cost clinical assistance in medical diagnosis, enabling SAM to adapt to the medical image domain without disrupting clinical workflows. We have released the full source code on https://github.com/Ruiz1026/GAM. Rongjun Ge, Ruiyi Li, Chong Wang 0011, Jean-Louis Coatrieux, Daoqiang Zhang, Yang Chen 0008, Shuo Li 0001, Yuting He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Sparse-View CT Reconstruction via Implicit Neural Representation Learning Powered by Dual-Domain Vision Foundation ModelsabstractSparse-view computed tomography (SVCT) offers the advantages of accelerated scanning and reduced X-ray radiation dose in different clinical applications. However, it faces a challenge due to incomplete data acquisition, resulting in streak artifacts in the analytically reconstructed CT images. Utilizing self-supervised learning, implicit neural representation (INR) recently has shown great promise in addressing inverse problems such as SVCT reconstruction. Nonetheless, given that the input of original INR only contains coordinate information, it is limited to represent one SVCT instance at a time, and its performance significantly declines when performing cross-instance reconstruction. In this study, we propose a novel self-supervised framework named VFMINR, which leverages generalizable representations extracted from the visual foundation models (VFMs) to tackle the cross-instance reconstruction issue of INR. Specifically, VFMINR first utilizes VFMs to effectively capture the spatial and frequency domain representations of sinograms, and then a fusion module is applied to fuse two domain features into complementary representations. This combination maximizes the utilization of local detail information from the spatial domain and the global structural information from the frequency domain. Subsequently, an adaptive cell decoding strategy is designed to map representations into variable resolution hybrid feature grids, which are integrated into the learning of the INR to enhance its generalizability for different SVCT instances. The VFMs and VFMINR are trained by using only SV sinogram data, and extensive results confirm that the proposed method can effectively handle the generalization problem of INR, while achieving superior performance in image fidelity and artifact suppression. The code is available at: https://github.com/nightastars/VFMINR-main. Yang Chen 0008, Yangchuan Liu, Zhongyi Wu, Hengyong Yu, Jian Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Airs-Net: Adversarial-Improved Reversible Steganography Network for CT Images in the Internet of Medical Things and TelemedicineabstractMedical imaging has developed from an auxiliary means of clinical examination into a significant method and intuitive basis for clinical diagnosis of diseases, providing all-around and full-cycle health protection for the people. The Internet of Medical Things (IoMT) allows medical equipment, intelligent terminals, medical infrastructure, and other elements of medical production to be interconnected, eliminating information silos and data fragmentation. Medical images disseminated in IoMT contain a wide diversity of sensitive patient information, which means protecting the patient's personal information is vital. In this work, an Adversarial-improved reversible steganography network (Airs-Net) for computed tomography (CT) images in the IoMT is presented. Specifically, the Airs-Net adopting the prediction-embedding strategy mainly consists of an image restoration network, an embedded pixel location network, and a discriminator. The image restoration network is effective in restoring the pixel prediction error of the restoration set in integer and non-integer scaled images of arbitrary size when information is concealed. The embedded information location network can automatically select pixel locations for information embedding based on the interpolated image features of the degraded image. The restored image, embedding location map, and embedding information are fed into the embedder for information embedding, and the subsequent secret-carrying image is continuously optimized for the quality of the information-embedded image by the discriminator. Quantitative results show that Airs-Net outperforms state-of-the-art methods in both PSNR and SSIM. Further, the qualitative and quantitative results and analyses under specific clinical application scenarios and in coping with multiple types of medical image information hiding demonstrate the excellent generalization performance and practical application capability of the Airs-Net. Kai Chen 0039, Mu Nie, Jean-Louis Coatrieux, Yang Chen 0008, Shipeng Xie |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | WOADNet: A Wavelet-Inspired Orientational Adaptive Dictionary Network for CT Metal Artifact ReductionabstractIn computed tomography (CT), metal artifacts pose a persistent challenge to achieving high-quality imaging. Despite advancements in metal artifact reduction (MAR) techniques, many existing approaches have not fully leveraged the intrinsic a priori knowledge related to metal artifacts, improved model interpretability, or addressed the complex texture of CT images effectively. To address these limitations, we propose a novel and interpretable framework, the wavelet-inspired oriented adaptive dictionary network (WOADNet). WOADNet builds on sparse coding with orientational information in the wavelet domain. By exploring the discriminative features of artifacts and anatomical tissues, we adopt a high-precision filter parameterization strategy that incorporates multiangle rotations. Furthermore, we integrate a reweighted sparse constraint framework into the convolutional dictionary learning process and employ a cross-space, multiscale attention mechanism to construct an adaptive convolutional dictionary unit for the artifact feature encoder. This innovative design allows for flexible adjustment of weights and convolutional representations, resulting in significant image quality improvements. The experimental results using synthetic and clinical datasets demonstrate that WOADNet outperforms both traditional and state-of-the-art MAR methods in terms of suppressing artifacts. Jin Liu 0019, Diandian Wang, Kun Wang 0021, Chenlong Miao, Yikun Zhang 0001, Dianlin Hu, Zhan Wu, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 9 |
| 2026 | Edge-Aware Diffusion Segmentation Model With Hessian Priors for Automated Diaphragm Thickness Measurement in Ultrasound ImagingabstractThe thickness of the diaphragm serves as a crucial biometric indicator, particularly in assessing rehabilitation and respiratory dysfunction. However, measuring diaphragm thickness from ultrasound images mainly depends on manual delineation of the fascia, which is subjective, time-consuming, and sensitive to the inherent speckle noise. In this study, we introduce an edge-aware diffusion segmentation model (ESADiff), which incorporates prior structural knowledge of the fascia to improve the accuracy and reliability of diaphragm thickness measurements in ultrasound imaging. We first apply a diffusion model, guided by annotations, to learn the image features while preserving edge details through an iterative denoising process. Specifically, we design an anisotropic edge-sensitive annotation refinement module that corrects inaccurate labels by integrating Hessian geometric priors with a backtracking shortest-path connection algorithm, further enhancing model accuracy. Moreover, a curvature-aware deformable convolution and edge-prior ranking loss function are proposed to leverage the shape prior knowledge of the fascia, allowing the model to selectively focus on relevant linear structures while mitigating the influence of noise on feature extraction. We evaluated the proposed model on an in-house diaphragm ultrasound dataset, a public calf muscle dataset, and an internal tongue muscle dataset to demonstrate robust generalization. Extensive experimental results demonstrate that our method achieves finer fascia segmentation and significantly improves the accuracy of thickness measurements compared to other state-of-the-art techniques, highlighting its potential for clinical applications. Chenlong Miao, Yikang He, Baike Shi, Zhongkai Bian, Wenxue Yu, Yang Chen 0008, Guangquan Zhou |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | ESIP: Explicit Surgical Instrument Prompting for Surgical Workflow RecognitionabstractSurgical workflow recognition (SWR) stands as a pivotal component in computer-assisted surgery and is dedicated to identifying phases from surgical videos. Many deep learning-based methods have been proposed for this task and achieved acceptable SWR results. However, these methods usually implicitly extract and aggregate spatio-temporal features, so that it is challenging for these methods to adequately use some spatial information that is strongly relevant to surgical phase in SWR task, such as the information from the surgical instruments. To address this issue, an Explicit Surgical Instrument Prompting (ESIP) approach is proposed for SWR task. ESIP leverages surgical instrument segmentation to generate instrument-specific visual prompts, which explicitly guide the extraction of crucial intra-frame spatial features through a frozen pre-trained backbone, then enable effective inter-frame spatio-temporal feature extraction and aggregation. Unlike multi-task approaches that jointly perform SWR with auxiliary tasks within a shared network framework, ESIP is a single-task SWR approach dedicated to optimize framework itself for more adequate feature extraction. Furthermore, to accomplish the segmentation prompting efficiently, this paper presents SAM-based segmentation with prompt tuning strategy to explicitly integrate segmentation features into spatial features. Experimental results on Cholec80, M2CAI and AutoLaparo datasets demonstrate that our ESIP method achieves the best performance in comparison with 16 SOTA methods, with a Precision of 91.8%, 89.5% and 89.6%, Recall of 92.2%, 89.5% and 76.9%, Jaccard of 83.3%, 77.0% and 67.3%, respectively. Mengxing Liu, Guangquan Zhou, Fei Lyu 0004, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | OrthoDetNet: An Enhanced YOLO-Based Framework for Detection of Orthopedic Surgical InstrumentsabstractAccurate detection of surgical instruments is critical for both routine surgical procedures and surgical robotics research. To the best of our knowledge, there is a notable lack of datasets and dedicated detection studies specifically addressing orthopedic surgical instruments. Detecting orthopedic surgical instruments presents particular challenges including significant size variations, highly similar shapes, and frequent, severe occlusions due to instrument intersections. To address these issues, we propose an orthopedic surgical instrument detection method (OrthoDetNet) incorporating three specialized modules. The FilterUnit mitigates occlusion effects via an adaptive feature filtering mechanism, that dynamically adjusts its filtering strategy based on context, prioritizing features from key regions while suppressing distracting interference features. The DEUnit enhances fine-grained feature discrimination in local regions to distinguish instruments with high shape similarity, and the BDFusion module improves multi-scale detection performance through bi-directional feature fusion between deep and shallow-level feature maps. A dataset for orthopedic surgical instrument detection is created, which is based on the proximal femoral nail antirotation (PFNA) instrument package manufactured by Shenzhen Mindray Bio-Medical Electronics Co., Ltd. Images were captured in a controlled, simulated experimental environment, ensuring no patient privacy or ethical concerns. We obtained explicit authorization from the manufacturer for instrument use. Experimental results on this dataset demonstrate the effectiveness of the OrthoDetNet and its constituent modules. Guangquan Zhou, Mengxing Liu, Chu Guo, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | TD-SAM: Temporal and Distance-Guided Adaptations of SAM for Accurate Surgical Instrument SegmentationabstractAccurate automatic surgical instrument segmentation plays a crucial role in robot-assisted surgery, but analyzing surgical videos remains challenging due to factors such as rapid instrument movements, high inter-category similarity, and frequent object occlusions. Current surgical instrument segmentation models struggle to capture both inter-frame variations and intra-frame details in complex surgical scenarios. The Segment Anything Model (SAM) has shown significant potential in various segmentation tasks. However, it has not fully addressed the unique challenges posed by surgical videos. To tackle these issues, we propose a Temporal and Distance-Guided SAM model (TD-SAM) for accurate surgical instrument segmentation. Specifically, we introduce a dynamic cross-frame attention module that effectively captures temporal information across frames, allowing the model to track the dynamic changes of surgical instruments and their environment, thus improving segmentation accuracy. In addition, we present a distance-guided instance refinement module, which enhances the model's ability to distinguish between similar categories, mitigating the class ambiguity caused by inter-category similarity. Extensive experiments on the EndoVis18 and EndoVis17 datasets show that the proposed TD-SAM model outperforms existing models, achieving state-of-the-art performance without using any prompts. Cheng Xue 0003, Danqiong Wang, Cheng Chen 0013, Guanyu Yang 0001, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | LADDA: Latent Diffusion-Based Domain-Adaptive Feature Disentangling for Unsupervised Multi-Modal Medical Image RegistrationabstractDeformable image registration (DIR) is critical for accurate clinical diagnosis and effective treatment planning. However, patient movement, significant intensity differences, and large breathing deformations hinder accurate anatomical alignment in multi-modal image registration. These factors exacerbate the entanglement of anatomical and modality-specific style information, thereby severely limiting the performance of multi-modal registration. To address this, we propose a novel LAtent Diffusion-based Domain-Adaptive feature disentangling (LADDA) framework for unsupervised multi-modal medical image registration, which explicitly addresses the representation disentanglement. First, LADDA extracts reliable anatomical priors from the Latent Diffusion Model (LDM), facilitating downstream content-style disentangled learning. A Domain-Adaptive Feature Disentangling (DAFD) module is proposed to promote anatomical structure alignment further. This module disentangles image features into content and style information, boosting the network to focus on cross-modal content information. Next, a Neighborhood-Preserving Hashing (NPH) is constructed to further perceive and integrate hierarchical content information through local neighbourhood encoding, thereby maintaining cross-modal structural consistency. Furthermore, a Unilateral-Query-Frozen Attention (UQFA) module is proposed to enhance the coupling between upstream prior and downstream content information. The feature interaction within intra-domain consistent structures improves the fine recovery of detailed textures. The proposed framework is extensively evaluated on large-scale multi-center datasets, demonstrating superior performance across diverse clinical scenarios and strong generalization on out-of-distribution (OOD) data. Jianmin Dong 0003, Wei Zhao 0029, Fei Lyu 0004, Cheng Xue 0003, Yudong Zhang 0001, Zhan Wu, Tianling Lyu, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 12 |
| 2026 | Conditional Virtual Imaging for Few-Shot Vascular Image SegmentationabstractIn the field of medical image processing, vascular image segmentation plays a crucial role in clinical diagnosis, treatment planning, prognosis, and medical decision-making. Accurate and automated segmentation of vascular images can assist clinicians in understanding the vascular network structure, leading to more informed medical decisions. However, manual annotation of vascular images is time-consuming and challenging due to the fine and low-contrast vascular branches, especially in the medical imaging domain where annotation requires specialized knowledge and clinical expertise. Data-driven deep learning models struggle to achieve good performance when only a small number of annotated vascular images are available. To address this issue, this paper proposes a novel Conditional Virtual Imaging (CVI) framework for few-shot vascular image segmentation learning. The framework combines limited annotated data with extensive unlabeled data to generate high-quality images, effectively improving the accuracy and robustness of segmentation learning. Our approach primarily includes two innovations: First, aligned image-mask pair generation, which leverages the powerful image generation capabilities of large pre-trained models to produce high-quality vascular images with complex structures using only a few training images; Second, the Dual-Consistency Learning (DCL) strategy, which simultaneously trains the generator and segmentation model, allowing them to learn from each other and maximize the utilization of limited data. Experimental results demonstrate that our CVI framework can generate high-quality medical images and effectively enhance the performance of segmentation models in few-shot scenarios. Our code will be made publicly available online. Yanglong He, Rongjun Ge, Mengqing Su, Jean-Louis Coatrieux, Huazhong Shu, Yang Chen 0008, Yuting He 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | Segmentation-Guided Accelerating Diffusion Model for Cardiac CT Motion Artifact Reduction via Limited-Angle Imaging
Dianlin Hu, Zhan Wu, Guotao Quan, Shangwen Yang, Yikun Zhang 0001, Huazhong Shu, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | UPMCL-Net: Unsupervised Projection-Domain Multiview Constraint Learning for CBCT Metal Artifact ReductionabstractCone-beam Computed Tomography (CBCT) provides real-time three-dimensional (3D) imaging support for intraoperative navigation. However, high-attenuation metal implants introduce severe metal artifacts in reconstructed CBCT images. These artifacts compromise image quality and therefore may affect diagnostic accuracy. Current CBCT metal artifact reduction (MAR) algorithms overlook the complementary information available across CBCT views, leading to inaccurate projection-domain interpolation and secondary artifacts in the reconstructed images. To tackle these challenges, we propose a novel Unsupervised Projection-domain Multiview Constraint Learning Network (UPMCL-Net), which directly learns from metal-affected data for CBCT MAR without ground truths. In addition, a transformer-based MultiView Consistency Module (MVCM) is constructed to interpolate the projection-domain metal region for cross-view consistency. Finally, a Hybrid Feature Attention Module (HFAM) is designed to adaptively fuse interview and intraview features. Comprehensive experiments conducted on real clinical datasets confirm the performance of UPMCL-Net, showcasing its potential as an efficient, accurate, and reliable approach for CBCT MAR in clinical intraoperative interventions. Zhan Wu, Yang Yang 0216, Yongjie Guo, Dayang Wang, Tianling Lyu, Yan Xi, Yang Chen 0008, Hengyong Yu |
IEEE Trans. Medical Imaging | 7 |
| 2026 | UPGRADE-Net: Unsupervised Sinogram-Domain Data-Consistent Network for Metal Artifact ReductionabstractComputed tomography (CT) scanners are widely used to obtain detailed internal images in clinical diagnosis. Highly attenuated metallic implants resulting from strong and energy-dependent attenuation cause metal artifacts in CT scanning. However, current supervised deep network-based metal artifact reduction (MAR) methods hardly generalize in clinical diagnosis and treatment because of difficult acquisition for the paired artifact-affected and artifact-free data. In addition, these deep model-based methods cannot ensure the sinogram-domain data consistency for the exact metal trace inpainting. To address the above problems, we propose an UnsuPervised sinoGRam-domAin Data-consistEnt network for MAR, i.e., UPGRADE-Net. First, UPGRADE-Net fully leverages the prior knowledge to guide the generative conditional diffusion model for fine-grained metal trace inpainting. Second, without the artifact-free ground truth, a deep unsupervised MAR framework in the reverse process is constructed to contextually learn the known background data distribution for the unknown metal trace restoration in sinogram-domain. Third, to further maintain the sinogram-domain data consistency, two physics-based consistency constraint loss functions, including conjugate-ray and accumulation-ray consistency loss, are designed for the conjugate point constraint and the accumulation constraint. The proposed UPGRADE-Net is trained and evaluated on a publicly available dataset and a clinical dataset. Extensive experimental results validate that the proposed method outperforms the state-of-the-art competing methods for MAR. Zhan Wu, Yikun Zhang 0001, Yongjie Guo, Huazhong Shu, Yan Xi, Yi Zhang 0018, Gouenou Coatrieux, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 10 |
| 2026 | A Physics-ASIC Architecture-Driven Deep Learning Photon-Counting Detector Model Under Limited DataabstractPhoton-counting computed tomography (PCCT) based on photon-counting detectors (PCDs) represents a cutting-edge CT technology, offering higher spatial resolution, reduced radiation dose, and advanced material decomposition capabilities. Accurately modeling complex and nonlinear PCDs under limited calibration data becomes one of the challenges hindering the widespread accessibility of PCCT. This paper introduces a physics-ASIC architecture-driven deep learning detector model for PCDs. This model adeptly captures the comprehensive response of the PCD, encompassing both sensor and ASIC responses. We present experimental results demonstrating the model's exceptional accuracy and robustness with limited calibration data. Key advancements include reduced calibration errors, reasonable physics-ASIC parameters estimation, and high-quality and high-accuracy material decomposition images. Qianyu Wu, Wenhui Qin, Mengqing Su, Jinglu Ma, Yikun Zhang 0001, Wenying Wang, Guotao Quan, Yanfeng Du, Yang Chen 0008, Xiaochun Lai |
IEEE Trans. Medical Imaging | 12 |
| 2026 | Delta-Net: Deep Dual-Domain Alternating Optimization Network for High Pitch Helical CT ReconstructionabstractHigh pitch helical Computed Tomography (CT) scanning significantly reduces radiation dose while improving temporal resolution, offering substantial clinical benefits. However, the incomplete scanning data commonly leads to artifacts in the reconstructed images, degrading image quality and potentially affecting clinical diagnosis. Existing high pitch reconstruction methods primarily operate within the image domain or combine image-domain networks with traditional iterative algorithms, yet their performance remains limited. To address such limitations, we propose Delta-Net, a deep dual-domain alternating iterative optimization network for high pitch helical CT reconstruction. We introduce a novel optimization objective and develop an alternating iterative optimization framework, where each sub iteration consists of projection domain correction and image domain refinement. To enhance generalization and robustness, deep neural networks are employed to learn domain-specific priors, which are incorporated as regularization terms, with all hyper-parameters automatically optimized during training. Specifically, the image domain residual refinement network (IRN) and projection domain consistency enhanced network (PCN) regularize the intermediate results across both domains. Additionally, to improve the capability of artifact suppression and structure restoration, a structure-aware joint loss is tailored for the optimization of Delta-Net. Quantitative and qualitative evaluations on clinical datasets demonstrate that Delta-Net outperforms other competitive methods in artifact suppression, fine structure recovery, and generalization. Xinyun Zhong, Guojun Zhu, Yikun Zhang 0001, Qianjin Feng 0001, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | DET-CPD: Dynamic Edge-Aware Transformer with Cross-Image Patch Dependency for Lesion Segmentation in Ultrasound ImagesabstractUltrasound image segmentation is critical for tumor screening but is hindered by noise, artifacts, and high variability in lesion appearance. Challenges like blurred boundaries and morphological similarities further complicate accurate delineation. To address this, we propose the Dynamic Edge-aware Transformer with Cross-image Patch Dependency (DET-CPD). Our model integrates two key modules: a Dynamic Difference Convolution Module (DDCM) to enhance edge representation for varied lesions, and a Cross-Scale Semantic Enhancement Module (CSEM) that leverages cross-scale channel information to distinguish tumors from surrounding tissue. Crucially, we introduce a novel Cross-image Patch Dependency Loss (CPDLoss) that captures semantic dependencies across different images in a batch, improving robustness. Extensive experiments on four public datasets (BUSI, DatasetB, DDTI, and TN3K) demonstrate that DET-CPD achieves state-of-the-art segmentation performance. Chufeng Jin, Tao Wang 0107, Baike Shi, Guangquan Zhou, Rongjun Ge, Qianjin Feng 0001, Yang Chen 0008, Jean-Louis Coatrieux |
BIBM | 9 |
| 2025 | Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language ModelabstractReducing radiation doses benefits patients, but the resultant low-dose computed tomography (LDCT) images often suffer from clinically unacceptable noise and artifacts. While deep learning (DL) has shown promise in LDCT reconstruction, it requires large-scale data collection from multiple clients, raising privacy concerns. Federated learning (FL) has been introduced to mitigate these privacy concerns; however, current methods are typically tailored to specific scanning protocols, which limits their generalizability and makes them less effective for unseen protocols. To address these issues, we propose SCANPhysFed, a novel SCanning- and ANatomy-level personalized Physics-Driven Federated learning paradigm for LDCT reconstruction. Since the noise distribution in LDCT data is closely tied to scanning protocols and anatomical structures, we propose a dual-level physics-informed way to address these challenges. Specifically, we incorporate physical and anatomical prompts into our physics-informed hypernetworks to capture scanning- and anatomy-specific information, enabling dual-level physics-driven personalization of imaging features. These prompts are derived from the scanning protocol and the radiology report generated by a medical large language model (MLLM). Subsequently, client-specific decoders project these dual-level personalized imaging features back into the image domain. Besides, to tackle the challenge of unseen data, we introduce a novel protocol vector-quantization strategy (PVQS), which ensures consistent performance across new clients by quantifying unseen scanning codes to the closest match in the scanning codebook. Extensive experimental results demonstrate the superior performance of SCAN-PhysFed on public datasets1. Ziyuan Yang 0001, Zhiwen Wang 0002, Hongming Shan, Yang Chen 0008, Yi Zhang 0018 |
CVPR | 5 |
| 2025 | IBS-Net: Advancing Implicit Boundary-Aware Segmentation for Diaphragm Ultrasound AnalysisabstractAccurate automated measurement of diaphragmatic thickness in ultrasound imaging is a critical challenging task for respiratory function assessment, primarily due to difficulties in precise fascial identification. And ultrasound visualization of the diaphragm is characterized by unique challenges, including discontinuous and blurred boundary delineations caused by imaging artifacts, as well as interference and influence from adjacent muscular reverberations. These problems are further compounded by subjects’ pose variations during image acquisition. To address these challenges, we introduce IBS-Net, an innovative triple-branch interactive segmentation network that synergistically combines boundary regression with auxiliary task learning to optimize feature representation in segmentation task. Moreover, Our framework incorporates two innovative module: an Adaptive Fusion Module (AFM) that enables multi-scale hierarchical feature refinement for precise boundary characterization, and a Cross Interactive Module (CIM) that employs parallel-encoded feature extraction to simultaneously achieve accurate fascial localization while preserving structural topology. These complementary mechanisms effectively resolve spatial feature inconsistencies, facilitating robust multi-level feature integration. Comprehensive experimental results demonstrate that IBS-Net achieves statistically significant improvements of 8.9% in Dice similarity coefficient and 8.05% in Jaccard index compared to conventional methods. Moreover, to verify the effectiveness of the proposed method, we extended it to other publicly available BUSI dataset for experimentation. The results demonstrate that our method is competitive in terms of both accuracy and completeness in the identification of fuzzy boundaries in ultrasound images. Baike Shi, Yikang He, Chenlong Miao, Tao Wang 0107, Jianmin Dong 0003, Rongjun Ge, Guangquan Zhou, Yang Chen 0008 |
ECAI | 11 |
| 2025 | Dual-energy CT metal artifact reduction by combined material decomposition and projection domain threshold segmentationabstractDual-energy CT exploits the different attenuation characteristics of substances under different energy X-rays and collects high- and low-energy data from the same area to differentiate and quantify specific substances, which is now widely used in clinical diagnosis, disease monitoring, and other fields. If the object scanned during dual-energy CT imaging contains metallic material, the reconstructed image will suffer from metal artifacts. Metal artifacts in dual-energy CT result in surrounding tissue structures presenting erroneous CT values, blurred image details, and unclear border demarcation lines, which seriously affects the quality of the images as well as the accuracy of clinical diagnosis. To reduce metal artifacts in dual-energy CT, we combine material decomposition techniques with metal artifact reduction to explore metal artifact reduction methods applied to dual-energy CT. Kai Chen 0039, Tianling Lyu, Jean-Louis Coatrieux, Yang Chen 0008 |
ICASSP | 4 |
| 2025 | Gaze-Assisted Human-Centric Domain Adaptation for Cardiac Ultrasound Image SegmentationabstractDomain adaptation (DA) for cardiac ultrasound image segmentation is clinically significant and valuable. However, previous domain adaptation methods are prone to be affected by the incomplete pseudo label and low-quality target to source images. Human-centric domain adaptation has great advantages of human cognitive guidance to help model adapt to target domain and reduce reliance on labels. Doctor gaze trajectories contains a large amount of cross-domain human guidance. To leverage gaze information and human cognition for guiding domain adaptation, we propose gaze-assisted human-centric domain adaptation (GAHCDA), which reliably guides the domain adaptation of cardiac ultrasound images. GAHCDA includes following modules: (1) Gaze Augment Alignment (GAA): GAA enables the model to obtain human cognition general features to recognize segmentation target in different domain of cardiac ultrasound images like humans. (2) Gaze Balance Loss (GBL): GBL fused gaze heatmap with outputs which makes the segmentation result structurally closer to the target domain. The experimental results show that our proposed framework is able to segment cardiac ultrasound images more effectively in the target domain than GAN-based methods and other self-train based methods and shown great potential in clinical application. Ruiyi Li, Yuting He 0001, Rongjun Ge, Chong Wang 0011, Daoqiang Zhang, Yang Chen 0008, Shuo Li 0001 |
ICASSP | 6 |
| 2025 | A Causality-Inspired Model for Intima-Media Thickening Assessment in Ultrasound Videos
Yang Chen 0008, Jingyang Zhang, Guangquan Zhou |
MICCAI (8) | 4 |
| 2025 | Think as Cardiac Sonographers: Marrying SAM with Left Ventricular Indicators Measurements According to Clinical Guidelines
Tuo Liu, Qinghan Yang, Rongjun Ge, Yang Chen 0008, Guangquan Zhou |
MICCAI (10) | 5 |
| 2025 | Three-dimensional reconstruction and fracture segmentation based on X-ray and computed tomography paired datasetabstractIn some orthopedic surgeries, the use of three-dimensional (3D) computed tomography (CT) scanning technology is not feasible due to scene limitations, leaving doctors to rely on two-dimensional (2D) X-ray images for real-time diagnosis. However, X-ray images lack 3D information, making accurate diagnosis challenging. Developing an algorithm to convert 2D X-ray images into 3D CT images, while simultaneously combining high-quality 3D reconstruction with precise fracture segmentation, offers a promising solution to the problem. In this study, we propose a novel artificial intelligence (AI)-driven framework named 3D reconstruction and segment anything model (3DRecSAM). The reconstruction image enhancer (RIE) is designed to achieve high-precision 3D reconstruction and provide high-quality feature initialization for fracture segmentation. Meanwhile, the mamba segment anything model (MSAM), based on the segment anything model (SAM) architecture, is developed for accurate fracture segmentation. We introduce a Kolmogorov–Arnold network (KAN)-based attention fusion module (KAF), which facilitates the joint optimization of the RIE reconstruction network and the MSAM segmentation network. Furthermore, the selective scanning mamba with KAN (SKM) is incorporated to enhance feature extraction for both RIE and MSAM. Mamba efficiently captures long-range dependencies and sequential patterns, while KAN’s learnable activation functions facilitate adaptive feature fusion and non-linear representation. To train and evaluate 3DRecSAM, we introduce the real X-ray and CT paired dataset (XCPData), which is publicly available on GitHub: https://github.com/YuanGao1201/XCPData . Yuan Gao 0033, Da Chen 0002, Mingle Zhou, Gang Li 0005, Yunbo Gu, Jean-Louis Coatrieux, Yang Chen 0008 |
Eng. Appl. Artif. Intell. | 9 |
| 2025 | Dynamic spectrum-driven hierarchical learning network for polyp segmentation
Kai-Ni Wang, Jie Hua 0004, Yang Chen 0008, Guangquan Zhou, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2025 | Topology-oriented foreground focusing network for semi-supervised coronary artery segmentation
Xiangxin Wang, Zhan Wu, Yujia Zhou 0001, Huazhong Shu, Jean-Louis Coatrieux, Yang Chen 0008 |
Medical Image Anal. | 7 |
| 2025 | TSdetector: Temporal-Spatial self-correction collaborative learning for colonoscopy video detection
Kai-Ni Wang, Guangquan Zhou, Ling Yang 0006, Yang Chen 0008, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2025 | Multi-modal clear cell renal cell carcinoma grading with the segment anything model
Yunbo Gu, Qianyu Wu, Junting Zou, Xiaoli Mai, Yang Chen 0008 |
Multim. Syst. | 7 |
| 2025 | Human gaze-based dual teacher guidance learning for semi-supervised medical image segmentation
Rongjun Ge, Chong Wang 0011, Chunqiang Lu, Cong Xia, Yehui Jiang, Fangyi Xu, Yinsu Zhu, Daoqiang Zhang, Chengyu Liu 0001, Yang Chen 0008, Shuo Li 0001, Yuting He 0001 |
Neural Networks | 11 |
| 2025 | Homeomorphism Prior for False Positive and Negative Problem in Medical Image Dense Contrastive Representation LearningabstractDense contrastive representation learning (DCRL) has greatly improved the learning efficiency for image dense prediction tasks, showing its great potential to reduce the large costs of medical image collection and dense annotation. However, the properties of medical images make unreliable correspondence discovery, bringing an open problem of large-scale false positive and negative (FP&N) pairs in DCRL. In this paper, we propose GEoMetric vIsual deNse sImilarity (GEMINI) learning which embeds the homeomorphism prior to DCRL and enables a reliable correspondence discovery for effective dense contrast. We proposes a deformable homeomorphism learning (DHL) which models the homeomorphism of medical images and learns to estimate a deformable mapping to predict the pixels' correspondence under the condition of topological preservation. It effectively reduces the searching space of pairing and drives an implicit and soft learning of negative pairs via gradient. We also proposes a geometric semantic similarity (GSS) which extracts semantic information in features to measure the alignment degree for the correspondence learning. It will promote the learning efficiency and performance of deformation, constructing positive pairs reliably. We implement two practical variants on two typical representation learning tasks in our experiments. Our promising results on seven datasets which outperform the existing methods show our great superiority. We will release our code at a companion website. Yuting He 0001, Boyu Wang 0004, Rongjun Ge, Yang Chen 0008, Guanyu Yang 0001, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | DAM: Degradation-Aware Model for Ultrasound Image Quality AssessmentabstractOne of the core challenges in ultrasound image quality assessment (IQA) is the entanglement of semantic content and quality-related information, such as blurring and shadows. Insufficient attention to the latter can easily lead to biased IQA results. Furthermore, fine-grained quality inconsistencies, i.e., subtle variations in ultrasound images that can impact quality interpretations, may further complicate the IQA tasks. To address these challenges, we propose a novel degradation-aware model (DAM) for the ultrasound IQA, which effectively perceives various and subtle variations of quality patterns, accurately assessing the quality of ultrasound images. The advanced degradation-derived augmentation (DDA) in DAM incorporates degradations that clinicians may focus on during IQA into the synthesis of appearance changes, promoting the disentanglement of quality-related representations from semantic contents. Subsequently, we present fine-grained degradation learning (FGDL), which encourages distinctions between image versions with diminishing quality inconsistencies, boosting the awareness of quality nuances from easy to hard for better ultrasound IQA performance. A universal boundary acquisition operator (UBAO) is also developed to suppress interferences from redundant information, achieving the standardization of ultrasound images from various devices. Extensive experimental results on an in-house ultrasound dataset demonstrate that DAM outperforms 14 baseline methods, achieving a PLCC of 0.760 and an SROCC of 0.766. The code can be available at this URL. Tuo Liu, Xiuzhu Ma, Xuejuan Wang, Yang Chen 0008, Guangquan Zhou, Faqin Lv |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Frequency-Phase Guided Attention Complex-Valued Network for Ultrasound Image SegmentationabstractUltrasound imaging has emerged as an effective tool for aiding diagnosis. The automatic segmentation of ultrasound images is crucial in identifying the lesion target and evaluating clinical indicators for accurate diagnosis and prognosis. However, the segmentation problems are challenging due to the inherent speckle noise interference and low contrast of ultrasound images. The complex-value-based neural network can directly deal with the phase components, offering a potential solution in a better-perceiving structure for ultrasound image segmentation. In this study, we develop a Frequency Phase-Guided Attention Network (FPGANet) for ultrasound image segmentation by exploring the properties of the complex-valued model under the guide of phase and frequency perspectives. First, our proposed method transforms images into a complex domain as the input to an advanced complex-value model consisting of pure complex-value convolutions and operations. Especially this model can then effectively scrutinize phase information to distinguish target areas from similar backgrounds better. Moreover, we introduce a complex hybrid attention module following complex convolution to selectively adjust the perception of phase components and the model's bias. Also, we designed a frequency-adaptive separation module to emphasize frequency features prioritized by the encoder and decoder using a combination of wavelet decomposition and frequency channel attention. We evaluate the proposed FPGANet on three publicly available ultrasound datasets of breast, cardiac and thyroid nodules and a private abdominal effusion ultrasound dataset. Comparative experiments were also conducted with state-of-the-art methods. The results demonstrate the superior performance of FPGANet, implying its potential for advancing ultrasound image segmentation. Wen-Bo Zhang, Yang Chen 0008, Guangquan Zhou |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | DPI-MoCo: Deep Prior Image Constrained Motion Compensation Reconstruction for 4D CBCTabstract4D cone-beam computed tomography (CBCT) plays a critical role in adaptive radiation therapy for lung cancer. However, extremely sparse sampling projection data will cause severe streak artifacts in 4D CBCT images. Existing deep learning (DL) methods heavily rely on large labeled training datasets which are difficult to obtain in practical scenarios. Restricted by this dilemma, DL models often struggle with simultaneously retaining dynamic motions, removing streak degradations, and recovering fine details. To address the above challenging problem, we introduce a Deep Prior Image Constrained Motion Compensation framework (DPI-MoCo) that decouples the 4D CBCT reconstruction into two sub-tasks including coarse image restoration and structural detail fine-tuning. In the first stage, the proposed DPI-MoCo combines the prior image guidance, generative adversarial network, and contrastive learning to globally suppress the artifacts while maintaining the respiratory movements. After that, to further enhance the local anatomical structures, the motion estimation and compensation technique is adopted. Notably, our framework is performed without the need for paired datasets, ensuring practicality in clinical cases. In the Monte Carlo simulation dataset, the DPI-MoCo achieves competitive quantitative performance compared to the state-of-the-art (SOTA) methods. Furthermore, we test DPI-MoCo in clinical lung cancer datasets, and experiments validate that DPI-MoCo not only restores small anatomical structures and lesions but also preserves motion information. Dianlin Hu, Xuanjia Fei, Yan Xi, Jin Liu 0019, Yikun Zhang 0001, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 10 |
| 2025 | Dual-Source CBCT for Large FoV Imaging Under Short-Scan TrajectoriesabstractCone-beam CT is extensively used in medical diagnosis and treatment. Despite its large longitudinal field of view (FoV), the horizontal FoV of CBCT systems is severely limited due to the detector width. Certain commercial CBCT systems increase the horizontal FoV by employing the offset detector method. However, this method necessitates 360° full circular scanning trajectory which increases the scanning time and is not compatible with specific CBCT system models. In this paper, we investigate the feasibility of large FoV imaging under short scan trajectories with an additional X-ray source. A dual-source CBCT geometry is proposed as well as two corresponding image reconstruction algorithms. The first one is based on cone-parallel rebinning and the subsequent employs a modified Parker weighting scheme. Theoretical calculations demonstrate that the proposed geometry achieves a wider horizontal FoV than the ${90}\%$ detector offset geometry (radius of ${214}.{83}\textit {mm}$ vs. ${198}.{99}\textit {mm}$ ) with a significantly reduced rotation angle (less than 230° vs. 360°). As demonstrated by experiments, the proposed geometry and reconstruction algorithms obtain comparable imaging qualities within the FoV to conventional CBCT imaging techniques. Implementing the proposed geometry is straightforward and does not substantially increase development expenses. It possesses the capacity to expand CBCT applications even further. Tianling Lyu, Xinyun Zhong, Zhan Wu, Yan Xi, Wei Zhao 0029, Yang Chen 0008, Yuanjing Feng, Wentao Zhu 0002 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | 2V-CBCT: Two-Orthogonal-Projection Based CBCT Reconstruction and Dose Calculation for Radiation Therapy Using Real Projection DataabstractThis work demonstrates the feasibility of two-orthogonal-projection-based CBCT (2V-CBCT) reconstruction and dose calculation for radiation therapy (RT) using real projection data, which is the first 2V-CBCT feasibility study with real projection data, to the best of our knowledge. RT treatments are often delivered in multiple fractions, for which on-board CBCT is desirable to calculate the delivered dose per fraction for the purpose of RT delivery quality assurance and adaptive RT. However, not all RT treatments/fractions have CBCT acquired, but two orthogonal projections are always available. The question to be addressed in this work is the feasibility of 2V-CBCT for the purpose of RT dose calculation. 2V-CBCT is a severely ill-posed inverse problem for which we propose a coarse-to-fine learning strategy. First, a 3D deep neural network that can extract and exploit the inter-slice and intra-slice information is adopted to predict the initial 3D volumes. Then, a 2D deep neural network is utilized to fine-tune the initial 3D volumes slice-by-slice. During the fine-tuning stage, a perceptual loss based on multi-frequency features is employed to enhance the image reconstruction. Dose calculation results from both photon and proton RT demonstrate that 2V-CBCT provides comparable accuracy with full-view CBCT based on real projection data. Yikun Zhang 0001, Dianlin Hu, Wangyao Li, Gaoyu Chen, Ronald C. Chen, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | GDP-Net: Global Dependency-Enhanced Dual-Domain Parallel Network for Ring Artifact RemovalabstractIn Computed Tomography (CT) imaging, the ring artifacts caused by the inconsistent detector response can significantly degrade the reconstructed images, having negative impacts on the subsequent applications. The new generation of CT systems based on photon-counting detectors are affected by ring artifacts more severely. The flexibility and variety of detector responses make it difficult to build a well-defined model to characterize the ring artifacts. In this context, this study proposes the global dependency-enhanced dual-domain parallel neural network for Ring Artifact Removal (RAR). First, based on the fact that the features of ring artifacts are different in Cartesian and Polar coordinates, the parallel architecture is adopted to construct the deep neural network so that it can extract and exploit the latent features from different domains to improve the performance of ring artifact removal. Besides, the ring artifacts are globally relevant whether in Cartesian or Polar coordinate systems, but convolutional neural networks show inherent shortcomings in modeling long-range dependency. To tackle this problem, this study introduces the novel Mamba mechanism to achieve a global receptive field without incurring high computational complexity. It enables effective capture of the long-range dependency, thereby enhancing the model performance in image restoration and artifact reduction. The experiments on the simulated data validate the effectiveness of the dual-domain parallel neural network and the Mamba mechanism, and the results on two unseen real datasets demonstrate the promising performance of the proposed RAR algorithm in eliminating ring artifacts and recovering image details. Yikun Zhang 0001, Guannan Liu 0002, Shipeng Xie, Jiabing Gu, Zujian Huang, Tianling Lyu, Yan Xi, Shouping Zhu, Jian Yang 0009, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 12 |
| 2025 | LowDDAWP-Net: Low-Resolution Double Deep Audio Waveform Prior Network for Audio Systems Reliability DefenceabstractImproving the information reliability of the audio system is critical to safeguarding the security of the audio system. Adversarial samples crafted by in-the-wild attackers by introducing perturbations to the audio become a severe threat to the trustworthiness of deep learning-based classifiers. To achieve dynamic defence against audio adversarial sample attacks, a low-resolution double deep audio waveform prior network (LowDDAWP-Net) for audio systems reliability defence is proposed. Specifically, LowDDAWP-Net consists of a noise audio prior extraction module ($\mathbf{DAWP}_{\mathbf{noise}}$), an speech prior extraction module ($\mathbf{DAWP}_{\mathbf{speech}}$), a low-resolution extraction module (LREM), and a voice activity detection module (VADM). The role of the VADM is to automatically detect voice activity signals and silent signals from the audio signal.$\mathbf{DAWP}_{\mathbf{speech}}$and$\mathbf{DAWP}_{\mathbf{noise}}$are encoder–decoders with the same architecture. The encoder extracts the superficial features of the input audio, and the decoder performs temporal fusion to form high-dimensional features and reconstructs them into waveform signals. A LREM is employed to extract low-resolution audio to facilitate the encoder–decoder to perform detail on low-resolution audio and to speed up the recovery of DAWP networks to high resolution. The adversarial samples generated by several diverse attack state-of-the-art on three different datasets and their corresponding benign samples form a novel private dataset. The qualitative and quantitative results of the novel private dataset demonstrate the effectiveness and superiority of LowDDAWP-Net. Kai Chen 0039, Yikun Zhang 0001, Jiasong Wu, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE Trans. Reliab. | 5 |
| 2024 | D-NAF: Dynamic Neural Attenuation Fields for 4D CBCT Reconstruction in Pulmonary ImagingabstractFour-dimensional Cone-beam CT (4D-CBCT) has already been integrated into many commercial radiotherapy systems to facilitate image-guided radiotherapy (IGRT). These techniques reconstruct a sequence of three-dimensional (3D) CBCT volumes based on respiratory phases, providing motion-compensated images. Nevertheless, current 4D-CBCT methods still suffer from poor imaging quality due to the limited number of views used for phase-resolved reconstruction, and deep learning-based methods require high-quality training data, which is commonly unavailable. In this paper, we propose a self-supervised 4D-CBCT imaging method based on dynamic neural attenuation fields (D-NAF). The dynamic pulmonary CBCT images are encoded into a 4D implicit representation incorporating both spatial information and temporal information. Moreover, we employed a separable spatial-temporal encoding method to reduce the dimensionality of the solution space, leading to enhanced convergence and reconstruction quality. The results demonstrate that the proposed method yields superior imaging quality in comparison to alternative approaches, with an RMSE value of 59.82 HU, PSNR of 35.33 dB and SSIM of 0.9358. This method is expected to be a new benchmark for self-supervised 4D-CBCT imaging. Yuxuan Long, Tianling Lyu, Fan Rao, Yang Chen 0008, Wentao Zhu 0002 |
BIBM | 5 |
| 2024 | Clinical Insight-Augmented Multi-View Learning for Alzheimer's Detection in Retinal OCTA ImagesabstractAlzheimer’s disease (AD) poses a significant global challenge, with a notable absence of accessible and cost-effective diagnostic tools for widespread AD detection. The retina, mirroring the brain in anatomy and physiology, has emerged as a potential avenue for rapid AD identification through retinal imaging. The current retinal image-based AD detection methods usually focus primarily on the macular area, but ignore the potential value that the optic disc region may have for the detection task. In this study, we leverage both macular- and disc-centered OCTA images and propose a multi-region fusion framework for AD detection. Based on clinical evidence, we integrate handcrafted features into the framework to improve model performance and interpretability. Specifically, vascular morphological parameters extracted from the macular and disc regions are used as input to a revalued KNN model to improve predictive capabilities. Furthermore, recognizing the significance of extracting and utilizing complementary information from the macular and optic disc regions, we propose an uncertainty-guided strategy based on Dempster-Shefer Theory (DST) to fuse knowledge from different regions. This approach considers each region’s forecast quality and significantly improves the effectiveness and robustness of the model. Through comparative analysis with existing methods, we have demonstrated that our method outperforms the state-of-the-art ones and provides more valuable pathological evidence for the association between retinal vascular changes and AD. Yuandi Zhang, Jinkui Hao, Botian Zheng, Yonghuai Liu, Yanda Meng, Jiong Zhang 0004, Yang Chen 0008, Yitian Zhao |
BIBM | 7 |
| 2024 | A Rotation-Invariant Texture ViT for Fine-Grained Recognition of Esophageal Cancer Endoscopic Ultrasound Images
Shuaishuai Zhuang, Jiacheng Nie, Yusheng Guo, Guangquan Zhou, Jean-Louis Coatrieux, Yang Chen 0008 |
ECCV (31) | 8 |
| 2024 | Material Decomposition in Photon-Counting CT: A Deep Learning Approach Driven by Detector Physics and ASIC Modeling
Qianyu Wu, Wenhui Qin, Mengqing Su, Jinglu Ma, Yikun Zhang 0001, Guotao Quan, Yang Chen 0008, Yanfeng Du, Xiaochun Lai |
MICCAI (7) | 10 |
| 2024 | Multi-grained contrastive representation learning for label-efficient lesion segmentation and onset time classification of acute ischemic stroke
Yuhao Liu 0001, Yan Xi, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008 |
Medical Image Anal. | 8 |
| 2024 | Domain adaptive noise reduction with iterative knowledge transfer and style generalization learning
Yufei Tang, Tianling Lyu, Haoyang Jin, Yang Chen 0008, Jian Zheng 0001 |
Medical Image Anal. | 8 |
| 2024 | TEST-Net: transformer-enhanced Spatio-temporal network for infectious disease prediction
Kai Chen 0039, Tianjiao Ji, Guanyu Yang 0001, Yang Chen 0008 |
Multim. Syst. | 5 |
| 2024 | Global texture sensitive convolutional transformer for medical image steganalysis
Zhengyuan Zhou, Kai Chen 0039, Dianlin Hu, Huazhong Shu, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008 |
Multim. Syst. | 7 |
| 2024 | Real-time endoscopy haze removal: a synthetical method
Xu Zhuo, Chenbin Yang, Yuming Gao, Yang Chen 0008 |
Multim. Tools Appl. | 5 |
| 2024 | A flow-based multi-scale learning network for single image stochastic super-resolution
Qianyu Wu, Zhongqian Hu, Aichun Zhu, Jiaxin Zou, Yan Xi, Yang Chen 0008 |
Signal Process. Image Commun. | 7 |
| 2024 | RED-Net: Residual and Enhanced Discriminative Network for Image Steganalysis in the Internet of Medical Things and TelemedicineabstractInternet of Medical Things (IoMT) and telemedicine technologies utilize computers, communications, and medical devices to facilitate off-site exchanges between specialists and patients, specialists, and medical staff. If the information communicated in IoMT is illegally steganography, tampered or leaked during transmission and storage, it will directly impact patient privacy or the consultation results with possible serious medical incidents. Steganalysis is of great significance for the identification of medical images transmitted illegally in IoMT and telemedicine. In this article, we propose a Residual and Enhanced Discriminative Network (RED-Net) for image steganalysis in the internet of medical things and telemedicine. RED-Net consists of a steganographic information enhancement module, a deep residual network, and steganographic information discriminative mechanism. Specifically, a steganographic information enhancement module is adopted by the RED-Net to boost the illegal steganographic signal in texturally complex high-dimensional medical image features. A deep residual network is utilized for steganographic feature extraction and compression. A steganographic information discriminative mechanism is employed by the deep residual network to enable it to recalibrate the steganographic features and drop high-frequency features that are mistaken for steganographic information. Experiments conducted on public and private datasets with data hiding payloads ranging from 0.1bpp/bpnzac-0.5bpp/bpnzac in the spatial and JPEG domain led to RED-Net's steganalysis error$P_{\mathrm{E}}$in the range of 0.0732-0.0010 and 0.231-0.026, respectively. In general, qualitative and quantitative results on public and private datasets demonstrate that the RED-Net outperforms 8 state-of-art steganography detectors. Kai Chen 0039, Zhengyuan Zhou, Jiasong Wu, Jean-Louis Coatrieux, Yang Chen 0008, Gouenou Coatrieux |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | LeSAM: Adapt Segment Anything Model for Medical Lesion SegmentationabstractThe Segment Anything Model (SAM) is a foundational model that has demonstrated impressive results in the field of natural image segmentation. However, its performance remains suboptimal for medical image segmentation, particularly when delineating lesions with irregular shapes and low contrast. This can be attributed to the significant domain gap between medical images and natural images on which SAM was originally trained. In this paper, we propose an adaptation of SAM specifically tailored for lesion segmentation termed LeSAM. LeSAM first learns medical-specific domain knowledge through an efficient adaptation module and integrates it with the general knowledge obtained from the pre-trained SAM. Subsequently, we leverage this merged knowledge to generate lesion masks using a modified mask decoder implemented as a lightweight U-shaped network design. This modification enables better delineation of lesion boundaries while facilitating ease of training. We conduct comprehensive experiments on various lesion segmentation tasks involving different image modalities such as CT scans, MRI scans, ultrasound images, dermoscopic images, and endoscopic images. Our proposed method achieves superior performance compared to previous state-of-the-art methods in 8 out of 12 lesion segmentation tasks while achieving competitive performance in the remaining 4 datasets. Additionally, ablation studies are conducted to validate the effectiveness of our proposed adaptation modules and modified decoder. Yunbo Gu, Qianyu Wu, Xiaoli Mai, Huazhong Shu, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Image Domain Multi-Material Decomposition Noise Suppression Through Basis Transformation and Selective FilteringabstractSpectral CT can provide material characterization ability to offer more precise material information for diagnosis purposes. However, the material decomposition process generally leads to amplification of noise which significantly limits the utility of the material basis images. To mitigate such problem, an image domain noise suppression method was proposed in this work. The method performs basis transformation of the material basis images based on a singular value decomposition. The noise variances of the original spectral CT images were incorporated in the matrix to be decomposed to ensure that the transformed basis images are statistically uncorrelated. Due to the difference in noise amplitudes in the transformed basis images, a selective filtering method was proposed with the low-noise transformed basis image as guidance. The method was evaluated using both numerical simulation and real clinical dual-energy CT data. Results demonstrated that compared with existing methods, the proposed method performs better in preserving the spatial resolution and the soft tissue contrast while suppressing the image noise. The proposed method is also computationally efficient and can realize real-time noise suppression for clinical spectral CT images. Xu Zhuo, Weilong Mao, Guotao Quan, Yan Xi, Tianling Lyu, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | SBCNet: Scale and Boundary Context Attention Dual-Branch Network for Liver Tumor SegmentationabstractAutomated segmentation of liver tumors in CT scans is pivotal for diagnosing and treating liver cancer, offering a valuable alternative to labor-intensive manual processes and ensuring the provision of accurate and reliable clinical assessment. However, the inherent variability of liver tumors, coupled with the challenges posed by blurred boundaries in imaging characteristics, presents a substantial obstacle to achieving their precise segmentation. In this paper, we propose a novel dual-branch liver tumor segmentation model, SBCNet, to address these challenges effectively. Specifically, our proposed method introduces a contextual encoding module, which enables a better identification of tumor variability using an advanced multi-scale adaptive kernel. Moreover, a boundary enhancement module is designed for the counterpart branch to enhance the perception of boundaries by incorporating contour learning with the Sobel operator. Finally, we propose a hybrid multi-task loss function, concurrently concerning tumors' scale and boundary features, to foster interaction across different tasks of dual branches, further improving tumor segmentation. Experimental validation on the publicly available LiTS dataset demonstrates the practical efficacy of each module, with SBCNet yielding competitive results compared to other state-of-the-art methods for liver tumor segmentation. Kai-Ni Wang, Shengxiao Li, Zhenyu Bu, Fuxing Zhao, Guangquan Zhou, Shoujun Zhou, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | PRECISION: A Physics-Constrained and Noise-Controlled Diffusion Model for Photon Counting Computed TomographyabstractRecently, the use of photon counting detectors in computed tomography (PCCT) has attracted extensive attention. It is highly desired to improve the quality of material basis image and the quantitative accuracy of elemental composition, particularly when PCCT data is acquired at lower radiation dose levels. In this work, we develop a physics-constrained and noise-controlled diffusion model, PRECISION in short, to address the degraded quality of material basis images and inaccurate quantification of elemental composition mainly caused by imperfect noise model and/or hand-crafted regularization of material basis images, such as local smoothness and/or sparsity, leveraged in the existing direct material basis image reconstruction approaches. In stark contrast, PRECISION learns distribution-level regularization to describe the feature of ideal material basis images via training a noise-controlled spatial-spectral diffusion model. The optimal material basis images of each individual subject are sampled from this learned distribution under the constraint of the physical model of a given PCCT and the measured data obtained from the subject. PRECISION exhibits the potential to improve the quality of material basis images and the quantitative accuracy of elemental composition for PCCT. Guotao Quan, Yanfeng Du, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | SC-SSL: Self-Correcting Collaborative and Contrastive Co-Training Model for Semi-Supervised Medical Image SegmentationabstractImage segmentation achieves significant improvements with deep neural networks at the premise of a large scale of labeled training data, which is laborious to assure in medical image tasks. Recently, semi-supervised learning (SSL) has shown great potential in medical image segmentation. However, the influence of the learning target quality for unlabeled data is usually neglected in these SSL methods. Therefore, this study proposes a novel self-correcting co-training scheme to learn a better target that is more similar to ground-truth labels from collaborative network outputs. Our work has three-fold highlights. First, we advance the learning target generation as a learning task, improving the learning confidence for unannotated data with a self-correcting module. Second, we impose a structure constraint to encourage the shape similarity further between the improved learning target and the collaborative network outputs. Finally, we propose an innovative pixel-wise contrastive learning loss to boost the representation capacity under the guidance of an improved learning target, thus exploring unlabeled data more efficiently with the awareness of semantic context. We have extensively evaluated our method with the state-of-the-art semi-supervised approaches on four public-available datasets, including the ACDC dataset, M&Ms dataset, Pancreas-CT dataset, and Task_07 CT dataset. The experimental results with different labeled-data ratios show our proposed method's superiority over other existing methods, demonstrating its effectiveness in semi-supervised medical image segmentation. Juzheng Miao, Siping Zhou, Guangquan Zhou, Kai-Ni Wang, Shoujun Zhou, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Pathological Asymmetry-Guided Progressive Learning for Acute Ischemic Stroke Infarct SegmentationabstractQuantitative infarct estimation is crucial for diagnosis, treatment and prognosis in acute ischemic stroke (AIS) patients. As the early changes of ischemic tissue are subtle and easily confounded by normal brain tissue, it remains a very challenging task. However, existing methods often ignore or confuse the contribution of different types of anatomical asymmetry caused by intrinsic and pathological changes to segmentation. Further, inefficient domain knowledge utilization leads to mis-segmentation for AIS infarcts. Inspired by this idea, we propose a pathological asymmetry-guided progressive learning (PAPL) method for AIS infarct segmentation. PAPL mimics the step-by-step learning patterns observed in humans, including three progressive stages: knowledge preparation stage, formal learning stage, and examination improvement stage. First, knowledge preparation stage accumulates the preparatory domain knowledge of the infarct segmentation task, helping to learn domain-specific knowledge representations to enhance the discriminative ability for pathological asymmetries by constructed contrastive learning task. Then, formal learning stage efficiently performs end-to-end training guided by learned knowledge representations, in which the designed feature compensation module (FCM) can leverage the anatomy similarity between adjacent slices from the volumetric medical image to help aggregate rich anatomical context information. Finally, examination improvement stage encourages improving the infarct prediction from the previous stage, where the proposed perception refinement strategy (RPRS) further exploits the bilateral difference comparison to correct the mis-segmentation infarct regions by adaptively regional shrink and expansion. Extensive experiments on public and in-house NCCT datasets demonstrated the superiority of the proposed PAPL, which is promising to help better stroke evaluation and treatment. Qiuxuan Li, Yuhao Liu 0001, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008, Jie Lu 0010 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image ReconstructionabstractSpectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction. Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Learning Better Registration to Learn Better Few-Shot Medical Image Segmentation: Authenticity, Diversity, and RobustnessabstractIn this work, we address the task of few-shot medical image segmentation (MIS) with a novel proposed framework based on the learning registration to learn segmentation (LRLS) paradigm. To cope with the limitations of lack of authenticity, diversity, and robustness in the existing LRLS frameworks, we propose the better registration better segmentation (BRBS) framework with three main contributions that are experimentally shown to have substantial practical merit. First, we improve the authenticity in the registration-based generation program and propose the knowledge consistency constraint strategy that constrains the registration network to learn according to the domain knowledge. It brings the semantic-aligned and topology-preserved registration, thus allowing the generation program to output new data with great space and style authenticity. Second, we deeply studied the diversity of the generation process and propose the space-style sampling program, which introduces the modeling of the transformation path of style and space change between few atlases and numerous unlabeled images into the generation program. Therefore, the sampling on the transformation paths provides much more diverse space and style features to the generated data effectively improving the diversity. Third, we first highlight the robustness in the learning of segmentation in the LRLS paradigm and propose the mix misalignment regularization, which simulates the misalignment distortion and constrains the network to reduce the fitting degree of misaligned regions. Therefore, it builds regularization for these regions improving the robustness of segmentation learning. Without any bells and whistles, our approach achieves a new state-of-the-art performance in few-shot MIS on two challenging tasks that outperform the existing LRLS-based few-shot methods. We believe that this novel and effective framework will provide a powerful few-shot benchmark for the field of medical image and efficiently reduce the costs of medical image research. All of our code will be made publicly available online. Yuting He 0001, Rongjun Ge, Xiaoming Qi, Yang Chen 0008, Jiasong Wu, Jean-Louis Coatrieux, Guanyu Yang 0001, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | SpMVNet: Spatial Multi-view Network for Head and Neck Organs at Risk Segmentation
Hongzhi Liu 0002, Qianjin Feng 0001, Yang Chen 0008 |
ADMA (2) | 4 |
| 2023 | Specific-Modal Spatial Guidance and Feature Enhancement for Multi-modal Brain Tumor SegmentationabstractMulti-modal information plays a pivotal role in the segmentation of brain tumors. However, previous studies have largely overlooked the distinctive characteristics of individual modalities, which are correlated with the target tumor region due to distinct imaging principles. In this paper, we harness the distinctive traits of individual modalities and introduce a brain tumor segmentation model called specific modality guided brain tumor segmentation model (SMG-BTS). Our SMG-BTS adopts a three-branch encoder-decoder architecture. The main branch utilizes full modalities fused at input-level, while the two affiliated branches operate in parallel to provide guidance to the main branch in acquiring a robust representation. We propose a specific modality spatial guidance (SMSG) module to guide the process of feature extraction. Spatial information is obtained from selected modalities and utilized to enhance features extracted from the main branch. A shared-specific feature enhancement(SSFE) module is proposed to enhance the shared features across modalities and utilizes modality-specific features to further supplement specific information of modalities. Experimental results on the BraTS2021 benchmark dataset demonstrate the effectiveness of our proposed SMG-BTS over state-of-the-art brain tumor segmentation methods. Junyang Han, Cheng Xue 0003, Hongzhi Liu 0002, Jiacheng Nie, Weixuan Wan, Wenxue Yu, Yang Chen 0008, Pinzheng Zhang, Jean-Louis Coatrieux |
BIBM | 10 |
| 2023 | Spinal Lesions Classification and Localization with ACAT-Net from X-ray ImagesabstractX-ray images play an important role in the diagnosis of spinal diseases because of their convenient collection and easy observation. But it is time-consuming and challenging for radiologists to examine the differences between the vertebrae to diagnose abnormalities and locate lesions. Many existing methods try to extract the global features of radiographs and do not make full use of adjacent vertebrae variations. In this paper, we propose a novel Axial-aware neural network with Consecutive Attention Transformer (CAT), namely ACAT-Net, which takes advantage of the convolutional neural network and transformer as a new deep learning framework. A deep convolutional network extracts features of anteroposterior and lateral X-ray images that may have abnormalities in them. The consecutive attention transformer block is then used to focus on the morphological differences of axial adjacent vertebrae on the spines. The ingenious structure we designed can significantly reduce the amount of network parameters. Extensive experiments on clinical and public datasets show that our method is remarkably superior to other existing approaches in the spine X-ray image analysis. Hongzhi Liu 0002, Xiaoli Mai, Junyang Han, Jiacheng Nie, Weixuan Wan, Pinzheng Zhang, Wenxue Yu, Cheng Xue 0003, Qianjin Feng 0001, Yang Chen 0008 |
BIBM | 14 |
| 2023 | Data-consistent Unsupervised Diffusion Model for Metal Artifact ReductionabstractComputed Tomography (CT) is an imaging technique widely used in clinical diagnosis. However, high-attenuation metallic implants result in the obstruction of low-energy Xrays and further lead to metal artifacts in the reconstructed CT images. Deep supervised model-based metal artifact reduction(MAR) approaches are limited in clinical applications due to the difficulty in obtaining paired artifact-affected and artifactfree data. Furthermore, these model-based methods lack the consideration of data consistency in the sinogram-domain to perform exact metal trace inpainting. To address these challenges, we propose a Data-consistent unsupErVised diffusiOn model for meTal artifact rEDuction, called DEVOTED-Net. First, DEVOTED-Net leverages prior knowledge to guide the conditional diffusion model for fine-grained metal trace inpainting. Second, an unsupervised MAR framework is designed in the reverse process for the unknown metal traces restoration in the sinogram domain. Third, to further enhance the sinogram-domain data consistency, physics-based consistency constraint loss including conjugateray consistency loss and accumulation-ray consistency loss is designed. Extensive experiments are carried out to verify the performance of our algorithm on the publicly available dataset and clinical experimental dataset. This efficient, accurate, and reliable MAR approach holds great potential in clinics. Zhan Tong, Zhan Wu, Yang Yang 0216, Weilong Mao, Yang Chen 0008 |
BIBM | 7 |
| 2023 | DUALWISE-MAR: Dual-domain Multi-view Weakly supervised Segmentation Network for CBCT Metal Artifact ReductionabstractIntraoperative Cone-Beam Computed Tomography (CBCT) provides fast multiplanar cross-sectional imaging and three-dimensional reconstructions in clinical scene. However, high-attenuation metallic implants result in the obstruction of low-energy X-rays and further lead to metal artifacts in the reconstructed CT images. Existing metal artifact reduction approaches do not consider automated fine-grained metallic implants segmentation. Imprecise metallic implants segmentation severely limits the clinical practical applicability of MAR approaches. To address the above challenge, in this study, we present a novel DUAL-domain multi-view Weak supervIsed SEgmentation network for Metal Artifact Reduction (DUALWISE-MAR) to coarse-to-fine segment metallic implants without segmentation ground-truth. First, an Iterative Segmentation Refinement Mechanism is constructed to progressively enhance initial coarse label and guide segmentation network close to the ideal metal mask. Second, a Multi-view Consistent Segmentation Network is constructed for data-consistent segmentation to capture the metal shape relationship between adjacent views in projection-domain. Third, a Segmentation Recalibration Module is proposed to calibrate the inaccurate prediction regions in the imagedomain. Extensive experiments have been conducted on the real clinical CBCT dataset and demonstrate that the proposed DUALWISE-MAR framework achieves excellent performance in metal segmentation compared to the state-of-the-art methods. Xinyun Zhong, Zhan Wu, Yang Yang 0216, Tianling Lyu, Wenxue Yu, Yan Xi, Yang Chen 0008 |
BIBM | 7 |
| 2023 | Geometric Visual Similarity Learning in 3D Medical Image Self-Supervised Pre-trainingabstractLearning inter-image similarity is crucial for 3D medical images self-supervised pre-training, due to their sharing of numerous same semantic regions. However, the lack of the semantic prior in metrics and the semantic-independent variation in 3D medical images make it challenging to get a reliable measurement for the inter-image similarity, hindering the learning of consistent representation for same semantics. We investigate the challenging problem of this task, i.e., learning a consistent representation between images for a clustering effect of same semantic features. We propose a novel visual similarity learning paradigm, Geometric Visual Similarity Learning, which embeds the prior of topological invariance into the measurement of the inter-image similarity for consistent representation of semantic regions. To drive this paradigm, we further construct a novel geometric matching head, the Z-matching head, to collaboratively learn the global and local similarity of semantic regions, guiding the efficient representation learning for different scale-level inter-image semantic features. Our experiments demonstrate that the pre-training with our learning of inter-image similarity yields more powerful inner-scene, inter-scene, and global-local transferring ability on four challenging 3D medical image tasks. Our codes and pre-trained models will be publicly available11https://github.com/YutingHe-list/GVSL. Yuting He 0001, Guanyu Yang 0001, Rongjun Ge, Yang Chen 0008, Jean-Louis Coatrieux, Boyu Wang 0004, Shuo Li 0001 |
CVPR | 4 |
| 2023 | O2M-UDA: Unsupervised dynamic domain adaptation for one-to-multiple medical image segmentation
Ziyue Jiang 0004, Yuting He 0001, Xiaomei Zhu, Yi Xu 0001, Yang Chen 0008, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001 |
Knowl. Based Syst. | 7 |
| 2023 | TIME-Net: Transformer-Integrated Multi-Encoder Network for limited-angle artifact removal in dual-energy CBCT
Yikun Zhang 0001, Dianlin Hu, Zhihong Yan, Qingxian Zhao, Guotao Quan, Shouhua Luo, Yi Zhang 0018, Yang Chen 0008 |
Medical Image Anal. | 8 |
| 2023 | Deformable multi-scale fusion network for non-uniform single image deblurring
Yang Chen 0008, Aichun Zhu, Hanxi Liu |
Multim. Tools Appl. | 2 |
| 2023 | Adaptive Frequency Learning Network With Anti-Aliasing Complex Convolutions for Colon Diseases SubtypesabstractThe automatic and dependable identification of colonic disease subtypes by colonoscopy is crucial. Once successful, it will facilitate clinically more in-depth disease staging analysis and the formulation of more tailored treatment plans. However, inter-class confusion and brightness imbalance are major obstacles to colon disease subtyping. Notably, the Fourier-based image spectrum, with its distinctive frequency features and brightness insensitivity, offers a potential solution. To effectively leverage its advantages to address the existing challenges, this article proposes a framework capable of thorough learning in the frequency domain based on four core designs: the position consistency module, the high-frequency self-supervised module, the complex number arithmetic model, and the feature anti-aliasing module. The position consistency module enables the generation of spectra that preserve local and positional information while compressing the spectral data range to improve training stability. Through band masking and supervision, the high-frequency autoencoder module guides the network to learn useful frequency features selectively. The proposed complex number arithmetic model allows direct spectral training while avoiding the loss of phase information caused by current general-purpose real-valued operations. The feature anti-aliasing module embeds filters in the model to prevent spectral aliasing caused by down-sampling and improve performance. Experiments are performed on the collected five-class dataset, which contains 4591 colorectal endoscopic images. The outcomes show that our proposed method produces state-of-the-art results with an accuracy rate of 89.82%. Kai-Ni Wang, Shuaishuai Zhuang, Juzheng Miao, Yang Chen 0008, Jie Hua 0004, Guangquan Zhou, Xiaopu He, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Multi-Task Learning for Pulmonary Arterial Hypertension Prognosis Prediction via Memory Drift and Prior Prompt Learning on 3D Chest CTabstractPulmonary arterial hypertension (PAH) prognosis prediction on 3D non-contrast CT images is one of the most important tasks for PAH treatment. It will help clinicians stratify patients into different groups for early diagnosis and timely intervention via automatically extracting the potential biomarkers of PAH to predict mortality. However, it is still a task of great challenges due to the large volume and low-contrast regions of interest in 3D chest CT images. In this paper, we propose the first multi-task learning-based PAH prognosis prediction framework, P$^{2}$-Net, which effectively optimizes the model and powerfully represents task-dependent features via our Memory Drift (MD) and Prior Prompt Learning (PPL) strategies. 1) Our MD maintains a large memory bank to provide a dense sampling of the deep biomarkers' distribution. Therefore, although the batch size is very small caused by our large volume, a reliable (negative log partial) likelihood loss is still able to be calculated on a representative probability distribution for robust optimization. 2) Our PPL simultaneously learns an additional manual biomarkers prediction task to embed clinical prior knowledge into our deep prognosis prediction task in hidden and explicit ways. Therefore, it will prompt the prediction of deep biomarkers and improve the perception of task-dependent features in our low-contrast regions. Our P$^{2}$-Net achieves a high prognostic correlation of the prediction and great generalization with the highest 70.19% C-index and 2.14 HR. Extensive experiments with promising results on our PAH prognosis prediction reveal powerful prognosis performance and great clinical significance in PAH treatment. All of our code will be made publicly available online. Guanyu Yang 0001, Yuting He 0001, Yang Chen 0008, Jean-Louis Coatrieux, Xiaoxuan Sun, Yongyue Wei, Shuo Li 0001, Yinsu Zhu |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | DREAM-Net: Deep Residual Error Iterative Minimization Network for Sparse-View CT ReconstructionabstractSparse-view Computed Tomography (CT) has the ability to reduce radiation dose and shorten the scan time, while the severe streak artifacts will compromise anatomical information. How to reconstruct high-quality images from sparsely sampled projections is a challenging ill-posed problem. In this context, we propose the unrolled Deep Residual Error iterAtive Minimization Network (DREAM-Net) based on a novel iterative reconstruction framework to synergize the merits of deep learning and iterative reconstruction. DREAM-Net performs constraints using deep neural networks in the projection domain, residual space, and image domain simultaneously, which is different from the routine practice in deep iterative reconstruction frameworks. First, a projection inpainting module completes the missing views to fully explore the latent relationship between projection data and reconstructed images. Then, the residual awareness module attempts to estimate the accurate residual image after transforming the projection error into the image space. Finally, the image refinement module learns a non-standard regularizer to further fine-tune the intermediate image. There is no need to empirically adjust the weights of different terms in DREAM-Net because the hyper-parameters are embedded implicitly in network modules. Qualitative and quantitative results have demonstrated the promising performance of DREAM-Net in artifact removal and structural fidelity. Yikun Zhang 0001, Dianlin Hu, Shilei Hao, Jin Liu 0019, Guotao Quan, Yi Zhang 0018, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | DSANet: Dual-Branch Shape-Aware Network for Echocardiography Segmentation in Apical ViewsabstractEchocardiography is an essential examination for cardiac disease diagnosis, from which anatomical structures segmentation is the key to assessing various cardiac functions. However, the obscure boundaries and large shape deformations due to cardiac motion make it challenging to accurately identify the anatomical structures in echocardiography, especially for automatic segmentation. In this study, we propose a dual-branch shape-aware network (DSANet) to segment the left ventricle, left atrium, and myocardium from the echocardiography. Specifically, the elaborate dual-branch architecture integrating shape-aware modules boosts the corresponding feature representation and segmentation performance, which guides the model to explore shape priors and anatomical dependence using an anisotropic strip attention mechanism and cross-branch skip connections. Moreover, we develop a boundary-aware rectification module together with a boundary loss to regulate boundary consistency, adaptively rectifying the estimation errors nearby the ambiguous pixels. We evaluate our proposed method on the publicly available and in-house echocardiography dataset. Comparative experiments with other state-of-the-art methods demonstrate the superiority of DSANet, which suggests its potential in advancing echocardiography segmentation. Guangquan Zhou, Wen-Bo Zhang, Zhong-Qing Shi, Zhan-Ru Qi, Kai-Ni Wang, Hong Song 0003, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Low-Dose CT Image Synthesis for Domain Adaptation Imaging Using a Generative Adversarial Network With Noise Encoding Transfer LearningabstractDeep learning (DL) based image processing methods have been successfully applied to low-dose x-ray images based on the assumption that the feature distribution of the training data is consistent with that of the test data. However, low-dose computed tomography (LDCT) images from different commercial scanners may contain different amounts and types of image noise, violating this assumption. Moreover, in the application of DL based image processing methods to LDCT, the feature distributions of LDCT images from simulation and clinical CT examination can be quite different. Therefore, the network models trained with simulated image data or LDCT images from one specific scanner may not work well for another CT scanner and image processing task. To solve such domain adaptation problem, in this study, a novel generative adversarial network (GAN) with noise encoding transfer learning (NETL), or GAN-NETL, is proposed to generate a paired dataset with a different noise style. Specifically, we proposed a method to perform noise encoding operator and incorporate it into the generator to extract a noise style. Meanwhile, with a transfer learning (TL) approach, the image noise encoding operator transformed the noise type of the source domain to that of the target domain for realistic noise generation. One public and two private datasets are used to evaluate the proposed method. Experiment results demonstrated the feasibility and effectiveness of our proposed GAN-NETL model in LDCT image synthesis. In addition, we conduct additional image denoising study using the synthesized clinical LDCT data, which verified the merit of the proposed synthesis in improving the performance of the DL based LDCT processing method. Yang Chen 0008, Yufei Tang, Zhongyi Wu, Yujin Qi, Haochuan Jiang, Jian Zheng 0001, Benjamin M. W. Tsui |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Unsharp Structure Guided Filtering for Self-Supervised Low-Dose CT ImagingabstractLow-dose computed tomography (LDCT) imaging faces great challenges. Although supervised learning has revealed great potential, it requires sufficient and high-quality references for network training. Therefore, existing deep learning methods have been sparingly applied in clinical practice. To this end, this paper presents a novel Unsharp Structure Guided Filtering (USGF) method, which can reconstruct high-quality CT images directly from low-dose projections without clean references. Specifically, we first employ low-pass filters to estimate the structure priors from the input LDCT images. Then, inspired by classical structure transfer techniques, deep convolutional networks are adopted to implement our imaging method which combines guided filtering and structure transfer. Finally, the structure priors serve as the guidance images to alleviate over-smoothing, as they can transfer specific structural characteristics to the generated images. Furthermore, we incorporate traditional FBP algorithms into self-supervised training to enable the transformation of projection domain data to the image domain. Extensive comparisons and analyses on three datasets demonstrate that the proposed USGF has achieved superior performance in terms of noise suppression and edge preservation, and could have a significant impact on LDCT imaging in the future. Qianyu Wu, Yunbo Gu, Guotao Quan, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 10 |
| 2022 | MNet: Rethinking 2D/3D Networks for Anisotropic Medical Image SegmentationabstractThe nature of thick-slice scanning causes severe inter-slice discontinuities of 3D medical images, and the vanilla 2D/3D convolutional neural networks (CNNs) fail to represent sparse inter-slice information and dense intra-slice information in a balanced way, leading to severe underfitting to inter-slice features (for vanilla 2D CNNs) and overfitting to noise from long-range slices (for vanilla 3D CNNs). In this work, a novel mesh network (MNet) is proposed to balance the spatial representation inter axes via learning. 1) Our MNet latently fuses plenty of representation processes by embedding multi-dimensional convolutions deeply into basic modules, making the selections of representation processes flexible, thus balancing representation for sparse inter-slice information and dense intra-slice information adaptively. 2) Our MNet latently fuses multi-dimensional features inside each basic module, simultaneously taking the advantages of 2D (high segmentation accuracy of the easily recognized regions in 2D view) and 3D (high smoothness of 3D organ contour) representations, thus obtaining more accurate modeling for target regions. Comprehensive experiments are performed on four public datasets (CT\&MR), the results consistently demonstrate the proposed MNet outperforms the other methods. The code and datasets are available at: https://github.com/zfdong-code/MNet Zhangfu Dong, Yuting He 0001, Xiaoming Qi, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux, Guanyu Yang 0001, Shuo Li 0001 |
IJCAI | 4 |
| 2022 | DDPNet: A Novel Dual-Domain Parallel Network for Low-Dose CT Reconstruction
Rongjun Ge, Yuting He 0001, Cong Xia, Hai-Long Sun, Yikun Zhang 0001, Dianlin Hu, Yang Chen 0008, Shuo Li 0001, Daoqiang Zhang |
MICCAI (6) | 8 |
| 2022 | X-CTRSNet: 3D cervical vertebra CT reconstruction and segmentation directly from 2D X-ray images
Rongjun Ge, Yuting He 0001, Cong Xia, Chenchu Xu, Weiya Sun, Guanyu Yang 0001, Hailing Yu, Daoqiang Zhang, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001, Yinsu Zhu |
Knowl. Based Syst. | 11 |
| 2022 | RE-3DLVNet: Refined estimation of the left ventricle volume via interactive 3D segmentation and reinforced quantification
Rongjun Ge, Cong Xia, Yuting He 0001, Hai-Long Sun, Daoqiang Zhang, Guanyu Yang 0001, Wentao Xiang, Jinjun Shi, Limin Luo 0001, Yinsu Zhu, Shuo Li 0001, Yang Chen 0008 |
Knowl. Based Syst. | 12 |
| 2022 | BKC-Net: Bi-Knowledge Contrastive Learning for renal tumor diagnosis on 3D CT images
Jindi Kong, Yuting He 0001, Xiaomei Zhu, Yi Xu 0001, Yang Chen 0008, Jean-Louis Coatrieux, Guanyu Yang 0001 |
Knowl. Based Syst. | 6 |
| 2022 | Projection network with Spatio-temporal information: 2D + time DSA to 2D aorta segmentation
Weiya Sun, Yuting He 0001, Rongjun Ge, Guanyu Yang 0001, Yang Chen 0008, Huazhong Shu |
Multim. Tools Appl. | 5 |
| 2022 | Online Hard Patch Mining Using Shape Models and Bandit Algorithm for Multi-Organ SegmentationabstractHard sample selection can effectively improve model convergence by extracting the most representative samples from a training set. However, due to the large capacity of medical images, existing sampling strategies suffer from insufficient exploitation for hard samples or high time cost for sample selection when adopted by 3D patch-based models in the field of multi-organ segmentation. In this paper, we present a novel and effective online hard patch mining (OHPM) algorithm. In our method, an average shape model that can be mapped with all training images is constructed to guide the exploration of hard patches and aggregate feedback from predicted patches. The process of hard mining is formalized as a multi-armed bandit problem and solved with bandit algorithms. With the shape model, OHPM requires negligible time consumption and can intuitively locate difficult anatomical areas during training. The employment of bandit algorithms ensures online and sufficient hard mining. We integrate OHPM with advanced segmentation networks and evaluate them on two datasets containing different anatomical structures. Comparative experiments with other sampling strategies demonstrate the superiority of OHPM in boosting segmentation performance and improving model convergence. The results in each dataset with each network suggest that OHPM significantly outperforms other sampling strategies by nearly 2% average Dice score. Jianan He, Guangquan Zhou, Shoujun Zhou, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | PRIOR: Prior-Regularized Iterative Optimization Reconstruction For 4D CBCTabstract4D cone-beam computed tomography (CBCT) is an important imaging modality in image-guided radiation therapy to address the motion-induced artifacts caused by organ movements during the respiratory process. However, due to the extremely sparse projection data for each temporal phase, 4D CBCT reconstructions will suffer from severe streaking artifacts. Therefore, to tackle the streak artifacts and provide high-quality images, we proposed a framework termed Prior-Regularized Iterative Optimization Reconstruction (PRIOR) for 4D CBCT. The PRIOR framework combines the physics-based model and data-driven method simultaneously, with powerful feature extracting capacity, significantly promoting the image quality compared to single model-based or deep learning-based methods. Besides, we designed a specialized deep learning model named PRIOR-Net, which can effectively excavate the static information in the prior image reconstructed from the fully-sampled projections at the encoding stage to improve the reconstruction performance for individual phase-resolved images. Both the simulated and clinical 4D CBCT datasets were performed to evaluate the performance of the PRIOR-Net and the PRIOR framework. Compared with the advanced 4D CBCT reconstruction methods, the proposed methods achieve promising results quantitatively and qualitatively in streak artifact suppression, soft tissue restoration, and tiny detail preservation. Dianlin Hu, Yikun Zhang 0001, Jin Liu 0019, Yi Zhang 0018, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | MVSGAN: Spatial-Aware Multi-View CMR Fusion for Accurate 3D Left Ventricular Myocardium SegmentationabstractThe accurate 3D left ventricular (LV) myocardium segmentation in short-axis (SAX) view of cardiac magnetic resonance (CMR) is challenged by the sparse spatial structure of CMR. The strategy of multi-view CMR fusion can provide fine-grained spatial structure for accurate segmentation. However, the large information misalignment and lack of dense 3D CMR as fusion target in multi-view CMR fusion, and the different spatial resolution between the fusion result and the ground truth in segmentation limit the strategy. In this study, we propose a multi-view spatial-aware adversarial network (MVSGAN). It studies the perception of fine-grained cardiac structure for accurate segmentation by the spatialaware multi-view CMR fusion. It consists of three modules: (1) A residual adversarial fusion (RAF) module takes inter-slices deep correlation and anatomical prior to refine the spatial structures by residual supplement and adversarial optimization. (2) A structural perception-aggregation (SPA) module establishes the spatial correlation between the dense cardiac model and sparse label for accurate CMR LV myocardium segmentation. (3) A joint training strategy utilizes the dense SAX volume as explicit and implicit goals to jointly optimize the framework. The experiments are applied on a public dataset and a clinical dataset to evaluate the performance of MVSGAN. The average Dice and Jaccard score of LV myocardium segmentation obtained by MVSGAN are highest among seven existing state-of-the-art methods, which are up to 0.92 and 0.75. It is concluded that the spatial-aware multi-view CMR fusion can provide meaningful spatial correlation for accurate LV myocardium segmentation. Xiaoming Qi, Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Jian Yang 0009, Wangyag Liu, Yinsu Zhu, Yi Xu 0001, Huazhong Shu, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Automatic Video Analysis Framework for Exposure Region Recognition in X-Ray Imaging AutomationabstractThe deep learning-based automatic recognition of the scanning or exposing region in medical imaging automation is a promising new technique, which can decrease the heavy workload of the radiographers, optimize imaging workflow and improve image quality. However, there is little related research and practice in X-ray imaging. In this paper, we focus on two key problems in X-ray imaging automation: automatic recognition of the exposure moment and the exposure region. Consequently, we propose an automatic video analysis framework based on the hybrid model, approaching real-time performance. The framework consists of three interdependent components: Body Structure Detection, Motion State Tracing, and Body Modeling. Body Structure Detection disassembles the patient to obtain the corresponding body keypoints and body Bboxes. Combining and analyzing the two different types of body structure representations is to obtain rich spatial location information about the patient body structure. Motion State Tracing focuses on the motion state analysis of the exposure region to recognize the appropriate exposure moment. The exposure region is calculated by Body Modeling when the exposure moment appears. A large-scale dataset for X-ray examination scene is built to validate the performance of the proposed method. Extensive experiments demonstrate the superiority of the proposed method in automatically recognizing the exposure moment and exposure region. This paradigm provides the first method that can enable automatically and accurately recognize the exposure region in X-ray imaging without the help of the radiographer. Zhan Wu, Zechen Yu, Huanji Chen, Changping Du, Juan Feng 0003, Gouenou Coatrieux, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 11 |
| 2022 | Masked Joint Bilateral Filtering via Deep Image Prior for Digital X-Ray Image DenoisingabstractMedical image denoising faces great challenges. Although deep learning methods have shown great potential, their efficiency is severely affected by millions of trainable parameters. The non-linearity of neural networks also makes them difficult to be understood. Therefore, existing deep learning methods have been sparingly applied to clinical tasks. To this end, we integrate known filtering operators into deep learning and propose a novel Masked Joint Bilateral Filtering (MJBF) via deep image prior for digital X-ray image denoising. Specifically, MJBF consists of a deep image prior generator and an iterative filtering block. The deep image prior generator produces plentiful image priors by a multi-scale fusion network. The generated image priors serve as the guidance for the iterative filtering block, which is utilized for the actual edge-preserving denoising. The iterative filtering block contains three trainable Joint Bilateral Filters (JBFs), each with only 18 trainable parameters. Moreover, a masking strategy is introduced to reduce redundancy and improve the understanding of the proposed network. Experimental results on the ChestX-ray14 dataset and real data show that the proposed MJBF has achieved superior performance in terms of noise suppression and edge preservation. Tests on the portability of the proposed method demonstrate that this denoising modality is simple yet effective, and could have a clinical impact on medical imaging in the future. Qianyu Wu, Hanxi Liu, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | DIOR: Deep Iterative Optimization-Based Residual-Learning for Limited-Angle CT ReconstructionabstractLimited-angle CT is a challenging problem in real applications. Incomplete projection data will lead to severe artifacts and distortions in reconstruction images. To tackle this problem, we propose a novel reconstruction framework termed Deep Iterative Optimization-based Residual-learning (DIOR) for limited-angle CT. Instead of directly deploying the regularization term on image space, the DIOR combines iterative optimization and deep learning based on the residual domain, significantly improving the convergence property and generalization ability. Specifically, the asymmetric convolutional modules are adopted to strengthen the feature extraction capacity in smooth regions for deep priors. Besides, in our DIOR method, the information contained in low-frequency and high-frequency components is also evaluated by perceptual loss to improve the performance in tissue preservation. Both simulated and clinical datasets are performed to validate the performance of DIOR. Compared with existing competitive algorithms, quantitative and qualitative results show that the proposed method brings a promising improvement in artifact removal, detail restoration and edge preservation. Dianlin Hu, Yikun Zhang 0001, Jin Liu 0019, Shouhua Luo, Yang Chen 0008 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Thin Semantics Enhancement via High-Frequency Priori Rule for Thin Structures SegmentationabstractReceptive field-based segmentation models represent features in receptive fields having weak perception for thin semantics in thin structures segmentation, due to the challenges in small local size and large global variation. High-frequency (HiFe) components have strong thin perception ability and is stable for global variation, but its weak adaptability limits its direct application. We propose a HiFe priori rule which enables the network to adaptively extract and fuse HiFe components, enhancing the thin semantics and making the network naturally prefer thin structures for their segmentation. We further propose High-Frequency Semantics Enhancement Network (HiFeNet) based on our HiFe priori rule, boosting the SOTA methods in thin structures segmentation: 1) Our Deep High Frequency (DHiFe) block learns to extract task-dependent HiFe components and adds them to feature maps, achieving great perception of thin structures. 2) Our Latent Residual Denoising (LRD) block progressively weakens task-independent features via hierarchical residuals and learns to fuse HiFe components back to feature maps, further enhancing the thin semantics and weakening the interference of global variation. Extensive experiments on the retinal vessel [1], [2], [3] and Massachusetts road [4] segmentation datasets show great superiority of our HiFeNet. Yuting He 0001, Rongjun Ge, Jiasong Wu, Jean-Louis Coatrieux, Huazhong Shu, Yang Chen 0008, Guanyu Yang 0001, Shuo Li 0001 |
ICDM | 6 |
| 2021 | Convolutional squeeze-and-excitation network for ECG arrhythmia detection
Rongjun Ge, Tengfei Shen, Chengyu Liu 0001, Benqiang Yang, Jean-Louis Coatrieux, Yang Chen 0008 |
Artif. Intell. Medicine | 9 |
| 2021 | Denoising auto-encoding priors in undecimated wavelet domain for MR image reconstruction
Junjie Lv, Zhuonan He, Dong Liang 0001, Yang Chen 0008, Qiegen Liu |
Neurocomputing | 5 |
| 2021 | Estimating dual-energy CT imaging from single-energy CT data with material decomposition convolutional neural network
Tianling Lyu, Wei Zhao 0029, Yinsu Zhu, Zhan Wu, Yikun Zhang 0001, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001, Lei Xing 0001 |
Medical Image Anal. | 6 |
| 2021 | Marginal loss and exclusion loss for partially supervised multi-organ segmentation
Gonglei Shi, Li Xiao 0005, Yang Chen 0008, Shaohua Kevin Zhou |
Medical Image Anal. | 3 |
| 2021 | ELNet: Automatic classification and segmentation for esophageal lesions using convolutional neural network
Zhan Wu, Rongjun Ge, Minli Wen, Gaoshuang Liu, Yang Chen 0008, Pinzheng Zhang, Xiaopu He, Jie Hua 0004, Limin Luo 0001, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2021 | ResNet-SCDA-50 for Breast Abnormality Classificationabstract(Aim) Breast cancer is the most common cancer in women and the second most common cancer worldwide. With the rapid advancement of deep learning, the early stages of breast cancer development can be accurately detected by radiologists with the help of artificial intelligence systems. (Method) Based on mammographic imaging, a mainstream clinical breast screening technique, we present a diagnostic system for accurate classification of breast abnormalities based on ResNet-50. To improve the proposed model, we created a new data augmentation framework called SCDA (Scaling and Contrast limited adaptive histogram equalization Data Augmentation). In its procedure, we first conduct the scaling operation to the original training set, followed by applying contrast limited adaptive histogram equalisation (CLAHE) to the scaled training set. By stacking the training set after SCDA with the original training set, we formed a new training set. The network trained by the augmented training set, was coined as ResNet-SCDA-50. Our system, which aims at a binary classification on mammographic images acquired from INbreast and MINI-MIAS, classifies masses, microcalcification as "abnormal", while normal regions are classified as "normal". (Results) We present the first attempt to use the image contrast enhancement method as the data augmentation method, resulting in an averaged 98.55 percent specificity and 92.83 percent sensitivity, which gives our best model an overall accuracy of 95.74 percent. (Conclusion) Our proposed method is effective in classifying breast abnormality. Cheng Kang, David S. Guttery, Seifedine Nimer Kadry, Yang Chen 0008, Yudong Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Homotopic Gradients of Generative Density Priors for MR Image ReconstructionabstractDeep learning, particularly the generative model, has demonstrated tremendous potential to significantly speed up image reconstruction with reduced measurements recently. Rather than the existing generative models that often optimize the density priors, in this work, by taking advantage of the denoising score matching, homotopic gradients of generative density priors (HGGDP) are exploited for magnetic resonance imaging (MRI) reconstruction. More precisely, to tackle the low-dimensional manifold and low data density region issues in generative density prior, we estimate the target gradients in higher-dimensional space. We train a more powerful noise conditional score network by forming high-dimensional tensor as the network input at the training phase. More artificial noise is also injected in the embedding space. At the reconstruction stage, a homotopy method is employed to pursue the density prior, such as to boost the reconstruction performance. Experiment results implied the remarkable performance of HGGDP in terms of high reconstruction accuracy. Only 10% of the k-space data can still generate image of high quality as effectively as standard MRI reconstructions with the fully sampled data. Cong Quan, Jinjie Zhou, Yuanzheng Zhu, Yang Chen 0008, Shanshan Wang 0002, Dong Liang 0001, Qiegen Liu |
IEEE Trans. Medical Imaging | 4 |
| 2021 | MAGIC: Manifold and Graph Integrative Convolutional Network for Low-Dose CT ReconstructionabstractLow-dose computed tomography (LDCT) scans, which can effectively alleviate the radiation problem, will degrade the imaging quality. In this paper, we propose a novel LDCT reconstruction network that unrolls the iterative scheme and performs in both image and manifold spaces. Because patch manifolds of medical images have low-dimensional structures, we can build graphs from the manifolds. Then, we simultaneously leverage the spatial convolution to extract the local pixel-level features from the images and incorporate the graph convolution to analyze the nonlocal topological features in manifold space. The experiments show that our proposed method outperforms both the quantitative and qualitative aspects of state-of-the-art methods. In addition, aided by a projection loss component, our proposed method also demonstrates superior performance for semi-supervised learning. The network can remove most noise while maintaining the details of only 10% (40 slices) of the training data labeled. Wenjun Xia, Yongqiang Huang 0003, Zuoqiang Shi, Yan Liu 0052, Hu Chen 0002, Yang Chen 0008, Jiliu Zhou, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | CLEAR: Comprehensive Learning Enabled Adversarial Reconstruction for Subtle Structure Enhanced Low-Dose CT ImagingabstractX-ray computed tomography (CT) is of great clinical significance in medical practice because it can provide anatomical information about the human body without invasion, while its radiation risk has continued to attract public concerns. Reducing the radiation dose may induce noise and artifacts to the reconstructed images, which will interfere with the judgments of radiologists. Previous studies have confirmed that deep learning (DL) is promising for improving low-dose CT imaging. However, almost all the DL-based methods suffer from subtle structure degeneration and blurring effect after aggressive denoising, which has become the general challenging issue. This paper develops the Comprehensive Learning Enabled Adversarial Reconstruction (CLEAR) method to tackle the above problems. CLEAR achieves subtle structure enhanced low-dose CT imaging through a progressive improvement strategy. First, the generator established on the comprehensive domain can extract more features than the one built on degraded CT images and directly map raw projections to high-quality CT images, which is significantly different from the routine GAN practice. Second, a multi-level loss is assigned to the generator to push all the network components to be updated towards high-quality reconstruction, preserving the consistency between generated images and gold-standard images. Finally, following the WGAN-GP modality, CLEAR can migrate the real statistical properties to the generated images to alleviate over-smoothing. Qualitative and quantitative analyses have demonstrated the competitive performance of CLEAR in terms of noise suppression, structural fidelity and visual perception improvement. Yikun Zhang 0001, Dianlin Hu, Qianlong Zhao, Guotao Quan, Jin Liu 0019, Qiegen Liu, Yi Zhang 0018, Gouenou Coatrieux, Yang Chen 0008, Hengyong Yu |
IEEE Trans. Medical Imaging | 9 |
| 2020 | One-Shot Learning for Long-Tail Visual Relation DetectionabstractThe aim of visual relation detection is to provide a comprehensive understanding of an image by describing all the objects within the scene, and how they relate to each other, in < object-predicate-object > form; for example, < person-lean on-wall > . This ability is vital for image captioning, visual question answering, and many other applications. However, visual relationships have long-tailed distributions and, thus, the limited availability of training samples is hampering the practicability of conventional detection approaches. With this in mind, we designed a novel model for visual relation detection that works in one-shot settings. The embeddings of objects and predicates are extracted through a network that includes a feature-level attention mechanism. Attention alleviates some of the problems with feature sparsity, and the resulting representations capture more discriminative latent features. The core of our model is a dual graph neural network that passes and aggregates the context information of predicates and objects in an episodic training scheme to improve recognition of the one-shot predicates and then generate the triplets. To the best of our knowledge, we are the first to center on the viability of one-shot learning for visual relation detection. Extensive experiments on two newly-constructed datasets show that our model significantly improved the performance of two tasks PredCls and SGCls from 2.8% to 12.2% compared with state-of-the-art baselines. Meng Wang 0009, Sen Wang 0001, Guodong Long, Lina Yao 0001, Guilin Qi, Yang Chen 0008 |
AAAI | 7 |
| 2020 | Deep Complementary Joint Model for Complex Scene Registration and Few-Shot Segmentation on Medical Images
Yuting He 0001, Guanyu Yang 0001, Youyong Kong, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux, Jean-Louis Dillenseger, Shuo Li 0001 |
ECCV (18) | 5 |
| 2020 | Multi-vertebrae Segmentation from Arbitrary Spine MR Images Under Global View
Heyou Chang, Yang Chen 0008, Shuo Li 0001 |
MICCAI (6) | 4 |
| 2020 | Memory-Based Network for Scene Graph with Unbalanced RelationsabstractThe scene graph which can be represented by a set of visual triples is composed of objects and the relations between object pairs. It is vital for image captioning, visual question answering, and many other applications. However, there is a long tail distribution on the scene graph dataset, and the tail relation cannot be accurately identified due to the lack of training samples. The problem of the nonstandard label and feature overlap on the scene graph affects the extraction of discriminative features and exacerbates the effect of data imbalance on the model. For these reasons, we propose a novel scene graph generation model that can effectively improve the detection of low-frequency relations. We use the method of memory features to realize the transfer of high-frequency relation features to low-frequency relation features. Extensive experiments on scene graph datasets show that our model significantly improved the performance of two evaluation metrics [email protected] and [email protected] compared with state-of-the-art baselines. Ruyang Liu, Meng Wang 0009, Sen Wang 0001, Xiaojun Chang, Yang Chen 0008 |
ACM Multimedia | 6 |
| 2020 | Coarse-to-fine classification for diabetic retinopathy grading using convolutional neural network
Zhan Wu, Gonglei Shi, Yang Chen 0008, Xinjian Chen 0001, Gouenou Coatrieux, Jian Yang 0009, Limin Luo 0001, Shuo Li 0001 |
Artif. Intell. Medicine | 3 |
| 2020 | Vessel Structure Extraction using Constrained Minimal Path Propagation
Guanyu Yang 0001, Tianling Lv, Yunpeng Shen, Shuo Li 0001, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux |
Artif. Intell. Medicine | 6 |
| 2020 | Automatic Seizure Detection using Fully Convolutional Nested LSTMabstractThe automatic seizure detection system can effectively help doctors to monitor and diagnose epilepsy thus reducing their workload. Many outstanding studies have given good results in the two-class seizure detection problems, but most of them are based on hand-wrought feature extraction. This study proposes an end-to-end automatic seizure detection system based on deep learning, which does not require heavy preprocessing on the EEG data or feature engineering. The fully convolutional network with three convolution blocks is first used to learn the expressive seizure characteristics from EEG data. Then these robust EEG features pertinent to seizures are presented as an input to the Nested Long Short-Term Memory (NLSTM) model to explore the inherent temporal dependencies in EEG signals. Lastly, the high-level features obtained from the NLSTM model are fed into the softmax layer to output predicted labels. The proposed method yields an accuracy range of 98.44-100% in 10 different experiments based on the Bonn University database. A larger EEG database is then used to evaluate the performance of the proposed method in real-life situations. The average sensitivity of 97.47%, specificity of 96.17%, and false detection rate of 0.487 per hour are yielded. For CHB-MIT Scalp EEG database, the proposed model also achieves a segment-level sensitivity of 94.07% with a false detection rate of 0.66 per hour. The excellent results obtained on three different EEG databases demonstrate that the proposed method has good robustness and generalization power under ideal and real-life conditions. Zuyi Yu, Yang Chen 0008, X. Allen Li |
Int. J. Neural Syst. | 3 |
| 2020 | Compressed sensing MR image reconstruction via a deep frequency-division network
Jiulou Zhang, Yunbo Gu, Youyong Kong, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux |
Neurocomputing | 6 |
| 2020 | Dense biased networks with deep priori anatomy and hard region adaptation: Semi-supervised learning for fine renal artery segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2020 | High-dimensional embedding network derived prior for compressive sensing MRI reconstruction
Jinjie Zhou, Yanjie Zhu, Shanshan Wang 0002, Dong Liang 0001, Yang Chen 0008, Qiegen Liu |
Medical Image Anal. | 7 |
| 2020 | Discriminative feature representation for Noisy image quality assessment
Yunbo Gu, Tianling Lv, Yang Chen 0008, Lu Zhang 0037, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux |
Multim. Tools Appl. | 4 |
| 2020 | K-Net: Integrate Left Ventricle Segmentation and Direct Quantification of Paired Echo SequenceabstractThe integration of segmentation and direct quantification on the left ventricle (LV) from the paired apical views(i.e., apical 4-chamber and 2-chamber together) echo sequence clinically achieves the comprehensive cardiac assessment: multiview segmentation for anatomical morphology, and multidimensional quantification for contractile function. Direct quantification of LV, i.e., to automatically quantify multiple LV indices directly from the image via task-aware feature representation and regression, avoids accumulative error from the inter-step target. This integration sequentially makes a stereoscopical reflection of cardiac activity jointly from the paired orthogonal cross views sequences, overcoming limited observation with a single plane. We propose a K-shaped Unified Network (K-Net), the first end-to-end framework to simultaneously segment LV from apical 4-chamber and 2-chamber views, and directly quantify LV from major- and minor-axis dimensions (1D), area (2D), and volume (3D), in sequence. It works via four components: 1) the K-Net architecture with the Attention Junction enables heterogeneous tasks learning of segmentation task of pixel-wise classification, and direct quantification task of image-wise regression, by interactively introducing the information from segmentation to jointly promote spatial attention map to guide quantification focusing on LV-related region, and transferring quantification feedback to make global constraint on segmentation; 2) the Bi-ResLSTMs distributed in K-Net layer-by-layer hierarchically extract spatial-temporal information in echo sequence, with bidirectional recurrent and short-cut connection to model spatial-temporal information among all frames; 3) the Information Valve tailing the Bi-ResLSTMs selectively exchanges information among multiple views, by stimulating complementary information and suppressing redundant information to make the efficient cross-flow for each view; 4) the Evolution Loss comprehensively guides sequential data learning, with static constraint for frame values, and dynamic constraint for inter-frame value changes. The experiments show that our K-Net gains high performance with a Dice coefficient up to 91.44% and a mean absolute error of the major-axis dimension down to 2.74mm, which reveal its clinical potential. Rongjun Ge, Guanyu Yang 0001, Yang Chen 0008, Limin Luo 0001, Junyi Ren, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2019 | An Efficient 3D-NAS Method for Video-Based Gesture Recognition
Ziheng Guo, Yang Chen 0008 |
ICANN (3) | 2 |
| 2019 | Learning to Hash for Efficient Search Over Incomplete Knowledge GraphsabstractKnowledge graph (KG) embedding techniques represent entities and relations as low-dimensional, continuous vectors, and thus enables machine learning models to be easily adapted to KG completion and querying tasks. However, learned dense vectors are inefficient for large-scale similarity computations. Learning-to-hash is to learn compact binary codes from high-dimensional input data and provides a promising way to accelerate efficiency by measuring Hamming distance instead of Euclidean distance or dot-product. Unfortunately, most of learning-to-hash methods cannot be directly applied to KG structure encoding. In this paper, we introduce a novel framework for encoding incomplete KGs and graph queries in Hamming space. To preserve KG structure information from embeddings to hash codes and address the ill-posed gradient issue in optimization, we utilize a continuation method with convergence guarantees to jointly encode queries and KG entities with geometric operations. The hashed embedding of a query can be utilized to discover target answers from incomplete KGs whilst the efficiency has been greatly improved.We compared our model with state-of-the-art methods on real-world KGs. Experimental results show that our framework not only significantly speeds up the searching process, but also provides good results for unanswerable queries caused by incomplete information. Meng Wang 0009, Haomin Shen, Sen Wang 0001, Lina Yao 0001, Yinlin Jiang, Guilin Qi, Yang Chen 0008 |
ICDM | 7 |
| 2019 | Stereo-Correlation and Noise-Distribution Aware ResVoxGAN for Dense Slices Reconstruction and Noise Reduction in Thick Low-Dose CT
Rongjun Ge, Guanyu Yang 0001, Chenchu Xu, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001 |
MICCAI (6) | 4 |
| 2019 | DPA-DenseBiasNet: Semi-supervised 3D Fine Renal Artery Segmentation with Dense Biased Network and Deep Priori Anatomy
Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
MICCAI (6) | 3 |
| 2019 | PV-LVNet: Direct left ventricle multitype indices estimation from 2D echocardiograms of paired apical views with deep neural networks
Rongjun Ge, Guanyu Yang 0001, Yang Chen 0008, Limin Luo 0001, Heye Zhang, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2019 | Denoising of 3D magnetic resonance images using a residual encoder-decoder Wasserstein generative adversarial network
Maosong Ran, Jinrong Hu, Yang Chen 0008, Hu Chen 0002, Huaiqiang Sun, Jiliu Zhou, Yi Zhang 0018 |
Medical Image Anal. | 3 |
| 2019 | Vessel segmentation using centerline constrained level set method
Tianling Lv, Guanyu Yang 0001, Yudong Zhang 0001, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Domain Progressive 3D Residual Convolution Network to Improve Low-Dose CT ImagingabstractThe wide applications of X-ray computed tomography (CT) bring low-dose CT (LDCT) into a clinical prerequisite, but reducing the radiation exposure in CT often leads to significantly increased noise and artifacts, which might lower the judgment accuracy of radiologists. In this paper, we put forward a domain progressive 3D residual convolution network (DP-ResNet) for the LDCT imaging procedure that contains three stages: sinogram domain network (SD-net), filtered back projection (FBP), and image domain network (ID-net). Though both are based on the residual network structure, the SD-net and ID-net provide complementary effect on improving the final LDCT quality. The experimental results with both simulated and real projection data show that this domain progressive deep-learning network achieves significantly improved performance by combing the network processing in the two domains. Xiangrui Yin, Jean-Louis Coatrieux, Qianlong Zhao, Jin Liu 0019, Wei Yang 0006, Jian Yang 0009, Guotao Quan, Yang Chen 0008, Huazhong Shu, Limin Luo 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2018 | PCANet: An energy perspective
Jiasong Wu, Shijie Qiu, Youyong Kong, Longyu Jiang, Yang Chen 0008, Wankou Yang, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 5 |
| 2018 | Structure-Adaptive Fuzzy Estimation for Random-Valued Impulse Noise SuppressionabstractNoise detection accuracy is crucial in suppressing random-valued impulse noise. Both false and miss detections determine the final estimation performance. Deterministic detection methods, which distinctly classify pixels into noisy or uncorrupted pixels, tend to increase the estimation error because some uncorrupted edge points are hard to discriminate from the random-valued impulse noise points. This paper proposes an iterative structure-adaptive fuzzy estimation (SAFE) for random-valued impulse noise suppression. This SAFE method is developed in the framework of Gaussian maximum likelihood estimation. The structure-adaptive fuzziness is reflected by two structure-adaptive metrics based on pixel reliability and patch similarity, respectively. The reliability metric for each pixel (as noise free) is estimated via a novel-minimal-path-based structure propagation to give full consideration of the spatially varying image structures. A robust iteration stopping strategy is also proposed by evaluating the reestimation error of the uncorrupted intensity information. The comparative experimental results show that the proposed structure-adaptive fuzziness can lead to effective restoration. An efficient implementation of this SAFE method is also realized via graphics-processing-unit-based parallelization. Yang Chen 0008, Yudong Zhang 0001, Huazhong Shu, Jian Yang 0009, Limin Luo 0001, Jean-Louis Coatrieux, Qianjin Feng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | 3D Feature Constrained Reconstruction for Low-Dose CT ImagingabstractLow-dose computed tomography (LDCT) images are often highly degraded by amplified mottle noise and streak artifacts. Maintaining image quality under low-dose scan protocols is a well-known challenge. Recently, sparse representation-based techniques have been shown to be efficient in improving such CT images. In this paper, we propose a 3D feature constrained reconstruction (3D-FCR) algorithm for LDCT image reconstruction. The feature information used in the 3D-FCR algorithm relies on a 3D feature dictionary constructed from available high quality standard-dose CT sample. The CT voxels and the sparse coefficients are sequentially updated using an alternating minimization scheme. The performance of the 3D-FCR algorithm was assessed through experiments conducted on phantom simulation data and clinical data. A comparison with previously reported solutions was also performed. Qualitative and quantitative results show that the proposed method can lead to a promising improvement of LDCT image quality. Jin Liu 0019, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001, Qianjing Feng, Zhiguo Gui, Gouenou Coatrieux |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Predicting CT Image From MRI Data Through Feature Matching With Learned Nonlinear Local DescriptorsabstractAttenuation correction for positron-emission tomography (PET)/magnetic resonance (MR) hybrid imaging systems and dose planning for MR-based radiation therapy remain challenging due to insufficient high-energy photon attenuation information. We present a novel approach that uses the learned nonlinear local descriptors and feature matching to predict pseudo computed tomography (pCT) images from T1-weighted and T2-weighted magnetic resonance imaging (MRI) data. The nonlinear local descriptors are obtained by projecting the linear descriptors into the nonlinear high-dimensional space using an explicit feature map and low-rank approximation with supervised manifold regularization. The nearest neighbors of each local descriptor in the input MR images are searched in a constrained spatial range of the MR images among the training dataset. Then the pCT patches are estimated through k-nearest neighbor regression. The proposed method for pCT prediction is quantitatively analyzed on a dataset consisting of paired brain MRI and CT images from 13 subjects. Our method generates pCT images with a mean absolute error (MAE) of 75.25 ± 18.05 Hounsfield units, a peak signal-to-noise ratio of 30.87 ± 1.15 dB, a relative MAE of 1.56 ± 0.5% in PET attenuation correction, and a dose relative structure volume difference of 0.055 ± 0.107% in , as compared with true CT. The experimental results also show that our method outperforms four state-of-the-art methods. Wei Yang 0006, Liming Zhong, Yang Chen 0008, Liyan Lin, Zhentai Lu, Shupeng Liu, Qianjin Feng 0003, Wufan Chen |
IEEE Trans. Medical Imaging | 3 |
| 2017 | MomentsNet: A simple learning-free method for binary image recognitionabstractIn this paper, we propose a new simple and learning-free deep learning network named MomentsNet, whose convolution layer, nonlinear processing layer and pooling layer are constructed by Moments kernels, binary hashing and block-wise histogram, respectively. Twelve typical moments (including geometrical moment, Zernike moment, Tchebichef moment, etc.) are used to construct the MomentsNet whose recognition performance for binary image is studied. The results reveal that MomentsNet has better recognition performance than its corresponding moments in almost all cases and ZernikeNet achieves the best recognition performance among MomentsNet constructed by twelve moments. ZernikeNet also shows better recognition performance on a binary image database than that of PCANet, which is a learning-based deep learning network. Jiasong Wu, Shijie Qiu, Youyong Kong, Yang Chen 0008, Lotfi Senhadji, Huazhong Shu |
ICIP | 4 |
| 2017 | Sparse-view X-ray CT reconstruction with Gamma regularization
Jian Yang 0009, Yang Chen 0008, Jean-Louis Coatrieux, Limin Luo 0001 |
Neurocomputing | 4 |
| 2017 | Low-Dose CT With a Residual Encoder-Decoder Convolutional Neural NetworkabstractGiven the potential risk of X-ray radiation to the patient, low-dose CT has attracted a considerable interest in the medical imaging field. Currently, the main stream low-dose CT methods include vendor-specific sinogram domain filtration and iterative reconstruction algorithms, but they need to access raw data, whose formats are not transparent to most users. Due to the difficulty of modeling the statistical characteristics in the image domain, the existing methods for directly processing reconstructed images cannot eliminate image noise very well while keeping structural details. Inspired by the idea of deep learning, here we combine the autoencoder, deconvolution network, and shortcut connections into the residual encoder-decoder convolutional neural network (RED-CNN) for low-dose CT imaging. After patch-based training, the proposed RED-CNN achieves a competitive performance relative to the-state-of-art methods in both simulated and clinical cases. Especially, our method has been favorably evaluated in terms of noise suppression, structural preservation, and lesion detection. Hu Chen 0002, Yi Zhang 0018, Mannudeep K. Kalra, Feng Lin 0010, Yang Chen 0008, Peixi Liao, Jiliu Zhou, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Discriminative Feature Representation to Improve Projection Data Inconsistency for Low Dose CT ImagingabstractIn low dose computed tomography (LDCT) imaging, the data inconsistency of measured noisy projections can significantly deteriorate reconstruction images. To deal with this problem, we propose here a new sinogram restoration approach, the sinogram- discriminative feature representation (S-DFR) method. Different from other sinogram restoration methods, the proposed method works through a 3-D representation-based feature decomposition of the projected attenuation component and the noise component using a well-designed composite dictionary containing atoms with discriminative features. This method can be easily implemented with good robustness in parameter setting. Its comparison to other competing methods through experiments on simulated and real data demonstrated that the S-DFR method offers a sound alternative in LDCT. Jin Liu 0019, Jianhua Ma 0001, Yi Zhang 0018, Yang Chen 0008, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux, Wei Yang 0006, Qianjin Feng 0004, Wufan Chen |
IEEE Trans. Medical Imaging | 4 |
| 2016 | Retinal vessel enhancement using multi-dictionary and sparse codingabstractA novel retinal vessel enhancement method based on multi-dictionary and sparse coding is proposed in this paper. Two dictionaries are utilized to gain the retinal vascular structures and details, one is the representation dictionary (RD) generated from the original retinal images, and another is the enhancement dictionary (ED) extracted from the corresponding label images. The proposed method represents the input image with RD to get the sparse coefficients via a sparse coding process. Then the enhanced retinal vessel image is obtained from the solved sparse coefficients and ED. Experimental results performed on the DRIVE and STARE databases indicate that the proposed method not only can effectively improve the image contrast but also enhance the details of the retinal vessels. Yang Chen 0008, Zhuhong Shao, Limin Luo 0001 |
ICASSP | 2 |
| 2016 | Blood vessel enhancement via multi-dictionary and sparse coding: Application to retinal vessel enhancing
Yang Chen 0008, Zhuhong Shao, Limin Luo 0001 |
Neurocomputing | 2 |
| 2016 | Color image classification via quaternion principal component analysis network
Jiasong Wu, Zhuhong Shao, Yang Chen 0008, Beijing Chen, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 4 |
| 2016 | Curve-Like Structure Extraction Using Minimal Path Propagation With BacktrackingabstractMinimal path techniques can efficiently extract geometrically curve-like structures by finding the path with minimal accumulated cost between two given endpoints. Though having found wide practical applications (e.g., line identification, crack detection, and vascular centerline extraction), minimal path techniques suffer from some notable problems. The first one is that they require setting two endpoints for each line to be extracted (endpoint problem). The second one is that the connection might fail when the geodesic distance between the two points is much shorter than the desirable minimal path (shortcut problem). In addition, when connecting two distant points, the minimal path connection might become inefficient as the accumulated cost increases over the propagation and results in leakage into some non-feature regions near the starting point (accumulation problem). To address these problems, this paper proposes an approach termed minimal path propagation with backtracking. We found that the information in the process of backtracking from reached points can be well utilized to overcome the above problems and improve the extraction performance. The whole algorithm is robust to parameter setting and allows a coarse setting of the starting point. Extensive experiments with both simulated and realistic data are performed to validate the performance of the proposed method. Yang Chen 0008, Yudong Zhang 0001, Jian Yang 0009, Guanyu Yang 0001, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux, Qianjing Feng |
IEEE Trans. Image Process. | 1 |
| 2015 | Segmentation of liver tumor via nonlocal active contoursabstractTo reduce the manual labor time and provide the accuracy of liver tumor segmentation in the treatment planning of radiofrequency ablation (RFA), a novel method for liver tumor image segmentation by nonlocal active contours is proposed in this paper. A multi Gabor feature map of the liver tumor image is computed to describe the homogeneity of patches in a nonlocal way, and the nonlocal comparisons between pairs of patches are used to calculate the active contour energy. The whole energy function is minimized via a level set method to give the final segmentation. The experimental results indicate that the proposed method leads to good liver tumor segmentation with a good robustness to initialization condition. Experiment results show the proposed method can provide segmentation close to manual results, with the mean overlap error (OE) less than 23.86%. Yang Chen 0008, Guanyu Yang 0001, Jingyu Meng, Limin Luo 0001 |
ICIP | 2 |
| 2014 | Artifact Suppressed Dictionary Learning for Low-Dose CT Image ProcessingabstractLow-dose computed tomography (LDCT) images are often severely degraded by amplified mottle noise and streak artifacts. These artifacts are often hard to suppress without introducing tissue blurring effects. In this paper, we propose to process LDCT images using a novel image-domain algorithm called "artifact suppressed dictionary learning (ASDL)." In this ASDL method, orientation and scale information on artifacts is exploited to train artifact atoms, which are then combined with tissue feature atoms to build three discriminative dictionaries. The streak artifacts are cancelled via a discriminative sparse representation operation based on these dictionaries. Then, a general dictionary learning processing is applied to further reduce the noise and residual artifacts. Qualitative and quantitative evaluations on a large set of abdominal and mediastinum CT images are carried out and the results show that the proposed method can be efficiently applied in most current CT systems. Yang Chen 0008, Luyao Shi, Qianjing Feng, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux, Wufan Chen |
IEEE Trans. Medical Imaging | 1 |
| 2007 | A Novel Way of Incorporating Large-Scale Knowledge into MRF Prior Model
Yang Chen 0008, Wufan Chen, Yanqiu Feng, Qingqi Wang |
AIME | 1 |