EDBT 2026 Demo / reviewers in the wild / expert
Shuo Li 0001
dblp:49/595-1
· DBLP profile ↗
281ranked-venue papers
8as first author
149since 2021 · last 2026
0000-0002-5184-3230ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 190 · 4 first-author · 101 since 2021Graphics, computer vision, multimedia, augmented reality and games · 106 · 4 first-author · 49 since 2021Artificial intelligence and machine learning · 70 · 2 first-author · 36 since 2021Computer networks · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ambiguity-aware Truncated Flow Matching for Ambiguous Medical Image SegmentationabstractA simultaneous enhancement of accuracy and diversity of predictions remains a challenge in ambiguous medical image segmentation (AMIS) due to the inherent trade-offs. While truncated diffusion probabilistic models (TDPMs) hold strong potential with a paradigm optimization, existing TDPMs suffer from entangled accuracy and diversity of predictions with insufficient fidelity and plausibility. To address the aforementioned challenges, we propose Ambiguity-aware Truncated Flow Matching (ATFM), which introduces a novel inference paradigm and dedicated model components. Firstly, we propose Data-Hierarchical Inference, a redefinition of AMIS-specific inference paradigm, which enhances accuracy and diversity at data-distribution and data-sample level, respectively, for an effective disentanglement. Secondly, Gaussian Truncation Representation (GTR) is introduced to enhance both fidelity of predictions and reliability of truncation distribution, by explicitly modeling it as a Gaussian distribution at Ttrunc instead of using sampling-based approximations. Thirdly, Segmentation Flow Matching (SFM) is proposed to enhance the plausibility of diverse predictions by extending semantic-aware flow transformation in Flow Matching (FM). Comprehensive evaluations on LIDC and ISIC3 datasets demonstrate that ATFM outperforms SOTA methods and simultaneously achieves a more efficient inference. ATFM improves GED and HM-IoU by up to 12% and 7.3% compared to advanced methods. Fanding Li, Xiangyu Li 0004, Xianghe Su, Xingyu Qiu, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Gongning Luo, Shuo Li 0001 |
AAAI | 9 |
| 2026 | DeepBooTS: Dual-Stream Residual Boosting for Drift-Resilient Time-Series ForecastingabstractTime-Series (TS) exhibits pronounced non-stationarity. Consequently, most forecasting methods display compromised robustness to concept drift, despite the prevalent application of instance normalization. We tackle this challenge by first analysing concept drift through a bias-variance lens and proving that weighted ensemble reduces variance without increasing bias. These insights motivate DeepBooTS, a novel end-to-end dual-stream residual-decreasing boosting method that progressively reconstructs the intrinsic signal. In our design, each block of a deep model becomes an ensemble of learners with an auxiliary output branch forming a highway to the final prediction. The block‑wise outputs correct the residuals of previous blocks, leading to a learning‑driven decomposition of both inputs and targets. This method enhances versatility and interpretability while substantially improving robustness to concept drift. Extensive experiments, including those on large-scale datasets, show that the proposed method outperforms existing methods by a large margin, yielding an average performance improvement of 15.8% across various datasets, establishing a new benchmark for TS forecasting. Daojun Liang, Jing Chen 0030, Yinglong Wang 0001, Shuo Li 0001 |
AAAI | 5 |
| 2026 | FedMSFD: Mistake-Aware and Structured Feature Distillation for Non-IID Federated Learning
Miaomiao He, Shuo Li 0001, Xuehua Bi |
ICIC (26) | 3 |
| 2026 | A high-performance lightweight network for real-time view recognition and quality control in transthoracic echocardiography
Peng Hong, Xueying Tan, Xiaoxue Fan, Lisheng Xu, Shuo Li 0001, Dongming Chen |
Expert Syst. Appl. | 6 |
| 2026 | Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
Sheng Lian, Jianlong Cai, Dengfeng Pan, Guang-Yong Chen, Fan Zhang 0045, Jialun Pei, Shuo Li 0001 |
Int. J. Comput. Vis. | 9 |
| 2026 | PLATO: ProbabiListic hierArchical mulTi-head mOdel for plug-and-play ambiguous medical image segmentation
Xiangyu Li 0004, Fanding Li, Yongfeng Yuan, Suyu Dong, Kuanquan Wang, Yi Shen 0001, Guohua Wang 0001, Gongning Luo, Shuo Li 0001 |
Knowl. Based Syst. | 9 |
| 2026 | SCULPT: Semantic-aware causal prompt tuning for out-of-distribution detection of whole slide images
Pengzhong Sun, Xiangyu Li 0004, Dong Liang 0001, Jun Liu 0080, Zhanshi Zhu, Xiaokun Li, Suyu Dong, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
Knowl. Based Syst. | 11 |
| 2026 | IUGC: A benchmark of landmark detection in end-to-end intrapartum ultrasound biometry
Jieyun Bai, Yitong Tang, Xiao Liu 0037, Jiale Hu, Yunda Li, Xufan Chen, Yunshu Li, Bowen Guo, Jing Jiao, Lifei Li, Yuzhang Ma, Xiaoxin Han, Haochen Shao, Qingchen Liu, Jingfan Kuang, Shanglin Song, Anirvan Krishna, Zaid Ahmed Khan, Zelan Li, Zhengyang Zhang, Hansen Zhang, Xuezhi Zhang, Lyuyang Tong, Bo Du 0004, Yu Chen 0099, Zilun Peng, Saeid Rezaei, Tom Weidong Cai, Fangyijie Wang, Kathleen M. Curran, Guénolé C. M. Silvestre, Isaac Khobo, Yaosheng Lu, Dong Ni 0001, Mohammad Yaqub, Jun Ma 0016, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 50 |
| 2026 | Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 63 |
| 2026 | VCC-DSA: A novel vascular consistency constrained DSA imaging model for motion artifact suppression
Rongjun Ge, Weilong Mao, Guanyu Yang 0001, Yang Chen 0008, Shuo Li 0001 |
Medical Image Anal. | 12 |
| 2026 | Anatomy-guided prompting with cross-modal self-alignment for whole-body PET-CT breast cancer segmentation
Jiaju Huang, Xinglong Liang, Shaobin Chen, Yue Sun 0001, Greta S. P. Mok, Shuo Li 0001, Tao Tan 0002 |
Medical Image Anal. | 7 |
| 2026 | Openness-aware multi-prototype learning for open set medical diagnosis
Mingyuan Liu 0002, Yuzhuo Gu, Jicong Zhang, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2026 | S2DENet: Shallow suppression and deep enhancement network for general ultrasound image segmentation
Xintao Pang, Jinlin Yang, Zhifan Gao, Chuan Lin 0003, Yue Sun 0001, Shuo Li 0001, Peter H. N. de With, Tao Tan 0002 |
Medical Image Anal. | 6 |
| 2026 | Diversity-driven MG-MAE: Multi-granularity representation learning for non-salient object segmentation
Chengjin Yu, Chenchu Xu, Dongsheng Ruan, Huafeng Liu 0003, Xiaohu Li, Shuo Li 0001 |
Medical Image Anal. | 8 |
| 2026 | Dynamical multi-order responses and global semantic-infused adversarial learning: A robust airway segmentation methodabstractAutomated airway segmentation in computerized tomography (CT) images is crucial for the accurate diagnosis of lung diseases. However, the scarcity of manual annotations hinders the efficacy of supervised learning, while unconstrained intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order responses and Global Semantic-infused Adversarial network (DMGSA), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to empower the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles; (3) we introduce the Adversarial Learning (AL) on the top of MONR module to discern nuances between real and fake images, focusing on capturing the textural features of terminal bronchioles. For the supervised branch, we propose an innovative Generalized Mean pooling based Global Semantic-infused (GMGS) module to ulteriorly improve the robustness. Ultimately, we have verified the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly. Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Yongkai Liu, Giorgos Papanastasiou, Zhifan Gao, Shuo Li 0001, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 8 |
| 2026 | Adaptation Follow Human Attention: Gaze-Assisted Medical Segment Anything ModelabstractSegment Anything Model (SAM) has demonstrated state-of-the-art performance in most segmentation tasks. However, due to insufficient training in the medical domain, SAM’s ability to generalize to medical images is limited. Although preliminary efforts have fine-tuned SAM for the medical domain, the fine-tuned model still struggles with variability in medical tasks. Some recent studies have explored weakly supervised learning to mitigate SAM’s performance degradation in the medical domain. However, the effectiveness of weakly supervised learning is heavily dependent on the quality of weakly supervised information, with performance significantly dropping as the quality declines. Doctors’ attention is closely related to the target area during diagnosis. Integrating gaze information into SAM’s adaptation process for medical image segmentation enhances efficiency and significantly improves performance in medical tasks. In this paper, we first propose a Gaze-assisted medical segment Anything Model (GAM), which utilizes gaze information to enable the adaptation of SAM in medical images following doctor’s attention. It has two innovations: 1) Feature-level adaptation: Gaze Alignment (GA) learning makes the feature-level adaptation follow the doctor’s attention which mines the human guidance from gaze heatmaps and guides model to extract general features for downstream tasks. 2) Output-level adaptation: Gaze-Balance (GB) learning makes the output-level adaptation follow the doctor’s attention which utilizes gaze heatmaps to enhance the human-focused area and solve the problem of over/under segmentation from the output-level. Our promising results on 7 tasks with 12 targets have demonstrated the powerful adaptation ability of our GAM in the medical domain. Our GAM demonstrates significant potential for low-cost clinical assistance in medical diagnosis, enabling SAM to adapt to the medical image domain without disrupting clinical workflows. We have released the full source code on https://github.com/Ruiz1026/GAM. Rongjun Ge, Ruiyi Li, Chong Wang 0011, Jean-Louis Coatrieux, Daoqiang Zhang, Yang Chen 0008, Shuo Li 0001, Yuting He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2026 | Source-Resilient Joint Learning Framework for Preserving Stable Generalization on Diverse Ultrasonic Source ScenariosabstractJoint learning on diverse ultrasonic source scenarios presents a challenge in preserving stable gen-eralization due to the combination of heterogeneity of different sources and the inconsistency of joint learning features. Previous joint learning studies, which are not source-resilient frameworks, may not preserve stable generalization when trained on diverse source scenarios. Furthermore, the limited variations insingle-source data and the interference from ultrasound imaging, which are common in ultrasonic source scenarios, further decrease generalization. To address these problems, we pro posed a source-resilient joint learning framework consisting of three stages: 1) Source transforming, where our 1-to-N transformation unifies diverse source scenarios for source-resiliency. 2) Our feature enhancement modules model the source-resilient joint learning network, including a manifold-constraint normalization module (MCNM) for addressing heterogeneity by minimizing manifold-based loss, a task-consistent attention module (TCAM) shares the multi-scale features with self-attention to address inconsistency, and an adaptive feature-shifting module (AFSM) for feature-level augmentation to overcome single-source data.3) Our ultrasound-hybrid linear mapping (USmapping) cascades speckle randomization and mask-guiding Monge-Kantorovitch linear mapping to achieve ultrasonic style randomization for addressing the interference of ultrasonic data. Our framework was evaluated on eight ultrasound datasets from various scanners at multiple center sand surpassed previous comparable studies in both segmentation (DSCWAvgof 75.7%) and classification (AUROCWAvgof 68.8%) tasks. Our framework has the potential to serve as a general framework for enhancing the performance of joint learning under diverse ultrasonic source scenarios. Bin Huang 0021, Zhong Liu 0004, Ziyue Xu 0001, S. C. Chan 0001, Huiying Wen, Qicai Huang, Meiqin Jiang, Changfeng Dong, Ruhai Zou, Bingsheng Huang, Xin Chen 0025, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 14 |
| 2026 | CalDiff: Calibrating Uncertainty and Accessing Reliability of Diffusion Models for Trustworthy Lesion SegmentationabstractLow reliability has consistently been a challenge in the application of deep learning models for high risk decision-making scenarios. In medical image segmentation, for instance, multiple expert annotations can be consulted to reduce subjective bias and reach a consensus, thereby enhancing the segmentation accuracy and reliability. To develop a reliable lesion segmentation model, we leverage the uncertainty introduced by multiple annotations, enabling the model to better capture real-world diagnostic variability and provide more informative predictions. Since a reliable model should produce calibrated uncertainty estimates that align with actual predictive performance, we propose CalDiff, a novel framework designed to calibrate model uncertainty in lesion segmentation and mitigate the risk of overconfident yet incorrect predictions. To harness the superior generative ability of diffusion mod els, a dual step-wise and sequence-aware calibration mechanism is proposed on the basis of the sequential nature of diffusion models. We evaluate the calibrated model through a comprehensive quantitative and visual analysis, thus ad dressing the previously overlooked challenge of assessing uncertainty calibration and model reliability in scenarios with multiple annotations and multiple predictions. Experimental results on two multi-annotated lesion segmentation datasets demonstrate that CalDiff produces uncertainty maps that can reflect informative low confidence areas, which can further indicate the false predictions potentially made by the model. By calibrating the uncertainty in the training phase, the uncertain areas produced from our model are more closely correlated with areas where the model has made errors in the inference. In summary, the uncertainty captured by our CalDiff framework can serve as apowerful indicator, which can help mitigate the risks of adopting model's outputs, allowing clinicians to prioritize reviewing areas or slices with higher uncertainty and enhancing the model's reliability and trustworthiness in real clinical practice. Mingrui Yang, Sercan Tosun, Kunio Nakamura, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Radiomics-Driven Diffusion Model and Monte Carlo Compression Sampling for Reliable Medical Image SynthesisabstractReliable medical image synthesis is crucial for clinical applications and downstream tasks, where high-quality anatomical structure and predictive confidence are essential. Existing studies have made significant progress by embedding prior conditional knowledge, such as conditional images or textual information, to synthesize natural images. However, medical image synthesis remains a challenging task due to: 1) Data scarcity: High-quality medical text prompt are extremely rare and require specialized expertise. 2) Insufficient uncertainty estimation: The uncertainty estimation is critical for evaluating the confidence of reliable medical image synthesis. This paper presents a novel approach for medical image synthesis, driven by radiomics prompts and combined with Monte Carlo Compression Sampling (MCCS) to ensure reliability. For the first time, our method leverages clinically focused radiomics prompts to condition the generation process, guiding the model to produce reliable medical images. Furthermore, the innovative MCCS algorithm employs Monte Carlo methods to randomly select and compress sampling steps within the denoising diffusion implicit models (DDIM), enabling efficient uncertainty quantification. Additionally, we introduce a MambaTrans architecture to model long-range dependencies in medical images and embed prior conditions (e.g., radiomics prompts). Extensive experiments on benchmark medical imaging datasets demonstrate that our approach significantly improves image quality and reliability, outperforming SoTA methods in both qualitative and quantitative evaluations. Jianfeng Zhao 0004, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Causality-Adjusted Data Augmentation for Domain Continual Medical Image SegmentationabstractIn domain continual medical image segmentation, distillation-based methods mitigate catastrophic forgetting by continuously reviewing old knowledge. However, these approaches often exhibit biases towards both new and old knowledge simultaneously due to confounding factors, which can undermine segmentation performance. To address these biases, we propose the Causality-Adjusted Data Augmentation (CauAug) framework, introducing a novel causal intervention strategy called the Texture-Domain Adjustment Hybrid-Scheme (TDAHS) alongside two causality-targeted data augmentation approaches: the Cross Kernel Network (CKNet) and the Fourier Transformer Generator (FTGen). (1) TDAHS establishes a domain-continual causal model that accounts for two types of knowledge biases by identifying irrelevant local textures (L) and domain-specific features (D) as confounders. It introduces a hybrid causal intervention that combines traditional confounder elimination with a proposed replacement approach to better adapt to domain shifts, thereby promoting causal segmentation. (2) CKNet eliminates confounder L to reduce biases in new knowledge absorption. It decreases reliance on local textures in input images, forcing the model to focus on relevant anatomical structures and thus improving generalization. (3) FTGen causally intervenes on confounder D by selectively replacing it to alleviate biases that impact old knowledge retention. It restores domain-specific features in images, aiding in the comprehensive distillation of old knowledge. Our experiments show that CauAug significantly mitigates catastrophic forgetting and surpasses existing methods in various medical image segmentation tasks. Zhanshi Zhu, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Guohua Wang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2026 | TKRL: Targeted Knowledge Rectification Learning Against Teacher-Originated Defects in Domain Continual SegmentationabstractKnowledge distillation can mitigate catastrophic forgetting in domain continual segmentation by transferring knowledge from the older model to the newer model. However, existing distillation-based methods primarily emphasize knowledge retention while overlooking inherent defects in the older teacher models. As a result, these teacher-originated defects, such as knowledge gaps or biases, are propagated and exacerbate forgetting. To address this challenge, we propose a Targeted Knowledge Rectification Learning framework (TKRL) to probe and correct teacher-originated defects. TKRL consists of two modules: 1) Probe-augmented Class Distillation, which generates gradient-driven "probes" to uncover underrepresented features in the older model, thereby bridging knowledge gaps by distilling hidden information into the new model; 2) Variance-guided Masked Autoencoder, which selectively masks and reconstructs critical high-uncertainty patches across multi-level semantic regions, thereby correcting biases inherited from the older model. Our experimental results show that TKRL effectively rectifies knowledge gaps and biases, thereby mitigating catastrophic forgetting and enhancing performance in domain continual segmentation. Zhanshi Zhu, Wenjian Gu, Xiangyu Li 0004, Qince Li, Yongfeng Yuan, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | A Trusted Lesion-assessment Network for Interpretable Diagnosis of Coronary Artery Disease in Coronary CT AngiographyabstractCoronary Artery Disease (CAD) poses a significant threat to cardiovascular patients worldwide, underscoring the critical importance of automated CAD diagnostic technologies in clinical practice. Previous technologies for lesion assessment in Coronary CT Angiography (CCTA) images have been insufficient in terms of interpretability, resulting in solutions that lack clinical reliability in both network architecture and prediction outcomes, even when diagnoses are accurate. To address the limitation of interpretability, we introduce the Trusted Lesion-Assessment Network (TLA-Net), which provides a clinically reliable solution for multi-view CAD diagnosis: (1) The causality-informed evidence collection constructs a causal graph for the diagnostic process and implements causal interventions, preventing confounders' interference and enhancing the transparency of the network architecture. (2) The clinically-aligned uncertainty integration hierarchically combines Dirichlet distributions from various views based on clinical priors, offering confidence coefficients for prediction outcomes that align with physicians' image analysis procedures. Experimental results on a dataset of 2,618 lesions demonstrate that TLA-Net, supported by its interpretable methodological design, exhibits superior performance with outstanding generalization, domain adaptability, and robustness. Xinghua Ma, Xinyan Fang, Mingye Zou, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Zhaowen Qiu, Xin Gao 0001, Shuo Li 0001 |
AAAI | 9 |
| 2025 | DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality pseudo labels. Vision-Language Model (VLM) has great potential to enhance pseudo labels by introducing text prompt guided multimodal supervision information. It nevertheless faces the cross-modal problem: the obtained messages tend to correspond to multiple targets. To address aforementioned problems, we propose a Dual Semantic Similarity-Supervised VLM (DuSSS) for SSMIS. Specifically, 1) a Dual Contrastive Learning (DCL) is designed to improve cross-modal semantic consistency by capturing intrinsic representations within each modality and semantic correlations across modalities. 2) To encourage the learning of multiple semantic correspondences, a Semantic Similarity-Supervision strategy (SSS) is proposed and injected into each contrastive learning process in DCL, supervising semantic similarity via the distribution-based uncertainty levels. Furthermore, a novel VLM-based SSMIS network is designed to compensate for the quality deficiencies of pseudo-labels. It utilizes the pretrained VLM to generate text prompt guided supervision information, refining the pseudo label for better consistency regularization. Experimental results demonstrate that our DuSSS achieves outstanding performance with Dice of 82.52%, 74.61% and 78.03% on three public datasets (QaTa-COV19, BM-Seg and MoNuSeg). Qingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji 0001, Shuo Li 0001 |
AAAI | 5 |
| 2025 | CLINav-GKD: Vision-Language Latent Hyperbolic Geometric Knowledge Distillation for Real-World 6-DOF Echocardiography Probe NavigationabstractVision-Language Models (VLMs) have great potential for advancing echocardiography (Echo) probe navigation, which is crucial for assisting sonographers in standardized view acquisition. However, the high clinical deployment costs and spurious correlations pose major challenges for VLM-based probe navigation. To address these challenges, we propose Contrastive Language-Image Navigator for Latent Hyperbolic Geometric Knowledge Distillation (CLINav-GKD), a novel VLM-based 6-DOF Echo probe navigation framework. Specifically, Contrastive Language-Image Navigator (CLN) proposes a lightweight VLM-based 6-DOF navigator, reducing deployment costs while improving sensitivity to quality variations. Latent Hy perbolic Geometric Distiller (LGD) models the global geometric-topology between samples, mitigating spurious correlations and enhancing robustness. We train CLINav-GKD on real-world data with probe motion trajectories. Experimental results show that CLINav-GKD outperforms other VLM-based distillation methods by 2.8%, 3.8%, and 3.6% in probe navigation, achieving a superior balance of accuracy, robustness, and deployability for real-world clinical use. Code and data are available at https://github.com/DaisyLi0516/CLINav-GKD. Yixuan Fan, Xiaoxiao Cui, Yuezhong Zhang, Jiaguang Song, Xifeng Hu, Kai Zheng 0001, Li-Zhen Cui 0001, Zhi Liu 0004, Shuo Li 0001 |
BIBM | 11 |
| 2025 | Finding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold NetworkabstractIn image generation, Schrödinger Bridge (SB)-based methods theoretically enhance the efficiency and quality compared to the diffusion models by finding the least costly path between two distributions. However, they are computationally expensive and time-consuming when applied to complex image data. The reason is that they focus on fitting globally optimal paths in high-dimensional spaces, directly generating images as next step on the path using complex networks through self-supervised training, which typically results in a gap with the global optimum. Meanwhile, most diffusion models are in the same path subspace generated by weights fA(t) and fB(t), as they follow the paradigm (xt= fA(t)xImg+ fB(t)ϵ). To address the limitations of SB-based methods, this paper proposes for the first time to find local Diffusion Schrödinger Bridges (LDSB) in the diffusion path subspace, which strengthens the connection between the SB problem and diffusion models. Specifically, our method optimizes the diffusion paths using Kolmogorov-Arnold Network (KAN), which has the advantage of resistance to forgetting and continuous output. The experiment shows that our LDSB significantly improves the quality and efficiency of image generation using the same pretrained denoising network and the KAN for optimising is only less than 0.1MB. The FID metric is reduced by more than 15%, especially with a reduction of 48.50% when NFE of DDIM is 5 for the CelebA dataset. Code is available at https://github.com/PerceptionComputingLab/LDSB. Xingyu Qiu, Mengying Yang, Xinghua Ma, Fanding Li, Dong Liang 0001, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
CVPR | 9 |
| 2025 | Gaze-Assisted Human-Centric Domain Adaptation for Cardiac Ultrasound Image SegmentationabstractDomain adaptation (DA) for cardiac ultrasound image segmentation is clinically significant and valuable. However, previous domain adaptation methods are prone to be affected by the incomplete pseudo label and low-quality target to source images. Human-centric domain adaptation has great advantages of human cognitive guidance to help model adapt to target domain and reduce reliance on labels. Doctor gaze trajectories contains a large amount of cross-domain human guidance. To leverage gaze information and human cognition for guiding domain adaptation, we propose gaze-assisted human-centric domain adaptation (GAHCDA), which reliably guides the domain adaptation of cardiac ultrasound images. GAHCDA includes following modules: (1) Gaze Augment Alignment (GAA): GAA enables the model to obtain human cognition general features to recognize segmentation target in different domain of cardiac ultrasound images like humans. (2) Gaze Balance Loss (GBL): GBL fused gaze heatmap with outputs which makes the segmentation result structurally closer to the target domain. The experimental results show that our proposed framework is able to segment cardiac ultrasound images more effectively in the target domain than GAN-based methods and other self-train based methods and shown great potential in clinical application. Ruiyi Li, Yuting He 0001, Rongjun Ge, Chong Wang 0011, Daoqiang Zhang, Yang Chen 0008, Shuo Li 0001 |
ICASSP | 7 |
| 2025 | Neuromanifold-Regularized KANs for Shape-fair Feature Representations
Mazlum Ferhat Arslan, Shuo Li 0001 |
ICCV | 3 |
| 2025 | Vector Contrastive Learning for Pixel-Wise Pretraining in Medical VisionabstractContrastive learning (CL) has become a cornerstone of self-supervised pretraining (SSP) in foundation models, however, extending CL to pixel-wise representation, crucial for medical vision, remains an open problem. Standard CL formulates SSP as a binary optimization problem (binary CL) where the excessive pursuit of feature dispersion leads to an over-dispersion problem, breaking pixel-wise feature correlation thus disrupting the intra-class distribution. Our vector CL reformulates CL as a vector regression problem, enabling dispersion quantification in pixel-wise pretraining via modeling feature distances in regressing displacement vectors. To implement this novel paradigm, we propose the COntrast in VEctor Regression (COVER) framework. COVER establishes an extendable vector-based self-learning, enforces a consistent optimization flow from vector regression to distance modeling, and leverages a vector pyramid architecture for granularity adaptation, thus preserving pixel-wise feature correlations in SSP. Extensive experiments across 8 tasks, spanning 2 dimensions and 4 modalities, show that COVER significantly improves pixel-wise SSP, advancing generalizable medical visual foundation models. Yuting He 0001, Shuo Li 0001 |
ICCV | 2 |
| 2025 | Anatomy-Based Self-supervised Pre-training for Scale-Robust Hierarchical Representations in Chest X-Rays
Surong Chu, Yan Qiang 0001, Guohua Ji, Xueting Ren, Baoping Jia, Yangyang Wei, Juanjuan Zhao 0002, Shuo Li 0001 |
MICCAI (10) | 9 |
| 2025 | Information Bottleneck-Based Causal Attention for Multi-label Medical Image Recognition
Xiaoxiao Cui, Shanzhi Jiang, Mengli Xue, Wentao Li 0001, Junhong Leng, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001 |
MICCAI (8) | 10 |
| 2025 | E-BayesSAM: Efficient Bayesian Adaptation of SAM with Self-optimizing KAN-Based Interpretation for Uncertainty-Aware Ultrasonic Segmentation
Bin Huang 0021, Zhong Liu 0004, Huiying Wen, Bingsheng Huang, Xin Chen 0025, Shuo Li 0001 |
MICCAI (14) | 6 |
| 2025 | Rethinking Multi-view Mammogram Representation Learning via Counterfactual Reasoning with Kolmogorov-Arnold Theorem
Benzheng Wei, Shuo Li 0001 |
MICCAI (8) | 3 |
| 2025 | Iterative Foundation-Dedicated Learning: Optimized Key Frames, Prompts and Memories for Semi-supervised Segmentation
Ziman Yin, Dong Nie, Shuo Li 0001, JunJun Pan, Zhenyu Tang 0002 |
MICCAI (8) | 3 |
| 2025 | RadKAM: Attention-Driven Kolmogorov-Arnold Model for Automatic Radiation-Induced Lymphopenia Prediction by Multimodal Learning
Rongchang Zhao, Zhangyue Wu, Jian Zhang 0048, Zijian Zhang 0004, Shuo Li 0001 |
MICCAI (15) | 5 |
| 2025 | Causality-Driven Spatio-Temporal Generator for Multi-phase Contrast-Enhanced CT Synthesis
Qikui Zhu, Shuo Li 0001 |
MICCAI (3) | 4 |
| 2025 | Advancing Fine-Grained Spine Segmentation Through Visual-Language Model with Omni- and Pixel-Level Semantic Enhancements
Jianlong Cai, Sheng Lian, Dengfeng Pan, Guang-Yong Chen, Lei Li 0048, Zhiming Luo, Shuo Li 0001 |
PRCV (14) | 7 |
| 2025 | Interactive prototype learning and self-learning for few-shot medical image segmentation
Yuhui Song, Chenchu Xu, Xiuquan Du, Jie Chen 0025, Yanping Zhang 0001, Shuo Li 0001 |
Artif. Intell. Medicine | 7 |
| 2025 | TASL-Net: Tri-attention selective learning network for intelligent diagnosis of bimodal ultrasound video
Chengqian Zhao, Zhao Yao, Zhaoyu Hu, Yuanxin Xie, Yafang Zhang, Yuanyuan Wang 0001, Shuo Li 0001, Jianqiao Zhou, Jinhua Yu 0003 |
Expert Syst. Appl. | 7 |
| 2025 | Cooperative metric learning-based hybrid transformer for automatic recognition of standard echocardiographic multi-views
Yankun Cao, Xiaoxiao Cui, Xifeng Hu, Yuezhong Zhang, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001 |
Future Gener. Comput. Syst. | 9 |
| 2025 | LFVDNet: Low-frequency variable-driven network for medical time series
Dengqun Sun, Lei Li 0048, Xiuquan Du, Shuo Li 0001 |
J. Biomed. Informatics | 6 |
| 2025 | LLM-GAODE: Large-language-model augmented neural ordinary differential equation network for video nystagmography classificationabstractBenign paroxysmal positional vertigo (BPPV), a common type of vertigo with complex etiologies, is traditionally diagnosed using video nystagmography (VNG). Current automated methods lack diagnostic precision owing to subjective interpretation of eye movement characteristics. To address these challenges, we introduce a l arge l anguage m odel-augmented G ram-based a ttentive neural o rdinary d ifferential e quation ( LLM-GAODE ), an innovative and data-driven framework integrating eye-tracking technology with a Gram-based attention mechanism and a neural ordinary differential equation network to improve BPPV classification. Furthermore, when the neural network exhibits low confidence in its predictions, an LLM can supplement the process with advanced reasoning in natural language. LLM-GAODE was evaluated using an extensive VNG dataset provided by a collaborative university hospital. Results suggest that LLM-GAODE significantly outperforms existing benchmarks in trajectory classification for BPPV diagnosis. The framework enhances BPPV diagnostic accuracy and achieves state-of-the-art performance in open-source trajectory classification benchmarks. The code is available at https://github.com/XiheQiu/LLM-GAODE . Xihe Qiu, Shaojie Shi, Bin Li 0091, Xiaoyu Tan, Yongbin Gao, Shuo Li 0001 |
Knowl. Based Syst. | 6 |
| 2025 | Uncertainty-guided and cross-modality attention network for liver tumor segmentation and quantification via integrating dynamic MRIabstractSegmentation and quantitative measurement of liver tumors, including hemangiomas and hepatocellular carcinoma (HCC), using dynamic Magnetic Resonance Imaging (MRI) sequences are crucial for effective treatment and prognosis. However, these tasks remain challenging due to two key issues: (1) the severe class imbalance between tumors and background, particularly for small HCC lesions, which complicates precise feature extraction; and (2) the diverse imaging features across dynamic MRI phases, making effective fusion of multi-phase information difficult. To address these challenges, this study proposes the Uncertainty-guided and Cross-modality Attention Network (UgCmA-Net). UgCmA-Net incorporates three innovative components: (1) a cross-modality attention pyramid module within a parallel attention-based encoder, enhancing tumor-specific feature extraction across dynamic phases; (2) a fusion Transformer (F-Trans), where the non-local Transformer captures long-range dependencies, and the phase-aware Transformer fuses multi-phase dynamic MRI features; and (3) an uncertainty-guided auxiliary-primary segmentor, which improves edge confidence and segmentation accuracy through uncertainty estimation. The UgCmA-Net was validated using dynamic MRI sequences (T1 pre-contrast MRI, arterial-phase, portal venous phase, and delay-phase contrast-enhanced MRI) from 265 clinical subjects. Experimental results show that the proposed UgCmA-Net achieves state-of-the-art performance, with a dice similarity coefficient of 85.44%, Hausdorff Distance of 2.28 mm, and mean absolute error values of 1.85 mm, 1.90 mm, 6.52 mm, and 97.27 mm 2 for multi-index quantification of center point, max-diameter, circumference, and area, respectively. Statistical analysis confirms that the improvements are statistically significant (p < 0.05), demonstrating the robustness of the proposed method. These findings demonstrate that UgCmA-Net is highly effective for liver tumor segmentation and quantification, indicating its potential clinical value in liver tumor analysis and treatment planning. Jianfeng Zhao 0004, Shuo Li 0001 |
Knowl. Based Syst. | 2 |
| 2025 | UM-Net: Rethinking ICGNet for polyp segmentation with uncertainty modeling
Xiuquan Du, Xuebin Xu, Jiajia Chen 0006, Lei Li 0048, Heng Liu 0003, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2025 | An orchestration learning framework for ultrasound imaging: Prompt-Guided Hyper-Perception and Attention-Matching Downstream Synchronization
Shuo Li 0001, Shanshan Wang 0010, Zhifan Gao, Yue Sun 0001, Chan-Tong Lam, Xindi Hu, Xin Yang 0009, Dong Ni 0001, Tao Tan 0002 |
Medical Image Anal. | 2 |
| 2025 | Knowledge-driven interpretative conditional diffusion model for contrast-free myocardial infarction enhancement synthesisabstractSynthesis of myocardial infarction enhancement (MIE) images without contrast agents (CAs) has shown great potential to advance myocardial infarction (MI) diagnosis and treatment. It provides results comparable to late gadolinium enhancement (LGE) images, thereby reducing the risks associated with CAs and streamlining clinical workflows. The existing knowledge-and-data-driven approach has made progress in addressing the complex challenges of synthesizing MIE images (i.e., invisible myocardial scars and high inter-individual variability) but still has limitations in the interpretability of kinematic inference, morphological knowledge integration, and kinematic-morphological fusion, thereby reducing the transparency and reliability of the model and causing information loss during synthesis. In this paper, we proposed a knowledge-driven interpretative conditional diffusion model (K-ICDM), which learns kinematic and morphological information from non-enhanced cardiac MR images (CINE sequence and T1 sequence) guided by cardiac knowledge, enabling the synthesis of MIE images. Importantly, our K-ICDM introduces three key innovations that address these limitations, thereby providing interpretability and improving synthesis quality. (1) A novel cardiac causal intervention that generates counterfactual strain to intervene in the inference process from motion maps to abnormal myocardial information, thereby establishing an explicit relationship and providing the clear causal interpretability. (2) A knowledge-driven cognitive combination strategy that utilizes cardiac signal topology knowledge to analyze T1 signal variations, enabling the model to understand how to learn morphological features, thus providing interpretability for morphology capture. (3) An information-specific adaptive fusion strategy that integrates kinematic and morphological information into the conditioning input of the diffusion model based on their specific contributions and adaptively learns their interactions, thereby preserving more detailed information. Experiments on a broad MI dataset with 315 patients show that our K-ICDM achieves state-of-the-art performance in contrast-free MIE image synthesis, improving structural similarity index measure (SSIM) by at least 2.1% over recent methods. These results demonstrate that our method effectively overcomes the limitations of existing methods in capturing the complex relationship between myocardial motion and scar distribution and integrating of static and dynamic sequences, thus enabling the accurate synthesis of subtle scar boundaries. Ronghui Qi, Chenchu Xu, Xiaohu Li, Siyuan Pan, Jie Chen 0025, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2025 | MeMGB-Diff: Memory-Efficient Multivariate Gaussian Bias Diffusion Model for 3D bias field correctionabstractBias fields inevitably degrade MRI that seriously interferes the diagnosis of physicians for accurate analysis, and removing it is a crucial image analysis task. Generative models (such as GANs) are used for bias field correction, and outperform traditional methods, however are hindered by the high cost of data annotation and instability during training. Recently, the diffusion-based methods have excelled over GANs in many applications, and they are powerful in removing noise from images, while the bias field can be regarded as a smooth noise. However, it is a challenge to directly apply to 3D bias field correction due to sampling inefficiency, the heavy computational demand, and implicit correction process. We propose a Memory-Efficient Multivariate Gaussian Bias Diffusion Model (MeMGB-Diff) that is an explicit, sampling, and memory both efficient diffusion model for 3D bias field correction without using clinical labels. MeMGB-Diff extends the diffusion models to multivariate Gaussian and models the bias field as a multivariate Gaussian variable, allowing direct diffusion and removal of the 3D bias fields without Gaussian noise. For memory efficiency, MeMGB-Diff performs diffusion model in smaller readable image domain at the expense of a negligible accuracy loss, based on the strong correlation among adjacent voxels of bias field. We also propose a loss function to mainly learn the intensity trend, which mainly causes the inhomogeneity of MRI, and effectively increases the correction accuracy. For comprehensive performance comparison, we propose a synthetic method for generating more varied bias fields during testing. Both quantitative and qualitative assessments on synthetic and clinical data confirm the high fidelity and uniform intensity of our results. MeMGB-Diff reduces data size by 64 times to use less memory, improves sampling efficiency by more than 10 times compared to other diffusion-based methods, and achieves optimal metrics, including SSIM, PSNR, COCO, and CV for various tissues. Hence, our MeMGB-Diff is a state-of-the-art (SOTA) method for 3D bias field correction. Xingyu Qiu, Dong Liang 0001, Gongning Luo, Xiangyu Li 0004, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2025 | Dynamic spectrum-driven hierarchical learning network for polyp segmentation
Kai-Ni Wang, Jie Hua 0004, Yang Chen 0008, Guangquan Zhou, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2025 | TSdetector: Temporal-Spatial self-correction collaborative learning for colonoscopy video detection
Kai-Ni Wang, Guangquan Zhou, Ling Yang 0006, Yang Chen 0008, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2025 | When evidence modeling meets knowledge distillation: Towards reliable contrast-enhanced knowledge distillation for non-contrast medical image segmentationabstractContrast-enhanced knowledge distillation promises to transform medical diagnostics and reveal promising approaches for tumor segmentation on non-contrast medical images. However, existing methods related to contrast-enhanced knowledge distillation still make it hard to distill reliable contrast-enhanced knowledge for tumor segmentation due to the limitations of (1) unable to quantify uncertainty information for reliable contrast-enhanced and non-contrast knowledge modeling, which leads to an over-confidence cross-domain adaptation for transferring contrast-enhanced knowledge; (2) using vision information only ignores rich semantic features in medical language, which make it hard to model complex tumor enhancement feature. In this study, we propose an evidence-guided and tumor-aware knowledge distillation (EGTA-KD) for transferring contrast-enhanced domain knowledge to non-contrast domain knowledge. Specifically, to achieve tumor-awareness by embedding semantic features from text, the tumor-aware cross-modal synchronizer (TACMS) is proposed to calculate tumor score maps for matching pixel wise image and text features. To achieve reliable cross-domain modeling for transferring contrast-enhanced knowledge, the innovative uncertainty-quantified evidence unit (UQEU) parameterizes the probability distribution within subjective logic to gather reliable evidence of contrast-enhanced knowledge while quantifying the uncertainty of prediction. Lastly, newly designed dual-level knowledge distillation (DLKD) minimizes tumor score map errors and matches evidence distribution for uncertainty-aware contrast-enhanced knowledge distillation. Extensive experiments of tumor segmentation on non-contrast medical images are performed using multi-modality medical image datasets (i.e., Brain MRI dataset, Liver MRI dataset, and Kidney CT dataset). Experimental results demonstrate the proposed EGTA-KD outperforms the other compared state-of-the-art methods, revealing its superiority of tumor segmentation on non-contrast medical images via uncertainty-aware contrast-enhanced knowledge distillation. Jianfeng Zhao 0004, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2025 | Human gaze-based dual teacher guidance learning for semi-supervised medical image segmentation
Rongjun Ge, Chong Wang 0011, Chunqiang Lu, Cong Xia, Yehui Jiang, Fangyi Xu, Yinsu Zhu, Daoqiang Zhang, Chengyu Liu 0001, Yang Chen 0008, Shuo Li 0001, Yuting He 0001 |
Neural Networks | 12 |
| 2025 | Homeomorphism Prior for False Positive and Negative Problem in Medical Image Dense Contrastive Representation LearningabstractDense contrastive representation learning (DCRL) has greatly improved the learning efficiency for image dense prediction tasks, showing its great potential to reduce the large costs of medical image collection and dense annotation. However, the properties of medical images make unreliable correspondence discovery, bringing an open problem of large-scale false positive and negative (FP&N) pairs in DCRL. In this paper, we propose GEoMetric vIsual deNse sImilarity (GEMINI) learning which embeds the homeomorphism prior to DCRL and enables a reliable correspondence discovery for effective dense contrast. We proposes a deformable homeomorphism learning (DHL) which models the homeomorphism of medical images and learns to estimate a deformable mapping to predict the pixels' correspondence under the condition of topological preservation. It effectively reduces the searching space of pairing and drives an implicit and soft learning of negative pairs via gradient. We also proposes a geometric semantic similarity (GSS) which extracts semantic information in features to measure the alignment degree for the correspondence learning. It will promote the learning efficiency and performance of deformation, constructing positive pairs reliably. We implement two practical variants on two typical representation learning tasks in our experiments. Our promising results on seven datasets which outperform the existing methods show our great superiority. We will release our code at a companion website. Yuting He 0001, Boyu Wang 0004, Rongjun Ge, Yang Chen 0008, Guanyu Yang 0001, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Pixel is All You Need: Adversarial Spatio-Temporal Ensemble Active Learning for Salient Object DetectionabstractAlthough weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial spatio-temporal ensemble active learning. Our contributions are four-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed spatio-temporal ensemble strategy not only achieves outstanding performance but significantly reduces the model's computational cost. 3) Our proposed relationship-aware diversity sampling can conquer oversampling while boosting model performance. 4) We provide theoretical proof for the existence of such a point-labeled dataset. Experimental results show that our approach can find such a point-labeled dataset, where a saliency model trained on it obtained 98%-99% performance of its fully-supervised version with only ten annotated points per image. Wei Wang 0169, Yacong Li, Fengmao Lv, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Multi-Domain Adversarial Variational Bayesian Inference for Domain GeneralizationabstractDomain generalization aims to learn common knowledge from multiple observed source domains and transfer it to unseen target domains, e.g. the object recognition in varieties of visual environments. Traditional domain generalization methods aim to learn the feature representation of the raw data with its distribution invariant across domains. This relies on the assumption that the two posterior distributions (the distributions of the label given the feature distribution and given the raw data) are stable in different domains. However, this does not always hold in many practical situations. In this paper, we relax the above assumption by permitting the posterior distribution of the label given the raw data changes in difference domains, and thus focuses on a more realistic learning problem that infers the conditional domain-invariant feature representation. Specifically, a multi-domain adversarial variational Bayesian inference approach is proposed to minimize the inter-domain discrepancy of the conditional distributions of the feature given the label. Besides, it is imposed by the constraints from the adversarial learning and feedback mechanism to enhance the condition invariant feature representation. The extensive experiments on two datasets demonstrate the effectiveness of our approach, as well as the state-of-the-art performance comparing with thirteen methods. Zhifan Gao, Saidi Guo, Chenchu Xu, Jinglin Zhang 0001, Mingming Gong, Javier Del Ser, Shuo Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | MedFILIP: Medical Fine-Grained Language-Image Pre-TrainingabstractMedical vision-language pretraining (VLP) that leverages naturally-paired medical image-report data is crucial for medical image analysis. However, existing methods struggle to accurately characterize associations between images and diseases, leading to inaccurate or incomplete diagnostic results. In this work, we propose MedFILIP, a fine-grained VLP model, introduces medical image-specific knowledge through contrastive learning, specifically: 1) An information extractor based on a large language model is proposed to decouple comprehensive disease details from reports, which excels in extracting disease deals through flexible prompt engineering, thereby effectively reducing text complexity while retaining rich information at a tiny cost. 2) A knowledge injector is proposed to construct relationships between categories and visual attributes, which help the model to make judgments based on image features, and fosters knowledge extrapolation to unfamiliar disease categories. 3) A semantic similarity matrix based on fine-grained annotations is proposed, providing smoother, information-richer labels, thus allowing fine-grained image-text alignment. 4) We validate MedFILIP on numerous datasets, e.g., RSNA-Pneumonia, NIH ChestX-ray14, VinBigData, and COVID-19. For single-label, multi-label, and fine-grained classification, our model achieves state-of-the-art performance, the classification accuracy has increased by a maximum of 6.69%. Xinjie Liang, Xiangyu Li 0004, Fanding Li, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Gongning Luo, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | MedKAFormer: When Kolmogorov-Arnold Theorem Meets Vision Transformer for Medical Image RepresentationabstractVision Transformers (ViTs) suffer from high parameter complexity because they rely on Multi-layer Perceptrons (MLPs) for nonlinear representation. This issue is particularly challenging in medical image analysis, where labeled data is limited, leading to inadequate feature representation. Existing methods have attempted to optimize either the patch embedding stage or the non-embedding stage of ViTs. Still, they have struggled to balance effective modeling, parameter complexity, and data availability. Recently, the Kolmogorov-Arnold Network (KAN) was introduced as an alternative to MLPs, offering a potential solution to the large parameter issue in ViTs. However, KAN cannot be directly integrated into ViT due to challenges such as handling 2D structured data and dimensionality catastrophe. To solve this problem, we propose MedKAFormer, the first ViT model to incorporate the Kolmogorov-Arnold (KA) theorem for medical image representation. It includes a Dynamic Kolmogorov-Arnold Convolution (DKAC) layer for flexible nonlinear modeling in the patch embedding stage. Additionally, it introduces a Nonlinear Sparse Token Mixer (NSTM) and a Nonlinear Dynamic Filter (NDF) in the non-embedding stage. These components provide comprehensive nonlinear representation while reducing model overfitting. MedKAFormer reduces parameter complexity by 85.61% compared to ViT-Base and achieves competitive results on 14 medical datasets across various imaging modalities and structures. Qikui Zhu, Chaoda Song, Benzheng Wei, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | VLD-Net: Localization and Detection of the Vertebrae From X-Ray Images by Reinforcement Learning With Adaptive Exploration Mechanism and Spine Anatomy InformationabstractAccurate and efficient vertebrae localization and detection in X-ray images are essential for diagnosing and treating spinal diseases. However, most existing methods struggle with the complexity of spine X-ray images, yielding inaccurate results due to insufficient utilization of spinal anatomy information and neglect of individual vertebra characteristics. In this paper, we propose an innovative Vertebrae Localization and Detection Network (VLD-Net) to accurately assist physicians in diagnosing spine-related diseases from X-ray images. Our VLD-Net, for the first time, defines vertebrae localization as a top-bottom sequential decision-making process, employing deep reinforcement learning (DRL) to fully leverage the anatomical information of the spine. Simultaneously, it also prioritizes the distinct characteristics of each vertebra for accurate detection. Specifically, VLD-Net combines three key components: 1) An advanced vertebrae localization module based on DRL is proposed, effectively leveraging anatomical information of the spine. 2) A novel adaptive exploration mechanism is coined to understand the behavior of the DRL agent during training, pinpointing how to effectively achieve the trade-off between exploration and exploitation. 3) An innovative vertebra-focused module is proposed to accurately detect vertebral landmarks, using the attention region of each vertebra as input to enhance focus on the target and reduce interference from surrounding tissue. Extensive experiments on two public spine datasets demonstrate that the VLD-Net outperforms the state-of-the-art methods in accuracy and robustness. Shun Xiang, Lei Zhang 0202, Yuanquan Wang 0001, Shoujun Zhou, Xing Zhao 0006, Tao Zhang 0131, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Common-Unique Decomposition Driven Diffusion Model for Contrast-Enhanced Liver MR Images Multi-Phase InterconversionabstractAll three contrast-enhanced (CE) phases (e.g., Arterial, Portal Venous, and Delay) are crucial for diagnosing liver tumors. However, acquiring all three phases is constrained due to contrast agents (CAs) risks, long imaging time, and strict imaging criteria. In this paper, we propose a novel Common-Unique Decomposition Driven Diffusion Model (CUDD-DM), capable of converting any two input phases in three phases into the remaining one, thereby reducing patient wait time, conserving medical resources, and reducing the use of CAs. 1) The Common-Unique Feature Decomposition Module, by utilizing spectral decomposition to capture both common and unique features among different inputs, not only learns correlations in highly similar areas between two input phases but also learns differences in different areas, thereby laying a foundation for the synthesis of remaining phase. 2) The Multi-scale Temporal Reset Gates Module, by bidirectional comparing lesions in current and multiple historical slices, maximizes reliance on previous slices when no lesions and minimizes this reliance when lesions are present, thereby preventing interference between consecutive slices. 3) The Diffusion Model-Driven Lesion Detail Synthesis Module, by employing a continuous and progressive generation process, accurately captures detailed features between data distributions, thereby avoiding the loss of detail caused by traditional methods (e.g., GAN) that overfocus on global distributions. Extensive experiments on a generalized CE liver tumor dataset have demonstrated that our CUDD-DM achieves state-of-the-art performance (improved the SSIM by at least 2.2% (lesions area 5.3%) comparing the seven leading methods). These results demonstrate that CUDD-DM advances CE liver tumor imaging technology. Chenchu Xu, Shijie Tian, Kemal Polat, Adi Alhudhaif, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Energy-Induced Explicit Quantification for Multi-modality MRI Fusion
Xiaoming Qi, Yuan Zhang 0019, Tong Wang 0022, Guanyu Yang 0001, Yueming Jin, Shuo Li 0001 |
ECCV (7) | 6 |
| 2024 | Class-consistent Contrastive Learning Driven Cross-dimensional Transformer for 3D Medical Image Classification
Qikui Zhu, Chuan Fu, Shuo Li 0001 |
IJCAI | 3 |
| 2024 | Single-source Domain Generalization in Deep Learning Segmentation via Lipschitz Regularization
Mazlum Ferhat Arslan, Shuo Li 0001 |
MICCAI (10) | 3 |
| 2024 | Multilevel Causality Learning for Multi-label Gastric Atrophy Diagnosis
Xiaoxiao Cui, Shanzhi Jiang, Baolin Sun, Yankun Cao, Zhen Li 0049, Chaoyang Lv, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001 |
MICCAI (3) | 10 |
| 2024 | CausCLIP: Causality-Adapting Visual Scoring of Visual Language Models for Few-Shot Learning in Portable Echocardiography Quality Assessment
Xiaoxiao Cui, Yankun Cao, Yuezhong Zhang, Li-Zhen Cui 0001, Zhi Liu 0004, Shuo Li 0001 |
MICCAI (1) | 8 |
| 2024 | Spatio-Temporal Contrast Network for Data-Efficient Learning of Coronary Artery Disease in Coronary CT Angiography
Xinghua Ma, Mingye Zou, Xinyan Fang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Zhaowen Qiu, Xin Gao 0001, Shuo Li 0001 |
MICCAI (11) | 10 |
| 2024 | Center-to-Edge Denoising Diffusion Probabilistic Models with Cross-domain Attention for Undersampled MRI Reconstruction
Jianfeng Zhao 0004, Shuo Li 0001 |
MICCAI (7) | 2 |
| 2024 | Client Selection Mechanism for Federated Learning Based on Class Imbalance
Congjie Lin, Zhangshuai Bie, Shuo Li 0001, Xuehua Bi |
PRCV (1) | 4 |
| 2024 | Long-short-view aware multi-agent reinforcement learning for signal snippet distillation in delirium movement detection
Qingtao Pan, Hao Wang 0255, Jingjiao Lou, Bing Ji 0001, Shuo Li 0001 |
Inf. Sci. | 6 |
| 2024 | Unified bi-encoder bispace-discriminator disentanglement for cross-domain echocardiography segmentation
Xiaoxiao Cui, Boyu Wang 0004, Shanzhi Jiang, Zhi Liu 0004, Hongji Xu, Li-Zhen Cui 0001, Shuo Li 0001 |
Knowl. Based Syst. | 7 |
| 2024 | An interpretable two-branch bi-coordinate network based on multi-grained domain knowledge for classification of thyroid nodules in ultrasound images
Ziyue Xu 0001, Weiwei Zhan, Jing Xiao 0006, Yiqing Hou, Bingsheng Huang, Lingyun Huang, Shuo Li 0001 |
Medical Image Anal. | 10 |
| 2024 | MIST: Multi-instance selective transformer for histopathological subtype prediction
Rongchang Zhao, Zijun Xi, Huanchi Liu, Xiangkun Jian, Jian Zhang 0048, Zijian Zhang 0004, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2024 | Dual domain distribution disruption with semantics preservation: Unsupervised domain adaptation for medical image segmentation
Boyun Zheng, Songhui Diao, Jingke Zhu, Yixuan Yuan, Jing Cai 0001, Shuo Li 0001, Wenjian Qin |
Medical Image Anal. | 8 |
| 2024 | Boosting knowledge diversity, accuracy, and stability via tri-enhanced distillation for domain continual medical image segmentation
Zhanshi Zhu, Xinghua Ma, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Lianming Wu, Gongning Luo, Guohua Wang 0001, Shuo Li 0001 |
Medical Image Anal. | 9 |
| 2024 | STANet: Spatio-Temporal Adaptive Network and Clinical Prior Embedding Learning for 3D+T CMR SegmentationabstractThe segmentation of cardiac structure in magnetic resonance images (CMR) is paramount in diagnosing and managing cardiovascular illnesses, given its 3D+Time (3D+T) sequence. The existing deep learning methods are constrained in their ability to 3D+T CMR segmentation, due to: (1) Limited motion perception. The complexity of heart beating renders the motion perception in 3D+T CMR, including the long-range and cross-slice motions. The existing methods' local perception and slice-fixed perception directly limit the performance of 3D+T CMR perception. (2) Lack of labels. Due to the expensive labeling cost of the 3D+T CMR sequence, the labels of 3D+T CMR only contain the end-diastolic and end-systolic frames. The incomplete labeling scheme causes inefficient supervision. Hence, we propose a novel spatio-temporal adaptation network with clinical prior embedding learning (STANet) to ensure efficient spatio-temporal perception and optimization on 3D+T CMR segmentation. (1) A spatio-temporal adaptive convolution (STAC) treats the 3D+T CMR sequence as a whole for perception. The long-distance motion correlation is embedded into the structural perception by learnable weight regularization to balance long-range motion perception. The structural similarity is measured by cross-attention to adaptively correlate the cross-slice motion. (2) A clinical prior embedding learning strategy (CPE) is proposed to optimize the partially labeled 3D+T CMR segmentation dynamically by embedding clinical priors into optimization. STANet achieves outstanding performance with Dice of 0.917 and 0.94 on two public datasets (ACDC and STACOM), which indicates STANet has the potential to be incorporated into computer-aided diagnosis tools for clinical application. Xiaoming Qi, Yuting He 0001, Yaolei Qi, Youyong Kong, Guanyu Yang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Automatic Delineation of the 3D Left Atrium From LGE-MRI: Actor-Critic Based Detection and Semi-Supervised SegmentationabstractAccurate and automatic delineation of the left atrium (LA) is crucial for computer-aided diagnosis of atrial fibrillation-related diseases. However, effective model training typically requires a large amount of labeled data, which is time-consuming and labor-intensive. In this study, we propose a novel LA delineation framework. The region of LA is first detected using an actor-critic based deep reinforcement learning method with a shape-adaptive detection strategy using only box-level annotations, bypassing the need for voxel-level labeling. With the effectively detected LA, the impacts of class-imbalance and interference from surrounding tissues are significantly reduced. Subsequently, a semi-supervised segmentation scheme is coined to precisely delineate the contour of LA in 3D volume. The scheme integrates two independent networks with distinct structures, enabling implicit consistency regularization, capturing more spatial features, and avoiding the error accumulation present in current mainstream semi-supervised frameworks. Specifically, one network is combined with Transformer to capture latent spatial features, while the other network is based on pure CNN to capture local features. The difference prediction between these two sub-networks is exploited to mutually provide high-quality pseudo-labels and correct the cognitive bias. Experimental results on two public datasets demonstrate that our proposed strategy outperforms several state-of-the-art methods in terms of accuracy and clinical convenience. Shun Xiang, Yuanquan Wang 0001, Shoujun Zhou, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Learning Better Registration to Learn Better Few-Shot Medical Image Segmentation: Authenticity, Diversity, and RobustnessabstractIn this work, we address the task of few-shot medical image segmentation (MIS) with a novel proposed framework based on the learning registration to learn segmentation (LRLS) paradigm. To cope with the limitations of lack of authenticity, diversity, and robustness in the existing LRLS frameworks, we propose the better registration better segmentation (BRBS) framework with three main contributions that are experimentally shown to have substantial practical merit. First, we improve the authenticity in the registration-based generation program and propose the knowledge consistency constraint strategy that constrains the registration network to learn according to the domain knowledge. It brings the semantic-aligned and topology-preserved registration, thus allowing the generation program to output new data with great space and style authenticity. Second, we deeply studied the diversity of the generation process and propose the space-style sampling program, which introduces the modeling of the transformation path of style and space change between few atlases and numerous unlabeled images into the generation program. Therefore, the sampling on the transformation paths provides much more diverse space and style features to the generated data effectively improving the diversity. Third, we first highlight the robustness in the learning of segmentation in the LRLS paradigm and propose the mix misalignment regularization, which simulates the misalignment distortion and constrains the network to reduce the fitting degree of misaligned regions. Therefore, it builds regularization for these regions improving the robustness of segmentation learning. Without any bells and whistles, our approach achieves a new state-of-the-art performance in few-shot MIS on two challenging tasks that outperform the existing LRLS-based few-shot methods. We believe that this novel and effective framework will provide a powerful few-shot benchmark for the field of medical image and efficiently reduce the costs of medical image research. All of our code will be made publicly available online. Yuting He 0001, Rongjun Ge, Xiaoming Qi, Yang Chen 0008, Jiasong Wu, Jean-Louis Coatrieux, Guanyu Yang 0001, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object DetectionabstractAlthough weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image. Wei Wang 0169, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
AAAI | 7 |
| 2023 | Synergistically Learning Class-specific Tokens for Multi-class Whole Slide Image ClassificationabstractThe application of transformer architecture in analyzing whole slide images (WSIs) has become increasingly popular due to its remarkable ability to learn complex associations. Nevertheless, a significant drawback emerges in the multiclass analysis of WSIs. The majority of the transformer-based methods available currently rely primarily on a single, class-agnostic token. This approach might not ideally capture the subtleties of class-discriminative information. To address this challenge, we present an innovative approach tailored for multi-class WSI analysis that harnesses the power of class-specific tokens. Central to our method is a novel attention mechanism designed to foster a synergistic learning relationship between patch and class tokens, enhancing the granularity of information captured and ensuring a more comprehensive representation of the WSI. Complementing this, we introduce a dynamic class-centric training strategy designed to optimize token representation learning, ensuring each token is informatively aligned with its corresponding class. Through extensive experimentation on three challenging multi-class WSI analysis datasets, our method consistently demonstrates superior performance, underscoring its potential as a robust solution for multi-class WSI analysis tasks. Pengzhong Sun, Wei Wang 0169, Xiangyu Li 0004, Suyu Dong, Shuo Li 0001, Kuanquan Wang, Gongning Luo |
BIBM | 5 |
| 2023 | Geometric Visual Similarity Learning in 3D Medical Image Self-Supervised Pre-trainingabstractLearning inter-image similarity is crucial for 3D medical images self-supervised pre-training, due to their sharing of numerous same semantic regions. However, the lack of the semantic prior in metrics and the semantic-independent variation in 3D medical images make it challenging to get a reliable measurement for the inter-image similarity, hindering the learning of consistent representation for same semantics. We investigate the challenging problem of this task, i.e., learning a consistent representation between images for a clustering effect of same semantic features. We propose a novel visual similarity learning paradigm, Geometric Visual Similarity Learning, which embeds the prior of topological invariance into the measurement of the inter-image similarity for consistent representation of semantic regions. To drive this paradigm, we further construct a novel geometric matching head, the Z-matching head, to collaboratively learn the global and local similarity of semantic regions, guiding the efficient representation learning for different scale-level inter-image semantic features. Our experiments demonstrate that the pre-training with our learning of inter-image similarity yields more powerful inner-scene, inter-scene, and global-local transferring ability on four challenging 3D medical image tasks. Our codes and pre-trained models will be publicly available11https://github.com/YutingHe-list/GVSL. Yuting He 0001, Guanyu Yang 0001, Rongjun Ge, Yang Chen 0008, Jean-Louis Coatrieux, Boyu Wang 0004, Shuo Li 0001 |
CVPR | 7 |
| 2023 | Knowledge Boosting: Rethinking Medical Contrastive Vision-Language Pre-training
Yuting He 0001, Cheng Xue 0003, Rongjun Ge, Shuo Li 0001, Guanyu Yang 0001 |
MICCAI (1) | 5 |
| 2023 | A Style Transfer-Based Augmentation Framework for Improving Segmentation and Classification Performance Across Different Sources in Ultrasound Images
Bin Huang 0021, Ziyue Xu 0001, S. C. Chan 0001, Zhong Liu 0004, Huiying Wen, Qicai Huang, Meiqin Jiang, Changfeng Dong, Ruhai Zou, Bingsheng Huang, Xin Chen 0025, Shuo Li 0001 |
MICCAI (6) | 14 |
| 2023 | Learning Reliability of Multi-modality Medical Images for Tumor Segmentation via Evidence-Identified Denoising Diffusion Probabilistic Models
Jianfeng Zhao 0004, Shuo Li 0001 |
MICCAI (4) | 2 |
| 2023 | DCAug: Domain-Aware and Content-Consistent Cross-Cycle Framework for Tumor Augmentation
Qikui Zhu, Yanxiang Cheng, Shuo Li 0001 |
MICCAI (5) | 6 |
| 2023 | Heuristic multi-modal integration framework for liver tumor detection from multi-modal non-enhanced MRIs
Dong Zhang 0009, Chenchu Xu, Shuo Li 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Context-aware network fusing transformer and V-Net for semi-supervised segmentation of 3D left atrium
Chenji Zhao, Shun Xiang, Yuanquan Wang 0001, Zhaoxi Cai, Jun Shen 0008, Shoujun Zhou, Weihua Su, Shijie Guo, Shuo Li 0001 |
Expert Syst. Appl. | 10 |
| 2023 | O2M-UDA: Unsupervised dynamic domain adaptation for one-to-multiple medical image segmentation
Ziyue Jiang 0004, Yuting He 0001, Xiaomei Zhu, Yi Xu 0001, Yang Chen 0008, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001 |
Knowl. Based Syst. | 9 |
| 2023 | Ambiguity-aware breast tumor cellularity estimation via self-ensemble label distribution learning
Xiangyu Li 0004, Xinjie Liang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2023 | Curriculum label distribution learning for imbalanced medical image segmentation
Xiangyu Li 0004, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2023 | Spatiotemporal knowledge teacher-student reinforcement learning to detect liver tumors without contrast agents
Chenchu Xu, Yuhong Song, Dong Zhang 0009, Leonardo Kayat Bittencourt, Sree Harsha Tirumani, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2023 | Attractive deep morphology-aware active contour network for vertebral body contour extraction with extensions to heterogeneous and semi-supervised scenarios
Yikang Wang, Hanying Zheng 0004, An Zeng, Fuxin Wei, Sadeer Al-Kindi, Shuo Li 0001 |
Medical Image Anal. | 10 |
| 2023 | From sMRI to task-fMRI: A unified geometric deep learning framework for cross-modal brain anatomo-functional mapping
Taicheng Huang, Zonglei Zhen, Boyu Wang 0004, Xia Wu 0001, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2023 | B-mode ultrasound based CAD for liver cancers via multi-view privileged information learning
Xiangmin Han, Bangming Gong, Lehang Guo, Jun Wang 0024, Shihui Ying, Shuo Li 0001, Jun Shi 0004 |
Neural Networks | 6 |
| 2023 | Towards Accurate and Robust Domain Adaptation Under Multiple Noisy EnvironmentsabstractIn many non-stationary environments, machine learning algorithms usually confront the distribution shift scenarios. Previous domain adaptation methods have achieved great success. However, they would lose algorithm robustness in multiple noisy environments where the examples of source domain become corrupted by label noise, feature noise, or open-set noise. In this paper, we report our attempt toward achieving noise-robust domain adaptation. We first give a theoretical analysis and find that different noises have disparate impacts on the expected target risk. To eliminate the effect of source noises, we propose offline curriculum learning minimizing a newly-defined empirical source risk. We suggest a proxy distribution-based margin discrepancy to gradually decrease the noisy distribution distance to reduce the impact of source noises. We propose an energy estimator for assessing the outlier degree of open-set-noise examples to defeat the harmful influence. We also suggest robust parameter learning to mitigate the negative effect further and learn domain-invariant feature representations. Finally, we seamlessly transform these components into an adversarial network that performs efficient joint optimization for them. A series of empirical studies on the benchmark datasets and the COVID-19 screening task show that our algorithm remarkably outperforms the state-of-the-art, with over 10% accuracy improvements in some transfer tasks. Zhongyi Han, Xian-Jin Gui, Haoliang Sun, Yilong Yin, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Carotid Lumen Diameter and Intima-Media Thickness Measurement via Boundary-Guided Pseudo-LabelingabstractThe carotid lumen diameter (CALD) and intima-media thickness (CIMT) are essential indicators for diagnosing and treating cardiovascular disease. Existing supervised methods rely on extensive labeled data, which is time-consuming and labor-intensive. This letter presents a boundary-guided pseudo-labeling (BGPL) method, which adopts prior anatomical knowledge to generate and select reliable pseudo-labels for unlabeled data to boost measurement performance. To improve the quality of pseudo-labels, we propose a fine self-attention (FSA) module and a boundary attention module (BAM) to force the feature extractor to highlight boundary information. The FSA maintains internal resolution and enhances non-linearity. The BAM injects boundary heatmaps from the pre-trained boundary regressor into the feature extractor. We specifically provide an alternate training strategy to increase the quality of pseudo-labels further. We evaluate the proposed method on challenging carotid ultrasound datasets. Experiments show that the proposed method outperforms the existing state-of-the-art algorithm. Shimeng Yang, Teng Li 0001, Yinping Lv, Shuo Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Dual Multiscale Mean Teacher Network for Semi-Supervised Infection Segmentation in Chest CT Volume for COVID-19abstractAutomated detecting lung infections from computed tomography (CT) data plays an important role for combating coronavirus 2019 (COVID-19). However, there are still some challenges for developing AI system: 1) most current COVID-19 infection segmentation methods mainly relied on 2-D CT images, which lack 3-D sequential constraint; 2) existing 3-D CT segmentation methods focus on single-scale representations, which do not achieve the multiple level receptive field sizes on 3-D volume; and 3) the emergent breaking out of COVID-19 makes it hard to annotate sufficient CT volumes for training deep model. To address these issues, we first build a multiple dimensional-attention convolutional neural network (MDA-CNN) to aggregate multiscale information along different dimension of input feature maps and impose supervision on multiple predictions from different convolutional neural networks (CNNs) layers. Second, we assign this MDA-CNN as a basic network into a novel dual multiscale mean teacher network (DM [Formula: see text]-Net) for semi-supervised COVID-19 lung infection segmentation on CT volumes by leveraging unlabeled data and exploring the multiscale information. Our DM [Formula: see text]-Net encourages multiple predictions at different CNN layers from the student and teacher networks to be consistent for computing a multiscale consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from multiple predictions of MDA-CNN. Third, we collect two COVID-19 segmentation datasets to evaluate our method. The experimental results show that our network consistently outperforms the compared state-of-the-art methods. Liansheng Wang 0002, Jiacheng Wang 0002, Lei Zhu 0003, Huazhu Fu, Ping Li 0016, Gary Cheng 0001, Shuo Li 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 8 |
| 2023 | Trajectory-Aware Adaptive Imaging Clue Analysis for Guidewire Artifact Removal in Intravascular Optical Coherence TomographyabstractGuidewire Artifact Removal (GAR) involves restoring missing imaging signals in areas of IntraVascular Optical Coherence Tomography (IVOCT) videos affected by guidewire artifacts. GAR helps overcome imaging defects and minimizes the impact of missing signals on the diagnosis of CardioVascular Diseases (CVDs). To restore the actual vascular and lesion information within the artifact area, we propose a reliable Trajectory-aware Adaptive imaging Clue analysis Network (TAC-Net) that includes two innovative designs: (i) Adaptive clue aggregation, which considers both texture-focused original (ORI) videos and structure-focused relative total variation (RTV) videos, and suppresses texture-structure imbalance with an active weight-adaptation mechanism; (ii) Trajectory-aware Transformer, which uses a novel attention calculation to perceive the attention distribution of artifact trajectories and avoid the interference of irregular and non-uniform artifacts. We provide a detailed formulation for the procedure and evaluation of the GAR task and conduct comprehensive quantitative and qualitative experiments. The experimental results demonstrate that TAC-Net reliably restores the texture and structure of guidewire artifact areas as expected by experienced physicians (e.g., SSIM: 97.23%). We also discuss the value and potential of the GAR task for clinical applications and computer-aided diagnosis of CVDs. Gongning Luo, Xinghua Ma, Jinwen Guo, Mingye Zou, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | ARR-GCN: Anatomy-Relation Reasoning Graph Convolutional Network for Automatic Fine-Grained Segmentation of Organ's Surgical AnatomyabstractAnatomical resection (AR) based on anatomical sub-regions is a promising method of precise surgical resection, which has been proven to improve long-term survival by reducing local recurrence. The fine-grained segmentation of an organ's surgical anatomy (FGS-OSA), i.e., segmenting an organ into multiple anatomic regions, is critical for localizing tumors in AR surgical planning. However, automatically obtaining FGS-OSA results in computer-aided methods faces the challenges of appearance ambiguities among sub-regions (i.e., inter-sub-region appearance ambiguities) caused by similar HU distributions in different sub-regions of an organ's surgical anatomy, invisible boundaries, and similarities between anatomical landmarks and other anatomical information. In this paper, we propose a novel fine-grained segmentation framework termed the "anatomic relation reasoning graph convolutional network" (ARR-GCN), which incorporates prior anatomic relations into the framework learning. In ARR-GCN, a graph is constructed based on the sub-regions to model the class and their relations. Further, to obtain discriminative initial node representations of graph space, a sub-region center module is designed. Most importantly, to explicitly learn the anatomic relations, the prior anatomic-relations among the sub-regions are encoded in the form of an adjacency matrix and embedded into the intermediate node representations to guide framework learning. The ARR-GCN was validated on two FGS-OSA tasks: i) liver segments segmentation, and ii) lung lobes segmentation. Experimental results on both tasks outperformed other state-of-the-art segmentation methods and yielded promising performances by ARR-GCN for suppressing ambiguities among sub-regions. Yinli Tian, Wenjian Qin, Ricardo Lambo, Meiyan Yue, Songhui Diao, Lequan Yu, Yaoqin Xie, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2023 | Adaptive Frequency Learning Network With Anti-Aliasing Complex Convolutions for Colon Diseases SubtypesabstractThe automatic and dependable identification of colonic disease subtypes by colonoscopy is crucial. Once successful, it will facilitate clinically more in-depth disease staging analysis and the formulation of more tailored treatment plans. However, inter-class confusion and brightness imbalance are major obstacles to colon disease subtyping. Notably, the Fourier-based image spectrum, with its distinctive frequency features and brightness insensitivity, offers a potential solution. To effectively leverage its advantages to address the existing challenges, this article proposes a framework capable of thorough learning in the frequency domain based on four core designs: the position consistency module, the high-frequency self-supervised module, the complex number arithmetic model, and the feature anti-aliasing module. The position consistency module enables the generation of spectra that preserve local and positional information while compressing the spectral data range to improve training stability. Through band masking and supervision, the high-frequency autoencoder module guides the network to learn useful frequency features selectively. The proposed complex number arithmetic model allows direct spectral training while avoiding the loss of phase information caused by current general-purpose real-valued operations. The feature anti-aliasing module embeds filters in the model to prevent spectral aliasing caused by down-sampling and improve performance. Experiments are performed on the collected five-class dataset, which contains 4591 colorectal endoscopic images. The outcomes show that our proposed method produces state-of-the-art results with an accuracy rate of 89.82%. Kai-Ni Wang, Shuaishuai Zhuang, Juzheng Miao, Yang Chen 0008, Jie Hua 0004, Guangquan Zhou, Xiaopu He, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | BMAnet: Boundary Mining With Adversarial Learning for Semi-Supervised 2D Myocardial Infarction SegmentationabstractAutomatic segmentation of myocardial infarction (MI) regions in late gadolinium-enhanced cardiac magnetic resonance images is an essential step in the computed diagnosis of myocardial infarction. Most of the current myocardial infarction region segmentation methods are based on fully supervised deep learning. However, cardiologists' annotation of myocardial infarction regions in cardiac magnetic resonance images during the diagnosis process is time-consuming and expensive. This paper proposes a semi-supervised myocardial infarction segmentation. It consists of two models: 1) a boundary mining model and 2) an adversarial learning model. The boundary mining model can solve the boundary ambiguity problem by enlarging the gap between the foreground and background features, thus segmenting the myocardial infarction region accurately. The adversarial learning model can make the boundary mining model learn from additional unlabeled data by evaluating the segmentation performance and providing pseudo supervision, which significantly increases the robustness of the boundary mining model. We conduct extensive experiments on an in-house myocardial magnetic resonance dataset. The experimental results on six evaluation metrics demonstrate that our method achieves excellent results in myocardial infarction segmentation and outperforms the state-of-the-art semi-supervised methods. Chenchu Xu, Dong Zhang 0009, Longfei Han, Yanping Zhang 0001, Jie Chen 0025, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | A Knowledge-Guided Framework for Fine-Grained Classification of Liver Lesions Based on Multi-Phase CT ImagesabstractAutomatic and accurate differentiation of liver lesions from multi-phase computed tomography imaging is critical for the early detection of liver cancer. Multi-phase data can provide more diagnostic information than single-phase data, and the effective use of multi-phase data can significantly improve diagnostic accuracy. Current fusion methods usually fuse multi-phase information at the image level or feature level, ignoring the specificity of each modality, therefore, the information integration capacity is always limited. In this paper, we propose a Knowledge-guided framework, named MCCNet, which adaptively integrates multi-phase liver lesion information from three different stages to fully utilize and fuse multi-phase liver information. Specifically, 1) a multi-phase self-attention module was designed to adaptively combine and integrate complementary information from three phases using multi-level phase features; 2) a cross-feature interaction module was proposed to further integrate multi-phase fine-grained features from a global perspective; 3) a cross-lesion correlation module was proposed for the first time to imitate the clinical diagnosis process by exploiting inter-lesion correlation in the same patient. By integrating the above three modules into a 3D backbone, we constructed a lesion classification network. The proposed lesion classification network was validated on an in-house dataset containing 3,683 lesions from 2,333 patients in 9 hospitals. Extensive experimental results and evaluations on real-world clinical applications demonstrate the effectiveness of the proposed modules in exploiting and fusing multi-phase information. Xingxin Xu, Qikui Zhu, Hanning Ying, Jiongcheng Li, Xiujun Cai, Shuo Li 0001, Yizhou Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Multi-Task Learning for Pulmonary Arterial Hypertension Prognosis Prediction via Memory Drift and Prior Prompt Learning on 3D Chest CTabstractPulmonary arterial hypertension (PAH) prognosis prediction on 3D non-contrast CT images is one of the most important tasks for PAH treatment. It will help clinicians stratify patients into different groups for early diagnosis and timely intervention via automatically extracting the potential biomarkers of PAH to predict mortality. However, it is still a task of great challenges due to the large volume and low-contrast regions of interest in 3D chest CT images. In this paper, we propose the first multi-task learning-based PAH prognosis prediction framework, P$^{2}$-Net, which effectively optimizes the model and powerfully represents task-dependent features via our Memory Drift (MD) and Prior Prompt Learning (PPL) strategies. 1) Our MD maintains a large memory bank to provide a dense sampling of the deep biomarkers' distribution. Therefore, although the batch size is very small caused by our large volume, a reliable (negative log partial) likelihood loss is still able to be calculated on a representative probability distribution for robust optimization. 2) Our PPL simultaneously learns an additional manual biomarkers prediction task to embed clinical prior knowledge into our deep prognosis prediction task in hidden and explicit ways. Therefore, it will prompt the prediction of deep biomarkers and improve the perception of task-dependent features in our low-contrast regions. Our P$^{2}$-Net achieves a high prognostic correlation of the prediction and great generalization with the highest 70.19% C-index and 2.14 HR. Extensive experiments with promising results on our PAH prognosis prediction reveal powerful prognosis performance and great clinical significance in PAH treatment. All of our code will be made publicly available online. Guanyu Yang 0001, Yuting He 0001, Yang Chen 0008, Jean-Louis Coatrieux, Xiaoxuan Sun, Yongyue Wei, Shuo Li 0001, Yinsu Zhu |
IEEE J. Biomed. Health Informatics | 9 |
| 2023 | Corrections to "Image Projection Network: 3D to 2D Image Segmentation in OCTA Images"
Mingchao Li 0002, Yerui Chen, Zexuan Ji, Keren Xie, Songtao Yuan, Qiang Chen 0004, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | A lightweight pose estimation network with multi-scale receptive field
Shuo Li 0001, Ju Dai, Zhangmeng Chen, JunJun Pan |
Vis. Comput. | 1 |
| 2022 | MNet: Rethinking 2D/3D Networks for Anisotropic Medical Image SegmentationabstractThe nature of thick-slice scanning causes severe inter-slice discontinuities of 3D medical images, and the vanilla 2D/3D convolutional neural networks (CNNs) fail to represent sparse inter-slice information and dense intra-slice information in a balanced way, leading to severe underfitting to inter-slice features (for vanilla 2D CNNs) and overfitting to noise from long-range slices (for vanilla 3D CNNs). In this work, a novel mesh network (MNet) is proposed to balance the spatial representation inter axes via learning. 1) Our MNet latently fuses plenty of representation processes by embedding multi-dimensional convolutions deeply into basic modules, making the selections of representation processes flexible, thus balancing representation for sparse inter-slice information and dense intra-slice information adaptively. 2) Our MNet latently fuses multi-dimensional features inside each basic module, simultaneously taking the advantages of 2D (high segmentation accuracy of the easily recognized regions in 2D view) and 3D (high smoothness of 3D organ contour) representations, thus obtaining more accurate modeling for target regions. Comprehensive experiments are performed on four public datasets (CT\&MR), the results consistently demonstrate the proposed MNet outperforms the other methods. The code and datasets are available at: https://github.com/zfdong-code/MNet Zhangfu Dong, Yuting He 0001, Xiaoming Qi, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux, Guanyu Yang 0001, Shuo Li 0001 |
IJCAI | 8 |
| 2022 | SAPJNet: Sequence-Adaptive Prototype-Joint Network for Small Sample Multi-sequence MRI Diagnosis
Yuqiang Gao, Guanyu Yang 0001, Xiaoming Qi, Yinsu Zhu, Shuo Li 0001 |
MICCAI (1) | 5 |
| 2022 | DDPNet: A Novel Dual-Domain Parallel Network for Low-Dose CT Reconstruction
Rongjun Ge, Yuting He 0001, Cong Xia, Hai-Long Sun, Yikun Zhang 0001, Dianlin Hu, Yang Chen 0008, Shuo Li 0001, Daoqiang Zhang |
MICCAI (6) | 9 |
| 2022 | ULTRA: Uncertainty-Aware Label Distribution Learning for Breast Tumor Cellularity Assessment
Xiangyu Li 0004, Xinjie Liang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001 |
MICCAI (3) | 6 |
| 2022 | Position-Prior Clustering-Based Self-attention Module for Knee Cartilage Segmentation
Dong Liang 0012, Jun Liu 0080, Kuanquan Wang, Gongning Luo, Wei Wang 0169, Shuo Li 0001 |
MICCAI (5) | 6 |
| 2022 | Contrastive Re-localization and History Distillation in Federated CMR Segmentation
Xiaoming Qi, Guanyu Yang 0001, Yuting He 0001, Wangyan Liu, Ali Islam, Shuo Li 0001 |
MICCAI (5) | 6 |
| 2022 | XMorpher: Full Transformer for Deformable Medical Image Registration via Cross Attention
Yuting He 0001, Youyong Kong, Jean-Louis Coatrieux, Huazhong Shu, Guanyu Yang 0001, Shuo Li 0001 |
MICCAI (6) | 7 |
| 2022 | FFCNet: Fourier Transform-Based Frequency Learning and Complex Convolutional Network for Colon Disease Classification
Kai-Ni Wang, Yuting He 0001, Shuaishuai Zhuang, Juzheng Miao, Xiaopu He, Guanyu Yang 0001, Guangquan Zhou, Shuo Li 0001 |
MICCAI (3) | 9 |
| 2022 | Contrast-Free Liver Tumor Detection Using Ternary Knowledge Transferred Teacher-Student Deep Reinforcement Learning
Chenchu Xu, Dong Zhang 0009, Yuhui Song, Leonardo Kayat Bittencourt, Sree Harsha Tirumani, Shuo Li 0001 |
MICCAI (5) | 6 |
| 2022 | SelfMix: A Self-adaptive Data Augmentation Method for Lesion Segmentation
Qikui Zhu, Jiancheng Yang, Shuo Li 0001 |
MICCAI (4) | 6 |
| 2022 | Synthetic Data Supervised Salient Object DetectionabstractAlthough deep salient object detection (SOD) has achieved remarkable progress, deep SOD models are extremely data-hungry, requiring large-scale pixel-wise annotations to deliver such promising results. In this paper, we propose a novel yet effective method for SOD, coined SODGAN, which can generate infinite high-quality image-mask pairs requiring only a few labeled data, and these synthesized pairs can replace the human-labeled DUTS-TR to train any off-the-shelf SOD model. Its contribution is three-fold. 1) Our proposed diffusion embedding network can address the manifold mismatch and is tractable for the latent code generation, better matching with the ImageNet latent space. 2) For the first time, our proposed few-shot saliency mask generator can synthesize infinite accurate image synchronized saliency masks with a few labeled data. 3) Our proposed quality-aware discriminator can select highquality synthesized image-mask pairs from noisy synthetic data pool, improving the quality of synthetic data. For the first time, our SODGAN tackles SOD with synthetic data directly generated from the generative model, which opens up a new research paradigm for SOD. Extensive experimental results show that the saliency model trained on synthetic data can achieve $98.4%$ F-measure of the saliency model trained on the DUTS-TR. Moreover, our approach achieves a new SOTA performance in semi/weakly-supervised methods, and even outperforms several fully-supervised SOTA methods. Code is available at https://github.com/wuzhenyubuaa/SODGAN Wei Wang 0169, Tengfei Shi, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
ACM Multimedia | 7 |
| 2022 | ResAttenGAN: Simultaneous segmentation of multiple spinal structures on axial lumbar MRI image using residual attention and adversarial learning
Jianhua Liu 0005, Bo Chen 0013, Shuo Li 0001 |
Artif. Intell. Medicine | 4 |
| 2022 | X-CTRSNet: 3D cervical vertebra CT reconstruction and segmentation directly from 2D X-ray images
Rongjun Ge, Yuting He 0001, Cong Xia, Chenchu Xu, Weiya Sun, Guanyu Yang 0001, Hailing Yu, Daoqiang Zhang, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001, Yinsu Zhu |
Knowl. Based Syst. | 13 |
| 2022 | RE-3DLVNet: Refined estimation of the left ventricle volume via interactive 3D segmentation and reinforced quantification
Rongjun Ge, Cong Xia, Yuting He 0001, Hai-Long Sun, Daoqiang Zhang, Guanyu Yang 0001, Wentao Xiang, Jinjun Shi, Limin Luo 0001, Yinsu Zhu, Shuo Li 0001, Yang Chen 0008 |
Knowl. Based Syst. | 11 |
| 2022 | Task relevance driven adversarial learning for simultaneous detection, size grading, and quantification of hepatocellular carcinoma via integrating multi-modality MRI
Xiaojiao Xiao, Jianfeng Zhao 0004, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2022 | MVFStain: Multiple virtual functional stain histopathology images generation based on specific domain mapping
Yankun Cao, Zhi Liu 0004, Jianye Wang, Xiaoyu Sui, Pengfei Zhang 0017, Li-Zhen Cui 0001, Shuo Li 0001 |
Medical Image Anal. | 11 |
| 2022 | Reasoning discriminative dictionary-embedded network for fully automatic vertebrae tumor diagnosis
Heyou Chang, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2022 | Diagnosing glaucoma on imbalanced data with self-ensemble dual-curriculum learning
Rongchang Zhao, Xuanlin Chen, Zailiang Chen 0001, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2022 | TRSA-Net: Task Relation Spatial Co-Attention for Joint Segmentation, Quantification and Uncertainty Estimation on Paired 2D EchocardiographyabstractClinical workflow of cardiac assessment on 2D echocardiography requires both accurate segmentation and quantification of the Left Ventricle (LV) from paired apical 4-chamber and 2-chamber. Moreover, uncertainty estimation is significant in clinically understanding the performance of a model. However, current research on 2D echocardiography ignores this vital task while joint segmentation with quantification, hence motivating the need for a unified optimization method. In this paper, we propose a multitask model with Task Relation Spatial co-Attention (referred as TRSA-Net) for joint segmentation, quantification, and uncertainty estimation on paired 2D echo. TRSA-Net achieves multitask joint learning by novelly exploring the spatial correlation between tasks. The task relation spatial co-attention learns the spatial mapping among task-specific features by non-local and co-excitation, which forcibly joints embedded spatial information in the segmentation and quantification. The Boundary-aware Structure Consistency (BSC) and Joint Indices Constraint (JIC) are integrated into the multitask learning optimization objective to guide the learning of segmentation and quantification paths. The BSC creatively promotes structural similarity of predictions, and JIC explores the internal relationship between three quantitative indices. We validate the efficacy of our TRSA-Net on the public CAMUS dataset. Extensive comparison and ablation experiments show that our approach can achieve competitive segmentation performance and highly accurate results on quantification. Xiaoxiao Cui, Yankun Cao, Zhi Liu 0004, Xiaoyu Sui, Yuezhong Zhang, Li-Zhen Cui 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2022 | Guest Editorial Generative Adversarial Networks in Biomedical Image ComputingabstractThe papers in this special section focus on generative adversarial networks in biomedical image computing. The field of biomedical imaging has obtained great progress from Roentgen’s original discovery of the X-ray to the current imaging tools, including Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Computed Tomography (CT), and Ultrasound (US). The benefits of using these non-invasive imaging technologies are to assess the current condition of an organ or tissue, which can be used to monitor a patient over time over time for accurate and timely diagnosis and treatment.With the development of imaging technologies, developing advanced artificial intelligence algorithms for automated image analysis has shown the potential to change many aspects of clinical applications within the next decade. Meanwhile, these advanced technologies have also brought new issues and challenges. Thus, there has been a growing demand for biomedical imaging computing to be a component of clinical trials and device improvement. Currently, Generative adversarial networks (GANs) have been attached growing interests in the computer vision community due to their capability of data generation or translation. GAN-based models are able to learn from a set of training data and generate new data with the same characteristics as the training ones, which have also proven to be the state of the art for generating sharp and realistic images. More importantly, GAN has been rapidly applied to many traditional and novel applications in the medical domain, such as image reconstruction, segmentation, diagnosis, synthesis, and so on. Despite GAN substantial progress in these areas, their application to medical image computing still faces challenges and unsolved problems remain. Huazhu Fu, Tao Zhou 0002, Shuo Li 0001, Alejandro F. Frangi |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Few-Shot Learning for Deformable Medical Image Registration With Perception-Correspondence Decoupling and Reverse TeachingabstractDeformable medical image registration estimates corresponding deformation to align the regions of interest (ROIs) of two images to a same spatial coordinate system. However, recent unsupervised registration models only have correspondence ability without perception, making misalignment on blurred anatomies and distortion on task-unconcerned backgrounds. Label-constrained (LC) registration models embed the perception ability via labels, but the lack of texture constraints in labels and the expensive labeling costs causes distortion internal ROIs and overfitted perception. We propose the first few-shot deformable medical image registration framework, Perception-Correspondence Registration (PC-Reg), which embeds perception ability to registration models only with few labels, thus greatly improving registration accuracy and reducing distortion. 1) We propose the Perception-Correspondence Decoupling which decouples the perception and correspondence actions of registration to two CNNs. Therefore, independent optimizations and feature representations are available avoiding interference of the correspondence due to the lack of texture constraints. 2) For few-shot learning, we propose Reverse Teaching which aligns labeled and unlabeled images to each other to provide supervision information to the structure and style knowledge in unlabeled images, thus generating additional training data. Therefore, these data will reversely teach our perception CNN more style and structure knowledge, improving its generalization ability. Our experiments on three datasets with only five labels demonstrate that our PC-Reg has competitive registration accuracy and effective distortion-reducing ability. Compared with LC-VoxelMorph( λ = 1), we achieve the 12.5%, 6.3% and 1.0% Reg-DSC improvements on three datasets, revealing our framework with great potential in clinical application. Yuting He 0001, Rongjun Ge, Jian Yang 0009, Youyong Kong, Huazhong Shu, Guanyu Yang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2022 | Hematoma Expansion Context Guided Intracranial Hemorrhage Segmentation and Uncertainty EstimationabstractAccurate segmentation of the Intracranial Hemorrhage (ICH) in non-contrast CT images is significant for computer-aided diagnosis. Although existing methods have achieved remarkable 1 1 The code will be available from https://github.com/JohnleeHIT/SLEX-Net. results, none of them incorporated ICH's prior information in their methods. In this work, for the first time, we proposed a novel SLice EXpansion Network (SLEX-Net), which incorporated hematoma expansion in the segmentation architecture by directly modeling the hematoma variation among adjacent slices. Firstly, a new module named Slice Expansion Module (SEM) was built, which can effectively transfer contextual information between two adjacent slices by mapping predictions from one slice to another. Secondly, to perceive contextual information from both upper and lower slices, we designed two information transmission paths: forward and backward slice expansion, and aggregated results from those paths with a novel weighing strategy. By further exploiting intra-slice and inter-slice context with the information paths, the network significantly improved the accuracy and continuity of segmentation results. Moreover, the proposed SLEX-Net enables us to conduct an uncertainty estimation with one-time inference, which is much more efficient than existing methods. We evaluated the proposed SLEX-Net and compared it with some state-of-the-art methods. Experimental results demonstrate that our method makes significant improvements in all metrics on segmentation performance and outperforms other existing uncertainty estimation methods in terms of several metrics. Xiangyu Li 0004, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Yue Gao 0002, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | MVSGAN: Spatial-Aware Multi-View CMR Fusion for Accurate 3D Left Ventricular Myocardium SegmentationabstractThe accurate 3D left ventricular (LV) myocardium segmentation in short-axis (SAX) view of cardiac magnetic resonance (CMR) is challenged by the sparse spatial structure of CMR. The strategy of multi-view CMR fusion can provide fine-grained spatial structure for accurate segmentation. However, the large information misalignment and lack of dense 3D CMR as fusion target in multi-view CMR fusion, and the different spatial resolution between the fusion result and the ground truth in segmentation limit the strategy. In this study, we propose a multi-view spatial-aware adversarial network (MVSGAN). It studies the perception of fine-grained cardiac structure for accurate segmentation by the spatialaware multi-view CMR fusion. It consists of three modules: (1) A residual adversarial fusion (RAF) module takes inter-slices deep correlation and anatomical prior to refine the spatial structures by residual supplement and adversarial optimization. (2) A structural perception-aggregation (SPA) module establishes the spatial correlation between the dense cardiac model and sparse label for accurate CMR LV myocardium segmentation. (3) A joint training strategy utilizes the dense SAX volume as explicit and implicit goals to jointly optimize the framework. The experiments are applied on a public dataset and a clinical dataset to evaluate the performance of MVSGAN. The average Dice and Jaccard score of LV myocardium segmentation obtained by MVSGAN are highest among seven existing state-of-the-art methods, which are up to 0.92 and 0.75. It is concluded that the spatial-aware multi-view CMR fusion can provide meaningful spatial correlation for accurate LV myocardium segmentation. Xiaoming Qi, Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Jian Yang 0009, Wangyag Liu, Yinsu Zhu, Yi Xu 0001, Huazhong Shu, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2022 | Guest Editorial Artificial Intelligence in Pre-DICOMabstractThe papers in this special section focus on artificial intelligence pre-DICOM medical imaging. AI for medical imaging is applied in three domains: pre-DICOM, pre-processing and clinical applications. Clinical applications mainly cover topics such as disease detection, classification, segmentation, registration. Pre-processing components are mainly designed for facilitating applications using image transformation such as image normalization, noise reduction, bias correction in MR. AI in the pre-DICOM domain is expected to improve imaging workflow, image protocol selection, imaging quality, imaging scanning time before images are converted into DICOM format for radiologists to review. The trends of AI publications in medical imaging have been gradually extended from clinical applications to pre-processing and, to pre-DICOM. The papers in this special section seek to present and highlight the latest development on applying advanced deep learning techniques in pre-DICOM space. The papers highlight the latest development on applying advanced deep learning techniques in pre-DICOM space. Tao Tan 0002, Ravi Soni, Jungong Han, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Regional Cardiac Motion Scoring With Multi-Scale Motion-Based Spatial AttentionabstractRegional cardiac motion scoring aims to classify the motion status of each myocardium segment into one of the four categories (normal, hypokinetic, akinetic, and dyskinetic) from multiple short-axis MR sequences. It is essential for prognosis and early diagnosis for various cardiac diseases. However, the complex motion procedure of the myocardium and the invisible pattern differences pose great challenges, leading to low performance for automatic methods. Most existing works mitigate the task by differentiating the normal motion patterns from the abnormal ones, without fine-grained motion scoring. We propose an effective method for the task of cardiac motion scoring by connecting a bottom-up and another top-down branch with a novel motion-based spatial attention module in multi-scale space. Specifically, we use the convolution blocks for low-level feature extraction that acts as a bottom-up mechanism, and the task of optical flow for explicit motion extraction that acts as a top-down mechanism for high-level allocation of spatial attention. To this end, a newly designed Multi-scale Motion-based Spatial Attention (MMSA) module is used as the pivot connecting the bottom-up part and the top-down part, and adaptively weight the low-level features according to the motion information. Experimental results on a newly constructed dataset of 1440 myocardium segments from 90 subjects demonstrate that the proposed MMSA can accurately analyze the regional myocardium motion, with accuracies of 79.3% for 4-way motion scoring, 89.0% for abnormality detection, and correlation of 0.943 for estimation of motion score index. This work has great potential for practical assessmentof cardiac motion function. Wufeng Xue, Zejian Chen, Tianfu Wang 0001, Shuo Li 0001, Dong Ni 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | SC2Net: A Novel Segmentation-Based Classification Network for Detection of COVID-19 in Chest X-Ray ImagesabstractThe pandemic of COVID-19 has become a global crisis in public health, which has led to a massive number of deaths and severe economic degradation. To suppress the spread of COVID-19, accurate diagnosis at an early stage is crucial. As the popularly used real-time reverse transcriptase polymerase chain reaction (RT-PCR) swab test can be lengthy and inaccurate, chest screening with radiography imaging is still preferred. However, due to limited image data and the difficulty of the early-stage diagnosis, existing models suffer from ineffective feature extraction and poor network convergence and optimisation. To tackle these issues, a segmentation-based COVID-19 classification network, namely SC2Net, is proposed for effective detection of the COVID-19 from chest x-ray (CXR) images. The SC2Net consists of two subnets: a COVID-19 lung segmentation network (CLSeg), and a spatial attention network (SANet). In order to supress the interference from the background, the CLSeg is first applied to segment the lung region from the CXR. The segmented lung region is then fed to the SANet for classification and diagnosis of the COVID-19. As a shallow yet effective classifier, SANet takes the ResNet-18 as the feature extractor and enhances high-level feature via the proposed spatial attention module. For performance evaluation, the COVIDGR 1.0 dataset is used, which is a high-quality dataset with various severity levels of the COVID-19. Experimental results have shown that, our SC2Net has an average accuracy of 84.23% and an average F1 score of 81.31% in detection of COVID-19, outperforming several state-of-the-art approaches. Huimin Zhao 0001, Zhenyu Fang, Jinchang Ren, Calum MacLellan, Yong Xia 0001, Shuo Li 0001, Meijun Sun, Kevin Ren |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Thin Semantics Enhancement via High-Frequency Priori Rule for Thin Structures SegmentationabstractReceptive field-based segmentation models represent features in receptive fields having weak perception for thin semantics in thin structures segmentation, due to the challenges in small local size and large global variation. High-frequency (HiFe) components have strong thin perception ability and is stable for global variation, but its weak adaptability limits its direct application. We propose a HiFe priori rule which enables the network to adaptively extract and fuse HiFe components, enhancing the thin semantics and making the network naturally prefer thin structures for their segmentation. We further propose High-Frequency Semantics Enhancement Network (HiFeNet) based on our HiFe priori rule, boosting the SOTA methods in thin structures segmentation: 1) Our Deep High Frequency (DHiFe) block learns to extract task-dependent HiFe components and adds them to feature maps, achieving great perception of thin structures. 2) Our Latent Residual Denoising (LRD) block progressively weakens task-independent features via hierarchical residuals and learns to fuse HiFe components back to feature maps, further enhancing the thin semantics and weakening the interference of global variation. Extensive experiments on the retinal vessel [1], [2], [3] and Massachusetts road [4] segmentation datasets show great superiority of our HiFeNet. Yuting He 0001, Rongjun Ge, Jiasong Wu, Jean-Louis Coatrieux, Huazhong Shu, Yang Chen 0008, Guanyu Yang 0001, Shuo Li 0001 |
ICDM | 8 |
| 2021 | Learning Consistency- and Discrepancy-Context for 2D Organ Segmentation
Lei Li 0048, Sheng Lian, Zhiming Luo, Shaozi Li, Beizhan Wang, Shuo Li 0001 |
MICCAI (1) | 6 |
| 2021 | CPNet: Cycle Prototype Network for Weakly-Supervised 3D Renal Compartments Segmentation on CT Images
Song Wang 0002, Yuting He 0001, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001 |
MICCAI (2) | 9 |
| 2021 | mfTrans-Net: Quantitative Measurement of Hepatocellular Carcinoma via Multi-Function Transformer Regression Network
Jianfeng Zhao 0004, Xiaojiao Xiao, Dengwang Li, Jaron Chong, Zahra Kassam, Bo Chen 0013, Shuo Li 0001 |
MICCAI (5) | 7 |
| 2021 | Applying Cross-Modality Data Processing for Infarction Learning in Medical Internet of ThingsabstractCross-modality data processing is critical for the Internet-of-Things (IoT) deployment in healthcare. It can convert the innumerable raw day-to-day medical big data from massive IoT-based medical devices to diagnostic valuable data so that they can be feed to clinical routine. In this article, we propose a novel spatiotemporal two-streams generative adversarial network (SpGAN) as a cross-modality data processing approach to deploy the medical IoT in infarction learning. Our SpGAN remotely converts diagnostic valuable contrast-enhanced images (the “gold standard” for infarction learning, but it requires the injection of contrast agents) directly from raw nonenhanced cine MR images. This converting allows physicians to remotely perform infarction observation and analysis to break through the limitations of time and space by building a cloud computing platform of IoT-based MRI devices. Importantly, this converting offers a low-risk IoT-based manner to eliminate the potential fatal risk caused by contrast agent injection in the current infarction learning workflow. Specifically, SpGAN consists of: 1) a spatiotemporal two-stream framework as an encoding–decoding model to achieve data converting and 2) a spatiotemporal pyramid network enhances those features that are responsible to the infarction learning during encoding to improve decoding performance. Real IoT-based remote diagnosis experiments performed on 230 patients demonstrate that SpGAN provides high-quality converted images for infarction learning and promotes the in-depth application and deployment of IoT in the medical field. Chenchu Xu, Zhifan Gao, Dong Zhang 0009, Jinglin Zhang 0003, Lei Xu 0037, Shuo Li 0001 |
IEEE Internet Things J. | 6 |
| 2021 | Weakly-Supervised teacher-Student network for liver tumor segmentation from non-enhanced images
Dong Zhang 0009, Bo Chen 0013, Jaron Chong, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2021 | Unifying neural learning and symbolic reasoning for spinal medical report generation
Zhongyi Han, Benzheng Wei, Xiaoming Xi, Bo Chen 0013, Yilong Yin, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2021 | Meta grayscale adaptive network for 3D integrated renal structures segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Rongjun Ge, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 12 |
| 2021 | APRIL: Anatomical prior-guided reinforcement learning for accurate carotid lumen diameter and intima-media thickness measurement
Sheng Lian, Zhiming Luo, Shaozi Li, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2021 | Estimating dual-energy CT imaging from single-energy CT data with material decomposition convolutional neural network
Tianling Lyu, Wei Zhao 0029, Yinsu Zhu, Zhan Wu, Yikun Zhang 0001, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001, Lei Xing 0001 |
Medical Image Anal. | 8 |
| 2021 | Evaluation and comparison of accurate automated spinal curvature estimation algorithms with spinal anterior-posterior X-Ray images: The AASCE2019 challenge
Liansheng Wang 0002, Kailin Chen, Dalong Cheng, Florian Dubost, Benjamin Collery, Bidur Khanal, Bishesh Khanal, Rong Tao, Shangliang Xu, Upasana Upadhyay Bharadwaj, Zhusi Zhong, Jie Li 0001, Shuo Li 0001 |
Medical Image Anal. | 17 |
| 2021 | ELNet: Automatic classification and segmentation for esophageal lesions using convolutional neural network
Zhan Wu, Rongjun Ge, Minli Wen, Gaoshuang Liu, Yang Chen 0008, Pinzheng Zhang, Xiaopu He, Jie Hua 0004, Limin Luo 0001, Shuo Li 0001 |
Medical Image Anal. | 10 |
| 2021 | Synthesis of gadolinium-enhanced liver tumors on nonenhanced liver MR images using pixel-level graph reinforcement learning
Chenchu Xu, Dong Zhang 0009, Jaron Chong, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2021 | Sequential conditional reinforcement learning for simultaneous vertebral body detection and segmentation with modeling the spine anatomy
Dong Zhang 0009, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2021 | OF-UMRN: Uncertainty-guided multitask regression network aided by optical flow for fully automated comprehensive analysis of carotid artery
Chengqian Zhao, Dengwang Li, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2021 | United adversarial learning for liver tumor segmentation and detection of multi-modality non-contrast MRI
Jianfeng Zhao 0004, Dengwang Li, Xiaojiao Xiao, Fabio Accorsi, Harry Marshall, Tyler Cossetto, Dongkeun Kim, Daniel McCarthy, Cameron Dawson, Stefan Knezevic, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 12 |
| 2021 | Automatic vertebrae recognition from arbitrary spine MRI images by a category-Consistent self-calibration detection framework
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2021 | Estimating Functional Connectivity by Integration of Inherent Brain Function Activity Pattern PriorsabstractBrain functional connectivity (FC) has shown great potential in becoming biomarkers of brain status. However, the problem of accurately estimating FC from complex-noisy fMRI time series remains unsolved. Usually, a regularization function is more appropriate in fitting the real inherent properties of the brain function activity pattern, which can further limit noise interference to improve the accuracy of the estimated result. Recently, the neuroscientists widely suggested that the inherent brain function activity pattern indicates sparse, modular and overlapping topology. However, previous studies have never considered this factual characteristic. Thus, we propose a novel method by integration of these inherent brain function activity pattern priors to estimate FC. Extensive experiments on synthetic data demonstrate that our method can more accurately estimate the FC than previous. Then, we applied the estimated FC to predict the symptom severity of depressed patients, the symptom severity is related to subtle abnormal changes in the brain function activity, a more accurate FC can more effectively capture the subtle abnormal brain function activity changes. As results, our method better than others with a higher correlation coefficient of 0.4201. Moreover, the overlapping probability of each brain region can be further explored by the proposed method. Zonglei Zhen, Xia Wu 0001, Shuo Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Thanka Mural Inpainting Based on Multi-Scale Adaptive Partial Convolution and Stroke-Like MaskabstractThanka murals are important cultural heritages of Tibet, but many precious murals were damaged during history. Thanka mural restoration is very important for the protection of Tibetan cultural heritage. Partial convolution has great potential for Thanka mural restoration due to its outstanding performance for inpainting irregular holes. However, three challenges prevent the existing partial convolution-based methods from solving Thanka restoration problems: 1) the features of multi-scale objects in Thanka murals cannot be extracted correctly because of single-scale partial convolution; 2) the stroke-like Thanka inpainting mode cannot be effectively simulated and learned by existing rectangular or arbitrary masks; and 3) the original content of damaged Thanka murals cannot be restored. To resolve these problems, we propose a Thanka mural inpainting method based on multi-scale adaptive partial convolution and stroke-like masks. The proposed method consists of three parts: 1) a kernel-level multi-scale adaptive partial convolution (MAPConv) to accurately discriminate valid pixels from invalid pixels, and to extract the features of multi-scale objects; 2) a parameter-configurable stroke-like mask generation method to simulate and learn the stroke-like Thanka inpainting mode; and 3) a 2-phase learning framework based on MAPConv Unet and different loss functions to restore the original content of Thanka murals. Experiments on both simulated and real damages of Thanka murals demonstrated that our approach works well on a small dataset (N=2780), generates realistic mural content, and restores the damaged Thanka murals with high speed (600 ms for multiple holes in 512×512 images). The proposed end-to-end method can be applied to other small datasets-based inpainting tasks. Nianyi Wang, Weilan Wang, Wenjin Hu 0001, Aaron Fenster, Shuo Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Quantifying Axial Spine Images Using Object-Specific Bi-Path NetworkabstractAutomatic estimation of indices from medical images is the main goal of computer-aided quantification (CADq), which speeds up diagnosis and lightens the workload of radiologists. Deep learning technique is a good choice for implementing CADq. Usually, to acquire high-accuracy quantification, specific network architecture needs to be designed for a given CADq task. In this study, considering that the target organs are the intervertebral disc and the dural sac, we propose an object-specific bi-path network (OSBP-Net) for axial spine image quantification. Each path of the OSBP-Net comprises a shallow feature extraction layer (SFE) and a deep feature extraction sub-network (DFE). The SFEs use different convolution strides because the two target organs have different anatomical sizes. The DFEs use average pooling for downsampling based on the observation that the target organs have lower intensity than the background. In addition, an inter-path dissimilarity constraint is proposed and applied to the output of the SFEs, taking into account that the activated regions in the feature maps of two paths should be different theoretically. An inter-index correlation regularization is introduced and applied to the output of the DFEs based on the observation that the diameter and area of the same object express an approximately linear relation. The prediction results of OSBP-Net are compared to several state-of-the-art machine learning-based CADq methods. The comparison reveals that the proposed methods precede other competing methods extensively, indicating its great potential for spine CADq. Liyan Lin, Wei Yang 0006, Shumao Pang, Zhihai Su, Shuo Li 0001, Qianjin Feng 0003, Bo Chen 0013 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Left Ventricle Quantification Challenge: A Comprehensive Comparison and Evaluation of Segmentation and Regression for Mid-Ventricular Short-Axis Cardiac MR DataabstractAutomatic quantification of the left ventricle (LV) from cardiac magnetic resonance (CMR) images plays an important role in making the diagnosis procedure efficient, reliable, and alleviating the laborious reading work for physicians. Considerable efforts have been devoted to LV quantification using different strategies that include segmentation-based (SG) methods and the recent direct regression (DR) methods. Although both SG and DR methods have obtained great success for the task, a systematic platform to benchmark them remains absent because of differences in label information during model learning. In this paper, we conducted an unbiased evaluation and comparison of cardiac LV quantification methods that were submitted to the Left Ventricle Quantification (LVQuan) challenge, which was held in conjunction with the Statistical Atlases and Computational Modeling of the Heart (STACOM) workshop at the MICCAI 2018. The challenge was targeted at the quantification of 1) areas of LV cavity and myocardium, 2) dimensions of the LV cavity, 3) regional wall thicknesses (RWT), and 4) the cardiac phase, from mid-ventricle short-axis CMR images. First, we constructed a public quantification dataset Cardiac-DIG with ground truth labels for both the myocardium mask and these quantification targets across the entire cardiac cycle. Then, the key techniques employed by each submission were described. Next, quantitative validation of these submissions were conducted with the constructed dataset. The evaluation results revealed that both SG and DR methods can offer good LV quantification performance, even though DR methods do not require densely labeled masks for supervision. Among the 12 submissions, the DR method LDAMT offered the best performance, with a mean estimation error of 301 mm2for the two areas, 2.15 mm for the cavity dimensions, 2.03 mm for RWTs, and a 9.5% error rate for the cardiac phase classification. Three of the SG methods also delivered comparable performances. Finally, we discussed the advantages and disadvantages of SG and DR methods, as well as the unsolved problems in automatic cardiac quantification for clinical practice applications. Wufeng Xue, Jiahui Li 0005, Eric Kerfoot, James R. Clough, Ilkay Öksüz, Vicente Grau, Fumin Guo, Matthew Ng, Xiang Li 0001, Quanzheng Li, Lihong Liu, Ilias Grinias, Georgios Tziritas, Angélica Atehortúa, Mireille Garreau, Yeonggul Jang, Alejandro Debus, Enzo Ferrante, Guanyu Yang 0001, Tiancong Hua, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 25 |
| 2021 | Multitask Learning for Estimating Multitype Cardiac Indices in MRI and CT Based on Adversarial Reverse MappingabstractThe estimation of multitype cardiac indices from cardiac magnetic resonance imaging (MRI) and computed tomography (CT) images attracts great attention because of its clinical potential for comprehensive function assessment. However, the most exiting model can only work in one imaging modality (MRI or CT) without transferable capability. In this article, we propose the multitask learning method with the reverse inferring for estimating multitype cardiac indices in MRI and CT. Different from the existing forward inferring methods, our method builds a reverse mapping network that maps the multitype cardiac indices to cardiac images. The task dependencies are then learned and shared to multitask learning networks using an adversarial training approach. Finally, we transfer the parameters learned from MRI to CT. A series of experiments were conducted in which we first optimized the performance of our framework via ten-fold cross-validation of over 2900 cardiac MRI images. Then, the fine-tuned network was run on an independent data set with 2360 cardiac CT images. The results of all the experiments conducted on the proposed adversarial reverse mapping show excellent performance in estimating multitype cardiac indices. Chengjin Yu, Zhifan Gao, Weiwei Zhang 0006, Guang Yang 0006, Shu Zhao 0005, Heye Zhang, Yanping Zhang 0001, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2020 | OF-MSRN: Optical Flow-Auxiliary Multi-Task Regression Network for Direct Quantitative Measurement, Segmentation and Motion EstimationabstractComprehensively analyzing the carotid artery is critically significant to diagnosing and treating cardiovascular diseases. The object of this work is to simultaneously achieve direct quantitative measurement and automated segmentation of the lumen diameter and intima-media thickness as well as the motion estimation of the carotid wall. No work has simultaneously achieved the comprehensive analysis of carotid artery due to three intractable challenges: 1) Tiny intima-media is more challenging to measure and segment; 2) Artifact generated by radial motion restrict the accuracy of measurement and segmentation; 3) Occlusions on diseased carotid walls generate dynamic complexity and indeterminacy. In this paper, we propose a novel optical flow-auxiliary multi-task regression network named OF-MSRN to overcome these challenges. We concatenate multi-scale features to a regression network to simultaneously achieve measurement and segmentation, which makes full use of the potential correlation between the two tasks. More importantly, we creatively explore an optical flow auxiliary module to take advantage of the co-promotion of segmentation and motion estimation to overcome the restrictions of the radial motion. Besides, we evaluate consistency between forward and backward optical flow to improve the accuracy of motion estimation of the diseased carotid wall. Extensive experiments on US sequences of 101 patients demonstrate the superior performance of OF-MSRN on the comprehensive analysis of the carotid artery by utilizing the dual optimization of the optical flow auxiliary module. Chengqian Zhao, Dengwang Li, Shuo Li 0001 |
AAAI | 4 |
| 2020 | Deep Complementary Joint Model for Complex Scene Registration and Few-Shot Segmentation on Medical Images
Yuting He 0001, Guanyu Yang 0001, Youyong Kong, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux, Jean-Louis Dillenseger, Shuo Li 0001 |
ECCV (18) | 9 |
| 2020 | EGDCL: An Adaptive Curriculum Learning Framework for Unbiased Glaucoma Diagnosis
Rongchang Zhao, Xuanlin Chen, Zailiang Chen 0001, Shuo Li 0001 |
ECCV (21) | 4 |
| 2020 | Multi-vertebrae Segmentation from Arbitrary Spine MR Images Under Global View
Heyou Chang, Yang Chen 0008, Shuo Li 0001 |
MICCAI (6) | 5 |
| 2020 | Segmentation of Paraspinal Muscles at Varied Lumbar Spinal Levels by Explicit Saliency-Aware Learning
Haotian Shen, Bo Chen 0013, Shuo Li 0001 |
MICCAI (6) | 5 |
| 2020 | Mt-UcGAN: Multi-task Uncertainty-Constrained GAN for Joint Segmentation, Quantification and Uncertainty Estimation of Renal Tumors on CT
Yanan Ruan, Dengwang Li, Harry Marshall, Timothy Miao, Tyler Cossetto, Ian Chan, Omar Daher, Fabio Accorsi, Aashish Goela, Shuo Li 0001 |
MICCAI (4) | 10 |
| 2020 | Temporal-Consistent Segmentation of Echocardiography with Co-learning from Appearance and Shape
Hongrong Wei, Yiqin Cao, Yongjin Zhou 0002, Wufeng Xue, Dong Ni 0001, Shuo Li 0001 |
MICCAI (2) | 7 |
| 2020 | Discriminative Dictionary-Embedded Network for Comprehensive Vertebrae Tumor Diagnosis
Heyou Chang, Xi Wu 0004, Shuo Li 0001 |
MICCAI (6) | 5 |
| 2020 | Damage Sensitive and Original Restoration Driven Thanka Mural Inpainting
Nianyi Wang, Weilan Wang, Wenjin Hu 0001, Aaron Fenster, Shuo Li 0001 |
PRCV (1) | 5 |
| 2020 | Coarse-to-fine classification for diabetic retinopathy grading using convolutional neural network
Zhan Wu, Gonglei Shi, Yang Chen 0008, Xinjian Chen 0001, Gouenou Coatrieux, Jian Yang 0009, Limin Luo 0001, Shuo Li 0001 |
Artif. Intell. Medicine | 9 |
| 2020 | Vessel Structure Extraction using Constrained Minimal Path Propagation
Guanyu Yang 0001, Tianling Lv, Yunpeng Shen, Shuo Li 0001, Jian Yang 0009, Yang Chen 0008, Huazhong Shu, Limin Luo 0001, Jean-Louis Coatrieux |
Artif. Intell. Medicine | 4 |
| 2020 | Recursive narrative alignment for movie narrating
Zhongyi Han, Hongbo Wu, Benzheng Wei, Yilong Yin, Shuo Li 0001 |
Sci. China Inf. Sci. | 5 |
| 2020 | Simultaneous left atrium anatomy and scar segmentations via deep learning in multiview information with attentionabstractThree-dimensional late gadolinium enhanced (LGE) cardiac MR (CMR) of left atrial scar in patients with atrial fibrillation (AF) has recently emerged as a promising technique to stratify patients, to guide ablation therapy and to predict treatment success. This requires a segmentation of the high intensity scar tissue and also a segmentation of the left atrium (LA) anatomy, the latter usually being derived from a separate bright-blood acquisition. Performing both segmentations automatically from a single 3D LGE CMR acquisition would eliminate the need for an additional acquisition and avoid subsequent registration issues. In this paper, we propose a joint segmentation method based on multiview two-task (MVTT) recursive attention model working directly on 3D LGE CMR images to segment the LA (and proximal pulmonary veins) and to delineate the scar on the same dataset. Using our MVTT recursive attention model, both the LA anatomy and scar can be segmented accurately (mean Dice score of 93% for the LA anatomy and 87% for the scar segmentations) and efficiently (∼0.27 s to simultaneously segment the LA anatomy and scars directly from the 3D LGE CMR dataset with 60–68 2D slices). Compared to conventional unsupervised learning and other state-of-the-art deep learning based methods, the proposed MVTT model achieved excellent results, leading to an automatic generation of a patient-specific anatomical model combined with scar segmentation for patients in AF. Guang Yang 0006, Jun Chen 0030, Zhifan Gao, Shuo Li 0001, Hao Ni 0001, Elsa D. Angelini, Tom Wong, Raad Mohiaddin, Eva Nyktari, Rick Wage, Lei Xu 0037, Yanping Zhang 0001, Xiuquan Du, Heye Zhang, David N. Firmin, Jennifer Keegan |
Future Gener. Comput. Syst. | 4 |
| 2020 | MMCL-Net: Spinal disease diagnosis in global mode using progressive multi-task joint learning
Yanfei Hong, Benzheng Wei, Zhongyi Han, Xiang Li 0114, Yuanjie Zheng, Shuo Li 0001 |
Neurocomputing | 6 |
| 2020 | Non-rigid retinal image registration using an unsupervised structure-driven regression network
Beiji Zou 0001, Zhiyou He, Rongchang Zhao, Chengzhang Zhu, Wangmin Liao, Shuo Li 0001 |
Neurocomputing | 6 |
| 2020 | Trustful Internet of Surveillance Things Based on Deeply Represented Visual Co-Saliency DetectionabstractTrustful Internet of Things (IoT) plays an important role in smart cities. The trust information in surveillance data motivates the analysis of images from numerous IoT devices. Saliency detection is a fundamental step in surveillance data analysis for providing help to the subsequent tasks, but unsuitable to IoT applications owing to the neglect of image similarity and difference from diverse IoT devices. To solve this problem, we enable the co-saliency detection in IoT, which detects the common and salient foreground regions in the group surveillance images. The main contributions include: 1) enable a multistage context perception scheme to efficiently extract the contextual information corresponding to different-size receptive fields in the single image; 2) construct a two-path information propagation to extract the interimage similarity and difference from the high-level image feature representations of the group images; and 3) propose the stage-wise refinement to allocate the label information to different parts of the network for helping the network to learn the enriched semantically common knowledge. The extensive experiments performed on three public data sets can demonstrate the effectiveness of our approach and its superiority to four state-of-the-art co-saliency detection methods. Zhifan Gao, Chenchu Xu, Heye Zhang, Shuo Li 0001, Victor Hugo C. de Albuquerque |
IEEE Internet Things J. | 4 |
| 2020 | Deep Atlas Network for Efficient 3D Left Ventricle Segmentation on Echocardiography
Suyu Dong, Gongning Luo, Clara M. Tam, Wei Wang 0169, Kuanquan Wang, Shaodong Cao, Bo Chen 0013, Henggui Zhang, Shuo Li 0001 |
Medical Image Anal. | 9 |
| 2020 | An integrated deep learning framework for joint segmentation of blood pool and myocardium
Xiuquan Du, Yuhui Song, Yueguo Liu, Yanping Zhang 0001, Heng Liu 0003, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2020 | Dense biased networks with deep priori anatomy and hard region adaptation: Semi-supervised learning for fine renal artery segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 14 |
| 2020 | Commensal correlation network between segmentation and direct area estimation for bi-ventricle quantification
Gongning Luo, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Shaodong Cao, Clara M. Tam, Henggui Zhang, Joanne Howey, Pavlo Ohorodnyk, Shuo Li 0001 |
Medical Image Anal. | 10 |
| 2020 | Dynamically constructed network with error correction for accurate ventricle volume estimation
Gongning Luo, Wei Wang 0169, Clara M. Tam, Kuanquan Wang, Shaodong Cao, Henggui Zhang, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 8 |
| 2020 | MB-FSGAN: Joint segmentation and quantification of kidney tumor on CT by the multi-branch feature sharing generative adversarial network
Yanan Ruan, Dengwang Li, Harry Marshall, Timothy Miao, Tyler Cossetto, Ian Chan, Omar Daher, Fabio Accorsi, Aashish Goela, Shuo Li 0001 |
Medical Image Anal. | 10 |
| 2020 | Holistic multitask regression network for multiapplication shape regression segmentation
Clara M. Tam, Dong Zhang 0009, Bo Chen 0013, Terry M. Peters, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2020 | SDAE-GAN: Enable high-dimensional pathological images in liver cancer survival prediction with a policy gradient based data augmentation method
Hejun Wu, Yeong Poh Sheng, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2020 | Segmentation and quantification of infarction without contrast agents via spatiotemporal generative adversarial learning
Chenchu Xu, Joanne Howey, Pavlo Ohorodnyk, Mike Roth, Heye Zhang, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2020 | Contrast agent-free synthesis and segmentation of ischemic heart disease images using progressive sequential causal GANs
Chenchu Xu, Lei Xu 0037, Pavlo Ohorodnyk, Mike Roth, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2020 | Multi-indices quantification of optic nerve head in fundus image via multitask collaborative learning
Rongchang Zhao, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2020 | Tripartite-GAN: Synthesizing liver contrast-enhanced MRI to improve tumor detection
Jianfeng Zhao 0004, Dengwang Li, Zahra Kassam, Joanne Howey, Jaron Chong, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2020 | Multiple Axial Spine Indices Estimation via Dense Enhancing Network With Cross-Space Distance-Preserving RegularizationabstractAutomatic estimation of axial spine indices is clinically desired for various spine computer aided procedures, such as disease diagnosis, therapeutic evaluation, pathophysiological understanding, risk assessment, and biomechanical modeling. Currently, the spine indices are manually measured by physicians, which is time-consuming and laborious. Even worse, the tedious manual procedure might result in inaccurate measurement. To deal with this problem, in this paper, we aim at developing an automatic method to estimate multiple indices from axial spine images. Inspired by the success of deep learning for regression problems and the densely connected network for image classification, we propose a dense enhancing network (DE-Net) which uses the dense enhancing blocks (DEBs) as its main body, where a feature enhancing layer is added to each of the bypass in a dense block. The DEB is designed to enhance discriminative feature embedding from the intervertebral disc and the dural sac areas. In addition, the cross-space distance-preserving regularization (CSDPR), which enforces consistent inter-sample distances between the output and the label spaces, is proposed to regularize the loss function of the DE-Net. To train and validate the proposed method, we collected 895 axial spine MRI images from 143 subjects and manually measured the indices as the ground truth. The results show that all deep learning models obtain very small prediction errors, and the proposed DE-Net with CSDPR acquires the smallest error among all methods, indicating that our method has great potential for spine computer aided procedures. Liyan Lin, Shumao Pang, Zhihai Su, Shuo Li 0001, Qianjin Feng 0003, Bo Chen 0013 |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | MRLN: Multi-Task Relational Learning Network for MRI Vertebral Localization, Identification, and SegmentationabstractMagnetic resonance imaging (MRI) vertebral localization, identification, and segmentation are important steps in the automatic analysis of spines. Due to the similar appearances of vertebrae, the accurate segmentation, localization, and identification of vertebrae remain challenging. Previous methods solved the three tasks independently, ignoring the intrinsic correlation among them. In this paper, we propose a multi-task relational learning network (MRLN) that utilizes both the relationships between vertebrae and the relevance of the three tasks. A dilation convolution group is used to expand the receptive field, and LSTM(Long Short-Term Memory) to learn the prior knowledge of the order relationship between the vertebral bodies. We introduce a co-attention module to learn the correlation information, localization-guided segmentation attention(LGSA) and segmentation-guided localization attention(SGLA), in the decoder stage of segmentation and localization tasks. Learning two tasks simultaneously as well as the correlation between tasks can not only avoid the overfitting of a single task but also correct each other. To avoids the cumbersome weight adjustment for different tasks loss functions, we formulated a novel XOR loss that provides a direct evaluation criterion for the localization relationship of the semantic location regression and semantic segmentation. This method was evaluated on a dataset which includes multiple MRI modalities (T1 and T2), various fields of view. Experimental results demonstrate that both of the co-attention and XOR loss work outperforms the most recent state of art. Xiaoyan Xiao, Zhi Liu 0004, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Direct Cup-to-Disc Ratio Estimation for Glaucoma Screening via Semi-Supervised LearningabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. The Cup-to-Disc Ratio (CDR) serves as the most important indicator for glaucoma screening and plays a significant role in clinical screening and early diagnosis of glaucoma. In general, obtaining CDR is subjected to measuring on manually or automatically segmented optic disc and cup. Despite great efforts have been devoted, obtaining CDR values automatically with high accuracy and robustness is still a great challenge due to the heavy overlap between optic cup and neuroretinal rim regions. In this paper, a direct CDR estimation method is proposed based on the well-designed semi-supervised learning scheme, in which CDR estimation is formulated as a general regression problem while optic disc/cup segmentation is cancelled. The method directly regresses CDR value based on the feature representation of optic nerve head via deep learning technique while bypassing intermediate segmentation. The scheme is a two-stage cascaded approach comprised of two phases: unsupervised feature representation of fundus image with a convolutional neural networks (MFPPNet) and CDR value regression by random forest regressor. The proposed scheme is validated on the challenging glaucoma dataset Direct-CSU and public ORIGA, and the experimental results demonstrate that our method can achieve a lower average CDR error of 0.0563 and a higher correlation of around 0.726 with measurement before manual segmentation of optic disc/cup by human experts. Our estimated CDR values are also tested for glaucoma screening, which achieves the areas under curve of 0.905 on dataset of 421 fundus images. The experiments show that the proposed method is capable of state-of-the-art CDR estimation and satisfactory glaucoma screening with calculated CDR value. Rongchang Zhao, Xuanlin Chen, Xiyao Liu 0001, Zailiang Chen 0001, Fan Guo 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Privileged Modality Distillation for Vessel Border Detection in Intracoronary ImagingabstractIntracoronary imaging is a crucial imaging technology in coronary disease diagnosis as it visualizes the internal tissue morphologies of coronary arteries. Vessel border detection in intracoronary images (VBDI) is desired because it can help the succeeding procedures of computer-aided disease diagnosis. However, existing VDBI methods suffer from the challenge of vessel-environment variability (i.e. high intra- and inter-subject diversity of vessels and their surrounding tissues appeared in images). This challenge leads to the ineffectiveness in the vessel region representation for hand-crafted features, in the receptive field extraction for deeply-represented features, as well as performance suppression derived from clinical data limitation. To solve this challenge, we propose a novel privileged modality distillation (PMD) framework for VBDI. PMD transforms the single-input-single-task (SIST) learning problem in the single-mode VBDI to a multiple-input-multiple-task (MIMT) problem by using the privileged image modality to help the learning model in the target modality. This learns the enriched high-level knowledge with similar semantics and generalizes PMD on diversity-increased low-level image features for improving the model adaptation to diverse vessel environments. Moreover, PMD refines MIMT to SIST by distilling the learned knowledge from multiple to one modality. This eliminates the reliance on privileged modality in the test phase, and thus enables the applicability to each of different intracoronary modalities. A structure-deformable neural network is proposed as an elaborately-designed implementation of PMD. It expands a conventional SIST network structure to the MIMT structure, and then recovers it to the final SIST structure. The PMD is validated on intravascular ultrasound imaging and optical coherence tomography imaging. One modality is the target, and the other one can be considered as the privileged modality owing to their semantic relatedness. The experiments show that our PMD is effective in VBDI (e.g. the Dice index is larger than 0.95), as well as superior to six state-of-the-art VBDI methods. Zhifan Gao, Jonathan Chung 0002, Mohamed Abdelrazek 0002, Stephanie Leung, William Kongto Hau, Zhanchao Xian, Heye Zhang, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2020 | K-Net: Integrate Left Ventricle Segmentation and Direct Quantification of Paired Echo SequenceabstractThe integration of segmentation and direct quantification on the left ventricle (LV) from the paired apical views(i.e., apical 4-chamber and 2-chamber together) echo sequence clinically achieves the comprehensive cardiac assessment: multiview segmentation for anatomical morphology, and multidimensional quantification for contractile function. Direct quantification of LV, i.e., to automatically quantify multiple LV indices directly from the image via task-aware feature representation and regression, avoids accumulative error from the inter-step target. This integration sequentially makes a stereoscopical reflection of cardiac activity jointly from the paired orthogonal cross views sequences, overcoming limited observation with a single plane. We propose a K-shaped Unified Network (K-Net), the first end-to-end framework to simultaneously segment LV from apical 4-chamber and 2-chamber views, and directly quantify LV from major- and minor-axis dimensions (1D), area (2D), and volume (3D), in sequence. It works via four components: 1) the K-Net architecture with the Attention Junction enables heterogeneous tasks learning of segmentation task of pixel-wise classification, and direct quantification task of image-wise regression, by interactively introducing the information from segmentation to jointly promote spatial attention map to guide quantification focusing on LV-related region, and transferring quantification feedback to make global constraint on segmentation; 2) the Bi-ResLSTMs distributed in K-Net layer-by-layer hierarchically extract spatial-temporal information in echo sequence, with bidirectional recurrent and short-cut connection to model spatial-temporal information among all frames; 3) the Information Valve tailing the Bi-ResLSTMs selectively exchanges information among multiple views, by stimulating complementary information and suppressing redundant information to make the efficient cross-flow for each view; 4) the Evolution Loss comprehensively guides sequential data learning, with static constraint for frame values, and dynamic constraint for inter-frame value changes. The experiments show that our K-Net gains high performance with a Dice coefficient up to 91.44% and a mean absolute error of the major-axis dimension down to 2.74mm, which reveal its clinical potential. Rongjun Ge, Guanyu Yang 0001, Yang Chen 0008, Limin Luo 0001, Junyi Ren, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Image Projection Network: 3D to 2D Image Segmentation in OCTA ImagesabstractWe present an image projection network (IPN), which is a novel end-to-end architecture and can achieve 3D-to-2D image segmentation in optical coherence tomography angiography (OCTA) images. Our key insight is to build a projection learning module (PLM) which uses a unidirectional pooling layer to conduct effective features selection and dimension reduction concurrently. By combining multiple PLMs, the proposed network can input 3D OCTA data, and output 2D segmentation results such as retinal vessel segmentation. It provides a new idea for the quantification of retinal indicators: without retinal layer segmentation and without projection maps. We tested the performance of our network for two crucial retinal image segmentation issues: retinal vessel (RV) segmentation and foveal avascular zone (FAZ) segmentation. The experimental results on 316 OCTA volumes demonstrate that the IPN is an effective implementation of 3D-to-2D segmentation networks, and the uses of multi-modality information and volumetric information make IPN perform better than the baseline methods. Mingchao Li 0002, Yerui Chen, Zexuan Ji, Keren Xie, Songtao Yuan, Qiang Chen 0004, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Direct Quantification of Coronary Artery Stenosis Through Hierarchical Attentive Multi-View LearningabstractQuantification of coronary artery stenosis on X-ray angiography (XRA) images is of great importance during the intraoperative treatment of coronary artery disease. It serves to quantify the coronary artery stenosis by estimating the clinical morphological indices, which are essential in clinical decision making. However, stenosis quantification is still a challenging task due to the overlapping, diversity and small-size region of the stenosis in the XRA images. While efforts have been devoted to stenosis quantification through low-level features, these methods have difficulty in learning the real mapping from these features to the stenosis indices. These methods are still cumbersome and unreliable for the intraoperative procedures due to their two-phase quantification, which depends on the results of segmentation or reconstruction of the coronary artery. In this work, we are proposing a hierarchical attentive multi-view learning model (HEAL) to achieve a direct quantification of coronary artery stenosis, without the intermediate segmentation or reconstruction. We have designed a multi-view learning model to learn more complementary information of the stenosis from different views. For this purpose, an intra-view hierarchical attentive block is proposed to learn the discriminative information of stenosis. Additionally, a stenosis representation learning module is developed to extract the multi-scale features from the keyframe perspective for considering the clinical workflow. Finally, the morphological indices are directly estimated based on the multi-view feature embedding. Extensive experiment studies on clinical multi-manufacturer dataset consisting of 228 subjects show the superiority of our HEAL against nine comparing methods, including direct quantification methods and multi-view learning methods. The experimental results demonstrate the better clinical agreement between the ground truth and the prediction, which endows our proposed method with a great potential for the efficient intraoperative treatment of coronary artery disease. Dong Zhang 0012, Guang Yang 0006, Shu Zhao 0005, Yanping Zhang 0001, Dhanjoo N. Ghista, Heye Zhang, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Weakly-Supervised Simultaneous Evidence Identification and Segmentation for Automated Glaucoma DiagnosisabstractEvidence identification, optic disc segmentation and automated glaucoma diagnosis are the most clinically significant tasks for clinicians to assess fundus images. However, delivering the three tasks simultaneously is extremely challenging due to the high variability of fundus structure and lack of datasets with complete annotations. In this paper, we propose an innovative Weakly-Supervised Multi-Task Learning method (WSMTL) for accurate evidence identification, optic disc segmentation and automated glaucoma diagnosis. The WSMTL method only uses weak-label data with binary diagnostic labels (normal/glaucoma) for training, while obtains pixel-level segmentation mask and diagnosis for testing. The WSMTL is constituted by a skip and densely connected CNN to capture multi-scale discriminative representation of fundus structure; a well-designed pyramid integration structure to generate high-resolution evidence map for evidence identification, in which the pixels with higher value represent higher confidence to highlight the abnormalities; a constrained clustering branch for optic disc segmentation; and a fully-connected discriminator for automated glaucoma diagnosis. Experimental results show that our proposed WSMTL effectively and simultaneously delivers evidence identification, optic disc segmentation (89.6% TP Dice), and accurate glaucoma diagnosis (92.4% AUC). This endows our WSMTL a great potential for the effective clinical assessment of glaucoma. Rongchang Zhao, Wangmin Liao, Beiji Zou 0001, Zailiang Chen 0001, Shuo Li 0001 |
AAAI | 5 |
| 2019 | Context-Aware Inductive Bias Learning for Vessel Border Detection in Multi-modal Intracoronary Imaging
Zhifan Gao, Shuo Li 0001 |
MICCAI (2) | 2 |
| 2019 | Stereo-Correlation and Noise-Distribution Aware ResVoxGAN for Dense Slices Reconstruction and Noise Reduction in Thick Low-Dose CT
Rongjun Ge, Guanyu Yang 0001, Chenchu Xu, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001 |
MICCAI (6) | 6 |
| 2019 | DPA-DenseBiasNet: Semi-supervised 3D Fine Renal Artery Segmentation with Dense Biased Network and Deep Priori Anatomy
Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
MICCAI (6) | 13 |
| 2019 | Recurrent Aggregation Learning for Multi-view Echocardiographic Sequences Segmentation
Ming Li 0005, Weiwei Zhang 0006, Guang Yang 0006, Chengjia Wang, Heye Zhang, Huafeng Liu 0003, Shuo Li 0001 |
MICCAI (2) | 8 |
| 2019 | A Deep Reinforcement Learning Framework for Frame-by-Frame Plaque Tracking on Intravascular Optical Coherence Tomography Image
Gongning Luo, Suyu Dong, Kuanquan Wang, Dong Zhang 0009, Yue Gao 0002, Xin Chen 0025, Henggui Zhang, Shuo Li 0001 |
MICCAI (1) | 8 |
| 2019 | Radiomics-guided GAN for Segmentation of Liver Tumor Without Contrast Agents
Xiaojiao Xiao, Juanjuan Zhao 0002, Yan Qiang 0001, Jaron Chong, Xiaotang Yang, Ntikurako Guy-Fernand Kazihise, Bo Chen 0013, Shuo Li 0001 |
MICCAI (2) | 8 |
| 2019 | Direct Quantification for Coronary Artery Stenosis Using Multiview Learning
Dong Zhang 0012, Guang Yang 0006, Shu Zhao 0005, Yanping Zhang 0001, Heye Zhang, Shuo Li 0001 |
MICCAI (2) | 6 |
| 2019 | Multi-index Optic Disc Quantification via MultiTask Ensemble Learning
Rongchang Zhao, Zailiang Chen 0001, Xiyao Liu 0001, Beiji Zou 0001, Shuo Li 0001 |
MICCAI (1) | 5 |
| 2019 | Automatic Vertebrae Recognition from Arbitrary Spine MRI Images by a Hierarchical Self-calibration Detection Framework
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
MICCAI (4) | 4 |
| 2019 | A spatial-aware joint optic disc and cup segmentation method
Qing Liu 0003, Xiaopeng Hong, Shuo Li 0001, Zailiang Chen 0001, Guoying Zhao 0001, Beiji Zou 0001 |
Neurocomputing | 3 |
| 2019 | Learning the implicit strain reconstruction in ultrasound elastography using privileged information
Zhifan Gao, Sitong Wu, Zhi Liu 0004, Jianwen Luo 0001, Heye Zhang, Mingming Gong, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2019 | PV-LVNet: Direct left ventricle multitype indices estimation from 2D echocardiograms of paired apical views with deep neural networks
Rongjun Ge, Guanyu Yang 0001, Yang Chen 0008, Limin Luo 0001, Heye Zhang, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2019 | Direct automated quantitative measurement of spine by cascade amplifier regression network with manifold regularization
Shumao Pang, Zhihai Su, Stephanie Leung, Ilanit Ben Nachum, Bo Chen 0013, Qianjin Feng 0003, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2019 | Accurate automated Cobb angles estimation using multi-view extrapolation net
Liansheng Wang 0002, Qiuhao Xu, Stephanie Leung, Jonathan Chung 0002, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2019 | Automatic spondylolisthesis grading from MRIs across modalities using faster adversarial recognition network
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2019 | Direct Segmentation-Based Full Quantification for Left Ventricle via Deep Multi-Task Regression Learning NetworkabstractQuantitative analysis of the heart is extremely necessary and significant for detecting and diagnosing heart disease, yet there are still some challenges. In this study, we propose a new end-to-end segmentation-based deep multi-task regression learning model (Indices-JSQ) to make a holonomic quantitative analysis of the left ventricle (LV), which contains a segmentation network (Img2Contour) and multi-task regression network (Contour2Indices). First, Img2Contour, which contains a deep convolutional encoder-decoder module, is designed to obtain the LV contour. Then, the predicted contour is fed as input to Contour2Indices for full quantification. On the whole, we take into account the relationship between different tasks, which can serve as a complementary advantage. Meanwhile, instead of using images directly from the original dataset, we creatively use the segmented contour of the original image to estimate the cardiac indices to achieve better and more accurate results. We make experiments on MR sequences of 145 subjects and gain the experimental results of 157 mm2, 2.43 mm, 1.29 mm, and 0.87 on areas, dimensions, regional wall thicknesses, and Dice Metric, respectively. It intuitively shows that the proposed method outperforms the other state-of-the-art methods and demonstrates that our method has a great potential in cardiac MR images segmentation, comprehensive clinical assessment, and diagnosis. Xiuquan Du, Renjun Tang, Susu Yin, Yanping Zhang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | Holistic and Deep Feature Pyramids for Saliency Detection
Shizhong Dong, Zhifan Gao, Shanhui Sun, Xin Wang 0045, Ming Li 0005, Heye Zhang, Guang Yang 0006, Huafeng Liu 0003, Shuo Li 0001 |
BMVC | 9 |
| 2018 | Deep Learning intra-image and inter-images features for Co-saliency detection
Shizhong Dong, Zhifan Gao, Xi Wu 0004, Heye Zhang, Guang Yang 0006, Shuo Li 0001 |
BMVC | 8 |
| 2018 | VoxelAtlasGAN: 3D Left Ventricle Segmentation on Echocardiography with Atlas Guided Generation and Voxel-to-Voxel Discrimination
Suyu Dong, Gongning Luo, Kuanquan Wang, Shaodong Cao, Ashley Mercado, Olga Shmuilovich, Henggui Zhang, Shuo Li 0001 |
MICCAI (4) | 8 |
| 2018 | Towards Automatic Report Generation in Spine Radiology Using Weakly Supervised Framework
Zhongyi Han, Benzheng Wei, Stephanie Leung, Jonathan Chung 0002, Shuo Li 0001 |
MICCAI (4) | 5 |
| 2018 | Direct Automated Quantitative Measurement of Spine via Cascade Amplifier Regression Network
Shumao Pang, Stephanie Leung, Ilanit Ben Nachum, Qianjin Feng 0003, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2018 | Direct Reconstruction of Ultrasound Elastography Using an End-to-End Deep Neural Network
Sitong Wu, Zhifan Gao, Zhi Liu 0004, Jianwen Luo 0001, Heye Zhang, Shuo Li 0001 |
MICCAI (1) | 6 |
| 2018 | MuTGAN: Simultaneous Segmentation and Quantification of Myocardial Infarction Without Contrast Agents via Joint Adversarial Learning
Chenchu Xu, Lei Xu 0037, Gary Brahm, Heye Zhang, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2018 | Cardiac Motion Scoring with Segment- and Subject-Level Non-local Modeling
Wufeng Xue, Gary Brahm, Stephanie Leung, Olga Shmuilovich, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2018 | Automated neural foraminal stenosis grading via task-aware structural representation learning
Xiaoxu He, Stephanie Leung, James Warrington, Olga Shmuilovich, Shuo Li 0001 |
Neurocomputing | 5 |
| 2018 | Sparse regression with output correlation for cardiac ejection fraction estimation
Bin Gu 0001, Yingying Shan, Victor S. Sheng, Yuhui Zheng, Shuo Li 0001 |
Inf. Sci. | 5 |
| 2018 | Spine-GAN: Semantic segmentation of multiple spinal structures
Zhongyi Han, Benzheng Wei, Ashley Mercado, Stephanie Leung, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2018 | Automated comprehensive Adolescent Idiopathic Scoliosis assessment using MVC-Net
Hongbo Wu, Parham Rasoulinejad, Shuo Li 0001 |
Medical Image Anal. | 4 |
| 2018 | Direct delineation of myocardial infarction without contrast agents using a joint motion feature learning architecture
Chenchu Xu, Lei Xu 0037, Zhifan Gao, Heye Zhang, Yanping Zhang 0001, Xiuquan Du, Shu Zhao 0005, Dhanjoo N. Ghista, Huafeng Liu 0003, Shuo Li 0001 |
Medical Image Anal. | 11 |
| 2018 | Full left ventricle quantification via deep multitask relationships learning
Wufeng Xue, Gary Brahm, Sachin Pandey, Stephanie Leung, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2018 | Multi-Target Regression via Robust Low-Rank LearningabstractMulti-target regression has recently regained great popularity due to its capability of simultaneously learning multiple relevant regression tasks and its wide applications in data mining, computer vision and medical image analysis, while great challenges arise from jointly handling inter-target correlations and input-output relationships. In this paper, we propose Multi-layer Multi-target Regression (MMR) which enables simultaneously modeling intrinsic inter-target correlations and nonlinear input-output relationships in a general framework via robust low-rank learning. Specifically, the MMR can explicitly encode inter-target correlations in a structure matrix by matrix elastic nets (MEN); the MMR can work in conjunction with the kernel trick to effectively disentangle highly complex nonlinear input-output relationships; the MMR can be efficiently solved by a new alternating optimization algorithm with guaranteed convergence. The MMR leverages the strength of kernel methods for nonlinear feature learning and the structural advantage of multi-layer learning architectures for inter-target correlation modeling. More importantly, it offers a new multi-layer learning paradigm for multi-target regression which is endowed with high generality, flexibility and expressive ability. Extensive experimental evaluation on 18 diverse real-world datasets demonstrates that our MMR can achieve consistently high performance and outperforms representative state-of-the-art algorithms, which shows its great effectiveness and generality for multivariate prediction. Xiantong Zhen, Mengyang Yu, Xiaofei He 0001, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Robust Segmentation of Intima-Media Borders With Different Morphologies and Dynamics During the Cardiac CycleabstractSegmentation of carotid intima-media (IM) borders from ultrasound sequences is challenging because of unknown image noise and varying IM border morphologies and/or dynamics. In this paper, we have developed a state-space framework to sequentially segment the carotid IM borders in each image throughout the cardiac cycle. In this framework, an ${\mathrm{H}}_{\mathrm{\infty }}$ filter is used to solve the state-space equations, and a grayscale-derivative constraint snake is used to provide accurate measurements for the ${\mathrm{H}}_{\mathrm{\infty }}$ filter. We have evaluated the performance of our approach by comparing our segmentation results to the manually traced contours of ultrasound image sequences of three synthetic models and 156 real subjects from four medical centers. The results show that our method has a small segmentation error (lumen intima, LI: 53 $\pm\, 67\;{\mathrm{\mu }}$m; media-adventitia, MA: 57 $\pm\, 63\;{\mathrm{\mu }}$m) for synthetic and real sequences of different image characteristics, and also agrees well with the manual segmentation (LI: bias = 1.44 ${\mathrm{\mu }}$m; MA: bias = $-$3.38 ${\mathrm{\mu }}$m). Our approach can robustly segment the carotid ultrasound sequences with various IM border morphologies, dynamics, and unknown image noise. These results indicate the potential of our framework to segment IM borders for clinical diagnosis. Zhifan Gao, Heye Zhang, Yaoqin Xie, Jianwen Luo 0001, Dhanjoo N. Ghista, Zhanghong Wei, Xiaojun Bi 0004, Huahua Xiong, Chenchu Xu, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 11 |
| 2018 | Motion Tracking of the Carotid Artery Wall From Ultrasound Image Sequences: a Nonlinear State-Space ApproachabstractThe motion of the common carotid artery (CCA) wall has been established to be useful in early diagnosis of atherosclerotic disease. However, tracking the CCA wall motion from ultrasound images remains a challenging task. In this paper, a nonlinear state-space approach has been developed to track CCA wall motion from ultrasound sequences. In this approach, a nonlinear state-space equation with a time-variant control signal was constructed from a mathematical model of the dynamics of the CCA wall. Then, the unscented Kalman filter (UKF) was adopted to solve the nonlinear state transfer function in order to evolve the state of the target tissue, which involves estimation of the motion trajectory of the CCA wall from noisy ultrasound images. The performance of this approach has been validated on 30 simulated ultrasound sequences and a real ultrasound dataset of 103 subjects by comparing the motion tracking results obtained in this study to those of three state-of-the-art methods and of the manual tracing method performed by two experienced ultrasound physicians. The experimental results demonstrated that the proposed approach is highly correlated with (intra-class correlation coefficient ≥ 0.9948 for the longitudinal motion and ≥ 0.9966 for the radial motion) and well agrees (the 95% confidence interval width is 0.8871 mm for the longitudinal motion and 0.4159 mm for the radial motion) with the manual tracing method on real data and also exhibits high accuracy on simulated data (0.1161 ~ 0.1260 mm). These results appear to demonstrate the effectiveness of the proposed approach for motion tracking of the CCA wall. Zhifan Gao, Jiayuan Yang, Huahua Xiong, Heye Zhang, Xin Liu 0023, Dong Liang 0001, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2018 | Multitarget Sparse Latent RegressionabstractMultitarget regression has recently generated intensive popularity due to its ability to simultaneously solve multiple regression tasks with improved performance, while great challenges stem from jointly exploring inter-target correlations and input-output relationships. In this paper, we propose multitarget sparse latent regression (MSLR) to simultaneously model intrinsic intertarget correlations and complex nonlinear input-output relationships in one single framework. By deploying a structure matrix, the MSLR accomplishes a latent variable model which is able to explicitly encode intertarget correlations via -norm-based sparse learning; the MSLR naturally admits a representer theorem for kernel extension, which enables it to flexibly handle highly complex nonlinear input-output relationships; the MSLR can be solved efficiently by an alternating optimization algorithm with guaranteed convergence, which ensures efficient multitarget regression. Extensive experimental evaluation on both synthetic data and six greatly diverse real-world data sets shows that the proposed MSLR consistently outperforms the state-of-the-art algorithms, which demonstrates its great effectiveness for multivariate prediction. Xiantong Zhen, Mengyang Yu, Feng Zheng 0001, Ilanit Ben Nachum, Mousumi Bhaduri, David T. Laidley, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2017 | Learning Deep Match Kernels for Image-Set ClassificationabstractImage-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this paper, we propose deep match kernels (DMK) to directly measure the similarity between image sets in the match kernel framework. Specifically, we build deep local match kernels between images upon arc-cosine kernels, which can faithfully characterize the similarity between images by mimicking deep neural networks, we introduce anchors to aggregate those deep local match kernels into a global match kernel between image sets, which is learned in a supervised way by kernel alignment and therefore more discriminative. The DMK provides the first match kernel framework for image-set classification, which removes specific assumptions usually required in previous approaches and is computationally more efficient. We conduct extensive experiments on four datasets for three diverse image-set classification tasks. The DMK achieves high performance and consistently surpasses state-of-the-art methods, showing its great effectiveness for image-set classification. Haoliang Sun, Xiantong Zhen, Yuanjie Zheng, Gongping Yang 0001, Yilong Yin, Shuo Li 0001 |
CVPR | 6 |
| 2017 | Automatic Landmark Estimation for Adolescent Idiopathic Scoliosis Assessment Using BoostNet
Hongbo Wu, Parham Rasoulinejad, Shuo Li 0001 |
MICCAI (1) | 4 |
| 2017 | Direct Detection of Pixel-Level Myocardial Infarction Areas via a Deep-Learning Algorithm
Chenchu Xu, Lei Xu 0037, Zhifan Gao, Heye Zhang, Yanping Zhang 0001, Xiuquan Du, Shu Zhao 0005, Dhanjoo N. Ghista, Shuo Li 0001 |
MICCAI (3) | 10 |
| 2017 | Full Quantification of Left Ventricle via Deep Multitask Learning Network Respecting Intra- and Inter-Task Relatedness
Wufeng Xue, Andrea Lum, Ashley Mercado, Mark Landis, James Warrington, Shuo Li 0001 |
MICCAI (3) | 6 |
| 2017 | Current trends in the development of intelligent unmanned autonomous systemsabstractIntelligent unmanned autonomous systems are some of the most important applications of artificial intelligence (AI). The development of such systems can significantly promote innovation in AI technologies. This paper introduces the trends in the development of intelligent unmanned autonomous systems by summarizing the main achievements in each technological platform. Furthermore, we classify the relevant technologies into seven areas, including AI technologies, unmanned vehicles, unmanned aerial vehicles, service robots, space robots, marine robots, and unmanned workshops/intelligent plants. Current trends and developments in each area are introduced. Tao Zhang 0006, Qing Li 0010, Changshui Zhang, Hua-wei Liang, Ping Li 0057, Tianmiao Wang, Shuo Li 0001, Yun-long Zhu |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2017 | Robust estimation of carotid artery wall motion using the elasticity-based state-space approach
Zhifan Gao, Huahua Xiong, Xin Liu 0023, Heye Zhang, Dhanjoo N. Ghista, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2017 | Unsupervised boundary delineation of spinal neural foramina using a multi-feature and adaptive spectral segmentation
Xiaoxu He, Heye Zhang, Mark Landis, Manas Sharma, James Warrington, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2017 | Direct and simultaneous estimation of cardiac four chamber volumes by multioutput sparse regression
Xiantong Zhen, Heye Zhang, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2017 | Cross Validation Through Two-Dimensional Solution Surface for Cost-Sensitive SVMabstractModel selection plays an important role in cost-sensitive SVM (CS-SVM). It has been proven that the global minimum cross validation (CV) error can be efficiently computed based on the solution path for one parameter learning problems. However, it is a challenge to obtain the global minimum CV error for CS-SVM based on one-dimensional solution path and traditional grid search, because CS-SVM is with two regularization parameters. In this paper, we propose a solution and error surfaces based CV approach (CV-SES). More specifically, we first compute a two-dimensional solution surface for CS-SVM based on a bi-parameter space partition algorithm, which can fit solutions of CS-SVM for all values of both regularization parameters. Then, we compute a two-dimensional validation error surface for each CV fold, which can fit validation errors of CS-SVM for all values of both regularization parameters. Finally, we obtain the CV error surface by superposing K validation error surfaces, which can find the global minimum CV error of CS-SVM. Experiments are conducted on seven datasets for cost sensitive learning and on four datasets for imbalanced learning. Experimental results not only show that our proposed CV-SES has a better generalization ability than CS-SVM with various hybrids between grid search and solution path methods, and than recent proposed cost-sensitive hinge loss SVM with three-dimensional grid search, but also show that CV-SES uses less running time. Bin Gu 0001, Victor S. Sheng, KengYeow Tay, Walter Romano, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Unsupervised shape discovery using synchronized spectral networks
Yunliang Cai, Andrea Lum, Ashley Mercado, Mark Landis, James Warrington, Shuo Li 0001 |
Pattern Recognit. | 6 |
| 2017 | Automated segmentation and area estimation of neural foramina with boundary regression model
Xiaoxu He, Andrea Lum, Manas Sharma, Gary Brahm, Ashley Mercado, Shuo Li 0001 |
Pattern Recognit. | 6 |
| 2017 | Direct Multitype Cardiac Indices Estimation via Joint Representation and Regression LearningabstractCardiac indices estimation is of great importance during identification and diagnosis of cardiac disease in clinical routine. However, estimation of multitype cardiac indices with consistently reliable and high accuracy is still a great challenge due to the high variability of cardiac structures and the complexity of temporal dynamics in cardiac MR sequences. While efforts have been devoted into cardiac volumes estimation through feature engineering followed by a independent regression model, these methods suffer from the vulnerable feature representation and incompatible regression model. In this paper, we propose a semi-automated method for multitype cardiac indices estimation. After the manual labeling of two landmarks for ROI cropping, an integrated deep neural network Indices-Net is designed to jointly learn the representation and regression models. It comprises two tightly-coupled networks, such as a deep convolution autoencoder for cardiac image representation, and a multiple output convolution neural network for indices regression. Joint learning of the two networks effectively enhances the expressiveness of image representation with respect to cardiac indices, and the compatibility between image representation and indices regression, thus leading to accurate and reliable estimations for all the cardiac indices. When applied with five-fold cross validation on MR images of 145 subjects, Indices-Net achieves consistently low estimation error for LV wall thicknesses (1.44 ± 0.71 mm) and areas of cavity and myocardium (204 ± 133 mm2). It outperforms, with significant error reductions, segmentation method (55.1% and 17.4%), and two-phase direct volume-only methods (12.7% and 14.6%) for wall thicknesses and areas, respectively. These advantages endow the proposed method a great potential in clinical cardiac function assessment. Wufeng Xue, Ali Islam, Mousumi Bhaduri, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Descriptor Learning via Supervised Manifold Regularization for Multioutput RegressionabstractMultioutput regression has recently shown great ability to solve challenging problems in both computer vision and medical image analysis. However, due to the huge image variability and ambiguity, it is fundamentally challenging to handle the highly complex input-target relationship of multioutput regression, especially with indiscriminate high-dimensional representations. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm for multioutput regression, which can establish discriminative and compact feature representations to improve the multivariate estimation performance. The SDL is formulated as generalized low-rank approximations of matrices with a supervised manifold regularization. The SDL is able to simultaneously extract discriminative features closely related to multivariate targets and remove irrelevant and redundant information by transforming raw features into a new low-dimensional space aligned to targets. The achieved discriminative while compact descriptor largely reduces the variability and ambiguity for multioutput regression, which enables more accurate and efficient multivariate estimation. We conduct extensive evaluation of the proposed SDL on both synthetic data and real-world multioutput regression tasks for both computer vision and medical image analysis. Experimental results have shown that the proposed SDL can achieve high multivariate estimation accuracy on all tasks and largely outperforms the algorithms in the state of the arts. Our method establishes a novel SDL framework for multioutput regression, which can be widely used to boost the performance in different applications. Xiantong Zhen, Mengyang Yu, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2016 | Carotid Artery Wall Motion Estimated from Ultrasound Imaging Sequences Using a Nonlinear State Space ApproachabstractIt is very challenge to investigate the motion of the carotid artery wall in ultrasound images, because of the high nonlinear dynamics of this motion. In our study, the nonlinear dynamics of carotid artery wall motion is first approximated by our nonlinear state-space approach driven by a mathematical model of the mechanical deformation of carotid artery wall. Then, the two-dimensional motion of carotid artery wall is computed by solving the nonlinear state-space approach using the unscented Kalman filter. We have then evaluated the performance of our approach by comparing it with the manual tracing method (the correlation coefficient equals 0.9897 for the radial motion and 0.9703 for the longitudinal motion) and three other state-of-the-art methods for 73 subjects. The results indicate the reliable applicability of our approach in tracking the motion of the carotid artery wall and its potential usefulness in routine clinical diagnosis. Zhifan Gao, Heye Zhang, Dhanjoo N. Ghista, Huahua Xiong, Xin Liu 0023, Yaoqin Xie, Shuo Li 0001 |
MICCAI (3) | 10 |
| 2016 | Automated Diagnosis of Neural Foraminal Stenosis Using Synchronized Superpixels Representation
Xiaoxu He, Yilong Yin, Manas Sharma, Gary Brahm, Ashley Mercado, Shuo Li 0001 |
MICCAI (2) | 6 |
| 2016 | Multi-task Shape Regression for Medical Image SegmentationabstractIn this paper, we propose a general segmentation framework of Multi-Task Shape Regression (MTSR) which formulates segmentation as multi-task learning to leverage its strength of jointly solving multiple tasks enhanced by capturing task correlations. The MTSR entirely estimates coordinates of all points on shape contours by multi-task regression, where estimation of each coordinate corresponds to a regression task; the MTSR can jointly handle nonlinear relationships between image appearance and shapes while capturing holistic shape information by encoding coordinate correlations, which enables estimation of highly variable shapes, even with vague edge or region inhomogeneity. The MTSR achieves a long-desired general framework without relying on any specific assumptions or initialization, which enables flexible and fully automatic segmentation of multiple objects simultaneously, for different applications irrespective of modalities. The MTSR is validated on six representative applications of diverse images, achieves consistently high performance with dice similarity coefficient (DSC) up to 0.93 and largely outperforms state of the arts in each application, which demonstrates its effectiveness and generality for medical image segmentation. Xiantong Zhen, Yilong Yin, Mousumi Bhaduri, Ilanit Ben Nachum, David T. Laidley, Shuo Li 0001 |
MICCAI (3) | 6 |
| 2016 | Multi-scale deep networks and regression forests for direct bi-ventricular volume estimation
Xiantong Zhen, Zhijie Wang 0003, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2016 | Unsupervised Freeview Groupwise Cardiac Segmentation Using Synchronized Spectral NetworkabstractThe diagnosis, comparative and population study of cardiac radiology data require heart segmentation on increasingly large amount of images from different modalities/chambers/patients under various imaging views. Most existing automatic cardiac segmentation methods are often limited to single image segmentation with regulated modality/region settings or well-cropped ROI areas, which is impossible for large datasets due to enormous device protocols and institutional differences. A pure data-driven unsupervised segmentation without regulated setting requirements is crucial in this scenario, and will significantly automate the manual work and adopt the various changes of modality, subject or view. In this paper, we propose a general unsupervised groupwise segmentation: a direct simultaneous segmentation for a group of multi-modality, multi-chamber, multi-subject ( M3) cardiac images from a freely chosen imaging view. The segmentation can directly perform not only on regulated two/four-chamber images, but also on non-regulated uncropped raw MR/CT scans. A new Synchronized Spectral Network (SSN) is developed for the simultaneous decomposing, synchronizing, and clustering the spectral features of free-view M3cardiac images. The SSN-based groupwise analysis of image spectral bases immediately leads to groupwise segmentation of M3freeview images. The segmentation is quantitatively evaluated by three datasets (MR and CT mixed) with more than 200 subjects. High dice metric ( ) is consistently achieved in validation. Our method provides a powerful tool for medical images under general imaging environment. Yunliang Cai, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2015 | Supervised descriptor learning for multi-output regressionabstractDescriptor learning has recently drawn increasing attention in computer vision, Existing algorithms are mainly developed for classification rather than for regression which however has recently emerged as a powerful tool to solve a broad range of problems, e.g., head pose estimation. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm to establish a discriminative and compact feature representation for multi-output regression. By formulating as generalized low-rank approximations of matrices with a supervised manifold regularization (SMR), the SDL removes irrelevant and redundant information from raw features by transforming into a low-dimensional space under the supervision of multivariate targets. The obtained discriminative while compact descriptor largely reduces the variability and ambiguity in multi-output regression, and therefore enables more accurate and efficient multivariate estimation. We demonstrate the effectiveness of the proposed SDL algorithm on a representative multi-output regression task: head pose estimation using the benchmark Pointing'04 dataset. Experimental results show that the SDL can achieve high pose estimation accuracy and significantly outperforms state-of-the-art algorithms by an error reduction up to 27.5%. The proposed SDL algorithm provides a general descriptor learning framework in a supervised way for multi-output regression which can largely boost the performance of existing multi-output regression tasks. Xiantong Zhen, Zhijie Wang 0003, Mengyang Yu, Shuo Li 0001 |
CVPR | 4 |
| 2015 | Bi-Parameter Space Partition for Cost-Sensitive SVM
Bin Gu 0001, Victor S. Sheng, Shuo Li 0001 |
IJCAI | 3 |
| 2015 | Maintaining constant towing tension between cable ship and burying system under sea waves by hybrid FUZZY P + ID controllerabstractIn this paper, we propose a hybrid FUZZY P + ID controller to stabilize the towing cable tension between a cable ship and a burying system. First, we develop the model of a winch system driven by valve-controlled hydraulic motors and evaluate the step responses yielded by the conventional PID and the proposed FUZZY P+ID controllers using simulations. The comparative studies show that the control performance yielded by the FUZZY P+ID controller is superior. We replace the existing PID controller implemented on the towing winch with the FUZZY P + ID controller for burying cable tasks at speed up to 1 knot under sea waves with significant variations (peak-to-peak) from 1.5 to 2.5 meters. The real applications demonstrate that the FUZZY P + ID controller is much more robust than the conventional PID controller. Wei Li 0006, Xiaohui Wang 0014, Shuo Li 0001, Bin Xian |
IROS | 5 |
| 2015 | Unsupervised Free-View Groupwise Segmentation for M3 Cardiac Images Using Synchronized Spectral Network
Yunliang Cai, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2015 | Direct and Simultaneous Four-Chamber Volume Estimation by Multi-Output Regression
Xiantong Zhen, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
MICCAI (1) | 5 |
| 2015 | Incremental learning for ν-Support Vector Regression
Bin Gu 0001, Victor S. Sheng, Zhijie Wang 0003, Derek Ho, Said Osman, Shuo Li 0001 |
Neural Networks | 6 |
| 2015 | Distribution Matching with the Bhattacharyya Similarity: A Bound Optimization FrameworkabstractWe present efficient graph cut algorithms for three problems: (1) finding a region in an image, so that the histogram (or distribution) of an image feature within the region most closely matches a given model; (2) co-segmentation of image pairs and (3) interactive image segmentation with a user-provided bounding box. Each algorithm seeks the optimum of a global cost function based on the Bhattacharyya measure, a convenient alternative to other matching measures such as the Kullback-Leibler divergence. Our functionals are not directly amenable to graph cut optimization as they contain non-linear functions of fractional terms, which make the ensuing optimization problems challenging. We first derive a family of parametric bounds of the Bhattacharyya measure by introducing an auxiliary labeling. Then, we show that these bounds are auxiliary functions of the Bhattacharyya measure, a result which allows us to solve each problem efficiently via graph cuts. We show that the proposed optimization procedures converge within very few graph cut iterations. Comprehensive and various experiments, including quantitative and comparative evaluations over two databases, demonstrate the advantages of the proposed algorithms over related works in regard to optimality, computational load, accuracy and flexibility. Ismail Ben Ayed, Kumaradevan Punithakumar, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Multi-Modality Vertebra Recognition in Arbitrary Views Using 3D Deformable Hierarchical ModelabstractComputer-aided diagnosis of spine problems relies on the automatic identification of spine structures in images. The task of automatic vertebra recognition is to identify the global spine and local vertebra structural information such as spine shape, vertebra location and pose. Vertebra recognition is challenging due to the large appearance variations in different image modalities/views and the high geometric distortions in spine shape. Existing vertebra recognitions are usually simplified as vertebrae detections, which mainly focuses on the identification of vertebra locations and labels but cannot support further spine quantitative assessment. In this paper, we propose a vertebra recognition method using 3D deformable hierarchical model (DHM) to achieve cross-modality local vertebra location+pose identification with accurate vertebra labeling, and global 3D spine shape recovery. We recast vertebra recognition as deformable model matching, fitting the input spine images with the 3D DHM via deformations. The 3D model-matching mechanism provides a more comprehensive vertebra location+pose+label simultaneous identification than traditional vertebra location+label detection, and also provides an articulated 3D mesh model for the input spine section. Moreover, DHM can conduct versatile recognition on volume and multi-slice data, even on single slice. Experiments show our method can successfully extract vertebra locations, labels, and poses from multi-slice T1/T2 MR and volume CT, and can reconstruct 3D spine model on different image views such as lumbar, cervical, even whole spine. The resulting vertebra information and the recovered shape can be used for quantitative diagnosis of spine problems and can be easily digitalized and integrated in modern medical PACS systems. Yunliang Cai, Said Osman, Manas Sharma, Mark Landis, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2015 | Guest Editorial Special Issue on Spine Imaging, Image-Based Modeling, and Image Guided InterventionabstractThis special issue consists of 12 research papers. The papers cover a wide range of important topics including novel imaging, computational modeling, automatic vertebra segmentation, as well as computer assisted intervention. Among the 12 papers, nine are related to medical image analysis and three are in the scope of computer assisted intervention. The papers cover a variety of imaging modalities, include CT, MR, X-ray, ), ultrasound, and multi-modality. Shuo Li 0001, Jianhua Yao 0001, Nassir Navab |
IEEE Trans. Medical Imaging | 1 |
| 2015 | Regression Segmentation for M3 Spinal ImagesabstractClinical routine often requires to analyze spinal images of multiple anatomic structures in multiple anatomic planes from multiple imaging modalities (M(3)). Unfortunately, existing methods for segmenting spinal images are still limited to one specific structure, in one specific plane or from one specific modality (S(3)). In this paper, we propose a novel approach, Regression Segmentation, that is for the first time able to segment M(3) spinal images in one single unified framework. This approach formulates the segmentation task innovatively as a boundary regression problem: modeling a highly nonlinear mapping function from substantially diverse M(3) images directly to desired object boundaries. Leveraging the advancement of sparse kernel machines, regression segmentation is fulfilled by a multi-dimensional support vector regressor (MSVR) which operates in an implicit, high dimensional feature space where M(3) diversity and specificity can be systematically categorized, extracted, and handled. The proposed regression segmentation approach was thoroughly tested on images from 113 clinical subjects including both disc and vertebral structures, in both sagittal and axial planes, and from both MRI and CT modalities. The overall result reaches a high dice similarity index (DSI) 0.912 and a low boundary distance (BD) 0.928 mm. With our unified and expendable framework, an efficient clinical tool for M(3) spinal image segmentation can be easily achieved, and will substantially benefit the diagnosis and treatment of spinal diseases. Zhijie Wang 0003, Xiantong Zhen, KengYeow Tay, Said Osman, Walter Romano, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2015 | Correction to "Regression Segmentation for M3 Spinal Images"
Zhijie Wang 0003, Xiantong Zhen, KengYeow Tay, Said Osman, Walter Romano, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2015 | Incremental Support Vector Learning for Ordinal RegressionabstractSupport vector ordinal regression (SVOR) is a popular method to tackle ordinal regression problems. However, until now there were no effective algorithms proposed to address incremental SVOR learning due to the complicated formulations of SVOR. Recently, an interesting accurate on-line algorithm was proposed for training ν -support vector classification (ν-SVC), which can handle a quadratic formulation with a pair of equality constraints. In this paper, we first present a modified SVOR formulation based on a sum-of-margins strategy. The formulation has multiple constraints, and each constraint includes a mixture of an equality and an inequality. Then, we extend the accurate on-line ν-SVC algorithm to the modified formulation, and propose an effective incremental SVOR algorithm. The algorithm can handle a quadratic formulation with multiple constraints, where each constraint is constituted of an equality and an inequality. More importantly, it tackles the conflicts between the equality and inequality constraints. We also provide the finite convergence analysis for the algorithm. Numerical experiments on the several benchmark and real-world data sets show that the incremental algorithm can converge to the optimal solution in a finite number of steps, and is faster than the existing batch and incremental SVOR algorithms. Meanwhile, the modified formulation has better accuracy than the existing incremental SVOR algorithm, and is as accurate as the sum-of-margins based formulation of Shashua and Levin. Bin Gu 0001, Victor S. Sheng, KengYeow Tay, Walter Romano, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | Direct Estimation of Cardiac Bi-ventricular Volumes with Regression Forests
Xiantong Zhen, Zhijie Wang 0003, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
MICCAI (2) | 6 |
| 2014 | Regional Assessment of Cardiac Left Ventricular Myocardial Function via MRI Statistical FeaturesabstractAutomating the detection and localization of segmental (regional) left ventricle (LV) abnormalities in magnetic resonance imaging (MRI) has recently sparked an impressive research effort, with promising performances and a breadth of techniques. However, despite such an effort, the problem is still acknowledged to be challenging, with much room for improvements in regard to accuracy. Furthermore, most of the existing techniques are labor intensive, requiring delineations of the endo- and/or epi-cardial boundaries in all frames of a cardiac sequence. The purpose of this study is to investigate a real-time machine-learning approach which uses some image features that can be easily computed, but that nevertheless correlate well with the segmental cardiac function. Starting from a minimum user input in only one frame in a subject dataset, we build for all the regional segments and all subsequent frames a set of statistical MRI features based on a measure of similarity between distributions. We demonstrate that, over a cardiac cycle, the statistical features are related to the proportion of blood within each segment. Therefore, they can characterize segmental contraction without the need for delineating the LV boundaries in all the frames. We first seek the optimal direction along which the proposed image features are most descriptive via a linear discriminant analysis. Then, using the results as inputs to a linear support vector machine classifier, we obtain an abnormality assessment of each of the standard cardiac segments in real-time. We report a comprehensive experimental evaluation of the proposed algorithm over 928 cardiac segments obtained from 58 subjects. Compared against ground-truth evaluations by experienced radiologists, the proposed algorithm performed competitively, with an overall classification accuracy of 86.09% and a kappa measure of 0.73. Mariam Afshin, Ismail Ben Ayed, Kumaradevan Punithakumar, Max W. K. Law, Ali Islam, Aashish Goela, Terry M. Peters, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2013 | Joint resource allocation for learning-based cognitive radio networks with MIMO-OFDM relay-aided transmissionsabstractIn this paper, we investigate the joint power allocation for dual-hop amplify-and-forward (AF) MIMO-OFDM cognitive radio (CR) networks. The considered AF MIMO-OFDM CR network coexists with a primary radio (PR) network through underlay spectrum sharing. In order to mitigate the interference to the PR network, environmental learning algorithm is adopted to blindly estimate the null space of the PR user, which are orthogonal to the PR communication channels. With necessary channel state information, under independent transmit power constraints as well as the interference constraints, the power allocation of CR source and relay and subcarrier pairing over two hops are optimized jointly to maximize the CR network throughput. Furthermore, the relay node implements an effective subcarrier permutation policy to enhance the performance further at the cost of affordable complexity. Finally, the performance advantages of the proposed algorithm are demonstrated by the simulation results. Shuo Li 0001, Bingquan Li, Chengwen Xing, Zesong Fei, Shaodan Ma |
WCNC | 1 |
| 2013 | Special issue on Shape Modeling in Medical Image Analysis
Wiro J. Niessen, Shuo Li 0001, Song Wang 0002 |
Comput. Vis. Image Underst. | 2 |
| 2013 | How to understand linear minimum mean-square-error transceiver design for multiple-input-multiple-output systems from quadratic matrix programmingabstractIn this study, a unified linear minimum mean‐square‐error (LMMSE) transceiver design framework is investigated, which is suitable for a wide range of wireless systems. The unified design is based on an elegant and powerful mathematical programming technology termed as quadratic matrix programming (QMP). Based on QMP it can be observed that for different wireless systems, there are certain common characteristics which can be exploited to design LMMSE transceivers, for example, the quadratic forms. It is also discovered that evolving from a point‐to‐point multiple‐input–multiple‐output (MIMO) system to various advanced wireless systems such as multi‐cell coordinated systems, multi‐user MIMO systems, MIMO cognitive radio systems, amplify‐and‐forward MIMO relaying systems and so on, the quadratic nature is always kept and the LMMSE transceiver designs can always be carried out via iteratively solving a number of QMP problems. A comprehensive framework on how to solve QMP problems is also given. The work presented in this study is likely to be the first shot for the transceiver design for the future ever‐changing wireless systems. Chengwen Xing, Shuo Li 0001, Zesong Fei, Jingming Kuang 0001 |
IET Commun. | 2 |
| 2013 | Intervertebral disc segmentation in MR images using anisotropic oriented flux
Max W. K. Law, KengYeow Tay, Andrew E. Leung, Gregory J. Garvin, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2013 | Regional heart motion abnormality detection: An information theoretic approach
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Aashish Goela, Ian G. Ross, Jaron Chong, Shuo Li 0001 |
Medical Image Anal. | 7 |
| 2013 | Robust Filter-and-forward Beamforming Design for Two-way Multi-antenna Relaying Networks
Zesong Fei, Niwei Wang, Chengwen Xing, Shuo Li 0001, Jiqing Ni, Jingming Kuang 0001 |
Mob. Networks Appl. | 4 |
| 2012 | Dilated Divergence Based Scale-Space Representation for Curve Analysis
Max W. K. Law, KengYeow Tay, Andrew E. Leung, Gregory J. Garvin, Shuo Li 0001 |
ECCV (2) | 5 |
| 2012 | Global Assessment of Cardiac Function Using Image Statistics in MRI
Mariam Afshin, Ismail Ben Ayed, Ali Islam, Aashish Goela, Terry M. Peters, Shuo Li 0001 |
MICCAI (2) | 6 |
| 2012 | Regional Heart Motion Abnormality Detection via Multiview Fusion
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Aashish Goela, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2012 | Max-flow segmentation of the left ventricle by recovering subject-specific distributions via a bound of the Bhattacharyya measure
Ismail Ben Ayed, Huamei Chen, Kumaradevan Punithakumar, Ian G. Ross, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2012 | A Convex Max-Flow Approach to Distribution-Based Figure-Ground SeparationabstractThis study investigates a convex relaxation approach to figure-ground separation with a global distribution matching prior evaluated by the Bhattacharyya measure. The problem amounts to finding a region that most closely matches a known model distribution. It has been previously addressed by curve evolution, which leads to suboptimal and computationally intensive algorithms, or by graph cuts, which result in metrication errors. Solving a sequence of convex subproblems, the proposed relaxation is based on a novel bound of the Bhattacharyya measure which yields an algorithm robust to initial conditions. Furthermore, we propose a novel flow configuration that accounts for labeling-function variations, unlike existing configurations. This leads to a new max-flow formulation which is dual to the convex relaxed subproblems we obtained. We further prove that such a formulation yields exact and global solutions to the original, nonconvex subproblems. A comprehensive experimental evaluation on the Microsoft GrabCut database demonstrates that our approach yields improvements in optimality and accuracy over related recent methods. Kumaradevan Punithakumar, Jing Yuan 0001, Ismail Ben Ayed, Shuo Li 0001, Yuri Boykov |
SIAM J. Imaging Sci. | 4 |
| 2011 | Assessment of Regional Myocardial Function via Statistical Features in MR Images
Mariam Afshin, Ismail Ben Ayed, Kumaradevan Punithakumar, Max W. K. Law, Ali Islam, Aashish Goela, Ian G. Ross, Terry M. Peters, Shuo Li 0001 |
MICCAI (3) | 9 |
| 2010 | Graph cut segmentation with a global constraint: Recovering region distribution via a bound of the Bhattacharyya measureabstractThis study investigates an efficient algorithm for image segmentation with a global constraint based on the Bhattacharyya measure. The problem consists of finding a region consistent with an image distribution learned a priori. We derive an original upper bound of the Bhattacharyya measure by introducing an auxiliary labeling. From this upper bound, we reformulate the problem as an optimization of an auxiliary function by graph cuts. Then, we demonstrate that the proposed procedure converges and give a statistical interpretation of the upper bound. The algorithm requires very few iterations to converge, and finds nearly global optima. Quantitative evaluations and comparisons with state-of-the-art methods on the Microsoft GrabCut segmentation database demonstrated that the proposed algorithm brings improvements in regard to segmentation accuracy, computational efficiency, and optimality. We further demonstrate the flexibility of the algorithm in object tracking. Ismail Ben Ayed, Huamei Chen, Kumaradevan Punithakumar, Ian G. Ross, Shuo Li 0001 |
CVPR | 5 |
| 2010 | Finding image distributions on active curvesabstractThis study investigates an active curve functional which measures a similarity between the distribution of an image feature on the curve and a model distribution learned a priori. The curve evolution equation resulting from the minimization of this contour-based functional can be viewed as a geodesic active contour with a variable stopping function. The variable stopping function depends on the distribution of image feature on the curve and, therefore, can deal with difficult cases where the desired boundary corresponds to very weak image transitions. We ran several experiments supported by quantitative performance evaluations over several examples of segmentation and tracking of the left ventricle inner and outer boundaries in cardiac magnetic resonance image sequences. The results are significantly more accurate than with region-based and edge-based functionals. Ismail Ben Ayed, Amar Mitiche, Mohamed Ben Salah, Shuo Li 0001 |
CVPR | 4 |
| 2010 | A Parameterization of Deformation Fields for Diffeomorphic Image Registration and Its Application to Myocardial Delineation
Huamei Chen, Aashish Goela, Gregory J. Garvin, Shuo Li 0001 |
MICCAI (1) | 4 |
| 2010 | Regional Heart Motion Abnormality Detection via Information Measures and Unscented Kalman Filtering
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Ian G. Ross, Shuo Li 0001 |
MICCAI (1) | 5 |
| 2010 | Detection of left ventricular motion abnormality via information measures and Bayesian filteringabstractWe present an original information theoretic measure of heart motion based on the Shannon's differential entropy (SDE), which allows heart wall motion abnormality detection. Based on functional images, which are subject to noise and segmentation inaccuracies, heart wall motion analysis is acknowledged as a difficult problem, and as such, incorporation of prior knowledge is crucial for improving accuracy. Given incomplete, noisy data and a dynamic model, the Kalman filter, a well-known recursive Bayesian filter, is devised in this study to the estimation of the left ventricular (LV) cavity points. However, due to similarity between the statistical information of normal and abnormal heart motions, detecting and classifying abnormality is a challenging problem, which we investigate with a global measure based on the SDE. We further derive two other possible information theoretic abnormality detection criteria, one is based on Rényi entropy and the other on Fisher information. The proposed methods analyze wall motion quantitatively by constructing distributions of the normalized radial distance estimates of the LV cavity. Using 269 x 20 segmented LV cavities of short-axis MRI obtained from 30 subjects, the experimental analysis demonstrates that the proposed SDE criterion can lead to a significant improvement over other features that are prevalent in the literature related to the LV cavity, namely, mean radial displacement and mean radial velocity. Kumaradevan Punithakumar, Ismail Ben Ayed, Ian G. Ross, Ali Islam, Jaron Chong, Shuo Li 0001 |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2009 | Tracking Endocardial Boundary and Motion via Graph Cut Distribution Matching and Multiple Model Filtering
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Ian G. Ross, Shuo Li 0001 |
ACCV (3) | 5 |
| 2009 | Left Ventricle Segmentation via Graph Cut Distribution Matching
Ismail Ben Ayed, Kumaradevan Punithakumar, Shuo Li 0001, Ali Islam, Jaron Chong |
MICCAI (1) | 3 |
| 2009 | Heart Motion Abnormality Detection via an Information Measure and Bayesian Filtering
Kumaradevan Punithakumar, Shuo Li 0001, Ismail Ben Ayed, Ian G. Ross, Ali Islam, Jaron Chong |
MICCAI (1) | 2 |
| 2009 | A Statistical Overlap Prior for Variational Image Segmentation
Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
Int. J. Comput. Vis. | 2 |
| 2009 | Embedding Overlap Priors in Variational Left Ventricle TrackingabstractWe propose to embed overlap priors in variational tracking of the left ventricle (LV) in cardiac magnetic resonance (MR) sequences. The method consists of evolving two curves toward the LV endo- and epicardium boundaries. We derive the curve evolution equations by minimizing two functionals each containing an original overlap prior constraint. The latter measures the conformity of the overlap between the nonparametric (kernel-based) intensity distributions within the three target regions--LV cavity, myocardium and background-to a prior learned from a given segmentation of the first frame. The Bhattacharyya coefficient is used as an overlap measure. Different from existing intensity-driven constraints, the proposed priors do not assume implicitly that the overlap between the intensity distributions within different regions has to be minimal. This prevents both the papillary muscles from being included erroneously in the myocardium and the curves from spilling into the background. Although neither geometric training nor preprocessing were used, quantitative evaluation of the similarities between automatic and independent manual segmentations showed that the proposed method yields a competitive score in comparison with existing methods. This allows more flexibility in clinical use because our solution is based only on the current intensity data, and consequently, the results are not bounded to the characteristics, variability, and mathematical description of a finite training set. We also demonstrate experimentally that the overlap measures are approximately constant over a cardiac sequence, which allows to learn the overlap priors from a single frame. Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
IEEE Trans. Medical Imaging | 2 |
| 2008 | Tracking distributions with an overlap priorabstractRecent studies have shown that embedding similarity/dissimilarity measures between distributions in the variational level set framework can lead to effective object segmentation/tracking algorithms. In this connection, existing methods assume implicitly that the overlap between the distributions of image data within the object and its background has to be minimal. Unfortunately, such assumption may not be valid in many important applications. This study investigates an overlap prior, which embeds knowledge about the overlap between the distributions of the object and the background in level set tracking. It consists of evolving a curve to delineate the target object in the current frame. The level set curve evolution equation is sought following the maximization of a functional containing three terms: (1) an original overlap prior which measures the conformity of overlap between the nonparametric (kernel-based) distributions within the object and the background to a learned description, (2) a term which measures the similarity between a model distribution of the object and the sample distribution inside the curve, and (3) a regularization term for smooth segmentation boundaries. The Bhattacharyya coefficient is used as an overlap measure. Apart from leading to a method which is more versatile than current ones, the overlap prior speeds up significantly the curve evolution. Comparisons and results demonstrate the advantages of the proposed prior over related methods, and its usefulness in important applications such as the left ventricle tracking in magnetic resonance (MR) images. Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
CVPR | 2 |
| 2008 | Left Ventricle Tracking Using Overlap Priors
Ismail Ben Ayed, Yingli Lu, Shuo Li 0001, Ian G. Ross |
MICCAI (1) | 3 |
| 2007 | Semi-automatic computer aided lesion detection in dental X-rays using variational level set
Shuo Li 0001, Thomas Fevens, Adam Krzyzak, Song Li 0003 |
Pattern Recognit. | 1 |
| 2007 | Motion learning-based framework for unarticulated shape animation
Thomas Fevens, Shuo Li 0001, Sudhir P. Mudur |
Vis. Comput. | 3 |
| 2006 | Fast and Robust Clinical Triple-Region Image Segmentation Using One Level Set Function
Shuo Li 0001, Thomas Fevens, Adam Krzyzak, Song Li 0003 |
MICCAI (2) | 1 |
| 2006 | Automatic clinical image segmentation using pathological modeling, PCA and SVM
Shuo Li 0001, Thomas Fevens, Adam Krzyzak, Song Li 0003 |
Eng. Appl. Artif. Intell. | 1 |
| 2005 | Toward Automatic Computer Aided Dental X-ray Analysis Using Level Set Method
Shuo Li 0001, Thomas Fevens, Adam Krzyzak, Song Li 0003 |
MICCAI | 1 |
| 2004 | Image Segmentation Adapted for Clinical Settings by Combining Pattern Classification and Level Sets
Shuo Li 0001, Thomas Fevens, Adam Krzyzak |
MICCAI (1) | 1 |