EDBT 2026 Demo / reviewers in the wild / expert
Lei Bi 0001
dblp:03/2981-1
· DBLP profile ↗
47ranked-venue papers
8as first author
32since 2021 · last 2027
0000-0001-9759-0200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Language-guided medical image segmentation with target-informed multi-level contrastive alignmentsabstractMedical image segmentation is a fundamental task in numerous medical applications. Recently, language-guided segmentation has shown promise in medical scenarios where textual clinical reports are readily available as semantic guidance. Clinical reports contain diagnostic information provided by clinicians, which can provide auxiliary textual semantics to guide segmentation. However, existing language-guided segmentation methods neglect the inherent pattern gaps between image and text modalities, resulting in sub-optimal visual-language integration. Contrastive learning is a well-recognized approach to align image-text patterns, but it has not been optimized for medical image segmentation, where clinically meaningful semantics are often concentrated in localized target regions rather than the entire image. In this study, we propose TMCA, a Target-informed Multi-level Contrastive Alignment framework to bridge image-text pattern gaps for medical language-guided segmentation. The core innovation is to reformulate image-text contrastive alignment from conventional instance-level matching to segmentation-oriented semantic matching, where image-text samples are aligned according to their segmentation targets rather than merely whether they come from the same patient. Specifically, TMCA enables target-informed image-text alignments and fine-grained textual guidance by introducing: (i) a target-sensitive semantic distance module that utilizes target information for more granular image-text alignment modeling, (ii) a multi-level contrastive alignment strategy that directs fine-grained textual guidance to multi-scale image details, and (iii) a language-guided target enhancement module that reinforces attention to critical regions based on the aligned image-text patterns. Extensive experiments on four public benchmarks, involving three medical imaging modalities with clinical reports, show that TMCA enabled superior performance over state-of-the-art language-guided medical segmentation methods. Mingjian Li, Mingyuan Meng, Shuchang Ye, Mingye Zou, Michael J. Fulham, Lei Bi 0001, Jinman Kim |
Expert Syst. Appl. | 6 |
| 2026 | Hierarchical Deep Decision Tree-Based Network for Odontogenic Cystic Lesion Classification in CBCT ImagesabstractOdontogenic cystic lesions (OCLs) are complex jaw abnormalities that require a precise diagnosis of the disease for treatment. Visual OCL diagnosis is commonly based on reviewing cone-beam computed tomography (CBCT) to identify morpho-pathological features associated with specific lesion types in a hierarchical manner. Current state-of-the-art methods focus on extracting features from the image without any guidance beyond the lesion diagnosis, and do not fully leverage the hierarchical relationship between the lesion diagnosis and morphological features. In this study, we propose a hierarchical deep decision tree network (H2DT-Net) with three modules: a deep decision tree-based hierarchical learning module (DHLM) to leverage inter-categorical relationships; a feature category embedding module (FCEM) to capture representations from both diagnostic and morpho-pathological domains and support the DHLM; and a lesion localised attention module (LLAM) to facilitate the feature extraction process by generating lesion-focused attention maps. Evaluated on 289 CBCT images, H2DT-Net achieved state-of-the-art performance in OCL classification. We further demonstrate that our method is effective in clinical settings, where it outperformed six maxillofacial clinicians in diagnostic assessment. Zimo Huang, Hao Wang 0143, Eduardo Delamare, Shengfu Huang, Lei Bi 0001, Jinman Kim |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | M2Net: Multimodal Multitask Mutual Learning for Anti-VEGF Efficacy PredictionabstractAge-related macular degeneration with abnormal blood vessel growth (neovascular AMD) is the leading cause of vision loss in elderly populations. While anti-VEGF injections are the standard treatment, they present financial burdens for patients and vary in effectiveness. Predicting treatment efficacy is therefore crucial for patient care. Current prediction methods fail to fully integrate information from different imaging techniques, typically focusing on either forecasting vision improvements or generating post-treatment images-but not both simultaneously. This approach overlooks the important relationship between these tasks. We present M2Net, a novel joint generation and classification network based on Multimodal Multitask Mutual learning, to simultaneously predict changes in visual acuity and generate post-treatment retinal images. M2Net employs a dual-branch structure that processes both fundus photographs and Optical Coherence Tomography (OCT) scans to improve prediction accuracy. Our framework includes two key innovations: the Multimodal Collaborative Treatment Efficacy Prediction module, which interacts the features between the two modalities and provides initial visual acuity change classification to guide the generation of post-treatment images; and the Pre-Post Treatment Image Joint Analysis module, which identifies both common and changing features between pre-treatment and post-treatment images to enhance prediction accuracy. To validate our approach, we created the dataset (MMPD) containing paired multimodal retinal images with corresponding visual acuity measurements. Experiments on the dataset demonstrate that M2Net achieves superior performance compared to existing methods, with a classification accuracy of 96.03%, an SSIM of 0.6377 on the OCT modality, and an SSIM of 0.8347 on the fundus modality. Our code will be available at https://github.com/zengying123/M2Net. Lei Bi 0001, Wuzhen Shi, Huazhu Fu, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | AutoFuse: Automatic fusion networks for deformable medical image registration
Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
Pattern Recognit. | 4 |
| 2025 | Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation ModelingabstractMedical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modeling the inter-matching relations, can learn improved medical image representations with better generalization capabilities. Mingjian Li, Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
IEEE Trans. Medical Imaging | 5 |
| 2025 | A Lightweight Depthwise Separable ConvNet with Frequency-domain Enhancement for Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is crucial in the diagnosis and treatment of various cardiovascular and eye diseases. Although current vessel segmentation methods have achieved impressive performance, some challenging issues still need to be addressed. For example, existing methods always cannot segment complex capillaries well because they may be interfered with or covered by other components in the retina, and they need to further improve the continuity and consistency of vessel segmentation results. Moreover, the excellent vessel segmentation methods are usually built on bulky and cumbersome models which greatly limit their application range. In this article, we propose a novel efficient depthwise separable convolution network with frequency-domain enhancement (dubbed RetiNeXt) for retinal vessel segmentation. Firstly, we design a lightweight vessel enhancement module to extract global fine topological structure features from the frequency domain to enhance the complex capillary vessel details. Secondly, we propose a global feature extraction block to fully capture the large-scale spatial information and global characterizations, which enables the model to maintain vessel structural coherence from a global perspective. Thirdly, we construct a local feature mixing block based on SimAM attention mechanism to highlight the tiny capillary topological structure features and optimize the segmentation of low-contrast blood vessels, thereby improving the integrity and continuity of complex capillaries. Comprehensive comparison experiments on three well-benchmarked retinal vessel segmentation datasets fully verify the effectiveness and superiority of the proposed RetiNeXt. To further demonstrate the universality of RetiNeXt for medical image segmentation, we also conduct sufficient comparative experiments on two classical coronary angiography datasets. Extensive quantitative and qualitative experiments fully show that RetiNeXt outperforms other state-of-the-art methods with only 0.4M of trainable parameters. Shunzhe Shen, Wuzhen Shi, Wenming Cao 0001, Lei Bi 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Deep learning-based binocular system for automated diabetic retinopathy grading with prior clinical knowledge integration
Saba Ghazanfar Ali, Xiangning Wang, Lei Bi 0001, Younhyun Jung, Tingli Chen, Haifang Zhang |
Vis. Comput. | 3 |
| 2025 | Deep contour attention learning for scleral deformation from OCT images
Hao Chen 0011, Yupeng Xu, Huating Li, Yuan Xie 0006, David Dagan Feng, Jinman Kim, Lei Bi 0001, Xiangui He, Bin Sheng 0001 |
Vis. Comput. | 9 |
| 2025 | HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001 |
Vis. Comput. | 17 |
| 2024 | Semi-Mamba: Improving Medical Image Segmentation via Semi-Automatic Mamba Network
Lei Bi 0001, Yige Peng, David Dagan Feng, Jinman Kim |
CGI (3) | 1 |
| 2024 | Correlation-aware Coarse-to-fine MLPs for Deformable Medical Image RegistrationabstractDeformable image registration is a fundamental step for medical image analysis. Recently, transformers have been used for registration and outperformed Convolutional Neural Networks (CNNs). Transformers can capture long-range dependence among image features, which have been shown beneficial for registration. However, due to the high computation/memory loads of self-attention, transformers are typically used at downsampled feature resolutions and cannot capture fine-grained long-range dependence at the full image resolution. This limits deformable registration as it necessitates precise dense correspondence between each image pixel. Multi-layer Perceptrons (MLPs) without self-attention are efficient in computation/memory usage, enabling the feasibility of capturing fine-grained long-range dependence at full resolution. Nevertheless, MLPs have not been extensively explored for image registration and are lacking the consideration of inductive bias crucial for medical registration tasks. In this study, we propose the first correlation-aware MLP-based registration network (CorrMLP) for deformable medical image registration. Our CorrMLP introduces a correlation-aware multi-window MLP block in a novel coarse-to-fine registration architecture, which captures fine-grained multi-range dependence to perform correlation-aware coarse-to-fine registration. Extensive experiments with seven public medical datasets show that our CorrMLP outperforms state-of-the-art deformable registration methods. Mingyuan Meng, David Dagan Feng, Lei Bi 0001, Jinman Kim |
CVPR | 3 |
| 2024 | 3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
Xiaoshuang Li, Mingyuan Meng, Zimo Huang, Lei Bi 0001, Eduardo Delamare, David Dagan Feng, Bin Sheng 0001, Jinman Kim |
MICCAI (7) | 4 |
| 2024 | A Transformer-Assisted Cascade Learning Network for Choroidal Vessel Segmentation
Lei Bi 0001, Wuzhen Shi, Yupeng Xu, Wenming Cao 0001, David Dagan Feng |
J. Comput. Sci. Technol. | 3 |
| 2024 | ClarityDiffuseNet: Enhancing fundus image quality under black shadows with diffusion model-based research
Jiadi Dong, Tianwei Qian, Yuxian Jiang, Lei Bi 0001, Jinman Kim, Lisheng Wang |
Pattern Recognit. Lett. | 4 |
| 2024 | Attention-based multi-scale feature fusion network for myopia grading using optical coherence tomography images
Gengyou Huang, Lei Bi 0001, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 4 |
| 2023 | Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (6) | 2 |
| 2023 | Non-iterative Coarse-to-Fine Transformer Networks for Joint Affine and Deformable Image Registration
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (10) | 2 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 40 |
| 2023 | Unsupervised Landmark Detection-Based Spatiotemporal Motion Estimation for 4-D Dynamic Medical ImagesabstractMotion estimation is a fundamental step in dynamic medical image processing for the assessment of target organ anatomy and function. However, existing image-based motion estimation methods, which optimize the motion field by evaluating the local image similarity, are prone to produce implausible estimation, especially in the presence of large motion. In addition, the correct anatomical topology is difficult to be preserved as the image global context is not well incorporated into motion estimation. In this study, we provide a novel motion estimation framework of dense-sparse-dense (DSD), which comprises two stages. In the first stage, we process the raw dense image to extract sparse landmarks to represent the target organ's anatomical topology, and discard the redundant information that is unnecessary for motion estimation. For this purpose, we introduce an unsupervised 3-D landmark detection network to extract spatially sparse but representative landmarks for the target organ's motion estimation. In the second stage, we derive the sparse motion displacement from the extracted sparse landmarks of two images of different time points. Then, we present a motion reconstruction network to construct the motion field by projecting the sparse landmarks' displacement back into the dense image domain. Furthermore, we employ the estimated motion field from our two-stage DSD framework as initialization and boost the motion estimation quality in light-weight yet effective iterative optimization. We evaluate our method on two dynamic medical imaging tasks to model cardiac motion and lung respiratory motion, respectively. Our method has produced superior motion estimation accuracy compared to the existing comparative methods. Besides, the extensive experimental results demonstrate that our solution can extract well-representative anatomical landmarks without any requirement of manual annotation. Our code is publicly available online: https://github.com/yyguo-sjtu/DSD-3D-Unsupervised-Landmark-Detection-Based-Motion-Estimation. Yuyu Guo 0002, Lei Bi 0001, Dongming Wei, Liyun Chen, Zhengbin Zhu, David Dagan Feng, Ruiyan Zhang, Qian Wang 0001, Jinman Kim |
IEEE Trans. Cybern. | 2 |
| 2023 | A Shortened Model for Logan Reference Plot Implemented via the Self-Supervised Neural Network for Parametric PET ImagingabstractDynamic PET imaging provides superior physiological information than conventional static PET imaging. However, the dynamic information is gained at the cost of a long scanning protocol; this limits the clinical application of dynamic PET imaging. We developed a modified Logan reference plot model to shorten the acquisition procedure in dynamic PET imaging by omitting the early-time information necessary for the conventional reference Logan model. The proposed model is accurate theoretically, but the straightforward approach raises the sampling problem in implementation and results in noisy parametric images. We then designed a self-supervised convolutional neural network to increase the noise performance of parametric imaging, with dynamic images of only a single subject for training. The proposed method was validated via simulated and real dynamic [Formula: see text]-fallypride PET data. Results showed that it accurately estimated the distribution volume ratio (DVR) in dynamic PET with a shortened scanning protocol, e.g., 20 minutes, where the estimations were comparable with those obtained from a standard dynamic PET study of 120 minutes of acquisition. Further comparisons illustrated that our method outperformed the shortened Logan model implemented with Gaussian filtering, regularization, BM4D and the 4D deep image prior methods in terms of the trade-off between bias and variance. Since the proposed method uses data acquired in a short period of time upon the equilibrium, it has the potential to add clinical values by providing both DVR and Standard Uptake Value (SUV) simultaneously. It thus promotes clinical applications of dynamic PET studies when neuronal receptor functions are studied. Wenxiang Ding, Qiaoqiao Ding, Kewei Chen 0001, Miao Zhang 0038, David Dagan Feng, Lei Bi 0001, Jinman Kim, Qiu Huang |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Non-iterative Coarse-to-Fine Registration Based on Single-Pass Deep Cumulative Learning
Mingyuan Meng, Lei Bi 0001, David Dagan Feng, Jinman Kim |
MICCAI (6) | 2 |
| 2022 | Deep multi-scale resemblance network for the sub-class differentiation of adrenal masses on computed tomography images
Lei Bi 0001, Jinman Kim, Tingwei Su, Michael J. Fulham, David Dagan Feng, Guang Ning |
Artif. Intell. Medicine | 1 |
| 2022 | Hyper-fusion network for semi-automatic segmentation of skin lesions
Lei Bi 0001, Michael J. Fulham, Jinman Kim |
Medical Image Anal. | 1 |
| 2022 | An attention-enhanced cross-task network to analyse lung nodule attributes in CT images
Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
Pattern Recognit. | 2 |
| 2022 | DeepMTS: Deep Multi-Task Learning for Survival Prediction in Patients With Advanced Nasopharyngeal Carcinoma Using Pretreatment PET/CTabstractNasopharyngeal Carcinoma (NPC) is a malignant epithelial cancer arising from the nasopharynx. Survival prediction is a major concern for NPC patients, as it provides early prognostic information to plan treatments. Recently, deep survival models based on deep learning have demonstrated the potential to outperform traditional radiomics-based survival prediction models. Deep survival models usually use image patches covering the whole target regions (e.g., nasopharynx for NPC) or containing only segmented tumor regions as the input. However, the models using the whole target regions will also include non-relevant background information, while the models using segmented tumor regions will disregard potentially prognostic information existing out of primary tumors (e.g., local lymph node metastasis and adjacent tissue invasion). In this study, we propose a 3D end-to-end Deep Multi-Task Survival model (DeepMTS) for joint survival prediction and tumor segmentation in advanced NPC from pretreatment PET/CT. Our novelty is the introduction of a hard-sharing segmentation backbone to guide the extraction of local features related to the primary tumors, which reduces the interference from non-relevant background information. In addition, we also introduce a cascaded survival network to capture the prognostic information existing out of primary tumors and further leverage the global tumor information (e.g., tumor size, shape, and locations) derived from the segmentation backbone. Our experiments with two clinical datasets demonstrate that our DeepMTS can consistently outperform traditional radiomics-based survival prediction models and existing deep survival models. Mingyuan Meng, Bingxin Gu, Lei Bi 0001, Shaoli Song, David Dagan Feng, Jinman Kim |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Improving Breast Tumor Segmentation in PET via Attentive Transformation Based NormalizationabstractPositron Emission Tomography (PET) has become a preferred imaging modality for cancer diagnosis, radiotherapy planning, and treatment responses monitoring. Accurate and automatic tumor segmentation is the fundamental requirement for these clinical applications. Deep convolutional neural networks have become the state-of-the-art in PET tumor segmentation. The normalization process is one of the key components for accelerating network training and improving the performance of the network. However, existing normalization methods either introduce batch noise into the instance PET image by calculating statistics on batch level or introduce background noise into every single pixel by sharing the same learnable parameters spatially. In this paper, we proposed an attentive transformation (AT)-based normalization method for PET tumor segmentation. We exploit the distinguishability of breast tumor in PET images and dynamically generate dedicated and pixel-dependent learnable parameters in normalization via the transformation on a combination of channel-wise and spatial-wise attentive responses. The attentive learnable parameters allow to re-calibrate features pixel-by-pixel to focus on the high-uptake area while attenuating the background noise of PET images. Our experimental results on two real clinical datasets show that the AT-based normalization method improves breast tumor segmentation performance when compared with the existing normalization methods. Xiaoya Qiao, Chunjuan Jiang, Panli Li, Yuan Yuan 0022, Qinglong Zeng, Lei Bi 0001, Shaoli Song, Jinman Kim, David Dagan Feng, Qiu Huang |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Graph-Based Intercategory and Intermodality Network for Multilabel Classification and Melanoma Diagnosis of Skin Lesions in Dermoscopy and Clinical ImagesabstractThe identification of melanoma involves an integrated analysis of skin lesion images acquired using clinical and dermoscopy modalities. Dermoscopic images provide a detailed view of the subsurface visual structures that supplement the macroscopic details from clinical images. Visual melanoma diagnosis is commonly based on the 7-point visual category checklist (7PC), which involves identifying specific characteristics of skin lesions. The 7PC contains intrinsic relationships between categories that can aid classification, such as shared features, correlations, and the contributions of categories towards diagnosis. Manual classification is subjective and prone to intra- and interobserver variability. This presents an opportunity for automated methods to aid in diagnostic decision support. Current state-of-the-art methods focus on a single image modality (either clinical or dermoscopy) and ignore information from the other, or do not fully leverage the complementary information from both modalities. Furthermore, there is not a method to exploit the 'intercategory' relationships in the 7PC. In this study, we address these issues by proposing a graph-based intercategory and intermodality network (GIIN) with two modules. A graph-based relational module (GRM) leverages intercategorical relations, intermodal relations, and prioritises the visual structure details from dermoscopy by encoding category representations in a graph network. The category embedding learning module (CELM) captures representations that are specialised for each category and support the GRM. We show that our modules are effective at enhancing classification performance using three public datasets (7PC, ISIC 2017, and ISIC 2018), and that our method outperforms state-of-the-art methods at classifying the 7PC categories and diagnosis. Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
IEEE Trans. Medical Imaging | 2 |
| 2021 | A Classification Network for Ocular Diseases Based on Structure Feature and Visual Attention
Yupeng Xu, Bin Sheng 0001, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 5 |
| 2021 | Multi-Stream Fusion Network for Multi-Distortion Image Super-Resolution
Yupeng Xu, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 5 |
| 2021 | High-parallelism Inception-like Spiking Neural Networks for Unsupervised Feature Learning
Mingyuan Meng, Lei Bi 0001, Jinman Kim, Shanlin Xiao, Zhiyi Yu |
Neurocomputing | 3 |
| 2021 | Unsupervised brain tumor segmentation using a symmetric-driven adversarial network
Xinheng Wu, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Luping Zhou, Jinman Kim |
Neurocomputing | 2 |
| 2021 | Multimodal Spatial Attention Module for Targeting Multimodal PET-CT Lung Tumor SegmentationabstractMultimodal positron emission tomography-computed tomography (PET-CT) is used routinely in the assessment of cancer. PET-CT combines the high sensitivity for tumor detection of PET and anatomical information from CT. Tumor segmentation is a critical element of PET-CT but at present, the performance of existing automated methods for this challenging task is low. Segmentation tends to be done manually by different imaging experts, which is labor-intensive and prone to errors and inconsistency. Previous automated segmentation methods largely focused on fusing information that is extracted separately from the PET and CT modalities, with the underlying assumption that each modality contains complementary information. However, these methods do not fully exploit the high PET tumor sensitivity that can guide the segmentation. We introduce a deep learning-based framework in multimodal PET-CT segmentation with a multimodal spatial attention module (MSAM). The MSAM automatically learns to emphasize regions (spatial areas) related to tumors and suppress normal regions with physiologic high-uptake from the PET input. The resulting spatial attention maps are subsequently employed to target a convolutional neural network (CNN) backbone for segmentation of areas with higher tumor likelihood from the CT image. Our experimental results on two clinical PET-CT datasets of non-small cell lung cancer (NSCLC) and soft tissue sarcoma (STS) validate the effectiveness of our framework in these different cancer types. We show that our MSAM, with a conventional U-Net backbone, surpasses the state-of-the-art lung tumor segmentation approach by a margin of 7.6% in Dice similarity coefficient (DSC). Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | A Spatiotemporal Volumetric Interpolation Network for 4D Dynamic Medical ImageabstractDynamic medical images are often limited in its application due to the large radiation doses and longer image scanning and reconstruction times. Existing methods attempt to reduce the volume samples in the dynamic sequence by interpolating the volumes between the acquired samples. However, these methods are limited to either 2D images and/or are unable to support large but periodic variations in the functional motion between the image volume samples. In this paper, we present a spatiotemporal volumetric interpolation network (SVIN) designed for 4D dynamic medical images. SVIN introduces dual networks: the first is the spatiotemporal motion network that leverages the 3D convolutional neural network (CNN) for unsupervised parametric volumetric registration to derive spatiotemporal motion field from a pair of image volumes; the second is the sequential volumetric interpolation network, which uses the derived motion field to interpolate image volumes, together with a new regression-based module to characterize the periodic motion cycles in functional organ structures. We also introduce an adaptive multi-scale architecture to capture the volumetric large anatomy motions. Experimental results demonstrated that our SVIN outperformed state-of-the-art temporal medical interpolation methods and natural video interpolation method that has been extended to support volumetric images. Code is available at [1]. Yuyu Guo 0002, Lei Bi 0001, Euijoon Ahn, David Dagan Feng, Qian Wang 0001, Jinman Kim |
CVPR | 2 |
| 2020 | Unsupervised Positron Emission Tomography Tumor Segmentation via GAN based Adversarial Auto-EncoderabstractFluorodeoxyglucose Positron emission tomography (FDG PET) is the imaging modality of choice for the diagnosis of lung cancer. The automated segmentation of tumors in PET images is a fundamental requirement for image analysis in computer aided diagnosis systems. Current tumor segmentation in PET generally relies on local features to discriminate tumor from the background. These methods are limited due to poor resolution, and subtle inter-class differences when there is normal FDG uptake region (i.e., in the heart and mediastinum) in the same field of view. We propose a new image based discriminative method to separate tumor regions from normal regions. We introduce a convolutional adversarial auto-encoder to learn a latent space which models normal (disease-free) variations of PET images, and then to compute a residual map that identifies where the PET image differs from this manifold due to anomalies, i.e., tumors. Our method is tolerant to normal intra-class variations among the PET images but is discriminative of the tumors with high sensitivity. Our experiments with a clinical lung cancer dataset show that our method outperformed the state-of-the-art unsupervised segmentation methods. We also achieved higher dice score (62.0%) and sensitivity (77.9%) than the supervised U-Net method (59.5% and 59.7%). Xinheng Wu, Lei Bi 0001, Michael J. Fulham, Jinman Kim |
ICARCV | 2 |
| 2020 | Malocclusion Treatment Planning via PointNet Based Spatial Transformation Network
Xiaoshuang Li, Lei Bi 0001, Jinman Kim, Tingyao Li, Peng Li 0079, Bin Sheng 0001, David Dagan Feng |
MICCAI (3) | 2 |
| 2020 | Multi-modality Information Fusion for Radiomics-Based Neural Architecture Search
Yige Peng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (7) | 2 |
| 2020 | Multi-Label classification of multi-modality skin lesion via hyper-connected convolutional neural network
Lei Bi 0001, David Dagan Feng, Michael J. Fulham, Jinman Kim |
Pattern Recognit. | 1 |
| 2019 | Optic Disc and Cup Segmentation Based on Enhanced SegNetabstractDue to imbalanced distributed and restricted medical resources, reliable analysis for medical images is hard to come by, and it is impractical to only rely on human beings to do all the analysis, which is time-consuming and not economic. Application of computer vision techniques in such fields emerges as the situation requires. In this paper, we use deep learning segmentation algorithm to segment the optic disc and the cup from each other and from the rest of the ophthalmoscopy photographs. For a better performance, we change the loss function and crop as a way of data augmentation. The segmentation results can be used to calculate the cup-to-disc ratio (CDR), which is further used to diagnose glaucoma. Challenges such as over-fitting, biased dataset, and poor generalization of the model exist in front of us. We illustrate our model and associated methods dealing with these challenges. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 6 |
| 2019 | Detect Glaucoma with Image Segmentation and Transfer LearningabstractIn this paper, we aim to automatically detect glaucoma via deep learning. To do that, we need to calculate the cup-to-disc ratio (CDR) on fine segmented retina images. To get precise segmentation, we implemented SegNet together with adversarial discriminative domain adaptation (ADDA), the former is a famous artificial neural network with encoder-decoder architecture used in image segmentation area and the latter is a transfer learning method for domain adaptation. We are the first to combine them together to detect glaucoma on test dataset which have different brightness from our training dataset. We thoroughly evaluated the proposed method with various loss functions, normal cross entropy loss, weighted cross entropy loss and dice coefficient loss included. And we show that dice loss is the best for this task. Last but not least, our experiments on transfer learning have shown that our ADDA method reduces the mean square error (MSE) between the CDR of our segmentation and annotations greatly. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 6 |
| 2019 | Deep Intrinsic Image Decomposition Using Joint Parallel Learning
Yuan Yuan 0022, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Enhua Wu |
CGI | 4 |
| 2019 | Deep Local-Global Refinement Network for Stent Analysis in IVOCT Images
Yuyu Guo 0002, Lei Bi 0001, Ashnil Kumar, Yue Gao 0002, Ruiyan Zhang, David Dagan Feng, Qian Wang 0001, Jinman Kim |
MICCAI (5) | 2 |
| 2019 | Step-wise integration of deep class-specific learning for dermoscopic image segmentation
Lei Bi 0001, Jinman Kim, Euijoon Ahn, Ashnil Kumar, David Dagan Feng, Michael J. Fulham |
Pattern Recognit. | 1 |
| 2018 | Dual-Path Adversarial Learning for Fully Convolutional Network (FCN)-Based Medical Image Segmentation
Lei Bi 0001, David Dagan Feng, Jinman Kim |
Vis. Comput. | 1 |
| 2017 | Saliency-Based Lesion Segmentation Via Background Detection in Dermoscopic ImagesabstractThe segmentation of skin lesions in dermoscopic images is a fundamental step in automated computer-aided diagnosis of melanoma. Conventional segmentation methods, however, have difficulties when the lesion borders are indistinct and when contrast between the lesion and the surrounding skin is low. They also perform poorly when there is a heterogeneous background or a lesion that touches the image boundaries; this then results in under- and oversegmentation of the skin lesion. We suggest that saliency detection using the reconstruction errors derived from a sparse representation model coupled with a novel background detection can more accurately discriminate the lesion from surrounding regions. We further propose a Bayesian framework that better delineates the shape and boundaries of the lesion. We also evaluated our approach on two public datasets comprising 1100 dermoscopic images and compared it to other conventional and state-of-the-art unsupervised (i.e., no training required) lesion segmentation methods, as well as the state-of-the-art unsupervised saliency detection methods. Our results show that our approach is more accurate and robust in segmenting lesions compared to other methods. We also discuss the general extension of our framework as a saliency optimization algorithm for lesion segmentation. Euijoon Ahn, Jinman Kim, Lei Bi 0001, Ashnil Kumar, ChangYang Li, Michael J. Fulham, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | Stacked fully convolutional networks with multi-channel learning: application to medical image segmentation
Lei Bi 0001, Jinman Kim, Ashnil Kumar, Michael J. Fulham, David Dagan Feng |
Vis. Comput. | 1 |
| 2014 | Multi-stage Thresholded Region Classification for Whole-Body PET-CT Lymphoma Studies
Lei Bi 0001, Jinman Kim, David Dagan Feng, Michael J. Fulham |
MICCAI (1) | 1 |
| 2013 | A web-based medical multimedia visualisation interface for personal health recordsabstractThe healthcare industry has begun to utilise web-based systems and cloud computing infrastructure to develop an increasing array of online personal health record (PHR) systems. Although these systems provide the technical capacity to store and retrieve medical data in various multimedia formats, including images, videos, voice, and text, individual patient use remains limited by the lack of intuitive data representation and visualisation techniques. As such, further research is necessary to better visualise and present these records, in ways that make the complex medical data more intuitive. In this study, we present a web-based PHR visualisation system, called the 3D medical graphical avatar (MGA), which was designed to explore web-based delivery of a wide array of medical data types including multi-dimensional medical images; medical videos; text-based data; and spatial annotations. Mapping information was extracted from each of the data types and was used to embed spatial and textual annotations, such as regions of interest (ROIs) and time-based video annotations. Our MGA itself is built from clinical patient imaging studies, when available. We have taken advantage of the emerging web technologies of HTML5 and WebGL to make our application available to a wider base of users and devices. We analysed the performance of our proof-of-concept prototype system on mobile and desktop consumer devices. Our initial experiments indicate that our system can render the medical data in a fashion that enables interactive navigation of the MGA. Michael de Ridder, Liviu Constantinescu, Lei Bi 0001, Younhyun Jung, Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham |
CBMS | 3 |