Hongjie Hu

dblp:89/7908 · DBLP profile ↗
← Back
43ranked-venue papers
1as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 15 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Combined myocardial motion and texture characterisation methods for the phenotyping of scarred myocardium
Yaming Wang, Daiguo Yang, Cailing Pu, Xiaowei Ruan, Chengjin Yu, Dongsheng Ruan, Mingfeng Jiang, Hongjie Hu, Huafeng Liu 0003
Expert Syst. Appl.9
2026 An improved multi-instance learning model with clinical-guided cross-attention for postoperative early recurrence prediction of hepatocellular carcinoma using histopathological images
Gan Zhan, Fang Wang 0030, Yinhao Li 0002, Rahul Kumar Jain 0001, Qingqing Chen 0001, Lanfen Lin, Hongjie Hu, C. Krishna Mohan, Yen-Wei Chen 0001
Neurocomputing9
2026 Multimodal Graph Learning With Multi-Hypergraph Reasoning Networks for Focal Liver Lesion Classification in Multimodal Magnetic Resonance Imaging
abstract
Multimodal magnetic resonance imaging (MRI) is instrumental in differentiating liver lesions. The major challenge involves modeling reliable connections and simultaneously learning complementary information across various MRI sequences. While previous studies have primarily focused on multimodal integration in a pair-wise manner using few modalities, our research seeks to advance a more comprehensive understanding of interaction modeling by establishing complex high-order correlations among the diverse modalities in multimodal MRI. In this paper, we introduce a multimodal graph learning with multi-hypergraph reasoning network to capture the full spectrum of both pair-wise and group-wise relationships among different modalities. Specifically, a weight-shared encoder extracts features from regions of interest (ROI) images across all modalities. Subsequently, a collection of uniform hypergraphs are constructed with varying vertex configurations, allowing for the modeling of not only pair-wise correlations but also the high-order collaborations for relational reasoning. Following information propagation through the hypergraph message passing, adaptive intra-modality fusion module is proposed to effectively fuse feature representations from different hypergraphs of the same modality. Finally, all refined features are concatenated to prepare for the classification task. Our experimental evaluations, including focal liver lesions classification using the LLD-MMRI2023 dataset and early recurrence prediction of hepatocellular carcinoma using our internal datasets, demonstrate that our method significantly surpasses the performance of existing approaches, indicating the effectiveness of our model in handling both pair-wise and group-wise interactions across multiple modalities.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Fang Wang 0030, Qingqing Chen 0001, Wenbin Ji, Yinhao Li 0002, Hongjie Hu, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics9
2025 2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images
abstract
Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instance selection. Thus, they cannot capture weak visual cues related to PD-L1 status on CT images. To address this, we propose a 2.5D top-K ranked multiple instance learning method. We design a 2.5D instance feature encoder, which takes advantage of knowledge from a 2D pre-trained model and has trainable parameters to learn information for 3D instances. In addition, we design a top-K ranked multiple instance learning strategy, which fully exploits the bag-level labels to select representative instances to eliminate the effect of atypical instances and guide the network to learn effective information. We demonstrate that our method can not only outperform state-of-the-art MIL methods on the PD-L1 status classification but also generalize well on a COVID-19 classification task.
Huadong Liu, Yongcen Li, Xinchen Ye, Hongkai Wang 0002, Yi Wang 0037, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu
ICASSP13
2024 Novelty Detection Based Discriminative Multiple Instance Feature Mining to Classify NSCLC PD-L1 Status on HE-Stained Histopathological Images
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yi Wang 0037, Hongkai Wang 0002, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu
MICCAI (4)13
2024 Segmentation Guided Crossing Dual Decoding Generative Adversarial Network for Synthesizing Contrast-Enhanced Computed Tomography Images
abstract
Although contrast-enhanced computed tomography (CE-CT) images significantly improve the accuracy of diagnosing focal liver lesions (FLLs), the administration of contrast agents imposes a considerable physical burden on patients. The utilization of generative models to synthesize CE-CT images from non-contrasted CT images offers a promising solution. However, existing image synthesis models tend to overlook the importance of critical regions, inevitably reducing their effectiveness in downstream tasks. To overcome this challenge, we propose an innovative CE-CT image synthesis model called Segmentation Guided Crossing Dual Decoding Generative Adversarial Network (SGCDD-GAN). Specifically, the SGCDD-GAN involves a crossing dual decoding generator including an attention decoder and an improved transformation decoder. The attention decoder is designed to highlight some critical regions within the abdominal cavity, while the improved transformation decoder is responsible for synthesizing CE-CT images. These two decoders are interconnected using a crossing technique to enhance each other's capabilities. Furthermore, we employ a multi-task learning strategy to guide the generator to focus more on the lesion area. To evaluate the performance of proposed SGCDD-GAN, we test it on an in-house CE-CT dataset. In both CE-CT image synthesis tasks-namely, synthesizing ART images and synthesizing PV images-the proposed SGCDD-GAN demonstrates superior performance metrics across the entire image and liver region, including SSIM, PSNR, MSE, and PCC scores. Furthermore, CE-CT images synthetized from our SGCDD-GAN achieve remarkable accuracy rates of 82.68%, 94.11%, and 94.11% in a deep learning-based FLLs classification task, along with a pilot assessment conducted by two radiologists.
Qingqing Chen 0001, Yinhao Li 0002, Fang Wang 0030, Xianhua Han, Yutaro Iwamoto, Jing Liu 0041, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics9
2024 Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level Teacher
abstract
Multiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based.
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
IEEE Trans. Medical Imaging6
2023 Pixel-Correlation-Based Scar Screening in Hypertrophic Myocardium
Cailing Pu, Chengjin Yu, Yuan-Ting Yan, Hongjie Hu, Huafeng Liu 0003
ICIG (5)5
2023 Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
MICCAI (6)6
2023 Adaptive Decomposition and Shared Weight Volumetric Transformer Blocks for Efficient Patch-Free 3D Medical Image Segmentation
abstract
High resolution (HR) 3D medical image segmentation is vital for an accurate diagnosis. However, in the field of medical imaging, it is still a challenging task to achieve a high segmentation performance with cost-effective and feasible computation resources. Previous methods commonly use patch-sampling to reduce the input size, but this inevitably harms the global context and decreases the model's performance. In recent years, a few patch-free strategies have been presented to deal with this issue, but either they have limited performance due to their over-simplified model structures or they follow a complicated training process. In this study, to effectively address these issues, we present Adaptive Decomposition (A-Decomp) and Shared Weight Volumetric Transformer Blocks (SW-VTB). A-Decomp can adaptively decompose features and reduce their spatial size, which greatly lowers GPU memory consumption. SW-VTB is able to capture long-range dependencies at a low cost with its lightweight design and cross-scale weight-sharing mechanism. Our proposed cross-scale weight-sharing approach enhances the network's ability to capture scale-invariant core semantic information in addition to reducing parameter numbers. By combining these two designs together, we present a novel patch-free segmentation framework named VolumeFormer. Experimental results on two datasets show that VolumeFormer outperforms existing patch-based and patch-free methods with a comparatively fast inference speed and relatively compact design.
Hongyi Wang 0002, Qingqing Chen 0001, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin
IEEE J. Biomed. Health Informatics6
2023 Multi-Modal Tumor Segmentation With Deformable Aggregation and Uncertain Region Inpainting
abstract
Multi-modal tumor segmentation exploits complementary information from different modalities to help recognize tumor regions. Known multi-modal segmentation methods mainly have deficiencies in two aspects: First, the adopted multi-modal fusion strategies are built upon well-aligned input images, which are vulnerable to spatial misalignment between modalities (caused by respiratory motions, different scanning parameters, registration errors, etc). Second, the performance of known methods remains subject to the uncertainty of segmentation, which is particularly acute in tumor boundary regions. To tackle these issues, in this paper, we propose a novel multi-modal tumor segmentation method with deformable feature fusion and uncertain region refinement. Concretely, we introduce a deformable aggregation module, which integrates feature alignment and feature aggregation in an ensemble, to reduce inter-modality misalignment and make full use of cross-modal information. Moreover, we devise an uncertain region inpainting module to refine uncertain pixels using neighboring discriminative features. Experiments on two clinical multi-modal tumor datasets demonstrate that our method achieves promising tumor segmentation results and outperforms state-of-the-art methods.
Yue Zhang 0042, Chengtao Peng, Ruofeng Tong 0001, Lanfen Lin, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging7
2022 Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions
abstract
Automatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has shown promising results in [1], its performance is relatively poor. We address this problem by proposing a pixel-level and affinity-level knowledge distillation method. It obtains a pre-trained teacher network with rich semantic knowledge of CT images by constructing and training an auto-encoder at first, and then trains a student network with the same architecture as the teacher by distilling the teacher’s knowledge only from normal CT images, and finally localizes COVID-19 lesions using the feature discrepancy between the teacher and the student networks. Besides, except for the traditional pixel-level distillation, we design the affinity-level distillation that takes into account the pairwise relationship of features to fully distill effective knowledge. We evaluate this method by using three different COVID-19 datasets and the experimental results show that the segmentation performance is largely improved when it is compared with the other existing unsupervised anomaly detection methods.
Rui Xu 0002, Xinchen Ye, Yen-Wei Chen 0001, Fangyi Xu, Wenchao Zhu, Hongjie Hu, Xiaofeng Qu, Shoji Kido, Noriyuki Tomiyama
ICASSP10
2022 An Accurate Unsupervised Liver Lesion Detection Method Using Pseudo-lesions
He Li 0042, Yutaro Iwamoto, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
MICCAI (8)5
2022 Mutual Information-Based Graph Co-Attention Networks for Multimodal Prior-Guided Magnetic Resonance Imaging Segmentation
abstract
Multimodal magnetic resonance imaging (MRI) provides complementary information about targets, and the segmentation of multimodal MRI is widely used as an essential preprocessing step for initial diagnosis, stage differentiation, and post-treatment efficacy evaluation in clinical situations. For the main modality or each of the modalities, it is important to enhance the visual information by modeling the connection and effectively fusing the features among them. However, the existing methods for multimodal segmentation have a drawback; they coincidentally drop information of individual modality during the fusion process. Recently, graph learning-based methods have been applied in segmentation, and these methods have achieved considerable improvements by modeling the relationships across feature regions and reasoning using global information. In this paper, we propose a graph learning-based approach to efficiently extract modality-specific features and establish regional correspondence effectively among all modalities. In detail, after projecting features into a graph domain and employing graph convolution to propagate information across all regions for learning global modality-specific features, we propose a mutual information-based graph co-attention module to learn the weight coefficients of one bipartite graph constructed by the fully connected graphs having different modalities in the graph domain and by selectively fusing the node features. Based on the deformation diagram between the spatial-graph space and our proposed graph co-attention module, we present a multimodal prior-guided segmentation framework, which uses two strategies for two clinical situations:Modality-Specific Learning StrategyandCo-Modality Learning Strategy. Besides, the improvedCo-Modality Learning Strategyis used with trainable weights in the multi-task loss for the optimization of the proposed framework. We validated our proposed modules and frameworks on two multimodal MRI datasets: our private liver lesion dataset and a public prostate zone dataset. Our experimental results on both datasets prove the superiority of our proposed approaches.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
IEEE Trans. Circuits Syst. Video Technol.7
2022 MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical Images
abstract
Organ segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods.
Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu
IEEE J. Biomed. Health Informatics12
2022 DeepRecS: From RECIST Diameters to Precise Liver Tumor Segmentation
abstract
Liver tumor segmentation (LiTS) is of primary importance in diagnosis and treatment of hepatocellular carcinoma. Known automated LiTS methods could not yield satisfactory results for clinical use since they were hard to model flexible tumor shapes and locations. In clinical practice, radiologists usually estimate tumor shape and size by a Response Evaluation Criteria in Solid Tumor (RECIST) mark. Inspired by this, in this paper, we explore a deep learning (DL) based interactive LiTS method, which incorporates guidance from user-provided RECIST marks. Our method takes a three-step framework to predict liver tumor boundaries. Under this architecture, we develop a RECIST mark propagation network (RMP-Net) to estimate RECIST-like marks in off-RECIST slices. We also devise a context-guided boundary-sensitive network (CGBS-Net) to distill tumors' contextual and boundary information from corresponding RECIST(-like) marks, and then predict tumor maps. To further refine the segmentation results, we process the tumor maps using a 3D conditional random field (CRF) algorithm and a morphology hole-filling operation. Verified on two clinical contrast-enhanced abdomen computed tomography (CT) image datasets, our proposed approach can produce promising segmentation results, and outperforms the state-of-the-art interactive segmentation methods.
Yue Zhang 0042, Chengtao Peng, Liying Peng, Lanfen Lin, Ruofeng Tong 0001, Zhiyi Peng, Xiongwei Mao, Hongjie Hu, Yen-Wei Chen 0001, Jingsong Li 0001
IEEE J. Biomed. Health Informatics9
2021 Corona Virus Disease (COVID-19) Detection in CT Images Using Synergic Deep Learning
Yiwei Gao, Hongjie Hu, Huafeng Liu 0003
ICIG (2)2
2021 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (2)4
2021 Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (1)3
2021 Multi-phase Liver Tumor Segmentation with Spatial Aggregation and Uncertain Region Inpainting
Yue Zhang 0042, Chengtao Peng, Liying Peng, Huimin Huang 0002, Ruofeng Tong 0001, Lanfen Lin, Jingsong Li 0001, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Zhiyi Peng
MICCAI (1)10
2021 A Tensor Sparse Representation-Based CBMIR System for Computer-Aided Diagnosis of Focal Liver Lesions and its Pilot Trial
abstract
Clinicians refer to diagnosed medical cases in order to make correct diagnosis and take appropriate treatments, due to the complexity of focal liver lesions. It's a heavy burden, however, for medical doctors to find out similar and meaningful cases from the accumulated extreme large medical datasets. Content based medical image retrieval (CBMIR) that searches for similar images in a large database has been attracting increasing research interest recently. A CBMIR system provides doctors the diagnosed cases to improve the diagnosis accuracy and confidence. This paper proposed a tensor sparse representation method to extract temporal and spatial features of multi-phase CT images, so as to provide doctors medical cases more relevant to the query one. The proposed tensor sparse representation method is applied to the retrieval of focal liver lesions (FLLs). Experiments show that the proposed method achieved better retrieval performance than conventional methods. Pilot trial was conducted and results show that diagnosis accuracy and confidence was improved significantly by the developed CBMIR system based on the proposed method.
Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
ICMR4
2021 M-DFNet: Multi-phase Discriminative Feature Network for Retrieval of Focal Liver Lesions
abstract
Content based medical image retrieval (CBMIR) plays a great role in computer aided diagnosis for assisting radiologists to detect and characterize focal liver lesions (FLLs). Deep learning has gained exciting performance on CBMIR. While the features generated by deep learning models trained using softmax loss are always separable but not discriminative enough, which is insufficient for retrieval task. In this paper, we propose a multi-phase discriminative feature network (M-DFNet) with a DeepExtracter and a feature refine module (FRModule) to learn discriminative and separable features under a joint supervision of center loss and softmax loss. The hybrid loss enables to minimize intra-class variations and enlarge inter-class differences as much as possible. The FRModule is proposed to recalibrate the deep features based on the learned class centers to tackle the complex imaging manifestations of FLLs and further enhance both the feature discrimination and generalization. Multi-phase computed tomography (CT) images contain pivotal information for diagnosis of FLLs. Thus the M-DFNet is designed to cope with multi-phase information and we explore an appropriate and effective method for multi-phase feature integration on limited data. Experimental results clearly demonstrate strong performance superiority by our proposed method.
Jing Liu 0041, Lanfen Lin, Hongjie Hu, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
ICMR4
2021 Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19
abstract
COVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy.
Titinunt Kitrungrotsakul, Qingqing Chen 0001, Huitao Wu, Yutaro Iwamoto, Hongjie Hu, Wenchao Zhu, Fangyi Xu, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics5
2021 Medical Image Segmentation With Deep Atlas Prior
abstract
Organ segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation.
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
IEEE Trans. Medical Imaging5
2020 UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
abstract
Recently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version.
Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001
ICASSP4
2020 Unsupervised Detection of Pulmonary Opacities for Computer-Aided Diagnosis of COVID-19 on CT Images
abstract
COVID-19 emerged towards the end of 2019 which was identified as a global pandemic by the world heath organization (WHO). With the rapid spread of COVID-19, the number of infected and suspected patients has increased dramatically. Chest computed tomography (CT) has been recognized as an efficient tool for the diagnosis of COVID-19. However, the huge CT data make it difficult for radiologist to fully exploit them on the diagnosis. In this paper, we propose a computer-aided diagnosis system that can automatically analyze CT images to distinguish the COVID-19 against to community-acquired pneumonia (CAP). The proposed system is based on an unsupervised pulmonary opacity detection method that locates opacity regions by a detector unsupervisedly trained from CT images with normal lung tissues. Radiomics based features are extracted insides the opacity regions, and fed into classifiers for classification. We evaluate the proposed CAD system by using 200 CT images collected from different patients in several hospitals. The accuracy, precision, recall, f1-score and AUC achieved are 95.5%, 100%, 91%, 95.1% and 95.9% respectively, exhibiting the promising capacity on the differential diagnosis of COVID-19 from CT images.
Rui Xu 0002, Xiao Cao, Yen-Wei Chen 0001, Xinchen Ye, Lin Lin 0008, Wenchao Zhu, Fangyi Xu, Hongjie Hu, Shoji Kido, Noriyuki Tomiyama
ICPR11
2020 Multimodal Priors Guided Segmentation of Liver Lesions in MRI Using Mutual Information Based Graph Co-Attention Networks
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
MICCAI (4)7
2020 Tensor-based sparse representations of multi-phase medical images for classification of focal liver lesions
Jian Wang 0004, Jing Li 0046, Xianhua Han, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yutaro Iwamoto, Yen-Wei Chen 0001
Pattern Recognit. Lett.5
2020 Semi-Supervised Learning for Semantic Segmentation of Emphysema With Partial Annotations
abstract
Segmentation and quantification of each subtype of emphysema is helpful to monitor chronic obstructive pulmonary disease. Due to the nature of emphysema (diffuse pulmonary disease), it is very difficult for experts to allocate semantic labels to every pixel in the CT images. In practice, partially annotating is a better choice for the radiologists to reduce their workloads. In this paper, we propose a new end-to-end trainable semi-supervised framework for semantic segmentation of emphysema with partial annotations, in which a segmentation network is trained from both annotated and unannotated areas. In addition, we present a new loss function, referred to as Fisher loss, to enhance the discriminative power of the model and successfully integrate it into our proposed framework. Our experimental results show that the proposed methods have superior performance over the baseline supervised approach (trained with only annotated areas) and outperform the state-of-the-art methods for emphysema segmentation.
Liying Peng, Lanfen Lin, Hongjie Hu, Yue Zhang 0042, Huali Li, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics3
2019 A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT Images
abstract
Automatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks.
Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICIP5
2019 Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct Images
abstract
Convolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1.
Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
ICIP6
2019 Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001
MICCAI (6)3
2019 Classification and Quantification of Emphysema Using a Multi-Scale Residual Network
abstract
Automated tissue classification is an essential step for quantitative analysis and treatment of emphysema. Although many studies have been conducted in this area, there still remain two major challenges. First, different emphysematous tissue appears in different scales, which we call "inter-class variations." Second, the intensities of CT images acquired from different patients, scanners or scanning protocols may vary, which we call "intra-class variations". In this paper, we present a novel multi-scale residual network with two channels of raw CT image and its differential excitation component. We incorporate multi-scale information into our networks to address the challenge of inter-class variations. In addition to the conventional raw CT image, we use its differential excitation component as a pair of inputs to handle intra-class variations. Experimental results show that our approach has superior performance over the state-of-the- art methods, achieving a classification accuracy of 93.74% on our original emphysema database. Based on the classification results, we also perform the quantitative analysis of emphysema in 50 subjects by correlating the quantitative results (the area percentage of each class) with pulmonary functions. We show that centrilobular emphysema (CLE) and panlobular emphysema (PLE) have strong correlation with the pulmonary functions and the sum of CLE and PLE can be used as a new and accurate measure of emphysema severity instead of the conventional measure (sum of all subtypes of emphysema). The correlations between the new measure and various pulmonary functions are up to |r| = 0.922 (r is correlation coefficient).
Liying Peng, Yen-Wei Chen 0001, Lanfen Lin, Hongjie Hu, Huali Li, Qingqing Chen 0001, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto
IEEE J. Biomed. Health Informatics4
2018 Classification of Pulmonary Emphysema in CT Images Based on Multi-Scale Deep Convolutional Neural Networks
abstract
In this work, we aim at classifying emphysema in computed tomography (CT) images of lungs. Most previous works are limited to extracting low-level features or mid-level features without enough high-level information. Moreover, these approaches do not take the characteristics (scales) of different emphysema into account, which are crucial for feature extraction. In contrast to previous works, we propose a novel deep learning method based on multiscale deep convolutional neural networks. There are three contributions for this paper. First, we propose to use a base residual network with 20 layers to extract more high-level information. To the best of our knowledge, this is the first deep learning method for classification of emphysema. Second, we incorporate multi-scale information into our deep neural networks so as to take full consideration of the characteristics of different emphysema. Finally, we established a high-quality emphysema dataset which contains 91 high-resolution computed tomography (HRCT) volumes, annotated manually by two experienced radiologists and checked by one experienced chest radiologist. A 92.68% classification accuracy is achieved on this dataset. The results show that (1) the multi-scale method is highly effective in comparison to the single scale setting; (2) the proposed approach is superior to the state-of-the-art techniques.
Liying Peng, Lanfen Lin, Hongjie Hu, Huali Li, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto, Yen-Wei Chen 0001
ICIP3
2018 Combining Convolutional and Recurrent Neural Networks for Classification of Focal Liver Lesions in Multi-phase CT Images
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
MICCAI (2)3
2018 Residual Convolutional Neural Networks with Global and Local Pathways for Classification of Focal Liver Lesions
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
PRICAI (1)3
2017 Joint weber-based rotation invariant uniform local ternary pattern for classification of pulmonary emphysema in CT images
abstract
In this paper, we present a novel image representation approach for classifying emphysema in computed tomography (CT) images of the lung. Our proposed method extends rotation invariant uniform local binary pattern (RIULBP) and local ternary pattern (LTP), which are extensively used in a variety of computer vision applications, into rotation invariant uniform local ternary pattern (RIULTP) with a human perception principle: Weber's law. In addition, by integrating the upper pattern and the lower pattern of the Weber-based RIULTP (WRIULTP), we further put forward the joint Weber-based rotation invariant uniform local ternary pattern (JWRIULTP), which allows for a much richer representation and also takes the comprehensive information of the image into account. The proposed methods are tested on the Outex database (texture database) and the Bruijne and Srensen database (emphysema database). The results show the superiority of the proposed approaches to the state-of-the-art techniques for emphysema classification including rotation invariant local binary pattern (RILBP) and texton-based approach.
Liying Peng, Lanfen Lin, Hongjie Hu, Xiaoli Ling, Xianhua Han, Yen-Wei Chen 0001
ICIP3
2017 Tensor Sparse Representation of Temporal Features for Content-Based Retrieval of Focal Liver Lesions Using Multi-phase Medical Images
abstract
Content Based Image Retrieval (CBIR) systems that search similar images in a large database are attracting more and more research interests recently, and have been applied to medical image characterization for expert's experience sharing. One challenging task in CBIR is how to extract features for effective image representation. Therein sparse coding technique has been proven to be an effective way to learn inherent structure features for image analysis. However, it is necessary to first vectorize the 2- or 3-dimensional spatial structure for analysis with sparse coding, and then destroy the spatial relation of nearby voxels. In this study, we propose a multilinear sparse coding method to learn features from multi-dimensional medical images. We regard high dimensional local structures as tensors and propose a K-CP (CANDECOMP/PARAFAC) algorithm to learn a tensor dictionary in an iterative way. With the learned tensor dictionary, sparse coefficients of tensor local structures are calculated by multilinear orthogonal matching pursuit (MOMP) algorithm, which is an extended multilinear version of the conventional linear OMP. The proposed multilinear sparse coding method is prospected to be more efficient and effective for inherent feature extraction compared with conventional linear methods. The proposed method is applied to a CBIR system for retrieval of focal liver lesions (FLLs) using a medical database consisting of contrast-enhanced multi-phase computer-tomography (CT) images. Experiments show that the constructed CBIR with multilinear sparse coding method can achieve promising retrieval performance.
Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Chongwu Jin, Yen-Wei Chen 0001
ISM5
2017 A Sub 6GHz Massive MIMO System for 5G New Radio
abstract
This paper introduces a Massive MIMO system for 5G New Radio (NR). Compared with LTE, which is initially designed for a single wide beam per cell, the system in the paper is natively designed for multiple narrow beams to fully utilize the benefits of Massive MIMO. Beside Massive MIMO, other 5G features, such as wide bandwidth and low latency, introduce additional challenges to hardware implantations. A prototype based on commercial base station platform is developed to support Massive MIMO, wide bandwidth and low latency at the same time. The prototype achieved high single user and multiple users' peak data rates in the field test. The work is helpful to understand the potential challenges in 5G air interface design and demonstrate Massive MIMO can be a viable solution for 5G NR.
Hongjie Hu, Zhongfeng Li, Youtuan Zhu
VTC Spring1
2016 Bag of temporal co-occurrence words for retrieval of focal liver lesions using 3D multiphase contrast-enhanced CT images
abstract
Computer-aided diagnosis (CAD) systems have been verified to have the potential to assist radiologists in clinical diagnosis to detect and characterize focal liver lesions (FLLs) based on single- or multiphase contrast-enhanced computed tomography (CT) images. Features extracted from multiphase contrast-enhanced CT images carry more important diagnostic information i.e. enhancement pattern and demonstrate much stronger discriminative ability compared to those of single-phase CT images. In this paper, we propose a new method for multiphase image feature generation called the bag of temporal co-occurrence words (BoTCoW). A temporal co-occurrence image connecting intensity from multiphase images is constructed. Then the bag of visual word (BoVW) model is employed on the temporal co-occurrence images to extract temporal features. The proposed method effectively captures temporal enhancement information and demonstrates the distribution of the evolution patterns. The effectiveness of this method is validated in a retrieval system using 132 FLLs with confirmed pathology type. The preliminary results show that the proposed BoTCoW method outperforms the previously proposed temporal features and multiphase features based on the BoVW model.
Lanfen Lin, Hongjie Hu, Yitao Liu, Jian Wang 0004, Xianhua Han, Yen-Wei Chen 0001
ICPR3
2014 MAP Based Iterative Channel Estimation for OFDM Systems: Approach, Convergence, and Performance Bound
abstract
Iterative channel estimation (ICE) usually exploits soft information of unknown data symbols as references to improve estimation performance. This paper investigates ICE for orthogonal frequency division multiplexing (OFDM) over wireless channels. The optimum ICE is derived in terms of maximum a posteriori (MAP) criterion, which can be solved using fixed-point iteration (FPI). Furthermore, the derived MAP ICE is closely related to the well-known expectation-maximization (EM) estimation. We also demonstrate that the MAP ICE converges within only one step when the signal-to-noise ratio (SNR) is large through analysis and simulation results.
Yinsheng Liu, Geoffrey Ye Li, Hongjie Hu, Zhenhui Tan
IEEE Trans. Wirel. Commun.3
2014 MAP-Based Iterative Channel Estimation for OFDM With Multiple Transmit Antennas Over Time-Varying Channels
abstract
This paper investigates iterative channel estimation (ICE) for orthogonal frequency-division multiplexing (OFDM) with multiple transmit antennas. To improve performance of channel estimation, we exploit the soft information of unknown data symbols on both the expected transmit antenna and the interfering transmit antenna. Maximum a posteriori (MAP)-based ICE is derived and is implemented using the fixed-point iteration (FPI). For an OFDM system with multiple transmit antennas, the proposed MAP-based ICE suggests a harmonic-average-based soft symbol on the expected transmit antenna while an arithmetic-average-based soft symbol on the interfering transmit antennas. Similar to an OFDM system with a single transmit antenna, MAP-based ICE can achieve the Cramer-Rao bound (CRB) within only one iteration, when the signal-to-noise ratio (SNR) is large enough.
Yinsheng Liu, Geoffrey Ye Li, Hongjie Hu, Zhenhui Tan
IEEE Trans. Wirel. Commun.3
2009 IBI Cancellation Based on Limited Channel Feedback for OFDM Systems over Channels with Large Delay Spreads
abstract
While cyclic prefix (CP) is no larger than the delay span of wireless channels in an orthogonal frequency division multiplexing (OFDM) system, inter-block interference (IBI) and inter-carrier interference (ICI) will occur and the performance of the system will be deteriorated. If OFDM is designed with long enough CP, the efficiency of OFDM modulation is significantly reduced. To effectively mitigate interference over channel with large delay spreads, pre-processing approaches at transmitters based on full channel state information (CSI) have been proposed. However, feeding full CSI back will occupy significant bandwidth and is impractical sometimes. Therefore, pre-processing optimization exploiting limited CSI feedback is investigated in this paper. An pre-processing optimization approaches for IBI and ICI mitigation has been developed and tested. Computer simulation result shows that the proposed optimization algorithm can effectively cancel IBI and ICI over channels with large delay spreads and significantly improve the performance for OFDM systems.
Geoffrey Ye Li, Hongjie Hu, Anthony C. K. Soong
GLOBECOM3