EDBT 2026 Demo / reviewers in the wild / expert
Yinhao Li 0002
dblp:210/7374-2
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-8924-5279ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An improved multi-instance learning model with clinical-guided cross-attention for postoperative early recurrence prediction of hepatocellular carcinoma using histopathological images
Gan Zhan, Fang Wang 0030, Yinhao Li 0002, Rahul Kumar Jain 0001, Qingqing Chen 0001, Lanfen Lin, Hongjie Hu, C. Krishna Mohan, Yen-Wei Chen 0001 |
Neurocomputing | 3 |
| 2026 | SPA: Leveraging the SAM With Spatial Priors Adapter for Enhanced Medical Image SegmentationabstractThe Segment Anything Model (SAM) has gained renown for its success in image segmentation, benefiting significantly from its pretraining on extensive datasets and its interactive prompt-based segmentation approach. Although highly effective in natural (real-world) image segmentation tasks, the SAM model encounters significant challenges in medical imaging due to the inherent differences between these two domains. To address these challenges, we propose the Spatial Prior Adapter (SPA) scheme, a parameter-efficient fine-tuning strategy that enhances SAM's adaptability to medical imaging tasks. SPA introduces two novel modules: the Spatial Prior Module (SPM), which captures localized spatial features through convolutional layers, and the Feature Communication Module (FCM), which integrates these features into SAM's image encoder via cross-attention mechanisms. Furthermore, we develop a Multiscale Feature Fusion Module (MSFFM) to enhance SAM's end-to-end segmentation capabilities by effectively aggregating multiscale contextual information. These lightweight modules require minimal computational resources while significantly boosting segmentation performance. Our approach demonstrates superior performance in both prompt-based and end-to-end segmentation scenarios through extensive experiments on publicly available medical imaging datasets. Performance highlights the potential of the proposed method to bridge the gap between foundation models and domain-specific medical imaging tasks. This advancement paves the way for more effective AI-assisted medical diagnostic systems. Jihong Hu, Yinhao Li 0002, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Multimodal Graph Learning With Multi-Hypergraph Reasoning Networks for Focal Liver Lesion Classification in Multimodal Magnetic Resonance ImagingabstractMultimodal magnetic resonance imaging (MRI) is instrumental in differentiating liver lesions. The major challenge involves modeling reliable connections and simultaneously learning complementary information across various MRI sequences. While previous studies have primarily focused on multimodal integration in a pair-wise manner using few modalities, our research seeks to advance a more comprehensive understanding of interaction modeling by establishing complex high-order correlations among the diverse modalities in multimodal MRI. In this paper, we introduce a multimodal graph learning with multi-hypergraph reasoning network to capture the full spectrum of both pair-wise and group-wise relationships among different modalities. Specifically, a weight-shared encoder extracts features from regions of interest (ROI) images across all modalities. Subsequently, a collection of uniform hypergraphs are constructed with varying vertex configurations, allowing for the modeling of not only pair-wise correlations but also the high-order collaborations for relational reasoning. Following information propagation through the hypergraph message passing, adaptive intra-modality fusion module is proposed to effectively fuse feature representations from different hypergraphs of the same modality. Finally, all refined features are concatenated to prepare for the classification task. Our experimental evaluations, including focal liver lesions classification using the LLD-MMRI2023 dataset and early recurrence prediction of hepatocellular carcinoma using our internal datasets, demonstrate that our method significantly surpasses the performance of existing approaches, indicating the effectiveness of our model in handling both pair-wise and group-wise interactions across multiple modalities. Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Fang Wang 0030, Qingqing Chen 0001, Wenbin Ji, Yinhao Li 0002, Hongjie Hu, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Clinical Data-Driven Retrieval-Augmented Model for Lung Nodule Malignancy Prediction
Ruibo Hou, Shurong Chai, Rahul Kumar Jain 0001, Yinhao Li 0002, Jiaqing Liu, Shiyu Teng, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (10) | 4 |
| 2025 | TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
Rahul Kumar Jain 0001, Yinhao Li 0002, Ruibo Hou, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Rui Xu 0002, Yen-Wei Chen 0001 |
MICCAI (6) | 3 |
| 2025 | PD-UniST: Prompt-Driven Universal Model for Unpaired H&E-to-IHC Stain Translation
Chujie Zhang, Yangyang Xie, Yinhao Li 0002, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (2) | 3 |
| 2025 | Multi-modal Medical SAM: An Adaptation Method of Segment Anything Model (SAM) for Glioma Segmentation Using Multi-modal MR ImagesabstractThe segmentation of glioma is crucial for early diagnosis, according to a World Health Organization (WHO) 2021 report. For glioma diagnosis, 3D multi-modal brain MRI/CT imaging has become an essential tool, offering detailed information. Nowadays, deep learning frameworks have been applied to various medical imaging problems, including brain glioma segmentation. Recently, foundation models like Segment Anything Model (SAM) have emerged as pivotal tools in computer vision tasks. These models are trained using large (real-world) datasets, offering a generalized understanding of visual data and semantic key features. Therefore, the effective utilization of foundation models in medical imaging is a significant area of current research. However, the differences in data distribution between multi-modal medical images and real-world images present challenges in directly applying foundation models to medical imaging. Additionally, utilizing multi-modal images to extract crucial information and its fusion poses further challenges. To address these issues, we propose a framework using foundation model and novel strategies for multi-modal fusion. Our fusion adapters effectively integrate the information from different modalities to enhance glioma segmentation in multi-modal MRI scans. Our method outperforms current state-of-the-art methods for accurate segmentation of the glioma using private and publicly available brain MRI datasets, proving the effectiveness of our approach across different datasets and imaging modalities. Rahul Kumar Jain 0001, Yinhao Li 0002, Shurong Chai, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Yen-Wei Chen 0001 |
ACM Trans. Comput. Heal. | 3 |
| 2025 | Accurate Tracking of Arabidopsis Root Cortex Cell Nuclei in 3D Time-Lapse Microscopy Images Based on Genetic AlgorithmabstractArabidopsis is a widely used model plant to study physiology and development. Live imaging is an important technique to visualize and quantify processes in plant growth and cell division, where accurate cell tracking is essential. The commonly used software TrackMate adopts a tracking-by-detection approach, applying Laplacian of Gaussian (LoG) for blob detection and a Linear Assignment Problem (LAP) tracker for tracking. However, its performance declines when cells are densely arranged. To overcome this limitation, we propose an accurate tracking method based on a Genetic Algorithm (GA) that incorporates knowledge of Arabidopsis root cellular patterns and spatial relationships among volumes. Our method follows a coarse-to-fine strategy: first performing relatively simple line-level tracking of nuclei, then refining associations based on the linear arrangement of cell files and their spatial relationships. We evaluated the method on long-term live imaging datasets of Arabidopsis root tips, and with minor manual correction, it achieved accurate nuclear tracking. To the best of our knowledge, this represents the first successful attempt to address a long-standing problem in time-lapse microscopy of the root meristem by providing an accurate tracking method for Arabidopsis root nuclei. Yu Song 0008, Tatsuaki Goh, Yinhao Li 0002, Jiahua Dong 0001, Shunsuke Miyashima, Yutaro Iwamoto, Yohei Kondo, Keiji Nakajima, Yenwei Chen |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | SAMA: A Self-and-Mutual Attention Network for Accurate Recurrence Prediction of Non-Small Cell Lung Cancer Using Genetic and CT DataabstractAccurate preoperative recurrence prediction for non-small cell lung cancer (NSCLC) is a challenging issue in the medical field. Existing studies primarily conduct image and molecular analyses independently or directly fuse multimodal information through radiomics and genomics, which fail to fully exploit and effectively utilize the highly heterogeneous cross-modal information at different levels and model the complex relationships between modalities, resulting in poor fusion performance and becoming the bottleneck of precise recurrence prediction. To address these limitations, we propose a novel unified framework, the Self-and-Mutual Attention (SAMA) Network, designed to efficiently fuse and utilize macroscopic CT images and microscopic gene data for precise NSCLC recurrence prediction, integrating handcrafted features, deep features, and gene features. Specifically, we design a Self-and-Mutual Attention Module that performs three-stage fusion: the self-enhancement stage enhances modality-specific features; the gene-guided and CT-guided cross-modality fusion stages perform bidirectional cross-guidance on the self-enhanced features, complementing and refining each modality, enhancing heterogeneous feature expression; and the optimized feature aggregation stage ensures the refined interactive features for precise prediction. Extensive experiments on both publicly available datasets from The Cancer Imaging Archive (TCIA) and The Cancer Genome Atlas (TCGA) demonstrate that our method achieves state-of-the-art performance and exhibits broad applicability to various cancers. Yang Ai, Jing Liu 0041, Yinhao Li 0002, Fang Wang 0030, Xiuju Du, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Ladder Fine-tuning Approach for SAM Integrating Complementary NetworkabstractRecently, foundation models have been introduced demonstrating various tasks in the field of computer vision. These models such as Segment Anything Model (SAM) are generalized models trained using huge datasets. Currently, ongoing research focuses on exploring the effective utilization of these generalized models for Specific domains, such as medical imaging. However, in medical imaging, the lack of training samples due to privacy concerns and other factors presents a major challenge for applying these generalized models to medical image segmentation task. To address this issue, the effective fine tuning of these models is crucial to ensure their optimal utilization. In this study, we propose to combine a complementary Convolutional Neural Network (CNN) along with the standard SAM network for medical image segmentation. To reduce the burden of fine tuning large foundation model and implement cost-efficient training scheme, we focus only on fine-tuning the additional CNN network and SAM decoder part. This strategy significantly reduces training time and achieves competitive results on publicly available dataset. The code is available at ">https://github.com/11yxk/SAM-LST . Shurong Chai, Rahul Kumar Jain 0001, Shiyu Teng, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Yen-Wei Chen 0001 |
KES | 5 |
| 2024 | A Novel Adaptive Hypergraph Neural Network for Enhancing Medical Image Segmentation
Shurong Chai, Rahul Kumar Jain 0001, Shaocong Mo, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (9) | 6 |
| 2024 | LGA: A Language Guide Adapter for Advancing the SAM Model's Capabilities in Medical Image Segmentation
Jihong Hu, Yinhao Li 0002, Hao Sun 0013, Yu Song 0008, Chujie Zhang, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (12) | 2 |
| 2024 | A motion-aware and temporal-enhanced Spatial-Temporal Graph Convolutional Network for skeleton-based human action segmentation
Shurong Chai, Rahul Kumar Jain 0001, Jiaqing Liu, Shiyu Teng, Tomoko Tateyama, Yinhao Li 0002, Yen-Wei Chen 0001 |
Neurocomputing | 6 |
| 2024 | Segmentation Guided Crossing Dual Decoding Generative Adversarial Network for Synthesizing Contrast-Enhanced Computed Tomography ImagesabstractAlthough contrast-enhanced computed tomography (CE-CT) images significantly improve the accuracy of diagnosing focal liver lesions (FLLs), the administration of contrast agents imposes a considerable physical burden on patients. The utilization of generative models to synthesize CE-CT images from non-contrasted CT images offers a promising solution. However, existing image synthesis models tend to overlook the importance of critical regions, inevitably reducing their effectiveness in downstream tasks. To overcome this challenge, we propose an innovative CE-CT image synthesis model called Segmentation Guided Crossing Dual Decoding Generative Adversarial Network (SGCDD-GAN). Specifically, the SGCDD-GAN involves a crossing dual decoding generator including an attention decoder and an improved transformation decoder. The attention decoder is designed to highlight some critical regions within the abdominal cavity, while the improved transformation decoder is responsible for synthesizing CE-CT images. These two decoders are interconnected using a crossing technique to enhance each other's capabilities. Furthermore, we employ a multi-task learning strategy to guide the generator to focus more on the lesion area. To evaluate the performance of proposed SGCDD-GAN, we test it on an in-house CE-CT dataset. In both CE-CT image synthesis tasks-namely, synthesizing ART images and synthesizing PV images-the proposed SGCDD-GAN demonstrates superior performance metrics across the entire image and liver region, including SSIM, PSNR, MSE, and PCC scores. Furthermore, CE-CT images synthetized from our SGCDD-GAN achieve remarkable accuracy rates of 82.68%, 94.11%, and 94.11% in a deep learning-based FLLs classification task, along with a pilot assessment conducted by two radiologists. Qingqing Chen 0001, Yinhao Li 0002, Fang Wang 0030, Xianhua Han, Yutaro Iwamoto, Jing Liu 0041, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (1) | 5 |
| 2021 | VolumeNet: A Lightweight Parallel Network for Super-Resolution of MR and CT Volumetric DataabstractDeep learning-based super-resolution (SR) techniques have generally achieved excellent performance in the computer vision field. Recently, it has been proven that three-dimensional (3D) SR for medical volumetric data delivers better visual results than conventional two-dimensional (2D) processing. However, deepening and widening 3D networks increases training difficulty significantly due to the large number of parameters and small number of training samples. Thus, we propose a 3D convolutional neural network (CNN) for SR of magnetic resonance (MR) and computer tomography (CT) volumetric data called ParallelNet using parallel connections. We construct a parallel connection structure based on the group convolution and feature aggregation to build a 3D CNN that is as wide as possible with a few parameters. As a result, the model thoroughly learns more feature maps with larger receptive fields. In addition, to further improve accuracy, we present an efficient version of ParallelNet (called VolumeNet), which reduces the number of parameters and deepens ParallelNet using a proposed lightweight building block module called the Queue module. Unlike most lightweight CNNs based on depthwise convolutions, the Queue module is primarily constructed using separable 2D cross-channel convolutions. As a result, the number of network parameters and computational complexity can be reduced significantly while maintaining accuracy due to full channel fusion. Experimental results demonstrate that the proposed VolumeNet significantly reduces the number of model parameters and achieves high precision results compared to state-of-the-art methods in tasks of brain MR image SR, abdomen CT image SR, and reconstruction of super-resolution 7T-like images from their 3T counterparts. Yinhao Li 0002, Yutaro Iwamoto, Lanfen Lin, Rui Xu 0002, Ruofeng Tong 0001, Yen-Wei Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Novel image restoration method based on multi-frame super-resolution for atmospherically distorted imagesabstractIn this study, the authors propose a novel multi‐frame super‐resolution method using frame selection and multiple fusions for atmospherically distorted, zoomed‐in, image‐quality enhancement. When a small part of the image captured by placing a target several kilometres away from the fixed camera is enlarged, the quality of the part becomes poor owing to low resolution, spatial deformations and noise that are mainly caused by long distance and atmospheric turbulence. Thus, the authors propose an adaptive frame selection method that selects only a few frames with small blur based on the corresponding images with relatively clear edges. Further, they propose multiple fusion schemes to reconstruct the selected frames, thereby suppressing the influence of deformation. By converting all the frames into high‐resolution based on each frame and integrating them, deformation and noise are effectively removed without high computation cost using the multiple fusion scheme. The proposed method, which enhances the quality of atmospherically distorted zoomed‐in images, exhibits superior performance than the state‐of‐the‐art image super‐resolution methods with regard to high accuracy, efficiency and ease of implementation, ensuring that the proposed method is suitable for enhancing the quality of an image captured using a general digital camera or a smartphone. Yinhao Li 0002, Katsuhisa Ogawa, Yutaro Iwamoto, Yen-Wei Chen 0001 |
IET Image Process. | 1 |