Huping Ye

dblp:197/2031 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-9114-205XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing 3D Medical Image Understanding With Pretraining Aided by 2D Multimodal Large Language Models
abstract
Understanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent advancements in multimodal large language models (MLLMs) provide a promising approach to enhance image understanding through text descriptions. To leverage these 2D MLLMs for improved 3D medical image understanding, we propose Med3DInsight, a novel pretraining framework that integrates 3D image encoders with 2D MLLMs via a specially designed plane-slice-aware transformer module. Additionally, our model employs a partial optimal transport based alignment, demonstrating greater tolerance to noise introduced by potential noises in LLM-generated content. Med3DInsight introduces a new paradigm for scalable multimodal 3D medical representation learning without requiring human annotations. Extensive experiments demonstrate our state-of-the-art performance on two downstream tasks, i.e., segmentation and classification, across various public datasets with CT and MRI modalities, outperforming current SSL methods. Med3DInsight can be seamlessly integrated into existing 3D medical image understanding networks, potentially enhancing their performance.
Qiuhui Chen, Xuancheng Yao, Huping Ye
IEEE J. Biomed. Health Informatics3
2025 Enhancing 3D Medical Image Understanding with 2D Multimodal Large Language Models
abstract
Understanding medical image volumes is crucial in healthcare, yet most current models for classification and segmentation often focus narrowly on task-specific features without capturing the broader medical context. To address this, we introduce Med3DInsight, a pre-training framework that enhances 3D image understanding by leveraging 2D multimodal large language models (MLLMs) through a Plane-Slice-Aware Transformer (PSAT) module. Med3DInsight connects 3D image encoders with 2D MLLMs, enhancing representation learning for downstream tasks. Extensive experiments on CT and MRI datasets demonstrate that Med3DInsight achieves state-of-the-art performance, surpassing 19 baseline methods and proving effective across diverse imaging modalities and anatomical structures. This framework can be seamlessly integrated into existing 3D medical imaging networks, significantly boosting their performance and adaptability. Our source code is publicly available at https://github.com/Qybc/Med3DInsight.
Qiuhui Chen, Xuancheng Yao, Huping Ye
ICASSP3
2025 MCAMamba: Multilevel Cross-Modal Attention-Guided State-Space Model for Multisource Remote Sensing Image Classification
abstract
Effective fusion of multi-source remote sensing data remains a fundamental challenge for Earth observation, as CNN and Transformer models suffer from limited receptive fields and high computational complexity. While State Space Models (SSM) like Mamba show promise in sequence modeling, they face three critical challenges in multi-source remote sensing: insufficient spatial-spectral coordination, cross-modal heterogeneity, and inadequate multi-scale feature integration. To address these limitations, this paper proposes MCAMamba: a Multi-Level Cross-Modal Attention-Guided Mamba framework for joint classification of hyperspectral images (HSI) and Light Detection and Ranging (LiDAR)/Synthetic Aperture Radar (SAR) data. MCAMamba introduces a novel three-stage feature fusion pipeline: 1) The FExt-Attention module enhances spatial structure and spectral information through parallel spatial-channel attention mechanisms. 2) The SSM-Attention module achieves deep cross-modal fusion by combining attention mechanisms with SSM for parametric interaction. 3) The FFus-Attention module performs adaptive multi-scale feature integration through global context modeling and cascaded attention. This hierarchical design enables superior feature representation with enhanced computational efficiency. Experiments on four public benchmark datasets (Houston2013, Houston2018, Augsburg, and Berlin) show that MCAMamba achieves Overall Accuracy (OA) of 94.75%, 93.35%, 92.46%, and 79.18%. The code will be available at https://github.com/Dmygithub/MCAMamba.
Mingyu Dou, Shi Qiu 0002, Ming Hu 0001, Xiaozhen Qiao, Huping Ye, Xiaohan Liao, Zhe Sun 0007
IEEE Trans. Geosci. Remote. Sens.5
2024 Epicardial Adipose Tissue Segmentation in MRIs Using Text-Prompted Pretraining Model
abstract
This paper addresses the challenge of segmenting the epicardial adipose tissue (EAT) in MRI scans, a critical aspect in cardiac-related clinical diagnosis, treatment, and evaluation of cardiovascular diseases. Despite the abundance of research focusing on CT images, studies on MRI segmentation are relatively scarce and are hindered by limited datasets for EAT segmentation. Existing models, such as the Segment Anything Model (SAM) family, provide a way to leverage existing related datasets and have shown promising results in various segmentation tasks. However, there is a significant gap between the general SAM models and our domain-specific task, resulting in unsatisfied performance on segmenting EAT. To address this, we adopt the SAM design but adapt it to our task using text prompts, resulting in a vision-language pretraining model for cardiac segmentation, especially for segmenting EAT. We utilize multiple publicly available cardiac-related medical datasets for pretraining and then fine-tune our model on a small EAT dataset. Additionally, we propose a pseudo-mask augmentation strategy to further mitigate the issue of limited EAT segmentation masks. Extensive experimental results demonstrate the effectiveness of our proposed model, achieving the best performance on EAT segmentation compared to recent baseline methods. Our source code and pre-trained model will be publicly available later.
Huping Ye, Yushan Deng
BIBM1
2024 Multinetwork Algorithm for Coastal Line Segmentation in Remote Sensing Images
abstract
The demarcation between the sea and the land, commonly referred to as the coastline, is of paramount importance for the dynamic monitoring of its alterations. This monitoring is essential for the effective utilization of marine resources and the conservation of the ecological environment. Addressing the challenges posed by the extensive expanse of coastal lines, which can complicate their acquisition and processing, this study utilizes remote sensing imagery to introduce an algorithm for coastal line segmentation. The algorithm integrates multiple networks to enhance its effectiveness. Innovations encompass the development of an extraction algorithm for coastal lines that are as follows. First, utilize an attention-guided conditional generative adversarial network (AC-GAN) model, which redefines the task of image segmentation by framing it as a style transformation problem. Second, a strategy for coastal line segmentation utilizes Dense Swin Transformer Unet (DSTUnet) to construct a densely structured model. This approach integrates Transformer to prioritize focal regions, thereby enhancing image and semantic interpretation. Third, a transfer learning framework is proposed to integrate multiple features, leveraging the strengths of different networks to achieve accurate segmentation of coastal lines. The study introduced two datasets, and the experimental results confirm that parallel network configurations and asymmetric weighting are superior in achieving optimal results, with an area overlap measure (AOM) score of 85%, outperforming the Unet by 5%.
Huping Ye, Shi Qiu 0002, Xiaohan Liao
IEEE Trans. Geosci. Remote. Sens.3
2023 Pretrain Once and Finetune Many Times: How Pretraining Benefits Brain MRI Segmentation
abstract
Brain MRI segmentation plays an important role in analyzing brain anatomical structures and understanding brain images. In this paper, we consider building a uniform 3D brain MRI segmentation framework using the pre-training and fine- tuning style to fully leverage existing public brain images and segmentation masks. Based on existing Transformer-based 3D image segmentation models, UNETR and Swin UNETR, we study the necessity and benefit of using pre-training, through pretraining on a big collection of over 6,000 brain scans from OASIS, ADNI, and CC359, and fine-tuning with limited segmentation masks to perform three downstream tasks, i.e., skull stripping, 4-structure segmentation, and 33-structure segmentation. Experimental results demonstrate that in most cases the pre-training can help reduce 90% of segmentation masks and half the time. Also, our method outperforms the recent method SynthSeg by a good margin. Our pre-trained model and source code are available online at https://github.com/AllanIverson/medical-segmentation.
Huping Ye
BIBM4
2023 Coastal Zone Extraction Algorithm Based on Multilayer Depth Features for Hyperspectral Images
abstract
The coastal zone is the most active natural area on the Earth’s surface and has the most favorable resources and environmental conditions. Therefore, it is of great significance to conduct research based on the coastal zone. Hyperspectral remote sensing images have spatial and spectral dimensions that reflect the spatial distribution and can be analyze the compositional information, which have been widely used for feature analysis and observation of ground objects. In this paper, we propose a coastal zone extraction algorithm based on multilayer depth features for hyperspectral images. The main contributions are as follows: 1) The Huanjing satellite hyperspectral coastal zone database is built for the first time, image composition is analyzed, and the noise removal algorithm is yielded. 2) 3D attention networks that are capable of carrying spatial and inter-spectral information are proposed. 3) A 3D CNN with SENet tandem structure is proposed to fully exploit detailed information, and a multi-layer feature extraction framework is built. We analyze four typical coastal zone patterns, and the experimental results show that our proposed algorithm can achieve coastal zone extraction with an average Kappa coefficient of 0.92, which is 0.06 higher than the mainstream algorithms. Our algorithm also shows good performance in complex environments. It provides a basis for further research on coastal zones.
Shi Qiu 0002, Huping Ye, Xiaohan Liao
IEEE Trans. Geosci. Remote. Sens.2
2020 A Contribution Algorithm from LDRI to HDRI
abstract
High dynamic range image (HDRI) which is combined with low dynamic range image (LDRI) needs to be mapped to a low dynamic area to display. In the process of mapping, it is impossible to determine the contribution of low dynamic image sequences in the display images, so that it results in a problem that the low dynamic images cannot be accurately selected. In this paper, for the first time, a contribution algorithm from LDRI to HDRI according to the corresponding response curve of the camera is proposed.
Junsong Luo, Shi Qiu 0002, Yizhang Jiang, Keyang Cheng, Huping Ye, Mingjin Zhang
Int. J. Pattern Recognit. Artif. Intell.5
2019 A CIE Color Purity Algorithm to Detect Black and Odorous Water in Urban Rivers Using High-Resolution Multispectral Remote Sensing Images
abstract
Urban black and odorous water (BOW) is a serious global environmental problem. Since these waters are often narrow rivers or small ponds, the detection of BOW waters using traditional satellite data and algorithms is limited both by a lack of spatial resolution and by imperfect retrieval algorithms. In this paper, we used the Chinese high-resolution remote sensing satellite Gaofen-2 (GF-2, 0.8 m). The atmospheric correction showed that the mean absolute percentage error of the derived remote sensing reflectance (Rrs) in visible bands is 25.19%. We first measured Rrsspectra of two classes of BOW [BOW with high concentrations of iron (II) sulfide, i.e., BOW1 and BOW with high concentrations of total suspended matter, i.e., BOW2] and ordinary water in Shenyang. Then, in situ Rrsdata were converted into Rrs corresponding to the wide GF-2 bands using the spectral response functions. We used the converted Rrsdata to calculate several band combinations, including the baseline height, [Rrs(green) - Rrs(red))/(Rrs(green) + Rrs(red)], and the color purity on a Commission Internationale de L'Eclairage (CIE) chromaticity diagram. The color purity was found to be the best index to extract BOW from ordinary water. Then, Rrs(645) was applied to categorize BOW into BOW1 and BOW2. We applied the algorithm to two synchronous GF-2 images. The recognition accuracy of BOW2 and ordinary water are both 100%. The extracted river water type near Weishanhu Road was BOW1, which agreed well with ground truth. The algorithm was further applied to other GF-2 data for Shenyang and Beijing.
Qian Shen 0003, Junsheng Li, Fangfang Zhang 0001, Shenglei Wang, Huping Ye, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.7