Shuaipeng Ding

dblp:384/5381 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wavelet-Driven Spatial and Frequency Domain Representation Learning for Medical Image Segmentation
Lanping Wang, Mingyong Li, Shuaipeng Ding
ICPR (3)3
2026 ITAdapter: Image-Tag adapter framework with retrieval knowledge enhancer for radiology report generation
Shuaipeng Ding, Jianan Shui, Mingyuan Ge, Mengnan Fan, Xin Li 0242, Mingyong Li
Expert Syst. Appl.1
2026 Latent preference mining and noise-aware for multimodal recommendation
Qinze Zhu, Shuaipeng Ding, Mingyong Li
Pattern Recognit. Lett.3
2025 MMKPL-Seg: Knowledge Proxy Learning with Missing Modalities for Medical Image Segmentation
abstract
Text-prompted methods have been proven effective in enhancing the performance of medical image segmentation tasks. However, text-prompted segmentation of pulmonary infection regions still faces significant challenges in practical applications: the high cost of comprehensive medical prompts leads to frequent occurrences of missing medical text modalities in clinical testing scenarios. To address this issue, we propose a novel Knowledge Proxy Learning with Missing Modalities for Medical Image Segmentation (MMKPL-Seg). This model employs a learnable Memory Bank to automatically extract and store textual information during the training phase. Then, we introduce the Knowledge Proxy Learning (KPL) for both complete and missing modality scenarios, it retrieves relevant knowledge from the Memory Bank and integrates it into the encoder. By iteratively training on both single-image and image-text paired samples, the model enables single-image inputs to retain potential semantic information during the testing phase. Extensive evaluations on the QaTa-COV19 and MosMedData+ datasets demonstrate that our model achieves state-of-the-art performance compared to both unimodal and previous multimodal methods. Notably, MMKPL-Seg maintains stable segmentation performance even when text modalities are completely absent, making it highly adaptable to the common occurrence of incomplete textual information in clinical practice. This characteristic underscores its substantial practical value.
Shuaipeng Ding, Mengnan Fan, Mingyong Li
BIBM1
2025 ITAdaptor: Image-Tag Adapter Framework with Knowledge Enhancement for Radiology Report Generation
Shuaipeng Ding, Mengnan Fan, Mingyong Li
MICCAI (6)1
2025 MG-UNet: A Memory-Guided UNet for Lesion Segmentation in Chest Images
Shuaipeng Ding, Mingyong Li
MICCAI (1)1
2025 Spatially-Aware Entity Relation Exploration for Remote Sensing Image-Text Retrieval
abstract
In recent years, remarkable progress has been made in remote sensing image-text retrieval (RSITR), which has transitioned from relying on compact global features to more fine-grained local features representing salient objects in images. However, existing methods typically concentrate only on significant entity information in remote sensing images as local features, overlooking the correlations between entities, resulting in isolated entity information. Moreover, there is ample room for exploring text local features.To address these issues, this paper presents an Entity Spatial Relation enhancement Network (ESRN), leveraging global and local entity features in remote sensing images and texts. For local feature processing, Graph Convolutional Network (GCN) is employed to aggregate the correlation between entity features and entity information in remote sensing images, enhancing entity information learning. The self-attention mechanism is used to model the remote dependency of entity keywords and spatial orientation semantic relations in text, strengthening the text representation ability. A strategy of proportionally adding different levels of features is proposed to enhance the representation of salient features and reduce noise interference.The approach was evaluated on two renowned remote sensing datasets, RSICD and RSITMD, validating the model's capacity to perceive the semantics of remote sensing images and text entities. Performance comparison, ablation experiments, and visualization analysis convincingly demonstrate the state-of-the-art performance of the ESRN method in the RSITR task.
Jianan Shui, Shuaipeng Ding, Mingyuan Ge, Mingyong Li
ICMR2
2025 Cross-modal Memory Alignment Framework With Disease Aware Contrastive Learning for Radiology Report Generation
Shuaipeng Ding, Mengnan Fan, Mingyong Li
PRCV (14)1
2024 Dual-Attention Fusion Network with Edge and Content Guidance for Remote Sensing Images Segmentation
Shuaipeng Ding, Jianan Shui, Xin Li 0242, Mingyong Li
ICPR (30)1