EDBT 2026 Demo / reviewers in the wild / expert
Mingyuan Meng
dblp:256/1106
· DBLP profile ↗
15ranked-venue papers
8as first author
13since 2021 · last 2027
0000-0002-9562-1613ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Language-guided medical image segmentation with target-informed multi-level contrastive alignmentsabstractMedical image segmentation is a fundamental task in numerous medical applications. Recently, language-guided segmentation has shown promise in medical scenarios where textual clinical reports are readily available as semantic guidance. Clinical reports contain diagnostic information provided by clinicians, which can provide auxiliary textual semantics to guide segmentation. However, existing language-guided segmentation methods neglect the inherent pattern gaps between image and text modalities, resulting in sub-optimal visual-language integration. Contrastive learning is a well-recognized approach to align image-text patterns, but it has not been optimized for medical image segmentation, where clinically meaningful semantics are often concentrated in localized target regions rather than the entire image. In this study, we propose TMCA, a Target-informed Multi-level Contrastive Alignment framework to bridge image-text pattern gaps for medical language-guided segmentation. The core innovation is to reformulate image-text contrastive alignment from conventional instance-level matching to segmentation-oriented semantic matching, where image-text samples are aligned according to their segmentation targets rather than merely whether they come from the same patient. Specifically, TMCA enables target-informed image-text alignments and fine-grained textual guidance by introducing: (i) a target-sensitive semantic distance module that utilizes target information for more granular image-text alignment modeling, (ii) a multi-level contrastive alignment strategy that directs fine-grained textual guidance to multi-scale image details, and (iii) a language-guided target enhancement module that reinforces attention to critical regions based on the aligned image-text patterns. Extensive experiments on four public benchmarks, involving three medical imaging modalities with clinical reports, show that TMCA enabled superior performance over state-of-the-art language-guided medical segmentation methods. Mingjian Li, Mingyuan Meng, Shuchang Ye, Mingye Zou, Michael J. Fulham, Lei Bi 0001, Jinman Kim |
Expert Syst. Appl. | 2 |
| 2026 | Dynamic Traceback Learning for Medical Report GenerationabstractAutomated medical report generation has demonstrated the potential to significantly reduce the workload associated with time-consuming medical reporting. Recent generative representation learning methods have shown promise in integrating vision and language modalities for medical report generation. However, when trained end-to-end and applied directly to medical image-to-text generation, they face two significant challenges: i) difficulty in accurately capturing subtle yet crucial pathological details, and ii) reliance on both visual and textual inputs during inference, leading to performance degradation in zero-shot inference when only images are available. To address these challenges, this study proposes a novel multimodal dynamic traceback learning framework (DTrace)1. Specifically, we introduce a traceback mechanism to supervise the semantic validity of generated content and a dynamic learning strategy to adapt to various proportions of image and text input, enabling text generation without strong reliance on the input from both modalities during inference. The learning of cross-modal knowledge is enhanced by supervising the model to recover masked semantic information from a complementary counterpart. Extensive experiments conducted on two benchmark datasets, IU-Xray and MIMIC-CXR, demonstrate that the proposedDTraceframework outperforms state-of-the-art methods for medical report generation. Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Usman Naseem, Jinman Kim |
IEEE Trans. Multim. | 2 |
| 2025 | Alleviating Textual Reliance in Medical Language-Guided Segmentation via Prototype-Driven Semantic ApproximationabstractMedical language-guided segmentation, integrating textual clinical reports as auxiliary guidance to enhance image segmentation, has demonstrated significant improvements over unimodal approaches. However, its inherent reliance on paired image-text input, which we refer to as ``textual reliance", presents two fundamental limitations: 1) many medical segmentation datasets lack paired reports, leaving a substantial portion of image-only data underutilized for training; and 2) inference is limited to retrospective analysis of cases with paired reports, limiting its applicability in most clinical scenarios where segmentation typically precedes reporting. To address these limitations, we propose ProLearn, the first Prototype-driven Learning framework for language-guided segmentation that fundamentally alleviates textual reliance. At its core, we introduce a novel Prototype-driven Semantic Approximation (PSA) module to enable approximation of semantic guidance from textual input. PSA initializes a discrete and compact prototype space by distilling segmentation-relevant semantics from textual reports. Once initialized, it supports a query-and-respond mechanism which approximates semantic guidance for images without textual input, thereby alleviating textual reliance. Extensive experiments on QaTa-COV19, MosMedData+ and Kvasir-SEG demonstrate that ProLearn outperforms state-of-the-art language-guided methods when limited text is available. Shuchang Ye, Usman Naseem, Mingyuan Meng, Jinman Kim |
ICCV | 3 |
| 2025 | AutoFuse: Automatic fusion networks for deformable medical image registration
Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
Pattern Recognit. | 1 |
| 2025 | Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation ModelingabstractMedical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modeling the inter-matching relations, can learn improved medical image representations with better generalization capabilities. Mingjian Li, Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Correlation-aware Coarse-to-fine MLPs for Deformable Medical Image RegistrationabstractDeformable image registration is a fundamental step for medical image analysis. Recently, transformers have been used for registration and outperformed Convolutional Neural Networks (CNNs). Transformers can capture long-range dependence among image features, which have been shown beneficial for registration. However, due to the high computation/memory loads of self-attention, transformers are typically used at downsampled feature resolutions and cannot capture fine-grained long-range dependence at the full image resolution. This limits deformable registration as it necessitates precise dense correspondence between each image pixel. Multi-layer Perceptrons (MLPs) without self-attention are efficient in computation/memory usage, enabling the feasibility of capturing fine-grained long-range dependence at full resolution. Nevertheless, MLPs have not been extensively explored for image registration and are lacking the consideration of inductive bias crucial for medical registration tasks. In this study, we propose the first correlation-aware MLP-based registration network (CorrMLP) for deformable medical image registration. Our CorrMLP introduces a correlation-aware multi-window MLP block in a novel coarse-to-fine registration architecture, which captures fine-grained multi-range dependence to perform correlation-aware coarse-to-fine registration. Extensive experiments with seven public medical datasets show that our CorrMLP outperforms state-of-the-art deformable registration methods. Mingyuan Meng, David Dagan Feng, Lei Bi 0001, Jinman Kim |
CVPR | 1 |
| 2024 | 3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
Xiaoshuang Li, Mingyuan Meng, Zimo Huang, Lei Bi 0001, Eduardo Delamare, David Dagan Feng, Bin Sheng 0001, Jinman Kim |
MICCAI (7) | 2 |
| 2024 | Enabling Text-Free Inference in Language-Guided Segmentation of Chest X-Rays via Self-guidance
Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Jinman Kim |
MICCAI (8) | 2 |
| 2023 | Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (6) | 1 |
| 2023 | Non-iterative Coarse-to-Fine Transformer Networks for Joint Affine and Deformable Image Registration
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (10) | 1 |
| 2022 | Non-iterative Coarse-to-Fine Registration Based on Single-Pass Deep Cumulative Learning
Mingyuan Meng, Lei Bi 0001, David Dagan Feng, Jinman Kim |
MICCAI (6) | 1 |
| 2022 | DeepMTS: Deep Multi-Task Learning for Survival Prediction in Patients With Advanced Nasopharyngeal Carcinoma Using Pretreatment PET/CTabstractNasopharyngeal Carcinoma (NPC) is a malignant epithelial cancer arising from the nasopharynx. Survival prediction is a major concern for NPC patients, as it provides early prognostic information to plan treatments. Recently, deep survival models based on deep learning have demonstrated the potential to outperform traditional radiomics-based survival prediction models. Deep survival models usually use image patches covering the whole target regions (e.g., nasopharynx for NPC) or containing only segmented tumor regions as the input. However, the models using the whole target regions will also include non-relevant background information, while the models using segmented tumor regions will disregard potentially prognostic information existing out of primary tumors (e.g., local lymph node metastasis and adjacent tissue invasion). In this study, we propose a 3D end-to-end Deep Multi-Task Survival model (DeepMTS) for joint survival prediction and tumor segmentation in advanced NPC from pretreatment PET/CT. Our novelty is the introduction of a hard-sharing segmentation backbone to guide the extraction of local features related to the primary tumors, which reduces the interference from non-relevant background information. In addition, we also introduce a cascaded survival network to capture the prognostic information existing out of primary tumors and further leverage the global tumor information (e.g., tumor size, shape, and locations) derived from the segmentation backbone. Our experiments with two clinical datasets demonstrate that our DeepMTS can consistently outperform traditional radiomics-based survival prediction models and existing deep survival models. Mingyuan Meng, Bingxin Gu, Lei Bi 0001, Shaoli Song, David Dagan Feng, Jinman Kim |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | High-parallelism Inception-like Spiking Neural Networks for Unsupervised Feature Learning
Mingyuan Meng, Lei Bi 0001, Jinman Kim, Shanlin Xiao, Zhiyi Yu |
Neurocomputing | 1 |
| 2020 | SPA: Stochastic Probability Adjustment for System Balance of Unsupervised SNNsabstractSpiking neural networks (SNNs) receive widespread attention because of their low-power hardware characteristic and brain-like signal response mechanism, but currently, the performance of SNNs is still behind Artificial Neural Networks (ANNs). We build an information theory-inspired system called Stochastic Probability Adjustment (SPA) system to reduce this gap. The SPA maps the synapses and neurons of SNNs into a probability space where a neuron and all connected pre-synapses are represented by a cluster. The movement of synaptic transmitter between different clusters is modeled as a Brownian-like stochastic process in which the transmitter distribution is adaptive at different firing phases. We experimented with a wide range of existing unsupervised SNN architectures and achieved consistent performance improvements. The improvements in classification accuracy have reached 1.99% and 6.29% on the MNIST and EMNIST datasets respectively. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
ICPR | 2 |
| 2020 | Spiking Inception Module for Multi-layer Unsupervised Spiking Neural NetworksabstractSpiking Neural Network (SNN), as a brain-inspired approach, is attracting attention due to its potential to produce ultra-high-energy-efficient hardware. Competitive learning based on Spike-Timing-Dependent Plasticity (STDP) is a popular method to train an unsupervised SNN. However, previous unsupervised SNNs trained through this method are limited to a shallow network with only one learnable layer and cannot achieve satisfactory results when compared with multi-layer SNNs. In this paper, we eased this limitation by: 1) We proposed a Spiking Inception (Sp-Inception) module, inspired by the Inception module in the Artificial Neural Network (ANN) literature. This module is trained through STDP-based competitive learning and outperforms the baseline modules on learning capability, learning efficiency, and robustness. 2)We proposed a Pooling-Reshape-Activate (PRA) layer to make the Sp-Inception module stackable. 3)We stacked multiple Sp-Inception modules to construct multilayer SNNs. Our algorithm outperforms the baseline algorithms on the hand-written digit classification task, and reaches state-of-the-art results on the MNIST dataset among the existing unsupervised SNNs. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
IJCNN | 1 |