Meidan Ding

dblp:311/1033 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0001-0520-020XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
abstract
Wenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung, Meidan Ding, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenxuan Wang 0001, Zizhan Ma, Yiu-Fai Cheung, Meidan Ding, Jie Liu 0044, Wenting Chen, LinLin Shen
ACL (1)5
2026 FProtoSeg: Fine-grained prototype alignment for Weakly Supervised Semantic Segmentation of histopathology images
Meidan Ding, Wenting Chen, Xiaoling Luo 0001, Haiqin Zhong, LinLin Shen
Pattern Recognit.1
2025 S³-Mamba: Small-Size-Sensitive Mamba for Lesion Segmentation
abstract
Small lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small lesions. To tackle the challenges, we propose a Small-Size-Sensitive Mamba (S³-Mamba), which promotes the sensitivity to small lesions across three dimensions: channel, spatial, and training strategy. Specifically, an Enhanced Visual State Space block is designed to focus on small lesions through multiple residual connections to preserve local features, and selectively amplify important details while suppressing irrelevant ones through channel-wise attention. A Tensor-based Cross-feature Multi-scale Attention is designed to integrate input image features and intermediate-layer features with edge features and exploit the attentive support of features across multiple scales, thereby retaining spatial details of small lesions at various granularities. Finally, we introduce a novel regularized curriculum learning to automatically assess lesion size and sample difficulty, and gradually focus from easy samples to hard ones like small lesions. Extensive experiments on three medical image segmentation datasets show the superiority of our S³-Mamba, especially in segmenting small lesions.
Gui Wang, Yuexiang Li, Wenting Chen, Meidan Ding, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, LinLin Shen
AAAI4
2025 EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language Model
abstract
Recent advancements in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses. While preference alignment methods have proven effective in general domains, acquiring high-quality preference data for pathology remains challenging due to limited expert resources and domain complexity. In this paper, we propose EAGLE (Expert-guided self-enhancement for preference Alignment in patholoGy Large vision-languagE model), a novel framework that systematically integrates medical expertise into preference alignment. EAGLE consists of three key stages: initialization through supervised fine-tuning, self-preference creation leveraging expert prompting and medical entity recognition, and iterative preference following-tuning. The self-preference creation stage uniquely combines expert-verified chosen sampling with expert-guided rejected sampling to generate high-quality preference data, while the iterative tuning process continuously refines both data quality and model performance. Extensive experiments demonstrate that EAGLE significantly outperforms existing pathological LVLMs, effectively reducing hallucination and bias while maintaining pathological accuracy. The source code is available at https://github.com/meidandz/EAGLE. © 2025 Association for Computational Linguistics.
Meidan Ding, Wenxuan Wang 0001, Haiqin Zhong, Xinheng Lyu, Wenting Chen, LinLin Shen
ACL (1)1
2025 FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
abstract
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench
Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen
CVPR4
2025 WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
abstract
Recent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy.
Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Xiangjian He, Xiaohan Xing, Sen Yang 0006, LinLin Shen
ICCV4
2025 FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and Editing
Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen
ICCV3
2025 🤖 WSI-Agents: A Collaborative Multi-agent System for Multi-modal Whole Slide Image Analysis
Xinheng Lyu, Yuci Liang, Wenting Chen, Meidan Ding, Guolin Huang, Daokun Zhang, Xiangjian He, LinLin Shen
MICCAI (5)4
2025 DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
abstract
Face parsing aims to segment facial images into key components such as eyes, lips, and eyebrows. While existing methods rely on dense pixel-level annotations, such annotations are expensive and labor-intensive to obtain. To reduce annotation cost, we introduce Weakly Supervised Face Parsing (WSFP), a new task setting that performs dense facial component segmentation using only weak supervision, such as image-level labels and natural language descriptions. WSFP introduces unique challenges due to the high co-occurrence and visual similarity of facial components, which lead to ambiguous activations and degraded parsing performance. To address this, we propose DisFaceRep, a representation disentanglement framework designed to separate co-occurring facial components through both explicit and implicit mechanisms. Specifically, we introduce a co-occurring component disentanglement strategy to explicitly reduce dataset-level bias, and a text-guided component disentanglement loss to guide component separation using language supervision implicitly. Extensive experiments on CelebAMask-HQ, LaPa, and Helen demonstrate the difficulty of WSFP and the effectiveness of DisFaceRep, which significantly outperforms existing weakly supervised semantic segmentation methods. The code will be released at https://github.com/CVI-SZU/DisFaceRep.
Xianxu Hou, Meidan Ding, Junliang Chen 0002, Kaijun Deng, Jinheng Xie, LinLin Shen
ACM Multimedia3
2025 MSMMIL: Multi-scan Mamba-based Multiple Instance Learning for whole slide image classification
Haiqin Zhong, Meidan Ding, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
Knowl. Based Syst.2
2023 An enhanced vision transformer with wavelet position embedding for histopathological image classification
Meidan Ding, Aiping Qu, Haiqin Zhong, Zhihui Lai 0001, Shuomin Xiao, Penghui He
Pattern Recognit.1
2021 A Transformer-based Network for Pathology Image Classification
abstract
Pathology image classification plays an important role in cancer diagnosis and precision treatment. Convolutional neural network has been widely employed in pathology image classification. Due to its convolution and pooling operation, it has a great advantage in extracting local features of small objects in images. However, it lacks the ability to extract the global contextual information contained in long-rang tissue structures in pathology images. Transformer, which adopts innate global self-attention mechanisms, has been obtained remarkable performance on large-scale datasets, but its localization abilities are limited because of insufficient low-level details. In this paper, we propose a transformer-based method which combines the advantages of CNN and Transformer for pathology image classification. CNN pays attention to extracting the local information of small objects, and Transformer focuses on digging out the global contextual information implied in the long-dependence tissue structures. The sparse interaction and weigh sharing inherited from CNN also allow the proposed method can be trained on small datasets. Experiments show that the proposed method achieves accuracy of 90.48 and 97.18 on PCam and NCT-CRC datasets, respectively, which is better than existing state-of-the-art methods.
Meidan Ding, Aiping Qu, Haiqin Zhong
BIBM1
2021 A Modified Convolutional Neural Network for Nuclei Classification in Histopathology Image
abstract
Classification of nuclei in histopathology images is an important step in pathology workflow. Although deep learning methods have been extensively employed in this task, automatic nuclei classification is still a challenging task because of the large inter-and intra-class variability as well as the serious clustered together. To address this challenge, we propose a modified automatic nuclei classification network for histopathology images. The proposed method adopts the encoding-decoding structure and leverages the segmented instance to guide the classification with the same feature maps from an encoding branch. We mainly pay attention to enhancing the feature representation. We introduce a new module composed of a modified multi-layer perceptron(MLP) module and a multi-kernel convolution(MKC) module. The MLP module focuses on mining global contextual information of nuclei and their micro-environments, and the MKC module focuses on extracting local information through different receptive fields. We validate the proposed method on two large publicly available datasets collected from different tissues. In comparison with other state-of-the-art methods, the proposed method obtains the best performance of Accuracy 88.1 and 81.6, respectively. We also verify that the proposed module can be easily incorporated into other networks for improving feature representation.
Haiqin Zhong, Aiping Qu, Meidan Ding
BIBM4