EDBT 2026 Demo / reviewers in the wild / expert
Ke Zou
dblp:88/10064
· DBLP profile ↗
31ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LMOD\(\boldsymbol{+}\): A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in OphthalmologyabstractThe rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. Artificial intelligence (AI) offers potential solutions. In particular, recent progress in foundation models and large language models-especially multimodal large language models (MLLMs)-has shown promise in medical image interpretation and automated clinical documentation. However, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets for development and evaluation. Most existing benchmarks were designed for earlier models, which focused on narrow tasks or specific disease conditions. These benchmarks typically provide outputs in the form of disease labels rather than free-text responses. As a result, they are less suitable for assessing emerging generative models. In this work, we present LMOD+, a large-scale multimodal ophthalmology benchmark dataset comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations. It supports primary ophthalmic applications such as anatomical structure recognition, disease screening, disease staging, and demographic prediction for potential performance bias evaluation. Alongside the dataset, we introduce a systematic and unified data curation pipeline that repurposes existing or new datasets for MLLM development. LMOD+ extends our preliminary LMOD benchmark-the first multimodal ophthalmology benchmark for MLLMs-with three major enhancements. First, we expanded the dataset by nearly 50% (from 21,933 to 32,633 instances). The color fundus photography (CFP) modality, the most accessible imaging modality in ophthalmology, was significantly enlarged to cover a broader range of pathological conditions. Second, we broadened task coverage to include (a) 12 binary disease diagnosis tasks for prevalent conditions such as diabetic retinopathy, age-related macular degeneration, and retinal vein occlusion; (b) multi-class ophthalmic disease diagnosis; (c) disease severity classification, including a diabetic retinopathy staging task, which uses two internationally adopted grading standards: the international clinical diabetic retinopathy classification and the Scottish diabetic retinopathy grading scheme classification; and (d) demographic prediction (age and sex) to assess potential model bias. Third, we systematically evaluated 24 state-of-the-art MLLMs, including recent models from the InternVL, Qwen, and DeepSeek families. Our evaluations highlight both the promise and limitations of current MLLMs in ophthalmology. For example, Qwen-7B and InternVL achieved accuracies of 58.26% and 57.83% in disease screening under a zero-shot setting with a single model-a considerably more challenging paradigm than traditional fine-tuning, where separate models are trained for each specific task. InternVL also demonstrated potential in anatomical recognition. Nonetheless, overall performance remained suboptimal and often close to random baselines for challenging tasks such as disease staging, underscoring the substantial gap between general-domain MLLMs and the specialized requirements of ophthalmology. We publicly release the dataset, curation pipeline, and leaderboard to encourage community-wide development and evaluation of MLLMs, with the goal of advancing ophthalmic applications and ultimately reducing the global burden of vision-threatening diseases through AI. The dataset website, benchmark leaderboard, and download link are available at https://kfzyqin.github.io/lmod_plus. Zhenyue Qin, Yang Liu 0249, Jinyu Ding, Anran Li 0001, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. Keenan, Emily Y. Chew, Zhiyong Lu, Ninghao Liu 0001, Xiuzhen Zhang 0001, Qingyu Chen 0001 |
ACM Trans. Comput. Heal. | 9 |
| 2026 | Federated semi-supervised calibrated efficient fine-tuning of foundation models for medical image classification
Along He, Yanlin Wu, LinLin Shen, Ke Zou, Huazhu Fu |
Knowl. Based Syst. | 5 |
| 2026 | MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAMabstractThe Medical Segment Anything Model (MedSAM) has demonstrated strong performance in medical image segmentation, attracting increasing attention in the medical imaging domain. However, as with many prompt-based segmentation models, its performance is highly sensitive to the type and location of input prompts. This sensitivity often leads to suboptimal segmentation outcomes and necessitates labor-intensive manual prompt tuning, which hampers both efficiency and robustness. To address this challenge, this paper proposes MedSAM-U, an uncertainty-guided framework designed to automatically refine prompt inputs and enhance segmentation reliability. Specifically, a Multi-Prompt Adapter is integrated into MedSAM, resulting in MPA-MedSAM, which enables the model to effectively accommodate diverse multi-prompt inputs. An uncertainty estimation module is then introduced to evaluate the reliability of the prompts and their initial segmentation results. Based on this, a novel uncertainty-guided prompt adaptation strategy is applied to automatically generate refined prompts and more accurate segmentation outputs. The proposed MedSAM-U framework is evaluated across multiple medical imaging modalities. Experimental results on five diverse datasets demonstrate that MedSAM-U achieves consistent performance improvements ranging from 1.7% to 20.5% over the baseline MedSAM, confirming its effectiveness and practicality for robust and efficient medical image segmentation. Ke Zou, Mengting Luo, Linchao He, Meng Wang 0038, Yi Zhang 0018, Hu Chen 0002, Huazhu Fu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisabstractMedical Visual Question Answering (Med-VQA) combines computer vision and natural language processing to automatically answer clinical inquiries about medical images. However, current Med-VQA datasets exhibit two significant limitations: (1) they often lack visual and textual explanations for answers, hindering comprehension for patients and junior doctors; (2) they typically offer a narrow range of question formats, inadequately reflecting the diverse requirements in practical scenarios. These limitations pose significant challenges to the development of a reliable and user-friendly Med-VQA system. To address these challenges, we introduce a large-scale, Groundable, and Explainable Medical VQA benchmark for chest X-ray diagnosis (GEMeX), featuring several innovative components: (1) a multi-modal explainability mechanism that offers detailed visual and textual explanations for each question-answer pair, thereby enhancing answer comprehensibility; (2) four question types, open-ended, closed-ended, single-choice, and multiple-choice, to better reflect practical needs. With 151,025 images and 1,605,575 questions, GEMeX is the currently largest chest X-ray VQA dataset. Evaluation of 12 representative large vision language models (LVLMs) on GEMeX reveals suboptimal performance, underscoring the dataset's complexity. Meanwhile, we propose a strong model by fine-tuning an existing LVLM on the GEMeX training set. The substantial performance improvement showcases the dataset's effectiveness. The benchmark is available at https://www.med-vqa.com/GEMeX. Bo Liu 0049, Ke Zou, Li-Ming Zhan, Chengqiang Xie, Jiannong Cao 0001, Xiao-Ming Wu 0003, Huazhu Fu |
ICCV | 2 |
| 2025 | RIFNet: Bridging Modalities for Accurate and Detailed Ocular Disease Analysis
Qingshan Hou, Peng Cao 0001, Jianguo Ju, Meng Wang 0001, Ke Zou, Huazhu Fu, Osmar R. Zaïane |
MICCAI (13) | 7 |
| 2025 | Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering
Zehui Liao, Shishuai Hu, Ke Zou, Huazhu Fu, Liangli Zhen, Yong Xia 0001 |
MICCAI (5) | 3 |
| 2025 | Uncertainty-Aware Medical Diagnostic Phrase Identification and GroundingabstractMedical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels. Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 5 |
| 2025 | Toward Reliable Medical Image Segmentation by Modeling Evidential Calibrated UncertaintyabstractMedical image segmentation is critical for disease diagnosis and treatment assessment. However, concerns regarding the reliability of segmentation regions persist among clinicians, mainly attributed to the absence of confidence assessment, robustness, and calibration to accuracy. To address this, we introduce deep evidential segmentation model (DEviS), an easily implementable foundational model that seamlessly integrates into various medical image segmentation networks. DEviS not only enhances the calibration and robustness of baseline segmentation accuracy but also provides high-efficiency uncertainty estimation for reliable predictions. By leveraging subjective logic theory, we explicitly model probability and uncertainty for medical image segmentation. Here, the Dirichlet distribution parameterizes the distribution of probabilities for different classes of the segmentation results. To generate calibrated predictions and uncertainty, we develop a trainable calibrated uncertainty penalty. Furthermore, DEviS incorporates an uncertainty-aware filtering (UAF) module, which designs the metric of uncertainty-calibrated error to filter out-of-distribution (OOD) data. We conducted validation studies on publicly available datasets, including ISIC2018, KiTS2021, LiTS2017, and BraTS2019, to assess the accuracy and robustness of different backbone segmentation models enhanced by DEviS, as well as the efficiency and reliability of uncertainty estimation. Additionally, two potential clinical trials were conducted using the UAF module. The clinical application conducted on the Johns Hopkins OCT and Duke OCT-DME datasets demonstrated the effectiveness of the model in filtering OOD data. The second trial evaluated its efficacy in filtering high-quality data on the FIVES datasets. At last, the proposed DEviS method was extended to semi-supervised medical image segmentation, where it exhibited strong robustness under noisy conditions. Our code has been released in https://github.com/Cocofeat/DEviS. Ke Zou, Ling Huang 0003, Xuedong Yuan, Xiaojing Shen, Meng Wang 0038, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Cybern. | 1 |
| 2025 | Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound DevicesabstractHandheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA. Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2024 | FRCNet: Frequency and Region Consistency for Semi-supervised Medical Image Segmentation
Along He, Tao Li 0022, Yanlin Wu, Ke Zou, Huazhu Fu |
MICCAI (8) | 4 |
| 2024 | Reliable Source Approximation: Source-Free Unsupervised Domain Adaptation for Vestibular Schwannoma MRI Segmentation
Hongye Zeng, Ke Zou, Zhihao Chen 0004, Huazhu Fu |
MICCAI (10) | 2 |
| 2024 | MFMAM: Image inpainting via multi-scale feature module with attention module
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Comput. Vis. Image Underst. | 4 |
| 2024 | Image inpainting algorithm based on inference attention module and two-stage network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | MICU: Image super-resolution via multi-level information compensation and U-net
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Expert Syst. Appl. | 4 |
| 2024 | Confidence-aware multi-modality learning for eye disease screening
Ke Zou, Tian Lin 0002, Zongbo Han, Meng Wang 0001, Xuedong Yuan, Haoyu Chen 0002, Changqing Zhang 0002, Xiaojing Shen, Huazhu Fu |
Medical Image Anal. | 1 |
| 2024 | MFFN: image super-resolution via multi-level features fusion network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Vis. Comput. | 4 |
| 2023 | Uncertainty-Informed Mutual Learning for Joint Medical Image Classification and Segmentation
Ke Zou, Xianjie Liu, Xuedong Yuan, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (4) | 2 |
| 2023 | Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
MICCAI (2) | 4 |
| 2023 | Reliable Multimodality Eye Disease Screening via Mixture of Student's t Distributions
Ke Zou, Tian Lin 0002, Xuedong Yuan, Haoyu Chen 0002, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (7) | 1 |
| 2023 | FFTI: Image inpainting algorithm via features fusion and two-steps inpainting
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Corrigendum to "FFTI: Image inpainting algorithm via features fusion and two-steps inpainting" [J. Visual Commun. Image Represent. 91 (2023) 103776]
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | DGCA: high resolution image inpainting via DR-GAN and contextual attention
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Multim. Tools Appl. | 4 |
| 2022 | TBraTS: Trusted Brain Tumor Segmentation
Ke Zou, Xuedong Yuan, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (8) | 1 |
| 2020 | A Methodology for Seamless Hot-Swap of Converters in DC MicrogridsabstractDC microgrids with paralleled sources require a seamless addition and removal of a converter module without shutting down the system. This replacement of a converter is known as a hot-swap event. In a democratic current sharing scheme, a critical issue is an overshoot in the grid voltage when a converter is suddenly added and an undershoot when a converter is suddenly removed. In this paper, by analyzing the start-up and shut-down operation of a converter module during the hot-swap, the reason for these excursions in the grid voltage is explained using an analytical model. It is shown that by changing the voltage controller's bandwidth, overshoot and undershoot can be reduced, but there's a trade-off between the voltage controller's bandwidth and the stability of the system. To overcome this trade-off, an algorithm that provides an additional degree of freedom is proposed in this paper. This algorithm without disturbing the system's stability prevents the overvoltage and the undervoltage by exponentially changing the current reference of the converter being added or removed. The stability analysis along with the circuit simulation results using the proposed algorithm are shown for a wide range of operating conditions and different system parameters. Shrivatsal Sharma, Vishnu Mahadeva Iyer, Subhashish Bhattacharya, Jun Kikuchi, Ke Zou, Mahima Gupta |
IECON | 5 |
| 2020 | Robust sensor fusion with heavy-tailed noises
Hao Zhu 0003, Ke Zou, Yongfu Li 0001, Henry Leung 0001 |
Signal Process. | 2 |
| 2019 | Non-rigid Image Feature Matching by Structure Constraints
Hao Zhu 0003, Ke Zou, Yongfu Li 0001, Henry Leung 0001 |
FUSION | 2 |
| 2018 | Distributed Evidential EM Algorithm for Gaussian Mixtures in Sensor Network with Uncertain DataabstractIn this paper, the problem of clustering in distributed sensor networks with uncertain measurements is considered. It is assumed that each node in the sensor network can be described as a mixture of some elementary conditions. Therefore, the measurements of the sensors can be modeled using a Gaussian mixture model, in which the uncertainty on the attributes is represented by the belief functions. We present a novel algorithm, called distributed evidential expectation maximization (DEEM) algorithm, for the estimation of the Gaussian components in the mixture model. The effectiveness of the proposed algorithm is demonstrated through simulations of sensor networks with uncertain data. Hao Zhu 0003, Ke Zou, Yongfu Li 0001 |
FUSION | 2 |
| 2018 | Vehicle Tracking at Nighttime by Kernelized Experts With Channel-Wise and Temporal Reliability EstimationabstractDespite the fact that in recent years, vision-based tracking approaches have made significant progress, the task of tracking vehicles at night still remains challenging. Visual information is strongly deteriorated or at least degraded due to poor illumination conditions. This reduces the perceptive ability of vision systems significantly and can even lead to target loss, resulting in false estimation and/or false prediction of object behavior. In this paper, we propose a novel online-learning method to track vehicles at night. Our method is based on the kernelized correlation filter and assembles different feature channels to kernelized experts. By estimating their reliabilities, we force the appearance model to focus on the most discriminative visual features to accomplish the classification. In addition, a temporal optimization step in conjunction with a memory model is used to remove outliers and keep the most reliable samples to train the tracker models. Experiments over various daytime and weather conditions show that our approach outperforms existing trackers at night and in case of bad weather while offering state-of-the-art performance in more favorable situations. As our tracker has only little computational cost, it is appropriate for use cases with real-time requirements like in automotive or industrial applications. Wei Tian 0001, Long Chen 0005, Ke Zou, Martin Lauer |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | A distributed range-free correction vector based localization refinement algorithm
Yingbiao Yao, Ke Zou, Xianyun Chen, Xiaorong Xu |
Wirel. Networks | 2 |
| 2010 | Intelligent sensor design in network based automatic controlabstractUntil recently, most oil pumping units (OPUs) have been using manual control in the oilfield. In this paper, network based automatic control is proposed for OPU management. This proposed network can sequentially realize automatic data sensing, automatic malfunction detection, remote data transmission, intelligent data organization and management, automatic malfunction warning, automatic stroke adjustment and state report via GSM short message service (SMS). Specifically, the proposed network based automatic control system contains four parts: 1) the first level sensors (FLS) for data sensing; 2) the developed intelligent sensors (ISs), which are the second level sensors, for data storage and elementary processing; 3) a communication protocol for data transmission; and 4) a software-based network center for data processing, data management, malfunction detection, remote stroke adjustment, GSM SMS and so on. In this paper, we focus on describing the IS design and reporting the elementary experiment results of the proposed system at the oil well HUACHI-38-122. Sandeep Chandana, Renlun He, Jiuqiang Han, Ke Zou |
IJCNN | 6 |