VLDB 2026 Research / reviewers in the wild / expert
Zhenbang Wu
dblp:315/0212
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RHealth: A R Toolkit for Deep Learning in HealthcareabstractMachine learning for electronic health records (EHR) is advancing rapidly and already underpins risk stratification, readmission and mortality prediction, and decision support, yet reliable translation still stalls on fragmented data pipelines, inconsistent medical-code handling, and hard-to-reproduce eval-uation-barriers that especially hinder R-centric clinical teams. Despite impressive methodological gains in temporal modeling, attention mechanisms, and strong classical baselines, most turnkey toolchains live in Python; as a result, many healthcare researchers and clinical data scientists working in R lack a single, integrated path from raw multi-table EHR to calibrated, auditable models. We address this gap with RHealth, an open-source, R-native toolkit that plays the role of an end-to-end conductor: from data harmonization and medical-code normal-ization to task specification, model training, and standardized reporting. Concretely, RHealth provides adapters for widely used public datasets (e.g., MIMIC-III/IV, eICU), utilities to traverse and map ICD-9/10 and CCS codes, task templates for common outcomes (mortality, 30-day readmission, length of stay), and a modeling stack that-at this development stage-supports standard recurrent baselines (e.g., RNN) and offers an extensible interface for user-defined architectures under active development, all evaluated with reproducible splits, AUROC/AUPRC, and cali-bration diagnostics. By packaging the full pipeline-from data to evaluation granularity-into modular, composable components, RHealth lowers the entry barrier for R users, reduces “glue code,” and promotes transparent, people-centric experimentation that can also serve as a trustworthy upstream substrate for LLM-enabled applications. To our knowledge, it is among the first comprehensive, integrated deep-learning toolkits for EHR in the R ecosystem. The code and documentation will be released after the double-blind review process. Ji Song, Zhixia Ren, Zhenbang Wu, John Wu, Chaoqi Yang, Yinghao Zhu, Wen Tang 0001, Jimeng Sun 0001, Ewen M. Harrison, Liantao Ma |
BIBM | 3 |
| 2025 | MEDS: Building Models and Tools in a Reproducible Health AI EcosystemabstractHealth AI suffers from a systemic reproducibility crisis that irreparably hinders research across both academia and industry [4,5].One key tool poised to solve this crisis is the Medical Event Data Standard (MEDS), a comprehensive data format and open-source ecosystem designed to enhance reproducibility and interoperability of AI research using longitudinal Electronic Health Records (EHR) [6].Currently adopted by over 15 institutions globally, MEDS encompasses various open-source tools, published models, and data processing pipelines, enabling streamlined model development and robust benchmarking.In this tutorial, participants will gain key hands-on experience in working with the MEDS format to perform efficient, reproducible, state-of-the-art AI research over real health data.Participants will transform data into the MEDS format, preprocess data, build predictive models, and contribute to the decentralized MEDS-DEV benchmarking platform.Interactive exercises using Jupyter notebooks will provide hands-on experience and practical skills for reproducible health AI research.Attendees will leave equipped Matthew B. A. McDermott, Justin Xu, Teya S. Bergamaschi, Hyewon Jeong, Simon A. Lee, Nassim Oufattole, Patrick Rockenschaub, Kamile Stankeviciute, Ethan Steinberg, Jimeng Sun 0001, Robin Van De Water, Michael Wornow, John Wu, Zhenbang Wu |
KDD (2) | 14 |
| 2025 | Reinforcement Learning for Out-of-Distribution Reasoning in LLMs: An Empirical Study on Diagnosis-Related Group CodingabstractDiagnosis-Related Group (DRG) codes are essential for hospital reimbursement and operations but require labor-intensive assignment. Large Language Models (LLMs) struggle with DRG coding due to the out-of-distribution (OOD) nature of the task: pretraining corpora rarely contain private clinical or billing data. We introduce DRG-Sapphire, which uses large-scale reinforcement learning (RL) for automated DRG coding from clinical notes. Built on Qwen2.5-7B and trained with Group Relative Policy Optimization (GRPO) using rule-based rewards, DRG-Sapphire introduces a series of RL enhancements to address domain-specific challenges not seen in previous mathematical tasks. Our model achieves state-of-the-art accuracy on the MIMIC-IV benchmark and generates physician-validated reasoning for DRG assignments, significantly enhancing explainability. Our study further sheds light on broader challenges of applying RL to knowledge-intensive, OOD tasks. We observe that RL performance scales approximately linearly with the logarithm of the number of supervised fine-tuning (SFT) examples, suggesting that RL effectiveness is fundamentally constrained by the domain knowledge encoded in the base model. For OOD tasks like DRG coding, strong RL performance requires sufficient knowledge infusion prior to RL. Consequently, scaling SFT may be more effective and computationally efficient than scaling RL alone for such tasks. Hanyin Wang, Zhenbang Wu, Gururaj Kolar, Hariprasad Korsapati, Brian Bartlett, Bryan Hull, Jimeng Sun 0001 |
NeurIPS | 2 |
| 2025 | Robust Multisource Forest Point Cloud Registration With Distribution Similarity AnalysisabstractAerial and terrestrial laser scanning (TLS) technologies offer complementary, high-precision 3-D data on forest structure. Registering point cloud data from multiplatform is crucial for achieving a more comprehensive understanding of forest structure. Currently, multisource forest point cloud registration remains challenging due to factors such as unstable point- and object-level features, varying observation perspectives, and nonstandardized processing pipelines. To address these challenges, this study introduces a unified and automated method for registering aerial and ground-based point clouds in forest areas. First, keypoints are extracted from multisource point clouds using fuzzy c-means (FCM) clustering. Local transformation is then derived from the keypoint sets using the gradient-based local convergence (GLC) algorithm, which is integrated with a nested branch and bound (BnB) structure for further optimization. We also developed a novel distribution similarity index (DSI) to evaluate the alignment of the keypoint sets and determine the initial transformation. Finally, this initial transformation is applied to the ground-based point cloud and refined using the GLC algorithm. Compared with existing methods that rely on point and object features, the proposed method does not depend on geometric descriptions or tree attributes (e.g., tree position and canopy structure). It also demonstrates robustness to variations in the initial position of the point clouds. Testing on 15 datasets with varying plot sizes and tree characteristics shows that the proposed method achieves comparable or superior performance to existing methods, with an average distance residual of 6.59 cm and an average runtime of 90.04 s. Xiangjiang Liu, Huabing Huang, Daile Wang, Zhenbang Wu, Peimin Chen, Xinlian Liang, Tianhong Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multimodal Patient Representation Learning with Missing Modalities and LabelsabstractMultimodal patient representation learning aims to integrate information from multiple modalities and generate comprehensive patient representations for subsequent clinical predictive tasks. However, many existing approaches either presuppose the availability of all modalities and labels for each patient or only deal with missing modalities. In reality, patient data often comes with both missing modalities and labels for various reasons (i.e., the missing modality and label issue). Moreover, multimodal models might over-rely on certain modalities, causing sub-optimal performance when these modalities are absent (i.e., the modality collapse issue). To address these issues, we introduce MUSE: a mutual-consistent graph contrastive learning method. MUSE uses a flexible bipartite graph to represent the patient-modality relationship, which can adapt to various missing modality patterns. To tackle the modality collapse issue, MUSE learns to focus on modality-general and label-decisive features via a mutual-consistent contrastive learning loss. Notably, the unsupervised component of the contrastive objective only requires self-supervision signals, thereby broadening the training scope to incorporate patients with missing labels. We evaluate MUSE on three publicly available datasets: MIMIC-IV, eICU, and ADNI. Results show that MUSE outperforms all baselines, and MUSE+ further elevates the absolute improvement to ~4% by extending the training scope to patients with absent labels. Zhenbang Wu, Anant Dadu, Nicholas J. Tustison, Brian B. Avants, Mike A. Nalls, Jimeng Sun 0001, Faraz Faghri |
ICLR | 1 |
| 2024 | Instruction Tuning Large Language Models to Understand Electronic Health RecordsabstractLarge language models (LLMs) have shown impressive capabilities in solving a wide range of tasks based on human instructions. However, developing a conversational AI assistant for electronic health record (EHR) data remains challenging due to (1) the lack of large-scale instruction-following datasets and (2) the limitations of existing model architectures in handling complex and heterogeneous EHR data.In this paper, we introduce MIMIC-Instr, a dataset comprising over 400K open-ended instruction-following examples derived from the MIMIC-IV EHR database. This dataset covers various topics and is suitable for instruction-tuning general-purpose LLMs for diverse clinical use cases. Additionally, we propose Llemr, a general framework that enables LLMs to process and interpret EHRs with complex data structures. Llemr demonstrates competitive performance in answering a wide range of patient-related questions based on EHR data.Furthermore, our evaluations on clinical predictive modeling benchmarks reveal that the fine-tuned Llemr achieves performance comparable to state-of-the-art (SOTA) baselines using curated features. The dataset and code are available at \url{https://github.com/zzachw/llemr}. Zhenbang Wu, Anant Dadu, Mike A. Nalls, Faraz Faghri, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsabstractArtificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing significant risks for future model deployment. In this paper, we introduce CARES and aim to comprehensively evaluate the Trustworthiness of Med-LVLMs across the medical domain. We assess the trustworthiness of Med-LVLMs across five dimensions, including trustfulness, fairness, safety, privacy, and robustness. CARES comprises about 41K question-answer pairs in both closed and open-ended formats, covering 16 medical image modalities and 27 anatomical regions. Our analysis reveals that the models consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies and failing to maintain fairness across different demographic groups. Furthermore, they are vulnerable to attacks and demonstrate a lack of privacy awareness. We publicly release our benchmark and code in https://github.com/richard-peng-xia/CARES. Peng Xia 0005, Juanxi Tian, Yangrui Gong, Ruibo Hou, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Zhaoyang Wang 0004, Xiao Wang 0044, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li 0010, Jimeng Sun 0001, ZongYuan Ge, Gang Li 0001, James Zou 0001, Huaxiu Yao |
NeurIPS | 7 |
| 2024 | MCCANet: A multispectral class-constraint attentional neural network for object detection in mining scenes
Zhenbang Wu, Hengkai Li, Beiping Long |
Expert Syst. Appl. | 1 |
| 2023 | MedLink: De-Identified Patient Health Record LinkageabstractA comprehensive patient health history is essential for patient care and healthcare research. However, due to the distributed nature of healthcare services, patient health records are often scattered across multiple systems. Existing record linkage approaches primarily rely on patient identifiers, which have inherent limitations such as privacy invasion and identifier discrepancies. To tackle this problem, we propose linking de-identified patient health records by matching health patterns without strictly relying on sensitive patient identifiers. Our model MedLink solves two challenges faced with the patient linkage task: (1) the challenge of identifying the same patients based on data collected in different timelines as disease progression makes the record matching difficult, and (2) the challenge of identifying distinct health patterns as common medical codes dominate health records and overshadow the more informative low-prevalence codes. To address these challenges, MedLink utilizes bi-directional health prediction to predict future codes forwardly and past codes backwardly, thus accounting for the health progression. MedLink also has a prevalence-aware retrieval design to focus more on the low-prevalence but informative codes during learning. MedLink can be trained end-to-end and is lightweight for efficient inference on large patient databases. We evaluate MedLink against leading baselines on real-world patient datasets, including the critical care dataset MIMIC-III and a large health claims dataset. Results show that MedLink outperforms the best baseline by 4% in top-1 accuracy with only 8% memory cost. Additionally, when combined with existing identifier-based linkage approaches, MedLink can improve their performance by up to 15%. Zhenbang Wu, Cao Xiao, Jimeng Sun 0001 |
KDD | 1 |
| 2023 | PyHealth: A Deep Learning Toolkit for Healthcare ApplicationsabstractDeep learning (DL) has emerged as a promising tool in healthcare applications. However, the reproducibility of many studies in this field is limited by the lack of accessible code implementations and standard benchmarks. To address the issue, we create PyHealth, a comprehensive library to build, deploy, and validate DL pipelines for healthcare applications. PyHealth supports various data modalities, including electronic health records (EHRs), physiological signals, medical images, and clinical text. It offers various advanced DL models and maintains comprehensive medical knowledge systems. The library is designed to support both DL researchers and clinical data scientists. Upon the time of writing, PyHealth has received 633 stars, 130 forks, and 15k+ downloads in total on GitHub. Chaoqi Yang, Zhenbang Wu, Patrick Jiang, Zhen Lin 0001, Benjamin P. Danek, Jimeng Sun 0001 |
KDD | 2 |
| 2023 | An Iterative Self-Learning Framework for Medical Domain GeneralizationabstractDeep learning models have been widely used to assist doctors with clinical decision-making. However, these models often encounter a significant performance drop when applied to data that differs from the distribution they were trained on. This challenge is known as the domain shift problem. Existing domain generalization algorithms attempt to address this problem by assuming the availability of domain IDs and training a single model to handle all domains. However, in healthcare settings, patients can be classified into numerous latent domains, where the actual domain categorizations are unknown. Furthermore, each patient domain exhibits distinct clinical characteristics, making it sub-optimal to train a single model for all domains. To overcome these limitations, we propose SLGD, a self-learning framework that iteratively discovers decoupled domains and trains personalized classifiers for each decoupled domain. We evaluate the generalizability of SLGD across spatial and temporal data distribution shifts on two real-world public EHR datasets: eICU and MIMIC-IV. Our results show that SLGD achieves up to 11% improvement in the AUPRC score over the best baseline. Zhenbang Wu, Huaxiu Yao, David M. Liebovitz, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2022 | MedCLIP: Contrastive Learning from Unpaired Medical Images and TextabstractExisting vision-text contrastive learning like CLIP (Radford et al., 2021) aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability and supports zero-shot prediction. However, medical image-text datasets are orders of magnitude below the general images and captions from the internet. Moreover, previous methods encounter many false negatives, i.e., images and reports from separate patients probably carry the same semantics but are wrongly treated as negatives. In this paper, we decouple images and texts for multimodal contrastive learning thus scaling the usable training data in a combinatorial magnitude with low cost. We also propose to replace the InfoNCE loss with semantic matching loss based on medical knowledge to eliminate false negatives in contrastive learning. We prove that MedCLIP is a simple yet effective framework: it outperforms state-of-the-art methods on zero-shot prediction, supervised classification, and image-text retrieval. Surprisingly, we observe that with only 20K pre-training data, MedCLIP wins over the state-of-the-art method (using ≈200K data). Zifeng Wang 0008, Zhenbang Wu, Dinesh Agarwal, Jimeng Sun 0001 |
EMNLP | 2 |
| 2022 | AutoMap: Automatic Medical Code Mapping for Clinical Prediction Model Deployment
Zhenbang Wu, Cao Xiao, Lucas Glass, David M. Liebovitz, Jimeng Sun 0001 |
ECML/PKDD (2) | 1 |