EDBT 2026 Demo / reviewers in the wild / expert
Zifeng Wang 0008
dblp:43/7716-8
· DBLP profile ↗
19ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-3026-9970ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compliance and factuality of large language models for clinical research document generationabstractOBJECTIVES: Large language models' (LLMs') performance in high-stakes, compliance-driven settings such as drafting clinical research documents remains underexplored. This study aims to build a benchmark and an evaluation framework for assessing LLMs' compliance and factuality in generating informed consent forms (ICFs) from clinical trial protocols. MATERIALS AND METHODS: We introduce InformBench, a benchmark comprising 900 clinical trial documents, and propose an evaluation framework grounded in regulatory guidelines and site-specific consent templates. We assess LLM performance on transforming trial protocols, often hundreds of pages, into concise, patient-facing ICFs. Additionally, we design InformGen, a retrieval-augmented, human-in-the-loop pipeline aimed at improving generation quality. RESULTS: Baseline LLMs such as GPT-4o achieved only 70%-80% compliance and exhibited factual errors in 18%-43% of cases. In contrast, InformGen substantially improved outputs, achieving nearly 100% regulatory compliance and over 90% factual accuracy, as validated by 5 domain-expert annotators. DISCUSSION: The study reveals critical limitations in current LLMs for clinical research document drafting, particularly in regulatory sensitivity and factual grounding. Our results highlight the need for domain-specific benchmarks and structured evaluations to support safe deployment in real-world clinical research workflows. CONCLUSION: LLMs offer value in clinical research document generation but must be adapted and rigorously evaluated for high-stakes applications. Our benchmark and framework provide a foundation for improving and assessing LLM-generated outputs in compliance-critical domains. Zifeng Wang 0008, Benjamin P. Danek, Brandon Theodorou, Ruba Shaik, Shivashankar Thati, Seunghyun Won, Jimeng Sun 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | s3: You Don't Need That Much Data to Train a Search Agent via RLabstractRetrieval-augmented generation (RAG) systems empower large language models (LLMs) to access external knowledge during inference.Recent advances have enabled LLMs to act as search agents via reinforcement learning (RL), improving information acquisition through multi-turn interactions with retrieval engines.However, existing approaches either optimize retrieval using search-only metrics (e.g., NDCG) that ignore downstream utility or fine-tune the entire LLM to jointly reason and retrieve-entangling retrieval with generation and limiting the real search utility and compatibility with frozen or proprietary models.In this work, we propose s3, a lightweight, modelagnostic framework that decouples the searcher from the generator and trains the searcher using a Gain Beyond RAG reward: the improvement in generation accuracy over naïve RAG.s3 requires only 2.4k training samples to outperform baselines trained on over 70× more data, consistently delivering stronger downstream performance across six general QA and five medical QA benchmarks. 1 Pengcheng Jiang, Xueqiang Xu, Jiacheng Lin, Jinfeng Xiao, Zifeng Wang 0008, Jimeng Sun 0001, Jiawei Han 0001 |
EMNLP | 5 |
| 2024 | MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language ModelsabstractLarge language models (LLMs) have achieved remarkable performance in natural language understanding and generation tasks.However, they often suffer from limitations such as difficulty in incorporating new knowledge, generating hallucinations, and explaining their reasoning process.To address these challenges, we propose a novel prompting pipeline, named MindMap, that leverages knowledge graphs (KGs) to enhance LLMs' inference and transparency.Our method enables LLMs to comprehend KG inputs and infer with a combination of implicit and external knowledge.Moreover, our method elicits the mind map of LLMs, which reveals their reasoning pathways based on the ontology of knowledge.We evaluate our method on diverse question & answering tasks, especially in medical domains, and show significant improvements over baselines.We also introduce a new hallucination evaluation benchmark and analyze the effects of different components of our method.Our results demonstrate the effectiveness and robustness of our method in merging knowledge from LLMs and KGs for combined inference.To reproduce our results and extend the framework further, we make our codebase available at https://github.com/wyl- willing/MindMap. Yilin Wen 0006, Zifeng Wang 0008, Jimeng Sun 0001 |
ACL (1) | 2 |
| 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge GraphsabstractFoundation models (FMs) learn from large volumes of unlabeled data to demonstrate superior performance across a wide range of tasks. However, FMs developed for biomedical domains have largely remained unimodal, i.e., independently trained and used for tasks on protein sequences alone, small molecule structures alone, or clinical data alone.
To overcome this limitation, we present BioBridge, a parameter-efficient learning framework, to bridge independently trained unimodal FMs to establish multimodal behavior. BioBridge achieves it by utilizing Knowledge Graphs (KG) to learn transformations between one unimodal FM and another without fine-tuning any underlying unimodal FMs.
Our results demonstrate that BioBridge can
beat the best baseline KG embedding methods (on average by ~ 76.3%) in cross-modal retrieval tasks. We also identify BioBridge demonstrates out-of-domain generalization ability by extrapolating to unseen modalities or relations. Additionally, we also show that BioBridge presents itself as a general-purpose retriever that can aid biomedical multimodal question answering as well as enhance the guided generation of novel drugs. Code is at https://github.com/RyanWangZf/BioBridge. Zifeng Wang 0008, Zichen Wang 0002, Vassilis N. Ioannidis, Huzefa Rangwala, Rishita Anubhai |
ICLR | 1 |
| 2024 | MediTab: Scaling Medical Tabular Data Predictors via Data Consolidation, Enrichment, and Refinement
Zifeng Wang 0008, Chufan Gao, Cao Xiao, Jimeng Sun 0001 |
IJCAI | 1 |
| 2024 | PILOT: Legal Case Outcome Prediction with Case LawabstractLang Cao, Zifeng Wang, Cao Xiao, Jimeng Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lang Cao, Zifeng Wang 0008, Cao Xiao, Jimeng Sun 0001 |
NAACL-HLT | 2 |
| 2024 | GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language ModelsabstractPengcheng Jiang, Jiacheng Lin, Zifeng Wang, Jimeng Sun, Jiawei Han. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Pengcheng Jiang, Jiacheng Lin, Zifeng Wang 0008, Jimeng Sun 0001, Jiawei Han 0001 |
NAACL-HLT | 3 |
| 2024 | TriSum: Learning Summarization Ability from Large Language Models with Structured RationaleabstractPengcheng Jiang, Cao Xiao, Zifeng Wang, Parminder Bhatia, Jimeng Sun, Jiawei Han. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Pengcheng Jiang, Cao Xiao, Zifeng Wang 0008, Parminder Bhatia, Jimeng Sun 0001, Jiawei Han 0001 |
NAACL-HLT | 3 |
| 2023 | AutoTrial: Prompting Language Models for Clinical Trial DesignabstractClinical trials are critical for drug development.Constructing the appropriate eligibility criteria (i.e., the inclusion/exclusion criteria for patient recruitment) is essential for the trial's success.Proper design of clinical trial protocols should consider similar precedent trials and their eligibility criteria to ensure sufficient patient coverage.In this paper, we present a method named AutoTrial to aid the design of clinical eligibility criteria using language models.It allows (1) controllable generation under instructions via a hybrid of discrete and neural prompting, (2) scalable knowledge incorporation via in-context learning, and (3) explicit reasoning chains to provide rationales for understanding the outputs.Experiments on over 70K clinical trials verify that AutoTrial generates high-quality criteria texts that are fluent and coherent and with high accuracy in capturing the relevant clinical concepts to the target trial.It is noteworthy that our method, with a much smaller parameter size, gains around 60% winning rate against the GPT-3.5 baselines via human evaluations. Zifeng Wang 0008, Cao Xiao, Jimeng Sun 0001 |
EMNLP | 1 |
| 2023 | TWIN: Personalized Clinical Trial Digital Twin GenerationabstractClinical trial digital twins are virtual patients that reflect personal characteristics in a high degree of granularity and can be used to simulate various patient outcomes under different conditions. With the growth of clinical trial databases captured by Electronic Data Capture (EDC) systems, there is a growing interest in using machine learning models to generate digital twins. This can benefit the drug development process by reducing the sample size required for participant recruitment, improving patient outcome predictive modeling, and mitigating privacy risks when sharing synthetic clinical trial data. However, prior research has mainly focused on generating Electronic Healthcare Records (EHRs), which often assume large training data and do not account for personalized synthetic patient record generation. In this paper, we propose a sample-efficient method TWIN for generating personalized clinical trial digital twins. TWIN can produce digital twins of patient-level clinical trial records with high fidelity to the targeting participant's record and preserves the temporal relations across visits and events. We compare our method with various baselines for generating real-world patient-level clinical trial data. The results show that TWIN generates synthetic trial data with high fidelity to facilitate patient outcome predictions in low-data scenarios and strong privacy protection against real patients from the trials. Trisha Das, Zifeng Wang 0008, Jimeng Sun 0001 |
KDD | 2 |
| 2022 | Finding Influential Instances for Distantly Supervised Relation ExtractionabstractDistant supervision (DS) is a strong way to expand the datasets for enhancing relation extraction (RE) models but often suffers from high label noise. Current works based on attention, reinforcement learning, or GAN are black-box models so they neither provide meaningful interpretation of sample selection in DS nor stability on different domains. On the contrary, this work proposes a novel model-agnostic instance sampling method for DS by influence function (IF), namely REIF. Our method identifies favorable/unfavorable instances in the bag based on IF, then does dynamic instance sampling. We design a fast influence sampling algorithm that reduces the computational complexity from \mathcal{O}(mn) to \mathcal{O}(1), with analyzing its robustness on the selected sampling function. Experiments show that by simply sampling the favorable instances during training, REIF is able to win over a series of baselines which have complicated architectures. We also demonstrate that REIF can support interpretable instance selection. Zifeng Wang 0008, Rui Wen 0001, Xi Chen 0003, Shao-Lun Huang, Ningyu Zhang 0001, Yefeng Zheng 0001 |
COLING | 1 |
| 2022 | MedCLIP: Contrastive Learning from Unpaired Medical Images and TextabstractExisting vision-text contrastive learning like CLIP (Radford et al., 2021) aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability and supports zero-shot prediction. However, medical image-text datasets are orders of magnitude below the general images and captions from the internet. Moreover, previous methods encounter many false negatives, i.e., images and reports from separate patients probably carry the same semantics but are wrongly treated as negatives. In this paper, we decouple images and texts for multimodal contrastive learning thus scaling the usable training data in a combinatorial magnitude with low cost. We also propose to replace the InfoNCE loss with semantic matching loss based on medical knowledge to eliminate false negatives in contrastive learning. We prove that MedCLIP is a simple yet effective framework: it outperforms state-of-the-art methods on zero-shot prediction, supervised classification, and image-text retrieval. Surprisingly, we observe that with only 20K pre-training data, MedCLIP wins over the state-of-the-art method (using ≈200K data). Zifeng Wang 0008, Zhenbang Wu, Dinesh Agarwal, Jimeng Sun 0001 |
EMNLP | 1 |
| 2022 | PromptEHR: Conditional Electronic Healthcare Records Generation with Prompt LearningabstractAccessing longitudinal multimodal Electronic Healthcare Records (EHRs) is challenging due to privacy concerns, which hinders the use of ML for healthcare applications. Synthetic EHRs generation bypasses the need to share sensitive real patient records. However, existing methods generate single-modal EHRs by unconditional generation or by longitudinal inference, which falls short of low flexibility and makes unrealistic EHRs. In this work, we propose to formulate EHRs generation as a text-to-text translation task by language models (LMs), which suffices to highly flexible event imputation during generation. We also design prompt learning to control the generation conditioned by numerical and categorical demographic features. We evaluate synthetic EHRs quality by two perplexity measures accounting for their longitudinal pattern (longitudinal imputation perplexity, lpl) and the connections cross modalities (cross-modality imputation perplexity, mpl). Moreover, we utilize two adversaries: membership and attribute inference attacks for privacy-preserving evaluation. Experiments on MIMIC-III data demonstrate the superiority of our methods on realistic EHRs generation (53.1% decrease of lpl and 45.3% decrease of mpl on average compared to the best baselines) with low privacy risks. Zifeng Wang 0008, Jimeng Sun 0001 |
EMNLP | 1 |
| 2022 | PAC-Bayes Information Bottleneck
Zifeng Wang 0008, Shao-Lun Huang, Ercan E. Kuruoglu, Jimeng Sun 0001, Xi Chen 0003, Yefeng Zheng 0001 |
ICLR | 1 |
| 2022 | TransTab: Learning Transferable Tabular Transformers Across TablesabstractTabular data (or tables) are the most widely used data format in machine learning (ML). However, ML models often assume the table structure keeps fixed in training and testing. Before ML modeling, heavy data cleaning is required to merge disparate tables with different columns. This preprocessing often incurs significant data waste (e.g., removing unmatched columns and samples). How to learn ML models from multiple tables with partially overlapping columns? How to incrementally update ML models as more columns become available over time? Can we leverage model pretraining on multiple distinct tables? How to train an ML model which can predict on an unseen table? To answer all those questions, we propose to relax fixed table structures by introducing a Transferable Tabular Transformer (TransTab) for tables. The goal of TransTab is to convert each sample (a row in the table) to a generalizable embedding vector, and then apply stacked transformers for feature encoding. One methodology insight is combining column description and table cells as the raw input to a gated transformer model. The other insight is to introduce supervised and self-supervised pretraining to improve model performance. We compare TransTab with multiple baseline methods on diverse benchmark datasets and five oncology clinical trial datasets. Overall, TransTab ranks 1.00, 1.00, 1.78 out of 12 methods in supervised learning, incremental feature learning, and transfer learning scenarios, respectively; and the proposed pretraining leads to 2.3\% AUC lift on average over the supervised learning. Zifeng Wang 0008, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2021 | Lifelong Learning Based Disease Diagnosis on Clinical Notes
Zifeng Wang 0008, Yifan Yang 0006, Rui Wen 0001, Xi Chen 0003, Shao-Lun Huang, Yefeng Zheng 0001 |
PAKDD (1) | 1 |
| 2021 | Online Disease Diagnosis with Inductive Heterogeneous Graph Convolutional NetworksabstractWe propose a Healthcare Graph Convolutional Network (HealGCN) to offer disease self-diagnosis service for online users based on Electronic Healthcare Records (EHRs). Two main challenges are focused in this paper for online disease diagnosis: (1) serving cold-start users via graph convolutional networks and (2) handling scarce clinical description via a symptom retrieval system. To this end, we first organize the EHR data into a heterogeneous graph that is capable of modeling complex interactions among users, symptoms and diseases, and tailor the graph representation learning towards disease diagnosis with an inductive learning paradigm. Then, we build a disease self-diagnosis system with a corresponding EHR Graph-based Symptom Retrieval System (GraphRet) that can search and provide a list of relevant alternative symptoms by tracing the predefined meta-paths. GraphRet helps enrich the seed symptom set through the EHR graph when confronting users with scarce descriptions, hence yield better diagnosis accuracy. At last, we validate the superiority of our model on a large-scale EHR dataset. Zifeng Wang 0008, Rui Wen 0001, Xi Chen 0003, Shilei Cao 0001, Shao-Lun Huang, Buyue Qian, Yefeng Zheng 0001 |
WWW | 1 |
| 2020 | Less Is Better: Unweighted Data Subsampling via Influence FunctionabstractIn the time of Big Data, training complex models on large-scale data sets is challenging, making it appealing to reduce data volume for saving computation resources by subsampling. Most previous works in subsampling are weighted methods designed to help the performance of subset-model approach the full-set-model, hence the weighted methods have no chance to acquire a subset-model that is better than the full-set-model. However, we question that how can we achieve better model with less data? In this work, we propose a novel Unweighted Influence Data Subsampling (UIDS) method, and prove that the subset-model acquired through our method can outperform the full-set-model. Besides, we show that overly confident on a given test set for sampling is common in Influence-based subsampling methods, which can eventually cause our subset-model's failure in out-of-sample test. To mitigate it, we develop a probabilistic sampling scheme to control the worst-case risk over all distributions close to the empirical distribution. The experiment results demonstrate our methods superiority over existed subsampling methods in diverse tasks, such as text classification, image classification, click-through prediction, etc. Zifeng Wang 0008, Hong Zhu 0003, Zhenhua Dong, Xiuqiang He 0001, Shao-Lun Huang |
AAAI | 1 |
| 2020 | Information Theoretic Counterfactual Learning from Missing-Not-At-Random FeedbackabstractCounterfactual learning for dealing with missing-not-at-random data (MNAR) is an intriguing topic in the recommendation literature, since MNAR data are ubiquitous in modern recommender systems. Instead, missing-at-random (MAR) data, namely randomized controlled trials (RCTs), are usually required by most previous counterfactual learning methods. However, the execution of RCTs is extraordinarily expensive in practice. To circumvent the use of RCTs, we build an information theoretic counterfactual variational information bottleneck (CVIB), as an alternative for debiasing learning without RCTs. By separating the task-aware mutual information term in the original information bottleneck Lagrangian into factual and counterfactual parts, we derive a contrastive information loss and an additional output confidence penalty, which facilitates balanced learning between the factual and counterfactual domains. Empirical evaluation on real-world datasets shows that our CVIB significantly enhances both shallow and deep models, which sheds light on counterfactual learning in recommendation that goes beyond RCTs. Zifeng Wang 0008, Xi Chen 0003, Rui Wen 0001, Shao-Lun Huang, Ercan E. Kuruoglu, Yefeng Zheng 0001 |
NeurIPS | 1 |