VLDB 2026 Research / reviewers in the wild / expert
Yilu Fang
dblp:298/1804
· DBLP profile ↗
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-2681-1931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A critical evaluation of generative query expansion on biomedical literature retrievalabstractOBJECTIVE: To evaluate the effectiveness of generative query expansion for biomedical literature retrieval. MATERIALS AND METHODS: We thoroughly examined eight generative query expansion methods using three large language models across five datasets for biomedical literature retrieval. We further performed a quantitative analysis, including performance comparisons, rank transition analysis, and article-type effect analysis. We also conducted a qualitative examination of representative cases, from which we derived an error taxonomy. RESULTS: On BioASQ-Y/N, GPT-4o-based query expansion shifts Recall@10 to 0.417-0.512 and nDCG@10 to 0.358-0.479, relative to a baseline of 0.491 and 0.456. For PubMedQA, Precision@1 ranges from 0.764 to 0.876 and nDCG@10 from 0.847 to 0.931, compared with baseline values of 0.893 and 0.935. For 2019-Trec-PM, query expansion yields Recall@100 of 0.217-0.256 and nDCG@100 of 0.272-0.312, versus a baseline of 0.227 and 0.274. Similarly, for 2018-TREC-PM, Recall@100 spans 0.169-0.227 and nDCG@100 spans 0.195-0.250, relative to baseline scores of 0.164 and 0.191. For 2017-TREC-PM, Recall@100 and nDCG@100 fall within 0.111-0.139 and 0.154-0.191 under query expansion, compared with baseline metrics of 0.102 and 0.147. Both general-purpose and domain-specific Llama-based models demonstrate similar performance to GPT-4o. DISCUSSION AND CONCLUSION: The impact of query expansion varies significantly by the expansion methods and type of evidence, but is relatively agnostic to backbone model choice. Notably, query expansion primarily affects article ranking but has a limited impact on the screening stage. Our findings underscore the unique challenges of biomedical literature retrieval and highlight the need to develop domain-specific information retrieval techniques. Yilu Fang, Fangyi Chen, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |
| 2026 | A data-driven method for research trend analysis in a scientific discipline: Application to the journal of biomedical informaticsabstractOBJECTIVE: Accurately characterizing research trends is critical for identifying cutting-edge scientific breakthroughs in their infancy and informing strategic priorities. This research contributes a pipeline that utilizes generative AI technologies to develop research topic taxonomies from publication keywords and analyze keyword evolution within topics, methodological and domain trends, and topic co-occurrences. We demonstrated the pipeline by conducting a retrospective analysis of biomedical informatics research trends in the Journal of Biomedical Informatics (JBI). METHODS: We identified the JBI publications with keywords available on PubMed, spanning 2011-2025. We downloaded all the keywords and categorized them into methodological innovations and health domains, identified topics, assigned topic names, and constructed their hierarchies, all using large-language models (LLMs). We introduced an automated method for evaluating topics, leveraging MeSH terminology as the underlying knowledge base. RESULTS: Using 6,930 unique keywords from 2,427 publications, we derived 1,028 distinct topics related to methodological innovations, with each topic associated with medians of four keywords (Q1: 2, Q3: 13) and six publications (Q1: 2, Q3: 19). We identified 904 topics related to health domains, with each topic associated with three keywords (Q1: 1, Q3: 11) and four publications (Q1: 1, Q3: 15). Based on the topics, we analyzed the prominent research areas, trends in publication volume, evolution of keyword distributions within each topic, and patterns of co-occurring topics. Among the 2,379 eligible publications, 2,009 (84.4%) exhibited overlap between the keyword-derived MeSH terms and the MeSH terms assigned to the publication by the National Library of Medicine. CONCLUSION: This study presents a method that leverages modern generative AI technologies for retrospective analysis of a scientific field to identify emerging topics and to detect shifts in scholarly focus. Illustrated by data for JBI and correlated with historical background events and policy changes, our findings demonstrate the effectiveness and utility of the methods while providing a powerful lens to understand the evolution of biomedical informatics research priorities in JBI. Yilu Fang, Samir Sanchez Tejada, Fangyi Chen, Edward H. Shortliffe, Vimla L. Patel, Mor Peleg, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2025 | Semi-supervised learning from small annotated data and large unlabeled data for fine-grained Participants, Intervention, Comparison, and Outcomes entity recognitionabstractOBJECTIVE: Extracting PICO elements-Participants, Intervention, Comparison, and Outcomes-from clinical trial literature is essential for clinical evidence retrieval, appraisal, and synthesis. Existing approaches do not distinguish the attributes of PICO entities. This study aims to develop a named entity recognition (NER) model to extract PICO entities with fine granularities. MATERIALS AND METHODS: Using a corpus of 2511 abstracts with PICO mentions from 4 public datasets, we developed a semi-supervised method to facilitate the training of a NER model, FinePICO, by combining limited annotated data of PICO entities and abundant unlabeled data. For evaluation, we divided the entire dataset into 2 subsets: a smaller group with annotations and a larger group without annotations. We then established the theoretical lower and upper performance bounds based on the performance of supervised learning models trained solely on the small, annotated subset and on the entire set with complete annotations, respectively. Finally, we evaluated FinePICO on both the smaller annotated subset and the larger, initially unannotated subset. We measured the performance of FinePICO using precision, recall, and F1. RESULTS: Our method achieved precision/recall/F1 of 0.567/0.636/0.60, respectively, using a small set of annotated samples, outperforming the baseline model (F1: 0.437) by more than 16%. The model demonstrates generalizability to a different PICO framework and to another corpus, which consistently outperforms the benchmark in diverse experimental settings (P-value < .001). DISCUSSION: We developed FinePICO to recognize fine-grained PICO entities from text and validated its performance across diverse experimental settings, highlighting the feasibility of using semi-supervised learning (SSL) techniques to enhance PICO entities extraction. Future work can focus on optimizing SSL algorithms to improve efficiency and reduce computational costs. CONCLUSION: This study contributes a generalizable and effective semi-supervised approach leveraging large unlabeled data together with small, annotated data for fine-grained PICO extraction. Fangyi Chen, Yilu Fang, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | A method for characterizing disease progression from acute kidney injury to chronic kidney disease
Yilu Fang, Jordan G. Nestor, Casey N. Ta, Jerard Kneifati-Hayek, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2025 | CLEAR: A vision to support clinical evidence lifecycle with continuous learningabstractHuman knowledge of diseases, treatments, and prevention techniques is constantly evolving. The generation of clinical evidence using randomized controlled trials on human subjects occurs notably slowly and inefficiently. The Learning Health System (LHS) has been proposed to facilitate the continuous improvement of individual and population health through a cycle of knowledge, practice, and data. However, the gap between the demand for high-quality evidence to support clinical decisions and the available evidence continues to enlarge. While the current LHS vision articulates the integration of Real-World Data (RWD), the rapid generation of RWD often outpaces the rate of effective evidence synthesis and implementation. Considering this, we propose a new framework that more effectively leverages RWD to support the entire clinical evidence lifecycle through a continuous learning mechanism. This framework, powered by modern data science and informatics, offers enhanced scalability and efficiency. In this vision, specifically, RWD is integrated into the clinical evidence lifecycle via four closed feedback loops: 1) guiding research prioritization and study design, 2) facilitating clinical guideline development, 3) assisting guideline evaluation, and 4) supporting shared decision-making. Our framework enables rapid responsiveness to emerging health data and evolving healthcare needs, timely development of clinical guidelines to optimize clinical recommendations, and sustained improvements in clinical practice and patient outcomes. This vision calls for informatics support for an efficient, scalable, and stakeholder-aware clinical evidence lifecycle. Yilu Fang, Fangyi Chen, George Hripcsak, Yifan Peng 0002, Patrick B. Ryan, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2025 | Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2024 | Knowledge-guided generative artificial intelligence for automated taxonomy learning from drug labelsabstractOBJECTIVES: To automatically construct a drug indication taxonomy from drug labels using generative Artificial Intelligence (AI) represented by the Large Language Model (LLM) GPT-4 and real-world evidence (RWE). MATERIALS AND METHODS: We extracted indication terms from 46 421 free-text drug labels using GPT-4, iteratively and recursively generated indication concepts and inferred indication concept-to-concept and concept-to-term subsumption relations by integrating GPT-4 with RWE, and created a drug indication taxonomy. Quantitative and qualitative evaluations involving domain experts were performed for cardiovascular (CVD), Endocrine, and Genitourinary system diseases. RESULTS: 2909 drug indication terms were extracted and assigned into 24 high-level indication categories (ie, initially generated concepts), each of which was expanded into a sub-taxonomy. For example, the CVD sub-taxonomy contains 242 concepts, spanning a depth of 11, with 170 being leaf nodes. It collectively covers a total of 234 indication terms associated with 189 distinct drugs. The accuracies of GPT-4 on determining the drug indication hierarchy exceeded 0.7 with "good to very good" inter-rater reliability. However, the accuracies of the concept-to-term subsumption relation checking varied greatly, with "fair to moderate" reliability. DISCUSSION AND CONCLUSION: We successfully used generative AI and RWE to create a taxonomy, with drug indications adequately consistent with domain expert expectations. We show that LLMs are good at deriving their own concept hierarchies but still fall short in determining the subsumption relations between concepts and terms in unregulated language from free-text drug labels, which is the same hard task for human experts. Yilu Fang, Patrick Ryan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Sociotechnical feasibility of natural language processing-driven tools in clinical trial eligibility prescreening for Alzheimer's disease and related dementiasabstractBACKGROUND: Alzheimer's disease and related dementias (ADRD) affect over 55 million globally. Current clinical trials suffer from low recruitment rates, a challenge potentially addressable via natural language processing (NLP) technologies for researchers to effectively identify eligible clinical trial participants. OBJECTIVE: This study investigates the sociotechnical feasibility of NLP-driven tools for ADRD research prescreening and analyzes the tools' cognitive complexity's effect on usability to identify cognitive support strategies. METHODS: A randomized experiment was conducted with 60 clinical research staff using three prescreening tools (Criteria2Query, Informatics for Integrating Biology and the Bedside [i2b2], and Leaf). Cognitive task analysis was employed to analyze the usability of each tool using the Health Information Technology Usability Evaluation Scale. Data analysis involved calculating descriptive statistics, interrater agreement via intraclass correlation coefficient, cognitive complexity, and Generalized Estimating Equations models. RESULTS: Leaf scored highest for usability followed by Criteria2Query and i2b2. Cognitive complexity was found to be affected by age, computer literacy, and number of criteria, but was not significantly associated with usability. DISCUSSION: Adopting NLP for ADRD prescreening demands careful task delegation, comprehensive training, precise translation of eligibility criteria, and increased research accessibility. The study highlights the relevance of these factors in enhancing NLP-driven tools' usability and efficacy in clinical research prescreening. CONCLUSION: User-modifiable NLP-driven prescreening tools were favorably received, with system type, evaluation sequence, and user's computer literacy influencing usability more than cognitive complexity. The study emphasizes NLP's potential in improving recruitment for clinical trials, endorsing a mixed-methods approach for future system evaluation and enhancements. Betina Ross S. Idnay, Jianfang Liu, Yilu Fang, Alex Hernandez, Shivani Kaw, Alicia Etwaru, Janeth Juarez Padilla, Sergio Ozoria Ramirez, Karen Marder, Chunhua Weng, Rebecca Schnall |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Promoting equity in clinical research: The role of social determinants of health
Betina Ross S. Idnay, Yilu Fang, Edward Stanley, Brenda Ruotolo, Wendy K. Chung, Karen Marder, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2024 | Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation
Jimyung Park, Yilu Fang, Casey N. Ta, Betina Ross S. Idnay, Fangyi Chen, Rebecca Shyu, Emily R. Gordon, Matthew E. Spotnitz, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2023 | Clinical and temporal characterization of COVID-19 subgroups using patient vector embeddings of electronic health recordsabstractOBJECTIVE: To identify and characterize clinical subgroups of hospitalized Coronavirus Disease 2019 (COVID-19) patients. MATERIALS AND METHODS: Electronic health records of hospitalized COVID-19 patients at NewYork-Presbyterian/Columbia University Irving Medical Center were temporally sequenced and transformed into patient vector representations using Paragraph Vector models. K-means clustering was performed to identify subgroups. RESULTS: A diverse cohort of 11 313 patients with COVID-19 and hospitalizations between March 2, 2020 and December 1, 2021 were identified; median [IQR] age: 61.2 [40.3-74.3]; 51.5% female. Twenty subgroups of hospitalized COVID-19 patients, labeled by increasing severity, were characterized by their demographics, conditions, outcomes, and severity (mild-moderate/severe/critical). Subgroup temporal patterns were characterized by the durations in each subgroup, transitions between subgroups, and the complete paths throughout the course of hospitalization. DISCUSSION: Several subgroups had mild-moderate severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections but were hospitalized for underlying conditions (pregnancy, cardiovascular disease [CVD], etc.). Subgroup 7 included solid organ transplant recipients who mostly developed mild-moderate or severe disease. Subgroup 9 had a history of type-2 diabetes, kidney and CVD, and suffered the highest rates of heart failure (45.2%) and end-stage renal disease (80.6%). Subgroup 13 was the oldest (median: 82.7 years) and had mixed severity but high mortality (33.3%). Subgroup 17 had critical disease and the highest mortality (64.6%), with age (median: 68.1 years) being the only notable risk factor. Subgroups 18-20 had critical disease with high complication rates and long hospitalizations (median: 40+ days). All subgroups are detailed in the full text. A chord diagram depicts the most common transitions, and paths with the highest prevalence, longest hospitalizations, lowest and highest mortalities are presented. Understanding these subgroups and their pathways may aid clinicians in their decisions for better management and earlier intervention for patients. Casey N. Ta, Jason Zucker 0001, Po-Hsiang Chiu, Yilu Fang, Karthik Natarajan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | A data-driven approach to optimizing clinical study eligibility criteria
Yilu Fang, Hao Liu 0054, Betina Ross S. Idnay, Casey N. Ta, Karen Marder, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2022 | Optimizing Clinical Research Eligibility Prescreening: An Iterative Usability Evaluation of an NLP-driven Cohort Identification Tool
Betina Ross S. Idnay, Yilu Fang, Caitlin N. Dreisbach, Karen Marder, Chunhua Weng, Rebecca Schnall |
AMIA | 2 |
| 2022 | Criteria2Query 2.0: Combining Machine Efficiency and Human Intelligence to Define a More Accurate and Feasible Cohort for Clinical Trial Recruitment
Betina Ross S. Idnay, Yilu Fang, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Rebecca Schnall, Chunhua Weng |
AMIA | 2 |
| 2022 | Combining human and machine intelligence for clinical trial eligibility queryingabstractOBJECTIVE: To combine machine efficiency and human intelligence for converting complex clinical trial eligibility criteria text into cohort queries. MATERIALS AND METHODS: Criteria2Query (C2Q) 2.0 was developed to enable real-time user intervention for criteria selection and simplification, parsing error correction, and concept mapping. The accuracy, precision, recall, and F1 score of enhanced modules for negation scope detection, temporal and value normalization were evaluated using a previously curated gold standard, the annotated eligibility criteria of 1010 COVID-19 clinical trials. The usability and usefulness were evaluated by 10 research coordinators in a task-oriented usability evaluation using 5 Alzheimer's disease trials. Data were collected by user interaction logging, a demographic questionnaire, the Health Information Technology Usability Evaluation Scale (Health-ITUES), and a feature-specific questionnaire. RESULTS: The accuracies of negation scope detection, temporal and value normalization were 0.924, 0.916, and 0.966, respectively. C2Q 2.0 achieved a moderate usability score (3.84 out of 5) and a high learnability score (4.54 out of 5). On average, 9.9 modifications were made for a clinical study. Experienced researchers made more modifications than novice researchers. The most frequent modification was deletion (5.35 per study). Furthermore, the evaluators favored cohort queries resulting from modifications (score 4.1 out of 5) and the user engagement features (score 4.3 out of 5). DISCUSSION AND CONCLUSION: Features to engage domain experts and to overcome the limitations in automated machine output are shown to be useful and user-friendly. We concluded that human-computer collaboration is key to improving the adoption and user-friendliness of natural language processing. Yilu Fang, Betina Ross S. Idnay, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Karen Marder, Hua Xu 0001, Rebecca Schnall, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |