VLDB 2026 Research / reviewers in the wild / expert
Robert Freeman
dblp:304/2598
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sociodemographic bias in large language model clinical trial screeningabstractBackground: Large language models (LLMs) are increasingly used in randomized clinical trial (RCT) screening, but their potential for sociodemographic bias remains unclear. Objective: To determine whether LLM-based trial screening judgments vary with patient sociodemographic characteristics when clinical details and eligibility criteria are held constant. Design Setting and Participants: Cross-sectional evaluation of Phase II-III RCT protocols from ClinicalTrials.gov (U.S. adult populations; 2023-2024). For each protocol, we created 15 physician-validated clinical vignettes rendered in 34 versions: one control (no identifiers) and 33 identity variants spanning gender, race/ethnicity, socioeconomic status, homelessness, unemployment, and sexual orientation. Exposures: Identity labels applied to otherwise identical vignettes, evaluated by nine contemporary LLMs. Main Outcomes and Measures: Primary: eligibility domain score (1-5 Likert scale) comparing identity variants versus control. Secondary: adherence, resources, risk-benefit, and trust/attitude domains. Mixed-effects models estimated adjusted mean differences with multiplicity-corrected P values; differences <.10 considered trivial. Results: Of 69 protocols, 58 met inclusion criteria. Analysis of 5,324,400 model evaluations showed eligibility judgments were largely stable: most identity-related differences fell within ±0.05 (transgender woman -.008 [95% CI -.04 to .02]; White male .036 [.01 to .07]). Only homelessness exceeded the trivial threshold (-.121 [-.15 to -.09], P<.001). Secondary domains revealed socioeconomic gradients, particularly for adherence (homeless -.595, P<.001) and resources (homeless -.715, P<.001), with smaller trust/attitude effects and negligible risk-benefit differences. Conclusions and Relevance: Bias in LLM-assisted trial screening is conditional. Within fixed criteria, models reason consistently; outside them, they echo the inequities of their data. Responsible deployment in clinical research depends on preserving that boundary so that automation strengthens fairness in trial access rather than inheriting distortion. Shelly Soffer, Mahmud Omar, Orly Efros, Donald Apakama, Aya Mudrik, Robert Freeman, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 6 |
| 2025 | A scalable framework for benchmark embedding models in semantic health-care tasksabstractOBJECTIVES: Text embeddings are promising for semantic tasks, such as retrieval augmented generation (RAG). However, their application in health care is underexplored due to a lack of benchmarking methods. We introduce a scalable benchmarking method to test embeddings for health-care semantic tasks. MATERIALS AND METHODS: We evaluated 39 embedding models across 7 medical semantic similarity tasks using diverse datasets. These datasets comprised real-world patient data (from the Mount Sinai Health System and MIMIC IV), biomedical texts from PubMed, and synthetic data generated with Llama-3-70b. We first assessed semantic textual similarity (STS) by correlating the model-generated similarity scores with noise levels using Spearman rank correlation. We then reframed the same tasks as retrieval problems, evaluated by mean reciprocal rank and recall at k. RESULTS: In total, evaluating 2000 text pairs per 7 tasks for STS and retrieval yielded 3.28 million model assessments. Larger models (>7b parameters), such as those based on Mistral-7b and Gemma-2-9b, consistently performed well, especially in long-context tasks. The NV-Embed-v1 model (7b parameters), although top in short tasks, underperformed in long tasks. For short tasks, smaller models such as b1ade-embed (335M parameters) performed on-par to the larger models. For long retrieval tasks, the larger models significantly outperformed the smaller ones. DISCUSSION: The proposed benchmarking framework demonstrates scalability and flexibility, offering a structured approach to guide the selection of embedding models for a wide range of health-care tasks. CONCLUSION: By matching the appropriate model with the task, the framework enables more effective deployment of embedding models, enhancing critical applications such as semantic search and retrieval-augmented generation (RAG). Shelly Soffer, Mahmud Omar, Moran Gendler, Benjamin S. Glicksberg, Patricia H. Kovatch, Orly Efros, Robert Freeman, Alexander Charney, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomabstractBACKGROUND: Artificial intelligence (AI) and large language models (LLMs) can play a critical role in emergency room operations by augmenting decision-making about patient admission. However, there are no studies for LLMs using real-world data and scenarios, in comparison to and being informed by traditional supervised machine learning (ML) models. We evaluated the performance of GPT-4 for predicting patient admissions from emergency department (ED) visits. We compared performance to traditional ML models both naively and when informed by few-shot examples and/or numerical probabilities. METHODS: We conducted a retrospective study using electronic health records across 7 NYC hospitals. We trained Bio-Clinical-BERT and XGBoost (XGB) models on unstructured and structured data, respectively, and created an ensemble model reflecting ML performance. We then assessed GPT-4 capabilities in many scenarios: through Zero-shot, Few-shot with and without retrieval-augmented generation (RAG), and with and without ML numerical probabilities. RESULTS: The Ensemble ML model achieved an area under the receiver operating characteristic curve (AUC) of 0.88, an area under the precision-recall curve (AUPRC) of 0.72 and an accuracy of 82.9%. The naïve GPT-4's performance (0.79 AUC, 0.48 AUPRC, and 77.5% accuracy) showed substantial improvement when given limited, relevant data to learn from (ie, RAG) and underlying ML probabilities (0.87 AUC, 0.71 AUPRC, and 83.1% accuracy). Interestingly, RAG alone boosted performance to near peak levels (0.82 AUC, 0.56 AUPRC, and 81.3% accuracy). CONCLUSIONS: The naïve LLM had limited performance but showed significant improvement in predicting ED admissions when supplemented with real-world examples to learn from, particularly through RAG, and/or numerical probabilities from traditional ML models. Its peak performance, although slightly lower than the pure ML model, is noteworthy given its potential for providing reasoning behind predictions. Further refinement of LLMs with real-world data is necessary for successful integration as decision-support tools in care settings. Benjamin S. Glicksberg, Prem Timsina, Ashwin Sawant, Akhil Vaid, Ganesh Raut, Alexander Charney, Donald Apakama, Brendan G. Carr, Robert Freeman, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 10 |
| 2024 | Local large language models for privacy-preserving accelerated review of historic echocardiogram reportsabstractOBJECTIVES: The study developed framework that leverages an open-source Large Language Model (LLM) to enable clinicians to ask plain-language questions about a patient's entire echocardiogram report history. This approach is intended to streamline the extraction of clinical insights from multiple echocardiogram reports, particularly in patients with complex cardiac diseases, thereby enhancing both patient care and research efficiency. MATERIALS AND METHODS: Data from over 10 years were collected, comprising echocardiogram reports from patients with more than 10 echocardiograms on file at the Mount Sinai Health System. These reports were converted into a single document per patient for analysis, broken down into snippets and relevant snippets were retrieved using text similarity measures. The LLaMA-2 70B model was employed for analyzing the text using a specially crafted prompt. The model's performance was evaluated against ground-truth answers created by faculty cardiologists. RESULTS: The study analyzed 432 reports from 37 patients for a total of 100 question-answer pairs. The LLM correctly answered 90% questions, with accuracies of 83% for temporality, 93% for severity assessment, 84% for intervention identification, and 100% for diagnosis retrieval. Errors mainly stemmed from the LLM's inherent limitations, such as misinterpreting numbers or hallucinations. CONCLUSION: The study demonstrates the feasibility and effectiveness of using a local, open-source LLM for querying and interpreting echocardiogram report data. This approach offers a significant improvement over traditional keyword-based searches, enabling more contextually relevant and semantically accurate responses; in turn showing promise in enhancing clinical decision-making and research by facilitating more efficient access to complex patient data. Akhil Vaid, Son Q. Duong, Joshua Lampert, Patricia H. Kovatch, Robert Freeman, Edgar Argulian, Lori Croft, Stamatios Lerakis, Martin Goldman, Rohan Khera, Girish N. Nadkarni |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Innovating in a crisis: a qualitative evaluation of a hospital and Google partnership to implement a COVID-19 inpatient video monitoring programabstractOBJECTIVE: To describe adaptations necessary for effective use of direct-to-consumer (DTC) cameras in an inpatient setting, from the perspective of health care workers. METHODS: Our qualitative study included semi-structured interviews and focus groups with clinicians, information technology (IT) personnel, and health system leaders affiliated with the Mount Sinai Health System. All participants either worked in a coronavirus disease 2019 (COVID-19) unit with DTC cameras or participated in the camera implementation. Three researchers coded the transcripts independently and met weekly to discuss and resolve discrepancies. Abiding by inductive thematic analysis, coders revised the codebook until they reached saturation. All transcripts were coded in Dedoose using the final codebook. RESULTS: Frontline clinical staff, IT personnel, and health system leaders (N = 39) participated in individual interviews and focus groups in November 2020-April 2021. Our analysis identified 5 areas for effective DTC camera use: technology, patient monitoring, workflows, interpersonal relationships, and infrastructure. Participants described adaptations created to optimize camera use and opportunities for improvement necessary for sustained use. Non-COVID-19 patients tended to decline participation. DISCUSSION: Deploying DTC cameras on inpatient units required adaptations in many routine processes. Addressing consent, 2-way communication issues, patient privacy, and messaging about video monitoring could help facilitate a nimble rollout. Implementation and dissemination of inpatient video monitoring using DTC cameras requires input from patients and frontline staff. CONCLUSIONS: Given the resources and time it takes to implement a usable camera solution, other health systems might benefit from creating task forces to investigate their use before the next crisis. Ksenia Gorbenko, Afrah Mohammed, Edward II Ezenwafor, Sydney Phlegar, Patrick Healy, Tamara Solly, Ingrid Nembhard, Lucy Xenophon, Cardinale Smith, Robert Freeman, David Reich, Madhu Mazumdar |
J. Am. Medical Informatics Assoc. | 10 |