VLDB 2026 Research / reviewers in the wild / expert
Danielle S. Bitterman
dblp:281/1619
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-0345-2232ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICsabstractCurrent ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report summaries. However, LLMs have demonstrated significantly varied performance across different languages in natural language question-answering tasks, potentially exacerbating healthcare disparities in Low and Middle-Income Countries (LMICs). This study introduces the first multilingual ophthalmological question-answering benchmark with manually curated questions parallel across languages, allowing for direct cross-lingual comparisons. Our evaluation of 6 popular LLMs across 7 different languages reveals substantial bias across different languages, highlighting risks for clinical deployment of LLMs in LMICs. Existing debiasing methods such as Translation Chain-of-Thought or Retrieval-augmented generation (RAG) by themselves fall short of closing this performance gap, often failing to improve performance across all languages and lacking specificity for the medical domain. To address this issue, We propose CLARA (Cross-Lingual Reflective Agentic system), a novel inference time de-biasing method leveraging retrieval augmented generation and self-verification. Our approach not only improves performance across all languages but also significantly reduces the multilingual bias gap, facilitating equitable LLM application across the globe. David S. Restrepo, Chenwei Wu 0006, Zhengxu Tang, Zitao Shuai, Thao Nguyen Minh Phan, Jun-En Ding, Cong-Tinh Dao, Jack Gallifant, Robyn Gayle Dychiao, Jose Carlo Artiaga, André Hiroshi Bando, Carolina Pelegrini Barbosa Gracitelli, Vincenz Ferrer, Leo A. Celi, Danielle S. Bitterman, Michael G. Morley, Luis Filipe Nakayama |
AAAI | 15 |
| 2025 | Do They Really Know? Evaluating Large Language Models' Ability to Reference and Cite Oncology Guidelines
Pietro Belligoli, Danielle S. Bitterman, Timothy A. Miller |
AIME (2) | 2 |
| 2025 | Sparse Autoencoder Features for Classifications and TransferabilityabstractSparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems.We systematically analyze SAE for interpretable feature extraction from LLMs in safety-critical classification tasks 1 .Our framework evaluates (1) model-layer selection and scaling properties, (2) SAE architectural configurations, including width and pooling strategies, and (3) the effect of binarizing continuous SAE activations.SAE-derived features achieve macro F1 > 0.8, outperforming hidden-state and BoW baselines while demonstrating cross-model transfer from Gemma 2 2B to 9B-IT models.These features generalize in a zero-shot manner to cross-lingual toxicity detection and visual classification tasks.Our analysis highlights the significant impact of pooling strategies and binarization thresholds, showing that binarization offers an efficient alternative to traditional feature selection while maintaining or improving performance.These findings establish new best practices for SAE-based interpretability and enable scalable, transparent deployment of LLMs in real-world applications. Jack Gallifant, Shan Chen 0004, Kuleen Sasse, Hugo J. W. L. Aerts, Thomas Hartvigsen, Danielle S. Bitterman |
EMNLP | 6 |
| 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language ModelsabstractCharacterizing a large language model's (LLM's) knowledge of a given question is challenging.
As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context.
However, this does not fully reflect how well the model knows the answer to the question.
In this paper, we first introduce a taxonomy of five knowledge statuses based on the consistency and correctness of LLM knowledge modes.
We then propose KScope, a hierarchical framework of statistical tests that progressively refines hypotheses about knowledge modes and characterizes LLM knowledge into one of these five statuses.
We apply KScope to nine LLMs across four datasets and systematically establish:
(1) Supporting context narrows knowledge gaps across models.
(2) Context features related to difficulty, relevance, and familiarity drive successful knowledge updates.
(3) LLMs exhibit similar feature preferences when partially correct or conflicted, but diverge sharply when consistently wrong.
(4) Context summarization constrained by our feature analysis, together with enhanced credibility, further improves update effectiveness and generalizes across LLMs. Yuxin Xiao, Shan Chen 0004, Jack Gallifant, Danielle S. Bitterman, Thomas Hartvigsen, Marzyeh Ghassemi |
NeurIPS | 4 |
| 2025 | Collaborative large language models for automated data extraction in living systematic reviewsabstractOBJECTIVE: Data extraction from the published literature is the most laborious step in conducting living systematic reviews (LSRs). We aim to build a generalizable, automated data extraction workflow leveraging large language models (LLMs) that mimics the real-world 2-reviewer process. MATERIALS AND METHODS: A dataset of 10 trials (22 publications) from a published LSR was used, focusing on 23 variables related to trial, population, and outcomes data. The dataset was split into prompt development (n = 5) and held-out test sets (n = 17). GPT-4-turbo and Claude-3-Opus were used for data extraction. Responses from the 2 LLMs were considered concordant if they were the same for a given variable. The discordant responses from each LLM were provided to the other LLM for cross-critique. Accuracy, ie, the total number of correct responses divided by the total number of responses, was computed to assess performance. RESULTS: In the prompt development set, 110 (96%) responses were concordant, achieving an accuracy of 0.99 against the gold standard. In the test set, 342 (87%) responses were concordant. The accuracy of the concordant responses was 0.94. The accuracy of the discordant responses was 0.41 for GPT-4-turbo and 0.50 for Claude-3-Opus. Of the 49 discordant responses, 25 (51%) became concordant after cross-critique, increasing accuracy to 0.76. DISCUSSION: Concordant responses by the LLMs are likely to be accurate. In instances of discordant responses, cross-critique can further increase the accuracy. CONCLUSION: Large language models, when simulated in a collaborative, 2-reviewer workflow, can extract data with reasonable performance, enabling truly "living" systematic reviews. Umair Ayub, Syed Arsalan Ahmed Naqvi, Kaneez Zahra Rubab Khakwani, Zaryab bin Riaz Sipra, Ammad Raina, Sihan Zhou, Amir Saeidi, Bashar Hasan, Robert Bryan Rumble, Danielle S. Bitterman, Jeremy L. Warner, Jia Zou 0001, Amye J. Tevaarwerk, Konstantinos Leventakos, Kenneth L. Kehl, Jeanne M. Palmer, Mohammad Hassan Murad, Chitta Baral, Irbaz Bin Riaz |
J. Am. Medical Informatics Assoc. | 12 |
| 2025 | LCD benchmark: long clinical document benchmark on mortality prediction for language modelsabstractOBJECTIVES: The application of natural language processing (NLP) in the clinical domain is important due to the rich unstructured information in clinical documents, which often remains inaccessible in structured data. When applying NLP methods to a certain domain, the role of benchmark datasets is crucial as benchmark datasets not only guide the selection of best-performing models but also enable the assessment of the reliability of the generated outputs. Despite the recent availability of language models capable of longer context, benchmark datasets targeting long clinical document classification tasks are absent. MATERIALS AND METHODS: To address this issue, we propose Long Clinical Document (LCD) benchmark, a benchmark for the task of predicting 30-day out-of-hospital mortality using discharge notes of Medical Information Mart for Intensive Care IV and statewide death data. We evaluated this benchmark dataset using baseline models, from bag-of-words and convolutional neural network to instruction-tuned large language models. Additionally, we provide a comprehensive analysis of the model outputs, including manual review and visualization of model weights, to offer insights into their predictive capabilities and limitations. RESULTS: Baseline models showed 28.9% for best-performing supervised models and 32.2% for GPT-4 in F1 metrics. Notes in our dataset have a median word count of 1687. DISCUSSION: Our analysis of the model outputs showed that our dataset is challenging for both models and human experts, but the models can find meaningful signals from the text. CONCLUSION: We expect our LCD benchmark to be a resource for the development of advanced supervised models, or prompting methods, tailored for clinical text. Wonjin Yoon, Shan Chen 0004, Yanjun Gao, Zhanzhan Zhao, Dmitriy Dligach, Danielle S. Bitterman, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model BiasabstractLarge language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data.In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessing biases and real world knowledge in LLMs, specifically focusing on the representation of disease prevalence across diverse demographic groups.We systematically evaluate how demographic biases embedded in pre-training corpora like $ThePile$ influence the outputs of LLMs.We expose and quantify discrepancies by juxtaposing these biases against actual disease prevalences in various U.S. demographic groups.Our results highlight substantial misalignment between LLM representation of disease prevalence and real disease prevalence rates across demographic subgroups, indicating a pronounced risk of bias propagation and a lack of real-world grounding for medical applications of LLMs.Furthermore, we observe that various alignment methods minimally resolve inconsistencies in the models' representation of disease prevalence across different languages.For further exploration and analysis, we make all data and a data visualization tool available at: \url{www.crosscare.net}. Shan Chen 0004, Jack Gallifant, Mingye Gao, Nikolaj Munch, Ajay Muthukkumar, Arvind Rajan, Jaya Kolluri, Amelia Fiske, Janna Hastings, Hugo J. W. L. Aerts, Brian Anthony 0001, Leo A. Celi, William G. La Cava, Danielle S. Bitterman |
NeurIPS | 15 |
| 2024 | Evaluating the ChatGPT family of models for biomedical reasoning and classificationabstractOBJECTIVE: Large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates ChatGPT family of models (GPT-3.5, GPT-4) in biomedical tasks beyond question-answering. MATERIALS AND METHODS: We evaluated model performance with 11 122 samples for two fundamental tasks in the biomedical domain-classification (n = 8676) and reasoning (n = 2446). The first task involves classifying health advice in scientific literature, while the second task is detecting causal relations in biomedical literature. We used 20% of the dataset for prompt development, including zero- and few-shot settings with and without chain-of-thought (CoT). We then evaluated the best prompts from each setting on the remaining dataset, comparing them to models using simple features (BoW with logistic regression) and fine-tuned BioBERT models. RESULTS: Fine-tuning BioBERT produced the best classification (F1: 0.800-0.902) and reasoning (F1: 0.851) results. Among LLM approaches, few-shot CoT achieved the best classification (F1: 0.671-0.770) and reasoning (F1: 0.682) results, comparable to the BoW model (F1: 0.602-0.753 and 0.675 for classification and reasoning, respectively). It took 78 h to obtain the best LLM results, compared to 0.078 and 0.008 h for the top-performing BioBERT and BoW models, respectively. DISCUSSION: The simple BoW model performed similarly to the most complex LLM prompting. Prompt engineering required significant investment. CONCLUSION: Despite the excitement around viral ChatGPT, fine-tuning for two fundamental biomedical natural language processing tasks remained the best strategy. Shan Chen 0004, Yingya Li, Hoang Van, Hugo J. W. L. Aerts, Guergana K. Savova, Danielle S. Bitterman |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | Classifying unstructured electronic consult messages to understand primary care physician specialty information needsabstractOBJECTIVE: Electronic consultation (eConsult) content reflects important information about referring clinician needs across an organization, but is challenging to extract. The objective of this work was to develop machine learning models for classifying eConsult questions for question type and question content. Another objective of this work was to investigate the ability to solve this task with constrained expert time resources. MATERIALS AND METHODS: Our data source is the San Francisco Health Network eConsult system, with over 700 000 deidentified questions from the years 2008-2017, from gastroenterology, urology, and neurology specialties. We develop classifiers based on Bidirectional Encoder Representations from Transformers, experimenting with multitask learning to learn when information can be shared across classifiers. We produce learning curves to understand when we may be able to reduce the amount of human labeling required. RESULTS: Multitask learning shows benefits only in the neurology-urology pair where they shared substantial similarities in the distribution of question types. Continued pretraining of models in new domains is highly effective. In the neurology-urology pair, near-peak performance is achieved with only 10% of the urology training data given all of the neurology data. DISCUSSION: Sharing information across classifier types shows little benefit, whereas sharing classifier components across specialties can help if they are similar in the balance of procedural versus cognitive patient care. CONCLUSION: We can accurately classify eConsult content with enough labeled data, but only in special cases do methods for reducing labeling effort apply. Future work should explore new learning paradigms to further reduce labeling effort. Xiyu Ding, Michael L. Barnett, Ateev Mehrotra, Delphine S. Tuot, Danielle S. Bitterman, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 5 |