VLDB 2026 Research / reviewers in the wild / expert
Alison Callahan
dblp:46/8463
· DBLP profile ↗
12ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0001-5163-380XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 78% Computational science and engineering · 22% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction following |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics
clinical assessment |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics › clinical text processing
clinical text generation |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics
biomedical natural language processing |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Computational science and engineering
multi-task learning |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
natural language generation metrics · 1.5large language model · 1.5language prompting · 0.6instruction tuning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical RecordsabstractThe ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on realistic text generation tasks for healthcare remains challenging. Existing question answering datasets for electronic health record (EHR) data fail to capture the complexity of information needs and documentation burdens experienced by clinicians. To address these challenges, we introduce MedAlign, a benchmark dataset of 983 natural language instructions for EHR data. MedAlign is curated by 15 clinicians (7 specialities), includes clinician-written reference responses for 303 instructions, and provides 276 longitudinal EHRs for grounding instruction-response pairs. We used MedAlign to evaluate 6 general domain LLMs, having clinicians rank the accuracy and quality of each LLM response. We found high error rates, ranging from 35% (GPT-4) to 68% (MPT-7B-Instruct), and 8.3% drop in accuracy moving from 32k to 2k context lengths for GPT-4. Finally, we report correlations between clinician rankings and automated natural language generation metrics as a way to rank LLMs without human review. MedAlign is provided under a research data use agreement to enable LLM evaluations on tasks aligned with clinician needs and preferences. Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Pontes Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak 0002, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay Chaudhari, Nima Aghaeepour, Christopher D. Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma Brunskill, Jason Alan Fries, Nigam H. Shah |
AAAI | 13 |
| 2023 | APLUS: A Python library for usefulness simulations of machine learning models in healthcareabstractDespite the creation of thousands of machine learning (ML) models, the promise of improving patient care with ML remains largely unrealized. Adoption into clinical practice is lagging, in large part due to disconnects between how ML practitioners evaluate models and what is required for their successful integration into care delivery. Models are just one component of care delivery workflows whose constraints determine clinicians' abilities to act on models' outputs. However, methods to evaluate the usefulness of models in the context of their corresponding workflows are currently limited. To bridge this gap we developed APLUS, a reusable framework for quantitatively assessing via simulation the utility gained from integrating a model into a clinical workflow. We describe the APLUS simulation engine and workflow specification language, and apply it to evaluate a novel ML-based screening pathway for detecting peripheral artery disease at Stanford Health Care. Michael Wornow, Elsie Gyang Ross, Alison Callahan, Nigam H. Shah |
J. Biomed. Informatics | 3 |
| 2023 | Principled estimation and evaluation of treatment effect heterogeneity: A case study application to dabigatran for patients with atrial fibrillationabstractOBJECTIVE: To apply the latest guidance for estimating and evaluating heterogeneous treatment effects (HTEs) in an end-to-end case study of the Long-term Anticoagulation Therapy (RE-LY) trial, and summarize the main takeaways from applying state-of-the-art metalearners and novel evaluation metrics in-depth to inform their applications to personalized care in biomedical research. METHODS: Based on the characteristics of the RE-LY data, we selected four metalearners (S-learner with Lasso, X-learner with Lasso, R-learner with random survival forest and Lasso, and causal survival forest) to estimate the HTEs of dabigatran. For the outcomes of (1) stroke or systemic embolism and (2) major bleeding, we compared dabigatran 150 mg, dabigatran 110 mg, and warfarin. We assessed the overestimation of treatment heterogeneity by the metalearners via a global null analysis and their discrimination and calibration ability using two novel metrics: rank-weighted average treatment effects (RATE) and estimated calibration error for treatment heterogeneity. Finally, we visualized the relationships between estimated treatment effects and baseline covariates using partial dependence plots. RESULTS: The RATE metric suggested that either the applied metalearners had poor performance of estimating HTEs or there was no treatment heterogeneity for either the stroke/SE or major bleeding outcome of any treatment comparison. Partial dependence plots revealed that several covariates had consistent relationships with the treatment effects estimated by multiple metalearners. The applied metalearners showed differential performance across outcomes and treatment comparisons, and the X- and R-learners yielded smaller calibration errors than the others. CONCLUSIONS: HTE estimation is difficult, and a principled estimation and evaluation process is necessary to provide reliable evidence and prevent false discoveries. We have demonstrated how to choose appropriate metalearners based on specific data properties, applied them using the off-the-shelf implementation tool survlearners, and evaluated their performance using recently defined formal metrics. We suggest that clinical implications should be drawn based on the common trends across the applied metalearners. Yizhe Xu, Katelyn K. Bechler, Alison Callahan, Nigam H. Shah |
J. Biomed. Informatics | 3 |
| 2022 | Real-World Data at the Point-of-Care: Exploring Education Needs Across All Stakeholders
Julian Z. Genkins, Dev Dash, Lillian Sung, Saurabh Gombar, Alison Callahan |
AMIA | 5 |
| 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language ProcessingabstractTraining and evaluating language models increasingly requires the construction of meta-datasets -- diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-shot generalization by transforming existing, supervised datasets into a variety of novel instruction tuning tasks, highlighting the benefits of meta-dataset curation. While successful in general-domain text, translating these data-centric approaches to biomedical language modeling remains challenging, as labeled biomedical datasets are significantly underrepresented in popular data hubs. To address this challenge, we introduce BigBio a community library of 126+ biomedical NLP datasets, currently covering 13 task categories and 10+ languages. BigBio facilitates reproducible meta-dataset curation via programmatic access to datasets and their metadata, and is compatible with current platforms for prompt engineering and end-to-end few/zero shot language model evaluation. We discuss our process for task schema harmonization, data auditing, contribution guidelines, and outline two illustrative use cases: zero-shot evaluation of biomedical prompts and large-scale, multi-task learning. BigBio is an ongoing community effort and is available at https://github.com/bigscience-workshop/biomedical Jason Alan Fries, Leon Weber-Genzel, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen H. Bach, Stella Biderman, Mario Sänger, Bo Wang 0044, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller 0002, Jenny Chim, José D. Posada, John M. Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, Benjamin Beilharz |
NeurIPS | 18 |
| 2021 | ACE: the Advanced Cohort Engine for searching longitudinal patient recordsabstractOBJECTIVE: To propose a paradigm for a scalable time-aware clinical data search, and to describe the design, implementation and use of a search engine realizing this paradigm. MATERIALS AND METHODS: The Advanced Cohort Engine (ACE) uses a temporal query language and in-memory datastore of patient objects to provide a fast, scalable, and expressive time-aware search. ACE accepts data in the Observational Medicine Outcomes Partnership Common Data Model, and is configurable to balance performance with compute cost. ACE's temporal query language supports automatic query expansion using clinical knowledge graphs. The ACE API can be used with R, Python, Java, HTTP, and a Web UI. RESULTS: ACE offers an expressive query language for complex temporal search across many clinical data types with multiple output options. ACE enables electronic phenotyping and cohort-building with subsecond response times in searching the data of millions of patients for a variety of use cases. DISCUSSION: ACE enables fast, time-aware search using a patient object-centric datastore, thereby overcoming many technical and design shortcomings of relational algebra-based querying. Integrating electronic phenotype development with cohort-building enables a variety of high-value uses for a learning health system. Tradeoffs include the need to learn a new query language and the technical setup burden. CONCLUSION: ACE is a tool that combines a unique query language for time-aware search of longitudinal patient records with a patient object datastore for rapid electronic phenotyping, cohort extraction, and exploratory data analyses. Alison Callahan, Vladimir Polony, José D. Posada, Juan M. Banda, Saurabh Gombar, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Measure what matters: Counts of hospitalized patients are a better metric for health system capacity planning for a reopeningabstractOBJECTIVE: Responding to the COVID-19 pandemic requires accurate forecasting of health system capacity requirements using readily available inputs. We examined whether testing and hospitalization data could help quantify the anticipated burden on the health system given shelter-in-place (SIP) order. MATERIALS AND METHODS: 16,103 SARS-CoV-2 RT-PCR tests were performed on 15,807 patients at Stanford facilities between March 2 and April 11, 2020. We analyzed the fraction of tested patients that were confirmed positive for COVID-19, the fraction of those needing hospitalization, and the fraction requiring ICU admission over the 40 days between March 2nd and April 11th 2020. RESULTS: We find a marked slowdown in the hospitalization rate within ten days of SIP even as cases continued to rise. We also find a shift towards younger patients in the age distribution of those testing positive for COVID-19 over the four weeks of SIP. The impact of this shift is a divergence between increasing positive case confirmations and slowing new hospitalizations, both of which affects the demand on health systems. CONCLUSION: Without using local hospitalization rates and the age distribution of positive patients, current models are likely to overestimate the resource burden of COVID-19. It is imperative that health systems start using these data to quantify effects of SIP and aid reopening planning. Sehj Kashyap, Saurabh Gombar, Steve Yadlowsky, Alison Callahan, Jason Alan Fries, Benjamin A. Pinsky, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Learning Across a Healthcare Data Network to Improve Model Robustness and Evidence Reliability
Noémie Elhadad, Iñigo Urteaga, Alison Callahan, Jenna Reps, Patrick B. Ryan |
AMIA | 3 |
| 2015 | An evidence-based approach to identify aging-related genes in Caenorhabditis elegansabstractBACKGROUND: Extensive studies have been carried out on Caenorhabditis elegans as a model organism to elucidate mechanisms of aging and the effects of perturbing known aging-related genes on lifespan and behavior. This research has generated large amounts of experimental data that is increasingly difficult to integrate and analyze with existing databases and domain knowledge. To address this challenge, we demonstrate a scalable and effective approach for automatic evidence gathering and evaluation that leverages existing experimental data and literature-curated facts to identify genes involved in aging and lifespan regulation in C. elegans. RESULTS: We developed a semantic knowledge base for aging by integrating data about C. elegans genes from WormBase with data about 2005 human and model organism genes from GenAge and 149 genes from GenDR, and with the Bio2RDF network of linked data for the life sciences. Using HyQue (a Semantic Web tool for hypothesis-based querying and evaluation) to interrogate this knowledge base, we examined 48,231 C. elegans genes for their role in modulating lifespan and aging. HyQue identified 24 novel but well-supported candidate aging-related genes for further experimental validation. CONCLUSIONS: We use semantic technologies to discover candidate aging genes whose effects on lifespan are not yet well understood. Our customized HyQue system, the aging research knowledge base it operates over, and HyQue evaluations of all C. elegans genes are freely available at http://hyque.semanticscience.org . Alison Callahan, Juan Cifuentes, Michel Dumontier |
BMC Bioinform. | 1 |
| 2013 | Bio2RDF Release 2: Improved Coverage, Interoperability and Provenance of Life Science Linked Data
Alison Callahan, Jose Cruz-Toledo, Peter Ansell, Michel Dumontier |
ESWC | 1 |
| 2012 | Evaluating Scientific Hypotheses Using the SPARQL Inferencing Notation
Alison Callahan, Michel Dumontier |
ESWC | 1 |
| 2010 | Contextual cocitation: Augmenting cocitation analysis and its applicationsabstractAbstract In this work, a novel method of cocitation analysis, coined “contextual cocitation analysis,” is introduced and described in comparison to traditional methods of cocitation analysis. Equations for quantifying contextual cocitation strength are introduced and their implications explored using theoretical examples alongside the application of contextual cocitation to a series of BioMed Central publications and their cited resources. Based on this work, the implications of contextual cocitation for understanding the granularity of the relationships created between cited published research and methods for its analysis are discussed. Future applications and improvements of this work, including its extended application to the published research of multiple disciplines, are then presented with rationales for their inclusion. Alison Callahan, Stephen Hockema, Gunther Eysenbach |
J. Assoc. Inf. Sci. Technol. | 1 |