VLDB 2026 Research / reviewers in the wild / expert
Jason Alan Fries
dblp:182/2122 · also Jason A. Fries
· DBLP profile ↗
16ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-9316-5768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Medical and health informatics · 94% Computational science and engineering · 6% | |
| Artificial intelligence
10 papers |
Language models and text generation · 26% Deep learning architectures and training · 25% Learning paradigms · 20% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 100% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics › clinical prediction
time-to-event prediction |
1.6 | 2 | 2025 | Time-to-Event Pretraining for 3D Medical Imaging · ICLR 2025 MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Medical and health informatics
clinical prediction |
1.5 | 2 | 2025 | Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR Data · ICLR 2025 EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
foundation model |
1.0 | 2 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models · NeurIPS 2023 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
0.9 | 1 | 2025 | Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR Data · ICLR 2025 |
Medical and health informatics › clinical prediction
clinical risk prediction |
0.9 | 1 | 2025 | Time-to-Event Pretraining for 3D Medical Imaging · ICLR 2025 |
Machine learning › Learning paradigms
weakly supervised learning |
0.8 | 2 | 2020 | Snorkel: rapid training data creation with weak supervision · VLDB J. 2020 Multi-Resolution Weak Supervision for Sequential Data · NeurIPS 2019 |
Natural language and speech › Language models and text generation
instruction following |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Machine learning › Deep learning architectures and training › foundation model
medical foundation model |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Medical and health informatics
clinical assessment |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics › clinical text processing
clinical text generation |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics
electronic health records |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Machine learning and data management › weak supervision
data programming |
0.7 | 2 | 2020 | Snorkel: rapid training data creation with weak supervision · VLDB J. 2020 Snorkel: Rapid Training Data Creation with Weak Supervision · Proc. VLDB Endow. 2017 |
Machine learning and data management › training data management
training data generation |
0.7 | 2 | 2020 | Snorkel: rapid training data creation with weak supervision · VLDB J. 2020 Snorkel: Rapid Training Data Creation with Weak Supervision · Proc. VLDB Endow. 2017 |
Machine learning and data management
weak supervision |
0.7 | 2 | 2020 | Snorkel: rapid training data creation with weak supervision · VLDB J. 2020 Snorkel: Rapid Training Data Creation with Weak Supervision · Proc. VLDB Endow. 2017 |
Medical and health informatics › clinical prediction
clinical outcome prediction |
0.7 | 1 | 2023 | INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary Embolisms · NeurIPS 2023 |
Medical and health informatics
multimodal clinical data |
0.7 | 1 | 2023 | INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary Embolisms · NeurIPS 2023 |
Machine learning › Learning paradigms
multi-task learning |
0.6 | 1 | 2022 | Multitask Prompted Training Enables Zero-Shot Task Generalization · ICLR 2022 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot task generalization |
0.6 | 1 | 2022 | Multitask Prompted Training Enables Zero-Shot Task Generalization · ICLR 2022 |
Medical and health informatics
biomedical natural language processing |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Computational science and engineering
multi-task learning |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Machine learning › Reinforcement learning
sample efficiency |
0.4 | 1 | 2020 | Snorkel: rapid training data creation with weak supervision · VLDB J. 2020 |
Machine learning and data management › weak supervision
labeling functions |
0.3 | 1 | 2017 | Snorkel: Rapid Training Data Creation with Weak Supervision · Proc. VLDB Endow. 2017 |
Computer vision › 3D vision
medical imaging |
0.3 | 1 | 2025 | Time-to-Event Pretraining for 3D Medical Imaging · ICLR 2025 |
Computer vision › Image recognition and object detection
medical image analysis |
0.2 | 1 | 2023 | INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary Embolisms · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 1.7pre-training · 1.7mamba · 1.7long-context modeling · 1.7survival analysis · 1.5self-supervised pretraining · 1.5natural language generation metrics · 1.5large language model · 1.5labeling functions · 1.4benchmark evaluation · 1.3multimodal fusion · 0.7denoising · 0.4data programming · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Systematic review of foundation models for structured electronic health recordsabstractPURPOSE: Foundation models pretrained on structured electronic health record (EHR) data promise improved predictive performance, sample efficiency and resilience to distribution shifts. However, model design, scale and use remain unclear. Objectives were to characterize foundation models pretrained on structured EHR data; examine temporal trends in model application and scale, architecture and design; and assess the extent to which publications omitted methodological details. METHODS: We searched MEDLINE and Embase (2018-October 2025) for foundation models pretrained on structured EHR data using self-supervised learning and applied to clinical prediction tasks. Study selection and data abstraction were performed in duplicate. Characteristics were summarized and stratified by median publication year. RESULTS: Fifty-three studies were included; publications increased over time. Most datasets (79%) originated from the United States. None pretrained exclusively on pediatric cohorts. Model architecture shifted towards transformers (P = .013) with longer context windows (P = .028), while application shifted from exclusively embedding-based toward generative or mixed use (P < .001). Choices regarding feature inclusion, temporal representation, self-supervised objective and downstream adaptation remained heterogeneous. Only 26% of studies evaluated transfer to external datasets, and none described clinical deployment. Key indicators of scale and compute were frequently unreported. CONCLUSIONS: EHR foundation models are proliferating and increasingly transformer-based and generative. Yet methodological choices and reporting remain fragmented, indicating design trade-offs and best practices for EHR foundation models have not yet been established. None describe clinical deployment. Future work should clarify which design choices improve performance, robustness and transferability, increase reporting transparency and identify if they can be implemented to improve patient-important outcomes. Lin Lawrence Guo, Santiago Eduardo Arciniegas, Adam Paul Yan, Jason Alan Fries, George Tomlinson, Lillian Sung |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Time-to-Event Pretraining for 3D Medical ImagingabstractWith the rise of medical foundation models and the growing availability of imaging data, scalable pretraining techniques offer a promising way to identify imaging biomarkers predictive of future disease risk. While current self-supervised methods for 3D medical imaging models capture local structural features like organ morphology, they fail to link pixel biomarkers with long-term health outcomes due to a missing context problem. Current approaches lack the temporal context necessary to identify biomarkers correlated with disease progression, as they rely on supervision derived only from images and concurrent text descriptions. To address this, we introduce time-to-event pretraining, a pretraining framework for 3D medical imaging models that leverages large-scale temporal supervision from paired, longitudinal electronic health records (EHRs). Using a dataset of 18,945 CT scans (4.2 million 2D images) and time-to-event distributions across thousands of EHR-derived tasks, our method improves outcome prediction, achieving an average AUROC increase of 23.7% and a 29.4% gain in Harrell’s C-index across 8 benchmark tasks. Importantly, these gains are achieved without sacrificing diagnostic classification performance. This study lays the foundation for integrating longitudinal EHR and 3D imaging data to advance clinical risk prediction. Zepeng Huo, Jason Alan Fries, Alejandro Lozano, Jeya Maria Jose Valanarasu, Ethan Steinberg, Louis Blankemeier, Akshay Chaudhari, Curt Langlotz, Nigam H. Shah |
ICLR | 2 |
| 2025 | Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR DataabstractFoundation Models (FMs) trained on Electronic Health Records (EHRs) have achieved state-of-the-art results on numerous clinical prediction tasks. However, prior EHR FMs typically have context windows of $<$1k tokens, which prevents them from modeling full patient EHRs which can exceed 10k's of events. For making clinical predictions, both model performance and robustness to the unique properties of EHR data are crucial. Recent advancements in subquadratic long-context architectures (e.g. Mamba) offer a promising solution. However, their application to EHR data has not been well-studied. We address this gap by presenting the first systematic evaluation of the effect of context length on modeling EHR data. We find that longer context models improve predictive performance -- our Mamba-based model surpasses the prior state-of-the-art on 9/14 tasks on the EHRSHOT prediction benchmark. Additionally, we measure robustness to three unique, previously underexplored properties of EHR data: (1) the prevalence of ``copy-forwarded" diagnoses which create artificial token repetition in EHR sequences; (2) the irregular time intervals between EHR events which can lead to a wide range of timespans within a context window; and (3) the natural increase in disease complexity over time which makes later tokens in the EHR harder to predict than earlier ones. Stratifying our EHRSHOT results, we find that higher levels of each property correlate negatively with model performance (e.g., a 14% higher Brier loss between the least and most irregular patients), but that longer context models are more robust to more extreme levels of these properties. Our work highlights the potential for using long-context architectures to model EHR data, and offers a case study on how to identify and quantify new challenges in modeling sequential data motivated by domains outside of natural language. We release all of our model checkpoints and code. Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Ré, Oluwasanmi Koyejo, Nigam H. Shah |
ICLR | 5 |
| 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical RecordsabstractThe ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on realistic text generation tasks for healthcare remains challenging. Existing question answering datasets for electronic health record (EHR) data fail to capture the complexity of information needs and documentation burdens experienced by clinicians. To address these challenges, we introduce MedAlign, a benchmark dataset of 983 natural language instructions for EHR data. MedAlign is curated by 15 clinicians (7 specialities), includes clinician-written reference responses for 303 instructions, and provides 276 longitudinal EHRs for grounding instruction-response pairs. We used MedAlign to evaluate 6 general domain LLMs, having clinicians rank the accuracy and quality of each LLM response. We found high error rates, ranging from 35% (GPT-4) to 68% (MPT-7B-Instruct), and 8.3% drop in accuracy moving from 32k to 2k context lengths for GPT-4. Finally, we report correlations between clinician rankings and automated natural language generation metrics as a way to rank LLMs without human review. MedAlign is provided under a research data use agreement to enable LLM evaluations on tasks aligned with clinician needs and preferences. Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Pontes Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak 0002, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay Chaudhari, Nima Aghaeepour, Christopher D. Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma Brunskill, Jason Alan Fries, Nigam H. Shah |
AAAI | 29 |
| 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical RecordsabstractWe present a self-supervised, time-to-event (TTE) foundation model called MOTOR (Many Outcome Time Oriented Representations) which is pretrained on timestamped sequences of events in electronic health records (EHR) and health insurance claims. TTE models are used for estimating the probability distribution of the time until a specific event occurs, which is an important task in medical settings. TTE models provide many advantages over classification using fixed time horizons, including naturally handling censored observations, but are challenging to train with limited labeled data. MOTOR addresses this challenge by pretraining on up to 55M patient records (9B clinical events). We evaluate MOTOR's transfer learning performance on 19 tasks, across 3 patient databases (a private EHR system, MIMIC-IV, and Merative claims data). Task-specific models adapted from MOTOR improve time-dependent C statistics by 4.6\% over state-of-the-art, improve label efficiency by up to 95\%, and are more robust to temporal distributional shifts. We further evaluate cross-site portability by adapting our MOTOR foundation model for six prediction tasks on the MIMIC-IV dataset, where it outperforms all baselines. MOTOR is the first foundation model for medical TTE predictions and we release a 143M parameter pretrained model for research use at https://huggingface.co/StanfordShahLab/motor-t-base. Ethan Steinberg, Jason Alan Fries, Yizhe Xu, Nigam H. Shah |
ICLR | 2 |
| 2023 | INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary EmbolismsabstractSynthesizing information from various data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly available, multimodal medical datasets. To address this limitation, we introduce INSPECT, which contains de-identified longitudinal records from a large cohort of pulmonary embolism (PE) patients, along with ground truth labels for multiple outcomes. INSPECT contains data from 19,402 patients, including CT images, sections of radiology reports, and structured electronic health record (EHR) data (including demographics, diagnoses, procedures, and vitals). Using our provided dataset, we develop and release a benchmark for evaluating several baseline modeling approaches on a variety of important PE related tasks. We evaluate image-only, EHR-only, and fused models. Trained models and the de-identified dataset are made available for non-commercial use under a data use agreement. To the best our knowledge, INSPECT is the largest multimodal dataset for enabling reproducible research on strategies for integrating 3D medical imaging and EHR data. Shih-Cheng Huang, Zepeng Huo, Ethan Steinberg, Chia-Chun Chiang, Curt Langlotz, Matthew P. Lungren, Serena Yeung-Levy, Nigam H. Shah, Jason Alan Fries |
NeurIPS | 9 |
| 2023 | EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation ModelsabstractWhile the general machine learning (ML) community has benefited from public datasets, tasks, and models, the progress of ML in healthcare has been hampered by a lack of such shared assets. The success of foundation models creates new challenges for healthcare ML by requiring access to shared pretrained models to validate performance benefits. We help address these challenges through three contributions. First, we publish a new dataset, EHRSHOT, which contains de-identified structured data from the electronic health records (EHRs) of 6,739 patients from Stanford Medicine. Unlike MIMIC-III/IV and other popular EHR datasets, EHRSHOT is longitudinal and not restricted to ICU/ED patients. Second, we publish the weights of CLMBR-T-base, a 141M parameter clinical foundation model pretrained on the structured EHR data of 2.57M patients. We are one of the first to fully release such a model for coded EHR data; in contrast, most prior models released for clinical data (e.g. GatorTron, ClinicalBERT) only work with unstructured text and cannot process the rich, structured data within an EHR. We provide an end-to-end pipeline for the community to validate and build upon its performance. Third, we define 15 few-shot clinical prediction tasks, enabling evaluation of foundation models on benefits such as sample efficiency and task adaptation. Our model and dataset are available via a research data use agreement from here: https://stanfordaimi.azurewebsites.net/. Code to reproduce our results is available here: https://github.com/som-shahlab/ehrshot-benchmark. Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Alan Fries, Nigam H. Shah |
NeurIPS | 4 |
| 2023 | Self-supervised machine learning using adult inpatient data produces effective models for pediatric clinical prediction tasksabstractOBJECTIVE: Development of electronic health records (EHR)-based machine learning models for pediatric inpatients is challenged by limited training data. Self-supervised learning using adult data may be a promising approach to creating robust pediatric prediction models. The primary objective was to determine whether a self-supervised model trained in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients, for pediatric inpatient clinical prediction tasks. MATERIALS AND METHODS: This retrospective cohort study used EHR data and included patients with at least one admission to an inpatient unit. One admission per patient was randomly selected. Adult inpatients were 18 years or older while pediatric inpatients were more than 28 days and less than 18 years. Admissions were temporally split into training (January 1, 2008 to December 31, 2019), validation (January 1, 2020 to December 31, 2020), and test (January 1, 2021 to August 1, 2022) sets. Primary comparison was a self-supervised model trained in adult inpatients versus count-based logistic regression models trained in pediatric inpatients. Primary outcome was mean area-under-the-receiver-operating-characteristic-curve (AUROC) for 11 distinct clinical outcomes. Models were evaluated in pediatric inpatients. RESULTS: When evaluated in pediatric inpatients, mean AUROC of self-supervised model trained in adult inpatients (0.902) was noninferior to count-based logistic regression models trained in pediatric inpatients (0.868) (mean difference = 0.034, 95% CI=0.014-0.057; P < .001 for noninferiority and P = .006 for superiority). CONCLUSIONS: Self-supervised learning in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients. This finding suggests transferability of self-supervised models trained in adult patients to pediatric patients, without requiring costly model retraining. Joshua Lemmon, Lin Lawrence Guo, Ethan Steinberg, Keith E. Morse, Scott L. Fleming, Catherine Aftandilian, Stephen Pfohl, José D. Posada, Nigam H. Shah, Jason Alan Fries, Lillian Sung |
J. Am. Medical Informatics Assoc. | 10 |
| 2022 | Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim 0002, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Mike Tian-Jian Jiang, Matteo Manica, Sheng Shen 0001, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf 0008, Alexander M. Rush |
ICLR | 34 |
| 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language ProcessingabstractTraining and evaluating language models increasingly requires the construction of meta-datasets -- diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-shot generalization by transforming existing, supervised datasets into a variety of novel instruction tuning tasks, highlighting the benefits of meta-dataset curation. While successful in general-domain text, translating these data-centric approaches to biomedical language modeling remains challenging, as labeled biomedical datasets are significantly underrepresented in popular data hubs. To address this challenge, we introduce BigBio a community library of 126+ biomedical NLP datasets, currently covering 13 task categories and 10+ languages. BigBio facilitates reproducible meta-dataset curation via programmatic access to datasets and their metadata, and is compatible with current platforms for prompt engineering and end-to-end few/zero shot language model evaluation. We discuss our process for task schema harmonization, data auditing, contribution guidelines, and outline two illustrative use cases: zero-shot evaluation of biomedical prompts and large-scale, multi-task learning. BigBio is an ongoing community effort and is available at https://github.com/bigscience-workshop/biomedical Jason Alan Fries, Leon Weber-Genzel, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen H. Bach, Stella Biderman, Mario Sänger, Bo Wang 0044, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller 0002, Jenny Chim, José D. Posada, John M. Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, Benjamin Beilharz |
NeurIPS | 1 |
| 2021 | Language models are an effective representation learning technique for electronic health record data
Ethan Steinberg, Kenneth Jung, Jason Alan Fries, Conor K. Corbin, Stephen Pfohl, Nigam H. Shah |
J. Biomed. Informatics | 3 |
| 2020 | Measure what matters: Counts of hospitalized patients are a better metric for health system capacity planning for a reopeningabstractOBJECTIVE: Responding to the COVID-19 pandemic requires accurate forecasting of health system capacity requirements using readily available inputs. We examined whether testing and hospitalization data could help quantify the anticipated burden on the health system given shelter-in-place (SIP) order. MATERIALS AND METHODS: 16,103 SARS-CoV-2 RT-PCR tests were performed on 15,807 patients at Stanford facilities between March 2 and April 11, 2020. We analyzed the fraction of tested patients that were confirmed positive for COVID-19, the fraction of those needing hospitalization, and the fraction requiring ICU admission over the 40 days between March 2nd and April 11th 2020. RESULTS: We find a marked slowdown in the hospitalization rate within ten days of SIP even as cases continued to rise. We also find a shift towards younger patients in the age distribution of those testing positive for COVID-19 over the four weeks of SIP. The impact of this shift is a divergence between increasing positive case confirmations and slowing new hospitalizations, both of which affects the demand on health systems. CONCLUSION: Without using local hospitalization rates and the age distribution of positive patients, current models are likely to overestimate the resource burden of COVID-19. It is imperative that health systems start using these data to quantify effects of SIP and aid reopening planning. Sehj Kashyap, Saurabh Gombar, Steve Yadlowsky, Alison Callahan, Jason Alan Fries, Benjamin A. Pinsky, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Snorkel: rapid training data creation with weak supervisionabstractLabeling training data is increasingly the largest bottleneck in deploying machine learning systems. We present Snorkel, a first-of-its-kind system that enables users to train state-of-the-art models without hand labeling any training data. Instead, users write labeling functions that express arbitrary heuristics, which can have unknown accuracies and correlations. Snorkel denoises their outputs without access to ground truth by incorporating the first end-to-end implementation of our recently proposed machine learning paradigm, data programming. We present a flexible interface layer for writing labeling functions based on our experience over the past year collaborating with companies, agencies, and research laboratories. In a user study, subject matter experts build models $$2.8\times $$ faster and increase predictive performance an average $$45.5\%$$ versus seven hours of hand labeling. We study the modeling trade-offs in this new setting and propose an optimizer for automating trade-off decisions that gives up to $$1.8\times $$ speedup per pipeline execution. In two collaborations, with the US Department of Veterans Affairs and the US Food and Drug Administration, and on four open-source text and image data sets representative of other deployments, Snorkel provides $$132\%$$ average improvements to predictive performance over prior heuristic approaches and comes within an average $$3.60\%$$ of the predictive performance of large hand-curated training sets. Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason Alan Fries, Sen Wu 0002, Christopher Ré |
VLDB J. | 4 |
| 2019 | Multi-Resolution Weak Supervision for Sequential DataabstractSince manually labeling training data is slow and expensive, recent industrial and scientific research efforts have turned to weaker or noisier forms of supervision sources. However, existing weak supervision approaches fail to model multi-resolution sources for sequential data, like video, that can assign labels to individual elements or collections of elements in a sequence. A key challenge in weak supervision is estimating the unknown accuracies and correlations of these sources without using labeled data. Multi-resolution sources exacerbate this challenge due to complex correlations and sample complexity that scales in the length of the sequence. We propose Dugong, the first framework to model multi-resolution weak supervision sources with complex correlations to assign probabilistic labels to training data. Theoretically, we prove that Dugong, under mild conditions, can uniquely recover the unobserved accuracy and correlation parameters and use parameter sharing to improve sample complexity. Our method assigns clinician-validated labels to population-scale biomedical video repositories, helping outperform traditional supervision by 36.8 F1 points and addressing a key use case where machine learning has been severely limited by the lack of expert labeled data. On average, Dugong improves over traditional supervision by 16.0 F1 points and existing weak supervision approaches by 24.2 F1 points across several video and sensor classification tasks. Paroma Varma, Frederic Sala, Shiori Sagawa, Jason Alan Fries, Daniel Y. Fu, Saelig Khattar, Ashwini Ramamoorthy, Kayvon Fatahalian, James Priest, Christopher Ré |
NeurIPS | 4 |
| 2017 | Snorkel: A System for Lightweight Extraction
Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason Alan Fries, Sen Wu 0002, Christopher Ré |
CIDR | 4 |
| 2017 | Snorkel: Rapid Training Data Creation with Weak SupervisionabstractLabeling training data is increasingly the largest bottleneck in deploying machine learning systems. We present Snorkel, a first-of-its-kind system that enables users to train state-of- the-art models without hand labeling any training data. Instead, users write labeling functions that express arbitrary heuristics, which can have unknown accuracies and correlations. Snorkel denoises their outputs without access to ground truth by incorporating the first end-to-end implementation of our recently proposed machine learning paradigm, data programming. We present a flexible interface layer for writing labeling functions based on our experience over the past year collaborating with companies, agencies, and research labs. In a user study, subject matter experts build models 2.8× faster and increase predictive performance an average 45.5% versus seven hours of hand labeling. We study the modeling tradeoffs in this new setting and propose an optimizer for automating tradeoff decisions that gives up to 1.8× speedup per pipeline execution. In two collaborations, with the U.S. Department of Veterans Affairs and the U.S. Food and Drug Administration, and on four open-source text and image data sets representative of other deployments, Snorkel provides 132% average improvements to predictive performance over prior heuristic approaches and comes within an average 3.60% of the predictive performance of large hand-curated training sets. Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason Alan Fries, Sen Wu 0002, Christopher Ré |
Proc. VLDB Endow. | 4 |