Steven Bedrick

dblp:52/7402 · also Stephen Bedrick · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-0163-9397ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts
abstract
Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant and long-standing challenges (e.g., privacy, anonymization, and data governance) which have led to the development of an increasing number of synthetic datasets. One increasingly important category of clinical dataset is that of clinical dialogues which are especially sensitive and difficult to collect, and as such are commonly synthesized. While such synthetic datasets have been shown to be sufficient in some situations, little theory exists to inform how they may be best used and generalized to new applications. In this paper, we provide an overview of how synthetic datasets are created, evaluated and being used for dialogue related tasks in the medical domain. Additionally, we propose a novel typology for use in classifying types and degrees of data synthesis, to facilitate comparison and evaluation.
Steven Bedrick, A. Seza Dogruöz, Sergiu Nisioi
LREC1
2026 Clinical document metadata extraction: A scoping review
abstract
OBJECTIVES: Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps warranting further investigation. METHODS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles from Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science and external sources that perform clinical document metadata extraction, either primarily as a methodology study, secondarily as a feature for a downstream application, or for analysis. We initially identified and screened 342 articles published between 2011 and 2025, then comprehensively reviewed 77 we deemed relevant to our study. RESULTS: Among the 77 articles included in our full text review, 49 were methodological, 22 used document metadata as features in a downstream application, and 6 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. DISCUSSION AND CONCLUSION: Clinical document metadata extraction research has accelerated over recent years. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows.
Kurt Miller, Qiuhao Lu, William R. Hersh, Kirk Roberts, Steven Bedrick, Andrew Wen
J. Biomed. Informatics5
2025 Dynamic few-shot prompting for clinical note section classification using lightweight, open-source large language models
abstract
OBJECTIVE: Unlocking clinical information embedded in clinical notes has been hindered to a significant degree by domain-specific and context-sensitive language. Identification of note sections and structural document elements has been shown to improve information extraction and dependent downstream clinical natural language processing (NLP) tasks and applications. This study investigates the viability of a dynamic example selection prompting method to section classification using lightweight, open-source large language models (LLMs) as a practical solution for real-world healthcare clinical NLP systems. MATERIALS AND METHODS: We develop a dynamic few-shot prompting approach to classifying sections where section samples are first embedded using a transformer-based model and deposited in a vector store. During inference, the embedded samples with the most similar contextual embeddings to a given input section text are retrieved from the vector store and inserted into the LLM prompt. We evaluate this technique on two datasets comprising two section schemas, including varying levels of context. We compare the performance to baseline zero-shot and randomly selected few-shot scenarios. RESULTS: The dynamic few-shot prompting experiments yielded the highest F1 scores in each of the classification tasks and datasets for all seven of the LLMs included in the evaluation, averaging a macro F1 increase of 39.3% and 21.1% in our primary section classification task over the zero-shot and static few-shot baselines, respectively. DISCUSSION AND CONCLUSION: Our results showcase substantial performance improvements imparted by dynamically selecting examples for few-shot LLM prompting, and further improvement by including section context, demonstrating compelling potential for clinical applications.
Kurt Miller, Steven Bedrick, Qiuhao Lu, Andrew Wen, William R. Hersh, Kirk Roberts
J. Am. Medical Informatics Assoc.2
2023 A Statistical Approach for Quantifying Group Difference in Topic Distributions Using Clinical Discourse Samples
abstract
Topic distribution matrices created by topic models are typically used for document classification or as features in a separate machine learning algorithm.Existing methods for evaluating these topic distributions include metrics such as coherence and perplexity; however, there is a lack of statistically grounded evaluation tools.We present a statistical method for investigating group difference in the documenttopic distribution vectors created by latent Dirichlet allocation (LDA).After transforming the vectors using Aitchison geometry, we use multivariate analysis of variance (MANOVA) to compare sample means and calculate effect size using partial eta-squared.We report the results of validating this method on a subset of the 20Newsgroup corpus.We also apply this method to a corpus of dialogues between Autistic and Typically Developing (TD) children and trained examiners.We found that the topic distributions of Autistic children differed from those of TD children when responding to questions about social difficulties.Furthermore, the examiners' topic distributions differed between the Autistic and TD groups when discussing emotions and social difficulties.These results support the use of topic modeling in studying clinically relevant features of social communication such as topic maintenance.
Grace Lawley, Peter A. Heeman, Jill K. Dolata, Eric Fombonne, Steven Bedrick
SIGDIAL5
2022 Digital tools to help parents screen their child for autism are of low quality
Benjamin Sanders, Steven Bedrick, Luis Andres Rivas Vazquez, Sarabeth Broder-Fingert, Shannon A Brown, Jill K. Dolata, Eric Fombonne, Plyce Fuchu, Katharine E. Zuckerman
AMIA2
2021 Comparing Scribed and Non-scribed Outpatient Progress Notes
Adam Rule, Sarah T. Florig, Steven Bedrick, Vishnu Mohan, Jeffrey Allen Gold, Michelle R. Hribar
AMIA3
2021 Refocusing on Relevance: Personalization in NLG
abstract
Many NLG tasks such as summarization, dialogue response, or open domain question answering focus primarily on a source text in order to generate a target response.This standard approach falls short, however, when a user's intent or context of work is not easily recoverable based solely on that source texta scenario that we argue is more of the rule than the exception.In this work, we argue that NLG systems in general should place a much higher level of emphasis on making use of additional context, and suggest that relevance (as used in Information Retrieval) be thought of as a crucial tool for designing user-oriented text-generating tasks.We further discuss possible harms and hazards around such personalization, and argue that value-sensitive design represents a crucial path forward through these challenges.
Shiran Dudy, Steven Bedrick, Bonnie L. Webber
EMNLP (1)2
2021 Searching for scientific evidence in a pandemic: An overview of TREC-COVID
abstract
We present an overview of the TREC-COVID Challenge, an information retrieval (IR) shared task to evaluate search on scientific literature related to COVID-19. The goals of TREC-COVID include the construction of a pandemic search test collection and the evaluation of IR methods for COVID-19. The challenge was conducted over five rounds from April to July 2020, with participation from 92 unique teams and 556 individual submissions. A total of 50 topics (sets of related queries) were used in the evaluation, starting at 30 topics for Round 1 and adding 5 new topics per round to target emerging topics at that state of the still-emerging pandemic. This paper provides a comprehensive overview of the structure and results of TREC-COVID. Specifically, the paper provides details on the background, task structure, topic structure, corpus, participation, pooling, assessment, judgments, results, top-performing systems, lessons learned, and benchmark datasets.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Biomed. Informatics3
2020 TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19
abstract
TREC-COVID is an information retrieval (IR) shared task initiated to support clinicians and clinical research during the COVID-19 pandemic. IR for pandemics breaks many normal assumptions, which can be seen by examining 9 important basic IR research questions related to pandemic situations. TREC-COVID differs from traditional IR shared task evaluations with special considerations for the expected users, IR modality considerations, topic development, participant requirements, assessment process, relevance criteria, evaluation metrics, iteration process, projected timeline, and the implications of data use as a post-task test collection. This article describes how all these were addressed for the particular requirements of developing IR systems under a pandemic situation. Finally, initial participation numbers are also provided, which demonstrate the tremendous interest the IR community has in this effort.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Am. Medical Informatics Assoc.3
2019 We Need to Talk about Standard Splits
abstract
It is standard practice in speech & language technology to rank systems according to performance on a test set held out for evaluation.However, few researchers apply statistical tests to determine whether differences in performance are likely to arise by chance, and few examine the stability of system ranking across multiple training-testing splits.We conduct replication and reproduction experiments with nine part-of-speech taggers published between 2000 and 2018, each of which reports state-of-the-art performance on a widely-used "standard split".We fail to reliably reproduce some rankings using randomly generated splits.We suggest that randomly generated splits should be used in system comparison.
Kyle Gorman, Steven Bedrick
ACL (1)2
2018 Benchmarking Information Retrieval for Precision Oncology: the TREC Precision Medicine Track
Kirk Roberts, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh, Steven Bedrick, Alexander J. Lazar, Shubham Pant
AMIA5
2018 Automatic analysis of pronunciations for children with speech sound disorders
Shiran Dudy, Steven Bedrick, Meysam Asgari, Alexander Kain
Comput. Speech Lang.2
2017 Evaluation of Clinical Text Segmentation to Facilitate Cohort Retrieval
Tracy Edinger, Dina Demner-Fushman, Aaron M. Cohen, Steven Bedrick, William R. Hersh
AMIA4
2017 Intrainstitutional EHR collections for patient-level information retrieval
abstract
Research in clinical information retrieval has long been stymied by the lack of open resources. However, both clinical information retrieval research innovation and legitimate privacy concerns can be served by the creation of intrainstitutional, fully protected resources. In this article, we provide some principles and tools for information retrieval resource‐building in the unique problem setting of patient‐level information retrieval, following the tradition of the Cranfield paradigm. We further include an analysis of parallel information retrieval resources at Oregon Health & Science University and Mayo Clinic that were built on these principles.
Stephen T. Wu, Sijia Liu 0002, Yanshan Wang, Tamara Timmons, Harsha Uppili, Steven Bedrick, William R. Hersh
J. Assoc. Inf. Sci. Technol.6
2016 Restoring line breaks in Epic-derived clinical notes
Stephen T. Wu, Allison Sliter, Meikun Wang, Tamara Timmons, Steven Bedrick
AMIA5
2016 On Developing Resources for Patient-level Information Retrieval
Stephen T. Wu, Tamara Timmons, Amy Yates, Meikun Wang, Steven Bedrick, William R. Hersh
LREC5
2016 Medical Information Search Workshop (MEDIR)
abstract
No abstract available.
Steven Bedrick, Lorraine Goeuriot, Gareth J. F. Jones, Anastasia Krithara, Henning Müller, Georgios Paliouras
SIGIR1
2012 Automated Detection of At-Risk Arabic Text in Faxed Medical Documents
Steven Bedrick, Homoud Al-Jalahma, Rashid Al-Ali, Shane R. Reti, Henry J. Feldman
AMIA1
2012 Barriers to Retrieving Patient Information from Electronic Health Record Data: Failure Analysis from the TREC Medical Records Track
Tracy Edinger, Aaron M. Cohen, Steven Bedrick, Kyle H. Ambert, William R. Hersh
AMIA3
2009 A Multi-Lingual Web Service for Drug Side-Effect Data
Steven Bedrick, Alejandro Mauro
AMIA1
2008 A Scientific Collaboration Tool Built on the Facebook Platform
Steven Bedrick, Dean F. Sittig
AMIA1