VLDB 2026 Research / reviewers in the wild / expert
Timothy A. Miller
dblp:117/6625 · also Tim Miller 0002, Timothy Miller 0002
· DBLP profile ↗
63ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0003-4513-403XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrievalabstract-hop retrieval on large KGs by building on symbolic KG formulations and executing traversal as hardware-efficient operations over decomposed subject, object, and relation representations. To scale to billion-edge graphs, LogosKG integrates degree-aware partitioning, cross-graph routing, and on-demand caching. Experiments show substantial efficiency gains over CPU and GPU baselines without loss of retrieval fidelity. With proven performance in KG retrieval, a downstream two-round KG-LLM interaction demonstrates how LogosKG enables large-scale, evidence-grounded analysis of how KG topology, such as hop distribution and connectivity, shapes the alignment between structured biomedical knowledge and LLM diagnostic reasoning, thereby opening the door for next-generation KG-LLM integration. The source code is publicly available at https://github.com/LARK-NLP-Lab/LogosKG, and an online demo is available at https://lark-nlp-lab-logoskg.hf.space/. He Cheng, Yifu Wu, Saksham Khatwani, Maya Kruse, Dmitriy Dligach, Timothy A. Miller, Majid Afshar, Yanjun Gao |
ACL (1) | 6 |
| 2026 | A Dataset of Psychiatric Hospital Notes with Temporal Information AnnotationsabstractTemporal information extraction is the task of identifying temporal entities in a text and relating them to each other. In medicine, electronic health records (EHRs) contain text that documents the sequence of events during an encounter with a patient, and sometimes the events prior to the encounter (e.g., psychosocial environment and history). Temporality is especially important for the specialty of psychiatry. In this work, we describe the updates to the guidelines that allowed us to create a corpus of temporally-annotated psychiatric discharge summaries and progress notes in English. These updated guidelines were used to create a corpus of over 18,000 events, 2,200 time expressions, and 13,000 temporal relations. Temporal information extraction performance with a baseline system trained on non-psychiatric data obtains an F1 score of 0.152 on relation extraction, indicating the importance of this new dataset for making progress on temporal information extraction in the psychiatric domain. Timothy A. Miller, Gaby Dinh, David Harris 0004, Wonjin Yoon, Spencer Thomas, Boyu Ren, Mei-Hua Hall, Guergana K. Savova |
LREC | 1 |
| 2025 | Do They Really Know? Evaluating Large Language Models' Ability to Reference and Cite Oncology Guidelines
Pietro Belligoli, Danielle S. Bitterman, Timothy A. Miller |
AIME (2) | 3 |
| 2025 | Aspect-Oriented Summarization for Psychiatric Short-Term Readmission PredictionabstractRecent progress in large language models (LLMs) has enabled the automated processing of lengthy documents even without supervised training on a task-specific dataset.Yet, their zero-shot performance in complex tasks as opposed to straightforward information extraction tasks remains suboptimal.One feasible approach for tasks with lengthy, complex input is to first summarize the document and then apply supervised fine-tuning to the summary.However, the summarization process inevitably results in some loss of information.In this study we present a method for processing the summaries of long documents aimed to capture different important aspects of the original document.We hypothesize that LLM summaries generated with different aspect-oriented prompts contain different information signals, and we propose methods to measure these differences.We introduce approaches to effectively integrate signals from these different summaries for supervised training of transformer models.We validate our hypotheses on a high-impact task -30-day readmission prediction from a psychiatric discharge -using real-world data from four hospitals, and show that our proposed method increases the prediction performance for the complex task of predicting patient outcome. Wonjin Yoon, Boyu Ren, Spencer Thomas, Chanhwi Kim, Guergana K. Savova, Mei-Hua Hall, Timothy A. Miller |
EMNLP | 7 |
| 2025 | Toward digital twins in the intensive care unit: a medication management case studyabstractOBJECTIVE: To evaluate the efficacy of digital twins developed using a large language model (LLaMA-3), fine-tuned with Low-Rank Adapters (LoRA) on intensive care units (ICU) physician notes, and to determine whether specialty-specific training enhances treatment recommendation accuracy compared to other ICU specialties or zero-shot baselines. MATERIALS AND METHODS: Digital twins were created using LLaMA-3 fine-tuned on discharge summaries from the Medical Information Mart for Intensive Care III dataset, where medications were masked to construct training and testing datasets. The medical ICU dataset (1000 notes) was used for evaluation, and performance was assessed using Bidirectional Encoder Representations from Transformers Score (BERTScore) and ROUGE-L. A zero-shot baseline model, relying solely on contextual instructions without training, was also evaluated. While our approach moves toward digital twin capabilities, it does not incorporate real-time, patient-specific electronic health records data and can be viewed as an ICU specialty-level language model adaptation. RESULTS: Models fine-tuned on medical ICU notes achieved the highest BERTScore (0.842), outperforming models trained on other specialties or mixed datasets. Zero-shot models showed the lowest performance, highlighting the importance of training. DISCUSSION: The findings demonstrate that specialty-specific training significantly improves treatment recommendation accuracy in digital twins compared to generalized or zero-shot approaches. Tailoring models to specific ICU domains strengthens their clinical decision-support capabilities. CONCLUSION: Context-specific fine-tuning of LLMs is crucial for developing effective digital twins, offering foundational insights for personalized clinical decision support. Behnaz Eslami, Majid Afshar, Samie Tootooni, Timothy A. Miller, Matthew M. Churpek, Yanjun Gao, Dmitriy Dligach |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Lessons learned on information retrieval in electronic health records: a comparison of embedding models and pooling strategiesabstractOBJECTIVES: Applying large language models (LLMs) to the clinical domain is challenging due to the context-heavy nature of processing medical records. Retrieval-augmented generation (RAG) offers a solution by facilitating reasoning over large text sources. However, there are many parameters to optimize in just the retrieval system alone. This paper presents an ablation study exploring how different embedding models and pooling methods affect information retrieval for the clinical domain. MATERIALS AND METHODS: Evaluating on 3 retrieval tasks on 2 electronic health record (EHR) data sources, we compared 7 models, including medical- and general-domain models, specialized encoder embedding models, and off-the-shelf decoder LLMs. We also examine the choice of embedding pooling strategy for each model, independently on the query and the text to retrieve. RESULTS: We found that the choice of embedding model significantly impacts retrieval performance, with BGE, a comparatively small general-domain model, consistently outperforming all others, including medical-specific models. However, our findings also revealed substantial variability across datasets and query text phrasings. We also determined the best pooling methods for each of these models to guide future design of retrieval systems. DISCUSSION: The choice of embedding model, pooling strategy, and query formulation can significantly impact retrieval performance and the performance of these models on other public benchmarks does not necessarily transfer to new domains. The high variability in performance across different query phrasings suggests that the choice of query may need to be tuned and validated for each task, or even for each institution's EHR. CONCLUSION: This study provides empirical evidence to guide the selection of models and pooling strategies for RAG frameworks in healthcare applications. Further studies such as this one are vital for guiding empirically-grounded development of retrieval frameworks, such as in the context of RAG, for the clinical domain. Skatje Myers, Timothy A. Miller, Yanjun Gao, Matthew M. Churpek, Anoop M. Mayampurath, Dmitriy Dligach, Majid Afshar |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | LCD benchmark: long clinical document benchmark on mortality prediction for language modelsabstractOBJECTIVES: The application of natural language processing (NLP) in the clinical domain is important due to the rich unstructured information in clinical documents, which often remains inaccessible in structured data. When applying NLP methods to a certain domain, the role of benchmark datasets is crucial as benchmark datasets not only guide the selection of best-performing models but also enable the assessment of the reliability of the generated outputs. Despite the recent availability of language models capable of longer context, benchmark datasets targeting long clinical document classification tasks are absent. MATERIALS AND METHODS: To address this issue, we propose Long Clinical Document (LCD) benchmark, a benchmark for the task of predicting 30-day out-of-hospital mortality using discharge notes of Medical Information Mart for Intensive Care IV and statewide death data. We evaluated this benchmark dataset using baseline models, from bag-of-words and convolutional neural network to instruction-tuned large language models. Additionally, we provide a comprehensive analysis of the model outputs, including manual review and visualization of model weights, to offer insights into their predictive capabilities and limitations. RESULTS: Baseline models showed 28.9% for best-performing supervised models and 32.2% for GPT-4 in F1 metrics. Notes in our dataset have a median word count of 1687. DISCUSSION: Our analysis of the model outputs showed that our dataset is challenging for both models and human experts, but the models can find meaningful signals from the text. CONCLUSION: We expect our LCD benchmark to be a resource for the development of advanced supervised models, or prompting methods, tailored for clinical text. Wonjin Yoon, Shan Chen 0004, Yanjun Gao, Zhanzhan Zhao, Dmitriy Dligach, Danielle S. Bitterman, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 8 |
| 2025 | Identifying task groupings for multi-task learning using pointwise V-usable information
Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Biomed. Informatics | 2 |
| 2024 | Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning modelsabstractOBJECTIVE: The timely stratification of trauma injury severity can enhance the quality of trauma care but it requires intense manual annotation from certified trauma coders. The objective of this study is to develop machine learning models for the stratification of trauma injury severity across various body regions using clinical text and structured electronic health records (EHRs) data. MATERIALS AND METHODS: Our study utilized clinical documents and structured EHR variables linked with the trauma registry data to create 2 machine learning models with different approaches to representing text. The first one fuses concept unique identifiers (CUIs) extracted from free text with structured EHR variables, while the second one integrates free text with structured EHR variables. Temporal validation was undertaken to ensure the models' temporal generalizability. Additionally, analyses to assess the variable importance were conducted. RESULTS: Both models demonstrated impressive performance in categorizing leg injuries, achieving high accuracy with macro-F1 scores of over 0.8. Additionally, they showed considerable accuracy, with macro-F1 scores exceeding or near 0.7, in assessing injuries in the areas of the chest and head. We showed in our variable importance analysis that the most important features in the model have strong face validity in determining clinically relevant trauma injuries. DISCUSSION: The CUI-based model achieves comparable performance, if not higher, compared to the free-text-based model, with reduced complexity. Furthermore, integrating structured EHR data improves performance, particularly when the text modalities are insufficiently indicative. CONCLUSIONS: Our multi-modal, multiclass models can provide accurate stratification of trauma injury severity and clinically relevant interpretations. Jifan Gao, Guanhua Chen 0002, Ann P. O'Rourke, John R. Caskey, Kyle A. Carey, Madeline Oguss, Anne Stey, Dmitriy Dligach, Timothy A. Miller, Anoop M. Mayampurath, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Cumulus: a federated electronic health record-based learning system powered by Fast Healthcare Interoperability Resources and artificial intelligenceabstractOBJECTIVE: To address challenges in large-scale electronic health record (EHR) data exchange, we sought to develop, deploy, and test an open source, cloud-hosted app "listener" that accesses standardized data across the SMART/HL7 Bulk FHIR Access application programming interface (API). METHODS: We advance a model for scalable, federated, data sharing and learning. Cumulus software is designed to address key technology and policy desiderata including local utility, control, and administrative simplicity as well as privacy preservation during robust data sharing, and artificial intelligence (AI) for processing unstructured text. RESULTS: Cumulus relies on containerized, cloud-hosted software, installed within a healthcare organization's security envelope. Cumulus accesses EHR data via the Bulk FHIR interface and streamlines automated processing and sharing. The modular design enables use of the latest AI and natural language processing tools and supports provider autonomy and administrative simplicity. In an initial test, Cumulus was deployed across 5 healthcare systems each partnered with public health. Cumulus output is patient counts which were aggregated into a table stratifying variables of interest to enable population health studies. All code is available open source. A policy stipulating that only aggregate data leave the institution greatly facilitated data sharing agreements. DISCUSSION AND CONCLUSION: Cumulus addresses barriers to data sharing based on (1) federally required support for standard APIs, (2) increasing use of cloud computing, and (3) advances in AI. There is potential for scalability to support learning across myriad network configurations and use cases. Andrew J. McMurry, Daniel Gottlieb 0001, Timothy A. Miller, James R. Jones, Ashish Atreja, Jennifer Crago, Pankaja M. Desai, Brian E. Dixon, Matthew Garber, Vlad Ignatov, Lyndsey A Kirchner, Philip R. O. Payne, Anil J. Saldanha, Prabhu R. V. Shankar, Yauheni Solad, Elizabeth Sprouse, Michael Terry, Adam B. Wilcox, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language ModelsabstractThe bias-variance tradeoff is the idea that learning methods need to balance model complexity with data size to minimize both under-fitting and over-fitting.Recent empirical work and theoretical analyses with over-parameterized neural networks challenge the classic bias-variance trade-off notion suggesting that no such trade-off holds: as the width of the network grows, bias monotonically decreases while variance initially increases followed by a decrease.In this work, we first provide a variance decomposition-based justification criteria to examine whether large pretrained neural models in a fine-tuning setting are generalizable enough to have low bias and variance.We then perform theoretical and empirical analysis using ensemble methods explicitly designed to decrease variance due to optimization.This results in essentially a two-stage fine-tuning algorithm that first ratchets down bias and variance iteratively, and then uses a selected fixed-bias model to further reduce variance due to optimization by ensembling.We also analyze the nature of variance change with the ensemble size in low-and high-resource classes.Empirical results show that this two-stage method obtains strong results on SuperGLUE tasks and clinical information extraction tasks.Code and settings are available: https://github.com/christa60/ bias-var-fine-tuning-plms.git Lijing Wang 0001, Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
ACL (1) | 3 |
| 2023 | Improving model transferability for clinical note section classification models using continued pretrainingabstractOBJECTIVE: The classification of clinical note sections is a critical step before doing more fine-grained natural language processing tasks such as social determinants of health extraction and temporal information extraction. Often, clinical note section classification models that achieve high accuracy for 1 institution experience a large drop of accuracy when transferred to another institution. The objective of this study is to develop methods that classify clinical note sections under the SOAP ("Subjective," "Object," "Assessment," and "Plan") framework with improved transferability. MATERIALS AND METHODS: We trained the baseline models by fine-tuning BERT-based models, and enhanced their transferability with continued pretraining, including domain-adaptive pretraining and task-adaptive pretraining. We added in-domain annotated samples during fine-tuning and observed model performance over a varying number of annotated sample size. Finally, we quantified the impact of continued pretraining in equivalence of the number of in-domain annotated samples added. RESULTS: We found continued pretraining improved models only when combined with in-domain annotated samples, improving the F1 score from 0.756 to 0.808, averaged across 3 datasets. This improvement was equivalent to adding 35 in-domain annotated samples. DISCUSSION: Although considered a straightforward task when performing in-domain, section classification is still a considerably difficult task when performing cross-domain, even using highly sophisticated neural network-based methods. CONCLUSION: Continued pretraining improved model transferability for cross-domain clinical note section classification in the presence of a small amount of in-domain labeled samples. Weipeng Zhou, Meliha Yetisgen, Majid Afshar, Yanjun Gao, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, John R. Caskey, Brihat Sharma, Matthew M. Churpek, Majid Afshar |
J. Biomed. Informatics | 3 |
| 2023 | Progress Note Understanding - Assessment and Plan Reasoning: Overview of the 2022 N2C2 Track 3 shared task
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Matthew M. Churpek, Özlem Uzuner, Majid Afshar |
J. Biomed. Informatics | 3 |
| 2023 | Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001 |
J. Biomed. Informatics | 8 |
| 2022 | Summarizing Patients' Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence ModelsabstractAutomatically summarizing patients’ main problems from daily progress notes using natural language processing methods helps to battle against information and cognitive overload in hospital settings and potentially assists providers with computerized diagnostic decision support. Problem list summarization requires a model to understand, abstract, and generate clinical documentation. In this work, we propose a new NLP task that aims to generate a list of problems in a patient’s daily care plan using input from the provider’s progress notes during hospitalization. We investigate the performance of T5 and BART, two state-of-the-art seq2seq transformer architectures, in solving this problem. We provide a corpus built on top of progress notes from publicly available electronic health record progress notes in the Medical Information Mart for Intensive Care (MIMIC)-III. T5 and BART are trained on general domain text, and we experiment with a data augmentation method and a domain adaptation pre-training method to increase exposure to medical vocabulary and knowledge. Evaluation methods include ROUGE, BERTScore, cosine similarity on sentence embedding, and F-score on medical concepts. Results show that T5 with domain adaptive pre-training achieves significant performance gains compared to a rule-based system and general domain pre-trained language models, indicating a promising direction for tackling the problem summarization task. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Dongfang Xu, Matthew M. Churpek, Majid Afshar |
COLING | 3 |
| 2022 | Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note UnderstandingabstractApplying methods in natural language processing on electronic health records (EHR) data has attracted rising interests. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there are a paucity of annotated corpus built to model clinical diagnostic thinking, a processing involving text understanding, domain knowledge abstraction and reasoning. In this work, we introduce a hierarchical annotation schema with three stages to address clinical text understanding, clinical reasoning and summarization. We create an annotated corpus based on a large collection of publicly available daily progress notes, a type of EHR that is time-sensitive, problem-oriented, and well-documented by the format of Subjective, Objective, Assessment and Plan (SOAP). We also define a new suite of tasks, Progress Note Understanding, with three tasks utilizing the three annotation stages. This new suite aims at training and evaluating future NLP models for clinical text understanding, clinical knowledge representation, inference and summarization. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Samuel Tesch, Ryan Laffin, Matthew M. Churpek, Majid Afshar |
LREC | 3 |
| 2022 | Classifying unstructured electronic consult messages to understand primary care physician specialty information needsabstractOBJECTIVE: Electronic consultation (eConsult) content reflects important information about referring clinician needs across an organization, but is challenging to extract. The objective of this work was to develop machine learning models for classifying eConsult questions for question type and question content. Another objective of this work was to investigate the ability to solve this task with constrained expert time resources. MATERIALS AND METHODS: Our data source is the San Francisco Health Network eConsult system, with over 700 000 deidentified questions from the years 2008-2017, from gastroenterology, urology, and neurology specialties. We develop classifiers based on Bidirectional Encoder Representations from Transformers, experimenting with multitask learning to learn when information can be shared across classifiers. We produce learning curves to understand when we may be able to reduce the amount of human labeling required. RESULTS: Multitask learning shows benefits only in the neurology-urology pair where they shared substantial similarities in the distribution of question types. Continued pretraining of models in new domains is highly effective. In the neurology-urology pair, near-peak performance is achieved with only 10% of the urology training data given all of the neurology data. DISCUSSION: Sharing information across classifier types shows little benefit, whereas sharing classifier components across specialties can help if they are similar in the balance of procedural versus cognitive patient care. CONCLUSION: We can accurately classify eConsult content with enough labeled data, but only in special cases do methods for reducing labeling effort apply. Future work should explore new learning paradigms to further reduce labeling effort. Xiyu Ding, Michael L. Barnett, Ateev Mehrotra, Delphine S. Tuot, Danielle S. Bitterman, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | A scoping review of publicly available language tasks in clinical natural language processingabstractOBJECTIVE: To provide a scoping review of papers on clinical natural language processing (NLP) shared tasks that use publicly available electronic health record data from a cohort of patients. MATERIALS AND METHODS: We searched 6 databases, including biomedical research and computer science literature databases. A round of title/abstract screening and full-text screening were conducted by 2 reviewers. Our method followed the PRISMA-ScR guidelines. RESULTS: A total of 35 papers with 48 clinical NLP tasks met inclusion criteria between 2007 and 2021. We categorized the tasks by the type of NLP problems, including named entity recognition, summarization, and other NLP tasks. Some tasks were introduced as potential clinical decision support applications, such as substance abuse detection, and phenotyping. We summarized the tasks by publication venue and dataset type. DISCUSSION: The breadth of clinical NLP tasks continues to grow as the field of NLP evolves with advancements in language systems. However, gaps exist with divergent interests between the general domain NLP community and the clinical informatics community for task motivation and design, and in generalizability of the data sources. We also identified issues in data preparation. CONCLUSION: The existing clinical NLP tasks cover a wide range of topics and the field is expected to grow and attract more attention from both general domain NLP and clinical informatics community. We encourage future work to incorporate multidisciplinary collaboration, reporting transparency, and standardization in data preparation. We provide a listing of all the shared task papers and datasets from this review in a GitLab repository. Yanjun Gao, Dmitriy Dligach, Leslie Christensen, Samuel Tesch, Ryan Laffin, Dongfang Xu, Timothy A. Miller, Özlem Uzuner, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | Confederated learning in healthcare: Training machine learning models using disconnected data separated by individual, data type and identity for Large-Scale health system Intelligence
Dianbo Liu, Kathe P. Fox, Griffin M. Weber, Timothy A. Miller |
J. Biomed. Informatics | 4 |
| 2022 | A simple neural vector space model for medical concept normalization using concept embeddings
Dongfang Xu, Timothy A. Miller |
J. Biomed. Informatics | 2 |
| 2021 | The SMART Cumulus Text-to-FHIR NLP Pipeline
Timothy A. Miller, Bin Mao, Daniel Gottlieb 0001, Kenneth D. Mandl |
AMIA | 1 |
| 2021 | Depth-Bounded Statistical PCFG Induction as a Model of Human Grammar AcquisitionabstractAbstract This article describes a simple PCFG induction model with a fixed category domain that predicts a large majority of attested constituent boundaries, and predicts labels consistent with nearly half of attested constituent labels on a standard evaluation data set of child-directed speech. The article then explores the idea that the difference between simple grammars exhibited by child learners and fully recursive grammars exhibited by adult learners may be an effect of increasing working memory capacity, where the shallow grammars are constrained images of the recursive grammars. An implementation of these memory bounds as limits on center embedding in a depth-specific transform of a recursive grammar yields a significant improvement over an equivalent but unbounded baseline, suggesting that this arrangement may indeed confer a learning advantage. Lifeng Jin, Lane Schwartz, Finale Doshi-Velez, Timothy A. Miller, William Schuler |
Comput. Linguistics | 4 |
| 2021 | Pre-training phenotyping classifiers
Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
J. Biomed. Informatics | 3 |
| 2021 | Deep representation learning of patient data from Electronic Health Records (EHR): A systematic review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
J. Biomed. Informatics | 5 |
| 2020 | Understanding eConsult Triage Behavior with Natural Language Processing
Xiyu Ding, Michael L. Barnett, Ateev Mehrotra, Timothy A. Miller |
AMIA | 4 |
| 2020 | Learning Patient Representations from Electronic Health Records across Diverse Data Types
Timothy A. Miller, Dmitriy Dligach, Fei Wang 0001, Yuqi Si, Funda Meric-Bernstam |
AMIA | 1 |
| 2020 | Deep Representation Learning of Patient Data from Electronic Health Records: A Systematic Review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
AMIA | 5 |
| 2020 | Learning Hierarchical Transformer-based Representations of Clinical Notes
Xin Su 0008, Timothy A. Miller, Majid Afshar, Dmitriy Dligach |
AMIA | 2 |
| 2020 | Does BERT need domain adaptation for clinical negation detection?abstractINTRODUCTION: Classifying whether concepts in an unstructured clinical text are negated is an important unsolved task. New domain adaptation and transfer learning methods can potentially address this issue. OBJECTIVE: We examine neural unsupervised domain adaptation methods, introducing a novel combination of domain adaptation with transformer-based transfer learning methods to improve negation detection. We also want to better understand the interaction between the widely used bidirectional encoder representations from transformers (BERT) system and domain adaptation methods. MATERIALS AND METHODS: We use 4 clinical text datasets that are annotated with negation status. We evaluate a neural unsupervised domain adaptation algorithm and BERT, a transformer-based model that is pretrained on massive general text datasets. We develop an extension to BERT that uses domain adversarial training, a neural domain adaptation method that adds an objective to the negation task, that the classifier should not be able to distinguish between instances from 2 different domains. RESULTS: The domain adaptation methods we describe show positive results, but, on average, the best performance is obtained by plain BERT (without the extension). We provide evidence that the gains from BERT are likely not additive with the gains from domain adaptation. DISCUSSION: Our results suggest that, at least for the task of clinical negation detection, BERT subsumes domain adaptation, implying that BERT is already learning very general representations of negation phenomena such that fine-tuning even on a specific corpus does not lead to much overfitting. CONCLUSION: Despite being trained on nonclinical text, the large training sets of models like BERT lead to large gains in performance for the clinical negation detection task. Chen Lin 0002, Steven Bethard, Dmitriy Dligach, Farig Sadeque, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Unsupervised Learning of PCFGs with Normalizing FlowabstractUnsupervised PCFG inducers hypothesize sets of compact context-free rules as explanations for sentences.These models not only provide tools for low-resource languages, but also play an important role in modeling language acquisition (Bannard et al., 2009;Abend et al., 2017).However, current PCFG induction models, using word tokens as input, are unable to incorporate semantics and morphology into induction, and may encounter issues of sparse vocabulary when facing morphologically rich languages.This paper describes a neural PCFG inducer which employs context embeddings (Peters et al., 2018) in a normalizing flow model (Dinh et al., 2015) to extend PCFG induction to use semantic and morphological information 1 .Linguistically motivated similarity penalty and categorical distance constraints are imposed on the inducer as regularization.Experiments show that the PCFG induction model with normalizing flow produces grammars with state-of-the-art accuracy on a variety of different languages.Ablation further shows a positive effect of normalizing flow, context embeddings and proposed regularizers. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, Lane Schwartz, William Schuler |
ACL (1) | 3 |
| 2019 | Towards a Universal Document-Level Clinical Text Encoder: Methods for Neural Network Pre-training with Applications to Substance Misuse
Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
AMIA | 3 |
| 2019 | Toward a clinical text encoder: pretraining for clinical natural language processing with applications to substance misuseabstractOBJECTIVE: Our objective is to develop algorithms for encoding clinical text into representations that can be used for a variety of phenotyping tasks. MATERIALS AND METHODS: Obtaining large datasets to take advantage of highly expressive deep learning methods is difficult in clinical natural language processing (NLP). We address this difficulty by pretraining a clinical text encoder on billing code data, which is typically available in abundance. We explore several neural encoder architectures and deploy the text representations obtained from these encoders in the context of clinical text classification tasks. While our ultimate goal is learning a universal clinical text encoder, we also experiment with training a phenotype-specific encoder. A universal encoder would be more practical, but a phenotype-specific encoder could perform better for a specific task. RESULTS: We successfully train several clinical text encoders, establish a new state-of-the-art on comorbidity data, and observe good performance gains on substance misuse data. DISCUSSION: We find that pretraining using billing codes is a promising research direction. The representations generated by this type of pretraining have universal properties, as they are highly beneficial for many phenotyping tasks. Phenotype-specific pretraining is a viable route for trading the generality of the pretrained encoder for better performance on a specific phenotyping task. CONCLUSIONS: We successfully applied our approach to many phenotyping tasks. We conclude by discussing potential limitations of our approach. Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Supervised methods to extract clinical events from cardiology reports in Italian
Natalia Viani, Timothy A. Miller, Carlo Napolitano, Silvia G. Priori, Guergana K. Savova, Riccardo Bellazzi, Lucia Sacchi |
J. Biomed. Informatics | 2 |
| 2018 | Depth-bounding is effective: Improvements and Evaluation of Unsupervised PCFG InductionabstractThere have been several recent attempts to improve the accuracy of grammar induction systems by bounding the recursive complexity of the induction model (Ponvert et al., 2011;Noji and Johnson, 2016;Shain et al., 2016;Jin et al., 2018).Modern depth-bounded grammar inducers have been shown to be more accurate than early unbounded PCFG inducers, but this technique has never been compared against unbounded induction within the same system, in part because most previous depthbounding models are built around sequence models, the complexity of which grows exponentially with the maximum allowed depth.The present work instead applies depth bounds within a chart-based Bayesian PCFG inducer (Johnson et al., 2007b), where bounding can be switched on and off, and then samples trees with and without bounding.1 Results show that depth-bounding is indeed significantly effective in limiting the search space of the inducer and thereby increasing the accuracy of the resulting parsing model.Moreover, parsing results on English, Chinese and German show that this bounded model with a new inference technique is able to produce parse trees more accurately than or competitively with state-ofthe-art constituency-based grammar induction models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
EMNLP | 3 |
| 2018 | Spotting Spurious Data with Neural NetworksabstractHadi Amiri, Timothy Miller, Guergana Savova. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
NAACL-HLT | 2 |
| 2018 | Unsupervised Grammar Induction with Depth-bounded PCFGabstractThere has been recent interest in applying cognitively- or empirically-motivated bounds on recursion depth to limit the search space of grammar induction models (Ponvert et al., 2011; Noji and Johnson, 2016; Shain et al., 2016). This work extends this depth-bounding approach to probabilistic context-free grammar induction (DB-PCFG), which has a smaller parameter space than hierarchical sequence models, and therefore more fully exploits the space reductions of depth-bounding. Results for this model on grammar acquisition from transcribed child-directed speech and newswire text exceed or are competitive with those of other models when evaluated on parse accuracy. Moreover, grammars acquired from this model demonstrate a consistent use of category labels, something which has not been demonstrated by other acquisition models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | Recurrent Neural Network Architectures for Event Extraction from Italian Medical Reports
Natalia Viani, Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Lucia Sacchi, Guergana K. Savova |
AIME | 2 |
| 2017 | DeepPhe - A Natural Language Processing System for Extracting Cancer Phenotypes from Clinical Records
Guergana K. Savova, Eugene Tseytlin, Sean Finan, Melissa Castine, Timothy A. Miller, Olga Medvedeva, David Harris 0004, Harry Hochheiser, Chen Lin 0002, Girish Chavan, Rebecca S. Jacobson |
AMIA | 5 |
| 2017 | Repeat before Forgetting: Spaced Repetition for Efficient and Effective Training of Neural NetworksabstractWe present a novel approach for training artificial neural networks.Our approach is inspired by broad evidence in psychology that shows human learners can learn efficiently and effectively by increasing intervals of time between subsequent reviews of previously learned materials (spaced repetition).We investigate the analogy between training neural models and findings in psychology about human memory model and develop an efficient and effective algorithm to train neural models.The core part of our algorithm is a cognitively-motivated scheduler according to which training instances and their "reviews" are spaced over time.Our algorithm uses only 34-50% of data per epoch, is 2.9-4.8 times faster than standard training, and outperforms competing state-of-the-art baselines.1 Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
EMNLP | 2 |
| 2017 | Towards generalizable entity-centric clinical coreference resolution
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Chen Lin 0002, Guergana K. Savova |
J. Biomed. Informatics | 1 |
| 2016 | Feature Portability in Cross-domain Clinical Coreference
Timothy A. Miller, Dmitriy Dligach, Chen Lin 0002, Steven Bethard, Guergana K. Savova |
AMIA | 1 |
| 2016 | Memory-Bounded Left-Corner Unsupervised Grammar Induction on Child-Directed InputabstractThis paper presents a new memory-bounded left-corner parsing model for unsupervised raw-text syntax induction, using unsupervised hierarchical hidden Markov models (UHHMM). We deploy this algorithm to shed light on the extent to which human language learners can discover hierarchical syntax through distributional statistics alone, by modeling two widely-accepted features of human language acquisition and sentence processing that have not been simultaneously modeled by any existing grammar induction algorithm: (1) a left-corner parsing strategy and (2) limited working memory capacity. To model realistic input to human language learners, we evaluate our system on a corpus of child-directed speech rather than typical newswire corpora. Results beat or closely match those of three competing systems. Cory Shain, William Bryce, Lifeng Jin, Victoria Krakovna, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
COLING | 6 |
| 2016 | Multilayered temporal modeling for the clinical domainabstractOBJECTIVE: To develop an open-source temporal relation discovery system for the clinical domain. The system is capable of automatically inferring temporal relations between events and time expressions using a multilayered modeling strategy. It can operate at different levels of granularity--from rough temporality expressed as event relations to the document creation time (DCT) to temporal containment to fine-grained classic Allen-style relations. MATERIALS AND METHODS: We evaluated our systems on 2 clinical corpora. One is a subset of the Temporal Histories of Your Medical Events (THYME) corpus, which was used in SemEval 2015 Task 6: Clinical TempEval. The other is the 2012 Informatics for Integrating Biology and the Bedside (i2b2) challenge corpus. We designed multiple supervised machine learning models to compute the DCT relation and within-sentence temporal relations. For the i2b2 data, we also developed models and rule-based methods to recognize cross-sentence temporal relations. We used the official evaluation scripts of both challenges to make our results comparable with results of other participating systems. In addition, we conducted a feature ablation study to find out the contribution of various features to the system's performance. RESULTS: Our system achieved state-of-the-art performance on the Clinical TempEval corpus and was on par with the best systems on the i2b2 2012 corpus. Particularly, on the Clinical TempEval corpus, our system established a new F1 score benchmark, statistically significant as compared to the baseline and the best participating system. CONCLUSION: Presented here is the first open-source clinical temporal relation discovery system. It was built using a multilayered temporal modeling strategy and achieved top performance in 2 major shared tasks. Chen Lin 0002, Dmitriy Dligach, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 3 |
| 2015 | Semi-supervised Learning for Phenotyping Tasks
Dmitriy Dligach, Timothy A. Miller, Guergana K. Savova |
AMIA | 2 |
| 2015 | Robust Sentence Segmentation for Clinical Text
Timothy A. Miller, Sean Finan, Dmitriy Dligach, Guergana K. Savova |
AMIA | 1 |
| 2015 | Automatic identification of methotrexate-induced liver toxicity in patients with rheumatoid arthritis from the electronic medical recordabstractOBJECTIVES: To improve the accuracy of mining structured and unstructured components of the electronic medical record (EMR) by adding temporal features to automatically identify patients with rheumatoid arthritis (RA) with methotrexate-induced liver transaminase abnormalities. MATERIALS AND METHODS: Codified information and a string-matching algorithm were applied to a RA cohort of 5903 patients from Partners HealthCare to select 1130 patients with potential liver toxicity. Supervised machine learning was applied as our key method. For features, Apache clinical Text Analysis and Knowledge Extraction System (cTAKES) was used to extract standard vocabulary from relevant sections of the unstructured clinical narrative. Temporal features were further extracted to assess the temporal relevance of event mentions with regard to the date of transaminase abnormality. All features were encapsulated in a 3-month-long episode for classification. Results were summarized at patient level in a training set (N=480 patients) and evaluated against a test set (N=120 patients). RESULTS: The system achieved positive predictive value (PPV) 0.756, sensitivity 0.919, F1 score 0.829 on the test set, which was significantly better than the best baseline system (PPV 0.590, sensitivity 0.703, F1 score 0.642). Our innovations, which included framing the phenotype problem as an episode-level classification task, and adding temporal information, all proved highly effective. CONCLUSIONS: Automated methotrexate-induced liver toxicity phenotype discovery for patients with RA based on structured and unstructured information in the EMR shows accurate results. Our work demonstrates that adding temporal features significantly improved classification results. Chen Lin 0002, Elizabeth W. Karlson, Dmitriy Dligach, Monica P. Ramirez, Timothy A. Miller, Huan Mo, Natalie S. Braggs, Andrew Cagan, Vivian S. Gainer, Joshua C. Denny, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | Discovering body site and severity modifiers in clinical textsabstractOBJECTIVE: To research computational methods for discovering body site and severity modifiers in clinical texts. METHODS: We cast the task of discovering body site and severity modifiers as a relation extraction problem in the context of a supervised machine learning framework. We utilize rich linguistic features to represent the pairs of relation arguments and delegate the decision about the nature of the relationship between them to a support vector machine model. We evaluate our models using two corpora that annotate body site and severity modifiers. We also compare the model performance to a number of rule-based baselines. We conduct cross-domain portability experiments. In addition, we carry out feature ablation experiments to determine the contribution of various feature groups. Finally, we perform error analysis and report the sources of errors. RESULTS: The performance of our method for discovering body site modifiers achieves F1 of 0.740-0.908 and our method for discovering severity modifiers achieves F1 of 0.905-0.929. DISCUSSION: Results indicate that both methods perform well on both in-domain and out-domain data, approaching the performance of human annotators. The most salient features are token and named entity features, although syntactic dependency features also contribute to the overall performance. The dominant sources of errors are infrequent patterns in the data and inability of the system to discern deeper semantic structures. CONCLUSIONS: We investigated computational methods for discovering body site and severity modifiers in clinical texts. Our best system is released open source as part of the clinical Text Analysis and Knowledge Extraction System (cTAKES). Dmitriy Dligach, Steven Bethard, Lee Becker, Timothy A. Miller, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Temporal Annotation in the Clinical DomainabstractThis article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, "the THYME Guidelines to ISO-TimeML (THYME-TimeML)". To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task. William F. Styler IV, Steven Bethard, Sean Finan, Martha Palmer, Sameer Pradhan, Piet C. de Groen, Bradley James Erickson, Timothy A. Miller, Chen Lin 0002, Guergana K. Savova, James Pustejovsky |
Trans. Assoc. Comput. Linguistics | 8 |
| 2013 | Discovering Body Site and Severity Modifiers in Clinical Texts
Dmitriy Dligach, Timothy A. Miller, Guergana K. Savova |
AMIA | 2 |
| 2013 | Automatic Prediction of Rheumatoid Arthritis Disease Activity from the Electronic Medical Records
Chen Lin 0002, Elizabeth W. Karlson, Helena Canhão, Timothy A. Miller, Dmitriy Dligach, Pei J. Chen, Raúl N. Pérez, Michael E. Weinblatt, Nancy A. Shadick, Robert M. Plenge, Guergana K. Savova |
AMIA | 4 |
| 2013 | Discovering Time Expressions in Clinical Text
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Sameer Pradhan, Chen Lin 0002, Guergana K. Savova |
AMIA | 1 |
| 2013 | Negation's Not Solved: Reconsidering Negation Annotation and Evaluation
Stephen T. Wu, Timothy A. Miller, James J. Masanz, Matthew Coarr, David Carrell, Scott R. Halgrim, David Harris 0004, Cheryl Clark |
AMIA | 2 |
| 2012 | A system for coreference resolution for the clinical narrativeabstractOBJECTIVE: To research computational methods for coreference resolution in the clinical narrative and build a system implementing the best methods. METHODS: The Ontology Development and Information Extraction corpus annotated for coreference relations consists of 7214 coreferential markables, forming 5992 pairs and 1304 chains. We trained classifiers with semantic, syntactic, and surface features pruned by feature selection. For the three system components--for the resolution of relative pronouns, personal pronouns, and noun phrases--we experimented with support vector machines with linear and radial basis function (RBF) kernels, decision trees, and perceptrons. Evaluation of algorithms and varied feature sets was performed using standard metrics. RESULTS: The best performing combination is support vector machines with an RBF kernel and all features (MUC score=0.352, B(3)=0.690, CEAF=0.486, BLANC=0.596) outperforming a traditional decision tree baseline. DISCUSSION: The application showed good performance similar to performance on general English text. The main error source was sentence distances exceeding a window of 10 sentences between markables. A possible solution to this problem is hinted at by the fact that coreferent markables sometimes occurred in predictable (although distant) note sections. Another system limitation is failure to fully utilize synonymy and ontological knowledge. Future work will investigate additional ways to incorporate syntactic features into the coreference problem. CONCLUSION: We investigated computational methods for coreference resolution in the clinical narrative. The best methods are released as modules of the open source Clinical Text Analysis and Knowledge Extraction System and Ontology Development and Information Extraction platforms. Jiaping Zheng, Wendy W. Chapman, Timothy A. Miller, Chen Lin 0002, Rebecca S. Jacobson, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | A Pronoun Anaphora Resolution System based on Factorial Hidden Markov Models
Dingcheng Li, Timothy A. Miller, William Schuler |
ACL | 2 |
| 2010 | Broad-Coverage Parsing Using Human-Like Memory ConstraintsabstractHuman syntactic processing shows many signs of taking place within a general-purpose short-term memory. But this kind of memory is known to have a severely constrained storage capacity—possibly constrained to as few as three or four distinct elements. This article describes a model of syntactic processing that operates successfully within these severe constraints, by recognizing constituents in a right-corner transformed representation (a variant of left-corner parsing) and mapping this representation to random variables in a Hierarchic Hidden Markov Model, a factored time-series model which probabilistically models the contents of a bounded memory store over time. Evaluations of the coverage of this model on a large syntactically annotated corpus of English sentences, and the accuracy of a a bounded-memory parsing strategy based on this model, suggest this model may be cognitively plausible. William Schuler, Samir AbdelRahman, Timothy A. Miller, Lane Schwartz |
Comput. Linguistics | 3 |
| 2009 | Word Buffering Models for Improved Speech Repair Parsing
Timothy A. Miller |
EMNLP | 1 |
| 2009 | Improved Syntactic Models for Parsing Speech with Repairs
Timothy A. Miller |
HLT-NAACL | 1 |
| 2008 | A Syntactic Time-Series Model for Parsing Fluent and Disfluent Speech
Timothy A. Miller, William Schuler |
COLING | 1 |
| 2008 | Toward a Psycholinguistically-Motivated Model of Language Processing
William Schuler, Samir AbdelRahman, Timothy A. Miller, Lane Schwartz |
COLING | 3 |
| 2007 | Elements of a spoken language programming interface for robotsabstractIn many settings, such as home care or mobile environments, demands on users' attention, or users' anticipated level of formal training, or other on-site conditions will make standard keyboard-and monitor-based robot programming interfaces impractical. In such cases, a spoken language interface may be preferable. However, the open-ended task of programming a machine is very different from the sort of closed-vocabulary, data-rich applications (e.g. call routing) for which most speaker-independent spoken language interfaces are designed. This paper will describe some of the challenges of designing a spoken language programming interface for robots, and will present an approach that uses these semantic-level resources as extensively as possible in order to address these challenges. Timothy A. Miller, Andrew Exley, William Schuler |
HRI | 1 |
| 2006 | Dynamic evidence models in a DBN phone recognizerabstractThis paper describes an implementation of a discriminative acoustical model – a Conditional Random Field (CRF) – within a Dynamic Bayes Net (DBN) formulation of a Hierarchic Hidden Markov Model (HHMM) phone recognizer. This CRF-DBN topology accounts for phone transition dynamics in conditional probability distributions over random variables associated with observed evidence, and therefore has less need for hidden variable states corresponding to transitions between phones, leaving more hypothesis space available for modeling higher-level linguistic phenomena such syntax and semantics. The model also has the interesting property that it explicitly represents likely formant trajectories and formant targets of modeled phones in its random variable distributions, making it more linguistically transparent than models based on traditional HMMs with conditionally independent evidence variables. Results on the standard TIMIT phone recognition task show this CRF evidence model, even with a relatively simple first-order feature set, is competitive with standard HMMs and DBN variants using static Gaussian mixture models on MFCC features. William Schuler, Timothy A. Miller, Stephen T. Wu, Andrew Exley |
INTERSPEECH | 2 |
| 2005 | Integrating denotational meaning into a DBN language modelabstractThis paper describes a dynamic Bayes net (DBN) language model which allows recognition decisions to be conditioned on features of entities in some environment, to which hypothesized directives might refer. The accuracy of this model is then evaluated on spoken directives in various domains. 1. William Schuler, Timothy A. Miller |
INTERSPEECH | 2 |