EDBT 2026 Demo / reviewers in the wild / expert
Trevor Cohen
dblp:10/4030 · also Trevor A. Cohen
· DBLP profile ↗
92ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0003-0159-6697ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 80 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation
Yue Guo 0007, Jae Ho Sohn, Gondy Leroy, Trevor Cohen |
J. Biomed. Informatics | 4 |
| 2025 | Mitigating Confounding in Speech-Based Dementia Detection through Weight MaskingabstractZhecheng Sheng, Xiruo Ding, Brian Hur, Changye Li, Trevor Cohen, Serguei V. S. Pakhomov. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhecheng Sheng, Xiruo Ding, Brian Hur, Changye Li 0001, Trevor Cohen, Serguei V. S. Pakhomov |
ACL (1) | 5 |
| 2025 | Coherence and comprehensibility: Large language models predict lay understanding of health-related content
Trevor Cohen, Weizhe Xu, Yue Guo 0007, Serguei V. S. Pakhomov, Gondy Leroy |
J. Biomed. Informatics | 1 |
| 2025 | Tailoring task arithmetic to address bias in models trained on multi-institutional datasets
Xiruo Ding, Zhecheng Sheng, Brian Hur, Justin Tauscher, Dror Ben-Zeev, Meliha Yetisgen, Serguei V. S. Pakhomov, Trevor Cohen |
J. Biomed. Informatics | 8 |
| 2025 | Perplexity and proximity: Large language model perplexity complements semantic distance metrics for the detection of incoherent speechabstractOBJECTIVE: Semantic coherence in speech is characterized by a logical, connected flow of ideas. A lack of coherence in speech may reflect disorganized thinking, a core feature of psychosis in schizophrenia spectrum disorders (SSDs). Developing tools that could help with automated assessment of semantic coherence in language could facilitate early detection of SSDs and improved monitoring of symptoms, enabling more timely intervention. Large language models (LLMs) have demonstrated strong capabilities on numerous language-centric tasks and have shown promise for analyzing semantic coherence due to the natural fit between their innate measures of language perplexity and the surprising turns that incoherent narrative often takes. This study aims to develop a novel representation and associated measure of semantic coherence using LLM-based perplexity metrics and to compare this measure with traditional vector distance-based coherence metrics. METHOD: We evaluated "bag" and "chain" models based on LLM perplexities as measures of semantic coherence. Regression models were trained using both single and paired combinations of perplexity- and proximity-based features to predict human ratings of semantic coherence using standardized instruments. Performance was evaluated on held-out examples from a training set of speeches from individuals experiencing psychotic symptoms and a test set of clinical interviews with patients diagnosed with SSDs, both with labels from human assessments of disorganized thinking severity. RESULTS: The best performance was achieved using a combination of perplexity and proximity features, yielding a Spearman correlation with human ratings of 0.61 (vs. 0.56 with proximity features alone) on leave-one-out cross-validation in the training set, and 0.54 (vs. 0.52 with proximity features alone) on the test set. CONCLUSION: We developed novel methods for assessing semantic coherence using LLM perplexities and found them complementary to proximity-based methods. Combined, these methods showed improved performance across two datasets, highlighting LLM's potential in enhancing automated diagnosis and monitoring of SSDs. Weizhe Xu, Serguei V. S. Pakhomov, Patrick Heagerty, Eric Horvitz, Ellen Bradley, Joshua Woolley, Andrew T. Campbell, Alex S. Cohen, Dror Ben-Zeev, Trevor Cohen |
J. Biomed. Informatics | 10 |
| 2024 | APPLS: Evaluating Evaluation Metrics for Plain Language SummarizationabstractWhile there has been significant development of models for Plain Language Summarization (PLS), evaluation remains a challenge. PLS lacks a dedicated assessment metric, and the suitability of text generation evaluation metrics is unclear due to the unique transformations involved (e.g., adding background explanations, removing jargon). To address these questions, our study introduces a granular meta-evaluation testbed, APPLS, designed to evaluate metrics for PLS. We identify four PLS criteria from previous work-informativeness, simplification, coherence, and faithfulness-and define a set of perturbations corresponding to these criteria that sensitive metrics should be able to detect. We apply these perturbations to the texts of two PLS datasets to create our testbed. Using APPLS, we assess performance of 14 metrics, including automated scores, lexical features, and LLM prompt-based evaluations. Our analysis reveals that while some current metrics show sensitivity to specific criteria, no single method captures all four criteria simultaneously. We therefore recommend a suite of automated metrics be used to capture PLS quality along all relevant criteria. This work contributes the first meta-evaluation testbed for PLS and a comprehensive evaluation of existing metrics. Yue Guo 0007, Tal August, Gondy Leroy, Trevor Cohen, Lucy Lu Wang |
EMNLP | 4 |
| 2024 | Personalized Jargon Identification for Enhanced Interdisciplinary CommunicationabstractScientific jargon can confuse researchers when they read materials from other domains. Identifying and translating jargon for individual researchers could speed up research, but current methods of jargon identification mainly use corpus-level familiarity indicators rather than modeling researcher-specific needs, which can vary greatly based on each researcher's background. We collect a dataset of over 10K term familiarity annotations from 11 computer science researchers for terms drawn from 100 paper abstracts. Analysis of this data reveals that jargon familiarity and information needs vary widely across annotators, even within the same sub-domain (e.g., NLP). We investigate features representing domain, subdomain, and individual knowledge to predict individual jargon familiarity. We compare supervised and prompt-based approaches, finding that prompt-based methods using information about the individual researcher (e.g., personal publications, self-defined subfield of research) yield the highest accuracy, though the task remains difficult and supervised approaches have lower false positive rates. This research offers insights into features and methods for the novel task of integrating personal data into scientific jargon identification. Yue Guo 0007, Joseph Chee Chang, Maria Antoniak, Erin Bransom, Trevor Cohen, Lucy Lu Wang, Tal August |
NAACL-HLT | 5 |
| 2024 | Large language models in biomedicine and health: current research landscape and future directionsabstractLarge language models in biomedicine and health: current research landscape and future directionsLarge language models (LLMs) are a specialized type of generative artificial intelligence (AI) focused on generating natural language text.These models are developed through extensive training on massive amounts of text data and use deep learning algorithms to generate new text that closely resembles human-generated text.Generative AI methods, including LLMs, are rapidly transforming various domains, including biomedicine and healthcare.[1][2][3][4][5][6] They have already demonstrated remarkable potential as a means to process and analyze large amounts of text, interpret natural language, and generate new content in these domains.For example, Nori et al reported that GPT-4 is able to correctly answer the majority of questions from medical practice licensing exams, comfortably obtaining a passing grade.7 Similarly, Stribling et al found that this model exceeded the average performance of students in the graduate medical sciences on the majority of examinations, including strong performance on short answer and essay questions.8 Even though passing the exam is not the same as applying the knowledge in a real-world setting, these results demonstrate that LLMs can generate appropriate multiple-choice and narrative responses to questions framed in natural language.ChatGPT, first released in November 2022, has garnered phenomenal attention from both the scientific community and a broader society.A keyword search of "large language models" OR "ChatGPT" in PubMed returned over 4500 articles that discuss the technology and its implications for various topics, including medical informatics, by the end of June 2024.In addition, LLM-based technologies have already been deployed in several healthcare systems and are offered as integrated products for use in the clinic within vendor electronic health record systems (for thoughts on initial evaluations of an early product, see Garcia et al 9 and Tai-Seale et al 10 ).This rapid adoption of LLMs like ChatGPT brings an unprecedented opportunity to use this novel AI technology to transform healthcare and medicine.Despite their potential benefits, LLMs can sometimes produce invalid and unsubstantiated responses, a phenomenon known as the "hallucination and confabulation issue" in the literature, or biased responses, due to the biases inherent in their training data.[11][12][13][14][15][16][17] With this great potential also comes the need for trustworthy and responsible development and use of technology.As we continue to explore the capabilities of ChatGPT and other LLMs, it is critical to address related ethical, legal, and social issues to ensure that the technology is used in ways that are safe, fair, trustworthy, and beneficial for all.In the context of biomedicine and healthcare, it is particularly important to engage stakeholders, such as AI researchers, developers of data-driven clinical decision support, care providers, and system implementers from both academic medical centers and industry, to ensure responsible use of LLMs for good.To accelerate research and development in this area, we issued a call for submissions in Summer 2023, specifically focusing on the intersection of biomedicine/health and LLMs, and invited contributions on all related aspects.We invited submissions that report on innovative informatics methods development and evaluation, as well as studies that demonstrate the effectiveness/limitations of LLMs methodologies in healthcare.We particularly encouraged submissions that address the challenges and opportunities of this intersection and offer new insights into how these fields can work together to advance healthcare.This editorial provides an overview of the papers accepted in this Focus Issue.We highlight major themes and unique aspects of the research papers in medical LLMs, discuss ongoing challenges, and recommend future research directions.Box 1 lists the relevant large language model terms and abbreviations used in this editorial. Overall statistics of the Focus IssueThis JAMIA Focus Issue on LLMs in biomedicine and health has drawn enthusiasm from many researchers across different research disciplines.In total, we received over 150 submissions from authors in 25 countries and regions across 6 continents worldwide.The rigorous JAMIA peer review process was applied to all submissions, 41 of which were ultimately accepted for publication in the Focus Issue (Table 1).The majority of the accepted papers were authored by authors in North America, followed by those in Asia and Europe (Figure 1A).The Focus Issue highlights the nature of multi-disciplinary collaboration in medical informatics research across the broad JAMIA community.The number of authors per paper varies from 1 to 23, with an average of 7.3.Many papers feature authors with diverse expertise from different departments and organizations.The authors' expertise spans a wide range of fields, including computer science, data science, informatics, statistics, medicine, nursing, clinical services, public health policies, and more.Several papers also demonstrate scientific collaborations across different sectors, including academia, government labs, research institutes, hospitals, and industry.Additionally, a few papers showcase international collaborations among authors. Zhiyong Lu, Yifan Peng 0002, Trevor Cohen, Marzyeh Ghassemi, Chunhua Weng, Shubo Tian |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Large language models for biomedicine: foundations, opportunities, challenges, and best practicesabstractOBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications. Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Retrieval augmentation of large language models for lay language generation
Yue Guo 0007, Wei Qiu 0006, Gondy Leroy, Trevor Cohen |
J. Biomed. Informatics | 5 |
| 2024 | Useful blunders: Can automated speech recognition errors improve downstream dementia classification?
Changye Li 0001, Weizhe Xu, Trevor Cohen, Serguei V. S. Pakhomov |
J. Biomed. Informatics | 3 |
| 2024 | Opportunities for incorporating intersectionality into biomedical informaticsabstractMany approaches in biomedical informatics (BMI) rely on the ability to define, gather, and manipulate biomedical data to support health through a cyclical research-practice lifecycle. Researchers within this field are often fortunate to work closely with healthcare and public health systems to influence data generation and capture and have access to a vast amount of biomedical data. Many informaticists also have the expertise to engage with stakeholders, develop new methods and applications, and influence policy. However, research and policy that explicitly seeks to address the systemic drivers of health would more effectively support health. Intersectionality is a theoretical framework that can facilitate such research. It holds that individual human experiences reflect larger socio-structural level systems of privilege and oppression, and cannot be truly understood if these systems are examined in isolation. Intersectionality explicitly accounts for the interrelated nature of systems of privilege and oppression, providing a lens through which to examine and challenge inequities. In this paper, we propose intersectionality as an intervention into how we conduct BMI research. We begin by discussing intersectionality's history and core principles as they apply to BMI. We then elaborate on the potential for intersectionality to stimulate BMI research. Specifically, we posit that our efforts in BMI to improve health should address intersectionality's five key considerations: (1) systems of privilege and oppression that shape health; (2) the interrelated nature of upstream health drivers; (3) the nuances of health outcomes within groups; (4) the problematic and power-laden nature of categories that we assign to people in research and in society; and (5) research to inform and support social change. Oliver J. Bear Don't Walk IV, Amandalynne Paullada, Avery R. Everhart, Reggie Casanova-Perez, Trevor Cohen, Tiffany C. Veinot |
J. Biomed. Informatics | 5 |
| 2023 | From benchmark to bedside: transfer learning from social media to patient-provider text messages for suicide risk predictionabstractOBJECTIVE: Compared to natural language processing research investigating suicide risk prediction with social media (SM) data, research utilizing data from clinical settings are scarce. However, the utility of models trained on SM data in text from clinical settings remains unclear. In addition, commonly used performance metrics do not directly translate to operational value in a real-world deployment. The objectives of this study were to evaluate the utility of SM-derived training data for suicide risk prediction in a clinical setting and to develop a metric of the clinical utility of automated triage of patient messages for suicide risk. MATERIALS AND METHODS: Using clinical data, we developed a Bidirectional Encoder Representations from Transformers-based suicide risk detection model to identify messages indicating potential suicide risk. We used both annotated and unlabeled suicide-related SM posts for multi-stage transfer learning, leveraging customized contemporary learning rate schedules. We also developed a novel metric estimating predictive models' potential to reduce follow-up delays with patients in distress and used it to assess model utility. RESULTS: Multi-stage transfer learning from SM data outperformed baseline approaches by traditional classification performance metrics, improving performance from 0.734 to a best F1 score of 0.797. Using this approach for automated triage could reduce response times by 15 minutes per urgent message. DISCUSSION: Despite differences in data characteristics and distribution, publicly available SM data benefit clinical suicide risk prediction when used in conjunction with contemporary transfer learning techniques. Estimates of time saved due to automated triage indicate the potential for the practical impact of such models when deployed as part of established suicide prevention interventions. CONCLUSIONS: This work demonstrates a pathway for leveraging publicly available SM data toward improving risk assessment, paving the way for better clinical care and improved clinical outcomes. Hannah A. Burkhardt, Xiruo Ding, Amanda Kerbrat, Katherine Anne Comtois, Trevor Cohen |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | Discerning conversational context in online health communities for personalized digital behavior change solutions using Pragmatics to Reveal Intent in Social Media (PRISM) frameworkabstractBACKGROUND: Online health communities (OHCs) have emerged as prominent platforms for behavior modification, and the digitization of online peer interactions has afforded researchers with unique opportunities to model multilevel mechanisms that drive behavior change. Existing studies, however, have been limited by a lack of methods that allow the capture of conversational context and socio-behavioral dynamics at scale, as manifested in these digital platforms. OBJECTIVE: We develop, evaluate, and apply a novel methodological framework, Pragmatics to Reveal Intent in Social Media (PRISM), to facilitate granular characterization of peer interactions by combining multidimensional facets of human communication. METHODS: We developed and applied PRISM to analyze peer interactions (N = 2.23 million) in QuitNet, an OHC for tobacco cessation. First, we generated a labeled set of peer interactions (n = 2,005) through manual annotation along three dimensions: communication themes (CTs), behavior change techniques (BCTs), and speech acts (SAs). Second, we used deep learning models to apply our qualitative codes at scale. Third, we applied our validated model to perform a retrospective analysis. Finally, using social network analysis (SNA), we portrayed large-scale patterns and relationships among the aforementioned communication dimensions embedded in peer interactions in QuitNet. RESULTS: Qualitative analysis showed that the themes of social support and behavioral progress were common. The most used BCTs were feedback and monitoring and comparison of behavior, and users most commonly expressed their intentions using SAs-expressive and emotion. With additional in-domain pre-training, bidirectional encoder representations from Transformers (BERT) outperformed other deep learning models on the classification tasks. Content-specific SNA revealed that users' engagement or abstinence status is associated with the prevalence of various categories of BCTs and SAs, which also was evident from the visualization of network structures. CONCLUSIONS: Our study describes the interplay of multilevel characteristics of online communication and their association with individual health behaviors. Tavleen Singh, Kirk Roberts, Trevor Cohen, Nathan K. Cobb, Amy Franklin, Sahiti Myneni |
J. Biomed. Informatics | 3 |
| 2022 | GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language ModelsabstractDeep learning (DL) techniques involving finetuning large numbers of model parameters have delivered impressive performance on the task of discriminating between language produced by cognitively healthy individuals, and those with Alzheimer's disease (AD).However, questions remain about their ability to generalize beyond the small reference sets that are publicly available for research.As an alternative to fitting model parameters directly, we propose a novel method by which a Transformer DL model (GPT-2) pre-trained on general English text is paired with an artificially degraded version of itself (GPT-D), to compute the ratio between these two models' perplexities on language from cognitively healthy and impaired individuals.This technique approaches state-ofthe-art performance on text data from a widely used "Cookie Theft" picture description task, and unlike established alternatives also generalizes well to spontaneous conversations.Furthermore, GPT-D generates text with characteristics known to be associated with AD, demonstrating the induction of dementia-related linguistic anomalies.Our study is a step toward better understanding of the relationships between the inner workings of generative neural language models, the language that they produce, and the deleterious effects of dementia on human speech and language characteristics. Changye Li 0001, David S. Knopman, Weizhe Xu, Trevor Cohen, Serguei V. S. Pakhomov |
ACL (1) | 4 |
| 2022 | Identifying opportunities for informatics-supported suicide prevention: the case of Caring Contacts
Hannah A. Burkhardt, Megan Laine, Amanda Kerbrat, Trevor Cohen, Katherine Anne Comtois, Andrea L. Hartzler |
AMIA | 4 |
| 2022 | Predicting Drug Blood-Brain Barrier Penetration with Adverse Event Report Embeddings
Justin Mower, Xiruo Ding, Oliver Li, Devika Subramanian, Trevor Cohen |
AMIA | 6 |
| 2022 | Fully automated detection of formal thought disorder with Time-series Augmented Representations for Detection of Incoherent Speech (TARDIS)
Weizhe Xu, Weichen Wang 0001, Jake Portanova, Ayesha Chander, Andrew T. Campbell, Serguei V. S. Pakhomov, Dror Ben-Zeev, Trevor Cohen |
J. Biomed. Informatics | 8 |
| 2021 | Automated Lay Language Summarization of Biomedical Scientific ReviewsabstractHealth literacy has emerged as a crucial factor in making appropriate health decisions and ensuring treatment outcomes. However, medical jargon and the complex structure of professional language in this domain make health information especially hard to interpret. Thus, there is an urgent unmet need for automated methods to enhance the accessibility of the biomedical literature to the general population. This problem can be framed as a type of translation problem between the language of healthcare professionals, and that of the general public. In this paper, we introduce the novel task of automated generation of lay language summaries of biomedical scientific reviews, and construct a dataset to support the development and evaluation of automated methods through which to enhance the accessibility of the biomedical literature. We conduct analyses of the various challenges in performing this task, including not only summarization of the key points but also explanation of background knowledge and simplification of professional language. We experiment with state-of-the-art summarization models as well as several data augmentation techniques, and evaluate their performance using both automated metrics and human assessment. Results indicate that automatically generated summaries produced using contemporary neural architectures can achieve promising quality and readability as compared with reference summaries developed for the lay public by experts (best ROUGE-L of 50.24 and Flesch-Kincaid readability score of 13.30). We also discuss the limitations of the current effort, providing insights and directions for future work. Yue Guo 0007, Wei Qiu 0006, Yizhong Wang, Trevor Cohen |
AAAI | 4 |
| 2021 | Linguistic indicators of Behavioral Activation in text-based therapy sessions anticipate changes in depression symptomatology
Hannah A. Burkhardt, George Alexopoulos, Michael D. Pullmann, Thomas D. Hull, Pat A. Areán, Trevor Cohen |
AMIA | 6 |
| 2021 | Quantum Mathematics in Artificial IntelligenceabstractIn the decade since 2010, successes in artificial intelligence have been at the forefront of computer science and technology, and vector space models have solidified a position at the forefront of artificial intelligence. At the same time, quantum computers have become much more powerful, and announcements of major advances are frequently in the news. The mathematical techniques underlying both these areas have more in common than is sometimes realized. Vector spaces took a position at the axiomatic heart of quantum mechanics in the 1930s, and this adoption was a key motivation for the derivation of logic and probability from the linear geometry of vector spaces. Quantum interactions between particles are modelled using the tensor product, which is also used to express objects and operations in artificial neural networks. This paper describes some of these common mathematical areas, including examples of how they are used in artificial intelligence (AI), particularly in automated reasoning and natural language processing (NLP). Techniques discussed include vector spaces, scalar products, subspaces and implication, orthogonal projection and negation, dual vectors, density matrices, positive operators, and tensor products. Application areas include information retrieval, categorization and implication, modelling word-senses and disambiguation, inference in knowledge bases, decision making, and and semantic composition. Some of these approaches can potentially be implemented on quantum hardware. Many of the practical steps in this implementation are in early stages, and some are already realized. Explaining some of the common mathematical tools can help researchers in both AI and quantum computing further exploit these overlaps, recognizing and exploring new directions along the way.This paper describes some of these common mathematical areas, including examples of how they are used in artificial intelligence (AI), particularly in automated reasoning and natural language processing (NLP). Techniques discussed include vector spaces, scalar products, subspaces and implication, orthogonal projection and negation, dual vectors, density matrices, positive operators, and tensor products. Application areas include information retrieval, categorization and implication, modelling word-senses and disambiguation, inference in knowledge bases, and semantic composition. Some of these approaches can potentially be implemented on quantum hardware. Many of the practical steps in this implementation are in early stages, and some are already realized. Explaining some of the common mathematical tools can help researchers in both AI and quantum computing further exploit these overlaps, recognizing and exploring new directions along the way. Dominic Widdows, Kirsty Kitto, Trevor Cohen |
J. Artif. Intell. Res. | 3 |
| 2021 | Augmenting aer2vec: Enriching distributed representations of adverse event report data with orthographic and lexical information
Xiruo Ding, Justin Mower, Devika Subramanian, Trevor Cohen |
J. Biomed. Informatics | 4 |
| 2021 | Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance
Scott A. Malec, Elmer V. Bernstam, Richard D. Boyce, Trevor Cohen |
J. Biomed. Informatics | 5 |
| 2020 | A Tale of Two Perplexities: Sensitivity of Neural Language Models to Lexical Retrieval Deficits in Dementia of the Alzheimer's TypeabstractIn recent years there has been a burgeoning interest in the use of computational methods to distinguish between elicited speech samples produced by patients with dementia, and those from healthy controls.The difference between perplexity estimates from two neural language models (LMs) -one trained on transcripts of speech produced by healthy participants and the other trained on transcripts from patients with dementia -as a single feature for diagnostic classification of unseen transcripts has been shown to produce state-of-the-art performance.However, little is known about why this approach is effective, and on account of the lack of case/control matching in the most widely-used evaluation set of transcripts (De-mentiaBank), it is unclear if these approaches are truly diagnostic, or are sensitive to other variables.In this paper, we interrogate neural LMs trained on participants with and without dementia using synthetic narratives previously developed to simulate progressive semantic dementia by manipulating lexical frequency.We find that perplexity of neural LMs is strongly and differentially associated with lexical frequency, and that a mixture model resulting from interpolating control and dementia LMs improves upon the current state-of-the-art for models trained on transcript text exclusively. Trevor Cohen, Serguei V. S. Pakhomov |
ACL | 1 |
| 2020 | Retrofitting Vector Representations of Adverse Event Reporting Data to Structured Knowledge to Improve Pharmacovigilance Signal Detection
Xiruo Ding, Trevor Cohen |
AMIA | 2 |
| 2020 | The Centroid Cannot Hold: Comparing Sequential and Global Estimates of Coherence as Indicators of Formal Thought Disorder
Weizhe Xu, Jake Portanova, Ayesha Chander, Dror Ben-Zeev, Trevor Cohen |
AMIA | 5 |
| 2020 | SCOR: A secure international informatics infrastructure to investigate COVID-19abstractGlobal pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale. Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux |
J. Am. Medical Informatics Assoc. | 29 |
| 2019 | Predicting Adverse Drug-Drug Interactions with Neural Embedding of Semantic Predications
Hannah A. Burkhardt, Devika Subramanian, Justin Mower, Trevor Cohen |
AMIA | 4 |
| 2019 | Complementing Observational Signal with Distributed Representations for Drug Side-effect Prediction
Justin Mower, Trevor Cohen, Devika Subramanian |
AMIA | 2 |
| 2019 | aer2vec: Distributed Representations of Adverse Event Reporting System Data as a Means to Identify Drug/Side-Effect Associations
Jake Portanova, Nathan Murray, Justin Mower, Devika Subramanian, Trevor Cohen |
AMIA | 5 |
| 2019 | Cost-aware active learning for named entity recognition in clinical textabstractOBJECTIVE: Active Learning (AL) attempts to reduce annotation cost (ie, time) by selecting the most informative examples for annotation. Most approaches tacitly (and unrealistically) assume that the cost for annotating each sample is identical. This study introduces a cost-aware AL method, which simultaneously models both the annotation cost and the informativeness of the samples and evaluates both via simulation and user studies. MATERIALS AND METHODS: We designed a novel, cost-aware AL algorithm (Cost-CAUSE) for annotating clinical named entities; we first utilized lexical and syntactic features to estimate annotation cost, then we incorporated this cost measure into an existing AL algorithm. Using the 2010 i2b2/VA data set, we then conducted a simulation study comparing Cost-CAUSE with noncost-aware AL methods, and a user study comparing Cost-CAUSE with passive learning. RESULTS: Our cost model fit empirical annotation data well, and Cost-CAUSE increased the simulation area under the learning curve (ALC) scores by up to 5.6% and 4.9%, compared with random sampling and alternate AL methods. Moreover, in a user annotation task, Cost-CAUSE outperformed passive learning on the ALC score and reduced annotation time by 20.5%-30.2%. DISCUSSION: Although AL has proven effective in simulations, our user study shows that a real-world environment is far more complex. Other factors have a noticeable effect on the AL method, such as the annotation accuracy of users, the tiredness of users, and even the physical and mental condition of users. CONCLUSION: Cost-CAUSE saves significant annotation cost compared to random sampling. Qiang Wei 0002, Yukun Chen 0001, Mandana Salimi, Joshua C. Denny, Qiaozhu Mei, Thomas A. Lasko, Qingxia Chen, Stephen Wu 0004, Amy Franklin, Trevor Cohen, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 10 |
| 2019 | Rapamycin-mTOR+BRAF=? Using relational similarity to find therapeutically relevant drug-gene relationships in unstructured text
Safa Fathiamini, Amber M. Johnson, Vijaykumar Holla, Nora S. Sanchez, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
J. Biomed. Informatics | 8 |
| 2018 | Clinical text annotation - what factors are associated with the cost of time?
Qiang Wei 0002, Amy Franklin, Trevor Cohen, Hua Xu 0001 |
AMIA | 3 |
| 2018 | Bringing Order to Neural Word Embeddings with Embeddings Augmented by Random Permutations (EARP)abstractWord order is clearly a vital part of human language, but it has been used comparatively lightly in distributional vector models.This paper presents a new method for incorporating word order information into word vector embedding models by combining the benefits of permutation-based order encoding with the more recent method of skip-gram with negative sampling.The new method introduced here is called Embeddings Augmented by Random Permutations (EARP).It operates by applying permutations to the coordinates of context vector representations during the process of training.Results show an 8% improvement in accuracy on the challenging Bigger Analogy Test Set, and smaller but consistent improvements on other analogy reference sets.These findings demonstrate the importance of order-based information in analogical retrieval tasks, and the utility of random permutations as a means to augment neural embeddings. Trevor Cohen, Dominic Widdows |
CoNLL | 1 |
| 2018 | DataMed - an open source discovery index for finding biomedical datasetsabstractOBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community. Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Learning predictive models of drug side-effect relationships from distributed representations of literature-derived semantic predicationsabstractObjective: The aim of this work is to leverage relational information extracted from biomedical literature using a novel synthesis of unsupervised pretraining, representational composition, and supervised machine learning for drug safety monitoring. Methods: Using ≈80 million concept-relationship-concept triples extracted from the literature using the SemRep Natural Language Processing system, distributed vector representations (embeddings) were generated for concepts as functions of their relationships utilizing two unsupervised representational approaches. Embeddings for drugs and side effects of interest from two widely used reference standards were then composed to generate embeddings of drug/side-effect pairs, which were used as input for supervised machine learning. This methodology was developed and evaluated using cross-validation strategies and compared to contemporary approaches. To qualitatively assess generalization, models trained on the Observational Medical Outcomes Partnership (OMOP) drug/side-effect reference set were evaluated against a list of ≈1100 drugs from an online database. Results: The employed method improved performance over previous approaches. Cross-validation results advance the state of the art (AUC 0.96; F1 0.90 and AUC 0.95; F1 0.84 across the two sets), outperforming methods utilizing literature and/or spontaneous reporting system data. Examination of predictions for unseen drug/side-effect pairs indicates the ability of these methods to generalize, with over tenfold label support enrichment in the top 100 predictions versus the bottom 100 predictions. Discussion and Conclusion: Our methods can assist the pharmacovigilance process using information from the biomedical literature. Unsupervised pretraining generates a rich relationship-based representational foundation for machine learning techniques to classify drugs in the context of a putative side effect, given known examples. Justin Mower, Devika Subramanian, Trevor Cohen |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | Information Retrieval for Biomedical Datasets: The 2016 bioCADDIE Challenge
Kirk Roberts, Anupama E. Gururaj, Saeid Pournejati, Trevor Cohen, William R. Hersh, Dina Demner-Fushman, Lucila Ohno-Machado, Hua Xu 0001 |
AMIA | 5 |
| 2017 | Measuring content overlap during handoff communication using distributional semantics: An exploratory study
Joanna Abraham, Thomas George Kannampallil, Vignesh Srinivasan, William L. Galanter, Gail Tagney, Trevor Cohen |
J. Biomed. Informatics | 6 |
| 2017 | Using Pathfinder networks to discover alignment between expert and consumer conceptual knowledge from online vaccine content
Muhammad Amith, Rachel Cunningham, Lara S. Savas, Julie Boom, Roger W. Schvaneveldt, Cui Tao, Trevor Cohen |
J. Biomed. Informatics | 7 |
| 2017 | Embedding of semantic predications
Trevor Cohen, Dominic Widdows |
J. Biomed. Informatics | 1 |
| 2016 | A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 9 |
| 2016 | Literature-Based Discovery of Confounding in Observational Clinical Data
Scott A. Malec, Hua Xu 0001, Elmer V. Bernstam, Sahiti Myneni, Trevor Cohen |
AMIA | 6 |
| 2016 | Semantic Relatedness and Similarity between Biomedical Concepts
Sungrim Moon, Trevor Cohen, Hua Xu 0001 |
AMIA | 2 |
| 2016 | Classification-by-Analogy: Using Vector Representations of Implicit Relationships to Identify Plausibly Causal Drug/Side-effect Relationships
Justin Mower, Devika Subramanian, Ning Shang 0004, Trevor Cohen |
AMIA | 4 |
| 2016 | Content-specific network analysis of peer-to-peer communication in an online community for smoking cessation
Sahiti Myneni, Nathan K. Cobb, Trevor Cohen |
AMIA | 3 |
| 2016 | Characterization of Temporal Semantic Shifts of Peer-to-peer Communication in a Health-related Online Community: Implications for Data-driven Health Promotion
Vishnupriya Sridharan, Trevor Cohen, Nathan K. Cobb, Sahiti Myneni |
AMIA | 2 |
| 2016 | Analyzing Similarities in Handoff Communication Content between Residents and Nurses
Vignesh Srinivasan, Thomas George Kannampallil, Trevor Cohen, Joanna Abraham |
AMIA | 3 |
| 2016 | A Study of Active Learning for Document Selection in Clinical Named Entity Recognition
Qiang Wei 0002, Yukun Chen 0001, Sungrim Moon, Trevor Cohen, Hua Xu 0001 |
AMIA | 4 |
| 2016 | Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 10 |
| 2016 | Extracting genetic alteration information for personalized cancer therapy from ClinicalTrials.govabstractOBJECTIVE: Clinical trials investigating drugs that target specific genetic alterations in tumors are important for promoting personalized cancer therapy. The goal of this project is to create a knowledge base of cancer treatment trials with annotations about genetic alterations from ClinicalTrials.gov. METHODS: We developed a semi-automatic framework that combines advanced text-processing techniques with manual review to curate genetic alteration information in cancer trials. The framework consists of a document classification system to identify cancer treatment trials from ClinicalTrials.gov and an information extraction system to extract gene and alteration pairs from the Title and Eligibility Criteria sections of clinical trials. By applying the framework to trials at ClinicalTrials.gov, we created a knowledge base of cancer treatment trials with genetic alteration annotations. We then evaluated each component of the framework against manually reviewed sets of clinical trials and generated descriptive statistics of the knowledge base. RESULTS AND DISCUSSION: The automated cancer treatment trial identification system achieved a high precision of 0.9944. Together with the manual review process, it identified 20 193 cancer treatment trials from ClinicalTrials.gov. The automated gene-alteration extraction system achieved a precision of 0.8300 and a recall of 0.6803. After validation by manual review, we generated a knowledge base of 2024 cancer trials that are labeled with specific genetic alteration information. Analysis of the knowledge base revealed the trend of increased use of targeted therapy for cancer, as well as top frequent gene-alteration pairs of interest. We expect this knowledge base to be a valuable resource for physicians and patients who are seeking information about personalized cancer therapy. Jun Xu 0007, Hee-Jin Lee, Yonghui Wu 0001, Yaoyun Zhang, Liang-Chin Huang, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Trevor Cohen, Funda Meric-Bernstam, Elmer V. Bernstam, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 10 |
| 2016 | Automated identification of molecular effects of drugs (AIMED)abstractINTRODUCTION: Genomic profiling information is frequently available to oncologists, enabling targeted cancer therapy. Because clinically relevant information is rapidly emerging in the literature and elsewhere, there is a need for informatics technologies to support targeted therapies. To this end, we have developed a system for Automated Identification of Molecular Effects of Drugs, to help biomedical scientists curate this literature to facilitate decision support. OBJECTIVES: To create an automated system to identify assertions in the literature concerning drugs targeting genes with therapeutic implications and characterize the challenges inherent in automating this process in rapidly evolving domains. METHODS: We used subject-predicate-object triples (semantic predications) and co-occurrence relations generated by applying the SemRep Natural Language Processing system to MEDLINE abstracts and ClinicalTrials.gov descriptions. We applied customized semantic queries to find drugs targeting genes of interest. The results were manually reviewed by a team of experts. RESULTS: Compared to a manually curated set of relationships, recall, precision, and F2 were 0.39, 0.21, and 0.33, respectively, which represents a 3- to 4-fold improvement over a publically available set of predications (SemMedDB) alone. Upon review of ostensibly false positive results, 26% were considered relevant additions to the reference set, and an additional 61% were considered to be relevant for review. Adding co-occurrence data improved results for drugs in early development, but not their better-established counterparts. CONCLUSIONS: Precision medicine poses unique challenges for biomedical informatics systems that help domain experts find answers to their research questions. Further research is required to improve the performance of such systems, particularly for drugs in development. Safa Fathiamini, Amber M. Johnson, Alejandro Araya, Vijaykumar Holla, Ann M. Bailey, Beate Litzenburger, Nora S. Sanchez, Yekaterina Khotskaya, Hua Xu 0001, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
J. Am. Medical Informatics Assoc. | 13 |
| 2016 | Improving the utility of MeSH® terms using the TopicalMeSH representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
J. Biomed. Informatics | 3 |
| 2015 | Real Time Active Learning Study for Clinical Named Entity Recognition
Yukun Chen 0001, Sungrim Moon, Thomas A. Lasko, Qiaozhu Mei, Trevor Cohen, Qingxia Chen, Joshua C. Denny, Hua Xu 0001 |
AMIA | 6 |
| 2015 | Evaluating the Effects of Cognitive Support on Interpreting ICU Patient Data
Peter V. Killoran, Swaroop Gantela, Sahiti Myneni, Khalid F. Almoosa, Bela Patel, Thomas George Kannampallil, Vimla L. Patel, Trevor Cohen |
AMIA | 8 |
| 2015 | Improving Retrieval of PubMed Articles Using the TopicalMeSH Representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
AMIA | 3 |
| 2014 | Exploring the Use of SemRep Predications to Help Identify Secondary Drug Targets for Personalized Cancer Therapy
Safa Fathiamini, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Lauren Brusco, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
AMIA | 9 |
| 2014 | A study of synonym extraction from clinical texts using semantic vector models
Sungrim Moon, Trevor Cohen, Hua Xu 0001 |
AMIA | 2 |
| 2014 | Facilitating Visual Exploration of System-generated Reasoning Pathways underlying Drug-Side effect Relations
Sahiti Myneni, Trevor Cohen |
AMIA | 2 |
| 2014 | Identifying Plausible Adverse Drug Reactions Using Knowledge Extracted from the Literature
Ning Shang 0004, Hua Xu 0001, Thomas C. Rindflesch, Trevor Cohen |
AMIA | 4 |
| 2014 | Evaluating the effects of cognitive support on psychiatric clinical comprehension
Venkata Vijaya Kumar Dalai, Sana Khalid, Dinesh Gottipati, Thomas George Kannampallil, Vineeth John, Brett Blatter, Vimla L. Patel, Trevor Cohen |
Artif. Intell. Medicine | 8 |
| 2014 | Identifying plausible adverse drug reactions using knowledge extracted from the literature
Ning Shang 0004, Hua Xu 0001, Thomas C. Rindflesch, Trevor Cohen |
J. Biomed. Informatics | 4 |
| 2013 | Characterizing the Effects of a Cognitive Support System for Psychiatric Clinical Comprehension Venkata V.K. Dalai, MBBS, MPH, Dinesh Gottipatti, MS. Thomas Kannampallil, MS. Vineeth John, MD, MBA. Trevor Cohen, MBChB, PhD. University of Texas School of Biomedical Informatics1. New York Academy of Medicine2. University of Texas Medical School at Houston3
Venkata Vijaya Kumar Dalai, Dinesh Gottipati, Thomas George Kannampallil, Trevor Cohen |
AMIA | 4 |
| 2013 | Reflective Random Indexing to Develop a Medication-Problem Knowledge Base
Safa Fathiamini, Trevor Cohen, Allison B. McCoy, Dean F. Sittig |
AMIA | 2 |
| 2013 | Word Sense Disambiguation of Clinical Abbreviations with Hyperdimensional Computing
Sungrim Moon, Bjoern-Toby Berster, Hua Xu 0001, Trevor Cohen |
AMIA | 4 |
| 2013 | Individual Attributes of Behavior Change in an Online Social Network
Sahiti Myneni, Amy Franklin, Nathan K. Cobb, Trevor Cohen |
AMIA | 4 |
| 2013 | Identifying Persuasive Qualities of Decentralized Peer-to-Peer Online Social Networks in Public Health
Sahiti Myneni, M. Sriram Iyengar, Nathan K. Cobb, Trevor Cohen |
PERSUASIVE | 4 |
| 2013 | Understanding the nature of information seeking behavior in critical care: Implications for the design of health information technology
Thomas George Kannampallil, Amy Franklin, Rashmi Mishra, Khalid F. Almoosa, Trevor Cohen, Vimla L. Patel |
Artif. Intell. Medicine | 5 |
| 2012 | Hyperdimensional Computing Approach to Word Sense Disambiguation
Bjoern-Toby Berster, Joshua Goodwin, Trevor Cohen |
AMIA | 3 |
| 2012 | Designing a Learning Technology for Community Health Workers
Claire Loe, Trevor Cohen, Maria Fernandez, Amy Franklin |
AMIA | 2 |
| 2012 | Reducing Cognitive Load: Exploring Knowledge Model-driven Clinical Information Displays
Dean F. Sittig, Allison B. McCoy, Adam Wright, Amy Franklin, Trevor Cohen |
AMIA | 5 |
| 2012 | Deterministic Binary Vectors for Efficient Automated Indexing of MEDLINE/PubMed Abstracts
Manuel Wahle, Dominic Widdows, Jorge R. Herskovic, Elmer V. Bernstam, Trevor Cohen |
AMIA | 5 |
| 2012 | Graph-based signal integration for high-throughput phenotypingabstractBACKGROUND: Electronic Health Records aggregated in Clinical Data Warehouses (CDWs) promise to revolutionize Comparative Effectiveness Research and suggest new avenues of research. However, the effectiveness of CDWs is diminished by the lack of properly labeled data. We present a novel approach that integrates knowledge from the CDW, the biomedical literature, and the Unified Medical Language System (UMLS) to perform high-throughput phenotyping. In this paper, we automatically construct a graphical knowledge model and then use it to phenotype breast cancer patients. We compare the performance of this approach to using MetaMap when labeling records. RESULTS: MetaMap's overall accuracy at identifying breast cancer patients was 51.1% (n=428); recall=85.4%, precision=26.2%, and F1=40.1%. Our unsupervised graph-based high-throughput phenotyping had accuracy of 84.1%; recall=46.3%, precision=61.2%, and F1=52.8%. CONCLUSIONS: We conclude that our approach is a promising alternative for unsupervised high-throughput phenotyping. Jorge R. Herskovic, Devika Subramanian, Trevor Cohen, Pamela A. Bozzo-Silva, Charles F. Bearden, Elmer V. Bernstam |
BMC Bioinform. | 3 |
| 2012 | Focus on information retrieval: Predicting biomedical document access as a function of past useabstractOBJECTIVE: To determine whether past access to biomedical documents can predict future document access. MATERIALS AND METHODS: The authors used 394 days of query log (August 1, 2009 to August 29, 2010) from PubMed users in the Texas Medical Center, which is the largest medical center in the world. The authors evaluated two document access models based on the work of Anderson and Schooler. The first is based on how frequently a document was accessed. The second is based on both frequency and recency. RESULTS: The model based only on frequency of past access was highly correlated with the empirical data (R²=0.932), whereas the model based on frequency and recency had a much lower correlation (R²=0.668). DISCUSSION: The frequency-only model accurately predicted whether a document will be accessed based on past use. Modeling accesses as a function of frequency requires storing only the number of accesses and the creation date for the document. This model requires low storage overheads and is computationally efficient, making it scalable to large corpora such as MEDLINE. CONCLUSION: It is feasible to accurately model the probability of a document being accessed in the future based on past accesses. J. Caleb Goodwin, Todd R. Johnson, Trevor Cohen, Jorge R. Herskovic, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Discovering discovery patterns with predication-based Semantic Indexing
Trevor Cohen, Dominic Widdows, Roger W. Schvaneveldt, Peter Davies 0002, Thomas C. Rindflesch |
J. Biomed. Informatics | 1 |
| 2012 | Enhancing clinical concept extraction with distributional semantics
Siddhartha Jonnalagadda, Trevor Cohen, Stephen T. Wu, Graciela Gonzalez-Hernandez |
J. Biomed. Informatics | 2 |
| 2012 | Avatar-based simulation in the evaluation of diagnosis and management of mental health disorders in primary care
Rachel M. Satter, Trevor Cohen, Pierina Ortiz, Kanav Kahol, James Mackenzie, Carol Olson, Mina Johnson-Glenberg, Vimla L. Patel |
J. Biomed. Informatics | 2 |
| 2011 | Making sense: Sensor-based investigation of clinician activities in complex critical care environments
Thomas George Kannampallil, Zhe Li 0053, Min Zhang 0001, Trevor Cohen, David J. Robinson, Amy Franklin, Vimla L. Patel |
J. Biomed. Informatics | 4 |
| 2011 | Considering complexity in healthcare systems
Thomas George Kannampallil, Guido F. Schauer, Trevor Cohen, Vimla L. Patel |
J. Biomed. Informatics | 3 |
| 2011 | Recovery at the edge of error: Debunking the myth of the infallible expert
Vimla L. Patel, Trevor Cohen, Tripti Murarka, Joanne Olsen, Srujana Kagita, Sahiti Myneni, Timothy G. Buchman, Vafa Ghaemmaghami |
J. Biomed. Informatics | 2 |
| 2011 | Toward automated workflow analysis and visualization in clinical environments
Mithra Vankipuram, Kanav Kahol, Trevor Cohen, Vimla L. Patel |
J. Biomed. Informatics | 3 |
| 2010 | A Distributional Semantics Approach to Simultaneous Recognition of Multiple Classes of Named Entities
Siddhartha Jonnalagadda, Robert Leaman, Trevor Cohen, Graciela Gonzalez-Hernandez |
CICLing | 3 |
| 2010 | Reflective Random Indexing and indirect inference: A scalable method for discovery of implicit connections
Trevor Cohen, Roger W. Schvaneveldt, Dominic Widdows |
J. Biomed. Informatics | 1 |
| 2010 | Reflective random indexing for semi-automatic indexing of the biomedical literature
Vidya Vasuki, Trevor Cohen |
J. Biomed. Informatics | 2 |
| 2009 | Predication-based Semantic Indexing: Permutations as a Means to Encode Predications in Semantic Space
Trevor Cohen, Roger W. Schvaneveldt, Thomas C. Rindflesch |
AMIA | 1 |
| 2009 | The Cognitive Basis of Effective Team Performance: Features of Failure and Success in Simulated Cardiac Resuscitation
Pallavi Shetty, Trevor Cohen, Bhavesh Patel, Vimla L. Patel |
AMIA | 2 |
| 2009 | Visualization and Analysis of Activities in Critical Care Environments
Mithra Vankipuram, Kanav Kahol, Trevor Cohen, Vimla L. Patel |
AMIA | 3 |
| 2009 | Empirical distributional semantics: Methods and biomedical applications
Trevor Cohen, Dominic Widdows |
J. Biomed. Informatics | 1 |
| 2008 | Exploring MEDLINE Space with Random Indexing and Pathfinder Networks
Trevor Cohen |
AMIA | 1 |
| 2008 | Simulating expert clinical comprehension: Adapting latent semantic analysis to accurately extract clinical concepts from psychiatric narrative
Trevor Cohen, Brett Blatter, Vimla L. Patel |
J. Biomed. Informatics | 1 |
| 2007 | Research Paper: Reevaluating Recovery: Perceived Violations and Preemptive Interventions on Emergency Psychiatry RoundsabstractOBJECTIVE: Contemporary error research suggests that the quest to eradicate error is misguided. Error commission, detection, and recovery are an integral part of cognitive work, even at the expert level. In collaborative workspaces, the perception of potential error is directly observable: workers discuss and respond to perceived violations of accepted practice norms. As perceived violations are captured and corrected preemptively, they do not fit Reason's widely accepted definition of error as "failure to achieve an intended outcome." However, perceived violations suggest the aversion of potential error, and consequently have implications for error prevention. This research aims to identify and describe perceived violations of the boundaries of accepted procedure in a psychiatric emergency department (PED), and how they are resolved in practice. DESIGN: Clinical discourse from fourteen PED patient rounds was audio-recorded. Excerpts from recordings suggesting perceived violations or incidents of miscommunication were extracted and analyzed using qualitative coding methods. The results are interpreted in relation to prior research on vulnerabilities to error in the PED. RESULTS: Thirty incidents of perceived violations or miscommunication are identified and analyzed. Of these, only one medication error was formally reported. Other incidents would not have been detected by a retrospective analysis. CONCLUSIONS: The analysis of perceived violations expands the data available for error analysis beyond occasional reported adverse events. These data are prospective: responses are captured in real time. This analysis supports a set of recommendations to improve the quality of care in the PED and other critical care contexts. Trevor Cohen, Brett Blatter, Vimla L. Patel |
J. Am. Medical Informatics Assoc. | 1 |
| 2006 | A cognitive blueprint of collaboration in context: Distributed cognition in the psychiatric emergency department
Trevor Cohen, Brett Blatter, Edward H. Shortliffe, Vimla L. Patel |
Artif. Intell. Medicine | 1 |
| 2005 | Exploring dangerous neighborhoods: Latent Semantic Analysis and computing beyond the bounds of the familiar
Trevor Cohen, Brett Blatter, Vimla L. Patel |
AMIA | 1 |