VLDB 2026 Research / reviewers in the wild / expert
Jyotishman Pathak
dblp:52/1723
· DBLP profile ↗
116ranked-venue papers
16as first author
22since 2021 · last 2025
0000-0002-4856-410XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 106 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 7Software engineering, systems software and programming languages · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language modelabstractOBJECTIVES: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented in narrative clinical notes rather than as structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of extraction of such information. MATERIALS AND METHODS: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n = 300) and Weill Cornell Medicine (WCM, n = 225) were annotated to create a gold-standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (eg, social network, instrumental support, and loneliness). RESULTS: For extracting SS/SI, the RBS obtained higher macroaveraged F1-scores than the LLM at both MSHS (0.89 versus 0.65) and WCM (0.85 versus 0.82). For extracting the subcategories, the RBS also outperformed the LLM at both MSHS (0.90 versus 0.62) and WCM (0.82 versus 0.81). DISCUSSION AND CONCLUSION: Unexpectedly, the RBS outperformed the LLMs across all metrics. An intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS was designed and refined to follow the same specific rules as the gold-standard annotations. Conversely, the LLM was more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages, although additional replication studies are warranted. Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Ardesheer Talati, Myrna Weissman, Mark Olfson, J. John Mann, Yiye Zhang, Alexander Charney, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 20 |
| 2024 | Visualizing machine learning-based predictions of postpartum depression risk for lay audiencesabstractOBJECTIVES: To determine if different formats for conveying machine learning (ML)-derived postpartum depression risks impact patient classification of recommended actions (primary outcome) and intention to seek care, perceived risk, trust, and preferences (secondary outcomes). MATERIALS AND METHODS: We recruited English-speaking females of childbearing age (18-45 years) using an online survey platform. We created 2 exposure variables (presentation format and risk severity), each with 4 levels, manipulated within-subject. Presentation formats consisted of text only, numeric only, gradient number line, and segmented number line. For each format viewed, participants answered questions regarding each outcome. RESULTS: Five hundred four participants (mean age 31 years) completed the survey. For the risk classification question, performance was high (93%) with no significant differences between presentation formats. There were main effects of risk level (all P < .001) such that participants perceived higher risk, were more likely to agree to treatment, and more trusting in their obstetrics team as the risk level increased, but we found inconsistencies in which presentation format corresponded to the highest perceived risk, trust, or behavioral intention. The gradient number line was the most preferred format (43%). DISCUSSION AND CONCLUSION: All formats resulted high accuracy related to the classification outcome (primary), but there were nuanced differences in risk perceptions, behavioral intentions, and trust. Investigators should choose health data visualizations based on the primary goal they want lay audiences to accomplish with the ML risk score. Pooja M. Desai, Sarah Harkins, Saanjaana Rahman, Shiveen Kumar, Alison Hermann, Rochelle Joly, Yiye Zhang, Jyotishman Pathak, Jessica Kim, Deborah D'angelo, Natalie C. Benda, Meghan Reading Turchioe |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Preparing for the bedside - optimizing a postpartum depression risk prediction model for clinical implementation in a health systemabstractOBJECTIVE: We developed and externally validated a machine-learning model to predict postpartum depression (PPD) using data from electronic health records (EHRs). Effort is under way to implement the PPD prediction model within the EHR system for clinical decision support. We describe the pre-implementation evaluation process that considered model performance, fairness, and clinical appropriateness. MATERIALS AND METHODS: We used EHR data from an academic medical center (AMC) and a clinical research network database from 2014 to 2020 to evaluate the predictive performance and net benefit of the PPD risk model. We used area under the curve and sensitivity as predictive performance and conducted a decision curve analysis. In assessing model fairness, we employed metrics such as disparate impact, equal opportunity, and predictive parity with the White race being the privileged value. The model was also reviewed by multidisciplinary experts for clinical appropriateness. Lastly, we debiased the model by comparing 5 different debiasing approaches of fairness through blindness and reweighing. RESULTS: We determined the classification threshold through a performance evaluation that prioritized sensitivity and decision curve analysis. The baseline PPD model exhibited some unfairness in the AMC data but had a fair performance in the clinical research network data. We revised the model by fairness through blindness, a debiasing approach that yielded the best overall performance and fairness, while considering clinical appropriateness suggested by the expert reviewers. DISCUSSION AND CONCLUSION: The findings emphasize the need for a thorough evaluation of intervention-specific models, considering predictive performance, fairness, and appropriateness before clinical implementation. Rochelle Joly, Meghan Reading Turchioe, Natalie C. Benda, Alison Hermann, Ashley Beecy, Jyotishman Pathak, Yiye Zhang |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias
Zehao Yu 0001, Cheng Peng 0009, Xi Yang 0015, Chong Dang, Prakash Adekkanattu, Braja Gopal Patra, Yifan Peng 0002, Jyotishman Pathak, Debbie L. Wilson, Ching-Yuan Chang, Wei-Hsuan Lo-Ciganic, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 8 |
| 2023 | CoRL: A Cost-Responsive Learning Optimizer for Neural NetworksabstractSelection of the optimal learning rate for training neural networks has often been a matter of concern for the machine learning community. The existing learning rates are dependent on multiple scaling factors. This paper proposes Cost-Responsive Learning (CoRL) which does not require manual hyper-parameter tuning. It maintains a linear relationship with the prediction error of the neural network. This is expected to offer the lowest learning rate at the global minima, and higher learning rates elsewhere. Hence, a number proportional to the prediction error is used as a learning rate, subject to the constraint that the number is within an acceptable range (here [0,1]). The derivation of an optimal learning rate from a given cost function is illustrated with the popular binary/categorical cross-entropy cost function(s). Experiments performed under multiple settings demonstrate that, with the CoRL optimizer needs no parameter tuning to obtain state-of-the-art results with significantly lower training time for equivalent performance. Reshma Kar, Vijay Kumar Reddy Voddi, Braja Gopal Patra, Jyotishman Pathak |
SMC | 4 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 21 |
| 2023 | An NLP approach to identify SDoH-related circumstance and suicide crisis from death investigation narrativesabstractOBJECTIVES: Suicide presents a major public health challenge worldwide, affecting people across the lifespan. While previous studies revealed strong associations between Social Determinants of Health (SDoH) and suicide deaths, existing evidence is limited by the reliance on structured data. To resolve this, we aim to adapt a suicide-specific SDoH ontology (Suicide-SDoHO) and use natural language processing (NLP) to effectively identify individual-level SDoH-related social risks from death investigation narratives. MATERIALS AND METHODS: We used the latest National Violent Death Report System (NVDRS), which contains 267 804 victim suicide data from 2003 to 2019. After adapting the Suicide-SDoHO, we developed a transformer-based model to identify SDoH-related circumstances and crises in death investigation narratives. We applied our model retrospectively to annotate narratives whose crisis variables were not coded in NVDRS. The crisis rates were calculated as the percentage of the group's total suicide population with the crisis present. RESULTS: The Suicide-SDoHO contains 57 fine-grained circumstances in a hierarchical structure. Our classifier achieves AUCs of 0.966 and 0.942 for classifying circumstances and crises, respectively. Through the crisis trend analysis, we observed that not everyone is equally affected by SDoH-related social risks. For the economic stability crisis, our result showed a significant increase in crisis rate in 2007-2009, parallel with the Great Recession. CONCLUSIONS: This is the first study curating a Suicide-SDoHO using death investigation narratives. We showcased that our model can effectively classify SDoH-related social risks through NLP approaches. We hope our study will facilitate the understanding of suicide crises and inform effective prevention strategies. Song Wang 0026, Yifang Dang, Zhaoyi Sun, Ying Ding 0001, Jyotishman Pathak, Cui Tao, Yunyu Xiao, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
J. Biomed. Informatics | 12 |
| 2022 | Developing a disease-specific symptom vocabulary for natural language processing
Meghan Reading Turchioe, Winston Guo, Alexander Volodarskiy, Brittany Taylor, Mollie Hobensack, David Slotwiner, Jyotishman Pathak |
AMIA | 7 |
| 2022 | Design and validation of a FHIR-based EHR-driven phenotyping toolboxabstractOBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error. Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 15 |
| 2022 | An architecture for research computing in health to support clinical and translational investigators with electronic patient dataabstractOBJECTIVE: Obtaining electronic patient data, especially from electronic health record (EHR) systems, for clinical and translational research is difficult. Multiple research informatics systems exist but navigating the numerous applications can be challenging for scientists. This article describes Architecture for Research Computing in Health (ARCH), our institution's approach for matching investigators with tools and services for obtaining electronic patient data. MATERIALS AND METHODS: Supporting the spectrum of studies from populations to individuals, ARCH delivers a breadth of scientific functions-including but not limited to cohort discovery, electronic data capture, and multi-institutional data sharing-that manifest in specific systems-such as i2b2, REDCap, and PCORnet. Through a consultative process, ARCH staff align investigators with tools with respect to study design, data sources, and cost. Although most ARCH services are available free of charge, advanced engagements require fee for service. RESULTS: Since 2016 at Weill Cornell Medicine, ARCH has supported over 1200 unique investigators through more than 4177 consultations. Notably, ARCH infrastructure enabled critical coronavirus disease 2019 response activities for research and patient care. DISCUSSION: ARCH has provided a technical, regulatory, financial, and educational framework to support the biomedical research enterprise with electronic patient data. Collaboration among informaticians, biostatisticians, and clinicians has been critical to rapid generation and analysis of EHR data. CONCLUSION: A suite of tools and services, ARCH helps match investigators with informatics systems to reduce time to science. ARCH has facilitated research at Weill Cornell Medicine and may provide a model for informatics and research leaders to support scientists elsewhere. Thomas R. Campion Jr., Evan Sholle, Jyotishman Pathak, Stephen B. Johnson, John P. Leonard, Curtis L. Cole |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Multi-site Evaluation of Longitudinal Changes in Ejection Fraction in Heart Failure Patients Through Data-driven Phenotyping
Prakash Adekkanattu, Jennifer A. Pacheco, Joseph Kabariti, Daniel J. Stone, Yue Yu 0012, Parag Goyal, Faraz S. Ahmad, Guoqian Jiang, Yuan Luo 0001, Luke V. Rasmussen, Pascal S. Brandt, Jie Xu 0012, Fei Wang 0001, Natalie C. Benda, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 17 |
| 2021 | Supporting EHR-based Cohort Discovery Through User-centered Design: Results of an Early Formative Usability Study
Natalie C. Benda, Pascal S. Brandt, Jessica S. Ancker, Jennifer A. Pacheco, Prakash Adekkanattu, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 7 |
| 2021 | Deep Significance Clustering (DICE) a Heterogenous Population
Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Alison Hermann, Rochelle Joly, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
AMIA | 10 |
| 2021 | Extracting Social Isolation Information From Psychiatric Notes in the Electronic Health Records
Lauren A. Lepow, Braja Gopal Patra, Isotta Landi, Prakash Adekkanattu, Jyotishman Pathak, Mark Olfson, J. John Mann, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Priya Wickramaratne, Myrna Weissman, Benjamin S. Glicksberg, Alexander Charney |
AMIA | 5 |
| 2021 | A Deep Learning Framework Using a Pre-trained BERT Model to Predict the Risk of Progression from Mild Cognitive Impairment to Alzheimer's Disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Fei Wang 0001, Richard Isaacson, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 8 |
| 2021 | Prescribing Pharmacogenomics Testing: Analyzing the Acceptance of Healthcare Providers through a Survey Study
Mohit Manoj Sharma, Yonaka Harris, Yiye Zhang, Jyotishman Pathak |
AMIA | 4 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 12 |
| 2021 | On Constraints and Considerations for Extending Support for Natural Language Processing-Based FHIR Resource Generation
Andrew Wen, Luke V. Rasmussen, Daniel J. Stone, Sijia Liu 0002, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Yuan Luo 0001, Fei Wang 0001, Jyotishman Pathak, Guoqian Jiang |
AMIA | 10 |
| 2021 | A Study of Social and Behavioral Determinants of Health in Lung Cancer Patients Using Transformers-based Natural Language Processing Models
Zehao Yu 0001, Xi Yang 0015, Chong Dang, Songzi Wu, Prakash Adekkanattu, Jyotishman Pathak, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
AMIA | 6 |
| 2021 | Deep significance clustering: a novel approach for identifying risk-stratified and predictive patient subgroupsabstractOBJECTIVE: Deep significance clustering (DICE) is a self-supervised learning framework. DICE identifies clinically similar and risk-stratified subgroups that neither unsupervised clustering algorithms nor supervised risk prediction algorithms alone are guaranteed to generate. MATERIALS AND METHODS: Enabled by an optimization process that enforces statistical significance between the outcome and subgroup membership, DICE jointly trains 3 components, representation learning, clustering, and outcome prediction while providing interpretability to the deep representations. DICE also allows unseen patients to be predicted into trained subgroups for population-level risk stratification. We evaluated DICE using electronic health record datasets derived from 2 urban hospitals. Outcomes and patient cohorts used include discharge disposition to home among heart failure (HF) patients and acute kidney injury among COVID-19 (Cov-AKI) patients, respectively. RESULTS: Compared to baseline approaches including principal component analysis, DICE demonstrated superior performance in the cluster purity metrics: Silhouette score (0.48 for HF, 0.51 for Cov-AKI), Calinski-Harabasz index (212 for HF, 254 for Cov-AKI), and Davies-Bouldin index (0.86 for HF, 0.66 for Cov-AKI), and prediction metric: area under the Receiver operating characteristic (ROC) curve (0.83 for HF, 0.78 for Cov-AKI). Clinical evaluation of DICE-generated subgroups revealed more meaningful distributions of member characteristics across subgroups, and higher risk ratios between subgroups. Furthermore, DICE-generated subgroup membership alone was moderately predictive of outcomes. DISCUSSION: DICE addresses a gap in current machine learning approaches where predicted risk may not lead directly to actionable clinical steps. CONCLUSION: DICE demonstrated the potential to apply in heterogeneous populations, where having the same quantitative risk does not equate with having a similar clinical profile. Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | Extracting social determinants of health from electronic health records using natural language processing: a systematic reviewabstractOBJECTIVE: Social determinants of health (SDoH) are nonclinical dispositions that impact patient health risks and clinical outcomes. Leveraging SDoH in clinical decision-making can potentially improve diagnosis, treatment planning, and patient outcomes. Despite increased interest in capturing SDoH in electronic health records (EHRs), such information is typically locked in unstructured clinical notes. Natural language processing (NLP) is the key technology to extract SDoH information from clinical text and expand its utility in patient care and research. This article presents a systematic review of the state-of-the-art NLP approaches and tools that focus on identifying and extracting SDoH data from unstructured clinical text in EHRs. MATERIALS AND METHODS: A broad literature search was conducted in February 2021 using 3 scholarly databases (ACL Anthology, PubMed, and Scopus) following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 6402 publications were initially identified, and after applying the study inclusion criteria, 82 publications were selected for the final review. RESULTS: Smoking status (n = 27), substance use (n = 21), homelessness (n = 20), and alcohol use (n = 15) are the most frequently studied SDoH categories. Homelessness (n = 7) and other less-studied SDoH (eg, education, financial problems, social isolation and support, family problems) are mostly identified using rule-based approaches. In contrast, machine learning approaches are popular for identifying smoking status (n = 13), substance use (n = 9), and alcohol use (n = 9). CONCLUSION: NLP offers significant potential to extract SDoH data from narrative clinical notes, which in turn can aid in the development of screening tools, risk prediction models, and clinical decision support systems. Braja Gopal Patra, Mohit Manoj Sharma, Veer Vekaria, Prakash Adekkanattu, Olga V. Patterson, Benjamin S. Glicksberg, Lauren A. Lepow, Euijung Ryu, Joanna M. Biernacka, Al'ona Furmanchuk, Thomas J. George, William R. Hogan, Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Myrna Weissman, Priya Wickramaratne, J. John Mann, Mark Olfson, Thomas R. Campion Jr., Mark G. Weiner, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 22 |
| 2020 | Feasibility of Cross-Platform EHR-Driven Phenotyping Using Clinical Quality Language
Pascal S. Brandt, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Faraz S. Ahmad, Jie Xu 0012, Jessica S. Ancker, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 13 |
| 2020 | Weak Supervision to Classify Unstructured Clinical Text for Current Suicidal Ideation
Marika M. Cusick, Prakash Adekkanattu, Thomas R. Campion Jr., Evan Sholle, Annie C. Myers, George Alexopoulos, Jyotishman Pathak |
AMIA | 7 |
| 2020 | Visual Rating Scales for Patient-Reported Outcome Measurement: A National Validation Study
Lisa Grossman Liu, Meghan Reading Turchioe, Annie C. Myers, Jyotishman Pathak, David K. Vawdrey, Ruth M. Masterson Creber |
AMIA | 4 |
| 2020 | Provider perspectives on the clinical utility of using a risk prediction tool for postpartum depression
Annie C. Myers, Fariha Ahsan, Rochelle Joly, Alison Hermann, Yiye Zhang, Michael Laskoff, Jyotishman Pathak, Meghan Reading Turchioe |
AMIA | 7 |
| 2020 | Evaluating Commercially Available Mobile Apps for Depression Self-Management
Annie C. Myers, Lewis Chesebrough, Ruixuan Hu, Meghan Reading Turchioe, Jyotishman Pathak, Ruth M. Masterson Creber |
AMIA | 5 |
| 2020 | Healthcare Contacts Among Patients with Psychiatric Hospitalization Admitted Through the Emergency Department
Wenna Xi, Samprit Banerjee, Robert B. Penfold, Greg E. Simon, George Alexopoulos, Jyotishman Pathak |
AMIA | 6 |
| 2020 | Identification of Alzheimer's Disease Subtypes from Electronic Health Records Using a Data-Driven Approach
Jie Xu 0012, Fei Wang 0001, Prakash Adekkanattu, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Yuan Luo 0001, Chengsheng Mao, Jennifer A. Pacheco, Luke V. Rasmussen, Yiye Zhang, Richard Isaacson, Jyotishman Pathak |
AMIA | 14 |
| 2020 | Identifying sub-phenotypes of acute kidney injury using structured and unstructured electronic health record data with memory networks
Jingyuan Chou, Xi Sheryl Zhang, Yuan Luo 0001, Tamara Isakova, Prakash Adekkanattu, Jessica S. Ancker, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Luke V. Rasmussen, Jyotishman Pathak, Fei Wang 0001 |
J. Biomed. Informatics | 12 |
| 2019 | Evaluating the Portability of an NLP System for Processing Echocardiograms: A Retrospective, Multi-site Observational Study
Prakash Adekkanattu, Guoqian Jiang, Yuan Luo 0001, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Richard C. Kiefer, Daniel J. Stone, Pascal S. Brandt, Yizhen Zhong, Fei Wang 0001, Jessica S. Ancker, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 17 |
| 2019 | Using Social Media to Study Mental Health Conditions - Challenges and Opportunities
Vasa Curcin, Elizabeth Ford, Jyotishman Pathak, Goran Nenadic |
AMIA | 3 |
| 2019 | Collecting Individual-Level Social Determinants of Health to Inform Patient-Centered Outcomes Research in Mental Health
Joseph DeFerio, Annie C. Myers, John M. Meddar, Judith Cukor, Jyotishman Pathak |
AMIA | 5 |
| 2019 | Considerations for Improving the Portability of Electronic Health Record-Based Phenotype Algorithms
Luke V. Rasmussen, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Jessica S. Ancker, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 10 |
| 2019 | Validation of a Computable Phenotype for Site-Specific Cancer Treatment
Evan Sholle, Orrin Belden, Jaclyn Rosenzweig, Jennifer Levine, Joseph Kabariti, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 7 |
| 2019 | A Survey of Clinicians' Perception of Pharmacogenomics Testing for Antidepressants
Yiye Zhang, Yonaka Harris, George Alexopoulos, Adam Stracher, Elizabeth Ross, Rainu Kaushal, Jyotishman Pathak |
AMIA | 7 |
| 2019 | Knowledge-aware Assessment of Severity of Suicide Risk for Early InterventionabstractMental health illness such as depression is a significant risk factor for suicide ideation, behaviors, and attempts. A report by Substance Abuse and Mental Health Services Administration (SAMHSA) shows that 80% of the patients suffering from Borderline Personality Disorder (BPD) have suicidal behavior, 5-10% of whom commit suicide. While multiple initiatives have been developed and implemented for suicide prevention, a key challenge has been the social stigma associated with mental disorders, which deters patients from seeking help or sharing their experiences directly with others including clinicians. This is particularly true for teenagers and younger adults where suicide is the second highest cause of death in the US. Prior research involving surveys and questionnaires (e.g. PHQ-9) for suicide risk prediction failed to provide a quantitative assessment of risk that informed timely clinical decision-making for intervention. Our interdisciplinary study concerns the use of Reddit as an unobtrusive data source for gleaning information about suicidal tendencies and other related mental health conditions afflicting depressed users. We provide details of our learning framework that incorporates domain-specific knowledge to predict the severity of suicide risk for an individual. Our approach involves developing a suicide risk severity lexicon using medical knowledge bases and suicide ontology to detect cues relevant to suicidal thoughts and actions. We also use language modeling, medical entity recognition and normalization and negation detection to create a dataset of 2181 redditors that have discussed or implied suicidal ideation, behavior, or attempt. Given the importance of clinical knowledge, our gold standard dataset of 500 redditors (out of 2181) was developed by four practicing psychiatrists following the guidelines outlined in Columbia Suicide Severity Rating Scale (C-SSRS), with the pairwise annotator agreement of 0.79 and group-wise agreement of 0.73. Compared to the existing four-label classification scheme (no risk, low risk, moderate risk, and high risk), our proposed C-SSRS-based 5-label classification scheme distinguishes people who are supportive, from those who show different severity of suicidal tendency. Our 5-label classification scheme outperforms the state-of-the-art schemes by improving the graded recall by 4.2% and reducing the perceived risk measure by 12.5%. Convolutional neural network (CNN) provided the best performance in our scheme due to the discriminative features and use of domain-specific knowledge resources, in comparison to SVM-L that has been used in the state-of-the-art tools over similar dataset. Manas Gaur, Amanuel Alambo, Joy Prakash Sain, Ugur Kursuncu, Krishnaprasad Thirunarayan, Ramakanth Kavuluru, Amit P. Sheth, Randy S. Welton, Jyotishman Pathak |
WWW | 9 |
| 2019 | Drug knowledge bases and their applications in biomedical informatics researchabstractRecent advances in biomedical research have generated a large volume of drug-related data. To effectively handle this flood of data, many initiatives have been taken to help researchers make good use of them. As the results of these initiatives, many drug knowledge bases have been constructed. They range from simple ones with specific focuses to comprehensive ones that contain information on almost every aspect of a drug. These curated drug knowledge bases have made significant contributions to the development of efficient and effective health information technologies for better health-care service delivery. Understanding and comparing existing drug knowledge bases and how they are applied in various biomedical studies will help us recognize the state of the art and design better knowledge bases in the future. In addition, researchers can get insights on novel applications of the drug knowledge bases through a review of successful use cases. In this study, we provide a review of existing popular drug knowledge bases and their applications in drug-related studies. We discuss challenges in constructing and using drug knowledge bases as well as future research directions toward a better ecosystem of drug knowledge bases. Yongjun Zhu 0001, Olivier Elemento, Jyotishman Pathak, Fei Wang 0001 |
Briefings Bioinform. | 3 |
| 2019 | Social determinants of health in mental health care and research: a case for greater inclusionabstractSocial determinants of health (SDOH) are known to influence mental health outcomes, which are independent risk factors for poor health status and physical illness. Currently, however, existing SDOH data collection methods are ad hoc and inadequate, and SDOH data are not systematically included in clinical research or used to inform patient care. Social contextual data are rarely captured prospectively in a structured and comprehensive manner, leaving large knowledge gaps. Extraction methods are now being developed to facilitate the collection, standardization, and integration of SDOH data into electronic health records. If successful, these efforts may have implications for health equity, such as reducing disparities in access and outcomes. Broader use of surveys, natural language processing, and machine learning methods to harness SDOH may help researchers and clinical teams reduce barriers to mental health care. Joseph DeFerio, Scott Breitinger, Dhruv Khullar, Amit P. Sheth, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Underserved populations with missing race ethnicity data differ significantly from those with structured race/ethnicity documentationabstractOBJECTIVE: We aimed to address deficiencies in structured electronic health record (EHR) data for race and ethnicity by identifying black and Hispanic patients from unstructured clinical notes and assessing differences between patients with or without structured race/ethnicity data. MATERIALS AND METHODS: Using EHR notes for 16 665 patients with encounters at a primary care practice, we developed rule-based natural language processing (NLP) algorithms to classify patients as black/Hispanic. We evaluated performance of the method against an annotated gold standard, compared race and ethnicity between NLP-derived and structured EHR data, and compared characteristics of patients identified as black or Hispanic using only NLP vs patients identified as such only in structured EHR data. RESULTS: For the sample of 16 665 patients, NLP identified 948 additional patients as black, a 26%increase, and 665 additional patients as Hispanic, a 20% increase. Compared with the patients identified as black or Hispanic in structured EHR data, patients identified as black or Hispanic via NLP only were older, more likely to be male, less likely to have commercial insurance, and more likely to have higher comorbidity. DISCUSSION: Structured EHR data for race and ethnicity are subject to data quality issues. Supplementing structured EHR race data with NLP-derived race and ethnicity may allow researchers to better assess the demographic makeup of populations and draw more accurate conclusions about intergroup differences in health outcomes. CONCLUSIONS: Black or Hispanic patients who are not documented as such in structured EHR race/ethnicity fields differ significantly from those who are. Relatively simple NLP can help address this limitation. Evan Sholle, Laura C. Pinheiro, Prakash Adekkanattu, Marcos Davila, Stephen B. Johnson, Jyotishman Pathak, Sanjai Sinha, Cassidie Li, Stasi A. Lubansky, Monika M. Safford, Thomas R. Campion Jr. |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Developing a FHIR-based EHR phenotyping framework: A case study for identification of patients with obesity and multiple comorbidities from discharge summaries
Na Hong, Andrew Wen, Daniel J. Stone, Shintaro Tsuji, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Prakash Adekkanattu, Fei Wang 0001, Yuan Luo 0001, Jyotishman Pathak, Guoqian Jiang |
J. Biomed. Informatics | 11 |
| 2018 | Ascertaining Depression Severity by Extracting Patient Health Questionnaire-9 (PHQ-9) Scores from Clinical Notes
Prakash Adekkanattu, Evan Sholle, Joseph DeFerio, Jyotishman Pathak, Stephen B. Johnson, Thomas R. Campion Jr. |
AMIA | 4 |
| 2018 | Methods for Integrating EHRs, Social Determinants of Health, and Built Environment Data for Patient-Centered Research
Joseph DeFerio, Evan Sholle, Kathleen Lee, Jyotishman Pathak |
AMIA | 4 |
| 2018 | Using EHRs to assess gender differences in the utilization of mental health services for patients diagnosed with major depression
Cody M. Phelps, Trisha R. Sanghvi, Ningrui Zhang, Joseph DeFerio, Jyotishman Pathak |
AMIA | 5 |
| 2018 | "Let Me Tell You About Your Mental Health!": Contextualized Classification of Reddit Posts to DSM-5 for Web-based InterventionabstractSocial media platforms are increasingly being used to share and seek advice on mental health issues. In particular, Reddit users freely discuss such issues on various subreddits, whose structure and content can be leveraged to formally interpret and relate subreddits and their posts in terms of mental health diagnostic categories. There is prior research on the extraction of mental health-related information, including symptoms, diagnosis, and treatments from social media; however, our approach can additionally provide actionable information to clinicians about the mental health of a patient in diagnostic terms for web-based intervention. Specifically, we provide a detailed analysis of the nature of subreddit content from domain expert's perspective and introduce a novel approach to map each subreddit to the best matching DSM-5 (Diagnostic and Statistical Manual of Mental Disorders - 5th Edition) category using multi-class classifier. Our classification algorithm analyzes all the posts of a subreddit by adapting topic modeling and word-embedding techniques, and utilizing curated medical knowledge bases to quantify relationship to DSM-5 categories. Our semantic encoding-decoding optimization approach reduces the false-alarm-rate from 30% to 2.5% over a comparable heuristic baseline, and our mapping results have been verified by domain experts achieving a kappa score of 0.84. Manas Gaur, Ugur Kursuncu, Amanuel Alambo, Amit P. Sheth, Raminta Daniulaityte, Krishnaprasad Thirunarayan, Jyotishman Pathak |
CIKM | 7 |
| 2018 | The potential value of social determinants of health in predicting health outcomesabstractDear Dr Ohno-Machado, As Kasthurirathne and colleagues1 point out in their recent JAMIA paper, it is well established that population health is affected by socioeconomic status and other social determinants of health (SDH). It is therefore reasonable to hypothesize that SDH data should have predictive power for clinical and healthcare utilization outcomes for individual patients, yet the authors found that adding SDH data to clinical data produced no significant improvement in the performance of algorithms predicting the need for social service referrals.1 It is important to be reminded that not all available data have utility for all purposes, and that many plausible hypotheses do not survive rigorous analysis. However, we would like to make sure that the informatics community does not interpret these findings more broadly to indicate that SDH data have no value. We would like to propose several possible explanations for these interesting and surprising findings, each of which might suggest future avenues of exploration. Correlations between predictors: It is known that SDH are correlated with the risk of many clinical conditions, and it seems possible that in this particular study, the SDH variables were strongly correlated with the clinical diagnoses. For example, social determinants (eg socioeconomic status, race, social support, etc.) are well established as strong predictors of cardiovascular disease.2 If, in the current study, the information contained in the SDH was already present in the clinical variables, then adding SDH would not improve model performance. Future work in different domains might show SDH to have more predictive power. For example, with certain congenital conditions, SDH might have little influence on the risk of disease but instead could be related to access to care or quality of life for people with the condition, and thus would contribute additional information to a model containing clinical diagnoses. In addition, it is likely that many of the SDH are correlated with each other. Community-level research, which often deals with collinear predictors, often addresses this problem by collapsing correlated measures into scales (such as the Centers for Disease Control’s Social Vulnerability Index [https://svi.cdc.gov]), which are then utilized as single metrics. Choice of outcome variables: A closely related potential explanation is that the specific social service referrals chosen in the current study were not well predicted by SDH variables because they were already too well predicted by the clinical ones. For example, referrals to dietitians (one of the study’s outcome variables) might be so strongly predicted by diagnosis of diabetes, hypertension, or congestive heart failure that additional data are not helpful. Similarly, referrals to mental health services might be sufficiently strongly predicted by the presence of psychiatric diagnoses. It is possible that SDH data may have predictive power for other sorts of healthcare utilization and health outcomes, just not for these particular social services. Variability in predictors: As the authors suggest in their discussion, it is possible that the patients of the safety net health system had insufficient variability in SDH. For example, if household income did not range much above or below the poverty line in the entire population of interest, this variable would provide little discriminatory power. It is possible that models built from a more socioeconomically diverse population data set might provide different results. Comprehensiveness of the clinical data: The authors had access to unusually comprehensive clinical data through the Indiana Network for Patient Care, a well-established health information exchange organization, which allowed the researchers to leverage data such as emergency and hospital admissions from other organizations. Healthcare organizations with access to only their own local clinical data may find that SDH variables contribute more to predictive models, precisely because the SDH data might serve as proxies for some of the missing clinical data. Ecological inferences: In the absence of individual-level SDH data, the authors used community-level (ZIP code and census tract level) measures as proxies. It is possible that for certain SDH variables, community estimates are either imprecise or biased. For example, if the within-community variance in education is extremely high, then individual educational attainment will not be predicted well by the neighborhood average, creating lack of precision. Alternately, if the patients who seek care at a safety net hospital tend to be less well educated than their close neighbors, then individual educational attainment will be systematically overestimated by the neighborhood average, creating bias.3 Future work might explore whether community-level data have more utility when they describe neighborhood characteristics (such as, in this study, data about local availability of well-lit walkways or grocery stores) rather than being used to infer individual-level characteristics (such as education). A useful family of methodological approaches to account for these ecological relationships is hierarchical or multilevel models, which explicitly account for the nested structure of the data (individuals within neighborhoods, in this case). Choice of predictor variables: In addition to the rich set of social and environmental factors used by Kasthurirathne and colleagues, it is possible that others not available to the researchers might have predictive power, such as social support and social capital, or detailed employment type.4 Inspecting variable importance measures5 in the models might suggest new hypotheses about which types of SDH have the most predictive utility, and whether these point to other SDH data to collect or obtain from other sources. Given the extensive public health literature on the relationship between social determinants and health, it is exciting to see the health informatics community begin to explore mergers of clinical and SDH data sets. We appreciate the contribution of Kasthurirathne et al. to this emerging literature and welcome additional exploration of the potential utility of social determinants of health in clinical care and predictive analytics. Conflict of interest statement. None declared. Jessica S. Ancker, Min-hyung Kim, Yiye Zhang, Yongkang Zhang 0004, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | A case study evaluating the portability of an executable computable phenotype algorithm across multiple institutions and electronic health record environmentsabstractElectronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges. Jennifer A. Pacheco, Luke V. Rasmussen, Richard C. Kiefer, Thomas R. Campion Jr., Peter Speltz, Robert J. Carroll, Sarah C. Stallings, Huan Mo, Monika Ahuja, Guoqian Jiang, Eric LaRose, Peggy L. Peissig, Ning Shang 0004, Barbara Benoit, Vivian S. Gainer, Kenneth Borthwick, Kathryn L. Jackson, Ambrish Sharma, Andy Yizhou Wu, Abel N. Kho, Dan M. Roden, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
J. Am. Medical Informatics Assoc. | 22 |
| 2018 | Association networks in a matched case-control design - Co-occurrence patterns of preexisting chronic medical conditions in patients with major depression versus their matched controls
Min-hyung Kim, Samprit Banerjee, Yize Zhao, Fei Wang 0001, Yiye Zhang, Yongjun Zhu 0001, Joseph DeFerio, Lauren Evans, Sang Min Park, Jyotishman Pathak |
J. Biomed. Informatics | 10 |
| 2017 | Leveraging Value Sets from the Value Set Authority Center (VSAC) in a Standards-Based Clinical Data Repository
Richard C. Kiefer, Luke V. Rasmussen, Jennifer A. Pacheco, Peter Speltz, Joshua C. Denny, William K. Thompson, Jyotishman Pathak, Guoqian Jiang |
AMIA | 7 |
| 2017 | A Framework for Data Quality Assessment in Clinical Research Datasets
Kathleen Lee, Nicole Gray Weiskopf, Jyotishman Pathak |
AMIA | 3 |
| 2017 | Secondary Use of Patients' Electronic Records (SUPER): An Approach for Meeting Specific Data Needs of Clinical and Translational Researchers
Evan Sholle, Joseph Kabariti, Stephen B. Johnson, John P. Leonard, Jyotishman Pathak, Vinay I. Varughese, Curtis L. Cole, Thomas R. Campion Jr. |
AMIA | 5 |
| 2017 | The Phenotype Execution and Modeling Architecture: A Roadmap Towards Next-generation Phenotyping Using EHRs
Peter Speltz, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, William K. Thompson, Guoqian Jiang, Jyotishman Pathak, Joshua C. Denny |
AMIA | 7 |
| 2017 | Semi-Supervised Approach to Monitoring Clinical Depressive Symptoms in Social MediaabstractWith the rise of social media, millions of people are routinely expressing their moods, feelings, and daily struggles with mental health issues on social media platforms like Twitter. Unlike traditional observational cohort studies conducted through questionnaires and self-reported surveys, we explore the reliable detection of clinical depression from tweets obtained unobtrusively. Based on the analysis of tweets crawled from users with self-reported depressive symptoms in their Twitter profiles, we demonstrate the potential for detecting clinical depression symptoms which emulate the PHQ-9 questionnaire clinicians use today. Our study uses a semi-supervised statistical model to evaluate how the duration of these symptoms and their expression on Twitter (in terms of word usage patterns and topical preferences) align with the medical findings reported via the PHQ-9. Our proactive and automatic screening tool is able to identify clinical depressive symptoms with an accuracy of 68% and precision of 72%. Amir Hossein Yazdavar, Hussein Al-Olimat, Monireh Ebrahimi, Goonmeet Bajaj, Tanvi Banerjee, Krishnaprasad Thirunarayan, Jyotishman Pathak, Amit P. Sheth |
ASONAM | 7 |
| 2017 | Temporal reflected logistic regression for probabilistic heart failure survival score predictionabstractHeart failure (HF) has a highly variable annual mortality rate and there is an urgent need of determining patient prognosis to enable informed decision-making about heart failure treatment strategies. Existing survival risk prediction models either require features that limit their applicability or pose difficulties for parameter estimation as physicians have to use a limited set of variables with known hazard ratios published in literature. We propose a new model to predict the probabilistic survival score after HF diagnosis based on all clinical variables derived from the electronic health record (EHR). We formalize the parameter estimation problem by using the maximum likelihood estimation (MLE) principle and devise an effective and efficient algorithm to solve the optimization problem. Experimental results using EHR data of 234 HF patients validate the superiority of this new model in predicting prognosis over the currently used Seattle Heart Failure Model. Mingjie Qian, Jyotishman Pathak, Naveen Pereira, ChengXiang Zhai |
BIBM | 2 |
| 2017 | Polyadic Regression and its Application to ChemogenomicsabstractWe study the problem of Polyadic Prediction, where the input consists of an ordered tuple of objects, and the goal is to predict a measurement associated with them. Many tasks can be naturally framed as Polyadic Prediction problems. In drug discovery, for instance, it is important to estimate the treatment effect of a drug on various tissue-specific diseases, as it is expressed over the available genes. Thus, we essentially predict the expression value measurements for several (drug, gene, tissue) triads. To tackle Polyadic Prediction problems, we propose a general framework, called Polyadic Regression, predicting measurements associated with multiple objects. Our framework is inductive, in the sense of enabling predictions for new objects, unseen during training. Our model is expressive, exploring high-order, polyadic interactions in an efficient manner. An alternating Proximal Gradient Descent procedure is proposed to fit our model. We perform an extensive evaluation using real-world chemogenomics data, where we illustrate the superior performance of Polyadic Regression over the prior art. Our method achieves an increase of 0.06 and 0.1 in Spearman correlation between the predicted and the actual measurement vectors, for predicting missing polyadic data and predicting polyadic data for new drugs, respectively. Ioakeim Perros, Fei Wang 0001, Ping Zhang 0016, Peter B. Walker, Richard W. Vuduc, Jyotishman Pathak, Jimeng Sun 0001 |
SDM | 6 |
| 2016 | An NLP Extension to the Quality Data Model for EHR-Driven Phenotype Algorithm Authoring and Execution
Guoqian Jiang, William K. Thompson, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, Huan Mo, Peter Speltz, Joshua C. Denny, Jyotishman Pathak |
AMIA | 9 |
| 2016 | Improving risk prediction for depression via Elastic Net regression - Results from Korea National Health Insurance Services Data
Min-hyung Kim, Samprit Banerjee, Sang Min Park, Jyotishman Pathak |
AMIA | 4 |
| 2016 | Predicting Prolonged Stay in the ICU Attributable to Bleeding in Patients Offered Plasma Transfusion
Che Ngufor, Dennis Murphree, Sudhindra Upadhyaya, Nageswar Madde, Jyotishman Pathak, Rickey E. Carter, Daryl J. Kor |
AMIA | 5 |
| 2016 | Standardized Representation of Clinical Study Data Dictionaries with CIMI Archetypes
Deepak K. Sharma, Harold R. Solbrig, Eric Prud'hommeaux, Jyotishman Pathak, Guoqian Jiang |
AMIA | 4 |
| 2016 | Clinical phenotyping in selected national networks: demonstrating the need for high-throughput, portable, and computational methods
Rachel L. Richesson, Jimeng Sun 0001, Jyotishman Pathak, Abel N. Kho, Joshua C. Denny |
Artif. Intell. Medicine | 3 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 9 |
| 2016 | Developing a data element repository to support EHR-driven phenotype algorithm authoring and execution
Guoqian Jiang, Richard C. Kiefer, Luke V. Rasmussen, Harold R. Solbrig, Huan Mo, Jennifer A. Pacheco, Jie Xu 0011, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
J. Biomed. Informatics | 12 |
| 2016 | Developing EHR-driven heart failure risk prediction models using CPXR(Log) with the probabilistic loss function
Vahid Taslimitehrani, Guozhu Dong, Naveen Pereira, Maryam Panahiazar, Jyotishman Pathak |
J. Biomed. Informatics | 5 |
| 2015 | A Heterogeneous Multi-Task Learning for Predicting RBC Transfusion and Perioperative Outcomes
Che Ngufor, Sudhindra Upadhyaya, Dennis Murphree, Nageswar Madde, Daryl J. Kor, Jyotishman Pathak |
AIME | 6 |
| 2015 | Developing Intermediary Medication Phenotypes via Metabolomics Biotechnology Using an Electronic Health Record-Linked Biobank
Matthew K. Breitenstein, Richard M. Weinshilboum, Jyotishman Pathak, Liewei Wang |
AMIA | 3 |
| 2015 | Harmonization of Quality Data Model with HL7 FHIR to Support EHR-driven Phenotype Authoring and Execution: A Pilot Study
Guoqian Jiang, Harold R. Solbrig, Richard C. Kiefer, Luke V. Rasmussen, Huan Mo, Jennifer A. Pacheco, Enid N. H. Montague, Jie Xu 0011, Peter Speltz, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
AMIA | 13 |
| 2015 | A Genome- and Phenome- Wide Study of Diverticulosis
Yoonjung Y. Joo, Jennifer A. Pacheco, Loren L. Armstrong, William K. Thompson, Robert J. Carroll, Joshua C. Denny, Peggy L. Peissig, James G. Linneman, Jyotishman Pathak, Girish N. Nadkarni, Laura Rasmussen-Torvik, M. Geoffrey Hayes, Abel N. Kho |
AMIA | 9 |
| 2015 | Translating Electronic Clinical Quality Measures to Executable, Portable, and Customizable Workflows in KNIME
Huan Mo, Jennifer A. Pacheco, Richard C. Kiefer, Luke V. Rasmussen, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
AMIA | 5 |
| 2015 | Usability of a phenotype builder prototype and lessons learned for the design of phenotyping tools
Enid N. H. Montague, Jie Xu 0011, Luke V. Rasmussen, Joshua C. Denny, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Peter Speltz, William K. Thompson, Jyotishman Pathak |
AMIA | 10 |
| 2015 | (Authoring) Rules, (Distributed Query) Tools, and Drools: The challenging new world of high throughput phenotyping
Jennifer A. Pacheco, Abel N. Kho, Jyotishman Pathak, Joshua C. Denny, Shawn N. Murphy |
AMIA | 3 |
| 2015 | PhEMA: Phenotype Modeling, Sharing and Execution Architecture
Jyotishman Pathak, Joshua C. Denny, William K. Thompson, Luke V. Rasmussen |
AMIA | 1 |
| 2015 | Multi-task learning with selective cross-task transfer for predicting bleeding and other important patient outcomesabstractIn blood transfusion studies, its is often desirable before a surgical procedure to estimate the likelihood of a patient bleeding, need for blood products, re-operation due to bleeding and other important patient outcomes. Such prediction rules are crucial in allowing for optimal planning, more efficient use of blood bank resources, and identification of high-risk patient cohort for specific perioperative interventions. The goal of this study is to present a simple and efficient algorithm that could estimate the risk of multiple outcomes simultaneously. Specifically, a heterogeneous multi-task learning method is presented for learning important surgical outcomes such as bleeding, intraoperative RBC transfusion, need for ICU care, length of stay and mortality. To improve the performance of the method, a post-learning strategy is implemented to further learn the relationship between the trained tasks by a simple “goodness of fit” measure. Specifically, two tasks are considered similar if the model parameters of one tasks improves predictive performance of the other. This strategy allows tasks to be grouped in clusters where selective cross-task transfer of knowledge is explicitly encouraged. To further improve prediction accuracy, a number of operative measurements or surgical outcomes whose predictions are not of direct interest are incorporated in the multi-task model as supplementary tasks to donate information and help the performance of relevant tasks. Results for predicting bleeding and need for blood transfusion for patients undergoing non-cardiac operations from an institutional transfusion datamart show that the proposed methods can improve prediction accuracy over standard single-tasks learning methods. Additional experiments on a real public available data set show that the method is accurate and competitive with some existing methods in the literature. Che Ngufor, Sudhindra Upadhyaya, Dennis Murphree, Daryl J. Kor, Jyotishman Pathak |
DSAA | 5 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 19 |
| 2015 | Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational researchabstractOBJECTIVE: To review and evaluate available software tools for electronic health record-driven phenotype authoring in order to identify gaps and needs for future development. MATERIALS AND METHODS: Candidate phenotype authoring tools were identified through (1) literature search in four publication databases (PubMed, Embase, Web of Science, and Scopus) and (2) a web search. A collection of tools was compiled and reviewed after the searches. A survey was designed and distributed to the developers of the reviewed tools to discover their functionalities and features. RESULTS: Twenty-four different phenotype authoring tools were identified and reviewed. Developers of 16 of these identified tools completed the evaluation survey (67% response rate). The surveyed tools showed commonalities but also varied in their capabilities in algorithm representation, logic functions, data support and software extensibility, search functions, user interface, and data outputs. DISCUSSION: Positive trends identified in the evaluation included: algorithms can be represented in both computable and human readable formats; and most tools offer a web interface for easy access. However, issues were also identified: many tools were lacking advanced logic functions for authoring complex algorithms; the ability to construct queries that leveraged un-structured data was not widely implemented; and many tools had limited support for plug-ins or external analytic software. CONCLUSIONS: Existing phenotype authoring tools could enable clinical researchers to work with electronic health record data more efficiently, but gaps still exist in terms of the functionalities of such tools. The present work can serve as a reference point for the future development of similar tools. Jie Xu 0011, Luke V. Rasmussen, Pamela L. Shaw, Guoqian Jiang, Richard C. Kiefer, Huan Mo, Jennifer A. Pacheco, Peter Speltz, Qian Zhu 0003, Joshua C. Denny, Jyotishman Pathak, William K. Thompson, Enid N. H. Montague |
J. Am. Medical Informatics Assoc. | 11 |
| 2014 | Analysis of Online Information Searching for Cardiovascular Diseases on a Consumer Health Information Portal
Ashutosh Jadhav, Amit P. Sheth, Jyotishman Pathak |
AMIA | 3 |
| 2014 | An Analysis of Mayo Clinic Search Query Logs for Cardiovascular Diseases
Ashutosh Jadhav, Amit P. Sheth, Jyotishman Pathak |
AMIA | 3 |
| 2014 | Using PhenotypePortal for Checking Clinical Guideline Recommendation Compliance
Lara Johnstun, Danielle Groat, Amol Bhalla, Kevin J. Peterson, Jyotishman Pathak, María Adela Grando |
AMIA | 5 |
| 2014 | Scalable and High-Throughput Execution of Clinical Quality Measures from Electronic Health Records using MapReduce and the JBoss(R) Drools Engine
Kevin J. Peterson, Jyotishman Pathak |
AMIA | 2 |
| 2014 | Evaluation of Existing Phenotype Authoring Tools for Clinical Research
Luke V. Rasmussen, Jie Xu 0011, Ruijue Liu, Qian Zhu 0003, Jennifer A. Pacheco, Jyotishman Pathak, William K. Thompson, Joshua C. Denny, Huan Mo, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague |
AMIA | 6 |
| 2014 | Qualitative evaluation of three phenotype information models to find methotrexate liver injury
Qian Zhu 0003, Huan Mo, Luke V. Rasmussen, Andrew R. Post, Jennifer A. Pacheco, Jie Xu 0011, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Jyotishman Pathak |
AMIA | 12 |
| 2014 | Empowering personalized medicine with big data and semantic web technology: Promises, challenges, and use casesabstractIn healthcare, big data tools and technologies have the potential to create significant value by improving outcomes while lowering costs for each individual patient. Diagnostic images, genetic test results and biometric information are increasingly generated and stored in electronic health records presenting us with challenges in data that is by nature high volume, variety and velocity, thereby necessitating novel ways to store, manage and process big data. This presents an urgent need to develop new, scalable and expandable big data infrastructure and analytical methods that can enable healthcare providers access knowledge for the individual patient, yielding better decisions and outcomes. In this paper, we briefly discuss the nature of big data and the role of semantic web and data analysis for generating "smart data" which offer actionable information that supports better decision for personalized medicine. In our view, the biggest challenge is to create a system that makes big data robust and smart for healthcare providers and patients that can lead to more effective clinical decision-making, improved health outcomes, and ultimately, managing the healthcare costs. We highlight some of the challenges in using big data and propose the need for a semantic data-driven environment to address them. We illustrate our vision with practical use cases, and discuss a path for empowering personalized medicine using big data and semantic web technology. Maryam Panahiazar, Vahid Taslimitehrani, Ashutosh Jadhav, Jyotishman Pathak |
IEEE BigData | 4 |
| 2014 | Research and applications: An electronic health record driven algorithm to identify incident antidepressant medication usersabstractOBJECTIVE: We validated an algorithm designed to identify new or prevalent users of antidepressant medications via population-based drug prescription records. PATIENTS AND METHODS: We obtained population-based drug prescription records for the entire Olmsted County, Minnesota, population from 2011 to 2012 (N=149,629) using the existing electronic medical records linkage infrastructure of the Rochester Epidemiology Project (REP). We selected electronically a random sample of 200 new antidepressant users stratified by age and sex. The algorithm required the exclusion of antidepressant use in the 6 months preceding the date of the first qualifying antidepressant prescription (index date). Medical records were manually reviewed and adjudicated to calculate the positive predictive value (PPV). We also manually reviewed the records of a random sample of 200 antihistamine users who did not meet the case definition of new antidepressant user to estimate the negative predictive value (NPV). RESULTS: 161 of the 198 subjects electronically identified as new antidepressant users were confirmed by manual record review (PPV 81.3%). Restricting the definition of new users to subjects who were prescribed typical starting doses of each agent for treating major depression in non-geriatric adults resulted in an increase in the PPV (90.9%). Extending the time windows with no antidepressant use preceding the index date resulted in only modest increases in PPV. The manual abstraction of medical records of 200 antihistamine users yielded an NPV of 98.5%. CONCLUSIONS: Our study confirms that REP prescription records can be used to identify prevalent and incident users of antidepressants in the Olmsted County, Minnesota, population. William V. Bobo, Jyotishman Pathak, Hilal M. Kremers, Barbara P. Yawn, Scott M. Brue, Cynthia J. Stoppel, Paul E. Croarkin, Jennifer L. St. Sauver, Mark A. Frye, Walter A. Rocca |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Design patterns for the development of electronic health record-driven phenotype extraction algorithms
Luke V. Rasmussen, William K. Thompson, Jennifer A. Pacheco, Abel N. Kho, David Carrell, Jyotishman Pathak, Peggy L. Peissig, Gerard Tromp, Joshua C. Denny, Justin Starren |
J. Biomed. Informatics | 6 |
| 2013 | PhenotypePortal: An Open-Source Library and Platform for Authoring, Executing and Visualization of Electronic Health Records Driven Phenotyping Algorithms
Jyotishman Pathak, Cory M. Endle, Dale Suesse, Kevin J. Peterson, Craig Stancle, Dingcheng Li, Christopher G. Chute |
AMIA | 1 |
| 2013 | Mining Drug-Drug Interaction Patterns from Linked Clinical Data
Jyotishman Pathak, Richard C. Kiefer, Christopher G. Chute |
AMIA | 1 |
| 2013 | Using standardized clinical data modeling and knowledge representation to compute pharmacogenomic data elements
Qian Zhu 0003, Jyotishman Pathak, Robert R. Freimuth, Christopher G. Chute |
AMIA | 2 |
| 2013 | Mining drug-drug interaction patterns from linked data: A case study for Warfarin, Clopidogrel, and SimvastatinabstractBy nature, healthcare data is highly complex and voluminous. While on one hand, it provides unprecedented opportunities to identify hidden and unknown relationships between patients and treatment outcomes, or drugs and allergic reactions for given individuals, representing and querying large network datasets poses significant technical challenges. In this research, we study the use of Semantic Web and Linked Data technologies for identifying potential drug-drug interaction (DDI) information from publicly available resources, and determining if such interactions were observed using real patient data. Specifically, we apply Linked Data principles and technologies for representing patient data from electronic health records (EHRs) at Mayo Clinic as Resource Description Framework (RDF) graphs, and identify potential DDIs for three widely prescribed cardiovascular drugs: Warfarin, Clopidogrel and Simvastatin. Our results from the proof-of-concept study demonstrate the potential of applying such a methodology to study patient health outcomes as well as enabling genome-guided drug therapies and treatment interventions. Jyotishman Pathak, Richard C. Kiefer, Christopher G. Chute |
BIBM | 1 |
| 2013 | A semantic-web oriented representation of the clinical element model for secondary use of electronic health records dataabstractThe clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data. Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 7 |
| 2013 | Terminology representation guidelines for biomedical ontologies in the semantic web notations
Cui Tao, Jyotishman Pathak, Harold R. Solbrig, Wei-Qi Wei, Christopher G. Chute |
J. Biomed. Informatics | 2 |
| 2013 | Harmonization and semantic annotation of data dictionaries from the Pharmacogenomics Research Network: A case study
Qian Zhu 0003, Robert R. Freimuth, Zonghui Lian, Scott Bauer, Jyotishman Pathak, Cui Tao, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 5 |
| 2013 | Disambiguation of PharmGKB drug-disease relations with NDF-RT and SPL
Qian Zhu 0003, Robert R. Freimuth, Jyotishman Pathak, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 3 |
| 2012 | Using Electronic Health Records to Identify Heart Failure Cohorts with Differentiation for Preserved and Reduced Ejection Fraction
Suzette J. Bielinski, Jyotishman Pathak, Sunghwan Sohn, Gail P. Jarvik, David Carrell, Naveen Pereira, Véronique L. Roger |
AMIA | 2 |
| 2012 | Using PheWAS to Assess Pleiotropy of Genetic Risk Scores for Rheumatoid Arthritis and Coronary Artery Disease in the eMERGE Network
Robert J. Carroll, Katherine P. Liao, Anne E. Eyler, Lisa Bastarache, Dana C. Crawford, Peggy L. Peissig, Jyotishman Pathak, David Carrell, Abel N. Kho, Rongling Li, Daniel R. Masys, Gail P. Jarvik, Christopher G. Chute, Rex L. Chisholm, Eric B. Larson, Catherine A. McCarty, Iftikhar J. Kullo |
AMIA | 7 |
| 2012 | Visualization and Reporting of Results for Electronic Health Records Driven Phenotyping using the Open-Source popHealth Platform
Cory M. Endle, Sahana Murthy, Dale Suesse, Craig Stancle, Dingcheng Li, Lacey Hart, Christopher G. Chute, Jyotishman Pathak |
AMIA | 8 |
| 2012 | Modeling and Executing Electronic Health Records Driven Phenotyping Algorithms using the NQF Quality Data Model and JBoss® Drools Engine
Dingcheng Li, Sahana Murthy, Davide Sottara, Christopher G. Chute, Stanley M. Huff, Jyotishman Pathak, Cory M. Endle, Dale Suesse, Craig Stancle |
AMIA | 6 |
| 2012 | Using Electronic Health Records to Identify Patient Cohorts for Drug-Induced Thrombocytopenia, Neutropenia and Liver Injury
Jyotishman Pathak, Aref Al-Kali, Jayant Talwalkar, Abel N. Kho, Joshua C. Denny, Sean P. Murphy, Kevin Bruce, Matthew J. Durski, Christopher G. Chute |
AMIA | 1 |
| 2012 | Mining the Human Phenome using Semantic Web Technologies: A Case Study for Type 2 Diabetes
Jyotishman Pathak, Richard C. Kiefer, Suzette J. Bielinski, Christopher G. Chute |
AMIA | 1 |
| 2012 | Mining Genotype-Phenotype Associations from Electronic Health Records and Biorepositories using Semantic Web Technologies
Jyotishman Pathak, Richard C. Kiefer, Robert R. Freimuth, Suzette J. Bielinski, Christopher G. Chute |
AMIA | 1 |
| 2012 | A Distributed Semantic Web Approach for Cohort Identification
Joseph Teagno, Richard C. Kiefer, Jyotishman Pathak, G. Q. Zhang, Satya Sanket Sahoo |
AMIA | 3 |
| 2012 | An Evaluation of the NQF Quality Data Model for Representing Electronic Health Record Driven Phenotyping Algorithms
William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Peggy L. Peissig, Joshua C. Denny, Abel N. Kho, Aaron W. Miller, Jyotishman Pathak |
AMIA | 8 |
| 2012 | Importance of multi-modal approaches to effectively identify cataract cases from electronic health recordsabstractOBJECTIVE: There is increasing interest in using electronic health records (EHRs) to identify subjects for genomic association studies, due in part to the availability of large amounts of clinical data and the expected cost efficiencies of subject identification. We describe the construction and validation of an EHR-based algorithm to identify subjects with age-related cataracts. MATERIALS AND METHODS: We used a multi-modal strategy consisting of structured database querying, natural language processing on free-text documents, and optical character recognition on scanned clinical images to identify cataract subjects and related cataract attributes. Extensive validation on 3657 subjects compared the multi-modal results to manual chart review. The algorithm was also implemented at participating electronic MEdical Records and GEnomics (eMERGE) institutions. RESULTS: An EHR-based cataract phenotyping algorithm was successfully developed and validated, resulting in positive predictive values (PPVs) >95%. The multi-modal approach increased the identification of cataract subject attributes by a factor of three compared to single-mode approaches while maintaining high PPV. Components of the cataract algorithm were successfully deployed at three other institutions with similar accuracy. DISCUSSION: A multi-modal strategy incorporating optical character recognition and natural language processing may increase the number of cases identified while maintaining similar PPVs. Such algorithms, however, require that the needed information be embedded within clinical documents. CONCLUSION: We have demonstrated that algorithms to identify and characterize cataracts can be developed utilizing data collected via the EHR. These algorithms provide a high level of accuracy even when implemented across multiple EHRs and institutional boundaries. Peggy L. Peissig, Luke V. Rasmussen, Richard L. Berg, James G. Linneman, Catherine A. McCarty, Carol Waudby, Joshua C. Denny, Russell A. Wilke, Jyotishman Pathak, David Carrell, Abel N. Kho, Justin Starren |
J. Am. Medical Informatics Assoc. | 10 |
| 2012 | Building a robust, scalable and standards-driven infrastructure for secondary use of EHR data: The SHARPn project
Susan Rea, Jyotishman Pathak, Guergana K. Savova, Thomas A. Oniki, Les Westberg, Calvin E. Beebe, Cui Tao, Craig G. Parker, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
J. Biomed. Informatics | 2 |
| 2011 | Letter: Further revamping VA's NDF-RT drug terminology for clinical researchabstractBiomedical terminology and vocabulary standards (for the purposes of this correspondence, we use the terms ‘terminology,’ ‘vocabulary’ and ‘ontology’ interchangeably.) play an important role in enabling consistent, comparable, and meaningful sharing of data within and across institutional boundaries, as well as ensuring semantic interoperability. An important domain for developing standardized vocabularies is medications, where existing standards structure and organize approved drug products and ingredients by various characteristics or properties to support a multitude of clinical and epidemiological research questions across the spectrum of health and disease. Veteran Affairs' National Drug File-Reference Terminology (NDF-RT; see figure 1) is a Federal Medication-recommended standardized terminology resource encompassing medications, ingredients, and high-level drug classes for Chemical Structure (eg, Acetanilides), Mechanism of Action (eg, Prostaglandin Receptor Antagonists), Physiological Effect (eg, Decreased Prostaglandin Production), drug–disease relationship describing the Therapeutic Intent (eg, Pain), and Pharmacokinetics describing the mechanisms of absorption and distribution of an administered drug within a body (eg, Hepatic Metabolism). Additionally, NDF-RT contains two independent lists of drug classes: Legacy VA classes and External Pharmacologic classes, where the former simply provides a shallow hierarchy of ‘clinically oriented’ classes (eg, β blockers), and the latter focuses on classifying drugs based on their chemical properties and functional groups (eg, H1 Receptor Antagonist). National Drug File-Reference Terminology drug-class hierarchy organization (adapted from Carter et al4). In the recent past, several research reports have highlighted important issues and challenges in using NDF-RT for clinical research and interoperability. Bodenreider et al1 focused on determining anticoagulation status of patients based on a list of medications prescribed using NDF-RT as the underlying drug-class terminology. In particular, this work concentrated on leveraging description logics (DL)-based representation of NDF-RT to infer additional information about drug-class relationships using the Legacy VA classes and External Pharmacologic classes. During this process, the authors not only had to make significant modifications and re-engineering to NDF-RT's DL representation, but also encountered several missing pieces of drug-class membership information in NDF-RT. In another study, Palchuk et al2 constructed a hierarchy of NDF-RT drug classes with drug and medication information from another standardized drug terminology, RxNorm, using data from the patient's electronic medical record. Similar to Bodenreider et al,1 here authors had to perform significant re-engineering to map, and subsequently classify, RxNorm drug products using Legacy VA classes from NDF-RT. The authors found this process to be extremely onerous, and proposed the evolution of RxNorm toward an interface terminology with hierarchical and categorical organization. In our own work published in JAMIA,3 we investigated similar issues in mapping and classifying drugs and medication products from RxNorm using NDF-RT's multiaxial classification. We found several issues where the mappings were incomplete and, in many occasions, semantically and clinically inconsistent. Based on these recent findings, it has become abundantly evident that to leverage NDF-RT continually for clinical and epidemiological research, it is vital to address the existing issues. In particular, we highlight the following issues for consideration: Alignment with RxNorm: Palchuk et al2 and our previous work3 illustrated several problematic examples where the relationships between NDF-RT and RxNorm drug concept entities were either missing or misrepresented due to curation problems. Given that RxNorm does not currently provide hierarchical classification of drug products, the mappings between drug concepts in RxNorm and the corresponding classes in NDF-RT are vital for research projects that use RxNorm for coding their medication data. Missing and inconsistent mappings can lead to incorrect conclusions. Relationships between drug products and Legacy VA classes: The Legacy VA classes that were derived from the VA-NDF,4 while deprecated, still continue to provide significant value and merit with respect to ‘clinically relevant’ drug classification. However, in its current formalism, NDF-RT only allows assignment of a single Legacy VA class to a particular drug product. For example, even though it is clinically appropriate to classify a drug as both an antihypertensive and a β-blocker, in reality a majority of drug products in NDF-RT are assigned a single Legacy VA class. We believe that this limitation needs to be addressed, since the Legacy VA classes, although derived from the legacy VA-NDF, have significant clinical implications for drug classifications. Relationships between drug products and External Pharmacologic classes: As illustrated by Bodenreider et al,1 NDF-RT in its current form does not contain any relationships between ingredients and External Pharmacologic classes. As an example, in NDF-RT, Clopidogrel and Platelet Aggregation Inhibitor are not related via any relationship—either direct or indirect. Arguably, this is a significant limitation and has implications with respect to drug classifications and querying. Metadata annotations: Finally, we believe that NDF-RT should follow best practices for vocabulary and terminology development.5 In particular, what is notably missing from NDF-RT are appropriate metadata annotations for different drug, ingredient, and drug-class entities. As an example, the Legacy VA class ‘Loop Diuretics’ and External Pharmacologic class ‘Loop Diuretic’ are distinguished only by a slight difference in the label name, without additional annotation indicating their differences, similarities, etc. Consequently, someone unfamiliar with NDF-RT multiaxial classification runs the risk of using the incorrect classification for her application. In summary, we hope that, via this correspondence, we have highlighted some of the important issues with NDF-RT, which arguably is emerging as one of the most important standardized public drug-classification terminologies. Our expectation is that addressing the above problems in future NDF-RT releases will significantly benefit the clinical research informatics community. This work was supported by National Human Genome Research Institute. None. Not commissioned; externally peer reviewed. Jyotishman Pathak, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Mapping clinical phenotype data elements to standardized metadata repositories and controlled terminologies: the eMERGE Network experienceabstractBACKGROUND: Systematic study of clinical phenotypes is important for a better understanding of the genetic basis of human diseases and more effective gene-based disease management. A key aspect in facilitating such studies requires standardized representation of the phenotype data using common data elements (CDEs) and controlled biomedical vocabularies. In this study, the authors analyzed how a limited subset of phenotypic data is amenable to common definition and standardized collection, as well as how their adoption in large-scale epidemiological and genome-wide studies can significantly facilitate cross-study analysis. METHODS: The authors mapped phenotype data dictionaries from five different eMERGE (Electronic Medical Records and Genomics) Network sites studying multiple diseases such as peripheral arterial disease and type 2 diabetes. For mapping, standardized terminological and metadata repository resources, such as the caDSR (Cancer Data Standards Registry and Repository) and SNOMED CT (Systematized Nomenclature of Medicine), were used. The mapping process comprised both lexical (via searching for relevant pre-coordinated concepts and data elements) and semantic (via post-coordination) techniques. Where feasible, new data elements were curated to enhance the coverage during mapping. A web-based application was also developed to uniformly represent and query the mapped data elements from different eMERGE studies. RESULTS: Approximately 60% of the target data elements (95 out of 157) could be mapped using simple lexical analysis techniques on pre-coordinated terms and concepts before any additional curation of terminology and metadata resources was initiated by eMERGE investigators. After curation of 54 new caDSR CDEs and nine new NCI thesaurus concepts and using post-coordination, the authors were able to map the remaining 40% of data elements to caDSR and SNOMED CT. A web-based tool was also implemented to assist in semi-automatic mapping of data elements. CONCLUSION: This study emphasizes the requirement for standardized representation of clinical research data using existing metadata and terminology resources and provides simple techniques and software for data element mapping using experiences from the eMERGE Network. Jyotishman Pathak, Janey Wang, Sudha Kashyap, Melissa A. Basford, Rongling Li, Daniel R. Masys, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Leveraging informatics for genetic studies: use of the electronic medical record to enable a genome-wide association study of peripheral arterial diseaseabstractBACKGROUND: There is significant interest in leveraging the electronic medical record (EMR) to conduct genome-wide association studies (GWAS). METHODS: A biorepository of DNA and plasma was created by recruiting patients referred for non-invasive lower extremity arterial evaluation or stress ECG. Peripheral arterial disease (PAD) was defined as a resting/post-exercise ankle-brachial index (ABI) less than or equal to 0.9, a history of lower extremity revascularization, or having poorly compressible leg arteries. Controls were patients without evidence of PAD. Demographic data and laboratory values were extracted from the EMR. Medication use and smoking status were established by natural language processing of clinical notes. Other risk factors and comorbidities were ascertained based on ICD-9-CM codes, medication use and laboratory data. RESULTS: Of 1802 patients with an abnormal ABI, 115 had non-atherosclerotic vascular disease such as vasculitis, Buerger's disease, trauma and embolism (phenocopies) based on ICD-9-CM diagnosis codes and were excluded. The PAD cases (66+/-11 years, 64% men) were older than controls (61+/-8 years, 60% men) but had similar geographical distribution and ethnic composition. Among PAD cases, 1444 (85.6%) had an abnormal ABI, 233 (13.8%) had poorly compressible arteries and 10 (0.6%) had a history of lower extremity revascularization. In a random sample of 95 cases and 100 controls, risk factors and comorbidities ascertained from EMR-based algorithms had good concordance compared with manual record review; the precision ranged from 67% to 100% and recall from 84% to 100%. CONCLUSION: This study demonstrates use of the EMR to ascertain phenocopies, phenotype heterogeneity and relevant covariates to enable a GWAS of PAD. Biorepositories linked to EMR may provide a relatively efficient means of conducting GWAS. Iftikhar J. Kullo, Jyotishman Pathak, Guergana K. Savova, Zeenat Ali, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Analyzing categorical information in two publicly available drug terminologies: RxNorm and NDF-RTabstractBACKGROUND: The RxNorm and NDF-RT (National Drug File Reference Terminology) are a suite of terminology standards for clinical drugs designated for use in the US federal government systems for electronic exchange of clinical health information. Analyzing how different drug products described in these terminologies are categorized into drug classes will help in their better organization and classification of pharmaceutical information. METHODS: Mappings between drug products in RxNorm and NDF-RT drug classes were extracted. Mappings were also extracted between drug products in RxNorm to five high-level NDF-RT categories: Chemical Structure; cellular or subcellular Mechanism of Action; organ-level or system-level Physiologic Effect; Therapeutic Intent; and Pharmacokinetics. Coverage for the mappings and the gaps were evaluated and analyzed algorithmically. RESULTS: Approximately 54% of RxNorm drug products (Semantic Clinical Drugs) were found not to have a correspondence in NDF-RT. Similarly, approximately 45% of drug products in NDF-RT are missing from RxNorm, most of which can be attributed to differences in dosage, strength, and route form. Approximately 81% of Chemical Structure classes, 42% of Mechanism of Action classes, 75% of Physiologic Effect classes, 76% of Therapeutic Intent classes, and 88% of Pharmacokinetics classes were also found not to have any RxNorm drug products classified under them. Finally, various issues regarding inconsistent mappings between drug concepts were identified in both terminologies. CONCLUSION: This investigation identified potential limitations of the existing classification systems and various issues in specification of correspondences between the concepts in RxNorm and NDF-RT. These proposals and methods provide the preliminary steps in addressing some of the requirements. Jyotishman Pathak, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Comparing and evaluating terminology services application programming interfaces: RxNav, UMLSKS and LexBIGabstractTo facilitate the integration of terminologies into applications, various terminology services application programming interfaces (API) have been developed in the recent past. In this study, three publicly available terminology services API, RxNav, UMLSKS and LexBIG, are compared and functionally evaluated with respect to the retrieval of information from one biomedical terminology, RxNorm, to which all three services provide access. A list of queries is established covering a wide spectrum of terminology services functionalities such as finding RxNorm concepts by their name, or navigating different types of relationships. Test data were generated from the RxNorm dataset to evaluate the implementation of the functionalities in the three API. The results revealed issues with various aspects of the API implementation (eg, handling of obsolete terms by LexBIG) and documentation (eg, navigational paths used in RxNav) that were subsequently addressed by the development teams of the three API investigated. Knowledge about such discrepancies helps inform the choice of an API for a given use case. Jyotishman Pathak, Lee B. Peters, Christopher G. Chute, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Implementation Brief: LexGrid: A Framework for Representing, Storing, and Querying Biomedical Terminologies from Simple to SublimeabstractMany biomedical terminologies, classifications, and ontological resources such as the NCI Thesaurus (NCIT), International Classification of Diseases (ICD), Systematized Nomenclature of Medicine (SNOMED), Current Procedural Terminology (CPT), and Gene Ontology (GO) have been developed and used to build a variety of IT applications in biology, biomedicine, and health care settings. However, virtually all these resources involve incompatible formats, are based on different modeling languages, and lack appropriate tooling and programming interfaces (APIs) that hinder their wide-scale adoption and usage in a variety of application contexts. The Lexical Grid (LexGrid) project introduced in this paper is an ongoing community-driven initiative, coordinated by the Mayo Clinic Division of Biomedical Statistics and Informatics, designed to bridge this gap using a common terminology model called the LexGrid model. The key aspect of the model is to accommodate multiple vocabulary and ontology distribution formats and support of multiple data stores for federated vocabulary distribution. The model provides a foundation for building consistent and standardized APIs to access multiple vocabularies that support lexical search queries, hierarchy navigation, and a rich set of features such as recursive subsumption (e.g., get all the children of the concept penicillin). Existing LexGrid implementations include the LexBIG API as well as a reference implementation of the HL7 Common Terminology Services (CTS) specification providing programmatic access via Java, Web, and Grid services. Jyotishman Pathak, Harold R. Solbrig, James D. Buntrock, Thomas M. Johnson, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Formalizing ICD coding rules using Formal Concept Analysis
Guoqian Jiang, Jyotishman Pathak, Christopher G. Chute |
J. Biomed. Informatics | 2 |
| 2008 | LexValueSets: An Approach for Context-Driven Value Sets Extraction
Jyotishman Pathak, Guoqian Jiang, Sridhar O. Dwarkanath, James D. Buntrock, Christopher G. Chute |
AMIA | 1 |
| 2007 | On Context-Specific Substitutability of Web ServicesabstractWeb service substitution refers to the problem of identifying a service that can replace another service in the context of a composition with a specified functionality. Existing solutions to this problem rely on detecting the functional and behavioral equivalence of a particular service to be replaced and candidate services that could replace it. We introduce the notion of context-specific substitutability, where context refers to the overall functionality of the composition that is required to be maintained after replacement of its constituents. Using the context information, we investigate two variants of the substitution problem, namely environment-independent and environment- dependent, where environment refers to the constituents of a composition and show how the substitutability criteria can be relaxed within this model. We provide a logical formulation of the resulting criteria based on model checking techniques as well as prove the soundness and completeness of the proposed approach. Jyotishman Pathak, Samik Basu 0001, Vasant G. Honavar |
ICWS | 1 |
| 2006 | Learning Classifiers from Distributed, Ontology-Extended Data Sources
Doina Caragea, Jun Zhang 0002, Jyotishman Pathak, Vasant G. Honavar |
DaWaK | 3 |
| 2006 | Modeling Web Services by Iterative Reformulation of Functional and Non-functional Requirements
Jyotishman Pathak, Samik Basu 0001, Vasant G. Honavar |
ICSOC | 1 |
| 2006 | Selecting and Composing Web Services through Iterative Reformulation of Functional SpecificationsabstractWe propose a specification-driven approach to Web service composition. The proposed framework allows users to start with a high-level, possibly incomplete specification of a desired (goal) service that is to be realized using a subset of the available component services. These services are represented by the system using transition systems augmented with guards over variables with infinite domains and are used to determine a strategy for their composition that would realize the goal service. In the event that the goal service cannot be realized using the available services, the system identifies the cause(s) for such failure which can then be used by the developer to reformulate the goal specification. Thus, the system supports Web service composition through iterative refinement of the functional specifications. We present a prototype implementation in tabled-logic programming environment that illustrates the key features of the proposed approach Jyotishman Pathak, Samik Basu 0001, Robyn R. Lutz, Vasant G. Honavar |
ICTAI | 1 |
| 2005 | Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar |
ALT | 4 |
| 2005 | Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar |
Discovery Science | 4 |