Indra Neil Sarkar

dblp:46/2509 · DBLP profile ↗
← Back
74ranked-venue papers
19as first author
7since 2021 · last 2025
0000-0003-2054-7356ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 74 · 19 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Learning health system linchpins: information exchange and a common data model
abstract
OBJECTIVE: To demonstrate the potential for a centrally managed health information exchange standardized to a common data model (HIE-CDM) to facilitate semantic data flow needed to support a learning health system (LHS). MATERIALS AND METHODS: The Rhode Island Quality Institute operates the Rhode Island (RI) statewide HIE, which aggregates RI health data for more than half of the state's population from 47 data partners. We standardized HIE data to the Observational Medical Outcomes Partnership (OMOP) CDM. Atherosclerotic cardiovascular disease (ASCVD) risk and primary prevention practices were selected to demonstrate LHS semantic data flow from 2013 to 2023. RESULTS: We calculated longitudinal 10-year ASCVD risk on 62,999 individuals. Nearly two-thirds had ASCVD risk factors from more than one data partner. This enabled granular tracking of individual ASCVD risk, primary prevention (ie, statin therapy), and incident disease. The population was on statins for fewer than half of the guideline-recommended days. We also found that individuals receiving care at Federally Qualified Health Centers were more likely to have unfavorable ASCVD risk profiles and more likely to be on statins. CDM transformation reduced data heterogeneity through a unified health record that adheres to defined terminologies per OMOP domain. DISCUSSION: We demonstrated the potential for an HIE-CDM to enable observational population health research. We also showed how to leverage existing health information technology infrastructure and health data best practices to break down LHS barriers. CONCLUSION: HIE-CDM facilitates knowledge curation and health system intervention development at the individual, health system, and population levels.
Aaron S. Eisman, Elizabeth S. Chen, Wen-Chih Wu, Karen Crowley, Dilum P. Aluthge, Katherine A. Brown, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.7
2025 The health data utility and the resurgence of health information exchanges as a national resource
abstract
OBJECTIVES: (1) Describe the evolution of Health Information Exchanges (HIEs) into Health Data Utilities (HDUs); (2) Provide motivation for HDUs as a public strategic investment target. MATERIALS AND METHODS: We examine trends in developing HIEs into HDUs and compare their criticality to that of the national highway system as an investment in the public good. RESULTS: We propose that investment in HDUs is essential for our nation's healthcare data ecosystem. This investment will address the increased need for healthcare delivery and public health data. DISCUSSION: HDUs can meet the current and future needs of healthcare delivery and public health surveillance. Their structure and capabilities will underpin their success to support data-driven decision-making. CONCLUSION: Transforming HIEs into HDUs is essential to realizing the vision of a distributed and connected healthcare data system. Public funding is critical for this model's success, similar to the continued investment in the national highway system.
Anjum Khurshid, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.2
2022 An Unsupervised Cluster Analysis of Post-COVID-19 Mental Health Outcomes and Associated Comorbidities
Katherine A. Brown, Indra Neil Sarkar, Karen Crowley, Dilum P. Aluthge, Elizabeth S. Chen
AMIA2
2022 GenBank as a source to monitor and analyze Host-Microbiome data
abstract
MOTIVATION: Microbiome datasets are often constrained by sequencing limitations. GenBank is the largest collection of publicly available DNA sequences, which is maintained by the National Center of Biotechnology Information (NCBI). The metadata of GenBank records are a largely understudied resource and may be uniquely leveraged to access the sum of prior studies focused on microbiome composition. Here, we developed a computational pipeline to analyze GenBank metadata, containing data on hosts, microorganisms and their place of origin. This work provides the first opportunity to leverage the totality of GenBank to shed light on compositional data practices that shape how microbiome datasets are formed as well as examine host-microbiome relationships. RESULTS: The collected dataset contains multiple kingdoms of microorganisms, consisting of bacteria, viruses, archaea, protozoa, fungi, and invertebrate parasites, and hosts of multiple taxonomical classes, including mammals, birds and fish. A human data subset of this dataset provides insights to gaps in current microbiome data collection, which is biased towards clinically relevant pathogens. Clustering and phylogenic analysis reveals the potential to use these data to model host taxonomy and evolution, revealing groupings formed by host diet, environment and coevolution. AVAILABILITY AND IMPLEMENTATION: GenBank Host-Microbiome Pipeline is available at https://github.com/bcbi/genbank_holobiome. The GenBank loader is available at https://github.com/bcbi/genbank_loader. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vivek Ramanan, Shanti Mechery, Indra Neil Sarkar
Bioinform.3
2022 Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborative
abstract
OBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require.
Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai
J. Am. Medical Informatics Assoc.74
2021 Clinical Note Section Detection Using a Hidden Markov Model of Unified Medical Language System Semantic Types
Aaron S. Eisman, Katherine A. Brown, Elizabeth S. Chen, Indra Neil Sarkar
AMIA4
2021 Public Health Informatics in a Global Pandemic: In the COVID response trenches
Jessica D. Tenenbaum, Theresa A. Cullen, Indra Neil Sarkar, Philip R. O. Payne, Brian E. Dixon
AMIA3
2020 Mental Health Comorbidity Analysis in Pediatric Patients with Autism Spectrum Disorder Using Rhode Island Medical Claims Data
Katherine A. Brown, Indra Neil Sarkar, Elizabeth S. Chen
AMIA2
2020 Extracting Angina Symptoms from Clinical Notes Using Pre-Trained Transformer Architectures
Aaron S. Eisman, Nishant R. Shah, Carsten Eickhoff, George Zerveas, Elizabeth S. Chen, Wen-Chih Wu, Indra Neil Sarkar
AMIA7
2020 Coding Free-Text Chief Complaints from a Health Information Exchange: A Preliminary Study
Sotiris Karagounis, Indra Neil Sarkar, Elizabeth S. Chen
AMIA2
2020 A Phylogenetic Approach to Analyze the Conservativeness of BRCA1 and BRCA2 Mutations
Jiaying Lai, Indra Neil Sarkar
AMIA2
2019 Solr-Plant: efficient extraction of plant names from text
abstract
BACKGROUND: The retrieval of plant-related information is a challenging task due to variations in species name mentions as well as spelling or typographical errors across data sources. Scalable solutions are needed for identifying plant name mentions from text and resolving them to accepted taxonomic names. RESULTS: An Apache Solr-based fuzzy matching system enhanced with the Smith-Waterman alignment algorithm ("Solr-Plant") was developed for mapping and resolution to a plant name and synonym thesaurus. Evaluation of Solr-Plant suggests promising results in terms of both accuracy and processing efficiency on misspelled species names from two benchmark datasets: (1) SALVIAS and (2) National Center for Biotechnology Information (NCBI) Taxonomy. Additional evaluation using S800 text corpus also reflects high precision and recall. The latest version of the source code is available at https://github.com/bcbi/SolrPlantAPI . A REST-compliant web interface and service for Solr-Plant is hosted at http://bcbi.brown.edu/solrplant . CONCLUSION: Automated techniques are needed for efficient and accurate identification of knowledge linked with biological scientific names. Solr-Plant complements the current state-of-the-art in terms of both efficiency and accuracy in identification of names restricted at species level. The approach can be extended to identify broader groups of organisms at different taxonomic levels. The results reflect potential utility of Solr-Plant as a data mining tool for extracting and correcting plant species names.
Vivekanand Sharma, Maria I. Restrepo 0002, Indra Neil Sarkar
BMC Bioinform.3
2018 Using Demographic Factors and Comorbidities to Develop a Predictive Model for ICU Mortality in Patients with Acute Exacerbation COPD
Sukrit S. Jain, Indra Neil Sarkar, Paul Stey, Rajsavi S. Anand, Dustin R. Biron, Elizabeth S. Chen
AMIA2
2018 Predicting Thromboembolism after Vaginal and Laparoscopic Hysterectomy in Gynecologic Oncology Patients using Machine Learning
Justin E. Kleiner, Dilum P. Aluthge, Ishan Sinha, Indra Neil Sarkar, Elizabeth S. Chen, Dario R. Roque
AMIA5
2018 Predicting the Costs of Medical Liability Using Artificial Neural Networks
Jonathan T. Vu, Dilum P. Aluthge, Ishan Sinha, Elizabeth S. Chen, Indra Neil Sarkar
AMIA5
2017 An Automated System for Categorizing Transthoracic Echocardiography Indications According to the Echocardiography Appropriate Use Criteria
Aaron S. Eisman, Rory B. Weiner, Elizabeth S. Chen, Paul Stey, Rishi K. Wadhera, Aaron P. Kithcart, Indra Neil Sarkar
AMIA7
2017 Comorbidity Miner: An Open Source Interactive Tool for Mining Disparate Electronic Health Data Sources
Ashley S. Lee, Indra Neil Sarkar, Genevieve B. Melton, Yuanqing Liu, Vivekanand Sharma, Elizabeth S. Chen
AMIA2
2017 Harnessing Biomedical Natural Language Processing Tools to Identify Medicinal Plant Knowledge from Historical Texts
Vivekanand Sharma, Wayne Law, Michael J. Balick, Indra Neil Sarkar
AMIA4
2017 Assessing Data Quality within Health Information Exchanges: A Case Study for Supporting Emergency Department Research
Margaret Thorsen, Indra Neil Sarkar, Elizabeth S. Chen, Elaine Fontaine, Gregory Walker, Megan Ranney
AMIA2
2017 Crossing the health IT chasm: considerations and policy recommendations to overcome current challenges and enable value-based care
abstract
While great progress has been made in digitizing the US health care system, today's health information technology (IT) infrastructure remains largely a collection of systems that are not designed to support a transition to value-based care. In addition, the pursuit of value-based care, in which we deliver better care with better outcomes at lower cost, places new demands on the health care system that our IT infrastructure needs to be able to support. Provider organizations pursuing new models of health care delivery and payment are finding that their electronic systems lack the capabilities needed to succeed. The result is a chasm between the current health IT ecosystem and the health IT ecosystem that is desperately needed.In this paper, we identify a set of focal goals and associated near-term achievable actions that are critical to pursue in order to enable the health IT ecosystem to meet the acute needs of modern health care delivery. These ideas emerged from discussions that occurred during the 2015 American Medical Informatics Association Policy Invitational Meeting. To illustrate the chasm and motivate our recommendations, we created a vignette from the multistakeholder perspectives of a patient, his provider, and researchers/innovators. It describes an idealized scenario in which each stakeholder's needs are supported by an integrated health IT environment. We identify the gaps preventing such a reality today and present associated policy recommendations that serve as a blueprint for critical actions that would enable us to cross the current health IT chasm by leveraging systems and information to routinely deliver high-value care.
Julia Adler-Milstein, Peter J. Embí, Blackford Middleton, Indra Neil Sarkar, Jeff Smith
J. Am. Medical Informatics Assoc.4
2016 New Pathways Into Biomedical Informatics: Educational Outreach Programs for High School Students
David Boone, John T. Finnell, Kim M. Unertl, Indra Neil Sarkar
AMIA4
2016 Mining and Visualizing Sequential Patterns in the Electronic Health Record: A Case Study for Asthma With and Without Mental Disorders
Elizabeth S. Chen, Genevieve B. Melton, Mark Howison, Erik Knoll, Ashley S. Lee, Indra Neil Sarkar
AMIA6
2016 Report from the ED-WG Frontline: Biomedical and Health Informatics Baccalaureate (BHIB) Course
Saif S. Khairat, Glynda Doyle, Indra Neil Sarkar, Jeff Williamson
AMIA3
2016 Developing new pathways into the biomedical informatics field: the AMIA High School Scholars Program
abstract
Increasing access to biomedical informatics experiences is a significant need as the field continues to face workforce challenges. Looking beyond traditional medical school and graduate school pathways into the field is crucial for expanding the number of individuals and increasing diversity in the field. This case report provides an overview of the development and initial implementation of the American Medical Informatics Association (AMIA) High School Scholars Program. Initiated in 2014, the program's primary goal was to provide dissemination opportunities for high school students engaged in biomedical informatics research. We discuss success factors including strong cross-institutional, cross-organizational collaboration and the high quality of high school student submissions to the program. The challenges encountered, especially around working with minors and communicating program expectations clearly, are also discussed. Finally, we present the path forward for the continued evolution of the AMIA High School Scholars Program.
Kim M. Unertl, John T. Finnell, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.3
2016 Harnessing next-generation informatics for personalizing medicine: a report from AMIA's 2014 Health Policy Invitational Meeting
abstract
The American Medical Informatics Association convened the 2014 Health Policy Invitational Meeting to develop recommendations for updates to current policies and to establish an informatics research agenda for personalizing medicine. In particular, the meeting focused on discussing informatics challenges related to personalizing care through the integration of genomic or other high-volume biomolecular data with data from clinical systems to make health care more efficient and effective. This report summarizes the findings (n = 6) and recommendations (n = 15) from the policy meeting, which were clustered into 3 broad areas: (1) policies governing data access for research and personalization of care; (2) policy and research needs for evolving data interpretation and knowledge representation; and (3) policy and research needs to ensure data integrity and preservation. The meeting outcome underscored the need to address a number of important policy and technical considerations in order to realize the potential of personalized or precision medicine in actual clinical contexts.
Laura K. Wiley, Peter Tarczy-Hornoch, Joshua C. Denny, Robert R. Freimuth, Casey Overby Taylor, Nigam H. Shah, Ross D. Martin, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.8
2015 Representation of Drug Use in Biomedical Standards, Clinical Text, and Research Measures
Elizabeth W. Carter, Indra Neil Sarkar, Genevieve B. Melton, Elizabeth S. Chen
AMIA2
2015 Mining and Visualizing Family History Associations in the Electronic Health Record: A Case Study for Pediatric Asthma
Elizabeth S. Chen, Genevieve B. Melton, Richard Wasserman, Paul Rosenau, Diantha B. Howard, Indra Neil Sarkar
AMIA6
2015 Surveying Problem List Perceptions and Use in the Electronic Health Record
Darlene Peterson, Timothy E. Burdick, Indra Neil Sarkar, Julie Lin, Dennis Plante, Elizabeth S. Chen
AMIA4
2015 Automated Extraction of Substance Use Information from Clinical Texts
Yan Wang 0025, Elizabeth S. Chen, Serguei V. S. Pakhomov, Elliot G. Arsoniadis, Elizabeth W. Carter, Elizabeth Lindemann, Indra Neil Sarkar, Genevieve B. Melton
AMIA7
2015 Multi-source development of an integrated model for family health history
abstract
OBJECTIVE: To integrate data elements from multiple sources for informing comprehensive and standardized collection of family health history (FHH). MATERIALS AND METHODS: Three types of sources were analyzed to identify data elements associated with the collection of FHH. First, clinical notes from multiple resources were annotated for FHH information. Second, questions and responses for family members in patient-facing FHH tools were examined. Lastly, elements defined in FHH-related specifications were extracted for several standards development and related organizations. Data elements identified from the notes, tools, and specifications were subsequently combined and compared. RESULTS: In total, 891 notes from three resources, eight tools, and seven specifications associated with four organizations were analyzed. The resulting Integrated FHH Model consisted of 44 data elements for describing source of information, family members, observations, and general statements about family history. Of these elements, 16 were common to all three source types, 17 were common to two, and 11 were unique. Intra-source comparisons also revealed common and unique elements across the different notes, tools, and specifications. DISCUSSION: Through examination of multiple sources, a representative and complementary set of FHH data elements was identified. Further work is needed to create formal representations of the Integrated FHH Model, standardize values associated with each element, and inform context-specific implementations. CONCLUSIONS: There has been increased emphasis on the importance of FHH for supporting personalized medicine, biomedical research, and population health. Multi-source development of an integrated model could contribute to improving the standardized collection and use of FHH information in disparate systems.
Elizabeth S. Chen, Elizabeth W. Carter, Tamara Winden, Indra Neil Sarkar, Yan Wang 0025, Genevieve B. Melton
J. Am. Medical Informatics Assoc.4
2015 Adapting simultaneous analysis phylogenomic techniques to study complex disease gene relationships
Joseph D. Romano, William G. Tharp, Indra Neil Sarkar
J. Biomed. Informatics3
2014 Examining the Use, Contents, and Quality of Free-Text Tobacco Use Documentation in the Electronic Health Record
Elizabeth S. Chen, Elizabeth W. Carter, Indra Neil Sarkar, Tamara Winden, Genevieve B. Melton
AMIA3
2014 Using Arden Syntax to Identify Registry-Eligible Very Low Birth Weight Neonates from the Electronic Health Record
Indra Neil Sarkar, Elizabeth S. Chen, Paul Rosenau, Matthew B. Storer, Beth Anderson, Jeffrey D. Horbar
AMIA1
2014 PubMedMiner: Mining and Visualizing MeSH-based Associations in PubMed
Yucan Zhang, Indra Neil Sarkar, Elizabeth S. Chen
AMIA2
2014 Drug repurposing: mining protozoan proteomes for targets of known bioactive compounds
abstract
OBJECTIVE: To identify potential opportunities for drug repurposing by developing an automated approach to pre-screen the predicted proteomes of any organism against databases of known drug targets using only freely available resources. MATERIALS AND METHODS: We employed a combination of Ruby scripts that leverage data from the DrugBank and ChEMBL databases, MySQL, and BLAST to predict potential drugs and their targets from 13 published genomes. Results from a previous cell-based screen to identify inhibitors of Cryptosporidium parvum growth were used to validate our in-silico prediction method. RESULTS: In-vitro validation of these results, using a cell-based C parvum growth assay, showed that the predicted inhibitors were significantly more likely than expected by chance to have confirmed activity, with 8.9-15.6% of predicted inhibitors confirmed depending on the drug target database used. This method was then used to predict inhibitors for the following 13 disease-causing protozoan parasites, including: C parvum, Entamoeba histolytica, Giardia intestinalis, Leishmania braziliensis, Leishmania donovani, Leishmania major, Naegleria gruberi (in proxy of Naegleria fowleri), Plasmodium falciparum, Plasmodium vivax, Toxoplasma gondii, Trichomonas vaginalis, Trypanosoma brucei and Trypanosoma cruzi. CONCLUSIONS: Although proteome-wide screens for drug targets have disadvantages, in-silico methods can be developed that are fast, broad, inexpensive, and effective. In-vitro validation of our results for C parvum indicate that the method presented here can be used to construct a library for more directed small molecule screening, or pipelined into structural modeling and docking programs to facilitate target-based drug development.
Adam Sateriale, Kovi Bessoff, Indra Neil Sarkar, Christopher D. Huston
J. Am. Medical Informatics Assoc.3
2014 Structural network analysis of biological networks for assessment of potential disease model organisms
Ahmed Ragab Nabhan, Indra Neil Sarkar
J. Biomed. Informatics2
2013 Development of a Comprehensive Family Health History Information Model
Elizabeth S. Chen, Elizabeth W. Carter, Tamara Winden, Indra Neil Sarkar, Genevieve B. Melton
AMIA4
2013 Resolving and Standardizing Providers within Administrative Data
Indra Neil Sarkar, Elizabeth S. Chen, Steven J. Kappel
AMIA1
2013 Bioinformatics opportunities for identification and study of medicinal plants
abstract
Plants have been used as a source of medicine since historic times and several commercially important drugs are of plant-based origin. The traditional approach towards discovery of plant-based drugs often times involves significant amount of time and expenditure. These labor-intensive approaches have struggled to keep pace with the rapid development of high-throughput technologies. In the era of high volume, high-throughput data generation across the biosciences, bioinformatics plays a crucial role. This has generally been the case in the context of drug designing and discovery. However, there has been limited attention to date to the potential application of bioinformatics approaches that can leverage plant-based knowledge. Here, we review bioinformatics studies that have contributed to medicinal plants research. In particular, we highlight areas in medicinal plant research where the application of bioinformatics methodologies may result in quicker and potentially cost-effective leads toward finding plant-based remedies.
Vivekanand Sharma, Indra Neil Sarkar
Briefings Bioinform.2
2013 Research and applications: Leveraging biodiversity knowledge for potential phyto-therapeutic applications
abstract
OBJECTIVE: To identify and highlight the feasibility, challenges, and advantages of providing a cross-domain pipeline that can link relevant biodiversity information for phyto-therapeutic assessment. MATERIALS AND METHODS: A public repository of clinical trials information (ClinicalTrials.gov) was explored to determine the state of plant-based interventions under investigation. RESULTS: The results showed that ≈ 15% of drug interventions in ClinicalTrials.gov were potentially plant related, with about 60% of them clustered within 10 taxonomic families. Further analysis of these plant-based interventions identified ≈ 3.7% of associated plant species as endangered as determined from the International Union for the Conservation of Nature Red List. DISCUSSION: The diversity of the plant kingdom has provided human civilization with life-sustaining food and medicine for centuries. There has been renewed interest in the investigation of botanicals as sources of new drugs, building on traditional knowledge about plant-based medicines. However, data about the plant-based biodiversity potential for therapeutics (eg, based on genetic or chemical information) are generally scattered across a range of sources and isolated from contemporary pharmacological resources. This study explored the potential to bridge biodiversity and biomedical knowledge sources. CONCLUSIONS: The findings from this feasibility study suggest that there is an opportunity for developing plant-based drugs and further highlight taxonomic relationships between plants that may be rich sources for bioprospecting.
Vivekanand Sharma, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.2
2013 Leveraging concept-based approaches to identify potential phyto-therapies
Vivekanand Sharma, Indra Neil Sarkar
J. Biomed. Informatics2
2012 Characterizing the Use and Contents of Free-Text Family History Comments in the Electronic Health Record
Elizabeth S. Chen, Genevieve B. Melton, Timothy E. Burdick, Paul Rosenau, Indra Neil Sarkar
AMIA5
2012 Social and Behavioral History Information in Public Health Datasets
Genevieve B. Melton, Sharad Manaktala, Indra Neil Sarkar, Elizabeth S. Chen
AMIA3
2012 Mining Disease Fingerprints From Within Genetic Pathways
Ahmed Ragab Nabhan, Indra Neil Sarkar
AMIA2
2012 Determining Compound Comorbidities for Heart Failure from Hospital Discharge Data
Indra Neil Sarkar, Elizabeth S. Chen
AMIA1
2012 The impact of taxon sampling on phylogenetic inference: a review of two decades of controversy
abstract
Over the past two decades, there has been a long-standing debate about the impact of taxon sampling on phylogenetic inference. Studies have been based on both real and simulated data sets, within actual and theoretical contexts, and using different inference methods, to study the impact of taxon sampling. In some cases, conflicting conclusions have been drawn for the same data set. The main questions explored in studies to date have been about the effects of using sparse data, adding new taxa, including more characters from genome sequences and using different (or concatenated) locus regions. These questions can be reduced to more fundamental ones about the assessment of data quality and the design guidelines of taxon sampling in phylogenetic inference experiments. This review summarizes progress to date in understanding the impact of taxon sampling on the accuracy of phylogenetic analysis.
Ahmed Ragab Nabhan, Indra Neil Sarkar
Briefings Bioinform.2
2012 A vector space model approach to identify genetically related diseases
abstract
OBJECTIVE: The relationship between diseases and their causative genes can be complex, especially in the case of polygenic diseases. Further exacerbating the challenges in their study is that many genes may be causally related to multiple diseases. This study explored the relationship between diseases through the adaptation of an approach pioneered in the context of information retrieval: vector space models. MATERIALS AND METHODS: A vector space model approach was developed that bridges gene disease knowledge inferred across three knowledge bases: Online Mendelian Inheritance in Man, GenBank, and Medline. The approach was then used to identify potentially related diseases for two target diseases: Alzheimer disease and Prader-Willi Syndrome. RESULTS: In the case of both Alzheimer Disease and Prader-Willi Syndrome, a set of plausible diseases were identified that may warrant further exploration. DISCUSSION: This study furthers seminal work by Swanson, et al. that demonstrated the potential for mining literature for putative correlations. Using a vector space modeling approach, information from both biomedical literature and genomic resources (like GenBank) can be combined towards identification of putative correlations of interest. To this end, the relevance of the predicted diseases of interest in this study using the vector space modeling approach were validated based on supporting literature. CONCLUSION: The results of this study suggest that a vector space model approach may be a useful means to identify potential relationships between complex diseases, and thereby enable the coordination of gene-based findings across multiple complex diseases.
Indra Neil Sarkar
J. Am. Medical Informatics Assoc.1
2012 Translating standards into practice: Experiences and lessons learned in biomedicine and health care
abstract
The value of standards for the representation, integration, and exchange of data, information, and knowledge across the spectrum of biomedicine and health care has been widely recognized for years. Recent initiatives further underscore the importance of standards (e.g., for certification and meaningful use of electronic health records) and establishment of systematic approaches for achieving semantic interoperability. There is accordingly a need for detailed, experience-based discussions pertaining to the adoption and implementation of the breadth of standards in biomedicine and health care. The goal of this special issue has been to provide a forum for describing advanced research and development in translating standards into practice. Each paper in this issue provides a comprehensive description of methodologies employed and challenges encountered during the process of implementing a specific standard or set of standards in a practical setting. The twenty-two papers (including one methodological review) represent a broad array of experiences with standards and are categorized into the following sections: (1) Terminology Standards, (2) Document Standards, (3) Decision Support Standards, (4) Standards-Based Infrastructure, and (5) Standards Adoption Processes.
Elizabeth S. Chen, Genevieve B. Melton, Indra Neil Sarkar
J. Biomed. Informatics3
2011 Translational bioinformatics: linking knowledge across biological and clinical realms
abstract
Nearly a decade since the completion of the first draft of the human genome, the biomedical community is positioned to usher in a new era of scientific inquiry that links fundamental biological insights with clinical knowledge. Accordingly, holistic approaches are needed to develop and assess hypotheses that incorporate genotypic, phenotypic, and environmental knowledge. This perspective presents translational bioinformatics as a discipline that builds on the successes of bioinformatics and health informatics for the study of complex diseases. The early successes of translational bioinformatics are indicative of the potential to achieve the promise of the Human Genome Project for gaining deeper insights to the genetic underpinnings of disease and progress toward the development of a new generation of therapies.
Indra Neil Sarkar, Atul J. Butte, Yves A. Lussier, Peter Tarczy-Hornoch, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.1
2011 Selected Papers from the 2011 Summit on Translational Bioinformatics
Indra Neil Sarkar
J. Biomed. Informatics1
2011 The joint summits on translational science: Crossing the translational chasm
abstract
The joint summits on translational science: Crossing the translational chasm From its beginnings nearly 60 years ago, the field that soon became known as ''medical informatics'' focused on the harnessing of computational and information science approaches, as well as socio-technical frameworks, in order to improve the quality, safety, and outcomes of clinical care.With the emergence of bioinformatics as a formal area of research approximately 30 years later, the biomedical and life sciences enterprises became positioned to gain deeper understanding of the underpinning phenomena associated with disease.The amalgamation of bioinformatics and medical informatics, formally becoming termed ''biomedical informatics'', suggested that there was a continuum of concepts that spanned from wetbench research to bedside medicine.The major thrust in research in biomedical informatics was thus to support a wide range of practitioners as they sought both to ask and answer complex questions spanning the biological, clinical, or public health domains.In many ways, biomedical informatics shifted from hypothesis-generation research to developing approaches to support the testing of specific biomedical hypotheses.By the early part of the 21st century, the scope and pace of foundational biomedical research was increasing rapidly, largely spurred on by the completion of the $2.7B Human Genome Project.Innovations in bioinformatics complemented by significant technological advances quickly helped the research community contemplate high-throughput 'omics' studies that would realize the promise of employing the blueprint of life in order to foster a new generation of diagnostic and therapeutic approaches.In a complementary manner, the clinical research community began to contemplate the necessary infrastructure and methodological innovations needed to bring bench-side innovations into the clinic in a more timely and efficient manner.Such emerging lines of research and development were largely motivated by the emergence of the United States' National Institutes of Health (NIH) Roadmap, which included specific objectives surrounding the ''re-engineering'' of the clinical research enterprise.These ''re-engineering'' efforts ultimately evolved into the launch of the Clinical and Translational Science Award (CTSA) program.Funded by the NIH through its National Center for Research Resources (and now via a new proposed Center that will focus on translational science), and ultimately seeking to identify and engage a nationwide network of 60 sites, the CTSA initiative focuses on creating academic and scholarly homes with accompanying infrastructure, services, and workforce development programs.The program is intended to catalyze rapid improvements in the state of clinical and translational research knowledge and practice and to decrease the time between basic science discoveries and the implementation of their clinical implications in routine practice.Given an increasing appreciation by the biomedical informatics community that research in bioinformatics and clinical informatics
Indra Neil Sarkar, Philip R. O. Payne
J. Biomed. Informatics1
2011 Enhancing phylogeography by improving geographical information from GenBank
abstract
Phylogeography is a field that focuses on the geographical lineages of species such as vertebrates or viruses. Here, geographical data, such as location of a species or viral host is as important as the sequence information extracted from the species. Together, this information can help illustrate the migration of the species over time within a geographical area, the impact of geography over the evolutionary history, or the expected population of the species within the area. Molecular sequence data from NCBI, specifically GenBank, provide an abundance of available sequence data for phylogeography. However, geographical data is inconsistently represented and sparse across GenBank entries. This can impede analysis and in situations where the geographical information is inferred, and potentially lead to erroneous results. In this paper, we describe the current state of geographical data in GenBank, and illustrate how automated processing techniques such as named entity recognition, can enhance the geographical data available for phylogeographic studies.
Matthew Scotch, Indra Neil Sarkar, Changjiang Mei, Robert Leaman, Kei-Hoi Cheung, Pierina Ortiz, Ashutosh Singraur, Graciela Gonzalez-Hernandez
J. Biomed. Informatics2
2010 Editorial: Bioinformatics education in the 21st century
abstract
Bioinformatics competency has become a key element in much of contemporary biology. The increasing potential that is to be realized with modern technologies, such as ‘next-generation’ sequencing, parallel and ‘cloud’ computing, and continually increasing data banks, requires a fundamental understanding of bioinformatics concepts and applications. On par with the complexity of biological inquiry, acquiring bioinformatics competency can become a complex endeavor of undecipherable jargon, mathematics and frustration. The requirement for bioinformatics training affects the full range of budding biologists to seasoned professionals. This special issue stems from a personal desire to coalesce the most recent ‘snapshot’ of current activities in bioinformatics education, with the goal of providing a resource for biologists and bioinformatics educators alike. This special issue begins with a pair of articles that captures the current state of bioinformatics education. First, Cummings and Temple explore the opportunities and challenges generally faced in bioinformatics education. Then, Schneider et al. present a discussion of approaches for providing adequate training for bioinformatics services. Both articles contextualize their observations relative to pragmatic training experiences, which provide context for the breadth of obstacles faced by modern bioinformatics education. Together, these opening articles help set the stage for understanding the range of educational experiences and approaches that are the focus of the rest of the articles in this special issue. The importance of bioinformatics has underscored the value of common resources that are available to the scientific community at large. Leading the way in providing publicly available bioinformatics education support are the European Molecular Biology Laboratory’s European Bioinformatics Institute (EBI) and the National Center for Biotechnology Information (NCBI) at the United States National Library of Medicine. Wright et al. describe the potential for ‘eLearning’ modalities to provide bioinformatics training, with specific attention to the resources available from the EBI. Cooper et al. then provide an overview of the bioinformatics education resources at the NCBI, introducing a newly redesigned Web-based portal to NCBI Educational resources. The next three articles explore the challenge of developing bioinformatics education approaches. Jungck et al. describe the BEDROCK project as an approach for supporting bioinformatics education, with a particular emphasis on undergraduate education. Appreciating that much of bioinformatics education must be pragmatic, Rother et al. describe ‘multi-stage learning aids’ as a way to provide detailed practical education on bioinformatics applications, such as PyMOL. Finally, Buttigieg presents a review of approaches for providing bioinformatics education that blends multimedia presentations, visual communication and practical pedagogy. Appreciating that much of bioinformatics education may not occur within the classroom setting, the next two articles describe educational approaches in ‘real-world’ settings. Williams et al. describe resources developed by OpenHelix that provide life-long learning opportunities for researchers to gain necessary bioinformatics expertise. Yamashita et al. then describe an approach for providing hands-on training of concepts necessary for bioinformatics application development through immersion in a practical setting (in this case, the Marine Biological Laboratory). The globalization of biological research implicates the need for bioinformatics education internationally. The last two articles of this special issue aim to provide an international perspective of bioinformatics education. First, Kilkarne-Kale et al. describe the history and future directions of bioinformatics in India. The special issue then concludes with a review of the Gulbenkian Training Programme in Bioinformatics in Portugal, which has provided a comprehensive bioinformatics educational experience with limited resources for nearly a decade. While not all biologists will necessarily specialize in developing bioinformatics innovations, their biology innovations will increasingly become dependent on bioinformatics. Bioinformatics education will thus become more essential in biology education. As exemplified through the articles in this special issue, the future for bioinformatics education has a significant underpinning of experiences and approaches for providing the necessary range of bioinformatics education. What is also highlighted through all the articles in this special issue is that there are still major obstacles that lie ahead for the development of a uniform approach for bioinformatics education. Nonetheless, it is my view that the inspired insights of those involved with bioinformatics education (such as those who contributed to this special issue) provide the hope that such obstacles are surmountable—the future of biology will depend on it.
Indra Neil Sarkar
Briefings Bioinform.1
2010 Evaluation of family history information within clinical documents and adequacy of HL7 clinical statement and clinical genomics family history models for its representation: a case report
abstract
Family history information has emerged as an increasingly important tool for clinical care and research. While recent standards provide for structured entry of family history, many clinicians record family history data in text. The authors sought to characterize family history information within clinical documents to assess the adequacy of existing models and create a more comprehensive model for its representation. Models were evaluated on 100 documents containing 238 sentences and 410 statements relevant to family history. Most statements were of family member plus disease or of disease only. Statement coverage was 91%, 77%, and 95% for HL7 Clinical Genomics Family History Model, HL7 Clinical Statement Model, and the newly created Merged Family History Model, respectively. Negation (18%) and inexact family member specification (9.5%) occurred commonly. Overall, both HL7 models could represent most family history statements in clinical reports; however, refinements are needed to represent the full breadth of family history data.
Genevieve B. Melton, Nandhini Raman, Elizabeth S. Chen, Indra Neil Sarkar, Serguei V. S. Pakhomov, Robert D. Madoff
J. Am. Medical Informatics Assoc.4
2010 MeSHing molecular sequences and clinical trials: A feasibility study
Elizabeth S. Chen, Indra Neil Sarkar
J. Biomed. Informatics2
2009 LigerCat: Using "MeSH Clouds" from Journal, Article, or Gene Citations to Facilitate the Identification of Relevant Biomedical Literature
Indra Neil Sarkar, Ryan Schenk, Holly Miller, Catherine N. Norton
AMIA1
2009 Selected proceedings of the First Summit on Translational Bioinformatics 2008
Atul J. Butte, Indra Neil Sarkar, Marco Ramoni, Yves A. Lussier, Olga G. Troyanskaya
BMC Bioinform.2
2009 Selected proceedings of the 2009 Summit on Translational Bioinformatics
Yves A. Lussier, Indra Neil Sarkar
BMC Bioinform.2
2009 Biodiversity Informatics: the emergence of a field
abstract
Recent years have seen great technological advances that have helped usher in a new generation of approaches to understand and share knowledge about the planet in which we live. A number of major initiatives that aim to catalyze necessary technological and biological advances synergistically have emerged globally. Of these enabling initiatives, the Encyclopedia of Life (EOL; http://www.eol.org) and the Barcode of Life (BOL; http://barcoding.si.edu) are projects that have collectively help lay a framework within which we will see the next generation of taxonomic innovation and discovery. The successes of both EOL and BOL have the potential to impact multiple facets of society - from the discovery of new species, to the development of conservation strategies for endangered life, to insights into infectious disease hosts and vectors, to the discovery of life-saving medicinal plants, to the piquing of general interest about life on Earth and our role in the complex web of life. Both the EOL and BOL initiatives are possible thanks to a number of significant advancements in knowledge discovery, integration, and management techniques. Collectively termed 'biodiversity informatics,' this new suite of methodologies and tools extends contemporary computer science and informatics principles within the context of biodiversity data. This supplement, made possible through funding from both EOL and BOL, brings forth some of the pioneering work from leading biodiversity informatics researchers. While nascent as a discipline, biodiversity informatics has proven to not only adopt, but also help significantly challenge and advance the most recent technological advances and computational approaches for managing complex data. In contrast to bioinformatics, which in primarily focused on managing relevant molecular biology data, biodiversity informatics requires frameworks and approaches that can accommodate the full range of biological information - from molecules to morphological features, to populations, to habitats - collectively developing the ultimate computational Web of knowledge about life on Earth. This supplement starts with a piece from Chavan and Ingwersen that goes through some of the fundamental principles for disseminating biodiversity information, discussing both the hindrances and opportunities [1]. Hill et al. then describe how one might leverage existing technological infrastructure for enabling georeferencing of biodiversity data [2]. Demonstrating how biodiversity informatics can often benefit from the latest advances in searching strategies, Hajibabaei and Singer describe an approach for making use of Google to identify relevant information with respect to DNA sequences [3]. Next, Page discusses the nuances of biological global unique identifiers, which may be necessary to link relevant biodiversity data across various spheres of knowledge [4]. In consideration of the necessary continual curation required for disparate biodiversity knowledge, Smith et al. describe a Drupal-based technology called Scratchpads for managing and sharing biodiversity knowledge [5]. In parallel to the emerging infrastructure for managing and disseminating biodiversity information, DNA Barcode analysis methods represent a crucial entry point into the realm of biodiversity knowledge. The last four articles thus focus on some recent advances in DNA Barcoding analytic approaches. The first of these articles, from Bertolazzi et al., presents a machine learning approach for classifying species according to DNA Barcode derived information [6]. Chu et al. then describe a 'composition vector' approach for making use of large datasets of DNA Barcodes for classification [7]. In light of molecular sequence alignment as an often rate-limiting step in many classification approaches, Kuksa and Pavlovic present an alignment-free approach for DNA Barcode data [8]. In light of the range of approaches associated with DNA Barcode based classification, Austerlitz et al. present an overview of common phylogenetic and statistical methods most commonly considered [9]. A unifying theme in the articles of this supplement is the diversity of issues that remain to be resolved going forward. It is my hope that this issue helps continue and inspire new dialogue across the full range of disciplines associated with this burgeoning field.
Indra Neil Sarkar
BMC Bioinform.1
2008 Automated simultaneous analysis phylogenetics (ASAP): an enabling tool for phlyogenomics
abstract
BACKGROUND: The availability of sequences from whole genomes to reconstruct the tree of life has the potential to enable the development of phylogenomic hypotheses in ways that have not been before possible. A significant bottleneck in the analysis of genomic-scale views of the tree of life is the time required for manual curation of genomic data into multi-gene phylogenetic matrices. RESULTS: To keep pace with the exponentially growing volume of molecular data in the genomic era, we have developed an automated technique, ASAP (Automated Simultaneous Analysis Phylogenetics), to assemble these multigene/multi species matrices and to evaluate the significance of individual genes within the context of a given phylogenetic hypothesis. CONCLUSION: Applications of ASAP may enable scientists to re-evaluate species relationships and to develop new phylogenomic hypotheses based on genome-scale data.
Indra Neil Sarkar, Mary G. Egan, Gloria M. Coruzzi, Ernest K. Lee, Robert DeSalle
BMC Bioinform.1
2007 Biodiversity informatics: organizing and linking information across the spectrum of life
abstract
Biological knowledge can be inferred from three major levels of information: molecules, organisms and ecologies. Bioinformatics is an established field that has made significant advances in the development of systems and techniques to organize contemporary molecular data; biodiversity informatics is an emerging discipline that strives to develop methods to organize knowledge at the organismal level extending back to the earliest dates of recorded natural history. Furthermore, while bioinformatics studies generally focus on detailed examinations of key 'model' organisms, biodiversity informatics aims to develop over-arching hypotheses that span the entire tree of life. Biodiversity informatics is presented here as a discipline that unifies biological information from a range of contemporary and historical sources across the spectrum of life using organisms as the linking thread. The present review primarily focuses on the use of organism names as a universal metadata element to link and integrate biodiversity data across a range of data sources.
Indra Neil Sarkar
Briefings Bioinform.1
2007 uBioRSS: Tracking taxonomic literature using RSS
abstract
UNLABELLED: Web content syndication through standard formats such as RSS and ATOM has become an increasingly popular mechanism for publishers, news sources and blogs to disseminate regularly updated content. These standardized syndication formats deliver content directly to the subscriber, allowing them to locally aggregate content from a variety of sources instead of having to find the information on multiple websites. The uBioRSS application is a 'taxonomically intelligent' service customized for the biological sciences. It aggregates syndicated content from academic publishers and science news feeds, and then uses a taxonomic Named Entity Recognition algorithm to identify and index taxonomic names within those data streams. The resulting name index is cross-referenced to current global taxonomic datasets to provide context for browsing the publications by taxonomic group. This process, called taxonomic indexing, draws upon services developed specifically for biological sciences, collectively referred to as 'taxonomic intelligence'. Such value-added enhancements can provide biologists with accelerated and improved access to current biological content. AVAILABILITY: http://names.ubio.org/rss/
Patrick R. Leary, David P. Remsen, Catherine N. Norton, Indra Neil Sarkar
Bioinform.5
2006 Literature Based Discovery of Gene Clusters Using Phylogenetic Methods
Indra Neil Sarkar, Abha Agrawal
AMIA1
2006 zMedline: A Medline-Based Content Management System
Chif N. Umejei, Indra Neil Sarkar
AMIA2
2006 OrthologID: automation of genome-scale ortholog identification within a parsimony framework
abstract
MOTIVATION: The determination of gene orthology is a prerequisite for mining and utilizing the rapidly increasing amount of sequence data for genome-scale phylogenetics and comparative genomic studies. Until now, most researchers use pairwise distance comparisons algorithms, such as BLAST, COG, RBH, RSD and INPARANOID, to determine gene orthology. In contrast, orthology determination within a character-based phylogenetic framework has not been utilized on a genomic scale owing to the lack of efficiency and automation. RESULTS: We have developed OrthologID, a Web application that automates the labor-intensive procedures of gene orthology determination within a character-based phylogenetic framework, thus making character-based orthology determination on a genomic scale possible. In addition to generating gene family trees and determining orthologous gene sets for complete genomes, OrthologID can also identify diagnostic characters that define each orthologous gene set, as well as diagnostic characters that are responsible for classifying query sequences from other genomes into specific orthology groups. The OrthologID database currently includes several complete plant genomes, including Arabidopsis thaliana, Oryza sativa, Populus trichocarpa, as well as a unicellular outgroup, Chlamydomonas reinhardtii. To improve the general utility of OrthologID beyond plant species, we plan to expand our sequence database to include the fully sequenced genomes of prokaryotes and other non-plant eukaryotes. AVAILABILITY: http://nypg.bio.nyu.edu/orthologid/
Joanna C. Chiu, Ernest K. Lee, Mary G. Egan, Indra Neil Sarkar, Gloria M. Coruzzi, Robert DeSalle
Bioinform.4
2006 Phylogenetics in the modern era
Indra Neil Sarkar
J. Biomed. Informatics1
2005 mILD: a tool for constructing and analyzing matrices of pairwise phylogenetic character incongruence tests
abstract
SUMMARY: Pairwise comparisons of disagreement in phylogenetic datasets offer a powerful tool for isolating historical incongruence for closer analysis. Statistically significant phylogenetic character incongruence may reflect important differences in evolutionary history, such as horizontal gene transfer. Such testing can also be used to specify possible combinations of datasets for further phylogenetic analysis. The process of comparing multiple datasets can be very time consuming, and it is sometimes unclear how to combine data partitions given the observed patterns of incongruence. Here we present an application that automates the process of making pairwise comparisons between large numbers of phylogenetic datasets using the Incongruence Length Difference (ILD) test. The application also implements strategies for data combination based on the patterns of incongruence observed in pairwise comparisons.
Paul J. Planet, Indra Neil Sarkar
Bioinform.2
2004 ORFcurator: molecular curation of genes and gene clusters in prokaryotic organisms
abstract
UNLABELLED: The ability to detect clusters of functionally related genes in multiple microbial genomes has enormous potential for enhancing studies on gene function and microbial evolution. The staggering amount of new genome sequence data presents a largely untapped resource for gene cluster discovery. To date, gene cluster analysis has not been fully automated, and one must rely on manual, tedious and time-consuming manipulation of sequences. To facilitate accurate and rapid identification of conserved gene clusters, we developed a database-driven web application, called ORFcurator. We used ORFcurator to find clusters containing any genes similar to those of the 14-gene Widespread Colonization Island of Actinobacillus actinomycetemcomitans. From 126 genomes, ORFcurator identified all 73 clusters previously determined by manual searching. AVAILABILITY: ORFcurator and all associated scripts are freely available as supplementary information. SUPPLEMENTARY INFORMATION: http://www.genomecurator.org/ORFcurator/
Jeffrey A. Rosenfeld, Indra Neil Sarkar, Paul J. Planet, David H. Figurski, Robert DeSalle
Bioinform.2
2002 An integrative model for in-silico clinical-genomics discovery science
Yves A. Lussier, Indra Neil Sarkar, Michael Cantor
AMIA2
2002 Exploring SNOMED Using Phylogenetic Analysis Tools
Indra Neil Sarkar, Michael N. Cantor, Robert DeSalle, Yves A. Lussier
AMIA1
2002 Discovering protein similarity using natural language processing
Indra Neil Sarkar, Thomas C. Rindflesch
AMIA1
2002 Implementation Brief: Desiderata for Personal Electronic Communication in Clinical Systems
abstract
Electronic communication among clinicians and patients is becoming an essential part of medical practice. Evaluation and selection of these electronic systems, called personal clinical electronic communication (PCEC) systems, can be a difficult task in institutions that have no prior experience with such systems. It is particularly difficult in the clinical context. To directly address this point, the authors consulted a group of potential users affiliated with a nationally recognized telemedicine project, to determine important characteristics of a hypothetical PCEC system. They compiled a list of these characteristics and produced a desiderata, or list of desired features, for PCEC systems. Two conventional e-mail implementations and three Web-based PCEC systems were evaluated with respect to the features. The Web-based systems all scored higher than conventional e-mail. It is the hope of the authors that this paper will initiate further discussions about the features of PCEC systems and how to evaluate them.
Indra Neil Sarkar, Justin Starren
J. Am. Medical Informatics Assoc.1
2002 Characteristic attributes in cancer microarrays
Indra Neil Sarkar, Paul J. Planet, T. E. Bael, S. E. Stanley, Mark Siddall, Robert DeSalle, David H. Figurski
J. Biomed. Informatics1
2001 Knowledge Aggregation from Organized Sets (KAOS) Applied to Clinical Data
Indra Neil Sarkar, Paul J. Planet, Robert DeSalle, David H. Figurski
AMIA1