EDBT 2026 Demo / reviewers in the wild / expert
Carol Friedman
dblp:68/6061
· DBLP profile ↗
115ranked-venue papers
18as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 109 · 16 first-authorArtificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% Medical and health informatics · 0% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 53% Knowledge representation and reasoning · 47% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
biomedical text mining |
0.3 | 4 | 2008 | Semantic reclassification of the UMLS concepts · Bioinform. 2008 Gene symbol disambiguation using knowledge-based profiles · Bioinform. 2007 Bio-Ontology and text: bridging the modeling gap · Bioinform. 2006 |
Bioinformatics and computational biology › biomedical text mining
gene name disambiguation |
0.1 | 2 | 2007 | Gene symbol disambiguation using knowledge-based profiles · Bioinform. 2007 Gene name ambiguity of eukaryotic nomenclatures · Bioinform. 2005 |
Bioinformatics and computational biology
biological data visualization |
0.1 | 1 | 2005 | Visualizing information across multidimensional post-genomic structured and textual databases · Bioinform. 2005 |
Bioinformatics and computational biology › biomedical text mining
gene name recognition |
0.1 | 1 | 2005 | Gene name ambiguity of eukaryotic nomenclatures · Bioinform. 2005 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference |
0.0 | 1 | 2004 | Probabilistic inference of molecular networks from noisy data sources · Bioinform. 2004 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.0 | 1 | 2004 | Probabilistic inference of molecular networks from noisy data sources · Bioinform. 2004 |
Bioinformatics and computational biology › biological network › network biology
molecular interaction network |
0.0 | 1 | 2002 | Of truth and pathways: chasing bits of information through myriads of articles · ISMB 2002 |
Bioinformatics and computational biology › systems biology
gene regulatory network modeling |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Bioinformatics and computational biology › ontology
ontology development |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology |
0.0 | 1 | 2008 | Semantic reclassification of the UMLS concepts · Bioinform. 2008 |
Natural language and speech › Information extraction and text analysis › lexical semantics › lexicon learning
semantic class induction |
0.0 | 1 | 2008 | Semantic reclassification of the UMLS concepts · Bioinform. 2008 |
Bioinformatics and computational biology › statistical genetics
phenotype-genotype integration |
0.0 | 1 | 2006 | Bio-Ontology and text: bridging the modeling gap · Bioinform. 2006 |
Bioinformatics and computational biology › data integration
database integration |
0.0 | 1 | 2005 | Visualizing information across multidimensional post-genomic structured and textual databases · Bioinform. 2005 |
Bioinformatics and computational biology
signal transduction |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Logic in computer science › semantics
computational semantics |
0.0 | 1 | 1989 | A General Computational Treatment of the Comparative · ACL 1989 |
Automata and formal languages
parsing |
0.0 | 1 | 1989 | A General Computational Treatment of the Comparative · ACL 1989 |
Medical and health informatics
clinical text processing |
0.0 | 1 | 1985 | Transporting the Linguistic String Project System from a Medical to a Navy Domain · ACM Trans. Inf. Syst. 1985 |
Methods — techniques the papers use, named apart from their topics
natural language processing · 0.2lexical feature analysis · 0.2contextual feature analysis · 0.2profile similarity ranking · 0.1information retrieval · 0.1string matching · 0.1yeast two-hybrid · 0.0probabilistic model · 0.0literature mining · 0.0stochastic modeling · 0.0syntactic analysis · 0.0sublanguage grammar · 0.0semantic analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu 0020, Casey N. Ta, James R. Rogers, Ziran Li, Alex M. Butler, Ning Shang 0004, Fabricio Sampaio Peres Kury, Liwei Wang 0010, Feichen Shen, Lyudmila Ena, Carol Friedman, Chunhua Weng |
J. Biomed. Informatics | 13 |
| 2018 | Validation of the Behavior of a Knowledge Base Implementing Clinical Guidelines for Point-of-Care Antiretroviral Toxicity Monitoring
William Ogallo, Carol Friedman, Andrew S. Kanter |
AMIA | 2 |
| 2018 | Detection of drug-drug interactions through data mining studies using clinical sources, scientific literature and social mediaabstractDrug-drug interactions (DDIs) constitute an important concern in drug development and postmarketing pharmacovigilance. They are considered the cause of many adverse drug effects exposing patients to higher risks and increasing public health system costs. Methods to follow-up and discover possible DDIs causing harm to the population are a primary aim of drug safety researchers. Here, we review different methodologies and recent advances using data mining to detect DDIs with impact on patients. We focus on data mining of different pharmacovigilance sources, such as the US Food and Drug Administration Adverse Event Reporting System and electronic health records from medical institutions, as well as on the diverse data mining studies that use narrative text available in the scientific biomedical literature and social media. We pay attention to the strengths but also further explain challenges related to these methods. Data mining has important applications in the analysis of DDIs showing the impact of the interactions as a cause of adverse effects, extracting interactions to create knowledge data sets and gold standards and in the discovery of novel and dangerous DDIs. Santiago Vilar, Carol Friedman, George Hripcsak |
Briefings Bioinform. | 2 |
| 2017 | Automated Metabolic Phenotyping of Cytochrome Polymorphisms Using PubMed Abstract Mining
Luoxin Chen, Carol Friedman, Joseph Finkelstein |
AMIA | 2 |
| 2017 | JASIST special issue on biomedical information retrieval
Robert Moskovitch, Fei Wang 0001, Jian Pei 0001, Carol Friedman |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2017 | Toward multimodal signal detection of adverse drug reactions
Rave Harpaz, William DuMouchel, Martijn J. Schuemie, Olivier Bodenreider, Carol Friedman, Eric Horvitz, Anna Ripple, Alfred Sorbello, Ryen W. White, Rainer Winnenburg, Nigam H. Shah |
J. Biomed. Informatics | 5 |
| 2015 | Implementing automated delivery of evidence-based medication safety information to the point of care
Joseph Finkelstein, Hayden Z. Adams, Qinlang Chen 0002, Kuo Lin, Carol Friedman |
AMIA | 5 |
| 2015 | Representation of Genetic Variants in Genomic Sequencing Reports
Evelyn Rustia, Chunhua Weng, Carol Friedman |
AMIA | 3 |
| 2015 | MAC Annotator: An interactive tool for translating medication appropriateness criteria into structured form
Hojjat Salmasian, Carol Friedman |
AMIA | 2 |
| 2015 | Medication-indication knowledge bases: a systematic review and critical appraisalabstractOBJECTIVE: Medication-indication information is a key part of the information needed for providing decision support for and promoting appropriate use of medications. However, this information is not readily available to end users, and a lot of the resources only contain this information in unstructured form (free text). A number of public knowledge bases (KBs) containing structured medication-indication information have been developed over the years, but a direct comparison of these resources has not yet been conducted. MATERIAL AND METHODS: We conducted a systematic review of the literature to identify all medication-indication KBs and critically appraised these resources in terms of their scope as well as their support for complex indication information. RESULTS: We identified 7 KBs containing medication-indication data. They notably differed from each other in terms of their scope, coverage for on- or off-label indications, source of information, and choice of terminologies for representing the knowledge. The majority of KBs had issues with granularity of the indications as well as with representing duration of therapy, primary choice of treatment, and comedications or comorbidities. DISCUSSION AND CONCLUSION: This is the first study directly comparing public KBs of medication indications. We identified several gaps in the existing resources, which can motivate future research. Hojjat Salmasian, Tran H. Tran, Herbert S. Chase, Carol Friedman |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortalityabstractOBJECTIVES: Drug repurposing, which finds new indications for existing drugs, has received great attention recently. The goal of our work is to assess the feasibility of using electronic health records (EHRs) and automated informatics methods to efficiently validate a recent drug repurposing association of metformin with reduced cancer mortality. METHODS: By linking two large EHRs from Vanderbilt University Medical Center and Mayo Clinic to their tumor registries, we constructed a cohort including 32,415 adults with a cancer diagnosis at Vanderbilt and 79,258 cancer patients at Mayo from 1995 to 2010. Using automated informatics methods, we further identified type 2 diabetes patients within the cancer cohort and determined their drug exposure information, as well as other covariates such as smoking status. We then estimated HRs for all-cause mortality and their associated 95% CIs using stratified Cox proportional hazard models. HRs were estimated according to metformin exposure, adjusted for age at diagnosis, sex, race, body mass index, tobacco use, insulin use, cancer type, and non-cancer Charlson comorbidity index. RESULTS: Among all Vanderbilt cancer patients, metformin was associated with a 22% decrease in overall mortality compared to other oral hypoglycemic medications (HR 0.78; 95% CI 0.69 to 0.88) and with a 39% decrease compared to type 2 diabetes patients on insulin only (HR 0.61; 95% CI 0.50 to 0.73). Diabetic patients on metformin also had a 23% improved survival compared with non-diabetic patients (HR 0.77; 95% CI 0.71 to 0.85). These associations were replicated using the Mayo Clinic EHR data. Many site-specific cancers including breast, colorectal, lung, and prostate demonstrated reduced mortality with metformin use in at least one EHR. CONCLUSIONS: EHR data suggested that the use of metformin was associated with decreased mortality after a cancer diagnosis compared with diabetic and non-diabetic cancer patients not on metformin, indicating its potential as a chemotherapeutic regimen. This study serves as a model for robust and inexpensive validation studies for drug repurposing signals using EHR data. Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Xiaoyang Ruan, Min Jiang 0007, Jamii St Julien, Jeremy L. Warner, Carol Friedman, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 15 |
| 2014 | Identification of Inflammatory Bowel Disease Patients with Steroid-induced Diabetes Mellitus Using an Electronic Health Record
Sivan Kinberg, Lyudmila Ena, Herbert S. Chase, Carol Friedman |
AMIA | 4 |
| 2014 | Combining Heterogeneous Databases to Detect Adverse Drug Reaction
Santiago Vilar, Carol Friedman |
AMIA | 4 |
| 2014 | Developing a Formal Representation for Medication Appropriateness Criteria
Hojjat Salmasian, Tran H. Tran, Carol Friedman |
AMIA | 3 |
| 2014 | A method for controlling complex confounding effects in the detection of adverse drug reactions using electronic health recordsabstractOBJECTIVE: Electronic health records (EHRs) contain information to detect adverse drug reactions (ADRs), as they contain comprehensive clinical information. A major challenge of using comprehensive information involves confounding. We propose a novel data-driven method to identify ADR signals accurately by adjusting for confounders. MATERIALS AND METHODS: We focused on two serious ADRs, rhabdomyolysis and pancreatitis, and used information in 264,155 unique patient records. We identified an ADR using established criteria, selected potential confounders, and then used penalized logistic regressions to estimate confounder-adjusted ADR associations. A reference standard was created to evaluate and compare the precision of the proposed method and four others. RESULTS: Precision was 83.3% for rhabdomyolysis and 60.8% for pancreatitis when using the proposed method, and we identified several drug safety signals that are interesting for further clinical review. DISCUSSION: The proposed method effectively estimated ADR associations after adjusting for confounders. A main cause of error was probably due to the nature of the dataset in that a substantial number of patients had a single visit only and, therefore, it was not possible to determine correctly the appropriate sequence of events for them. It is likely that performance will be improved with use of EHR data that contain more longitudinal records. CONCLUSIONS: This data-driven method is effective in controlling for confounding, resulting in either a higher or similar precision when compared with four comparators, has the unique ability to provide insight into confounders for each specific medication-ADR pair, and can be easily adapted to other EHR systems. Hojjat Salmasian, Santiago Vilar, Herbert S. Chase, Carol Friedman |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | HospSum: Integrating physician discharge notes with coded nursing care data to generate patient-centric summaries
Barbara Di Eugenio, Camillo Lugaresi, Gail M. Keenan, Yves A. Lussier, Jianrong Li, Mike D. Burton, Carol Friedman, Andrew D. Boyd |
AMIA | 7 |
| 2013 | Deriving comorbidities from medical records using Natural Language Processing
Hojjat Salmasian, Daniel Freedberg, Carol Friedman |
AMIA | 3 |
| 2013 | Combing signals from spontaneous reports and electronic health records for detection of adverse drug reactionsabstractOBJECTIVE: Data-mining algorithms that can produce accurate signals of potentially novel adverse drug reactions (ADRs) are a central component of pharmacovigilance. We propose a signal-detection strategy that combines the adverse event reporting system (AERS) of the Food and Drug Administration and electronic health records (EHRs) by requiring signaling in both sources. We claim that this approach leads to improved accuracy of signal detection when the goal is to produce a highly selective ranked set of candidate ADRs. MATERIALS AND METHODS: Our investigation was based on over 4 million AERS reports and information extracted from 1.2 million EHR narratives. Well-established methodologies were used to generate signals from each source. The study focused on ADRs related to three high-profile serious adverse reactions. A reference standard of over 600 established and plausible ADRs was created and used to evaluate the proposed approach against a comparator. RESULTS: The combined signaling system achieved a statistically significant large improvement over AERS (baseline) in the precision of top ranked signals. The average improvement ranged from 31% to almost threefold for different evaluation categories. Using this system, we identified a new association between the agent, rasburicase, and the adverse event, acute pancreatitis, which was supported by clinical review. CONCLUSIONS: The results provide promising initial evidence that combining AERS with EHRs via the framework of replicated signaling can improve the accuracy of signal detection for certain operating scenarios. The use of additional EHR data is required to further evaluate the capacity and limits of this system and to extend the generalizability of these results. Rave Harpaz, Santiago Vilar, William DuMouchel, Hojjat Salmasian, Krystl Haerian, Nigam H. Shah, Herbert S. Chase, Carol Friedman |
J. Am. Medical Informatics Assoc. | 8 |
| 2013 | Natural language processing: State of the art and prospects for significant progress, a workshop sponsored by the National Library of MedicineabstractNatural language processing (NLP) is crucial for advancing healthcare because it is needed to transform relevant information locked in text into structured data that can be used by computer processes aimed at improving patient care and advancing medicine. In light of the importance of NLP to health, the National Library of Medicine (NLM) recently sponsored a workshop to review the state of the art in NLP focusing on text in English, both in biomedicine and in the general language domain. Specific goals of the NLM-sponsored workshop were to identify the current state of the art, grand challenges and specific roadblocks, and to identify effective use and best practices. This paper reports on the main outcomes of the workshop, including an overview of the state of the art, strategies for advancing the field, and obstacles that need to be addressed, resulting in recommendations for a research agenda intended to advance the field. Carol Friedman, Thomas C. Rindflesch, Milton Corn |
J. Biomed. Informatics | 1 |
| 2012 | Methods for Identifying Suicide or Suicidal Ideation in EHRs
Krystl Haerian, Hojjat Salmasian, Carol Friedman |
AMIA | 3 |
| 2012 | Identifying overuse of medications using natural language processing and electronic health records
Hojjat Salmasian, Julian Abrams, Daniel Freedberg, Carol Friedman |
AMIA | 4 |
| 2012 | Electronic health record data suggests metformin improves cancer survival: A new model for drug repurposing studies
Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Carol Friedman, Joshua C. Denny |
AMIA | 10 |
| 2012 | Combining Corpus-derived Sense Profiles with Estimated Frequency Information to Disambiguate Clinical Abbreviations
Hua Xu 0001, Peter D. Stetson, Carol Friedman |
AMIA | 3 |
| 2012 | Drug-drug interaction through molecular structure similarity analysisabstractBACKGROUND: Drug-drug interactions (DDIs) are responsible for many serious adverse events; their detection is crucial for patient safety but is very challenging. Currently, the US Food and Drug Administration and pharmaceutical companies are showing great interest in the development of improved tools for identifying DDIs. METHODS: We present a new methodology applicable on a large scale that identifies novel DDIs based on molecular structural similarity to drugs involved in established DDIs. The underlying assumption is that if drug A and drug B interact to produce a specific biological effect, then drugs similar to drug A (or drug B) are likely to interact with drug B (or drug A) to produce the same effect. DrugBank was used as a resource for collecting 9454 established DDIs. The structural similarity of all pairs of drugs in DrugBank was computed to identify DDI candidates. RESULTS: The methodology was evaluated using as a gold standard the interactions retrieved from the initial DrugBank database. Results demonstrated an overall sensitivity of 0.68, specificity of 0.96, and precision of 0.26. Additionally, the methodology was also evaluated in an independent test using the Micromedex/Drugdex database. CONCLUSION: The proposed methodology is simple, efficient, allows the investigation of large numbers of drugs, and helps highlight the etiology of DDI. A database of 58 403 predicted DDIs with structural evidence is provided as an open resource for investigators seeking to analyze DDIs. Santiago Vilar, Rave Harpaz, Eugenio Uriarte, Lourdes Santana, Raul Rabadan, Carol Friedman |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | A new clustering method for detecting rare senses of abbreviations in clinical notes
Hua Xu 0001, Yonghui Wu 0001, Noémie Elhadad, Peter D. Stetson, Carol Friedman |
J. Biomed. Informatics | 5 |
| 2011 | Facilitating adverse drug event detection in pharmacovigilance databases using molecular structure similarity: application to rhabdomyolysisabstractBACKGROUND: Adverse drug events (ADE) cause considerable harm to patients, and consequently their detection is critical for patient safety. The US Food and Drug Administration maintains an adverse event reporting system (AERS) to facilitate the detection of ADE in drugs. Various data mining approaches have been developed that use AERS to detect signals identifying associations between drugs and ADE. The signals must then be monitored further by domain experts, which is a time-consuming task. OBJECTIVE: To develop a new methodology that combines existing data mining algorithms with chemical information by analysis of molecular fingerprints to enhance initial ADE signals generated from AERS, and to provide a decision support mechanism to facilitate the identification of novel adverse events. RESULTS: The method achieved a significant improvement in precision in identifying known ADE, and a more than twofold signal enhancement when applied to the ADE rhabdomyolysis. The simplicity of the method assists in highlighting the etiology of the ADE by identifying structurally similar drugs. A set of drugs with strong evidence from both AERS and molecular fingerprint-based modeling is constructed for further analysis. CONCLUSION: The results demonstrate that the proposed methodology could be used as a pharmacovigilance decision support tool to facilitate ADE detection. Santiago Vilar, Rave Harpaz, Herbert S. Chase, Stefano Costanzi, Raul Rabadan, Carol Friedman |
J. Am. Medical Informatics Assoc. | 6 |
| 2011 | Deriving a probabilistic syntacto-semantic grammar for biomedicine based on domain-specific terminologies
Carol Friedman |
J. Biomed. Informatics | 2 |
| 2010 | Mining multi-item drug adverse effect associations in spontaneous reporting systemsabstractBACKGROUND: Multi-item adverse drug event (ADE) associations are associations relating multiple drugs to possibly multiple adverse events. The current standard in pharmacovigilance is bivariate association analysis, where each single drug-adverse effect combination is studied separately. The importance and difficulty in the detection of multi-item ADE associations was noted in several prominent pharmacovigilance studies. In this paper we examine the application of a well established data mining method known as association rule mining, which we tailored to the above problem, and demonstrate its value. The method was applied to the FDAs spontaneous adverse event reporting system (AERS) with minimal restrictions and expectations on its output, an experiment that has not been previously done on the scale and generality proposed in this work. RESULTS: Based on a set of 162,744 reports of suspected ADEs reported to AERS and published in the year 2008, our method identified 1167 multi-item ADE associations. A taxonomy that characterizes the associations was developed based on a representative sample. A significant number (67% of the total) of potential multi-item ADE associations identified were characterized and clinically validated by a domain expert as previously recognized ADE associations. Several potentially novel ADEs were also identified. A smaller proportion (4%) of associations were characterized and validated as known drug-drug interactions. CONCLUSIONS: Our findings demonstrate that multi-item ADEs are present and can be extracted from the FDA's adverse effect reporting system using our methodology, suggesting that our method is a valid approach for the initial identification of multi-item ADEs. The study also revealed several limitations and challenges that can be attributed to both the method and quality of data. Rave Harpaz, Herbert S. Chase, Carol Friedman |
BMC Bioinform. | 3 |
| 2010 | Selecting information in electronic health records for knowledge acquisition
Herbert S. Chase, Marianthi Markatou, George Hripcsak, Carol Friedman |
J. Biomed. Informatics | 5 |
| 2009 | Discovering Novel Adverse Drug Events Using Natural Language Processing and Mining of the Electronic Health Record
Carol Friedman |
AIME | 1 |
| 2009 | Generating quality word sense disambiguation test sets based on MeSH indexing
Carol Friedman |
AMIA | 2 |
| 2009 | PhenoGO: an integrated resource for the multiscale mining of clinical and biological dataabstractThe evolving complexity of genome-scale experiments has increasingly centralized the role of a highly computable, accurate, and comprehensive resource spanning multiple biological scales and viewpoints. To provide a resource to meet this need, we have significantly extended the PhenoGO database with gene-disease specific annotations and included an additional ten species. This a computationally-derived resource is primarily intended to provide phenotypic context (cell type, tissue, organ, and disease) for mining existing associations between gene products and GO terms specified in the Gene Ontology Databases Automated natural language processing (BioMedLEE) and computational ontology (PhenOS) methods were used to derive these relationships from the literature, expanding the database with information from ten additional species to include over 600,000 phenotypic contexts spanning eleven species from five GO annotation databases. A comprehensive evaluation evaluating the mappings (n = 300) found precision (positive predictive value) at 85%, and recall (sensitivity) at 76%. Phenotypes are encoded in general purpose ontologies such as Cell Ontology, the Unified Medical Language System, and in specialized ontologies such as the Mouse Anatomy and the Mammalian Phenotype Ontology. A web portal has also been developed, allowing for advanced filtering and querying of the database as well as download of the entire dataset http://www.phenogo.org. Lee T. Sam, Eneida A. Mendonça, Jianrong Li, Judith A. Blake, Carol Friedman, Yves A. Lussier |
BMC Bioinform. | 5 |
| 2009 | Characterizing environmental and phenotypic associations using information theory and electronic health recordsabstractBACKGROUND: The availability of up-to-date, executable, evidence-based medical knowledge is essential for many clinical applications, such as pharmacovigilance, but executable knowledge is costly to obtain and update. Automated acquisition of environmental and phenotypic associations in biomedical and clinical documents using text mining has showed some success. The usefulness of the association knowledge is limited, however, due to the fact that the specific relationships between clinical entities remain unknown. In particular, some associations are indirect relations due to interdependencies among the data. RESULTS: In this work, we develop methods using mutual information (MI) and its property, the data processing inequality (DPI), to help characterize associations that were generated based on use of natural language processing to encode clinical information in narrative patient records followed by statistical methods. Evaluation based on a random sample consisting of two drugs and two diseases indicates an overall precision of 81%. CONCLUSION: This preliminary study demonstrates that the proposed method is effective for helping to characterize phenotypic and environmental associations obtained from clinical reports. George Hripcsak, Carol Friedman |
BMC Bioinform. | 3 |
| 2009 | Research Paper: Syndromic Surveillance Using Ambulatory Electronic Health RecordsabstractOBJECTIVE: To assess the performance of electronic health record data for syndromic surveillance and to assess the feasibility of broadly distributed surveillance. DESIGN: Two systems were developed to identify influenza-like illness and gastrointestinal infectious disease in ambulatory electronic health record data from a network of community health centers. The first system used queries on structured data and was designed for this specific electronic health record. The second used natural language processing of narrative data, but its queries were developed independently from this health record. Both were compared to influenza isolates and to a verified emergency department chief complaint surveillance system. MEASUREMENTS: Lagged cross-correlation and graphs of the three time series. RESULTS: For influenza-like illness, both the structured and narrative data correlated well with the influenza isolates and with the emergency department data, achieving cross-correlations of 0.89 (structured) and 0.84 (narrative) for isolates and 0.93 and 0.89 for emergency department data, and having similar peaks during influenza season. For gastrointestinal infectious disease, the structured data correlated fairly well with the emergency department data (0.81) with a similar peak, but the narrative data correlated less well (0.47). CONCLUSIONS: It is feasible to use electronic health records for syndromic surveillance. The structured data performed best but required knowledge engineering to match the health record data to the queries. The narrative data illustrated the potential performance of a broadly disseminated system and achieved mixed results. George Hripcsak, Nicholas D. Soulakis, Li Li 0062, Frances P. Morrison, Albert M. Lai, Carol Friedman, Neil S. Calman, Farzad Mostashari |
J. Am. Medical Informatics Assoc. | 6 |
| 2009 | Research Paper: Active Computerized Pharmacovigilance Using Natural Language Processing, Statistics, and Electronic Health Records: A Feasibility StudyabstractOBJECTIVE It is vital to detect the full safety profile of a drug throughout its market life. Current pharmacovigilance systems still have substantial limitations, however. The objective of our work is to demonstrate the feasibility of using natural language processing (NLP), the comprehensive Electronic Health Record (EHR), and association statistics for pharmacovigilance purposes. DESIGN Narrative discharge summaries were collected from the Clinical Information System at New York Presbyterian Hospital (NYPH). MedLEE, an NLP system, was applied to the collection to identify medication events and entities which could be potential adverse drug events (ADEs). Co-occurrence statistics with adjusted volume tests were used to detect associations between the two types of entities, to calculate the strengths of the associations, and to determine their cutoff thresholds. Seven drugs/drug classes (ibuprofen, morphine, warfarin, bupropion, paroxetine, rosiglitazone, ACE inhibitors) with known ADEs were selected to evaluate the system. RESULTS One hundred thirty-two potential ADEs were found to be associated with the 7 drugs. Overall recall and precision were 0.75 and 0.31 for known ADEs respectively. Importantly, qualitative evaluation using historic roll back design suggested that novel ADEs could be detected using our system. CONCLUSIONS This study provides a framework for the development of active, high-throughput and prospective systems which could potentially unveil drug safety profiles throughout their entire market life. Our results demonstrate that the framework is feasible although there are some challenging issues. To the best of our knowledge, this is the first study using comprehensive unstructured data from the EHR for pharmacovigilance. George Hripcsak, Marianthi Markatou, Carol Friedman |
J. Am. Medical Informatics Assoc. | 4 |
| 2009 | Research Paper: Methods for Building Sense Inventories of Abbreviations in Clinical NotesabstractOBJECTIVE: To develop methods for building corpus-specific sense inventories of abbreviations occurring in clinical documents. DESIGN: A corpus of internal medicine admission notes was collected and instances of each clinical abbreviation in the corpus were clustered to different sense clusters. One instance from each cluster was manually annotated to generate a final list of senses. Two clustering-based methods (Expectation Maximization--EM and Farthest First--FF) and one random sampling method for sense detection were evaluated using a set of 12 clinical abbreviations. MEASUREMENTS: The clustering-based sense detection methods were evaluated using a set of clinical abbreviations that were manually sense annotated. "Sense Completeness" and "Annotation Cost" were used to measure the performance of different methods. Clustering error rates were also reported for different clustering algorithms. RESULTS: A clustering-based semi-automated method was developed to build corpus-specific sense inventories for abbreviations in hospital admission notes. Evaluation demonstrated that this method could largely reduce manual annotation cost and increase the completeness of sense inventories when compared with a manual annotation method using random samples. CONCLUSION: The authors developed an effective clustering-based method for building corpus-specific sense inventories for abbreviations in a clinical corpus. To the best of the authors knowledge, this is the first time clustering technologies have been used to help building sense inventories of abbreviations in clinical text. The results demonstrated that the clustering-based method performed better than the manual annotation method using random samples for the task of building sense inventories of clinical abbreviations. Hua Xu 0001, Peter D. Stetson, Carol Friedman |
J. Am. Medical Informatics Assoc. | 3 |
| 2008 | Word Sense Disambiguation via Semantic Type Classification
Carol Friedman |
AMIA | 2 |
| 2008 | Comparing ICD9-Encoded Diagnoses and NLP-Processed Discharge Summaries for Clinical Trials Pre-Screening: A Case Study
Herbert S. Chase, Chintan Patel, Carol Friedman, Chunhua Weng |
AMIA | 4 |
| 2008 | Automated Knowledge Acquisition from Clinical Narrative Reports
Amy E. Chused, Noémie Elhadad, Carol Friedman, Marianthi Markatou |
AMIA | 4 |
| 2008 | Methods for Building Sense Inventories of Abbreviations in Clinical Notes
Hua Xu 0001, Peter D. Stetson, Carol Friedman |
AMIA | 3 |
| 2008 | Semantic reclassification of the UMLS conceptsabstractAbstract Summary: Accurate semantic classification is valuable for text mining and knowledge-based tasks that perform inference based on semantic classes. To benefit applications using the semantic classification of the Unified Medical Language System (UMLS) concepts, we automatically reclassified the concepts based on their lexical and contextual features. The new classification is useful for auditing the original UMLS semantic classification and for building biomedical text mining applications. Availability: http://www.dbmi.columbia.edu/~juf7002/reclassify_production Contact: [email protected] Supplementary information: Supplementary data is available at http://www.dbmi.columbia.edu/~juf7002/reclassify_production. Carol Friedman |
Bioinform. | 2 |
| 2008 | Research Paper: Automated Acquisition of Disease-Drug Knowledge from Biomedical and Clinical Documents: An Initial StudyabstractOBJECTIVE: Explore the automated acquisition of knowledge in biomedical and clinical documents using text mining and statistical techniques to identify disease-drug associations. DESIGN: Biomedical literature and clinical narratives from the patient record were mined to gather knowledge about disease-drug associations. Two NLP systems, BioMedLEE and MedLEE, were applied to Medline articles and discharge summaries, respectively. Disease and drug entities were identified using the NLP systems in addition to MeSH annotations for the Medline articles. Focusing on eight diseases, co-occurrence statistics were applied to compute and evaluate the strength of association between each disease and relevant drugs. RESULTS: Ranked lists of disease-drug pairs were generated and cutoffs calculated for identifying stronger associations among these pairs for further analysis. Differences and similarities between the text sources (i.e., biomedical literature and patient record) and annotations (i.e., MeSH and NLP-extracted UMLS concepts) with regards to disease-drug knowledge were observed. CONCLUSION: This paper presents a method for acquiring disease-specific knowledge and a feasibility study of the method. The method is based on applying a combination of NLP and statistical techniques to both biomedical and clinical documents. The approach enabled extraction of knowledge about the drugs clinicians are using for patients with specific diseases based on the patient record, while it is also acquired knowledge of drugs frequently involved in controlled trials for those same diseases. In comparing the disease-drug associations, we found the results to be appropriate: the two text sources contained consistent as well as complementary knowledge, and manual review of the top five disease-drug associations by a medical expert supported their correctness across the diseases. Elizabeth S. Chen, George Hripcsak, Hua Xu 0001, Marianthi Markatou, Carol Friedman |
J. Am. Medical Informatics Assoc. | 5 |
| 2007 | Detection of Practice Pattern Trends through Natural Language Processing of Clinical Narratives and Biomedical Literature
Elizabeth S. Chen, Peter D. Stetson, Yves A. Lussier, Marianthi Markatou, George Hripcsak, Carol Friedman |
AMIA | 6 |
| 2007 | Combining Contextual and Lexical Features to Classify UMLS Concepts
Carol Friedman |
AMIA | 2 |
| 2007 | Information-Theoretic Classification of SNOMED Improves the Organization of Context-Sensitive Excerpts from Cochrane Reviews
Lee T. Sam, Tara Borlawsky, Ying Tao, Jianrong Li, Carol Friedman, Barry Smith 0001, Yves A. Lussier |
AMIA | 5 |
| 2007 | A Study of Abbreviations in Clinical Notes
Hua Xu 0001, Peter D. Stetson, Carol Friedman |
AMIA | 3 |
| 2007 | Gene symbol disambiguation using knowledge-based profilesabstractMOTIVATION: The ambiguity of biomedical entities, particularly of gene symbols, is a big challenge for text-mining systems in the biomedical domain. Existing knowledge sources, such as Entrez Gene and the MEDLINE database, contain information concerning the characteristics of a particular gene that could be used to disambiguate gene symbols. RESULTS: For each gene, we create a profile with different types of information automatically extracted from related MEDLINE abstracts and readily available annotated knowledge sources. We apply the gene profiles to the disambiguation task via an information retrieval method, which ranks the similarity scores between the context where the ambiguous gene is mentioned, and candidate gene profiles. The gene profile with the highest similarity score is then chosen as the correct sense. We evaluated the method on three automatically generated testing sets of mouse, fly and yeast organisms, respectively. The method achieved the highest precision of 93.9% for the mouse, 77.8% for the fly and 89.5% for the yeast. AVAILABILITY: The testing data sets and disambiguation programs are available at http://www.dbmi.columbia.edu/~hux7002/gsd2006 Hua Xu 0001, Jungwei Fan 0001, George Hripcsak, Eneida A. Mendonça, Marianthi Markatou, Carol Friedman |
Bioinform. | 6 |
| 2007 | Using contextual and lexical features to restructure and validate the classification of biomedical conceptsabstractBACKGROUND: Biomedical ontologies are critical for integration of data from diverse sources and for use by knowledge-based biomedical applications, especially natural language processing as well as associated mining and reasoning systems. The effectiveness of these systems is heavily dependent on the quality of the ontological terms and their classifications. To assist in developing and maintaining the ontologies objectively, we propose automatic approaches to classify and/or validate their semantic categories. In previous work, we developed an approach using contextual syntactic features obtained from a large domain corpus to reclassify and validate concepts of the Unified Medical Language System (UMLS), a comprehensive resource of biomedical terminology. In this paper, we introduce another classification approach based on words of the concept strings and compare it to the contextual syntactic approach. RESULTS: The string-based approach achieved an error rate of 0.143, with a mean reciprocal rank of 0.907. The context-based and string-based approaches were found to be complementary, and the error rate was reduced further by applying a linear combination of the two classifiers. The advantage of combining the two approaches was especially manifested on test data with sufficient contextual features, achieving the lowest error rate of 0.055 and a mean reciprocal rank of 0.969. CONCLUSION: The lexical features provide another semantic dimension in addition to syntactic contextual features that support the classification of ontological concepts. The classification errors of each dimension can be further reduced through appropriate combination of the complementary classifiers. Jungwei Fan 0001, Hua Xu 0001, Carol Friedman |
BMC Bioinform. | 3 |
| 2007 | Research Paper: Semantic Classification of Biomedical Concepts Using Distributional SimilarityabstractOBJECTIVE: To develop an automated, high-throughput, and reproducible method for reclassifying and validating ontological concepts for natural language processing applications. DESIGN: We developed a distributional similarity approach to classify the Unified Medical Language System (UMLS) concepts. Classification models were built for seven broad biomedically relevant semantic classes created by grouping subsets of the UMLS semantic types. We used contextual features based on syntactic properties obtained from two different large corpora and used alpha-skew divergence as the similarity measure. MEASUREMENTS: The testing sets were automatically generated based on the changes by the National Library of Medicine to the semantic classification of concepts from the UMLS 2005AA to the 2006AA release. Error rates were calculated and a misclassification analysis was performed. RESULTS: The estimated lowest error rates were 0.198 and 0.116 when considering the correct classification to be covered by our top prediction and top 2 predictions, respectively. CONCLUSION: The results demonstrated that the distributional similarity approach can recommend high level semantic classification suitable for use in natural language processing. Carol Friedman |
J. Am. Medical Informatics Assoc. | 2 |
| 2007 | Natural language processing and visualization in the molecular imaging domain
P. Karina Tulipano, Ying Tao, William S. Millar, Pat Zanzonico, Katherine Kolbert, Hua Xu 0001, Hong Yu 0001, Lifeng Chen, Yves A. Lussier, Carol Friedman |
J. Biomed. Informatics | 10 |
| 2006 | Generating Executable Knowledge for Evidence-Based Medicine Using Natural Language and Semantic Processing
Tara Borlawsky, Carol Friedman, Yves A. Lussier |
AMIA | 2 |
| 2006 | Disseminating Natural Language Processed Clinical Narratives
Elizabeth S. Chen, George Hripcsak, Carol Friedman |
AMIA | 3 |
| 2006 | Exploiting Semantic Relations for Literature-Based Discovery
Dimitar Hristovski, Carol Friedman, Thomas C. Rindflesch, Borut Peterlin |
AMIA | 2 |
| 2006 | ZebraHunter: Searching Rare Medical Diagnoses and Retrieving Relevant Citations
Eric Silfen, Chintan Patel, Eneida A. Mendonça, Carol Friedman |
AMIA | 4 |
| 2006 | A Natural Language Processing (NLP) Tool to Assist in the Curation Of the Laboratory Mouse Tumor Biology Database
Hua Xu 0001, Debra M. Krupke, Judith A. Blake, Carol Friedman |
AMIA | 4 |
| 2006 | Bio-Ontology and text: bridging the modeling gapabstractMOTIVATION: Natural language processing (NLP) techniques are increasingly being used in biology to automate the capture of new biological discoveries in text, which are being reported at a rapid rate. Yet, information represented in NLP data structures is classically very different from information organized with ontologies as found in model organisms or genetic databases. To facilitate the computational reuse and integration of information buried in unstructured text with that of genetic databases, we propose and evaluate a translational schema that represents a comprehensive set of phenotypic and genetic entities, as well as their closely related biomedical entities and relations as expressed in natural language. In addition, the schema connects different scales of biological information, and provides mappings from the textual information to existing ontologies, which are essential in biology for integration, organization, dissemination and knowledge management of heterogeneous phenotypic information. A common comprehensive representation for otherwise heterogeneous phenotypic and genetic datasets, such as the one proposed, is critical for advancing systems biology because it enables acquisition and reuse of unprecedented volumes of diverse types of knowledge and information from text. RESULTS: A novel representational schema, PGschema, was developed that enables translation of phenotypic, genetic and their closely related information found in textual narratives to a well-defined data structure comprising phenotypic and genetic concepts from established ontologies along with modifiers and relationships. Evaluation for coverage of a selected set of entities showed that 90% of the information could be represented (95% confidence interval: 86-93%; n = 268). Moreover, PGschema can be expressed automatically in an XML format using natural language techniques to process the text. To our knowledge, we are providing the first evaluation of a translational schema for NLP that contains declarative knowledge about genes and their associated biomedical data (e.g. phenotypes). AVAILABILITY: http://zellig.cpmc.columbia.edu/PGschema Carol Friedman, Tara Borlawsky, Lyudmila Shagina, H. Rosie Xing, Yves A. Lussier |
Bioinform. | 1 |
| 2006 | Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issuesabstractBACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology. Hua Xu 0001, Marianthi Markatou, Rositsa Dimova, Carol Friedman |
BMC Bioinform. | 5 |
| 2006 | Research Paper: Human and Automated Coding of Rehabilitation Discharge Summaries According to the International Classification of Functioning, Disability, and HealthabstractOBJECTIVE: The International Classification of Functioning, Disability, and Health (ICF) is designed to provide a common language and framework for describing health and health-related states. The goal of this research was to investigate human and automated coding of functional status information using the ICF framework. DESIGN: The authors extended an existing natural language processing (NLP) system to encode rehabilitation discharge summaries according to the ICF. MEASUREMENTS: The authors conducted a formal evaluation, comparing the coding performed by expert coders, non-expert coders, and the NLP system. RESULTS: Automated coding can be used to assign codes using the ICF, with results similar to those obtained by human coders, at least for the selection of ICF code and assignment of the performance qualifier. Coders achieved high agreement on ICF code assignment. CONCLUSION: This research is a key next step in the development of the ICF as a sensitive and universal classification of functional status information. It is worthwhile to continue to investigate automated ICF coding. Rita Kukafka, Michael E. Bales, Ann Burkhardt, Carol Friedman |
J. Am. Medical Informatics Assoc. | 4 |
| 2006 | Research Paper: Quantitative Assessment of Dictionary-based Protein Named Entity TaggingabstractOBJECTIVE: Natural language processing (NLP) approaches have been explored to manage and mine information recorded in biological literature. A critical step for biological literature mining is biological named entity tagging (BNET) that identifies names mentioned in text and normalizes them with entries in biological databases. The aim of this study was to provide quantitative assessment of the complexity of BNET on protein entities through BioThesaurus, a thesaurus of gene/protein names for UniProt knowledgebase (UniProtKB) entries that was acquired using online resources. METHODS: We evaluated the complexity through several perspectives: ambiguity (i.e., the number of genes/proteins represented by one name), synonymy (i.e., the number of names associated with the same gene/protein), and coverage (i.e., the percentage of gene/protein names in text included in the thesaurus). We also normalized names in BioThesaurus and measures were obtained twice, once before normalization and once after. RESULTS: The current version of BioThesaurus has over 2.6 million names or 2.1 million normalized names covering more than 1.8 million UniProtKB entries. The average synonymy is 3.53 (2.86 after normalization), ambiguity is 2.31 before normalization and 2.32 after, while the coverage is 94.0% based on the BioCreAtive data set comprising MEDLINE abstracts containing genes/proteins. CONCLUSION: The study indicated that names for genes/proteins are highly ambiguous and there are usually multiple names for the same gene or protein. It also demonstrated that most gene/protein names appearing in text can be found in BioThesaurus. Zhang-Zhi Hu, Manabu Torii, Cathy H. Wu, Carol Friedman |
J. Am. Medical Informatics Assoc. | 5 |
| 2006 | Terminology model discovery using natural language processing and visualization techniques
Li Zhou 0007, Ying Tao, James J. Cimino, Elizabeth S. Chen, Yves A. Lussier, George Hripcsak, Carol Friedman |
J. Biomed. Informatics | 8 |
| 2005 | Extending a Medical Language Processing System to the Functional Status Domain
Michael E. Bales, Rita Kukafka, Ann Burkhardt, Carol Friedman |
AMIA | 4 |
| 2005 | Natural Language Processing in the Molecular Imaging Domain
P. Karina Tulipano, Ying Tao, Pat Zanzonico, Katherine Kolbert, Yves A. Lussier, Carol Friedman |
AMIA | 6 |
| 2005 | System Architecture for Temporal Information Extraction, Representationand Reasoning in Clinical Narrative Reports
Li Zhou 0007, Carol Friedman, Simon Parsons, George Hripcsak |
AMIA | 2 |
| 2005 | Gene name ambiguity of eukaryotic nomenclaturesabstractMOTIVATION: With more and more scientific literature published online, the effective management and reuse of this knowledge has become problematic. Natural language processing (NLP) may be a potential solution by extracting, structuring and organizing biomedical information in online literature in a timely manner. One essential task is to recognize and identify genomic entities in text. 'Recognition' can be accomplished using pattern matching and machine learning. But for 'identification' these techniques are not adequate. In order to identify genomic entities, NLP needs a comprehensive resource that specifies and classifies genomic entities as they occur in text and that associates them with normalized terms and also unique identifiers so that the extracted entities are well defined. Online organism databases are an excellent resource to create such a lexical resource. However, gene name ambiguity is a serious problem because it affects the appropriate identification of gene entities. In this paper, we explore the extent of the problem and suggest ways to address it. RESULTS: We obtained gene information from 21 organisms and quantified naming ambiguities within species, across species, with English words and with medical terms. When the case (of letters) was retained, official symbols displayed negligible intra-species ambiguity (0.02%) and modest ambiguities with general English words (0.57%) and medical terms (1.01%). In contrast, the across-species ambiguity was high (14.20%). The inclusion of gene synonyms increased intra-species ambiguity substantially and full names contributed greatly to gene-medical-term ambiguity. A comprehensive lexical resource that covers gene information for the 21 organisms was then created and used to identify gene names by using a straightforward string matching program to process 45,000 abstracts associated with the mouse model organism while ignoring case and gene names that were also English words. We found that 85.1% of correctly retrieved mouse genes were ambiguous with other gene names. When gene names that were also English words were included, 233% additional 'gene' instances were retrieved, most of which were false positives. We also found that authors prefer to use synonyms (74.7%) to official symbols (17.7%) or full names (7.6%) in their publications. CONTACT: [email protected] Lifeng Chen, Carol Friedman |
Bioinform. | 3 |
| 2005 | Visualizing information across multidimensional post-genomic structured and textual databasesabstractMOTIVATION: Visualizing relationships among biological information to facilitate understanding is crucial to biological research during the post-genomic era. Although different systems have been developed to view gene-phenotype relationships for specific databases, very few have been designed specifically as a general flexible tool for visualizing multidimensional genotypic and phenotypic information together. Our goal is to develop a method for visualizing multidimensional genotypic and phenotypic information and a model that unifies different biological databases in order to present the integrated knowledge using a uniform interface. RESULTS: We developed a novel, flexible and generalizable visualization tool, called PhenoGenesviewer (PGviewer), which in this paper was used to display gene-phenotype relationships from a human-curated database (OMIM) and from an automatic method using a Natural Language Processing tool called BioMedLEE. Data obtained from multiple databases were first integrated into a uniform structure and then organized by PGviewer. PGviewer provides a flexible query interface that allows dynamic selection and ordering of any desired dimension in the databases. Based on users' queries, results can be visualized using hierarchical expandable trees that present views specified by users according to their research interests. We believe that this method, which allows users to dynamically organize and visualize multiple dimensions, is a potentially powerful and promising tool that should substantially facilitate biological research. AVAILABILITY: PhenogenesViewer as well as its support and tutorial are available at http://www.dbmi.columbia.edu/pgviewer/ CONTACT: [email protected]. Ying Tao, Carol Friedman, Yves A. Lussier |
Bioinform. | 2 |
| 2005 | Extracting information on pneumonia in infants using natural language processing of radiology reports
Eneida A. Mendonça, Janet Haas, Lyudmila Shagina, Elaine Larson, Carol Friedman |
J. Biomed. Informatics | 5 |
| 2004 | Probabilistic inference of molecular networks from noisy data sourcesabstractInformation on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources. Ivan Iossifov, Michael Krauthammer, Carol Friedman, Vasileios Hatzivassiloglou, Joel S. Bader, Kevin P. White, Andrey Rzhetsky |
Bioinform. | 3 |
| 2004 | Research Paper: Automated Encoding of Clinical Documents Based on Natural Language ProcessingabstractOBJECTIVE: The aim of this study was to develop a method based on natural language processing (NLP) that automatically maps an entire clinical document to codes with modifiers and to quantitatively evaluate the method. METHODS: An existing NLP system, MedLEE, was adapted to automatically generate codes. The method involves matching of structured output generated by MedLEE consisting of findings and modifiers to obtain the most specific code. Recall and precision applied to Unified Medical Language System (UMLS) coding were evaluated in two separate studies. Recall was measured using a test set of 150 randomly selected sentences, which were processed using MedLEE. Results were compared with a reference standard determined manually by seven experts. Precision was measured using a second test set of 150 randomly selected sentences from which UMLS codes were automatically generated by the method and then validated by experts. RESULTS: Recall of the system for UMLS coding of all terms was .77 (95% CI.72-.81), and for coding terms that had corresponding UMLS codes recall was .83 (.79-.87). Recall of the system for extracting all terms was .84 (.81-.88). Recall of the experts ranged from .69 to .91 for extracting terms. The precision of the system was .89 (.87-.91), and precision of the experts ranged from .61 to .91. CONCLUSION: Extraction of relevant clinical information and UMLS coding were accomplished using a method based on NLP. The method appeared to be comparable to or better than six experts. The advantage of the method is that it maps text to codes along with other related information, rendering the coded output suitable for effective retrieval. Carol Friedman, Lyudmila Shagina, Yves A. Lussier, George Hripcsak |
J. Am. Medical Informatics Assoc. | 1 |
| 2004 | Research Paper: A Multi-aspect Comparison Study of Supervised Word Sense DisambiguationabstractOBJECTIVE: The aim of this study was to investigate relations among different aspects in supervised word sense disambiguation (WSD; supervised machine learning for disambiguating the sense of a term in a context) and compare supervised WSD in the biomedical domain with that in the general English domain. METHODS: The study involves three data sets (a biomedical abbreviation data set, a general biomedical term data set, and a general English data set). The authors implemented three machine-learning algorithms, including (1) naïve Bayes (NBL) and decision lists (TDLL), (2) their adaptation of decision lists (ODLL), and (3) their mixed supervised learning (MSL). There were six feature representations (various combinations of collocations, bag of words, oriented bag of words, etc.) and five window sizes (2, 4, 6, 8, and 10). RESULTS: Supervised WSD is suitable only when there are enough sense-tagged instances with at least a few dozens of instances for each sense. Collocations combined with neighboring words are appropriate selections for the context. For terms with unrelated biomedical senses, a large window size such as the whole paragraph should be used, while for general English words a moderate window size between 4 and 10 should be used. The performance of the authors' implementation of decision list classifiers for abbreviations was better than that of traditional decision list classifiers. However, the opposite held for the other two sets. Also, the authors' mixed supervised learning was stable and generally better than others for all sets. CONCLUSION: From this study, it was found that different aspects of supervised WSD depend on each other. The experiment method presented in the study can be used to select the best supervised WSD classifier for each ambiguous term. Virginia Teller, Carol Friedman |
J. Am. Medical Informatics Assoc. | 3 |
| 2004 | Introduction: named entity recognition in biomedicine
Sophia Ananiadou, Carol Friedman, Jun'ichi Tsujii |
J. Biomed. Informatics | 2 |
| 2003 | Natural Language Processing Challenges in HIV/AIDS Clinic Notes
Sookyung Hyun, Suzanne Bakken, Carol Friedman, Stephen B. Johnson |
AMIA | 3 |
| 2003 | A Native XML Database Design for Clinical Document Research
Stephen B. Johnson, David A. Campbell, Michael Krauthammer, P. Karina Tulipano, Eneida A. Mendonça, Carol Friedman, George Hripcsak |
AMIA | 6 |
| 2003 | Facilitating Research in Pathology using Natural Language Processing
Hua Xu 0001, Carol Friedman |
AMIA | 2 |
| 2003 | A vocabulary development and visualization tool based on natural language processing and the mining of textual patient reports
Carol Friedman, Lyudmila Shagina |
J. Biomed. Informatics | 1 |
| 2002 | A comparison of the Charlson comorbidities derived from medical language processing and administrative data
Jen-Hsiang Chuang, Carol Friedman, George Hripcsak |
AMIA | 2 |
| 2002 | Representing nested semantic information in a linear string of text using XML
Michael Krauthammer, Stephen B. Johnson, George Hripcsak, David A. Campbell, Carol Friedman |
AMIA | 5 |
| 2002 | A study of abbreviations in MEDLINE abstracts
Alan R. Aronson, Carol Friedman |
AMIA | 3 |
| 2002 | Automatic extraction of gene and protein synonyms from MEDLINE and journal articles
Hong Yu 0001, Vasileios Hatzivassiloglou, Carol Friedman, Andrey Rzhetsky, W. John Wilbur |
AMIA | 3 |
| 2002 | Of truth and pathways: chasing bits of information through myriads of articlesabstractKnowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data. Michael Krauthammer, Pauline Kra, Ivan Iossifov, Shawn M. Gomez, George Hripcsak, Vasileios Hatzivassiloglou, Carol Friedman, Andrey Rzhetsky |
ISMB | 7 |
| 2002 | Research Paper: Automatic Resolution of Ambiguous Terms Based on Machine Learning and Conceptual Relations in the UMLSabstractUNLABELLED: Motivation. The UMLS has been used in natural language processing applications such as information retrieval and information extraction systems. The mapping of free-text to UMLS concepts is important for these applications. To improve the mapping, we need a method to disambiguate terms that possess multiple UMLS concepts. In the general English domain, machine-learning techniques have been applied to sense-tagged corpora, in which senses (or concepts) of ambiguous terms have been annotated (mostly manually). Sense disambiguation classifiers are then derived to determine senses (or concepts) of those ambiguous terms automatically. However, manual annotation of a corpus is an expensive task. We propose an automatic method that constructs sense-tagged corpora for ambiguous terms in the UMLS using MEDLINE abstracts. METHODS: For a term W that represents multiple UMLS concepts, a collection of MEDLINE abstracts that contain W is extracted. For each abstract in the collection, occurrences of concepts that have relations with W as defined in the UMLS are automatically identified. A sense-tagged corpus, in which senses of W are annotated, is then derived based on those identified concepts. The method was evaluated on a set of 35 frequently occurring ambiguous biomedical abbreviations using a gold standard set that was automatically derived. The quality of the derived sense-tagged corpus was measured using precision and recall. RESULTS: The derived sense-tagged corpus had an overall precision of 92.9% and an overall recall of 47.4%. After removing rare senses and ignoring abbreviations with closely related senses, the overall precision was 96.8% and the overall recall was 50.6%. CONCLUSIONS: UMLS conceptual relations and MEDLINE abstracts can be used to automatically acquire knowledge needed for resolving ambiguity when mapping free-text to UMLS concepts. Stephen B. Johnson, Carol Friedman |
J. Am. Medical Informatics Assoc. | 3 |
| 2002 | Research Paper: Mapping Abbreviations to Full Forms in Biomedical ArticlesabstractOBJECTIVE: To develop methods that automatically map abbreviations to their full forms in biomedical articles. METHODS: The authors developed two methods of mapping defined and undefined abbreviations (defined abbreviations are paired with their full forms in the articles, whereas undefined ones are not). For defined abbreviations, they developed a set of pattern-matching rules to map an abbreviation to its full form and implemented the rules into a software program, AbbRE (for "abbreviation recognition and extraction"). Using the opinions of domain experts as a reference standard, they evaluated the recall and precision of AbbRE for defined abbreviations in ten biomedical articles randomly selected from the ten most frequently cited medical and biological journals. They also measured the percentage of undefined abbreviations in the same set of articles, and they investigated whether they could map undefined abbreviations to any of four public abbreviation databases (GenBank LocusLink, SWISSPROT, LRABR of the UMLS Specialist Lexicon, and BioABACUS). RESULTS: AbbRE had an average 0.70 recall and 0.95 precision for the defined abbreviations. The authors found that an average of 25 percent of abbreviations were defined in biomedical articles and that of a randomly selected subset of undefined abbreviations, 68 percent could be mapped to any of four abbreviation databases. They also found that many abbreviations are ambiguous (i.e., they map to more than one full form in abbreviation databases). CONCLUSION: AbbRE is efficient for mapping defined abbreviations. To couple AbbRE with abbreviation databases for the mapping of undefined abbreviations, not only exhaustive abbreviation databases but also a method to resolve the ambiguity of abbreviations in the databases are needed. Hong Yu 0001, George Hripcsak, Carol Friedman |
J. Am. Medical Informatics Assoc. | 3 |
| 2002 | Editorial
Carol Friedman |
J. Biomed. Informatics | 1 |
| 2002 | Two biomedical sublanguages: a description based on the theories of Zellig Harris
Carol Friedman, Pauline Kra, Andrey Rzhetsky |
J. Biomed. Informatics | 1 |
| 2001 | Evaluating the UMLS as a source of lexical knowledge for medical language processing
Carol Friedman, Lyudmila Shagina, Stephen B. Johnson, George Hripcsak |
AMIA | 1 |
| 2001 | Linking Protein Interaction Data to the MESH Hierarchy
Michael Krauthammer, Pauline Kra, Carol Friedman |
AMIA | 3 |
| 2001 | A study of abbreviations in the UMLS
Yves A. Lussier, Carol Friedman |
AMIA | 3 |
| 2001 | Automating SNOMED coding using medical language understanding: a feasibility study
Yves A. Lussier, Lyudmila Shagina, Carol Friedman |
AMIA | 3 |
| 2001 | Disambiguating Ambiguous Biomedical Terms in Biomedical Narrative Text: An Unsupervised Method
Yves A. Lussier, Carol Friedman |
J. Biomed. Informatics | 3 |
| 2000 | Limited parsing of notational text visit notes: ad-hoc vs. NLP approaches
Randolph C. Barrows Jr., M. Busuioc, Carol Friedman |
AMIA | 3 |
| 2000 | A broad-coverage natural language processing system
Carol Friedman |
AMIA | 1 |
| 2000 | Using BLAST, A DNA and Protein Sequence Comparison Tool, for Finding Gene and Protein Names in Journal Articles
Michael Krauthammer, Andrey Rzhetsky, Pavel Morozov, Carol Friedman |
AMIA | 4 |
| 2000 | A method for vocabulary development and visualization based on medical language processing and XML
Carol Friedman |
AMIA | 2 |
| 2000 | Automating ICD-9-CM Encoding Using Medical Language Processing: A Feasibility Study
Yves A. Lussier, Lyudmila Shagina, Carol Friedman |
AMIA | 3 |
| 2000 | Knowledge-driven Highlighting of Clinical Texts: Does it Help or Distract?
Irina Shablinsky, Justin Starren, Carol Friedman |
AMIA | 3 |
| 2000 | A knowledge model for analysis and simulation of regulatory networksabstractMOTIVATION: In order to aid in hypothesis-driven experimental gene discovery, we are designing a computer application for the automatic retrieval of signal transduction data from electronic versions of scientific publications using natural language processing (NLP) techniques, as well as for visualizing and editing representations of regulatory systems. These systems describe both signal transduction and biochemical pathways within complex multicellular organisms, yeast, and bacteria. This computer application in turn requires the development of a domain-specific ontology, or knowledge model. RESULTS: We introduce an ontological model for the representation of biological knowledge related to regulatory networks in vertebrates. We outline a taxonomy of the concepts, define their 'whole-to-part' relationships, describe the properties of major concepts, and outline a set of the most important axioms. The ontology is partially realized in a computer system designed to aid researchers in biology and medicine in visualizing and editing a representation of a signal transduction system. Andrey Rzhetsky, Tomohiro Koike, Sergey Kalachikov, Shawn M. Gomez, Michael Krauthammer, Sabina H. Kaplan, Pauline Kra, James J. Russo, Carol Friedman |
Bioinform. | 9 |
| 2000 | Coding Neuroradiology Reports for the Northern Manhattan Stroke Study: A Comparison of Natural Language Processing and Manual Review
Jacob S. Elkins, Carol Friedman, Bernadette Boden-Albala, Ralph L. Sacco, George Hripcsak |
Comput. Biomed. Res. | 2 |
| 1999 | Automating a severity score guideline for community-acquired pneumonia employing medical language processing of discharge summaries
Carol Friedman, Charles Knirsch, Lyudmila Shagina, George Hripcsak |
AMIA | 1 |
| 1999 | A Knowledge Model for Analysis and Simulation of Regulatory Based on Information in Electronic Publications
Andrey Rzhetsky, Tomohiro Koike, Sergey Kalachikov, Pauline Kra, Carol Friedman |
AMIA | 5 |
| 1999 | What do ER physicians really want? A method for elucidating ER information needs
Irina Shablinsky, Justin Starren, Carol Friedman |
AMIA | 3 |
| 1999 | Representing genomic knowledge in the UMLS semantic network
Hong Yu 0001, Carol Friedman, Andrey Rzhetsky, Pauline Kra |
AMIA | 2 |
| 1999 | Research Paper: Representing Information in Patient Reports Using Natural Language Processing and the Extensible Markup LanguageabstractOBJECTIVE: To design a document model that provides reliable and efficient access to clinical information in patient reports for a broad range of clinical applications, and to implement an automated method using natural language processing that maps textual reports to a form consistent with the model. METHODS: A document model that encodes structured clinical information in patient reports while retaining the original contents was designed using the extensible markup language (XML), and a document type definition (DTD) was created. An existing natural language processor (NLP) was modified to generate output consistent with the model. Two hundred reports were processed using the modified NLP system, and the XML output that was generated was validated using an XML validating parser. RESULTS: The modified NLP system successfully processed all 200 reports. The output of one report was invalid, and 199 reports were valid XML forms consistent with the DTD. CONCLUSIONS: Natural language processing can be used to automatically create an enriched document that contains a structured component whose elements are linked to portions of the original textual report. This integrated document model provides a representation where documents containing specific information can be accurately and efficiently retrieved by querying the structured components. If manual review of the documents is desired, the salient information in the original reports can also be identified and highlighted. Using an XML model of tagging provides an additional benefit in that software tools that manipulate XML documents are readily available. Carol Friedman, George Hripcsak, Lyudmila Shagina |
J. Am. Medical Informatics Assoc. | 1 |
| 1999 | Research Paper: A Reliability Study for Evaluating Information Extraction from Radiology ReportsabstractGOAL: To assess the reliability of a reference standard for an information extraction task. SETTING: Twenty-four physician raters from two sites and two specialties judged whether clinical conditions were present based on reading chest radiograph reports. METHODS: Variance components, generalizability (reliability) coefficients, and the number of expert raters needed to generate a reliable reference standard were estimated. RESULTS: Per-rater reliability averaged across conditions was 0.80 (95% CI, 0.79-0.81). Reliability for the nine individual conditions varied from 0.67 to 0.97, with central line presence and pneumothorax the most reliable, and pleural effusion (excluding CHF) and pneumonia the least reliable. One to two raters were needed to achieve a reliability of 0.70, and six raters, on average, were required to achieve a reliability of 0.95. This was far more reliable than a previously published per-rater reliability of 0.19 for a more complex task. Differences between sites were attributable to changes to the condition definitions. CONCLUSION: In these evaluations, physician raters were able to judge very reliably the presence of clinical conditions based on text reports. Once the reliability of a specific rater is confirmed, it would be possible for that rater to create a reference standard reliable enough to assess aggregate measures on a system. Six raters would be needed to create a reference standard sufficient to assess a system on a case-by-case basis. These results should help evaluators design future information extraction studies for natural language processors and other knowledge-based systems. George Hripcsak, Gilad J. Kuperman, Carol Friedman, Daniel F. Heitjan |
J. Am. Medical Informatics Assoc. | 3 |
| 1998 | An evaluation of natural language processing methodologies
Carol Friedman, George Hripcsak, Irina Shablinsky |
AMIA | 1 |
| 1998 | Continuous-Speech Structured Reporting
David Rosenthal, Carol Friedman |
AMIA | 2 |
| 1998 | Natural Language as a Tool in the Development of a Controlled Vocabulary
Adam B. Wilcox, Carol Friedman, George Hripcsak |
AMIA | 2 |
| 1997 | Towards a comprehensive medical language processing system: methods and issues
Carol Friedman |
AMIA | 1 |
| 1997 | Identification of findings suspicious for breast cancer based on natural language processing of mammogram reports
Nilesh L. Jain, Carol Friedman |
AMIA | 2 |
| 1995 | Research Paper: The Canon Group's Effort: Working Toward a Merged ModelabstractOBJECTIVE: To develop a representational schema for clinical data for use in exchanging data and applications, using a collaborative approach. DESIGN: Representational models for clinical radiology were independently developed manually by several Canon Group members who had diverse application interests, using sample reports. These models were merged into one common model through an iterative process by means of workshops, meetings, and electronic mail. RESULTS: A core merged model for radiologic findings present in a set of reports that subsumed the models that were developed independently. CONCLUSIONS: The Canon Group's modeling effort focused on a collaborative approach to developing a representational schema for clinical concepts, using chest radiography reports as the initial experiment. This effort resulted in a core model that represents a consensus. Further efforts in modeling will extend the representational coverage and will also address issues such as scalability, automation, evaluation, and support of the collaborative effort. Carol Friedman, Stanley M. Huff, William R. Hersh, Edward Pattison-Gordon, James J. Cimino |
J. Am. Medical Informatics Assoc. | 1 |
| 1995 | Natural language processing in an operational clinical information systemabstractAbstract This paper describes a natural language text extraction system, called MEDLEE, that has been applied to the medical domain. The system extracts, structures, and encodes clinical information from textual patient reports. It was integrated with the Clinical Information System (CIS), which was developed at Columbia-Presbyterian Medical Center (CPMC) to help improve patient care. MEDLEE is currently used on a daily basis to routinely process radiological reports of patients at CPMC. In order to describe how the natural language system was made compatible with the existing CIS, this paper will also discuss engineering issues which involve performance, robustness, and accessibility of the data from the end users' viewpoint. Also described are the three evaluations that have been performed on the system. The first evaluation was useful primarily for further refinement of the system. The two other evaluations involved an actual clinical application which consisted of retrieving reports that were associated with specified diseases. Automated queries were written by a medical expert based on the structured output forms generated as a result of text processing. The retrievals obtained by the automated system were compared to the retrievals obtained by independent medical experts who read the reports manually to determine whether they were associated with the specified diseases. MEDLEE was shown to perform comparably to the experts. The technique used to perform the last two evaluations was found to be a realistic evaluation technique for a natural language processor. Carol Friedman, George Hripcsak, William DuMouchel, Stephen B. Johnson, Paul D. Clayton |
Nat. Lang. Eng. | 1 |
| 1994 | Research Paper: A General Natural-language Text Processor for Clinical RadiologyabstractOBJECTIVE: Development of a general natural-language processor that identifies clinical information in narrative reports and maps that information into a structured representation containing clinical terms. DESIGN: The natural-language processor provides three phases of processing, all of which are driven by different knowledge sources. The first phase performs the parsing. It identifies the structure of the text through use of a grammar that defines semantic patterns and a target form. The second phase, regularization, standardizes the terms in the initial target structure via a compositional mapping of multi-word phrases. The third phase, encoding, maps the terms to a controlled vocabulary. Radiology is the test domain for the processor and the target structure is a formal model for representing clinical information in that domain. MEASUREMENTS: The impression sections of 230 radiology reports were encoded by the processor. Results of an automated query of the resultant database for the occurrences of four diseases were compared with the analysis of a panel of three physicians to determine recall and precision. RESULTS: Without training specific to the four diseases, recall and precision of the system (combined effect of the processor and query generator) were 70% and 87%. Training of the query component increased recall to 85% without changing precision. Carol Friedman, Philip O. Alderson, John H. M. Austin, James J. Cimino, Stephen B. Johnson |
J. Am. Medical Informatics Assoc. | 1 |
| 1994 | Research Paper: A Schema for Representing Medical Language Applied to Clinical RadiologyabstractOBJECTIVE: Develop a representational schema for clinical concepts and apply it to the task of encoding radiology reports of the chest. DESIGN: The schema was developed following a manual analysis of sample reports from the domain. The schema has two main components: the Medical Entities Dictionary (MED), which specifies the formal representation of the concepts in the domain and of their structures, and the natural-language processor, which specifies the linguistic expressions of the concepts. The schema was evaluated by applying it to a test set of 7,500 reports. Two-hundred reports from the test set were manually analyzed by a medical expert to determine the accuracy and success rate of the system. RESULTS: 82% of the 7,500 reports that contained relevant clinical information were successfully structured automatically. For the smaller set of 200 reports, 80% were structured successfully with an accuracy rate of 97%. CONCLUSIONS: The schema is a formal representation for clinical concepts in radiology reports, and provides domain coverage that is particularly well-suited for natural-language processing of radiology for use in a decision support system. Carol Friedman, James J. Cimino, Stephen B. Johnson |
J. Am. Medical Informatics Assoc. | 1 |
| 1989 | A General Computational Treatment of the ComparativeabstractWe present a general treatment of the comparative that is based on more basic linguistic elements so that the underlying system can be effectively utilized: in the syntactic analysis phase, the comparative is treated the same as similar structures; in the syntactic regularization phase, the comparative is transformed into a standard form so that subsequent proceasing is basically unaffected by it.The scope of quantifiers under the comparative is also integrated into the system in a general way. Carol Friedman |
ACL | 1 |
| 1985 | Processing Free-Text Input to Obtain a Database of Medical InformationabstractThe Linguistic String Project of New York University has developed computer programs that convert the information in free-text documents of a technical specialty into a structured form suitable for mapping into a relational database. The processing is based upon the restrictions on the use of language that are characteristic of the subject matter and the document type. These restrictions are summarized in a “sublanguage grammar” that provides a set of word classes and formulas corresponding to the objects and relations of interest in the domain. The programs are independent of the particular sublanguage grammar employed. The application to narrative patient records will be described and the applicability of the methods to other domains discussed. Emile C. Chi, Carol Friedman, Naomi Sager, Margaret S. Lyman |
SIGIR | 2 |
| 1985 | Transporting the Linguistic String Project System from a Medical to a Navy DomainabstractThe Linguistic String Project (LSP) natural language processing system has been developed as a domain-independent natural language processing system. Initially utilized for processing sets of medical messages and other texts in the medical domain, it has been used at the Naval Research Laboratory for processing Navy messages about shipboard equipment failures. This paper describes the structure of the LSP system and the features that make it transportable from one domain to another. The processing procedures encourage the isolation of domain-specific information, yet take advantage of the syntactic and semantic similarities between the medical and Navy domains. From our experience in transporting the LSP system, we identify the features that are required for transportable natural language systems. Elaine Marsh, Carol Friedman |
ACM Trans. Inf. Syst. | 2 |
| 1982 | Natural Language Interfaces Using Limited Semantic Information
Ralph Grishman, Lynette Hirschman, Carol Friedman |
COLING | 3 |